Reference doc for running evals in preparation for filling in the readme's sizzler table.
We need to create well-founded eval data for a skill vs. no-skill run for each of our most behavior-shaping skills, which we want to advertise in the readme. We want to feel confident that these are good, reliable numbers we're reporting, and that they actually validate what we're trying to prove - our skills actually, factually improve agent performance.
All existing baseline data can be considered outdated, and isn't necessary to preserve (or delete, if we don't need to). The backing project for eval runs, eval-magic, has recently introduced better eval run isolation, which should hopefully avoid a lot of common confounds in the resulting data. Previous data was all generated using a version of eval magic that lacked that isolation, and in general was generated during runs that were more about pressure testing eval-magic, not generating data for these skills.
Eval-magic should now be more stable, but we'll still consider it to be in a pre-release state. We should keep an eye out for points of friction or issues encountered during these runs. That said, the goal will still be to explicitly generate data for these skills, showing that they provide real, verifiable improvement over a no-skill case.
Eval data in readme
See #244 for background
Reference doc for running evals in preparation for filling in the readme's sizzler table.
We need to create well-founded eval data for a skill vs. no-skill run for each of our most behavior-shaping skills, which we want to advertise in the readme. We want to feel confident that these are good, reliable numbers we're reporting, and that they actually validate what we're trying to prove - our skills actually, factually improve agent performance.
All existing baseline data can be considered outdated, and isn't necessary to preserve (or delete, if we don't need to). The backing project for eval runs, eval-magic, has recently introduced better eval run isolation, which should hopefully avoid a lot of common confounds in the resulting data. Previous data was all generated using a version of eval magic that lacked that isolation, and in general was generated during runs that were more about pressure testing eval-magic, not generating data for these skills.
Eval-magic should now be more stable, but we'll still consider it to be in a pre-release state. We should keep an eye out for points of friction or issues encountered during these runs. That said, the goal will still be to explicitly generate data for these skills, showing that they provide real, verifiable improvement over a no-skill case.
Eval data in readme
See #244 for background