Skip to content

Creating eval baseline data #250

Description

@slowdini

Reference doc for running evals in preparation for filling in the readme's sizzler table.

We need to create well-founded eval data for a skill vs. no-skill run for each of our most behavior-shaping skills, which we want to advertise in the readme. We want to feel confident that these are good, reliable numbers we're reporting, and that they actually validate what we're trying to prove - our skills actually, factually improve agent performance.

All existing baseline data can be considered outdated, and isn't necessary to preserve (or delete, if we don't need to). The backing project for eval runs, eval-magic, has recently introduced better eval run isolation, which should hopefully avoid a lot of common confounds in the resulting data. Previous data was all generated using a version of eval magic that lacked that isolation, and in general was generated during runs that were more about pressure testing eval-magic, not generating data for these skills.

Eval-magic should now be more stable, but we'll still consider it to be in a pre-release state. We should keep an eye out for points of friction or issues encountered during these runs. That said, the goal will still be to explicitly generate data for these skills, showing that they provide real, verifiable improvement over a no-skill case.

Eval data in readme

  • hardening-plans
  • investigating-bugs
  • test-driven-development
  • verifying-development-work

See #244 for background

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions