Skip to content

test(runner): establish compatibility baseline - #214

Merged
BjRo merged 2 commits into
mainfrom
test/211-eval-runner-baseline
Sep 19, 2026
Merged

BjRo merged 2 commits into
mainfrom
test/211-eval-runner-baseline

Conversation

@BjRo

@BjRo BjRo commented Sep 19, 2026

Copy link
Copy Markdown
Owner

Why

Closes #211. Darrow needs executable behavior-parity evidence before generic evaluation infrastructure moves to Sevro under ADR-0010.

What changed

Adds a configurable command-level suite that runs the current eval runner in an isolated project with a credential-free synthetic adapter. The baseline covers case selection, project and result roots, success and failure categories, result and diagnostic artifacts, cancellation, and retained trial evidence without asserting terminal wording or runner module structure. It also documents bun run test:eval-runner-compatibility and the JSON-argv override for a future Sevro launcher.

Verification

  • bun run test:eval-runner-compatibility — 5 passed
  • bun test evals/runner/run-cli.test.ts evals/runner/run-selection.test.ts evals/runner/run-artifact.test.ts evals/runner/run-recovery.test.ts — 43 passed
  • bun test evals/runner/suite.test.ts — 8 passed on isolated rerun
  • bun run typecheck — passed
  • bun run lint:ts — passed
  • bun run lint — passed
  • bun run lint:shell — passed
  • bun run check:docs — passed
  • bun run check:decisions — passed
  • bun run check:python — passed
  • bun test — 635 passed; one unchanged discovery-eval test failed reproducibly, and one suite test timed out under the full parallel load but passed on the isolated rerun above

Review notes

The alternate-runner seam is a JSON argv array with fixture-path placeholders; the configured launcher owns synthetic-adapter wiring for the four documented scenarios. This slice does not move runner source, change Sevro, publish a package, or switch Darrow evaluation commands. The full-suite discovery failure is in unchanged eval data and also fails when run alone.

Checklist

  • I have read and followed CONTRIBUTING.md, including the contribution
    licensing terms.
  • I added or updated the applicable invariant before implementation, or
    this change does not affect a capability invariant.
  • I added or updated colocated evals, or this change does not affect skill
    behavior.
  • I confirmed that each changed plugin remains self-contained, or this
    change does not affect plugin content.
  • I ran bun run check:python, or this change does not affect registered
    Python packages or their repository quality infrastructure.

@BjRo
BjRo merged commit f739bb4 into main Sep 19, 2026
29 checks passed
@BjRo
BjRo deleted the test/211-eval-runner-baseline branch September 19, 2026 14:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Establish a black-box compatibility baseline for the eval runner

1 participant