Smallest readable coding-agent harness that scores on benchmarks. ~970 lines, 59.6% on Terminal-Bench 2.0.
-
Updated
Jul 21, 2026 - Python
Smallest readable coding-agent harness that scores on benchmarks. ~970 lines, 59.6% on Terminal-Bench 2.0.
Defensive evaluation of refusal calibration when LLMs interpret science foundation model outputs, with reproducible harnesses and 24.3K outcome records.
Add a description, image, and links to the frontier-model-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the frontier-model-evaluation topic, visit your repo's landing page and select "manage topics."