Evaluation-driven autonomous development: a harness for agent loops whose acceptance criterion is a metric, not a test suite. Held-out metrics, sealed scoring, integrity checks.
-
Updated
Jul 30, 2026 - Python
Evaluation-driven autonomous development: a harness for agent loops whose acceptance criterion is a metric, not a test suite. Held-out metrics, sealed scoring, integrity checks.
Auditable DPO evaluation framework for LLM-as-judge reliability, held-out validation, and claim-boundary reporting.
Add a description, image, and links to the held-out-validation topic page so that developers can more easily learn about it.
To associate your repository with the held-out-validation topic, visit your repo's landing page and select "manage topics."