Skip to content

Add the offline evaluation-study harness (synthetic engineering evidence) - #1

Merged
ammarphp merged 1 commit into
mainfrom
evaluation-study-harness
Sep 27, 2026
Merged

ammarphp merged 1 commit into
mainfrom
evaluation-study-harness

Conversation

@ammarphp

@ammarphp ammarphp commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

What this adds

  • benchmarks/governance/, the offline evaluation-study harness. It contains:
    • evaluation contracts and campaign-manifest verification;
    • an independent, standard-library counting-likelihood oracle and a synthetic development task family;
    • deny-default Seatbelt subject isolation with a fail-closed process census;
    • the custody broker and delivery guard over RAVEL's stage workers;
    • fake, Claude Code and Codex host adapters (the real hosts are exercised only with hand-written fixtures);
    • treatment manifests with behavioural identity checks;
    • a mechanical evaluator, prespecified analysis and cost planning;
    • the campaign runner and CLI.
  • tests/governance/, about 2,600 tests.
  • docs/development/evaluation-study/, the study's working records: slice design, decisions, blockers, plan, the oracle review packet, the G1 record and the 2026-09-25 incident record.

What it is not

No model has been run through the harness yet. Every campaign so far is a synthetic engineering campaign that uses the fake host. Nothing here is an agent result, a treatment effect or a physics measurement.

Why a pull request

tests/governance has never run on Linux; only macOS runs and a macOS simulation exist. This PR lets the test suite job, the first real Linux run, finish before main moves. Per docs/development/distribution.md, main is then fast-forwarded to this commit rather than merged in the web interface.

Checks run on the source commit (553db5c) and on the staged export, on macOS:

  • the exporter's own checks (evidence 17/17, agent surface, publication);
  • the staged tree's full test suite, in chunks: 5029 collected, 0 failing once a git checkout and hooks are present, as in CI.

…nce)

Adds benchmarks/governance/: evaluation contracts and campaign manifest
verification, an independent counting-likelihood oracle with a synthetic
development task family, a deny-default Seatbelt isolation layer with a
fail-closed process census, the custody broker and delivery guard over the
RAVEL kernel's stage workers, fake/Claude Code/Codex host adapters (the real
hosts are exercised only with hand-written fixtures), treatment manifests
with behavioural identity checks, a mechanical evaluator, prespecified
analysis and cost planning, and the campaign runner and CLI. Adds
tests/governance/ and the evaluation-study records under
docs/development/evaluation-study/.

No model has been run through the harness; every campaign so far is a
synthetic engineering campaign with the fake host. Source commit 553db5c.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ammarphp
ammarphp merged commit 6d8b6e3 into main Sep 27, 2026
@ammarphp
ammarphp deleted the evaluation-study-harness branch September 28, 2026 15:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant