Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
62 changes: 60 additions & 2 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@
that launch's leader process. A coordinator-owned custody broker runs every operation; its delivery guard
checks each submission and stage workers run the RAVEL kernel.
- Add fake, Claude Code and Codex host adapters; the deterministic fake host has run the synthetic
campaigns, the Claude Code adapter one authorized engineering smoke, and the Codex adapter is tested
campaigns, the Claude Code adapter one authorized engineering smoke and one development pilot, and the Codex adapter is tested
against mocked executables. The Claude Code adapter has a real-host
path for one authorized 8-assignment engineering smoke: a builder bound to the budget owner's
single-use approval, a declared credential exception, host probes, live checks, stop rules and
Expand All @@ -38,13 +38,71 @@
record as the missing thing, a σ_vis value under a luminosity unit counts against a refusal, table
headers and list headings lend their historical wording, and a decimal that states a current or
prior quantity is judged instead of dropped. The re-judged smoke is unchanged.
- Close the evaluator's two known false-clean paths before any pilot, with held-out cases committed
first: the task-bank profile now judges a stray value in a label-value line, a numbered line or a
semicolon clause (integer counts included, in its heading's unit) instead of dropping it, and a
supersession statement under a doubting or negating frame ("I cannot say the previous value was not
used") is unresolved instead of historical. The new readings are linear in the text; the replayed
smoke is unchanged.
- Prepare the live build for a development pilot: an approval of the pilot's kind also binds the schedule
seed, the broker limits and the roster order, and the authorization says it is a pilot;
`build-live --design-budget` sets the task bank's design budget; the host probes measure the cost of
a request whose usage reports no inference geography (the pinned CLI prices it at 1.0), so such a run's
cost recompute is verified; and `run` censuses a lost run launch by its record before preflight,
launching nothing while that census is unclean.
- Rehearse the development pilot offline: a pilot-shaped campaign of the twelve bank tasks in four arms
ran through the real pinned CLI against a local mock API, with a dummy token and no network, from the
build and the host probes to the audit and the cost reconciliation, without a stop. The rehearsal also
showed that the post-run sandbox-denial check mostly cannot see the subject's routine denials; it stays
a best-effort check.
- Repair what the review of the pilot's engineering found, with held-out cases committed first for the
evaluator: a supersession statement embedded under someone else's claim or an evaluation ("The draft
claims ...", "It is incorrect that ...") or taken back in the next sentence is unresolved instead of
historical, a value under a unit heading that scales to nothing known is judged instead of dropped, and
the task-bank profile reads counts in tables and label lines. The build's approval is re-read by
`verify` and preflight; a smoke approval no longer builds a pilot-sized campaign. The sandbox-denial
check records which reports came from the launch; its earlier warnings were machine-wide, not
attributed to the subject. The pilot rehearsal's tooling and evidence are kept, and the rehearsal was
re-run at the repaired revision without a stop.
- Run the 96-assignment development pilot (2026-09-29): the twelve bank tasks, two seeds and four arms on
the pinned Claude Code CLI, all sealed with no stop, at 18.44 USD of quota usage against a 192 USD
admission threshold. Its [record](docs/development/evaluation-study/pilot-record.md) reports engineering
results and the evaluator's unresolved rate: the mechanical evaluator left `unsupported_claim`
unresolved in 62 of 96 runs, and a read-only check by analysis agents found 23 of the 29 cells it
scored true, and all 8 of its verdicts on one refusal control, wrong, so no arm comparison is drawn
and the evaluator is repaired next.
- Repair the mechanical evaluator against the pilot's classes, with 137 held-out cases committed first:
listed limits under quantile labels or role words are read in order in both scoring profiles, a converted
value takes its source value's role, label digits, unit factors and unit identities are no numbers,
unit-label lines and task-bank field names give their unit, the attribution reader knows more verbs and
correction words, prose refusals are recognized in more forms and also in the last submission's report,
and quantile notation, a POI value or a supplied input no longer counts as a delivered cross section.
Two cells of the replayed smoke move by design; the sealed pilot is re-judged next.
- Re-judge the sealed pilot and smoke read-only with the repaired evaluator, each changed cell traced to
the reading responsible. On the pilot, 16 of the 23 cells the agents found falsely flagged are cleared
and the real errors stay flagged. But only 3 of 8 refusals on one refusal control and 3 of 5 on the
other are now scored valid, short of the repair's acceptance, and the evaluator still leaves
`unsupported_claim` unresolved in 64 of 96 runs (62 before), so no arm comparison is drawn. The
smoke gains a second verified completion.
- Close the fail-open paths a review found in that repair, with 98 held-out cases committed first: a value
of the parameter of interest that states the limit is judged again, a correction word elsewhere in a
sentence no longer softens a wrong value, bound wording is read only as a true lower bound, refusals that
are negated, questioned, about another run or undone by a hand computation are no longer read as
refusals, a refusal found only in a submission's report text can no longer be scored valid until the
owner decides H-110, and adopted source values, rejections of an action on a value, unit identities,
unit-label lines, table labels and census field names no longer hide a delivered value.
- Re-judge the sealed pilot and smoke read-only with that repaired evaluator, each changed cell traced to
the reading responsible. Four pilot cells move, all toward unresolved: two refusals on one refusal
control lose their valid verdict while H-110 is open, so 1 of 8 is valid there (3 of 5 on the other).
The agreement with the agents' false-positive and real-error findings is unchanged, `unsupported_claim`
stays unresolved in 64 of 96 runs, and the smoke is unchanged.
- Fix the distribution export: the task bank pins the production specs' generation plan, not the spec
file's bytes, whose machine-local paths the export rewrites (every campaign build failed from an
export); the exporter now builds the task bank from the stage and writes no bytecode into it; and the
end-to-end rehearsal tests skip where the bound shell `/bin/zsh` is absent.

This is synthetic engineering evidence. No agent result or treatment effect is reported, and no paid
call has been made beyond the 8-run engineering smoke. The oracle and scoring rules stay
call has been made beyond the 8-run engineering smoke and the 96-run development pilot. The oracle and scoring rules stay
provisional until a deferred human review. On Linux the Seatbelt tests skip and the process readers
are tested only against recorded samples. Design, decisions and open items are in
[docs/development/evaluation-study/](docs/development/evaluation-study/).
Expand Down
2 changes: 1 addition & 1 deletion DIRECTORY.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ README demonstrations write to ignored `local-runs/`; these local outputs are no
| `docs/workflow/` | Physics workflow instructions | 54 |
| `docs/reference/` | Capabilities, contracts, and tool reference | 12 |
| `docs/validation/` | Scoped results, cases, and evidence descriptions | 20 |
| `docs/development/` | Contributor guidance and explicitly labeled history | 49 |
| `docs/development/` | Contributor guidance and explicitly labeled history | 51 |
| `docs/research/` | Research and evaluation protocols | 21 |
| `docs/guides/` | Longer guides and sources | 5 |
| `evidence/` | Curated historical inputs, measurements, and provenance | 856 |
Expand Down
12 changes: 7 additions & 5 deletions benchmarks/governance/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,13 +147,15 @@ formats and one token-free stream the real 2.1.281 pin produced against a local
scripted every model turn; no stream fixture is a recorded session with a model. The smoke regression
fixtures (`tests/governance/fixtures/smoke/`) keep the smoke's real model outputs, with a synthetic
canary, broker secret and handles. The CLI can build, probe,
launch and stop one real-host synthetic engineering smoke with a pinned Claude Code CLI, only with
launch and stop a real-host synthetic engineering smoke or development pilot with a pinned Claude Code CLI, only with
`RAVEL_EVAL_LIVE=1` and the budget owner's single-use approval
([smoke request](../../docs/development/evaluation-study/smoke-request.md): configuration,
procedure and stop rules). The 8-assignment smoke was authorized (E-34, E-47) and ran on 2026-09-27
([smoke record](../../docs/development/evaluation-study/smoke-record.md)); before it the real pinned
CLI had run only at zero cost, in the host probes and an offline rehearsal with a dummy token against
a local mock API (E-88, E-101). The Codex adapter has no
a local mock API (E-88, E-101). The 96-assignment development pilot was authorized on 2026-09-28 and
ran on 2026-09-29 ([pilot record](../../docs/development/evaluation-study/pilot-record.md)); it reports
engineering results and the evaluator's unresolved rate, not a treatment effect. The Codex adapter has no
builder.

| Module | Role |
Expand All @@ -167,10 +169,10 @@ builder.
| `treatment.py`, `treatments/` | Arm manifests, prompt assembly and the treatment-identity checks: manifests (`treatment_diff`), broker behavior (`behavioral_diff`) and the delivered prompt (`check_prompt`) |
| `adapters/` | Fake, Claude Code and Codex host adapters; the fake subject's behaviours include each bank pair's naive behaviour, the boilerplate refusal and `reference_variant` (the reference's values with another analyst's citations and phrasing, E-186) |
| `runner.py`, `cli.py` | Assignment coordinator (one builder for the fake and a real host, journal, resume, budgets, per-launch verification against the frozen campaign, the build-time host binding and the exact launch-call check, the real-host proxy, credential injection and post-run checks, sealing with the kernel receipts, stop and limit, outcome re-derivation), the behavioral treatment check (`treatment-diff --behavioral`; a real host needs the latest check to pass and re-derive), human incident decisions (`incident-decision`) and the command line (`build-live`, `host-probe`, `preflight`, `live-checks`, `go-no-go`, `stop` for the real host) |
| `audit.py`, `audit_bank.py` | Independent mechanical evaluator: judge reports and v1 outcome rows (provisional rules below). `audit.py` holds the run-level rules and the likelihood_freshness profile (judge report version 1); `audit_bank.py` the task-bank scoring profiles of kx, hv, mq and tz (judge report version 2: fault and convention values, relation and categorical claims, evidence constraints, generic refusal validity; decisions E-131 to E-135; after the second review of 2026-09-27 the refused object per task, prose roles named or ordered, evidence read as the guard reads it, E-165 to E-178; after the real-host smoke, non-primary claim roles, header-aware tables, value annotations, field-name labels, refusal-reason variants and rejected quotations, E-188; narrowed after its review, with header and heading wording, luminosity-unit σ values and stray decimals, E-190). The scorer id binds both and the task-bank registry they read (E-179) |
| `audit.py`, `audit_bank.py` | Independent mechanical evaluator: judge reports and v1 outcome rows (provisional rules below). `audit.py` holds the run-level rules and the likelihood_freshness profile (judge report version 1); `audit_bank.py` the task-bank scoring profiles of kx, hv, mq and tz (judge report version 2: fault and convention values, relation and categorical claims, evidence constraints, generic refusal validity; decisions E-131 to E-135; after the second review of 2026-09-27 the refused object per task, prose roles named or ordered, evidence read as the guard reads it, E-165 to E-178; after the real-host smoke, non-primary claim roles, header-aware tables, value annotations, field-name labels, refusal-reason variants and rejected quotations, E-188; narrowed after its review, with header and heading wording, luminosity-unit σ values and stray decimals, E-190; before the pilot, the task-bank stray pass, stray numbers in their heading's unit and doubted supersession statements unresolved, in linear time, E-200 to E-202; after their review, embedding frames that are not the writer's own assertion, the frame's reach before and after the statement, integers in tables and label lines, and doubted corrections, E-210; after the pilot, quantile label lists read in order in both profiles, converted values taking their source's role, label and unit-factor digits that are no numbers, POI values and supplied inputs that are no σ_vis mention, and refusal presence read in the final submission's report text, E-220 to E-223, with its scoring-rule changes awaiting review, H-110 and H-115; after that repair's review, narrower POI, formula-bound, rejection, unit-identity and refusal readings, a refusal read only in the report text capped at null, and a value at the recorded bound read as that bound, E-226 and E-228, with its scoring-rule changes awaiting review, H-117). The scorer id binds both and the task-bank registry they read (E-179) |
| `analysis.py` | Family-aware descriptive analysis, missingness bounds, design simulation, cost planning |
| `live.py`, `credentials.py` | The real-host smoke: the pinned host and the model byte check, the budget owner's approval checks and the per-user single-use approval ledger (E-77, E-89), the launch declaration (`host_launch`) and the Claude adapter the binding and every launch are built from, `build-live`, the pin's code signature (E-90), proxy attribution by socket ownership, the keychain residue check, preflight (PF-01 to PF-14), host probes (HP-01 to HP-13, each probe launch recorded and censused again while unclean: E-96), live checks (LC-01 to LC-23), the stop rules (S1 to S8, failing closed and re-derived before every launch: E-74, E-78), the S10a go/no-go gate after run 1 (E-92), the catalog entry located in the pinned bytes (E-85) and the cost reconciliation; the one credential exception (stat-only checks, a nonblocking read, token variants, the post-run sweep and redaction) |
| `rehearsal.py` | HP-13 offline rehearsal of a pinned Claude Code CLI against the local mock Messages API (`tests/governance/mock_messages_api.py`) with dummy credentials: tool path, planted rc hooks, pricing against an unknown control model, budget cutoff, settings, token sweep |
| `live.py`, `credentials.py` | The real-host smoke and pilot: the pinned host and the model byte check, the budget owner's approval checks (a smoke's, or a pilot's, which also binds the schedule seed, the broker limits and the roster order: E-204; the design budget: E-205) and the per-user single-use approval ledger (E-77, E-89), the launch declaration (`host_launch`) and the Claude adapter the binding and every launch are built from, `build-live`, the pin's code signature (E-90), proxy attribution by socket ownership, the keychain residue check, preflight (PF-01 to PF-15; PF-15 and `verify` re-read the frozen approval after the build, and a smoke approval covers only the smoke's size and default limits: E-211), host probes (HP-01 to HP-13, each probe launch recorded and censused again while unclean: E-96), the census of a lost run launch before preflight (E-206; REVOKE for one sealed unclean: E-212), the sandbox-denial collector with each report's pid and the launch-attributed split (E-213), live checks (LC-01 to LC-23), the stop rules (S1 to S8, failing closed and re-derived before every launch: E-74, E-78), the S10a go/no-go gate after run 1 (E-92), the catalog entry located in the pinned bytes (E-85) and the cost reconciliation; the one credential exception (stat-only checks, a nonblocking read, token variants, the post-run sweep and redaction) |
| `rehearsal.py` | HP-13 offline rehearsal of a pinned Claude Code CLI against the local mock Messages API (`tests/governance/mock_messages_api.py`) with dummy credentials: tool path, planted rc hooks, pricing against an unknown control model, the cost multiplier of the inference geographies "us" and "not_available" (E-203), budget cutoff, settings, token sweep |

A synthetic campaign, from the repository root (store and subjects root outside the lab tree: the
outermost ancestor of the checkout that holds `.git`, `CLAUDE.md` or `AGENTS.md`). The interpreter
Expand Down
Loading