Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 29 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,34 @@
# Changelog

## Unreleased — offline evaluation-study harness

- Add an offline harness in `benchmarks/governance/` for a planned controlled study: does an
evidence-bound delivery guard improve agent outcomes beyond the same tools plus instructions?
It wraps the unchanged v1 registry with strict contracts, a campaign manifest bound to each
campaign's kind, freeze and sealed run records, an independent pyhf-free counting oracle and a
synthetic four-variant development family whose answers are public and used only for development.
- Run subjects in fresh workspaces under a deny-default macOS Seatbelt profile, with admission
checks and System V IPC cleanup. An allowlist proxy class is tested against local targets, but no
launch starts it yet. A coordinator-owned custody broker runs every operation; its delivery guard
checks each submission and stage workers run the RAVEL kernel.
- Add fake, Claude Code and Codex host adapters; only the deterministic fake host has run, and the
other two are tested against mocked executables. Add treatment manifests with behavioral identity
checks, a mechanical evaluator that re-derives its judgments from sealed evidence, analysis and
cost planning, and a runner/CLI with journal, resume, budgets and per-launch verification.
- During development, on 2026-09-25, the harness process census killed every user process on the
development Mac twice: with a marker file missing, its membership test matched almost every
process. The census now fails closed and signals only processes proven to belong to the launch
and started after it; tests run under a session signal guard. See the
[incident record](docs/development/evaluation-study/incident-2026-09-25.md).
- The G1 engineering acceptance run at `0b3d251` passed `pytest tests` (4,776 passed, 0 failed)
and two 16-assignment synthetic campaigns through the CLI.

This is synthetic engineering evidence. No model has been evaluated, no agent result or treatment
effect is reported, and the real-host path has not run. The oracle and scoring rules stay
provisional until a deferred human review. On Linux the Seatbelt tests skip and the process readers
are tested only against recorded samples. Design, decisions and open items are in
[docs/development/evaluation-study/](docs/development/evaluation-study/).

## Unreleased — supplied-data scientific studies

- Add approved, bounded `analyze`, `quantities`, `measurement` and `domain` studies
Expand Down
6 changes: 4 additions & 2 deletions DIRECTORY.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,15 +40,17 @@ README demonstrations write to ignored `local-runs/`; these local outputs are no
| `tests/unit/` | Focused regression tests | 117 |
| `tests/adversarial/` | Adversarial workflow scenarios | 37 |
| `tests/fixtures/` | Immutable test inputs | 7 |
| `benchmarks/` | Benchmark and capability registries | 21 |
| `tests/governance/` | Evaluation-study harness tests | 37 |
| `benchmarks/` | Benchmark and capability registries | 65 |
| `benchmarks/governance/` | Offline evaluation-study harness (counted in benchmarks/ too) | 46 |
| `native/src/` | Native C++ source | 3 |
| `native/scripts/` | Native build and execution scripts | 9 |
| `environment/` | Simulation environment setup | 7 |
| `scripts/` | Maintenance, documentation, and export commands | 24 |
| `docs/workflow/` | Physics workflow instructions | 54 |
| `docs/reference/` | Capabilities, contracts, and tool reference | 12 |
| `docs/validation/` | Scoped results, cases, and evidence descriptions | 20 |
| `docs/development/` | Contributor guidance and explicitly labeled history | 36 |
| `docs/development/` | Contributor guidance and explicitly labeled history | 46 |
| `docs/research/` | Research and evaluation protocols | 21 |
| `docs/guides/` | Longer guides and sources | 5 |
| `evidence/` | Curated historical inputs, measurements, and provenance | 856 |
Expand Down
6 changes: 6 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -272,3 +272,9 @@ with [upstream acknowledgements](docs/reference/third-party.md). Licensed under
For supported-analysis expansion, use the [public routine landscape](docs/research/2026-09-26-analysis-landscape.md)
and `ravel analyses summary`. The census distinguishes public availability,
adaptation requirements and demonstrated scientific support.

The [evaluation-study harness](benchmarks/governance/README.md) is offline tooling for a planned
controlled study of evidence-bound delivery safeguards for coding agents. So far it has run only
synthetic campaigns with a deterministic fake host. It reports no agent or model results, and its
real-host path has not been exercised. Design, decisions and open reviews are in
[docs/development/evaluation-study/](docs/development/evaluation-study/).
76 changes: 71 additions & 5 deletions benchmarks/governance/README.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,12 @@
# Governance experiment registry

`governance_experiment.py` freezes a complete 2×2 assignment roster and scores
`experiment.py` freezes a complete 2×2 assignment roster and scores
independently adjudicated outcomes. It launches no agents or simulation jobs,
contains no physics oracle, and provides descriptive accounting only. The
prospective scientific protocol is in
[`docs/research/2026-09-05-competitive-design-and-validation.md`](../../docs/research/2026-09-05-competitive-design-and-validation.md).
The evaluation-study harness built around this unchanged v1 contract is documented in
[`docs/development/evaluation-study/`](../../docs/development/evaluation-study/).

The four arms are fixed: `baseline` (no additional instructions, no experimental
enforcement), `instructions` (instructions only), `enforcement` (enforcement only),
Expand Down Expand Up @@ -122,13 +124,77 @@ self-drive. Those require the separate experiment described in the protocol.

## Software verification

Run pytest from outside this checkout to avoid its legacy `py.py` shadowing
pytest's dependency:
From the repository root:

```sh
cd /tmp
python3 -m pytest /absolute/path/to/hep-agentic-pipeline/tests/unit/test_governance_experiment.py -q
python -m pytest tests/unit/test_governance_experiment.py -q
```

The fixtures are explicitly synthetic tests of accounting failures, not empirical
agent outcomes. No campaign results ship in this directory.

## Evaluation harness around v1 (synthetic engineering only)

The other modules in this directory implement the offline evaluation slice described in
[`slice-design.md`](../../docs/development/evaluation-study/slice-design.md). They are standard
library only, except the broker's stage workers (`stages/`), which run the RAVEL kernel and import
pyhf and `ravel` under the stage interpreter. They never modify `experiment.py`, and they add
sidecar records linked to v1 rows by run id and digest. They launch no paid model call: the only
host that runs in tests or the CLI is the deterministic fake adapter, and every record it produces
is labeled synthetic. The Claude Code and Codex adapters are exercised only against mocked
executables and hand-written synthetic fixtures in the hosts' documented stream formats; no fixture
is a recorded host session. An 8-assignment Claude Code engineering smoke is authorized but cannot
run until the real-host pieces it needs exist
([smoke request](../../docs/development/evaluation-study/smoke-request.md)).

| Module | Role |
|---|---|
| `canonical.py`, `contracts.py`, `schemas/` | Canonical JSON, strict loading, exact-key sidecar validators (schemas are documentation mirrors) |
| `campaign_manifest.py` | One v1 spec and registry per host configuration and campaign kind (the kind is bound into the spec), a write-once approval record, the frozen family definitions, byte-level verification of every sealed run record the runner would trust (`verify(sealed_runs=...)` limits that to named runs), and provenance checks |
| `oracle/`, `tasks/development/` | Independent pyhf-free counting oracle and the provisional `likelihood_freshness` development family |
| `isolation.py`, `allowlist_proxy.py` | Fresh-byte workspaces, admission checks, deny-default Seatbelt profile (no terminals, no POSIX shared memory or named semaphores), sandboxed launcher with its process census and System V IPC cleanup; the allowlist proxy is tested but nothing starts it yet |
| `broker.py`, `guard.py`, `stages/`, `client/` | Coordinator-owned operation broker over the RAVEL kernel, the delivery guard, and the subject's `ravel-task` client with its neutral tool guide |
| `treatment.py`, `treatments/` | Arm manifests, prompt assembly and the treatment-identity checks: manifests (`treatment_diff`), broker behavior (`behavioral_diff`) and the delivered prompt (`check_prompt`) |
| `adapters/` | Fake, Claude Code and Codex host adapters |
| `runner.py`, `cli.py` | Assignment coordinator (journal, resume, budgets, per-launch verification against the frozen campaign, sealing, outcome re-derivation), the behavioral treatment check (`treatment-diff --behavioral`), human incident decisions (`incident-decision`) and the command line |
| `audit.py` | Independent mechanical evaluator: judge reports and v1 outcome rows (provisional rules below) |
| `analysis.py` | Family-aware descriptive analysis, missingness bounds, design simulation, cost planning |

A synthetic campaign, from the repository root (store and subjects root outside the lab tree: the
outermost ancestor of the checkout that holds `.git`, `CLAUDE.md` or `AGENTS.md`). The interpreter
must import pyhf and the kernel's dependencies, because the build probes it as the broker's stage
interpreter; the kernel itself is always imported from this checkout's `src/` through the stage
environment's pinned `PYTHONPATH`, and the broker refuses an interpreter that would import `ravel` from
anywhere else. `.venv-dev/bin/python` qualifies, as does an interpreter with `requirements-replay.lock`
installed (the CI setup); a system `python3` without pyhf fails at that probe:

```sh
PY=.venv-dev/bin/python
$PY benchmarks/governance/cli.py build-synthetic --store STORE --campaign-id demo \
--created-utc 2026-09-25T00:00:00Z --seed 11 --schedule-seed 7 --subjects-root SUBJECTS
$PY benchmarks/governance/cli.py treatment-diff --behavioral --campaign STORE/synthetic/demo
$PY benchmarks/governance/cli.py run --campaign STORE/synthetic/demo
$PY benchmarks/governance/cli.py audit --campaign STORE/synthetic/demo
$PY benchmarks/governance/cli.py report --campaign STORE/synthetic/demo
$PY benchmarks/governance/cli.py verify --campaign STORE/synthetic/demo
```

The evaluator's verdicts include `historical` (a superseded value that its own clause marks as
superseded) and `retracted_after_delivery` (a delivered finding that a later positive retraction
withdrew). Its invalid-claim quantities count distinct conclusions: `attempted_invalid` covers
everything put to the gate (every submission, blocked or accepted, and forged output files, before any
later withdrawal), `delivered_invalid` what was delivered and still stands (accepted submissions, the
final message, forged files), and neither bounds the other; `false_block` and `repaired_after_block`
are null when a finding's validity is unknown (slice design §4.8, §11).

The oracle, task variants, tolerance and the audit's claim-classification rules are
provisional until the human reviews listed in
[`decisions.md`](../../docs/development/evaluation-study/decisions.md), which are deferred to one
consolidated review (E-34). A synthetic campaign tests the harness; it is not evidence about any
agent, arm or scientific claim. Run the harness
tests with `.venv-dev/bin/python -m pytest tests/governance -q`; sandbox tests skip where
`sandbox-exec` is unavailable. On Linux (the public CI) there is no Seatbelt: those tests skip, a
campaign runs only as `none_test_only`, and the launcher reads processes from `/proc`
(`isolation._linux_proc_read`, tested on every host against recorded Linux samples). The tests launch and kill subject processes: keep the session signal
guard in `tests/governance/conftest.py`, and run one full suite at a time on a host (the prevention
rules in the [incident record](../../docs/development/evaluation-study/incident-2026-09-25.md)).
6 changes: 6 additions & 0 deletions benchmarks/governance/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
"""Evaluation-study harness around the strict v1 governance registry.

Standard library only (the broker's RAVEL stage workers run under the pinned
replay interpreter). Nothing in this package launches paid model calls, event
generation or training by default. See docs/development/evaluation-study/.
"""
12 changes: 12 additions & 0 deletions benchmarks/governance/adapters/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
"""Host adapters for the evaluation harness (slice design §9, §13).

- ``base``: AdapterResult, the Adapter base class, launch and raw-stream helpers.
- ``fake`` / ``fake_subject``: the synthetic fake host and its stdlib subject script (WP06).
- ``claude_cli``: Claude Code CLI adapter (WP09); ``codex_cli``: Codex CLI adapter (WP08).

Standard library only. No adapter test calls a model; the CLI adapters are exercised with
mocked executables and synthetic recorded-format streams.
"""
from .base import Adapter, AdapterResult, HostDriftError

__all__ = ["Adapter", "AdapterResult", "HostDriftError"]
Loading