Repository navigation
Audit and harden the eval harness; unblock the sandboxed independent critic - #2
Conversation
- Count PMIDs given to the skill (development/seed/protocol seeds) as seen, so they never inflate never-reviewed recall. - Refuse a non-empty --run-dir so stale artifacts are not scored. - Verify the critic round-1 snapshot against the evidence-bundle hash and resolve critic/ledger evidence against the live agent workspace before it is deleted (absolute paths); report an undefined ablation delta as null. - Require a passing completion_gate.json before the suite scores a generated strategy; keep undefined-recall topics out of means and diffs. - Treat a Codex timeout as a failed run instead of a crash. - Create the agent workspace root with an inherited ACL: mkdtemp's owner-only ACL (Python 3.13+ on Windows) left sandbox-written files unreadable to the harness. Adds EVAL_HARNESS_AUDIT.md with findings, E2E results, and known gaps. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- run_child launches the critic/re-screen child in its own process group and, on timeout, kills the whole tree (taskkill /T on Windows, killpg on POSIX). subprocess.run killed only the codex.cmd wrapper, so communicate() blocked on pipes held by node/codex.exe and --timeout never fired. Timeout errors now carry the child's last stdout/stderr. - isolated_runner.py preflight checks the child CLI is signed in for the current account and makes one trivial bounded call. Inside a Codex sandbox the child runs as a separate sandbox user with no login, which made every nested critic hang; preflight now reports that in 0.2 s. - SKILL.md runs the preflight before record work; press-critic.md gains runner troubleshooting. Tests retarget run_child patches. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Inside a Codex sandbox the critic/re-screen child runs as a separate sandbox user with no Codex login and no usable root-certificate store, so nested children hung, then failed TLS once signed in. When `codex login status` fails for the current account, the codex-cli runner copies the host auth.json into a private owner-only home for that one child (outside the staged workspace), exports the machine's trusted roots to a CA bundle unless one is configured, and removes the copied credentials before the home when the child exits. The execution record notes codex_home host/private-copy; preflight reports it. Verified nested in the eval sandbox configuration: preflight ok in 7.3 s, critic-shaped call ok in 8 s, no credential copy left on success or failure. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- manifest_tool add refuses --kind mesh for outputs that are not mesh_tool artifacts (the gate rejected term-diff outputs only at handoff). - manifest_tool add reports the gate's hash-binding findings as binding_warnings as soon as an artifact is edited after being bound. - audit-scaffold fills reporting_notes limits/filters decisions from the locked protocol, or lists them as placeholders instead of failing only at render. - mesh_tool explains shell-split multi-word arguments; the PowerShell guidance warns against Start-Process -ArgumentList and points to --variants-file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- generate.py scores final_strategy.txt only when its text matches the strategy input of the final topic-only search the gate bound (exit 3). - generate.py scans the transcript for repository/fixture/qrel references and gold PMIDs used before any tool output showed them; the result is recorded as leakage_scan and a hit exits 5 (not a measurement). - The Codex driver relaunches a transient failure from a clean run dir (only the staged prompt and protocol are kept). - run_suite finds generate.py's <topic>/run-<UTC>/ layout, uses the latest run (which must pass the gate and have a clean scorecard), and reports never-reviewed recall per generated row and per source. - --runs > 1 is refused without --no-cache; default --timeout is 7200 s (GUI too) and elapsed_seconds is recorded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI (Linux) failed test_a_sandboxed_child_gets_a_private_login_that_is_removed_afterwards: it asserted a CA bundle is always exported, but the missing root store is a Windows sandbox-user problem and write_ca_bundle correctly returns None on a Linux runner with no default cafile. Stub the export to test the wiring, add a test that a configured CA bundle is never overridden, and keep a Windows-only test for the actual export. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e2280e936d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if pmid not in shown and pattern.search(action): | ||
| undiscovered_gold.append({"pmid": pmid, "action": action[:200]}) | ||
| shown.add(pmid) | ||
| shown.update(pmid for pmid, pattern in patterns.items() if pattern.search(output)) |
There was a problem hiding this comment.
Restrict PMID discovery to trusted retrieval output
Do not mark gold PMIDs as legitimately discovered from every command's output. If the agent runs a generic command such as find /workspace -type f -exec cat {} +, the command contains none of the forbidden markers, but fixture PMIDs in its output are added to shown; their subsequent use is therefore exempt and the run receives a clean leakage scan. This can silently admit answer-key-contaminated runs, so only output from known PubMed/retrieval commands should establish discovery.
Useful? React with 👍 / 👎.
| if not pem: | ||
| cafile = ssl.get_default_verify_paths().cafile | ||
| if cafile and Path(cafile).is_file(): | ||
| return Path(cafile) | ||
| return None |
There was a problem hiding this comment.
Fall back when certificate environment variables are empty
Handle defined-but-empty certificate variables without passing them through to ssl.get_default_verify_paths(). In that environment Python treats the empty SSL_CERT_FILE as an override and returns no cafile, so this function returns None and codex_child_home() leaves both certificate variables empty; the nested Codex child then still cannot establish TLS, defeating the provisioning path intended to unblock the critic. The newly added sandbox-home test reproduces this failure.
Useful? React with 👍 / 👎.
| flat = generated_root / topic | ||
| if (flat / "final_strategy.txt").is_file(): | ||
| return flat | ||
| runs = sorted(path for path in flat.glob("run-*") if path.is_dir()) if flat.is_dir() else [] |
There was a problem hiding this comment.
Prefer the newest nested run over a legacy flat copy
Check for run-* directories before returning the flat topic directory. When a legacy flat run containing final_strategy.txt coexists with newer generate.py runs, this early return permanently selects the legacy result, so even a newer failed rerun is hidden despite the function and README explicitly promising that the latest run wins.
Useful? React with 👍 / 👎.
| try: | ||
| os.kill(pid, 0) | ||
| except ProcessLookupError: | ||
| return False | ||
| return True |
There was a problem hiding this comment.
Treat zombie descendants as terminated in the timeout test
Avoid using os.kill(pid, 0) alone as the POSIX liveness check because it returns success for zombie processes. After killpg() terminates the test grandchild, container environments whose PID 1 does not promptly reap orphans retain it as a zombie, making this new test wait ten seconds and fail even though the descendant is no longer running; this occurs in the repository's current test environment.
Useful? React with 👍 / 👎.
cb02566
into
protocol-first-empirical-search-builder
Summary
This PR audits the eval harness for correctness before expensive benchmarks and fixes the bugs the audit demonstrated. It also works through Phases 0–2 of the resulting remediation plan, which covers why a real end-to-end generated build (CD011926) failed its completion gate. Every fix has a regression test that fails on the pre-change code.
Details:
EVAL_HARNESS_AUDIT.md(findings and E2E evidence) andEVAL_REMEDIATION_PLAN.md(plan and per-phase status).Type of Change
Related Issue
None.
Changes Made
Eval harness audit (
evals/)--run-diris refused.null, not0.0.mkdtemp's owner-only ACL (Python 3.13+ on Windows) made every file written by the sandboxed agent unreadable to the harness. The workspace root now inherits its parent's ACL.Phase 0: independent critic and re-screen inside the Codex sandbox (
scripts/isolated_runner.py)On timeout, the child's whole process tree is killed. Before, killing only the
codex.cmdwrapper leftcommunicate()blocked on the grandchildren.New
isolated_runner.py preflightcommand, run at intake perSKILL.md, so a runner that cannot start stops the build in seconds rather than 45 minutes in.Inside a Codex sandbox the child runs as a separate sandbox user with no Codex login and no usable root-certificate store. When
codex login statusfails, the runner:auth.jsoninto a private, owner-only home for that one child;This was verified live, nested in the eval sandbox configuration. Both failing and successful children leave no token copy behind.
Phase 1: fail fast on build recording errors
manifest_tool addrefuses--kind meshfor outputs that are not frommesh_tool.manifest_tool addreports the gate's hash-binding findings asbinding_warningswhen they arise.audit-scaffoldfills the required limits/filters notes from the locked protocol, or lists them as placeholders.mesh_toolexplains shell-split multi-word arguments. The PowerShell guidance warns againstStart-Process -ArgumentList.Phase 2: harness known gaps
final_strategy.txtis scored only if it matches the strategy input of the final search the gate bound (exit 3).leakage_scanof the transcript runs on every scored run (exit 5 on a hit).generate.py's<topic>/run-<UTC>/layout and uses the latest run;--runs > 1without--no-cache.--timeoutis 7200 s, andelapsed_secondsis recorded.CD010657 pilot fixes (added after opening)
noon any required criterion now excludes, regardless of criterion order. Before, an earlierunclearreturneduncertainfirst, contrary toreferences/candidate-screening.md.add --output P --supersedes Pre-binds P only when P is a fill-in-place artifact (a screening worksheet or a provenance-blinded pilot round). A broader draft of this change let an edited final QA be laundered past its hash check, and there is now a regression test for that. The pilot manifest validates cleanly under the narrowed rule.Testing
python -m pytest tests -qgives 973 passed.python scripts/pubmed_tool.py doctorreportsok: true.evals/generate.py CD011926attempts.Checklist
pubmed_tool.py doctorReviewer notes
run_suite.resolve_strategynow returns a 4-tuple;completion_gate.jsonplus a cleanscorecard.json;generate.pyhas new exit codes 3 (scored file is not the gated strategy) and 5 (leakage);codex_child_homeso they don't depend on the machine's login;run_childreplacessubprocess.runas the patch point for isolated-runner tests.auth.json. If that child refreshes the token, the host login may need signing in again. Theclaude-code-clirunner is not provisioned this way.generate.pyrun with a live critic ablation, has not been run yet.protocol-first-empirical-search-builder, notmain. The work builds on protocol-first code, and the two lines are intentionally kept separate.🤖 Generated with Claude Code