Repository navigation
fix(evaluation): register a Scope for each arm before running Codex - #1723
Merged
Merged
Conversation
Since oceanbase#1401 the Server generates Scope IDs and rejects an explicit Scope that does not exist. The console still passed the literal `eval:{run_id}:{arm}` to the Codex plugin, so its UserPromptSubmit hook got a 404 while resolving the Scope and exited without capturing or recalling. The ON arm then behaved like the OFF arm and failed treatment validation. Each arm now creates its Scope through `POST /v1/scopes` once the Server is ready, using `eval:{run_id}:{arm}` as the idempotency key, and passes the returned ID to Codex. The evidence query and treatment validation use that ID, and the evidence records the key as `scope_key` so reports can tie it to the run. Evidence recorded before this change carries the key as its Scope ID and still validates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The report loader accepted `scope_key=""` and, because the check used `or`, fell back to `scope_id` as if the evidence predated per-arm Scopes. The harness never writes an empty key, since its own evidence model rejects one, but the report layer re-validates stored evidence independently and should enforce the same rule. Only an absent key now selects the earlier format. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The console parses treatment evidence with a strict schema, and the API now returns `scope_key` for every run: `null` for evidence recorded before arms registered their own Scope, and the run arm key otherwise. Without the field in the schema, the console rejected the evidence of every report, including earlier runs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scope creation is part of preparing the Server for an arm, like the readiness gate, but its failure was raised as a plain invalid treatment, which the worker treats as terminal. A transient failure now raises a readiness failure with its own reason, so the task is retried with a fresh runtime. Also rename the key helper to `arm_scope_key`, so it no longer reads like the evidence field of the same name. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue or RFC does this PR close?
Closes #1722.
Rationale for this change
Since #1401, the Server generates Scope IDs and rejects an explicit Scope that does not exist. The SWE-bench Pro console still passed the literal
eval:{run_id}:{arm}to the Codex plugin and never created that Scope. The plugin'sUserPromptSubmithook got a 404 while resolving the Scope and exited without capturing or recalling. On current master, the ON arm therefore behaves like the OFF arm and fails treatment validation. The linked issue has the reproduction across4e0f78ad, #1401 and master.What changes are included in this PR?
POST /v1/scopesinside the container. The API requires an idempotency key; the arm useseval:{run_id}:{arm}.validate_treatment. The container-levelPOWERCONTEXT_CODEX_SCOPE_IDis removed: the ID only exists after the Server starts, and onlycodex execruns the plugin hooks.scope_not_created). The worker already retries readiness failures with a fresh runtime, and it does so before Codex runs.scope_keyin evidence and reports. Treatment evidence records the key in a new optionalscope_keyfield next toscope_id. The report check matchesscope_keywhen present. When it is absent, which is the case for evidence recorded before this change (whose Scope ID was the key itself), the check falls back toscope_id, so existing runs keep validating. An emptyscope_keyis rejected, matching the harness's own evidence model.scope_key. The API returnsscope_keyfor every run, includingnullfor earlier runs, so without this the console would reject the evidence of every report.This PR does not change how
mcp_requestsis counted. The issue explains why that count can't be observed from the Server logs; it is left for a separate decision.Are there any user-facing changes?
There are no CLI changes.
scope_keyfield, and the console accepts it.treatment.json,scope_idis now a Server-generatedscp_…ID.How was this change tested?
make checkandmake unit-testpass (2841 passed, 58 skipped).uv run pytestinevaluation/gives 1052 passed. Five tests intests/web/test_worker.pyfail both on this branch and on unmodifiedupstream/master(1e90376a); they are unrelated to this change.evaluation/web/,npx tsc -bandnpm testpass (96 tests). The report contract test now uses the evidence shapes the API returns (anullkey for an earlier run, a registered Scope for a new one). It failed before the schema change.scope_not_createdreadiness failure before Codex runs.scope_keyis rejected; with the empty-key rejection reverted, that case fails.scp_…ID./runtime/pc-env/bin/python), loopback address and proxy environment (NO_PROXYincludes loopback) as the existing readiness probe and evidence query. A full run would still stop at the report'smcp_requests > 0check described in the issue.AI usage statement
This PR was developed with Claude Code (Claude Opus 5.5), which reproduced the regression, wrote the change and tests, and ran the validation above. The author reviewed the change.
🤖 Generated with Claude Code