Skip to content

fix(evaluation): register a Scope for each arm before running Codex - #1723

Merged
PsiACE merged 4 commits into
oceanbase:masterfrom
Fengzdadi:fix/evaluation-arm-scope
Sep 23, 2026
Merged

PsiACE merged 4 commits into
oceanbase:masterfrom
Fengzdadi:fix/evaluation-arm-scope

Conversation

@Fengzdadi

@Fengzdadi Fengzdadi commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Which issue or RFC does this PR close?

Closes #1722.

Rationale for this change

Since #1401, the Server generates Scope IDs and rejects an explicit Scope that does not exist. The SWE-bench Pro console still passed the literal eval:{run_id}:{arm} to the Codex plugin and never created that Scope. The plugin's UserPromptSubmit hook got a 404 while resolving the Scope and exited without capturing or recalling. On current master, the ON arm therefore behaves like the OFF arm and fails treatment validation. The linked issue has the reproduction across 4e0f78ad, #1401 and master.

What changes are included in this PR?

  • Per-arm Scope. Once the Server is ready and before Codex starts, each arm creates its Scope through POST /v1/scopes inside the container. The API requires an idempotency key; the arm uses eval:{run_id}:{arm}.
  • One Scope ID end to end. The returned Scope ID is passed to Codex and is also used by the treatment evidence query and validate_treatment. The container-level POWERCONTEXT_CODEX_SCOPE_ID is removed: the ID only exists after the Server starts, and only codex exec runs the plugin hooks.
  • Retryable Scope creation failure. If the Server does not create the Scope, the arm fails with a readiness failure (scope_not_created). The worker already retries readiness failures with a fresh runtime, and it does so before Codex runs.
  • scope_key in evidence and reports. Treatment evidence records the key in a new optional scope_key field next to scope_id. The report check matches scope_key when present. When it is absent, which is the case for evidence recorded before this change (whose Scope ID was the key itself), the check falls back to scope_id, so existing runs keep validating. An empty scope_key is rejected, matching the harness's own evidence model.
  • Console schema. The console's strict evidence schema and TypeScript type accept scope_key. The API returns scope_key for every run, including null for earlier runs, so without this the console would reject the evidence of every report.

This PR does not change how mcp_requests is counted. The issue explains why that count can't be observed from the Server logs; it is left for a separate decision.

Are there any user-facing changes?

There are no CLI changes.

  • The console API's treatment evidence gains a scope_key field, and the console accepts it.
  • In treatment.json, scope_id is now a Server-generated scp_… ID.
  • A failed Scope creation shows up as a retryable readiness failure with the summary "PowerContext Server did not create the arm Scope."

How was this change tested?

  • Checks: make check and make unit-test pass (2841 passed, 58 skipped).
  • Evaluation tests: uv run pytest in evaluation/ gives 1052 passed. Five tests in tests/web/test_worker.py fail both on this branch and on unmodified upstream/master (1e90376a); they are unrelated to this change.
  • Console tests: in evaluation/web/, npx tsc -b and npm test pass (96 tests). The report contract test now uses the evidence shapes the API returns (a null key for an earlier run, a registered Scope for a new one). It failed before the schema change.
  • New contract tests:
    • Codex runs in the Scope the Server created for the arm, and that Scope is created before Codex starts.
    • A Scope creation failure raises a scope_not_created readiness failure before Codex runs.
    • With the Codex environment reverted to the literal key, the Scope test and the updated pair test fail.
  • Worker test: a Scope the Server did not create is a retryable readiness failure with a fixed summary.
  • Reporting tests:
    • Evidence for registered Scopes is accepted.
    • A mismatched or empty scope_key is rejected; with the empty-key rejection reverted, that case fails.
    • The existing legacy-shape fixtures still pass.
  • Manual check against a real master Server:
    • The Scope creation script returned an scp_… ID.
    • Repeating it with the same key returned the same ID.
    • The Codex hook captured one Source into that Scope.
  • Not run:
    • A full console run in Docker with a SWE-bench Pro image and a real Codex session. Inside the container, Scope creation uses the same interpreter (/runtime/pc-env/bin/python), loopback address and proxy environment (NO_PROXY includes loopback) as the existing readiness probe and evidence query. A full run would still stop at the report's mcp_requests > 0 check described in the issue.
    • The console Playwright suite. The local Playwright cache has a different browser build than the pinned version requires.

AI usage statement

This PR was developed with Claude Code (Claude Opus 5.5), which reproduced the regression, wrote the change and tests, and ran the validation above. The author reviewed the change.

🤖 Generated with Claude Code

Since oceanbase#1401 the Server generates Scope IDs and rejects an explicit Scope that
does not exist. The console still passed the literal `eval:{run_id}:{arm}` to
the Codex plugin, so its UserPromptSubmit hook got a 404 while resolving the
Scope and exited without capturing or recalling. The ON arm then behaved like
the OFF arm and failed treatment validation.

Each arm now creates its Scope through `POST /v1/scopes` once the Server is
ready, using `eval:{run_id}:{arm}` as the idempotency key, and passes the
returned ID to Codex. The evidence query and treatment validation use that ID,
and the evidence records the key as `scope_key` so reports can tie it to the
run. Evidence recorded before this change carries the key as its Scope ID and
still validates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@CLAassistant

CLAassistant commented Sep 23, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

Fengzdadi and others added 3 commits September 23, 2026 00:54
The report loader accepted `scope_key=""` and, because the check used `or`,
fell back to `scope_id` as if the evidence predated per-arm Scopes. The harness
never writes an empty key, since its own evidence model rejects one, but the
report layer re-validates stored evidence independently and should enforce the
same rule. Only an absent key now selects the earlier format.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The console parses treatment evidence with a strict schema, and the API now
returns `scope_key` for every run: `null` for evidence recorded before arms
registered their own Scope, and the run arm key otherwise. Without the field in
the schema, the console rejected the evidence of every report, including
earlier runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scope creation is part of preparing the Server for an arm, like the readiness
gate, but its failure was raised as a plain invalid treatment, which the worker
treats as terminal. A transient failure now raises a readiness failure with its
own reason, so the task is retried with a fresh runtime.

Also rename the key helper to `arm_scope_key`, so it no longer reads like the
evidence field of the same name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@PsiACE PsiACE left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@PsiACE
PsiACE merged commit ed54cc5 into oceanbase:master Sep 23, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug(evaluation): ON arm stops capturing after #1401 because the literal eval: scope does not exist

3 participants