Skip to content

refactor(evaluation): organize suites with independent commands - #1855

Open
Teingi wants to merge 5 commits into
oceanbase:masterfrom
Teingi:codex/unify-evaluation-layout
Open

Teingi wants to merge 5 commits into
oceanbase:masterfrom
Teingi:codex/unify-evaluation-layout

Conversation

@Teingi

@Teingi Teingi commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Which issue or RFC does this PR close?

No linked issue; this is a repository evaluation-layout cleanup.

Rationale for this change

Benchmarks and evaluation tools were split between benchmark/ and a shared evaluation package. Group them by what they measure, with each suite owning its commands, tests, documentation, and execution tools.

What changes are included in this PR?

  • Consolidate suites under evaluation/coding, evaluation/memory, evaluation/performance, and evaluation/skills. Keep the SWE console, worker, and deployment files inside SWE-bench Pro. Include the work-continuity benchmark from current master under coding.
  • Replace the shared Python package and command with independent SWE-bench Pro, LongMemEval-V2, and work-continuity packages and commands. Update build metadata, CI, ownership rules, imports, and documentation.
  • Fix the SWE console attempt-history decoder to accept the backend retry metadata. Align browser fixtures with batch creation, source resolution, quota gating, and automatic retry behavior.
  • Fix LoCoMo Scope registration and persisted Scope reuse, and preserve inference configuration between ingestion and evaluation. Make LoCoMo-Plus use configured settings.database with a local SQLite fallback; persist backend identity and reject cross-database resumes.
  • Add explicit HTTP opt-in for LongMemEval model endpoints, retaining HTTPS by default and rejecting redirects and URL credentials.
  • Support isolated OceanBase databases for SWE OFF/ON arms, record the actual backend, preserve backend metadata through Web report reads (including legacy evidence), and redact database credentials. Install only Server/CLI runtime dependencies and include timezone data. Keep the upstream SQLite import behavior and install the complete locked dependency set for both database backends. Reject SWE OceanBase URL query options that can override credentials or the selected database target before preparing a run; connection credentials and host/port/database belong in the URL authority and path.

Are there any user-facing changes?

Evaluation paths, Python imports, and commands change. Replace powercontext-eval ... with the documented swebench-pro, longmemeval-v2, or work-continuity command; LoCoMo and capacity modules now run under evaluation.*. There is no compatibility wrapper. Existing dataset contents and result schema identifiers are preserved. LoCoMo resumes require a valid saved Scope mapping, and LoCoMo-Plus resumes must match the recorded backend. OceanBase resume identity uses the effective database target and excludes authentication values. Saved OceanBase runs with an older fingerprint format remain replayable but require a new output directory for execution.

How was this change tested?

On the merged branch:

  • make check: passed, including hooks, lock consistency, type checks, generated contracts, and 30 integration-manifest tests.
  • Evaluation pytest -m "not live": 1,530 passed. make evaluation-unit-test also passed all 705 cases with GitHub Actions ANSI colors enabled; CLI assertions normalize formatting while retaining error and credential-redaction checks.
  • LoCoMo-Plus: 73 tests passed, including offline run_benchmark password rotation, effective target changes, and replay/resume guards using the installed OceanBase dialect. These checks do not connect to OceanBase or call an external model.
  • Runtime and SQLite vector regression tests: 14 passed; focused LoCoMo/Plus, Scope resume, and access-control checks passed; Bub harbor tests: 43 passed. Bub nested lock metadata includes the root timezone dependency, with all 122 locked package identities preserved. Locked sync, make harness-check (150 tests), and make harness-compose-check passed. The integrated upstream harness tests are preserved, with answer-directory protection targeting evaluation/. All 10 mount-protection cases passed and rejected an injected evaluation-directory mount; the proxy regression also passed with existing uppercase and lowercase HTTP/HTTPS proxy settings.
  • make docs-test with Node 22: passed, including static export and internal links across 875 public pages.
  • Skill-up pin and fixture validation plus 26 unit tests: passed.
  • SWE CLI, runner phases, and Codex contracts: 255 passed, including rejected credential/target query overrides and accepted encoded authority credentials. SWE console: 101 frontend tests, 3 Chromium end-to-end scenarios, and production build passed.
  • Evaluation Ruff, formatting, and type checks; sdist-to-wheel build; clean installation outside the checkout; all three CLI entry points and fixed work-continuity fixture validation/run/compare: passed.

Bounded live samples used runtime base 5cf66be6 plus a separate runtime-fix snapshot. The successful Alpine OceanBase startup samples used lazy SQLite extension imports; this PR retains upstream top-level imports and requires the complete locked dependency set, so those samples do not establish Alpine compatibility of this branch. Memory-capacity operations passed on real OceanBase. LoCoMo/Plus persisted real data; LongMemEval completed 10/10 FTS retrieval and 7/10 Reader answers. SWE Gold passed and both OFF and a separate ON-only attempt executed real model-generated commands against isolated OceanBase databases; ON ingestion persisted a Source. Provider account refusals and other upstream response errors prevented full scoring and a completed SWE pair. These samples do not establish complete benchmark accuracy or a successful OFF/ON comparison. Skill-up used a real host/model with mocked MCP. Temporary databases and credentials were cleaned up.

AI usage statement

Implemented and validated with OpenAI Codex (GPT-6), including parallel agents for migration and validation.

Teingi added 2 commits October 6, 2026 02:55
Consolidate coding, memory, capacity, and skill suites under evaluation and
keep each suite's commands, tests, documentation, and execution tools local.
Include the work-continuity benchmark from the latest master branch.

Fix issues found by bounded live smoke tests: Scope reuse, configured
database identity, HTTP model endpoints, and isolated OceanBase SWE runs.
Preserve backend evidence through report loading and align console retry
history with the current API and automatic retry behavior.
@Teingi
Teingi marked this pull request as ready for review October 5, 2026 22:53
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@Teingi Teingi left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 828ad2a. Three correctness issues remain in the new database configuration handling: credential protection, OFF/ON isolation, and resume identity. I recommend addressing the inline findings before merging.

Comment thread evaluation/memory/locomo_plus/runner.py
task = load_tasks(_REPOSITORY / "e2e" / "bub" / manifest)[0]
protected = [_REPOSITORY / "e2e" / "bub" / name for name in ("harbor-tasks", "paired-tasks", "tasks")]
protected.append(_REPOSITORY / "benchmark")
protected.append(_REPOSITORY / "evaluation")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Heads with #1853: e2e/bub/tests/test_harbor_job_config.py conflicts, and one side is a security-relevant path

This file is the only conflict between the two branches, but it is worth resolving deliberately rather than by taking one side. git merge-tree 673d44d6 <this head> 799ea12c reports exactly one changed in both:

  base   100644 b9fe8124...  e2e/bub/tests/test_harbor_job_config.py
  our    100644 293e50a6...  e2e/bub/tests/test_harbor_job_config.py
  their  100644 06a40cbe...  e2e/bub/tests/test_harbor_job_config.py

This PR changes one line there, and it is the one that matters:

-    protected.append(_REPOSITORY / "benchmark")
+    protected.append(_REPOSITORY / "evaluation")

That is the protected list handed to the agent container in test_agent_container_cannot_read_workload_answers. Since this PR moves benchmark/ to evaluation/, dropping that line would leave the answers directory readable by the agent under test — the assertion would still be testing something, but not the protection.

#1853 rewrites the same file (about 120 added lines in test_pi_runs_without_a_saved_session, plus imports and five other tests), so a textual conflict is unavoidable; the two changes do not conflict semantically. Whichever side lands first, please keep both: the protected entry pointing at evaluation, and #1853's new tests.

Also worth re-checking after the merge: #1853 has an open review comment on this file (4192361330) whose line anchors were computed against its own head.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in eb5b658 by merging upstream master. All of #1853's tests are preserved; the only difference from the upstream test file is the protected path changing from benchmark to evaluation. I also verified that the mount-configuration checks reject an injected evaluation/ mount and that the existing-proxy regression passes with uppercase and lowercase HTTP/HTTPS proxy variables set.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants