Add a complete agent-axis run (Fable 5 / Cursor agent × TensorCircuit-NG): 12/12 official reward 1.0, runtime-ratio protocol, and two infrastructure fixes - #4
Draft
QingyunQian wants to merge 46 commits into
Conversation
…, verify tools Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…d stamp provenance Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
… cloud precheck Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…(01, 02) Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…gyunQian/ORBIT-Q into cursor/fable5-benchmark-7148 Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…time ratio; challenge-03 stamp in summary Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…est environment Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…co integrated); re-measure all candidates and references Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…udit rejected Kraus-ladder v1 Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…k, runtime ratio 1.19 Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…ltiplier from reward (doc changed in ea088db, code never followed); template and all 12 task copies patched identically (upstream task generator source unavailable in this workspace) Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…04's stamped reward Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…DE evolution), precheck, ratio 1.09 Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…ltiplier from reward (doc changed in ea088db, code never followed); template and all 12 task copies patched identically (upstream task generator source unavailable in this workspace) Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
… 1.28; record 04 re-stamp and 06 stamp Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…io 2.64; record 07 stamp Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…ing-7148 Fix verifier/doc desync: remove deprecated runtime_score multiplier from reward in score_submission.py (policy declared in ea088db)
…analytic max), ratio 3.63; record 08 stamp Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…), ratio 18.06; record 09 stamp Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…, ratio 0.74; record 10 stamp Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
… record 11 stamp Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…summary figures and full report Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…l 1.11x); annotate summary and report Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…y post-freeze check); 0.74 is a matched-precision win; document per-task dtype matchups Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…ts the reference (0.76); c04's method requires double precision; report and figures updated Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…ntime-ratio protocol, matched-precision controls, report and figures Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…el control diamonds; user-approved push to main Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
Member
There was a problem hiding this comment.
this figure has no extra info than the other one?
…le agent resource-use figure (solve wall time + artifact runtime); record solve-time data with method note Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
…ol / Fable 5) with dataset and script; restore ratio-bars title Co-authored-by: QingyunQian <QingyunQian@users.noreply.github.com>
Author
|
As it is Cursor agent ,which means different harness for the task, so transfer to draft and dont merge now |
QingyunQian
marked this pull request as draft
August 3, 2026 06:14
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this PR does
1. A complete agent-axis benchmark run (results/fable5/)
Adds a full 12-task agent-axis data row: solver = Fable 5 via Cursor agent, framework fixed at TensorCircuit-NG. Every task passes the unmodified official verifier — functional evaluator, static policy, and Codex LLM audit (model
gpt-5.6-sol, recorded per task) — with reward = 1.0 on all 12 tasks. Per-task artifacts include the solution, verifier outputs, audit provenance, and a same-machine runtime comparison against the publication reference. A full report (results/fable5/REPORT.md) documents the pipeline, per-task observations, and reproducible summary figures.Harness caveat, stated upfront: the solver ran outside the Harbor agent container (Cursor agent operating on the workspace, verification through the documented verifier-only candidate-check path). Artifact-side metrics (pass components, runtime, T/T_ref) follow the standard verifier and are comparable; agent-side metrics (tokens, solve wall time) are not comparable with in-container rows. Solving-integrity protocol: only instruction.md and the framework prompt were read; tests/ and solution/ were never opened during solving; references were executed for timing only, after each candidate was frozen.
2. Fix: reward formula desync between docs and scorer
Since ea088db, AGENTS.md/README declare
reward = functional × static_policy × llm_audit("runtime … must not reduce pass reward"), butscore_submission.py(template + all 12 task copies) still multipliedruntime_scoreintoreward. Observed impact: a run with all three gating components at 1.0 and runtime_sec = 196.25 scored reward = 0.8645833… — exactly (300−196.25)/(300−180). This PR syncs the code with the declared policy (runtime_score/runtime_sec still computed and reported unchanged). Also available as a standalone atomic PR if preferred.3. Fix: framework pin predates the omeco integration
frameworks/tensorcircuit/requirements.txtpinnedtensorcircuit-nightly==1.7.0.dev20260618, which predatesset_contractor("omeco")landing in master, so the publication references for tasks 01/05 crash on the public image (the omeco package itself was already listed). Bumped to1.8.0.dev20260726: all 12 references now run unmodified; the four candidates already measured were re-verified on the new pin (all still functional 1.0).What we found
Note for maintainers
gpt-5togpt-5.6-solduring the run (the ChatGPT-auth catalog no longer serves the baregpt-5slug). Keep or revert as you prefer.results/fable5/placement and naming are proposals; happy to move/rename if you want a different results layout.