|
| 1 | +# GPT-6 Astra evaluation status — updated 2026-09-08 |
| 2 | + |
| 3 | +**No complete Astra OFF/ON benchmark has been executed. No token-saving claim is supported.** |
| 4 | + |
| 5 | +This directory records preflight facts, not benchmark results. Product integration tests and a |
| 6 | +successful model request do not demonstrate improved task quality or lower token use. |
| 7 | + |
| 8 | +The active experiment is [SWE-bench Pro](pro/README.md). Its complete 731-task manifest, |
| 9 | +protocol and first official reference-validation artifacts are published there. HumanEval+ |
| 10 | +preparation was withdrawn before any scored model inference. |
| 11 | + |
| 12 | +## Frozen public benchmark inputs |
| 13 | + |
| 14 | +- Dataset: [ScaleAI/SWE-bench_Pro](https://huggingface.co/datasets/ScaleAI/SWE-bench_Pro), |
| 15 | + revision `7ab5114912baf22bb098818e604c02fe7ad2c11f`, complete `test` split: **731 tasks**. |
| 16 | +- Evaluator: [scaleapi/SWE-bench_Pro-os](https://github.com/scaleapi/SWE-bench_Pro-os), |
| 17 | + commit `ca10a60a5fcae51e6948ffe1485d4153d421e6c5`. |
| 18 | +- A single complete OFF/ON comparison requires **1,462 independent task executions**, followed |
| 19 | + by official grading. It is not equivalent to 1,462 unit tests or grading empty patches. |
| 20 | +- Intended model: `gpt-6-astra`, with identical reasoning effort, prompts, tools, limits and |
| 21 | + independent state in both conditions. No task has been selected based on its model outcome. |
| 22 | + |
| 23 | +## Observed preflight |
| 24 | + |
| 25 | +| Check | Observation | |
| 26 | +| --- | --- | |
| 27 | +| Codex 0.147.0 + Astra | Provider rejected the request and required a newer Codex version | |
| 28 | +| Codex 0.153.4 + Astra | Minimal no-tool request completed successfully | |
| 29 | +| Successful request usage | 16,474 input; 11,520 cached input; 7 output; 0 reasoning output | |
| 30 | +| Docker | Existing local benchmark VM started; Docker 29.5.2, 6 CPUs, about 12 GiB RAM | |
| 31 | +| Existing task images | Three historical Lite smoke images retained; first Pro image downloaded | |
| 32 | +| First deterministic Pro image | Pinned in [Pro preflight](pro/preflight.json); actual OS is Alpine 3.18.3 | |
| 33 | +| Official Pro grading | First reference patch passed; 1,350.20 seconds under x86 QEMU | |
| 34 | +| Scored Astra task trajectories | **0 baseline / 0 MARGINAL** | |
| 35 | + |
| 36 | +The successful no-tool request is an API/CLI compatibility probe only. Its input token count |
| 37 | +includes the Codex scaffold and cannot be extrapolated into a reliable per-task cost estimate. |
| 38 | +Cached input and reasoning output are subsets, not additional tokens to sum twice. |
| 39 | + |
| 40 | +## Required before results can be published |
| 41 | + |
| 42 | +1. Adapt and validate the runner against current Codex and official Pro task environments. |
| 43 | + The historical Lite adapter is not a verified Pro execution backend. |
| 44 | +2. Pin images and prompts; validate the evaluator on reference patches in isolated containers. |
| 45 | +3. Publish the executable protocol before scored inference. The dataset pins above alone are |
| 46 | + **not** a complete preregistration. |
| 47 | +4. Run all 731 tasks in each condition without exposing reference patches or grading tests |
| 48 | + to the solving agent. Preserve failures, timeouts and every task ID. |
| 49 | +5. Publish sanitized predictions, per-task usage, official grading, intervention traces, |
| 50 | + environment hashes and analysis. Report correctness and governance overhead together. |
| 51 | + |
| 52 | +The sprint reserves 20% of the short Codex quota window for validation and publication. A |
| 53 | +quota stop must remain visibly incomplete; it cannot become a fabricated full benchmark, |
| 54 | +a smaller benchmark labelled Pro, or a claim that default Shadow Mode saves reasoning tokens. |
| 55 | + |
| 56 | +HumanEval+ was considered as a lower-cost alternative but has not been run or substituted |
| 57 | +for the requested real-repository comparison. No external leaderboard certification is claimed. |
0 commit comments