Skip to content

Commit bbe47d2

Browse files
Add GPT-6 Astra memory support and reproducible SWE-bench Pro infrastructure (#84)
* Add Astra model identity and operation memory * Package Astra support and document benchmark preflight honestly * Add resumable paired benchmark orchestration * Restore SWE-bench Pro target and record full task protocol * Publish deterministic 731-task SWE-bench Pro manifest * Add minimal SWE-bench Pro container lane * Reject nested Git history in Pro lanes * Bind benchmark inputs to protocol metadata * Use musl runtimes for Pro task image * Use musl runtimes for Pro task images * Publish official SWE-bench Pro reference validation evidence * Use Alpine Node in Pro runtime * Link active Pro experiment and pin evaluator dependencies * Clarify Astra support and distinguish Pro reference checks from model scores * Use packaged native Codex in Pro solver * Include Pro Codex launcher in solver context * Format Pro solver regression test * Record verified Pro runtime and incomplete benchmark closeout * Fix SWE-bench Pro instance ID validation * Pin corrected Pro execution module and pre-inference invocation * Preserve ignored tracked files in Pro snapshots * Publish exact-tree verification and explicit incomplete Pro coverage * Use literal pathspecs for Pro snapshot staging * Avoid unnecessary contention for atomic outbox writes * Preregister full paired SWE-bench Lite campaign
1 parent 5ef1106 commit bbe47d2

56 files changed

Lines changed: 3390 additions & 49 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.dockerignore‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -9,5 +9,9 @@ benchmark/*
99
!benchmark/__init__.py
1010
!benchmark/codex_adapter/
1111
!benchmark/codex_adapter/**
12+
!benchmark/astra/
13+
benchmark/astra/*
14+
!benchmark/astra/__init__.py
15+
!benchmark/astra/pro_lane.py
1216
!benchmark/container/
1317
!benchmark/container/**
Lines changed: 49 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,49 @@
1+
# SWE-bench Pro Container Lane Report
2+
3+
Implemented the minimal in-container Pro solver lane without running model inference or changing
4+
the benchmark protocol.
5+
6+
## Delivered
7+
8+
- `benchmark/astra/pro_lane.py`: validates the instance ID, exact lowercase base commit, condition,
9+
fixed input paths, repository root, and non-existing output before mutation. It archives the base
10+
tree, retains ignored installed dependencies, removes future/current tracked content and original
11+
Git metadata, creates a deterministic detached single-commit snapshot, verifies tree equality and
12+
history size, rejects base trees containing gitlinks before mutation, purges nested Git metadata,
13+
forwards the unchanged public prompt to `run_task`, and records original/snapshot commit and tree
14+
hashes.
15+
- `benchmark/astra/pro/Dockerfile.solver`: layers on a controller-supplied task image, uses pinned
16+
amd64 Alpine Node and uv builder manifests, installs Codex 0.153.4 with the existing wrapper,
17+
adds an independent musl Python 3.12.11 runtime, installs MARGINAL/benchmark packages into that
18+
runtime's site-packages, records source/task/tool labels, and invokes the lane CLI from `/app`.
19+
- `.dockerignore`: admits only the Astra package files required by the solver image.
20+
- `tests/benchmark/test_astra_pro_lane.py`: focused repository-isolation, validation, overwrite, and
21+
forwarding tests.
22+
23+
## Verification
24+
25+
```text
26+
$ .venv/bin/python -m pytest -q tests/benchmark/test_astra_pro_lane.py
27+
...... [100%]
28+
7 passed in 2.24s
29+
```
30+
31+
```text
32+
$ .venv/bin/ruff check benchmark/astra/pro_lane.py tests/benchmark/test_astra_pro_lane.py
33+
All checks passed!
34+
```
35+
36+
The image build uses task image
37+
`jefzda/sweap-images@sha256:d902632d1374cf0282a4ea301b82c296e13a41127308da0204aca87a4ba62c02`.
38+
Despite earlier evaluator metadata describing Ubuntu 20.04/Python 3.8, this exact image is Alpine
39+
Linux 3.18.3. The solver therefore uses musl-compatible Node and standalone Python binaries. Its
40+
final image ID and no-inference executable smoke checks are recorded below once complete.
41+
42+
## Concerns
43+
44+
- Repository sanitization is intentionally destructive and is safe only because the official task
45+
container is disposable.
46+
- Ignored dependency retention depends on the base tree's ignore rules. The verified snapshot tree
47+
contains only tracked base-commit files; ignored dependency content remains outside Git.
48+
- Base commits containing gitlinks are rejected as unsupported infrastructure before mutation.
49+
- No solver/model request was made during implementation or verification.

‎README.md‎

Lines changed: 21 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -68,6 +68,27 @@ marginal install codex --autopilot-consent
6868

6969
## How MARGINAL earns authority
7070

71+
### GPT-6 Astra
72+
73+
The exact model ID `gpt-6-astra` has its own evidence namespace, `openai/gpt-6-astra`.
74+
When Codex hooks identify that model, decisions and outcomes are recorded locally as work
75+
happens. Existing unlabelled evidence is not reassigned to Astra. Commons sharing remains opt-in.
76+
77+
A real Astra request was verified with Codex CLI **0.153.4**; **0.147.0** was rejected by the
78+
provider. This is a tested version, not a claim that every intermediate version is unsupported.
79+
80+
**Astra token savings are not yet demonstrated.** Shadow Mode observes work; it does not
81+
reduce the model's internal reasoning budget. The complete public OFF/ON evaluation is still
82+
outstanding. See [Astra evaluation status](benchmark/astra/README.md) for exact preflight facts
83+
and what remains before a performance claim is justified.
84+
85+
The [SWE-bench Pro protocol and evidence](benchmark/astra/pro/README.md) include all 731
86+
planned task IDs and the first passing official reference check. That reference check is
87+
not a model score. A complete OFF/ON run requires 1,462 model executions; no complete
88+
performance result is claimed until those executions and their grading are available.
89+
90+
### Evidence gates
91+
7192
1. **Observe** — collect derived action, outcome, coverage, state and evidence signals locally.
7293
2. **Verify** — bind decisions, policy identity, trust state and governance cost into Decision Receipts.
7394
3. **Earn** — require representative local evidence, clean coverage and explicit promotion.

‎benchmark/astra/README.md‎

Lines changed: 57 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,57 @@
1+
# GPT-6 Astra evaluation status — updated 2026-09-08
2+
3+
**No complete Astra OFF/ON benchmark has been executed. No token-saving claim is supported.**
4+
5+
This directory records preflight facts, not benchmark results. Product integration tests and a
6+
successful model request do not demonstrate improved task quality or lower token use.
7+
8+
The active experiment is [SWE-bench Pro](pro/README.md). Its complete 731-task manifest,
9+
protocol and first official reference-validation artifacts are published there. HumanEval+
10+
preparation was withdrawn before any scored model inference.
11+
12+
## Frozen public benchmark inputs
13+
14+
- Dataset: [ScaleAI/SWE-bench_Pro](https://huggingface.co/datasets/ScaleAI/SWE-bench_Pro),
15+
revision `7ab5114912baf22bb098818e604c02fe7ad2c11f`, complete `test` split: **731 tasks**.
16+
- Evaluator: [scaleapi/SWE-bench_Pro-os](https://github.com/scaleapi/SWE-bench_Pro-os),
17+
commit `ca10a60a5fcae51e6948ffe1485d4153d421e6c5`.
18+
- A single complete OFF/ON comparison requires **1,462 independent task executions**, followed
19+
by official grading. It is not equivalent to 1,462 unit tests or grading empty patches.
20+
- Intended model: `gpt-6-astra`, with identical reasoning effort, prompts, tools, limits and
21+
independent state in both conditions. No task has been selected based on its model outcome.
22+
23+
## Observed preflight
24+
25+
| Check | Observation |
26+
| --- | --- |
27+
| Codex 0.147.0 + Astra | Provider rejected the request and required a newer Codex version |
28+
| Codex 0.153.4 + Astra | Minimal no-tool request completed successfully |
29+
| Successful request usage | 16,474 input; 11,520 cached input; 7 output; 0 reasoning output |
30+
| Docker | Existing local benchmark VM started; Docker 29.5.2, 6 CPUs, about 12 GiB RAM |
31+
| Existing task images | Three historical Lite smoke images retained; first Pro image downloaded |
32+
| First deterministic Pro image | Pinned in [Pro preflight](pro/preflight.json); actual OS is Alpine 3.18.3 |
33+
| Official Pro grading | First reference patch passed; 1,350.20 seconds under x86 QEMU |
34+
| Scored Astra task trajectories | **0 baseline / 0 MARGINAL** |
35+
36+
The successful no-tool request is an API/CLI compatibility probe only. Its input token count
37+
includes the Codex scaffold and cannot be extrapolated into a reliable per-task cost estimate.
38+
Cached input and reasoning output are subsets, not additional tokens to sum twice.
39+
40+
## Required before results can be published
41+
42+
1. Adapt and validate the runner against current Codex and official Pro task environments.
43+
The historical Lite adapter is not a verified Pro execution backend.
44+
2. Pin images and prompts; validate the evaluator on reference patches in isolated containers.
45+
3. Publish the executable protocol before scored inference. The dataset pins above alone are
46+
**not** a complete preregistration.
47+
4. Run all 731 tasks in each condition without exposing reference patches or grading tests
48+
to the solving agent. Preserve failures, timeouts and every task ID.
49+
5. Publish sanitized predictions, per-task usage, official grading, intervention traces,
50+
environment hashes and analysis. Report correctness and governance overhead together.
51+
52+
The sprint reserves 20% of the short Codex quota window for validation and publication. A
53+
quota stop must remain visibly incomplete; it cannot become a fabricated full benchmark,
54+
a smaller benchmark labelled Pro, or a claim that default Shadow Mode saves reasoning tokens.
55+
56+
HumanEval+ was considered as a lower-cost alternative but has not been run or substituted
57+
for the requested real-repository comparison. No external leaderboard certification is claimed.

‎benchmark/astra/__init__.py‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
1+
"""Astra public-benchmark orchestration."""
Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,11 @@
1+
FROM ganler/evalplus@sha256:26b118098bef281fe8dfe999bf05f1d5b45374b4e6c00161ec0f30592aef4740
2+
3+
COPY evaluator-requirements.txt /tmp/evaluator-requirements.txt
4+
RUN python -m pip install --no-deps --require-hashes -r /tmp/evaluator-requirements.txt \
5+
&& python -c 'import importlib.metadata; assert importlib.metadata.version("evalplus") == "0.3.1"'
6+
7+
ENV PYTHONDONTWRITEBYTECODE=1 \
8+
XDG_CACHE_HOME=/tmp/evalplus-cache \
9+
HUMANEVAL_OVERRIDE_PATH=/data/HumanEvalPlus.jsonl
10+
11+
ENTRYPOINT ["evalplus.evaluate"]
Lines changed: 51 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,51 @@
1+
# Astra × MARGINAL — full HumanEval+ protocol
2+
3+
**Withdrawn before scored inference.** The user requires SWE-bench Pro. No model samples
4+
were generated for this suite. The reference-only evaluator check passed 163/164 tasks;
5+
its remaining failure was not investigated after the scope correction. This directory is
6+
retained for traceability, not as a performance result or an active preregistration.
7+
8+
This is a **function-level** OFF/ON experiment, not SWE-bench Pro. No result is claimed
9+
until all 164 tasks have both a real model run and official grading.
10+
11+
## What is being tested
12+
13+
The same Codex 0.153.4 agent and GPT-6 Astra (`low` reasoning) implement every HumanEval+
14+
v0.1.10 function twice, in separate clean workspaces. Baseline has no MARGINAL benchmark
15+
adapter; ON uses the existing balanced/diminishing-return experimental enforcement adapter.
16+
The installed plugin's default Shadow Mode is **not** this intervention.
17+
18+
Each task starts with only its public function prompt in `solution.py`. Reference solutions
19+
and official base/extra test inputs are not provided to the solver. The solver may perform
20+
local checks, but may not use the network or files outside its workspace. Execution traces
21+
must be audited for violations before any results are admitted. No extra repetition is
22+
requested to make MARGINAL look effective.
23+
24+
## Frozen conditions
25+
26+
The machine-readable [protocol](protocol.json) records task/data hashes, order, model,
27+
reasoning, timeout, concurrency, analysis and stop rules. Runner source revision, evaluator
28+
image digest and exact prompt hash are recorded and published before scored inference.
29+
All original tasks and all extra tests are included; neither `mini` nor `noextreme` is used.
30+
31+
Correctness is primary: official base **and** extra tests must pass. Failed, timed-out,
32+
missing-usage and infrastructure-invalid lanes remain visible. They are never silently
33+
replaced with another attempt. The complete experiment requires 328 recorded lanes.
34+
35+
Token accounting uses provider-reported totals and includes all calls, including recovery
36+
work. Cached input is part of input, and reasoning output is part of output. They are not
37+
added twice. Subscription quota is not a dollar bill. Unavailable costs stay unavailable.
38+
39+
## Reading the eventual result
40+
41+
Report paired correctness, token use, tool calls, elapsed time, intervention counts and
42+
governance overhead together. Report overall results and the both-solve subset separately.
43+
A zero-regression sample is an observation, not a statistical proof of equivalence. One
44+
sample per task cannot establish production reliability; a public function benchmark may
45+
also be saturated or present in model training data.
46+
47+
Negative, zero-effect and inconclusive results are valid. This experiment cannot justify
48+
claiming that installing the default plugin reduces Astra's internal reasoning tokens.
49+
50+
Sources: [EvalPlus v0.3.1](https://github.com/evalplus/evalplus/tree/e5d0ed0bab96280b60b637ec7f15b5e4841b0cb2),
51+
[HumanEval+ v0.1.10](https://github.com/evalplus/humanevalplus_release/releases/tag/v0.1.10).
Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
1+
evalplus==0.3.1 --hash=sha256:cd601debb67419113d10ac5c3317689d847f27de5d8cf3837975f3cab571b75d
Lines changed: 41 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,41 @@
1+
{
2+
"schema_version": 1,
3+
"experiment": "astra-humaneval-plus-20260907",
4+
"status": "withdrawn_before_scored_inference_user_requires_swe_bench_pro",
5+
"dataset": "HumanEval+",
6+
"dataset_version": "v0.1.10",
7+
"dataset_sha256": "42526ec0e7d5f3ee0b06d6ced98f8c8bae3d76519151bfb3d36f79010645bd7f",
8+
"task_inputs_sha256": "8fbac146a932a0e1bdbccf8b0596061d43312dc20143960c48b66029d2dbf425",
9+
"task_count": 164,
10+
"conditions": ["baseline", "marginal"],
11+
"repetitions": 1,
12+
"model": "gpt-6-astra",
13+
"model_snapshot": "public alias; immutable provider weight snapshot unavailable",
14+
"codex_version": "0.153.4",
15+
"evaluator_version": "evalplus==0.3.1",
16+
"evaluator_image_id": "sha256:a90b5cd691de7334d203b64435aaf1ef10701acf90eeb4455703df3724b5e6aa",
17+
"reasoning_effort": "low",
18+
"timeout_seconds": 180,
19+
"concurrency": 2,
20+
"task_order": "ascending sha256(task_id UTF-8)",
21+
"condition_order": "baseline first at even task positions, marginal first at odd positions",
22+
"baseline": "Codex without MARGINAL benchmark hooks or daemon",
23+
"marginal": "existing benchmark adapter: balanced policy, default diminishing-return detector, experimental enforce, unlimited hard budgets",
24+
"native_shadow_equivalent": false,
25+
"reference_tests_available_to_solver": false,
26+
"solutions_per_task_condition": 1,
27+
"quality_primary": "official HumanEval+ base AND extra tests pass@1",
28+
"failures": "retain every task; no re-solving after inspecting grading; timeouts and invalid infrastructure reported separately",
29+
"quality_gate": "no observed baseline-only regressions; descriptive criterion, not a non-inferiority proof",
30+
"efficiency": "paired total input+output tokens, cached-input subset, reasoning-output subset, tool calls, wall time, governance overhead and interventions",
31+
"cost_usd": "unavailable unless directly reported; no invented zero or API bill for subscription usage",
32+
"analysis": "per-task paired results and 95% task-bootstrap interval for token savings; report both-solve separately; do not infer equivalence from non-significance",
33+
"stop_rules": "stop new lanes on missing usage, hook failures, invalid isolation or short quota >=80%; retain partial artifacts and label incomplete",
34+
"claim_limits": [
35+
"Function-level coding tasks, not real-repository SWE-bench Pro",
36+
"Experimental benchmark enforcement, not default installed Shadow Mode",
37+
"One stochastic sample per task per condition",
38+
"Public benchmark contamination and saturation are possible",
39+
"Self-published evaluation, not external leaderboard certification"
40+
]
41+
}

0 commit comments

Comments
 (0)