Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
6a7a841
Add Astra model identity and operation memory
SignalLayerLabs Sep 7, 2026
821eb2b
Package Astra support and document benchmark preflight honestly
SignalLayerLabs Sep 7, 2026
10764b8
Add resumable paired benchmark orchestration
SignalLayerLabs Sep 7, 2026
579fa1a
Restore SWE-bench Pro target and record full task protocol
SignalLayerLabs Sep 7, 2026
6d5f3f3
Publish deterministic 731-task SWE-bench Pro manifest
SignalLayerLabs Sep 7, 2026
fb3c4a3
Add minimal SWE-bench Pro container lane
SignalLayerLabs Sep 7, 2026
2dc8fb3
Reject nested Git history in Pro lanes
SignalLayerLabs Sep 7, 2026
529f41b
Bind benchmark inputs to protocol metadata
SignalLayerLabs Sep 7, 2026
1d46b85
Use musl runtimes for Pro task image
SignalLayerLabs Sep 7, 2026
4ad6fc4
Use musl runtimes for Pro task images
SignalLayerLabs Sep 7, 2026
8743d3c
Publish official SWE-bench Pro reference validation evidence
SignalLayerLabs Sep 7, 2026
2097391
Use Alpine Node in Pro runtime
SignalLayerLabs Sep 7, 2026
1cf8145
Link active Pro experiment and pin evaluator dependencies
SignalLayerLabs Sep 7, 2026
91b5ff8
Clarify Astra support and distinguish Pro reference checks from model…
SignalLayerLabs Sep 8, 2026
6dcc81d
Use packaged native Codex in Pro solver
SignalLayerLabs Sep 8, 2026
a167ab0
Include Pro Codex launcher in solver context
SignalLayerLabs Sep 8, 2026
6470794
Format Pro solver regression test
SignalLayerLabs Sep 8, 2026
1d22d7f
Record verified Pro runtime and incomplete benchmark closeout
SignalLayerLabs Sep 8, 2026
e416154
Fix SWE-bench Pro instance ID validation
SignalLayerLabs Sep 9, 2026
f80f36e
Pin corrected Pro execution module and pre-inference invocation
SignalLayerLabs Sep 9, 2026
d9b167f
Preserve ignored tracked files in Pro snapshots
SignalLayerLabs Sep 9, 2026
5a1a94c
Publish exact-tree verification and explicit incomplete Pro coverage
SignalLayerLabs Sep 9, 2026
553c2ac
Use literal pathspecs for Pro snapshot staging
SignalLayerLabs Sep 9, 2026
71c8eae
Avoid unnecessary contention for atomic outbox writes
SignalLayerLabs Sep 9, 2026
64d59cf
Preregister full paired SWE-bench Lite campaign
SignalLayerLabs Sep 9, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .dockerignore
Original file line number Diff line number Diff line change
Expand Up @@ -9,5 +9,9 @@ benchmark/*
!benchmark/__init__.py
!benchmark/codex_adapter/
!benchmark/codex_adapter/**
!benchmark/astra/
benchmark/astra/*
!benchmark/astra/__init__.py
!benchmark/astra/pro_lane.py
!benchmark/container/
!benchmark/container/**
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
# SWE-bench Pro Container Lane Report

Implemented the minimal in-container Pro solver lane without running model inference or changing
the benchmark protocol.

## Delivered

- `benchmark/astra/pro_lane.py`: validates the instance ID, exact lowercase base commit, condition,
fixed input paths, repository root, and non-existing output before mutation. It archives the base
tree, retains ignored installed dependencies, removes future/current tracked content and original
Git metadata, creates a deterministic detached single-commit snapshot, verifies tree equality and
history size, rejects base trees containing gitlinks before mutation, purges nested Git metadata,
forwards the unchanged public prompt to `run_task`, and records original/snapshot commit and tree
hashes.
- `benchmark/astra/pro/Dockerfile.solver`: layers on a controller-supplied task image, uses pinned
amd64 Alpine Node and uv builder manifests, installs Codex 0.153.4 with the existing wrapper,
adds an independent musl Python 3.12.11 runtime, installs MARGINAL/benchmark packages into that
runtime's site-packages, records source/task/tool labels, and invokes the lane CLI from `/app`.
- `.dockerignore`: admits only the Astra package files required by the solver image.
- `tests/benchmark/test_astra_pro_lane.py`: focused repository-isolation, validation, overwrite, and
forwarding tests.

## Verification

```text
$ .venv/bin/python -m pytest -q tests/benchmark/test_astra_pro_lane.py
...... [100%]
7 passed in 2.24s
```

```text
$ .venv/bin/ruff check benchmark/astra/pro_lane.py tests/benchmark/test_astra_pro_lane.py
All checks passed!
```

The image build uses task image
`jefzda/sweap-images@sha256:d902632d1374cf0282a4ea301b82c296e13a41127308da0204aca87a4ba62c02`.
Despite earlier evaluator metadata describing Ubuntu 20.04/Python 3.8, this exact image is Alpine
Linux 3.18.3. The solver therefore uses musl-compatible Node and standalone Python binaries. Its
final image ID and no-inference executable smoke checks are recorded below once complete.

## Concerns

- Repository sanitization is intentionally destructive and is safe only because the official task
container is disposable.
- Ignored dependency retention depends on the base tree's ignore rules. The verified snapshot tree
contains only tracked base-commit files; ignored dependency content remains outside Git.
- Base commits containing gitlinks are rejected as unsupported infrastructure before mutation.
- No solver/model request was made during implementation or verification.
21 changes: 21 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,27 @@ marginal install codex --autopilot-consent

## How MARGINAL earns authority

### GPT-6 Astra

The exact model ID `gpt-6-astra` has its own evidence namespace, `openai/gpt-6-astra`.
When Codex hooks identify that model, decisions and outcomes are recorded locally as work
happens. Existing unlabelled evidence is not reassigned to Astra. Commons sharing remains opt-in.

A real Astra request was verified with Codex CLI **0.153.4**; **0.147.0** was rejected by the
provider. This is a tested version, not a claim that every intermediate version is unsupported.

**Astra token savings are not yet demonstrated.** Shadow Mode observes work; it does not
reduce the model's internal reasoning budget. The complete public OFF/ON evaluation is still
outstanding. See [Astra evaluation status](benchmark/astra/README.md) for exact preflight facts
and what remains before a performance claim is justified.

The [SWE-bench Pro protocol and evidence](benchmark/astra/pro/README.md) include all 731
planned task IDs and the first passing official reference check. That reference check is
not a model score. A complete OFF/ON run requires 1,462 model executions; no complete
performance result is claimed until those executions and their grading are available.

### Evidence gates

1. **Observe** — collect derived action, outcome, coverage, state and evidence signals locally.
2. **Verify** — bind decisions, policy identity, trust state and governance cost into Decision Receipts.
3. **Earn** — require representative local evidence, clean coverage and explicit promotion.
Expand Down
57 changes: 57 additions & 0 deletions benchmark/astra/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
# GPT-6 Astra evaluation status — updated 2026-09-08

**No complete Astra OFF/ON benchmark has been executed. No token-saving claim is supported.**

This directory records preflight facts, not benchmark results. Product integration tests and a
successful model request do not demonstrate improved task quality or lower token use.

The active experiment is [SWE-bench Pro](pro/README.md). Its complete 731-task manifest,
protocol and first official reference-validation artifacts are published there. HumanEval+
preparation was withdrawn before any scored model inference.

## Frozen public benchmark inputs

- Dataset: [ScaleAI/SWE-bench_Pro](https://huggingface.co/datasets/ScaleAI/SWE-bench_Pro),
revision `7ab5114912baf22bb098818e604c02fe7ad2c11f`, complete `test` split: **731 tasks**.
- Evaluator: [scaleapi/SWE-bench_Pro-os](https://github.com/scaleapi/SWE-bench_Pro-os),
commit `ca10a60a5fcae51e6948ffe1485d4153d421e6c5`.
- A single complete OFF/ON comparison requires **1,462 independent task executions**, followed
by official grading. It is not equivalent to 1,462 unit tests or grading empty patches.
- Intended model: `gpt-6-astra`, with identical reasoning effort, prompts, tools, limits and
independent state in both conditions. No task has been selected based on its model outcome.

## Observed preflight

| Check | Observation |
| --- | --- |
| Codex 0.147.0 + Astra | Provider rejected the request and required a newer Codex version |
| Codex 0.153.4 + Astra | Minimal no-tool request completed successfully |
| Successful request usage | 16,474 input; 11,520 cached input; 7 output; 0 reasoning output |
| Docker | Existing local benchmark VM started; Docker 29.5.2, 6 CPUs, about 12 GiB RAM |
| Existing task images | Three historical Lite smoke images retained; first Pro image downloaded |
| First deterministic Pro image | Pinned in [Pro preflight](pro/preflight.json); actual OS is Alpine 3.18.3 |
| Official Pro grading | First reference patch passed; 1,350.20 seconds under x86 QEMU |
| Scored Astra task trajectories | **0 baseline / 0 MARGINAL** |

The successful no-tool request is an API/CLI compatibility probe only. Its input token count
includes the Codex scaffold and cannot be extrapolated into a reliable per-task cost estimate.
Cached input and reasoning output are subsets, not additional tokens to sum twice.

## Required before results can be published

1. Adapt and validate the runner against current Codex and official Pro task environments.
The historical Lite adapter is not a verified Pro execution backend.
2. Pin images and prompts; validate the evaluator on reference patches in isolated containers.
3. Publish the executable protocol before scored inference. The dataset pins above alone are
**not** a complete preregistration.
4. Run all 731 tasks in each condition without exposing reference patches or grading tests
to the solving agent. Preserve failures, timeouts and every task ID.
5. Publish sanitized predictions, per-task usage, official grading, intervention traces,
environment hashes and analysis. Report correctness and governance overhead together.

The sprint reserves 20% of the short Codex quota window for validation and publication. A
quota stop must remain visibly incomplete; it cannot become a fabricated full benchmark,
a smaller benchmark labelled Pro, or a claim that default Shadow Mode saves reasoning tokens.

HumanEval+ was considered as a lower-cost alternative but has not been run or substituted
for the requested real-repository comparison. No external leaderboard certification is claimed.
1 change: 1 addition & 0 deletions benchmark/astra/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
"""Astra public-benchmark orchestration."""
11 changes: 11 additions & 0 deletions benchmark/astra/humaneval-plus/Dockerfile.evaluator
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
FROM ganler/evalplus@sha256:26b118098bef281fe8dfe999bf05f1d5b45374b4e6c00161ec0f30592aef4740

COPY evaluator-requirements.txt /tmp/evaluator-requirements.txt
RUN python -m pip install --no-deps --require-hashes -r /tmp/evaluator-requirements.txt \
&& python -c 'import importlib.metadata; assert importlib.metadata.version("evalplus") == "0.3.1"'

ENV PYTHONDONTWRITEBYTECODE=1 \
XDG_CACHE_HOME=/tmp/evalplus-cache \
HUMANEVAL_OVERRIDE_PATH=/data/HumanEvalPlus.jsonl

ENTRYPOINT ["evalplus.evaluate"]
51 changes: 51 additions & 0 deletions benchmark/astra/humaneval-plus/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# Astra × MARGINAL — full HumanEval+ protocol

**Withdrawn before scored inference.** The user requires SWE-bench Pro. No model samples
were generated for this suite. The reference-only evaluator check passed 163/164 tasks;
its remaining failure was not investigated after the scope correction. This directory is
retained for traceability, not as a performance result or an active preregistration.

This is a **function-level** OFF/ON experiment, not SWE-bench Pro. No result is claimed
until all 164 tasks have both a real model run and official grading.

## What is being tested

The same Codex 0.153.4 agent and GPT-6 Astra (`low` reasoning) implement every HumanEval+
v0.1.10 function twice, in separate clean workspaces. Baseline has no MARGINAL benchmark
adapter; ON uses the existing balanced/diminishing-return experimental enforcement adapter.
The installed plugin's default Shadow Mode is **not** this intervention.

Each task starts with only its public function prompt in `solution.py`. Reference solutions
and official base/extra test inputs are not provided to the solver. The solver may perform
local checks, but may not use the network or files outside its workspace. Execution traces
must be audited for violations before any results are admitted. No extra repetition is
requested to make MARGINAL look effective.

## Frozen conditions

The machine-readable [protocol](protocol.json) records task/data hashes, order, model,
reasoning, timeout, concurrency, analysis and stop rules. Runner source revision, evaluator
image digest and exact prompt hash are recorded and published before scored inference.
All original tasks and all extra tests are included; neither `mini` nor `noextreme` is used.

Correctness is primary: official base **and** extra tests must pass. Failed, timed-out,
missing-usage and infrastructure-invalid lanes remain visible. They are never silently
replaced with another attempt. The complete experiment requires 328 recorded lanes.

Token accounting uses provider-reported totals and includes all calls, including recovery
work. Cached input is part of input, and reasoning output is part of output. They are not
added twice. Subscription quota is not a dollar bill. Unavailable costs stay unavailable.

## Reading the eventual result

Report paired correctness, token use, tool calls, elapsed time, intervention counts and
governance overhead together. Report overall results and the both-solve subset separately.
A zero-regression sample is an observation, not a statistical proof of equivalence. One
sample per task cannot establish production reliability; a public function benchmark may
also be saturated or present in model training data.

Negative, zero-effect and inconclusive results are valid. This experiment cannot justify
claiming that installing the default plugin reduces Astra's internal reasoning tokens.

Sources: [EvalPlus v0.3.1](https://github.com/evalplus/evalplus/tree/e5d0ed0bab96280b60b637ec7f15b5e4841b0cb2),
[HumanEval+ v0.1.10](https://github.com/evalplus/humanevalplus_release/releases/tag/v0.1.10).
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
evalplus==0.3.1 --hash=sha256:cd601debb67419113d10ac5c3317689d847f27de5d8cf3837975f3cab571b75d
41 changes: 41 additions & 0 deletions benchmark/astra/humaneval-plus/protocol.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
{
"schema_version": 1,
"experiment": "astra-humaneval-plus-20260907",
"status": "withdrawn_before_scored_inference_user_requires_swe_bench_pro",
"dataset": "HumanEval+",
"dataset_version": "v0.1.10",
"dataset_sha256": "42526ec0e7d5f3ee0b06d6ced98f8c8bae3d76519151bfb3d36f79010645bd7f",
"task_inputs_sha256": "8fbac146a932a0e1bdbccf8b0596061d43312dc20143960c48b66029d2dbf425",
"task_count": 164,
"conditions": ["baseline", "marginal"],
"repetitions": 1,
"model": "gpt-6-astra",
"model_snapshot": "public alias; immutable provider weight snapshot unavailable",
"codex_version": "0.153.4",
"evaluator_version": "evalplus==0.3.1",
"evaluator_image_id": "sha256:a90b5cd691de7334d203b64435aaf1ef10701acf90eeb4455703df3724b5e6aa",
"reasoning_effort": "low",
"timeout_seconds": 180,
"concurrency": 2,
"task_order": "ascending sha256(task_id UTF-8)",
"condition_order": "baseline first at even task positions, marginal first at odd positions",
"baseline": "Codex without MARGINAL benchmark hooks or daemon",
"marginal": "existing benchmark adapter: balanced policy, default diminishing-return detector, experimental enforce, unlimited hard budgets",
"native_shadow_equivalent": false,
"reference_tests_available_to_solver": false,
"solutions_per_task_condition": 1,
"quality_primary": "official HumanEval+ base AND extra tests pass@1",
"failures": "retain every task; no re-solving after inspecting grading; timeouts and invalid infrastructure reported separately",
"quality_gate": "no observed baseline-only regressions; descriptive criterion, not a non-inferiority proof",
"efficiency": "paired total input+output tokens, cached-input subset, reasoning-output subset, tool calls, wall time, governance overhead and interventions",
"cost_usd": "unavailable unless directly reported; no invented zero or API bill for subscription usage",
"analysis": "per-task paired results and 95% task-bootstrap interval for token savings; report both-solve separately; do not infer equivalence from non-significance",
"stop_rules": "stop new lanes on missing usage, hook failures, invalid isolation or short quota >=80%; retain partial artifacts and label incomplete",
"claim_limits": [
"Function-level coding tasks, not real-repository SWE-bench Pro",
"Experimental benchmark enforcement, not default installed Shadow Mode",
"One stochastic sample per task per condition",
"Public benchmark contamination and saturation are possible",
"Self-published evaluation, not external leaderboard certification"
]
}
Loading
Loading