Skip to content

Add the development pilot record, pilot-readiness engineering and evaluator repairs - #3

Merged
ammarphp merged 2 commits into
mainfrom
evaluation-study-3
Oct 4, 2026
Merged

ammarphp merged 2 commits into
mainfrom
evaluation-study-3

Conversation

@ammarphp

@ammarphp ammarphp commented Oct 4, 2026

Copy link
Copy Markdown
Owner

What this adds

  • Development pilot record (docs/development/evaluation-study/pilot-record.md). 96 assignments: 12 task-bank tasks × 2 seeds × 4 arms, on one pinned host and one model.
    • Every assignment was sealed, scored and verified with no stop, about 18.4 USD of CLI-computed subscription quota against a 192 USD admission threshold.
    • A separate adjudication found that the subjects mostly behaved correctly. All 8 refusals on the out-of-range task were valid, and no swapped-legend case was a real role error. Meanwhile the mechanical evaluator left about two thirds of unsupported_claim outcomes unresolved, and most of its invalid verdicts were wrong.
  • Pilot-readiness engineering:
    • a pilot approval kind that binds the schedule seed and broker limits;
    • the design budget;
    • a cost check for usage that reports no inference geography;
    • a census of lost launches before preflight;
    • per-process attribution of sandbox-denial reports.
  • Evaluator repairs. Two rounds, each with held-out cases committed before the change, and read-only re-judgments of the sealed smoke and pilot into separate directories. Neither round met its acceptance bar; the record states this.

What it is not

Development evidence only. It gives no treatment-effect estimate, the per-cell samples are tiny, and the scoring rules remain provisional pending review.

Checks

On macOS, the staged export's full suite collected 7206 tests: 7139 passed, 53 skipped (optional dependencies and dev-only state), 14 xfailed, 0 failed. A Linux-like governance run, with Seatbelt reported unavailable, has 0 failures.

The exporter's evidence, agent-surface and publication checks passed, and the task bank builds from the staged tree. main will be fast-forwarded only after this PR's Linux test suite job is green.

…luator repairs

Adds the record of the 96-assignment development pilot on the WP12 task bank
(docs/development/evaluation-study/pilot-record.md): every assignment
sealed, scored and verified with no stop, about 18.4 USD of CLI-computed
subscription quota against a 192 USD admission threshold, and a separate
adjudication by analysis agents that finds the subjects mostly behaved
correctly while the mechanical evaluator left about two thirds of outcomes
unresolved and produced most of its invalid verdicts in error.

Also adds the pilot approval kind and design budget, the absent-geography
cost check, the pre-preflight lost-launch census, the per-process LC-18
attribution, and two rounds of held-out-first evaluator repairs with
read-only re-judgments of the sealed smoke and pilot. Development evidence
only: no treatment-effect estimate. Source commit 02fc758.
…dget

The ubuntu-24.04 CI run failed three parameters of the task-bank megabyte
test at 21 s of process CPU against a 20 s budget; the same reads take 6 to
11 s on the development machine, so the budget measured runner speed, not
linearity. Both megabyte tests now require the whole text to cost under
eight times a quarter of it plus 2 s (linear is about four times, quadratic
about sixteen), with a 120 s ceiling. Source commit 11596f5.
@ammarphp
ammarphp merged commit 3613e6a into main Oct 4, 2026
@ammarphp
ammarphp deleted the evaluation-study-3 branch October 6, 2026 00:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant