Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 16 additions & 5 deletions .dev/STATE.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,11 +7,22 @@ Roadmap detail lives in the work-plan tracker **#257**, not here (PLAN.md predat

## In flight

- **v0.29.2 release cut 2026-09-21** (payload #317 `62504c1`, on top of #315 `cd41558`) — patch:
the scope sentence on the `ml` exercise rule that v0.29.1's failed gate called for, carrying
with it the editor's round-3 answers that v0.29.1 never delivered. §4a gate status is
recorded on the release PR; scenario 17 on the `.ml` lane is the one to read first. After the
floating tags move, round 4 (`numpy`, lecture-python-programming.ml#23) is regenerated at `@v0`.
- **v0.29.2 released 2026-09-21** (release PR #318 `5f74d74`; payload #317 `62504c1`, on top of
#315 `cd41558`; §4a gate **completed**: 84/84 sync runs, 28/28 delivery + 28/28
`engineVersion: 0.29.2` verdicts per lane, scenario 17 on `.ml` delivered; `v0.29` = `v0` =
`5f74d74`; GitHub release published; `@v0` smoke `engineRef: v0` on all three lanes — tally
on #318) — patch: the scope sentence on the `ml` exercise rule that v0.29.1's failed gate
called for, carrying the editor's round-3 answers that v0.29.1 never delivered. **Round 4
(`numpy`, lecture-python-programming.ml#23) regenerated at `@v0` the same day**: three draws,
draw 1 sent (bare endings 1 / 37 / 2), one `ml_repair.py` comma applied and disclosed — the
first repaired seed; arm `experiments/ml-benchmark/arms/2026-09-21-round4-numpy-v0.29.2/`.
Harness note: the reset left one PR from the previous gate open on `.ml` (closed by hand).
W1 (#259) still targets v0.30.0. **2026-09-23 close-out**: the reset survivor's cause was
`gh pr list` with no `--limit` — fixed with a post-reset check on **#322** (closes #321); the
gate's one-draw blindness (a ~40% defect passed three gates) is **#320**, decided as a local
N-draw rate check — `tool-test-action-on-github/rate-check.sh`, release step 4b, on **#323**
(v0.29.2 reads 0/12, v0.29.0 4/12). **Next engine-side: the rule-2 arm** (rewrite around the
editor's everyday-speech test, judged held-out before it ships).
- **v0.29.1 is tagged but NOT released (2026-09-21)** — its §4a gate came back 83/84: scenario 17
on the `.ml` lane failed twice (the model wrapped a plain `## Exercises` list in
`{exercise-start}`; structural parity refused the file). `v0` = `v0.29` = `a6fda54` (v0.29.0)
Expand Down
4 changes: 4 additions & 0 deletions .dev/log/2026-09-21-v0291-gate-scenario17.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,3 +24,7 @@ The first six draws per arm read 1/6 against 5/6 and looked like a clean regress
**Merged, release cut.** #317 merged as `62504c1` (Copilot: two release-status wording comments — STATE entries reworded, CHANGELOG marks 0.29.1 `[YANKED]` with a note, `cc7c6ab`); release PR for v0.29.2 opened on `release-v0.29.2`.

**Next (as written before the merge)**: PR → merge → **v0.29.2** (v0.29.1 stays a tag that was never released; the CHANGELOG says so) → §4a gate on the new tag → floating tags → smoke → release → regenerate round 4 at `@v0`.

**Addendum — v0.29.2 released.** #318 merged `5f74d74`; tagged `v0.29.2`; §4a gate on the tag: census 9/9 at `@v0.29.2`, 28 source PRs, **84/84 sync success**, 28/28 delivery and **28/28 verdicts at `engineVersion: 0.29.2` per lane**; scenario 17 on `.ml` delivered (test-translation-sync.ml#242). The `.ml` lane briefly read 29 PRs: the reset had left test-translation-sync.ml#192 open from the v0.29.1 gate — closed by hand, and worth a look in the script's close-all step. `v0.29` and `v0` → `5f74d74`; `@v0` smoke (scenario 01) 3/3, `engineRef: v0` on all lanes; GitHub release published (title = tag; notes cover #315 and #317 and say not to pin `v0.29.1`). Tally on #318.

**Round 4 regenerated at `@v0`.** Three draws of `numpy`, draw 1 chosen by lint (bare endings 1 / 37 / 2), `ml_repair.py` applied — one comma, `Columns-ഉം, rows-ഉം` — and disclosed; force-updated lecture-python-programming.ml#23 (`889d001`), retitled, description rewritten, editor told it is ready. He had not started on the v0.29.0 draft. Arm: `experiments/ml-benchmark/arms/2026-09-21-round4-numpy-v0.29.2/`.
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# Arm: round 4 (`numpy`) regenerated — three draws at v0.29.2, chosen by lint, 2026-09-21

The round-4 calibration seed for lecture-python-programming.ml was first generated at v0.29.0
(`../2026-09-18-round4-numpy-v0.29.0/`, draw 3 sent). The editor then answered the round-3
questions (lecture-python-programming.ml#22) before starting on it, so the seed was regenerated
at the release that carries his answers — the round-3 precedent at v0.28.0. Three draws at
v0.29.2 (`5f74d74` = `@v0`) from `lecture-python-programming@b0b0b56` (`numpy.md` unchanged
since `1706cea`), `init -f numpy.md --localize none -m claude-sonnet-5`.

| Draw | bare endings before a cell/list | paragraph without punctuation | lowercase-initial | banned | hortative watch | `-ഉം` pair, no comma | *For example* → ഉദാഹരണത്തിന് | `provide ചെയ്യ…` | exercise blocks | headings / code cells |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 (**sent**, after repair) | **1** | 0 | 0 | 1 | 1 | 1 → 0 | 0 | 0 | 8/8 | identical |
| 2 | **37** | 1 | 0 | 2 | 0 | 1 | 6 | 0 | 8/8 | identical |
| 3 | **2** | 8 | 2 | 2 | 2 | 1 | 0 | 0 | 8/8 | identical |
| v0.29.0 draw 3 (first sent) | 3 | 0 | 0 | 3 | 1 | 2 | 0 | 6 | 8/8 | identical |

`ml_repair.py` was applied to draw 1 and **disclosed on the PR** — the first seed to go to the
editor repaired. It made one change: `Columns-ഉം rows-ഉം` → `Columns-ഉം, rows-ഉം` (line 280).
Every draw here leaves exactly one hyphenated pair without its comma — that pair in draws 1 and
3, `Step 1-ഉം 2-ഉം` in draw 2 — and those are the same two lines the 2026-09-21 answers arm
missed: the prompt rule does not reach them, the script does. Left as
generated and disclosed: line 878 (ends on a closing bracket before a cell), line 898 (*In
fact* → വാസ്തവത്തിൽ, a banned rendering), line 1092 (a genuine future statement the hortative
watch flags). Whether he touches the repaired line is the first field evidence for #260.

Terminal punctuation remains the draw-dependent class (1 / 37 / 2 here; 37 / 34 / 3 at
v0.29.0): best-of-N by lint is still what keeps it away from the editor.

One more observation, raised in review of the PR that archived this arm: draw 2, line 177,
carries a garbage token — `2x2 array undegerbestellen ഉണ്ടാക്കാൻ` — a Latin-script string that
is neither English, code nor a glossary term, emitted by the model and absent from the other
two draws and from the draft sent. It is left exactly as generated: the draws are the record.
It is also a lint gap — nothing in `ml_metrics.py` flags a Latin token that is not English,
code or a pinned term (#301) — and draw 2 was already the one not chosen, on bare endings.

Loading
Loading