Written 2026-08-23, before the training set it governs finished collecting and before any number from it was seen. It exists because the 32 seeds every arm has been scored on are no longer a test set: dozens of decisions have been made after looking at them, and an effect the size of the learned probe's +449 cannot survive that kind of reuse. Large effects — the +4590 of an averaged-tail oracle — are not seriously threatened. Small ones are.
Everything below is fixed now. If any of it changes, this file changes with a dated note saying what and why, and the run does not count as confirmatory.
A student trained to imitate which candidate an averaged-tail oracle picks plays further than the reactive policy, and further than the best fixed template.
Every previous student regressed the plan's value and took an argmax of its own estimate. This one is given the choice directly.
- Development: 0-31. Everything measured so far. Model selection, hyper-parameters, sanity checks, anything that involves looking.
- Training data: 9100 and up (
draw_matrix_big.npz), 5000-5041 (the DAgger collection), 9000-9003 (the first matrices). Disjoint from both other blocks. - Confirmatory: 4000-4031, thirty-two seeds, never used by any arm in this project. They are looked at once, after everything below is frozen.
- Architecture:
plan_probe.Probe, thestripinput variant — the hero crop plus the forward band, hero velocity as the vector part. Unchanged from every previous probe, so the comparison is about the target and not the net. - Output: six logits, one per candidate (
bc,run,jump now,jump later,wait,back off). - Target:
p_i = P(i = argmax_j Q_j), estimated by 200 bootstrap resamples of the sixteen saved returns per candidate. Death is terminal: a draw that dies scores the worst surviving return in that point minus 200 px. - Loss: cross-entropy against
p_i, each point weighted by1 - H(p_i)/log 6, floored at 0.1. A point whose label is a coin is not taught as a fact. - Optimiser: Adam, lr 1e-3, batch 128, 40 epochs, seed 0.
- Draws: sixteen, with common random numbers — one stream per draw shared across candidates. Measured to shrink the standard deviation of a paired difference by 18-22% against independent sampling.
On the development data only — the held-out runs of the training matrix, not
the confirmatory seeds — exactly one choice is made: soft target against hard
one-hot (--hard), and weighted against unweighted (--no-weight). Four
configurations, judged by regret under the sixteen-draw mean on held-out runs.
The winner is the single model that goes to the confirmatory seeds. Ties go to
the soft, weighted version, as the pre-registered default.
Nothing else is tuned. If the winner's regret is worse than the best constant template's on the same held-out runs, the confirmatory run is not made at all and the result is reported as a failure at the development stage.
Three arms on seeds 4000-4031, 3000 frames, commit 16, horizon 48, paired:
bc— the reactive policy.always jump now— the best fixed template, with the scoring loop actually skipped.- the student.
Metric: progress = level * 4000 + x, the best reached in the run.
The claim is confirmed only if both hold:
- the student beats
bc— paired permutation test on the per-seed differences, 10000 permutations, two-sided, p < 0.05; - the student beats
always jump nowby the same test, p < 0.05.
Reported alongside, not as criteria: the median and mean progress, a paired bootstrap 95% interval on each difference, deaths per arm, and McNemar's exact test on 1-1 completions.
If only the first holds, the honest statement is "better than the policy, not shown to be better than a fixed habit". If neither holds, the direction is reported as failed, and the failure is not re-analysed into a success by subsetting seeds, changing the metric, or extending the frame budget.
-
Looking at 4000-4031 before the model is frozen.
-
Running the confirmatory arms more than once and reporting the better run.
-
Choosing the comparison template after seeing the results — it is
jump now, chosen now, because it carries 52% of the full candidate set's value on the development matrix and had the lowest constant regret there.Noted the same day, before the confirmatory run: on the development game metric the strongest constant is
run(-268 px against the policy) rather thanjump now(-292 px). The arm staysjump now— it was chosen on the matrix, which is the development artefact the student is trained from, and switching now would be selection on the game numbers. The 24 px between them is far inside the seed spread, and both lose heavily to the policy, so the binding half of the criterion is beatingbceither way. -
Changing the metric, the frame budget or the commitment after the fact.
Two of the campaigns run in September were unmeasurable before they were
launched, and that was visible in advance. With paired runs the smallest
effect a comparison can resolve is the two-sided 95% t bound, t·sd/√n, and
the spreads this project actually observes are known:
| sd of the difference | 138 | 171 | 200 | 263 | 309 | 394 |
|---|---|---|---|---|---|---|
| resolvable with 6 runs | 145 | 179 | 210 | 276 | 324 | 413 |
| resolvable with 13 runs | 84 | 104 | 121 | 159 | 187 | 239 |
| resolvable with 20 runs | 65 | 80 | 94 | 123 | 145 | 184 |
The DAgger question — is +169 px real — sat under the six-run line from the start. "Unmeasurable at this size" is a different sentence from "did not replicate", and only the first one saves the machine time.
Before a campaign, therefore: state the effect worth detecting, take the sd from the nearest comparison already on disk, and read off the runs needed. If that number is unaffordable, the campaign does not answer the question and should be redesigned rather than run.
Which to buy, seeds or runs. The spread between runs mixes the training
lottery with the finite evaluation sample, and the two have opposite remedies.
nes_player.evaluation.stats.variance_split separates them from logs already
on disk, at no cost. Measured here: the evaluation sample is 84% of the spread
for the attention comparison but only 23% for DAgger, so extra evaluation seeds
would sharpen the first and do almost nothing for the second.
At the cost of the current recipe — about 160 s to train one epoch, about 95 s to evaluate 32 seeds — extra training runs win in both cases, and by a wide margin where training noise dominates. That ratio is a property of the recipe, not a law: under the three-epoch schedule used until 10 September, training cost twelve minutes and buying evaluation seeds was the better trade. Recompute it when the recipe changes.
More evaluation seeds are still worth buying for a reason the variance arithmetic does not see: 32 seeds sample the game's phases thinly, and this project has already been fooled once by seeds that were not independent.