This was written agentically; verify its assertions:
Slice 8 of epic #152. Design: docs/superpowers/specs/2026-08-13-token-benchmark-design.md, section "Staged plan" stage 2.
Why
The highest-value comparison is new@low vs old@high: effort tiers move tokens by multiples while the refactor moves them by percents, so a dominance result survives a noise floor that a percentage delta cannot (repeat-run token spend varies up to ~30×; arXiv:2604.22750).
What
- 5 paired trials (10 runs),
new@low vs old@high, A,B,B,A alternation so provider drift cancels; each round includes the frozen 06d18cf drift-control trial per the pre-registration.
- Adjudicate strictly by the pre-registered dominance rule (median acceptance ≥ AND median blended USD <, non-overlapping IQRs); publish the reference-hit-rate analysis (the ≥90%/0% siting rule) regardless of the dominance outcome.
- Go/no-go record for stage 3.
Acceptance
🤖 Co-authored by Claude Fable 5.
This was written agentically; verify its assertions:
Slice 8 of epic #152. Design:
docs/superpowers/specs/2026-08-13-token-benchmark-design.md, section "Staged plan" stage 2.Why
The highest-value comparison is
new@lowvsold@high: effort tiers move tokens by multiples while the refactor moves them by percents, so a dominance result survives a noise floor that a percentage delta cannot (repeat-run token spend varies up to ~30×; arXiv:2604.22750).What
new@lowvsold@high,A,B,B,Aalternation so provider drift cancels; each round includes the frozen06d18cfdrift-control trial per the pre-registration.Acceptance
🤖 Co-authored by Claude Fable 5.