Skip to content

test(bench): stage 2 dominance probe — new@low vs old@high #330

Description

@thewrz

This was written agentically; verify its assertions:

Slice 8 of epic #152. Design: docs/superpowers/specs/2026-08-13-token-benchmark-design.md, section "Staged plan" stage 2.

Why

The highest-value comparison is new@low vs old@high: effort tiers move tokens by multiples while the refactor moves them by percents, so a dominance result survives a noise floor that a percentage delta cannot (repeat-run token spend varies up to ~30×; arXiv:2604.22750).

What

  • 5 paired trials (10 runs), new@low vs old@high, A,B,B,A alternation so provider drift cancels; each round includes the frozen 06d18cf drift-control trial per the pre-registration.
  • Adjudicate strictly by the pre-registered dominance rule (median acceptance ≥ AND median blended USD <, non-overlapping IQRs); publish the reference-hit-rate analysis (the ≥90%/0% siting rule) regardless of the dominance outcome.
  • Go/no-go record for stage 3.

Acceptance

  • 10 trial records + drift controls in the ledger; void-trial rule applied as pre-registered
  • Dominance verdict stated against the pre-registered rule, not a post-hoc one
  • Reference-siting findings reported (misfiled / dead-weight references named)
  • Go/no-go decision note committed

🤖 Co-authored by Claude Fable 5.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/testsThe suite and its gatesenhancementNew feature or requestp2Wanted before public release

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions