Skip to content

test(bench): stage 1 pilot — two trials, go/no-go #329

Description

@thewrz

This was written agentically; verify its assertions:

Slice 7 of epic #152. Design: docs/superpowers/specs/2026-08-13-token-benchmark-design.md, section "Staged plan" stage 1.

Why

Per-trial cost (USD, wall-clock, GraphQL) is unknown until measured; the design stages all spend behind go/no-go gates. The pilot is that measurement.

What

  • Run 2 trials: old@high (06d18cf), new@high (53e7e8c), alternating arm order per the design.
  • Outputs: real per-trial cost, wall clock, GraphQL consumption; sets the trial timeout for later stages; validates harness, oracle injection, and rollout parsing end to end on real data.
  • Go/no-go record: a short committed note (with ledger references) deciding whether stage 2 proceeds and at what budget.

Acceptance

  • 2 complete trial records in the ledger, schema-valid, is_drift_control handling exercised
  • Cost/wall-clock/GraphQL figures recorded; timeout derived and committed
  • Go/no-go decision note committed
  • No fixture edits (fixture_version unchanged)

🤖 Co-authored by Claude Fable 5.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/testsThe suite and its gatesenhancementNew feature or requestp2Wanted before public release

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions