Skip to content

[bug] the harness cannot distinguish a real change from a noisy sweep #49

Description

@MarcusKainth

Where

bench CLI / harness

What happened

The harness has no way to tell "the system under test changed" from "the rig was
noisy during this sweep", and it publishes either as a result.

Run 20260823-224744-32671669475-194524202c60e64b9931fb50 is the case in point.
spate:native reported 6,122,279–7,871,064 rows/s across three repetitions — a
28% spread — and every one of those readings was published. Nothing in the run
establishes what spread the rig produces when nothing changes, so there is no
basis on which to call 28% acceptable or alarming.

The nearest thing to a noise floor we have is the 0.1.0 baseline, whose six
records agree to five significant figures. That looks like a quiet rig but is
not: the arm was pinned against its 6-core cap (5.71–5.82 cores used, 79–86 s
throttled), so its throughput was its own CPU ceiling and shared-infrastructure
noise could not reach the number. At 0.2.0 the arm draws 2.37–4.03 cores and is
throttled under 3 s, so it no longer masks the rig. We have never measured this
rig's noise floor in the regime the suite now operates in.

docs/reproduce.md and .github/ISSUE_TEMPLATE/bug.yml both quote 14.5%
run-to-run spread on throughput. That figure predates the current environment
and is not reproducible from anything committed here.

What you expected instead

An A/A control that runs inside an ordinary sweep and calibrates the rig against
itself:

  • Measure one entrant twice under two labels, on the same corpus, in the same
    interleave as any other pair.
  • If the two labels differ by more than a declared threshold, mark the sweep
    inconclusive and downgrade its timing verdicts, rather than publishing them.
  • Record the observed A/A delta on every record from that sweep, so a reader can
    see the noise floor the number was measured against.

Two properties make this worth the box time. It cannot be forgotten, because it
runs as part of the sweep rather than as a separate exercise someone remembers
to do. And an A/A run reporting a difference is a harness or environment bug
rather than a finding, which makes it a self-checking gate instead of another
number to interpret.

A deliberate calibration mode (bench run <entrant> --aa) is the companion
piece: run when the box shape, the infrastructure caps, the ClickHouse or broker
version, or the corpus size changes — each of which invalidates whatever floor
was measured before.

Environment

$ bench validate
entrants: 6 descriptor(s) valid
environments: 1 profile(s) valid, 1 with a ceiling that may be gated against
results: 36 record(s) in 5 file(s) valid
harness v1, dataset d2-60d7e5bb2a82

Environment c8g-8xl-ec2-docker, infra_digest d867b69febf3.

Output

# spate:native, run 20260823-224744-..., three repetitions
rows_per_s      7,871,064 / 6,122,279 / 6,482,244      spread 28%
cpu_us_per_row      0.476 /     0.621 /     0.621      spread 30%

# the same arm at 0.1.0, six records, CPU-capped and therefore self-limiting
rows_per_s      3,554,846 - 3,555,075                  spread 0.006%

Related: #47 and #48 both propose explanations for that spread — background
merges and measurement-window resolution respectively. Neither can be tested
without knowing what the rig does when nothing changes, which is what this
issue asks for.

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions