Where
bench CLI / harness
What happened
The harness has no way to tell "the system under test changed" from "the rig was
noisy during this sweep", and it publishes either as a result.
Run 20260823-224744-32671669475-194524202c60e64b9931fb50 is the case in point.
spate:native reported 6,122,279–7,871,064 rows/s across three repetitions — a
28% spread — and every one of those readings was published. Nothing in the run
establishes what spread the rig produces when nothing changes, so there is no
basis on which to call 28% acceptable or alarming.
The nearest thing to a noise floor we have is the 0.1.0 baseline, whose six
records agree to five significant figures. That looks like a quiet rig but is
not: the arm was pinned against its 6-core cap (5.71–5.82 cores used, 79–86 s
throttled), so its throughput was its own CPU ceiling and shared-infrastructure
noise could not reach the number. At 0.2.0 the arm draws 2.37–4.03 cores and is
throttled under 3 s, so it no longer masks the rig. We have never measured this
rig's noise floor in the regime the suite now operates in.
docs/reproduce.md and .github/ISSUE_TEMPLATE/bug.yml both quote 14.5%
run-to-run spread on throughput. That figure predates the current environment
and is not reproducible from anything committed here.
What you expected instead
An A/A control that runs inside an ordinary sweep and calibrates the rig against
itself:
- Measure one entrant twice under two labels, on the same corpus, in the same
interleave as any other pair.
- If the two labels differ by more than a declared threshold, mark the sweep
inconclusive and downgrade its timing verdicts, rather than publishing them.
- Record the observed A/A delta on every record from that sweep, so a reader can
see the noise floor the number was measured against.
Two properties make this worth the box time. It cannot be forgotten, because it
runs as part of the sweep rather than as a separate exercise someone remembers
to do. And an A/A run reporting a difference is a harness or environment bug
rather than a finding, which makes it a self-checking gate instead of another
number to interpret.
A deliberate calibration mode (bench run <entrant> --aa) is the companion
piece: run when the box shape, the infrastructure caps, the ClickHouse or broker
version, or the corpus size changes — each of which invalidates whatever floor
was measured before.
Environment
$ bench validate
entrants: 6 descriptor(s) valid
environments: 1 profile(s) valid, 1 with a ceiling that may be gated against
results: 36 record(s) in 5 file(s) valid
harness v1, dataset d2-60d7e5bb2a82
Environment c8g-8xl-ec2-docker, infra_digest d867b69febf3.
Output
# spate:native, run 20260823-224744-..., three repetitions
rows_per_s 7,871,064 / 6,122,279 / 6,482,244 spread 28%
cpu_us_per_row 0.476 / 0.621 / 0.621 spread 30%
# the same arm at 0.1.0, six records, CPU-capped and therefore self-limiting
rows_per_s 3,554,846 - 3,555,075 spread 0.006%
Related: #47 and #48 both propose explanations for that spread — background
merges and measurement-window resolution respectively. Neither can be tested
without knowing what the rig does when nothing changes, which is what this
issue asks for.
Where
bench CLI / harness
What happened
The harness has no way to tell "the system under test changed" from "the rig was
noisy during this sweep", and it publishes either as a result.
Run
20260823-224744-32671669475-194524202c60e64b9931fb50is the case in point.spate:nativereported 6,122,279–7,871,064 rows/s across three repetitions — a28% spread — and every one of those readings was published. Nothing in the run
establishes what spread the rig produces when nothing changes, so there is no
basis on which to call 28% acceptable or alarming.
The nearest thing to a noise floor we have is the 0.1.0 baseline, whose six
records agree to five significant figures. That looks like a quiet rig but is
not: the arm was pinned against its 6-core cap (5.71–5.82 cores used, 79–86 s
throttled), so its throughput was its own CPU ceiling and shared-infrastructure
noise could not reach the number. At 0.2.0 the arm draws 2.37–4.03 cores and is
throttled under 3 s, so it no longer masks the rig. We have never measured this
rig's noise floor in the regime the suite now operates in.
docs/reproduce.mdand.github/ISSUE_TEMPLATE/bug.ymlboth quote 14.5%run-to-run spread on throughput. That figure predates the current environment
and is not reproducible from anything committed here.
What you expected instead
An A/A control that runs inside an ordinary sweep and calibrates the rig against
itself:
interleave as any other pair.
inconclusive and downgrade its timing verdicts, rather than publishing them.
see the noise floor the number was measured against.
Two properties make this worth the box time. It cannot be forgotten, because it
runs as part of the sweep rather than as a separate exercise someone remembers
to do. And an A/A run reporting a difference is a harness or environment bug
rather than a finding, which makes it a self-checking gate instead of another
number to interpret.
A deliberate calibration mode (
bench run <entrant> --aa) is the companionpiece: run when the box shape, the infrastructure caps, the ClickHouse or broker
version, or the corpus size changes — each of which invalidates whatever floor
was measured before.
Environment
$ bench validate entrants: 6 descriptor(s) valid environments: 1 profile(s) valid, 1 with a ceiling that may be gated against results: 36 record(s) in 5 file(s) valid harness v1, dataset d2-60d7e5bb2a82Environment
c8g-8xl-ec2-docker,infra_digestd867b69febf3.Output
Related: #47 and #48 both propose explanations for that spread — background
merges and measurement-window resolution respectively. Neither can be tested
without knowing what the rig does when nothing changes, which is what this
issue asks for.