Where
bench CLI / harness
What happened
A sweep brings the infrastructure up once and then reuses the same ClickHouse
process for every arm, every repetition, and every subsequent run on the
same box.
RunOptions::fresh_infra defaults to false, so driver.rs calls
infra::bring_up(&env, !opts.fresh_infra) with reuse = true. The AWS payload
(.github/aws/run-bench.sh) never passes --fresh-infra, so every published
record carries Flag::ReusedInfra.
What is reset per arm per repetition is the table, not the server:
TRUNCATE TABLE sensor_events
wait_until_settled(...) # polls until
# count(system.parts WHERE active) + count(system.merges) == 0
# for that table; 250ms poll, 120s cap, 2s quiet sleep
That is sound and it does what its comment says: merges started by one
repetition cannot be charged to the next. It says nothing about the server
process, which keeps its page cache, mark and uncompressed caches, background
pool state and allocator state across every arm in the sweep and across runs on
the box.
The observation that prompted this. On run
20260826-230331-33021878215-…, spate:native moved to 1048576-row batches
(#58) and spate:rowbinary was re-measured unchanged alongside it. RowBinary's
knobs did not move, and its published median fell 8.1%:
|
before |
after |
delta |
| rows/s |
11,140,161 |
10,235,913 |
-8.1% |
| rows/s per core |
1,675,264 |
1,661,840 |
-0.8% |
| cores used |
6.65 |
6.15 |
-7.6% |
| ClickHouse merge duration |
7.74e9 us |
8.35e9 us |
+8.0% |
Per-core efficiency is unchanged, so the arm was paced rather than slowed, and
every repetition is classified infra_bound. This is not proof that
process carryover caused it — the arm has no A/A control of its own, and 8.1%
could be run-to-run spread on an arm riding the target's plateau. It is the
observation that showed the isolation is not there to rule it out with.
The reason it is worth closing regardless: Native now hands that server 65 MB
parts where it previously handed it 14 MB ones, and the arm measured next
inherits whatever that left behind. A published number should not depend on
which arm ran before it.
What you expected instead
Each run gets a ClickHouse server in a known state, so a number cannot depend
on what a previous arm or a previous run left in the process.
The capability is already most of the way there. infra::bring_up decides
reuse per container, and the comment above it explains that the split
exists precisely so recreating ClickHouse does not destroy the broker's
corpus:
let reused_broker = reuse && running(BROKER);
let reused_clickhouse = reuse && running(CLICKHOUSE);
Both still read one flag. --fresh-infra therefore recreates both, and the
corpus lives in the broker, which is why bench ceiling refuses the flag
outright and why a sweep cannot use it either.
Proposed: expose the split that already exists, as --fresh-clickhouse
(recreate ClickHouse, keep the broker and its corpus). Open questions for
whoever picks it up:
- Granularity. Per run is cheap and kills cross-run carryover. Per arm
bounds "which arm ran before me". Per repetition is the strongest and costs a
container restart plus a readiness wait on every drain — on a 9-drain sweep
that is real time, and bring_up's readiness wait already needs one paused
retry when a freshly recreated ClickHouse replays its data directory.
- Whether it should be the default, and therefore whether
ReusedInfra
stops appearing on published records.
- Whether the flag also belongs on
bench ceiling, which today refuses
--fresh-infra for a reason that only applies to the broker.
Note for scheduling: this touches harness/*, which affected-entrants.sh
maps to all=true, so landing it proposes a re-measurement of every arm on
every environment rather than of one entrant.
Environment
$ bench validate
entrants: 6 descriptor(s) valid
environments: 1 profile(s) valid, 1 with a ceiling that may be gated against
results: 25 record(s) in 5 file(s) valid
harness v2, dataset d2-60d7e5bb2a82
environment: c8gd-metal-24xl-ec2-docker (c8gd.metal-24xl, Ubuntu 24.04.4,
aarch64, Docker); ClickHouse 26.3.23.7 capped at 32 cpus / 32g
Output
# harness/src/driver.rs — reuse is the default at both call sites
let (ep, _infra, _flags) = infra::bring_up(&env, !opts.fresh_infra)?;
let (ep, infra, base_flags) = infra::bring_up(&env, !opts.fresh_infra)?;
# harness/src/bin/bench.rs — the default, and the only flag that changes it
fresh_infra: false,
"--fresh-infra" => o.fresh_infra = true,
# harness/src/bin/bench.rs — why the flag cannot simply be turned on
REFUSED: --fresh-infra recreates the broker, and the prefilled corpus lives
inside it. The consume ceiling is measured against the corpus's own messages,
so destroying it would leave nothing to measure
# every record of the run
"flags": ["reused_infra"]
# and the driver says so on the way past
reusing the running infrastructure. Caps are still read back and asserted
below, so a container started under a different envelope will fail the run
rather than quietly produce a number.
Where
bench CLI / harness
What happened
A sweep brings the infrastructure up once and then reuses the same ClickHouse
process for every arm, every repetition, and every subsequent run on the
same box.
RunOptions::fresh_infradefaults tofalse, sodriver.rscallsinfra::bring_up(&env, !opts.fresh_infra)withreuse = true. The AWS payload(
.github/aws/run-bench.sh) never passes--fresh-infra, so every publishedrecord carries
Flag::ReusedInfra.What is reset per arm per repetition is the table, not the server:
That is sound and it does what its comment says: merges started by one
repetition cannot be charged to the next. It says nothing about the server
process, which keeps its page cache, mark and uncompressed caches, background
pool state and allocator state across every arm in the sweep and across runs on
the box.
The observation that prompted this. On run
20260826-230331-33021878215-…,spate:nativemoved to 1048576-row batches(#58) and
spate:rowbinarywas re-measured unchanged alongside it. RowBinary'sknobs did not move, and its published median fell 8.1%:
Per-core efficiency is unchanged, so the arm was paced rather than slowed, and
every repetition is classified
infra_bound. This is not proof thatprocess carryover caused it — the arm has no A/A control of its own, and 8.1%
could be run-to-run spread on an arm riding the target's plateau. It is the
observation that showed the isolation is not there to rule it out with.
The reason it is worth closing regardless: Native now hands that server 65 MB
parts where it previously handed it 14 MB ones, and the arm measured next
inherits whatever that left behind. A published number should not depend on
which arm ran before it.
What you expected instead
Each run gets a ClickHouse server in a known state, so a number cannot depend
on what a previous arm or a previous run left in the process.
The capability is already most of the way there.
infra::bring_updecidesreuse per container, and the comment above it explains that the split
exists precisely so recreating ClickHouse does not destroy the broker's
corpus:
Both still read one flag.
--fresh-infratherefore recreates both, and thecorpus lives in the broker, which is why
bench ceilingrefuses the flagoutright and why a sweep cannot use it either.
Proposed: expose the split that already exists, as
--fresh-clickhouse(recreate ClickHouse, keep the broker and its corpus). Open questions for
whoever picks it up:
bounds "which arm ran before me". Per repetition is the strongest and costs a
container restart plus a readiness wait on every drain — on a 9-drain sweep
that is real time, and
bring_up's readiness wait already needs one pausedretry when a freshly recreated ClickHouse replays its data directory.
ReusedInfrastops appearing on published records.
bench ceiling, which today refuses--fresh-infrafor a reason that only applies to the broker.Note for scheduling: this touches
harness/*, whichaffected-entrants.shmaps to
all=true, so landing it proposes a re-measurement of every arm onevery environment rather than of one entrant.
Environment
Output