Which document
methodology/measurement.md — how you are measured
The passage
From the caveats following the metrics table:
Server-side cost excludes background merges, which live in
system.part_log and are arm-dependent: 25,000-row batches make far more
parts than 262,144-row ones, and that cost lands nowhere in this figure.
The two readings
Reading A — merges are absent from the Server-side cost row and from
nothing else. The caveat sits with that row, names that figure, and the metrics
table's other rows are unaffected.
Reading B — merges are unaccounted work that runs on the same cores and the
same disk during the measurement window, so they also depress the Throughput
row, which is wall-clock count() over the sampler's window. Under this reading
the exclusion is not only a gap in one figure's attribution; it is an
uncontrolled variable in the headline number.
The document states A. Only A is written down, and B is neither asserted nor
excluded.
What turns on the answer
It changes a number, and it has just started changing one materially.
Run 20260823-224744-32671669475-194524202c60e64b9931fb50 moved the Spate arm
from spate 0.1.0 to 0.2.0. Compared against the six 0.1.0 records in
results/c8g-8xl-ec2-docker/spate/2026-07.jsonl:
|
0.1.0 |
0.2.0 |
cores_used (6-core cap) |
5.71–5.82 |
2.37–4.03 |
throttled_us |
79.2–86.1 s |
0.6–2.8 s |
ch_cpu_wait_us |
16.5–342.7 ms |
2.1–94.6 s |
rows_per_s, native |
3,554,846–3,555,075 |
6,122,279–7,871,064 |
At 0.1.0 the arm was pinned against its own CPU cap, so its throughput was
reproducible to five significant figures — the suite was measuring the arm's
envelope, and infrastructure noise could not reach the number. At 0.2.0 the arm
draws two-thirds of the cap and is barely throttled, so whatever now limits it
is outside its envelope, and the native spread is 28%.
That is the general case rather than a Spate-specific one: as any arm stops
saturating its envelope, the published figure absorbs more of the shared
infrastructure's state, and background merges are the one component of that
state the contract explicitly declines to account for.
Whether merges are in fact the cause here is not settled by this run, and the
data cuts both ways:
- For — native rep 1 runs on the freshest infrastructure and is both the
fastest (7,871,064) and the lowest-wait (2.1 s); reps 2 and 3 fall to
6,122,279 and 6,482,244 with wait rising to 5.4 and 5.9 s. Every record
carries reused_infra, so a merge backlog accumulates across a sweep.
- Against — rowbinary rep 3 has the highest wait of the run (94.6 s) and also
the highest rowbinary throughput (4,791,395). ch_cpu_wait_us is an aggregate
over concurrent inserters, so it grows with rows pushed and is not an
independent congestion signal.
One structural detail is worth stating whichever way it resolves: interleaving
protects against ordering effects between arms, but reps still run in sequence
on reused infrastructure, so a monotonic drift systematically favours whichever
arm runs first. Native rep 1 is first in this sweep and is the outlier.
Three things that would settle it, cheapest first:
- Record per-window merge activity from
system.part_log into the result, so
drift is visible rather than inferred. bench ceiling --measure already
captures parts.merges and parts.rows_merged; the run path does not.
- Quiesce merges, or reset the table, between reps rather than reusing.
- Counterbalance rep order so drift does not always favour the same position.
If the answer is A, the document should say so explicitly, because the natural
reading of an exclusion is that the excluded thing is gone rather than merely
uncharged. If it is B, methodology/comparability.md has a stake too: two
numbers measured at different points in a sweep's merge backlog are not
straightforwardly comparable.
Which document
methodology/measurement.md — how you are measured
The passage
From the caveats following the metrics table:
The two readings
Reading A — merges are absent from the Server-side cost row and from
nothing else. The caveat sits with that row, names that figure, and the metrics
table's other rows are unaffected.
Reading B — merges are unaccounted work that runs on the same cores and the
same disk during the measurement window, so they also depress the Throughput
row, which is wall-clock
count()over the sampler's window. Under this readingthe exclusion is not only a gap in one figure's attribution; it is an
uncontrolled variable in the headline number.
The document states A. Only A is written down, and B is neither asserted nor
excluded.
What turns on the answer
It changes a number, and it has just started changing one materially.
Run
20260823-224744-32671669475-194524202c60e64b9931fb50moved the Spate armfrom
spate 0.1.0to0.2.0. Compared against the six 0.1.0 records inresults/c8g-8xl-ec2-docker/spate/2026-07.jsonl:cores_used(6-core cap)throttled_usch_cpu_wait_usrows_per_s, nativeAt 0.1.0 the arm was pinned against its own CPU cap, so its throughput was
reproducible to five significant figures — the suite was measuring the arm's
envelope, and infrastructure noise could not reach the number. At 0.2.0 the arm
draws two-thirds of the cap and is barely throttled, so whatever now limits it
is outside its envelope, and the native spread is 28%.
That is the general case rather than a Spate-specific one: as any arm stops
saturating its envelope, the published figure absorbs more of the shared
infrastructure's state, and background merges are the one component of that
state the contract explicitly declines to account for.
Whether merges are in fact the cause here is not settled by this run, and the
data cuts both ways:
fastest (7,871,064) and the lowest-wait (2.1 s); reps 2 and 3 fall to
6,122,279 and 6,482,244 with wait rising to 5.4 and 5.9 s. Every record
carries
reused_infra, so a merge backlog accumulates across a sweep.the highest rowbinary throughput (4,791,395).
ch_cpu_wait_usis an aggregateover concurrent inserters, so it grows with rows pushed and is not an
independent congestion signal.
One structural detail is worth stating whichever way it resolves: interleaving
protects against ordering effects between arms, but reps still run in sequence
on reused infrastructure, so a monotonic drift systematically favours whichever
arm runs first. Native rep 1 is first in this sweep and is the outlier.
Three things that would settle it, cheapest first:
system.part_loginto the result, sodrift is visible rather than inferred.
bench ceiling --measurealreadycaptures
parts.mergesandparts.rows_merged; the run path does not.If the answer is A, the document should say so explicitly, because the natural
reading of an exclusion is that the excluded thing is gone rather than merely
uncharged. If it is B,
methodology/comparability.mdhas a stake too: twonumbers measured at different points in a sweep's merge backlog are not
straightforwardly comparable.