Which document
methodology/measurement.md — how you are measured
The passage
From What the instrument can resolve:
The drain's window comes from the sampler's own timestamps, and the sampler
runs at 1 Hz. So a drain is measured in whole seconds, and a throughput
difference smaller than one sampler tick is not a difference this instrument
can see. On a 26-second drain one tick is 3.8%; on a 40-second one it is
2.5%.
The two readings
Reading A — the resolution is a fixed property of the rig, in the 2.5–3.8%
band the passage works through, and a reader can treat that band as the
instrument's floor.
Reading B — the resolution is a function of how fast the arm is, because
the window is corpus / throughput over a fixed corpus. It has no lower bound,
degrades as arms improve, and the 26–40 s band describes only the arms that
existed when the passage was written.
B is correct, and the document does not say so. Nothing bounds the window from
below, and no record carries the resolution it was measured at — so a reader
cannot tell a 3.8% reading from a 7.1% one.
What turns on the answer
It changes a number, and it already has.
Run 20260823-224744-32671669475-194524202c60e64b9931fb50 took the Spate arm
from 0.1.0 to 0.2.0. The corpus is a fixed 110,250,000 rows, so rows_per_s is
exactly rows / window and the drain window halved as the arm got faster — from
a uniform 31 s at 0.1.0 to 14–18 s for spate:native. One tick at 14 s is 7.1%,
outside the band the passage works through.
Sorting this run's six records by window shows the effect on cpu_us_per_row,
which is otherwise the low-variance metric:
| window |
variant |
cpu µs/row |
| 14.0 s |
native |
0.476 |
| 17.0 s |
native |
0.621 |
| 18.0 s |
native |
0.621 |
| 23.0 s |
rowbinary |
0.604 |
| 26.0 s |
rowbinary |
0.598 |
| 28.0 s |
rowbinary |
0.602 |
Five of six agree to ~1% within variant. The single outlier is the shortest
window. Excluding it, the two variants converge on the same figure — which is
what a saving upstream of encoding should produce, since it must land equally on
both wire formats:
|
all 3 reps |
excluding the 14 s rep |
native cpu_us_per_row vs 0.1.0 |
−64.5% |
−61.5% |
rowbinary cpu_us_per_row vs 0.1.0 |
−62.0% |
−62.0% |
native rows_per_s vs 0.1.0 |
1.92x |
1.77x |
So the published mean for this run is pulled by its least precisely measured
repetition, and the effect is systematic rather than random: the faster the arm,
the shorter the window, the coarser the reading, and short windows read high
because the fixed row count divides by a smaller number.
Three things worth deciding, roughly in order of cost:
- A minimum window. The batch count is a variant knob (
variant.batches),
and dataset_version hashes the corpus rules, the .avsc, the DDL and the
workload spec — not the batch count. So lengthening the drain moves no digest
and invalidates no ceiling. This looks like the cheap fix.
- Record the resolution on the record, so a reader can see that a given
reading was taken at 7.1% rather than 3.8%, and so a short_window flag can
exist alongside reused_infra and cpu_cap_throttled.
- Raise the sampler rate, which narrows the tick directly but does not stop
the window shrinking.
Whatever is decided, the passage should say that the window is a function of arm
speed rather than a property of the rig, because the current wording invites a
reader to assume a floor that is not there.
Related: #47. That issue asks whether background merges depress throughput, and
this run cannot separate the two hypotheses — reps run in sequence on reused
infrastructure, so the first rep is both the freshest and the shortest, and
both explanations predict the same ordering. Settling either one needs the other
controlled.
Which document
methodology/measurement.md — how you are measured
The passage
From What the instrument can resolve:
The two readings
Reading A — the resolution is a fixed property of the rig, in the 2.5–3.8%
band the passage works through, and a reader can treat that band as the
instrument's floor.
Reading B — the resolution is a function of how fast the arm is, because
the window is
corpus / throughputover a fixed corpus. It has no lower bound,degrades as arms improve, and the 26–40 s band describes only the arms that
existed when the passage was written.
B is correct, and the document does not say so. Nothing bounds the window from
below, and no record carries the resolution it was measured at — so a reader
cannot tell a 3.8% reading from a 7.1% one.
What turns on the answer
It changes a number, and it already has.
Run
20260823-224744-32671669475-194524202c60e64b9931fb50took the Spate armfrom 0.1.0 to 0.2.0. The corpus is a fixed 110,250,000 rows, so
rows_per_sisexactly
rows / windowand the drain window halved as the arm got faster — froma uniform 31 s at 0.1.0 to 14–18 s for
spate:native. One tick at 14 s is 7.1%,outside the band the passage works through.
Sorting this run's six records by window shows the effect on
cpu_us_per_row,which is otherwise the low-variance metric:
Five of six agree to ~1% within variant. The single outlier is the shortest
window. Excluding it, the two variants converge on the same figure — which is
what a saving upstream of encoding should produce, since it must land equally on
both wire formats:
cpu_us_per_rowvs 0.1.0cpu_us_per_rowvs 0.1.0rows_per_svs 0.1.0So the published mean for this run is pulled by its least precisely measured
repetition, and the effect is systematic rather than random: the faster the arm,
the shorter the window, the coarser the reading, and short windows read high
because the fixed row count divides by a smaller number.
Three things worth deciding, roughly in order of cost:
variant.batches),and
dataset_versionhashes the corpus rules, the.avsc, the DDL and theworkload spec — not the batch count. So lengthening the drain moves no digest
and invalidates no ceiling. This looks like the cheap fix.
reading was taken at 7.1% rather than 3.8%, and so a
short_windowflag canexist alongside
reused_infraandcpu_cap_throttled.the window shrinking.
Whatever is decided, the passage should say that the window is a function of arm
speed rather than a property of the rig, because the current wording invites a
reader to assume a floor that is not there.
Related: #47. That issue asks whether background merges depress throughput, and
this run cannot separate the two hypotheses — reps run in sequence on reused
infrastructure, so the first rep is both the freshest and the shortest, and
both explanations predict the same ordering. Settling either one needs the other
controlled.