Skip to content

[contract] the measurement window shrinks as arms get faster, and nothing bounds it #48

Description

@MarcusKainth

Which document

methodology/measurement.md — how you are measured

The passage

From What the instrument can resolve:

The drain's window comes from the sampler's own timestamps, and the sampler
runs at 1 Hz. So a drain is measured in whole seconds, and a throughput
difference smaller than one sampler tick is not a difference this instrument
can see
. On a 26-second drain one tick is 3.8%; on a 40-second one it is
2.5%.

The two readings

Reading A — the resolution is a fixed property of the rig, in the 2.5–3.8%
band the passage works through, and a reader can treat that band as the
instrument's floor.

Reading B — the resolution is a function of how fast the arm is, because
the window is corpus / throughput over a fixed corpus. It has no lower bound,
degrades as arms improve, and the 26–40 s band describes only the arms that
existed when the passage was written.

B is correct, and the document does not say so. Nothing bounds the window from
below, and no record carries the resolution it was measured at — so a reader
cannot tell a 3.8% reading from a 7.1% one.

What turns on the answer

It changes a number, and it already has.

Run 20260823-224744-32671669475-194524202c60e64b9931fb50 took the Spate arm
from 0.1.0 to 0.2.0. The corpus is a fixed 110,250,000 rows, so rows_per_s is
exactly rows / window and the drain window halved as the arm got faster — from
a uniform 31 s at 0.1.0 to 14–18 s for spate:native. One tick at 14 s is 7.1%,
outside the band the passage works through.

Sorting this run's six records by window shows the effect on cpu_us_per_row,
which is otherwise the low-variance metric:

window variant cpu µs/row
14.0 s native 0.476
17.0 s native 0.621
18.0 s native 0.621
23.0 s rowbinary 0.604
26.0 s rowbinary 0.598
28.0 s rowbinary 0.602

Five of six agree to ~1% within variant. The single outlier is the shortest
window. Excluding it, the two variants converge on the same figure — which is
what a saving upstream of encoding should produce, since it must land equally on
both wire formats:

all 3 reps excluding the 14 s rep
native cpu_us_per_row vs 0.1.0 −64.5% −61.5%
rowbinary cpu_us_per_row vs 0.1.0 −62.0% −62.0%
native rows_per_s vs 0.1.0 1.92x 1.77x

So the published mean for this run is pulled by its least precisely measured
repetition, and the effect is systematic rather than random: the faster the arm,
the shorter the window, the coarser the reading, and short windows read high
because the fixed row count divides by a smaller number.

Three things worth deciding, roughly in order of cost:

  1. A minimum window. The batch count is a variant knob (variant.batches),
    and dataset_version hashes the corpus rules, the .avsc, the DDL and the
    workload spec — not the batch count. So lengthening the drain moves no digest
    and invalidates no ceiling. This looks like the cheap fix.
  2. Record the resolution on the record, so a reader can see that a given
    reading was taken at 7.1% rather than 3.8%, and so a short_window flag can
    exist alongside reused_infra and cpu_cap_throttled.
  3. Raise the sampler rate, which narrows the tick directly but does not stop
    the window shrinking.

Whatever is decided, the passage should say that the window is a function of arm
speed rather than a property of the rig, because the current wording invites a
reader to assume a floor that is not there.

Related: #47. That issue asks whether background merges depress throughput, and
this run cannot separate the two hypotheses — reps run in sequence on reused
infrastructure, so the first rep is both the freshest and the shortest, and
both explanations predict the same ordering. Settling either one needs the other
controlled.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions