Skip to content

[contract] the background-merge exclusion is scoped to server-side cost, but merges reach throughput #47

Description

@MarcusKainth

Which document

methodology/measurement.md — how you are measured

The passage

From the caveats following the metrics table:

Server-side cost excludes background merges, which live in
system.part_log and are arm-dependent: 25,000-row batches make far more
parts than 262,144-row ones, and that cost lands nowhere in this figure.

The two readings

Reading A — merges are absent from the Server-side cost row and from
nothing else. The caveat sits with that row, names that figure, and the metrics
table's other rows are unaffected.

Reading B — merges are unaccounted work that runs on the same cores and the
same disk during the measurement window, so they also depress the Throughput
row, which is wall-clock count() over the sampler's window. Under this reading
the exclusion is not only a gap in one figure's attribution; it is an
uncontrolled variable in the headline number.

The document states A. Only A is written down, and B is neither asserted nor
excluded.

What turns on the answer

It changes a number, and it has just started changing one materially.

Run 20260823-224744-32671669475-194524202c60e64b9931fb50 moved the Spate arm
from spate 0.1.0 to 0.2.0. Compared against the six 0.1.0 records in
results/c8g-8xl-ec2-docker/spate/2026-07.jsonl:

0.1.0 0.2.0
cores_used (6-core cap) 5.71–5.82 2.37–4.03
throttled_us 79.2–86.1 s 0.6–2.8 s
ch_cpu_wait_us 16.5–342.7 ms 2.1–94.6 s
rows_per_s, native 3,554,846–3,555,075 6,122,279–7,871,064

At 0.1.0 the arm was pinned against its own CPU cap, so its throughput was
reproducible to five significant figures — the suite was measuring the arm's
envelope, and infrastructure noise could not reach the number. At 0.2.0 the arm
draws two-thirds of the cap and is barely throttled, so whatever now limits it
is outside its envelope, and the native spread is 28%.

That is the general case rather than a Spate-specific one: as any arm stops
saturating its envelope, the published figure absorbs more of the shared
infrastructure's state
, and background merges are the one component of that
state the contract explicitly declines to account for.

Whether merges are in fact the cause here is not settled by this run, and the
data cuts both ways:

  • For — native rep 1 runs on the freshest infrastructure and is both the
    fastest (7,871,064) and the lowest-wait (2.1 s); reps 2 and 3 fall to
    6,122,279 and 6,482,244 with wait rising to 5.4 and 5.9 s. Every record
    carries reused_infra, so a merge backlog accumulates across a sweep.
  • Against — rowbinary rep 3 has the highest wait of the run (94.6 s) and also
    the highest rowbinary throughput (4,791,395). ch_cpu_wait_us is an aggregate
    over concurrent inserters, so it grows with rows pushed and is not an
    independent congestion signal.

One structural detail is worth stating whichever way it resolves: interleaving
protects against ordering effects between arms, but reps still run in sequence
on reused infrastructure, so a monotonic drift systematically favours whichever
arm runs first. Native rep 1 is first in this sweep and is the outlier.

Three things that would settle it, cheapest first:

  1. Record per-window merge activity from system.part_log into the result, so
    drift is visible rather than inferred. bench ceiling --measure already
    captures parts.merges and parts.rows_merged; the run path does not.
  2. Quiesce merges, or reset the table, between reps rather than reusing.
  3. Counterbalance rep order so drift does not always favour the same position.

If the answer is A, the document should say so explicitly, because the natural
reading of an exclusion is that the excluded thing is gone rather than merely
uncharged. If it is B, methodology/comparability.md has a stake too: two
numbers measured at different points in a sweep's merge backlog are not
straightforwardly comparable.

Activity

  1. MarcusKainth commented on Aug 24, 2026

    @MarcusKainth
    CollaboratorAuthor

    The 0.2.0 run gives this a second candidate explanation that it cannot be separated from, so recording it here.

    rows_per_s is exactly corpus / window over a fixed 110,250,000 rows, and the drain window halved as the arm got faster — a uniform 31 s at 0.1.0, 14–18 s for spate:native at 0.2.0. At 14 s one sampler tick is 7.1%, outside the 2.5–3.8% band methodology/measurement.md works through. Filed as #48.

    The two hypotheses predict the same ordering in this run, which is why neither is settled by it: repetitions run in sequence on reused infrastructure, so the first repetition is both the freshest (least merge backlog) and the fastest (shortest window). spate:native rep 1 is first, shortest, highest throughput and lowest CPU per row — consistent with either story.

    Separating them needs one controlled while the other varies. The per-window merge census proposed above does that for this issue: with merge activity on the record, a short-window reading and a merge-loaded reading stop being the same observation.

    Also relevant: #49, on the absence of any A/A noise floor for the regime the suite now runs in. Both this issue and #48 are attempts to explain a spread whose baseline is unmeasured.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions