Repository navigation
test(ingestor): make the backfill write-hold test robust under CI load (#267) - #272
Conversation
TestResolvedPathBackfill_WriteHoldUnderBudget_188 asserted a wall-clock hold under 250 ms per 500-row batch and failed under CI load. Per-phase timing under -cover and parallel load puts about three quarters of the hold in COMMIT (the WAL fsync, synchronous=FULL); resolution already runs before the transaction and nothing inside it can move out. The timing check also missed resolution moved into the transaction and all rows deferred to one transaction. The test now asserts the bound itself: one WriterTx per batch, at most batchSize rows updated per transaction (seen by SQLite triggers calling a Go probe from inside the transaction), and no resolution while writerMu is held (through a resolver seam). The hold is logged, and BenchmarkResolvedPathBackfillBatch reports it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The triggers that report the backfill's writes run a Go callback per row inside the transaction, so the hold the test logs is a little higher than production. BenchmarkResolvedPathBackfillBatch measures it without them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rapport — CS-pve-agent2 PR#272 #267 — head bc04203Status: Draft ready for review. The cause is found and the test now asserts the bound structurally. CI is green; the PR is not marked ready. Evidence tags:
Cause
Fix
Measurements, before and after
Mutants
Local checks
CI per job (run 37349971997, head bc04203)
Remaining
|
Review — CS-pve-agent3 PR#272 — head bc04203Dom: APPROVE with nits Evidence tags:
Independent, read-only review. Head checked with Findings
1. Cause[T] I timed each phase again with a temporary patch in a scratch copy, not pushed. Setup: head tree,
2. The structural assertion[K] The test asserts:
The probe function is registered once, before the store opens its connection. The seam and the probe pointer are restored in Mutants, each applied to
Both regression shapes named in the issue (m1, and m2/m3) now fail deterministically. Before this PR, only m2 was caught. 3. Stability under load[T]
4. Production code
Acceptance criteria (issue #267)
Invariants
Tests on the merged tree (
|
Relates to #267
Cause
The 264 ms overshoot is disk-flush latency in
COMMIT, not work that could move out of the write transaction.I timed each phase of
resolvedPathBackfillBatchwith a temporary patch that is not part of this PR. Setup:0572e7f9,-cover, two to fourgo test ./...loops oncmd/serverrunning alongside, load average 5–6 on 4 cores, 200 batches of 500 rows:PRAGMA synchronousis 2 (FULL), the driver default.PRAGMA synchronous=OFFon the same load, the hold drops to a 3.1 ms median and a 12.8 ms max (COMMIT median 0.12 ms).BEGIN. The UPDATEs, the watermark upsert and COMMIT must stay atomic, so nothing in the transaction can move out without changing semantics. The production code is unchanged, apart from one test seam.The old timing assertion also did not protect what it was meant to protect. These three mutants passed the old test on master:
Even 5000 rows in one transaction took only 45–160 ms. The old test caught only the "one batch reads everything" mutant, and only through its
Batches == 10check.Fix
TestResolvedPathBackfill_WriteHoldUnderBudget_188now asserts the bound structurally and deterministically:WriterTxcount for theresolved_path_backfillcomponent (from the writer stats) equalsres.Batches.batchSizerows per transaction: test-only SQLite triggers onobservations.resolved_path(NULL → value) and onresolved_path_backfill_statecall a Go function,backfill_probe_267, from inside the transaction. The writes are split at each watermark write, which a batch makes last in its transaction. The test asserts one watermark per batch, ≤ 500 rows per transaction, 5000 rows in all, no row update without a following watermark, and every write made withwriterMuheld.resolveObservationPathis now called throughresolvedPathBackfillResolve, a variable, likereadProcSelfIOFn. The test's wrapper counts resolutions made whilewriterMuis held. The count must be 0, out of 5000.No timing assertion remains. Why:
The hold is still logged.
BenchmarkResolvedPathBackfillBatchreports it ashold-ms/batchwithout the probe triggers. The test's logged hold includes the probe's per-row callback, so it reads about 10–20 ms higher.The test fixtures
backfillFixture188andprimeIndexAndGraph188now taketesting.TB, so the benchmark can use them.Measurements
-count=50 -cover, two load loops: 50/50 pass, hold 17.8 / 29.4 / 89.5 ms (min / median / max).go test -count=50 -run TestResolvedPathBackfill_WriteHoldUnderBudget_188 -cover ./..., four load loops:BenchmarkResolvedPathBackfillBatch: 17.5 hold-ms/batch idle, 36.7 hold-ms/batch under load.Mutants (new test)
WriterTxclosureChecks
cd cmd/ingestor && go test -race -count=1 -timeout 90m -v ./...: ok in 1643 s (963 PASS lines), no data race reported. A run with the default 10 min timeout timed out inTestConcurrentWrites, which takes 111 s alone under-raceon master; CI uses-timeout 20m.go vet ./...incmd/ingestor: cleangofmt -lon the two touched files: cleansh test-all.sh: 218 of 219 files pass.test-channels-client-state-152.jsfailed once (R4-3 S2, "Decrypting…" pane). It is unrelated, since no JS is touched. Standalone it passed 9 of 10 runs while other tests loaded the machine, so it is a pre-existing timing flake.Writes stay in the ingestor. No new
map[string]interface{}. No workflow files are touched: fork guards are 9 indeploy.ymland 1 inrelease-fast-path.yml.🤖 Generated with Claude Code