Problem
TestResolvedPathBackfill_WriteHoldUnderBudget_188 in cmd/ingestor/resolved_path_backfill_188_test.go failed on master CI. The run was 37321760306 for 0572e7f9, in "Build and test Go ingestor (with coverage)":
max write hold per 500-row batch: 263.738039ms (budget 250ms)
max write hold 263.738039ms exceeds the budget 250ms
The test measures the wall-clock write-transaction hold of the slowest of 10 batches (5000 rows, batch size 500) and asserts it is under a fixed 250 ms. The result depends on runner load and coverage instrumentation, not only on the code. The commit under test changed only two unrelated test files, and the same test passed on master earlier the same day.
Proposed approach
Keep the guarantee the test protects: one backfill batch must not hold the single write connection for long enough to stall live ingest. Make the assertion robust:
- Assert the structural property deterministically: rows per write transaction ≤ the batch size, one transaction per batch, and no work done while holding the transaction that could have been done before it (for example, resolution happens before
BEGIN).
- If a timing assertion stays, make it relative and tolerant. For example, compare the max hold against the median batch, or against a per-row budget calibrated in the same run. Alternatively, skip or relax it under
-cover and -race, and keep a separate benchmark (BenchmarkResolvedPathBackfillBatch) that reports the hold.
- Investigate first whether the 264 ms is real work inside the transaction that could be moved out, or only scheduling noise. Name the cause in the PR.
Acceptance
Problem
TestResolvedPathBackfill_WriteHoldUnderBudget_188incmd/ingestor/resolved_path_backfill_188_test.gofailed on master CI. The run was 37321760306 for0572e7f9, in "Build and test Go ingestor (with coverage)":The test measures the wall-clock write-transaction hold of the slowest of 10 batches (5000 rows, batch size 500) and asserts it is under a fixed 250 ms. The result depends on runner load and coverage instrumentation, not only on the code. The commit under test changed only two unrelated test files, and the same test passed on master earlier the same day.
Proposed approach
Keep the guarantee the test protects: one backfill batch must not hold the single write connection for long enough to stall live ingest. Make the assertion robust:
BEGIN).-coverand-race, and keep a separate benchmark (BenchmarkResolvedPathBackfillBatch) that reports the hold.Acceptance
go test -count=50 -run TestResolvedPathBackfill_WriteHoldUnderBudget_188 -cover ./...incmd/ingestorunder parallel CPU load.