Repository navigation
fix(ingestor): recompute route_mask inline when merging an uncomputed side (#287) - #295
Conversation
…287) A NULL transmissions.route_mask means "not computed yet", not "no routes". mergeTransmissions OR-ed COALESCE(route_mask, 0) from each side, so when only one side had a computed mask the result was non-NULL. The route_mask backfill only recomputes NULL rows, so it then skipped the merged row and the uncomputed side's route bits were lost permanently. Keep the merged route_mask NULL whenever either side is NULL, so the backfill recomputes it from the survivor's full (merged) set of observations. In the non-NULL branch both sides are known, so a plain OR is exact. The other merged columns were reviewed: first_seen/last_seen are NOT NULL (MIN/MAX), and the fill columns use COALESCE(survivor, loser), which fills an unknown survivor value from a known loser value and never fabricates a sentinel that would block a backfill. Only route_mask had the pattern. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rapport — CS-MacBook PR#295 #287 — head a558c8cStatus: Done — fix + 2 tests pushed, draft PR open, all CI jobs green (no flakes, no re-run). Requirements
Note on requirement 2's testReq 2 is an investigation whose conclusion is "no other column needs the change", so a red-before test is not possible for it. The provided test is a guard (green before and after) that locks the invariant and kills a plausible sentinel-injecting mutant. All other points have a genuine red→green test. [A] CI — per job (run 37404626968)
No job failed; the known-flaky #271 did not trigger, so no re-run was needed. [T] Guardrails [T]/[A]
Remainders
|
Review — CS-pve-agent3 PR#295 — head a558c8cDom: REQUEST CHANGES The fix fixes the reported repro, and the guardrails are clean. But deferring to the backfill adds a new permanent route-bit loss of the same kind as #287: a stored, known bit of the computed side is thrown away when the backfill cannot rebuild it from the observations that survive the merge. The reverse direction (survivor NULL, loser computed) is also untested, and a mutant that breaks it survives the whole suite. Evidence tags: [T] = I ran it; [A] = analysis/code reading; [K] = taken from the PR/issue/CI as stated. Findings
Suggested fix for 1 (and 3). This is the issue's second option. In the exactly-one-NULL branch, compute the mask inside the merge transaction: Answers to the review points
Tests and mutantsTrees.
The PR adds no E2E. That is fine here: the change is an ingestor-only migration path, and the server never runs it during E2E. [A] Mutants. Each was applied to the merged tree. "PR+existing" runs
Reviewer edge tests (scratch only, not pushed):
Guardrails
Not verified
|
…mputed (#287) PR #295 review (CS-pve-agent3, REQUEST CHANGES). The first fix kept the merged route_mask NULL whenever either side was NULL and let the #89 backfill rebuild it. That still loses a route bit: the backfill rebuilds the mask only from the survivor's route_type and the observations that survive the merge, so a bit the computed side stored from a frame that no surviving observation can rebuild (an observation idx_observations_dedup drops during the move, or a frame a later reception overwrote) is thrown away for good — the same loss as #287, in the other direction. mergeTransmissions now recomputes the mask inline, in the merge transaction, whenever either side is NULL: it ORs both sides' stored masks (keeping every bit either side recorded) with the #89 lower bound the backfill would produce (both sides' route_type plus the surviving observation headers). The result is non-NULL, so the backfill leaves the row alone and the route_mask_changes row is logged as for any grown mask, which also tells a running server (finding 3). When both sides are already computed the merge still takes the exact OR. The per-observation header parse the backfill and the merge share is extracted into routeMaskBitFromHeader (finding 4). The file header is corrected to describe the inline recompute rather than a backfill that "holds both sides'" observations, which is not true for a dedup-dropped frame. Tests (red before the fix): the loser-NULL and the previously untested survivor-NULL directions both recompute the union inline and log the change (findings 2, 3); a dedup-dropped frame's stored bit is kept across merge and backfill (finding 1). The both-NULL fill-column guard is unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rapport — CS-MacBook PR#295 runde 2 — head 5e95d93Review feedback addressed (commit The round-1 fix kept the merged Findings
The per-observation header parse the backfill and the merge share is extracted Tests
Mutants (one per finding)
Each mutant was applied to the merged tree and reverted after the run. Guardrails
CI — per job (run 37423779406, head
|
| Job | Result |
|---|---|
| ✅ Go Build & Test | pass (22m49s) |
| 🎭 Playwright E2E Tests | pass (23m34s) |
| 🏗️ Build & Publish Docker Image | pass (57s) |
| 📦 Release Artifacts | skipping (fork guard) |
| 🚀 Deploy Staging / 📝 Publish Badges & Summary | skipping (fork guard) |
No job failed. The known-flaky #271/#301 did not trigger, so no re-run was needed. [T][K]
Review — CS-Macmini PR#295 — head 5e95d93Dom: APPROVE med nits Round 2 resolves the round-1 blocking finding. I reproduced the round-1 repro on three trees and get the numbers the previous reviewer predicted: master Evidence tags: [T] = I ran it; [A] = code reading/analysis; [K] = taken from the PR, the issue or CI. Findings
Answers to the review points1. Round-1 finding 1 (Medium, blocking) — resolved. I rebuilt the reviewer's repro as an independent test that asserts only the end state (after merge and backfill), so it is comparable across trees: survivor
The stored DIRECT bit survives because the caller ORs both stored masks into the recomputed lower bound. The PR's own 2. Both directions. Both have a dedicated test, and I broke each direction separately:
The round-1 mutant that survived the whole suite (guard only one side) has no surviving analogue: the code is symmetric — both masks are tested with 3. Both sides NULL. The merge no longer leaves NULL; it writes the recomputed union. Two cases, both verified against master: [T]
A non-NULL 4. Not too broad. No counterexample exists, by construction. [A] The merged mask is a pure OR of five terms and nothing else: Every term is evidence of a frame actually received for this content hash: a transmission row exists only because an observation created it, and One asymmetry worth naming, not a defect: the merged row's own 5. Performance and locks. Per merge, head reads: 2 point Measured, because the whole batch runs in one
≈ +16 µs per merge, i.e. +9 % on a hold that is already well inside the #89 batch budget. The worst realistic case scales with observations per transmission, not with batch size squared. No long lock held in a loop; the transaction shape is unchanged. The change-row count is finding 4. 6. Requirements and tests.
All four 7. Scope. Tests and mutantsTrees.
E2E against a local Go server on a freshened
The PR adds no E2E. Correct here: this is an ingestor-only one-time migration path that the server never runs during E2E, and the route-mask consumers read masks seeded directly by SQL. [A] Mutants. Each applied to the merged tree and reverted after the run. "PR + existing" =
Reviewer probes (scratch only, never pushed):
RX5 is a second case this PR improves over master beyond the reported repro. [T] Guardrails
Not verified
|
…ge-route-mask-null Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review round 3 (findings 1 and 3). Three guards for code shipped in round 2, red under the reviewer's mutants: - the only DIRECT evidence is a surviving observation header (kills MX2 no scan, MX3 scan the loser's rows, MX4 header bit 0); - the only DIRECT evidence is the loser's route_type (kills MX1); - both sides NULL, with and without evidence: the merge writes a known mask and logs exactly one route_mask_changes row (kills a both-NULL merge that stays NULL, and a merge that only logs for a known survivor). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rapport — CS-pve-agent3 PR#295 runde 3 — head c0d1d51Review feedback addressed (commits Evidence tags: [T] = I ran it; [A] = code reading/analysis; [K] = taken from the PR, earlier reports or CI.
The production diff this round is the comment only. Red → greenFindings 1 and 3 pin code that already shipped in round 2, so the new tests are green on head. Their red state is shown against the mutants and against master's merge code, each applied to a throwaway copy of the merged tree. [T] Tests
The E2E against a local Go server (port 13900) built from head, on a scratch copy of
No E2E is affected by the change. It is an ingestor-only, one-time migration/merge path that the server never runs during E2E. [A] MutantsEach mutant was applied to a copy of the merged tree and run against the full
Guardrails
CI — per job (run 37441403893, head
|
| Job | Result |
|---|---|
| ✅ Go Build & Test | pass (19m18s) |
| 🎭 Playwright E2E Tests | pass (26m01s) |
| 🏗️ Build & Publish Docker Image | pass (49s) |
| 📦 Release Artifacts | skipped (fork guard) |
| 🚀 Deploy Staging / 📝 Publish Badges & Summary | skipped (fork guard) |
The run concluded success and headSha matches. Neither known flake (#271, #301) triggered, so no re-run was needed. [T][K]
Review — CS-Minimax PR#295 — head c0d1d51Dom: APPROVE med nits Independent round-3 re-review. All four round-2 findings are addressed and I reproduced each of Evidence tags: [T] = I ran it; [A] = code reading/analysis; [K] = taken from the PR, Findings
No blocking finding. The shipped behaviour is correct everywhere I probed it, including the three Answers to the review points1. Round-2 finding 1 (test gap) — addressed. I re-ran MX1–MX4 against head myself; all four now
The two new tests are the right shapes and each isolates exactly one source: in 2. Round-2 finding 2 (stale PR body) — addressed. The body now describes the inline recompute, 3. Round-2 finding 3 (both-NULL unpinned) — addressed. 4. Round-2 finding 4 (the change-log row comment) — addressed, and the comment matches the 5. No production behaviour change since round 2 beyond the comment — confirmed. 6. Requirements and tests — each #287 criterion is red before and green after. I dropped
The three round-3 tests guard behaviour that shipped in round 2, so they are green on head; the 7. Not too broad. The merged mask is a pure OR of five terms, every one of them evidence of a Tests and mutantsTrees.
E2E against a local Go server built from the merged tree on port 13700, on a scratch copy of
That one failure is a local timing flake, not this PR: Mutants. Each applied to its own throwaway copy and run against the full
Reviewer probe RZ1 (scratch only, never pushed): survivor computed FLOOD /
So RZ1 is a fourth case this PR improves over master beyond the reported repro, and the one the Guardrails
CI — per job (run 37441403893, head
|
| Job | Result |
|---|---|
| ✅ Go Build & Test | success (19m18s) |
| 🎭 Playwright E2E Tests | success (26m01s) |
| 🏗️ Build & Publish Docker Image | success (49s) |
| 📦 Release Artifacts | skipped (fork guard) |
| 🚀 Deploy Staging | skipped (fork guard) |
| 📝 Publish Badges & Summary | skipped (fork guard) |
Run conclusion success, and it is the only run on this head. Neither known flake (#256 Hash Stats
sort, #267 backfill write-hold) triggered, so no re-run was needed. [T][K]
Not verified
- The performance claim in the body (89 ms vs 97 ms for 500 both-NULL merges, ≈ +16 µs per merge).
I did not re-measure it; I only checked the query shape by reading, which matches the body's
description: the both-known path is query-neutral and the NULL path adds 2 pointSELECT route_typeplus oneidx_observations_transmission_id-driven scan, with no per-observation
query. [A][K] - The server-side consumer of the new change row was read, not exercised against a running server:
RefreshRouteMaskChanges→mergeKnownRouteMaskORs the logged mask and a0row flips
routeMaskKnown. Code reading only. [A] - My
cmd/server,sh test-all.shandnode test-frontend-helpers.jsruns were against the merge
withorigin/masterat30c7de46.origin/masteradvanced to2820d18dduring the review; I
re-merged against it and the result is still conflict-free with the same three files, and none of
the new master commits touchcmd/ingestororinternal/, so the ingestor results carry over
unchanged. The newer master does touchcmd/server,public/and two root node tests, which I
did not re-run. [T] - How often the dedup-drop case and the RZ1 shape actually occur on real data. No staging or
production access was used. [A] - Long-run migration behaviour at production scale. [A]
Relates to #287
Problem
mergeTransmissionsincmd/ingestor/hash_migrate.gofolded the loser'sroute_maskinto the survivor withCOALESCE(route_mask, 0) | COALESCE(loser.route_mask, 0), keepingNULLonlywhen both sides were
NULL. But aNULLroute_maskmeans "not computedyet", not "no routes". When only one side was computed,
COALESCE(…, 0)injected
0for the uncomputed side and the result was non-NULL. Theroute_mask backfill only recomputes
NULLrows (WHERE route_mask IS NULL),so it then skipped the merged row and the uncomputed side's route bits were lost
for good. The issue's repro ended with
DIRECTonly, although the packet wasalso heard as
FLOOD.Fix: recompute the mask inline, in the merge transaction
Both sides known: each stored mask already encodes that side's
route_typeand every observation it ever held, so the merged mask is theplain OR
oldMask | loserMask. There is noCOALESCEon this path, so theoriginal bug cannot come back through it.
Either side
NULL(including both): the merge computes the mask in Go,inside the same write transaction, and writes it as one bound value:
The first three terms are the Same advert seen on flood and zero-hop routes keeps the route of the first inserted observation #89 lower bound that the backfill would
produce, plus the loser's
route_type. They are computed by the newrecomputeMergedRouteMask. The stored masks are ORed in by the caller.The result is always non-
NULL, so the backfill leaves the row alone. If thesurvivor's mask changed, or was
NULLbefore the merge, the merge logs aroute_mask_changesrow, which lets a running server pick up the new maskwithout a restart.
Why not keep the mask
NULLand defer to the backfill? That was round 1,and the round-1 review showed it still loses a bit in the other direction. The
backfill rebuilds only from the survivor's
route_typeand the observationsthat survive the merge. A bit that the computed side stored from a frame no
surviving observation can rebuild would be thrown away. One example is an
observation that
idx_observations_dedupdrops while the loser's observationsare re-parented. Keeping both stored masks in the OR keeps those bits.
The per-observation header parse that the backfill and the merge share is
extracted into
routeMaskBitFromHeader(cmd/ingestor/route_mask_backfill.go),so both read a frame header the same way.
Not too broad. Every term is evidence of a frame actually received for this
content hash.
packetpath.RouteMaskBitreturns 0 outside route types 0..3, anda missing or non-hex header contributes nothing.
|is idempotent. The mergedmask is a superset of both stored masks, so a merge never shrinks a stored mask.
Change-log rows. The row is logged whenever the survivor's mask was
NULL,even when the merged mask is a known
0. On a pre-#89 database, where everymask is
NULL, the one-time migration therefore logs one row per merge. Masterlogged none. The count is bounded (one per merge), the server consumer reads at
most
routeMaskChangesBatchrows per tick, and the rows are pruned.Cost. For each merge, the both-known path still issues 2 point
SELECT route_maskqueries, the same number as before; the post-UPDATEre-read is gone. The
NULLpath adds 2 pointSELECT route_typequeries andone indexed scan of the survivor's observations
(
idx_observations_transmission_id). There is no per-observation query. Theround-2 reviewer measured 500 both-
NULLmerges in one batch at 89 ms onmaster and 97 ms on this branch (≈ +16 µs per merge).
Requirement 2: other columns in the merge
I reviewed every column
mergeTransmissionswrites:first_seen,last_seenMIN/MAXNOT NULL, never unknownroute_maskNULLroute_type,payload_type,payload_version,decoded_json,from_pubkey,channel_hash,scope_name)COALESCE(survivor, loser)The fill columns use
COALESCE(survivor, loser): they fill an unknown (NULL)survivor value from a known loser value and never fabricate a sentinel. A
column whose NULL means "pending" therefore keeps that meaning: if both sides
are NULL, the result stays NULL. Only
route_maskwas merged as a union with aCOALESCE(…, 0)sentinel while also being owned by a NULL-gated backfill, so itwas the only affected column. No other column changed.
Repairing already-merged rows on existing deployments
content_hash_formula_v1has already run on live deployments and is recorded asdone, so it will not re-run. There is no automatic repair, because a safe,
cheap, targeted one does not exist:
mask that is indistinguishable from a legitimately complete one. Nothing
selects only the affected rows.
route_maskback toNULLand re-running the backfill would shrink masks that live ingestlegitimately grew. The backfill is explicitly a lower bound: it does not
re-invent a route variant whose frame no longer exists in any
observations.raw_hex. In live ingest,route_maskis monotonic: bits areOR-ed in and never cleared. A blanket recompute would trade the hash migration: merging a transmission with a NULL route_mask turns it into 0, so route flags are lost permanently #287 loss for
a broader one. It would also rewrite the whole table, so it is not cheap.
during the one-time migration and had exactly one side NULL.
The fix prevents the loss on any deployment that has not run the migration yet.
If the formula ever changes again, the migration gets a new name and re-runs,
and merges are then correct. Writes still live only in the ingestor, and
cmd/serveris untouched.Tests
All seven tests are in
cmd/ingestor/hash_migrate_route_mask_287_test.go. Eachone runs the real
migrateContentHashesmerge and, where relevant, the realbackfillTxRouteMaskafterwards.TestContentHashMigration_MergeRecomputesUnionWhenLoserUnknown_287DIRECT, loserNULLheardFLOOD(the issue's repro)DIRECT|FLOODinline, oneroute_mask_changesrow with the full mask, unchanged by the backfillTestContentHashMigration_MergeRecomputesUnionWhenSurvivorUnknown_287NULLheardFLOOD, loser computedDIRECTDIRECT|FLOODinlineTestContentHashMigration_MergeKeepsStoredBitsOfDedupDroppedObservation_287DIRECTbit whose only frame is dedup-dropped in the moveDIRECT|FLOOD, kept through the backfillTestContentHashMigration_MergeKeepsFillColumnsNullWhenBothUnknown_287NULLNULLTestContentHashMigration_MergeRecomputesBitFromSurvivingObservationHeader_287DIRECTevidence is a surviving observation headerDIRECT|FLOODinline and after the backfillTestContentHashMigration_MergeRecomputesBitFromLoserRouteType_287DIRECTevidence is the loser'sroute_type(its observation has noraw_hex)DIRECT|FLOODinline and after the backfillTestContentHashMigration_MergeBothRouteMasksUnknown_287(subtestswith evidence,without evidence)NULLDIRECT|FLOOD/ known0written by the merge, exactly one change row, unchanged by the backfillTests 1–3 were added red-first in rounds 1–2. Test 4 is a guard and was green
before and after, because requirement 2 needs no code change. Tests 5–7 were
added in round 3 to pin the code shipped in round 2. They guard existing
behaviour, so their red state is shown by the mutants below. Against master's
merge code, tests 1, 2, 3, 5, 6 and 7 fail and test 4 passes.
Mutants
Each mutant was applied to a copy of the merged tree, the tests were run, and
the copy was discarded. The round-3 rows, plus master's merge code, were run
against the full
cmd/ingestorsuite (go test ./...). The round-1 andround-2 rows are taken from the reports of those rounds.
COALESCE(route_mask, 0) | COALESCE(…, 0)mergeCOALESCE(scope_name, ''))NULLsideroute_mask_changesrow..._ConvergesInTheIngestor_215)route_typerouteMaskBitFromHeaderalways returns 0route_mask_backfilltests)NULLmerge staysNULL(master's behaviour)NULL🤖 Generated with Claude Code