Repository navigation
fix(store): move content-hash migration writes to the ingestor and merge duplicates in memory (#215) - #222
Conversation
…e return (#215) Red on master: the server migration writes on the mode=ro handle (every statement fails and is logged as a collision), leaves ghost duplicates that share a hash, keeps nodeHashes keys under the old hash, and leaves stale charges. The ingestor has no content-hash migration yet (these tests do not compile until it exists). Also pins perFallbackRelayBytes in the default run (the =0 mutant survived every default test, #209 review). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…rge duplicates in memory (#215) cmd/server ran UPDATE/DELETE on transmissions and observations on its mode=ro handle. Every write failed, was logged as a collision, and was retried at every start; in memory the colliding rows stayed as ghosts sharing one hash. Ingestor: new async migration content_hash_formula_v1 rehashes stale rows in bounded batches and merges collisions into the lowest id (observations are re-parented, duplicates of the dedup index dropped, last_seen/route_mask folded, ping_triggers follow, the duplicate row is deleted). It is recorded as done, so it runs once. Server: the migration only rehashes in memory and merges the same way (lowest id survives, observations move, the duplicate leaves every index and its charge is credited once). nodeHashes keys are renamed with the hash, and the charge is refreshed for the new hash length. The path-change index refresh of IngestNewObservations is extracted as reindexTxPath so the merge reuses it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ite probe and perf checks (#215) Starting the migration from OpenStore raced every test that opens a store and seeds its own rows with made-up hashes (it rehashed and merged them under the test): TestScopeNameMigration, the prune tests. Like the route_mask backfill it is now started by main once the ingest buffer is draining and cancelled on shutdown; a test pins that wiring. Server: lower the write-SQL ratchet for hash_migrate.go from 3 to 0, and add a probe on a recording connector that fails on any statement other than a read. Env-gated perf checks for the migration (server rehash of 5000 rows, ingestor merge of 1000 pairs and rehash of 5000 rows). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… entries are removed (#215) The path-hop sweep finds a duplicate's resolved relay keys through the hops of its observed paths. The merge cleared the duplicate's observation list first, so when the duplicate had itself taken over a later duplicate's observations (a chain of three rows across batches) a resolved key survived in byPathHop after migration and eviction. Found by comparing byPathHop, spIndex and spTxIndex against a store that never held the rows; the test fixture now has a relay path that changes the survivor's best path, and relay paths consistent with their resolved pubkeys. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ow count (#215) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Rapport — CS-MacBook PR#222 #215 — head f211164Status: Draft PR open, one PR (no split needed), all local tests green, 26 of 26 mutants killed, CI pending (one check queued, not polled by me). Evidence: [T] = ran it, [A] = read in source or diff, [K] = known from earlier work or another report, not re-verified here. Base Design and why
Acceptance criteria
The baseline test found a real bug in my first version [T]: with three duplicates across batches, a middle duplicate that had taken over a later one's observations left a resolved relay key in Tests run [T]
Benchmark [T]Interleaved master and PR test binaries, same fixtures, 5 rounds, median of 7 runs per round. The tests are in the PR and env-gated (
CIPending: the app's PR monitor shows 1 check queued, 0 passing, 0 failing, PR mergeable, review required [T]. I did not poll it. Nothing in CI has been seen yet. Remaining
|
Review — CS-Macmini PR#222 hash-migrate — head f211164Dom: APPROVE med nits Evidence: [T] = ran it, [A] = read in source or diff, [K] = known from earlier work or another report, not re-verified here. Base: head Summary. The data-safety core holds up. With live inserts running alongside the migration, there were no UNIQUE errors and no observations were lost. A cancelled run resumes to exactly the same DB state as an uninterrupted one. No foreign-key violations or orphans appeared, and the server issues no writes [T]. There are no blocking defects. Before ready, I recommend N1 (cancel/resume test), N2 (the Findings
1. Ingestor migrationData loss on a collision [A, T]:
Other references to
Transactions and concurrency [A, T]:
Crash or stop mid-run [T]:
Run once [A]:
Start placement and load [A, T]:
2. ServerNo SQL [A, T]:
Writes the guards miss [A]:
In-memory merge, indexes cleared [A, T]:
Order between server and ingestor [A, T]:
3. Accounting — probes P1–P4 repeated on the merged tree [T]
4. Perf
Round 3 ran while I was compiling mutants and should be ignored [T].
5. Tests, mutantsFull runs on the merged tree (= head), once each [T]:
Mutants (11, each restored and checked byte-for-byte with
CI (run
6. Rules [T]
7.
|
…h-migrate-readonly
…t_seen, dedup key, rechecks (#215) Review round 2. Red on the previous head: the cancelled-run log line, the merge keeping only the survivor's columns (first_seen, scope_name, channel_hash, from_pubkey), the rename walking the whole of nodeHashes per batch, and main starting the in-memory pass right after the first load chunk. Pins what the previous head did right but no test saw: a run cancelled half way is not recorded as done and the restart converges to the uninterrupted state; rows deleted or rehashed between the scan and the batch are skipped; the server applies a stale batch without resurrecting evicted rows; the dedup key keeps the same observer's different paths. Adds two behaviour-neutral seams: contentHashMigrationHook in the ingestor and the hashRekeySweeps counter in the server. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…, whole-store in-memory pass (#215) Review round 2. Ingestor: a cancel is checked at the top of every batch and returned as an error, so a run stopped at shutdown stays unfinished and resumes (a test pins that the restarted run ends in the uninterrupted state). A cancelled run is logged as "cancelled (will resume)", not FAILED. A merge now also takes the earliest first_seen and fills every nullable column the survivor has no value for (read from the table, so later columns are covered); the survivor's copy of an observation that collides on observer and path still wins, documented. Server: the nodeHashes rename and the byNode removal of a merged duplicate find the keys through the transmission (decoded pubkeys, fallback relays, resolved relays via the resolved-pubkey index) instead of walking all of nodeHashes and byNode per batch; the walk stays only as the fallback with the index off. The merge takes the earliest first_seen, fills scope and route type, and puts s.packets, byPayloadType and byNode back in order. main starts the in-memory pass after the whole startup load (StartupLoadDone), not after the first chunk. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Rapport — CS-MacBook PR#222 runde 2 — head 7051289Status: Draft PR, N1–N5 and N8 fixed, N3 solved (cheap), N7 covered, N6 left as noted; all local tests green, 39 of 39 mutants killed including R2b, R3, R4, S1 and S4; CI for the new head not read by me (not polled). Evidence: [T] = ran it, [A] = read in source or diff, [K] = known from the review or earlier work, not re-verified here. Review feedback addressed (commit
PR text: Tests run [T]
Benchmark [T]Interleaved test binaries (previous head
CIFor the new head: the app's PR monitor is showing PR 211 for this session, so I could not read the status of 222 from it, and I did not poll. Nothing in CI has been seen for Remaining
|
Review — CS-Macmini PR#222 runde 2 — head 7051289Dom: APPROVE med nits Evidence: [T] = ran it, [A] = read in source or diff, [K] = known from earlier work or another report, not re-verified here. Scope and tree:
Summary:
N1–N8 status
New findings
1. N1 cancel/resume
2. N2, targeted renameCompleteness [A]:
Probes on the merged tree, index on and off [T]:
Interleaved benchmark [T]. Test binaries: master + the PR's perf file (A),
The resolved-relay fixture covers the one-time-per-batch hash→pubkey map, which the author's fixture does not exercise. 3. N3, whole-store passOrdering and locks [A, T]:
Write-lock hold per batch: see NEW-1. It is O(store) when a batch merges, and about O(batch) otherwise. The 50k no-collision row in the table works out to about 5 ms per 5000 rows. P5 repeated [T]: chunked load with chunk size 3, migration started the way 4. N4 MIN/COALESCE
Server re-sort [A, T]:
5. MutantsThe review's round-1 mutants against the PR's own tests [T]:
New mutants, focused on N2 completeness and N3 concurrency (each restored and checked with
6. Deploy plan in the PR textCorrect and complete for one container under supervisord, matching my round-1 point 8: snapshot (with the dedup-index check), deploy, wait for Is the restart still needed after N3? Yes, it is required. The reason given in step 3 should change:
Two additions:
7. Rules [T]
TestsFull runs on the merged tree
CI, run
Ingestor perf repeated [T]: 120,000 tx and 600,000 observations, batch 2000, file DB, local SSD:
About +7 % from the column fill. The bound per batch is unchanged. RP-I1 is unchanged: 1500 concurrent inserts, 0 errors, observations exact (6000 − 600 dedup + 1500). RP-I4 (no dedup index) is unchanged: both copies are kept [T]. Not verified
Recommendation for the staging round
|
…summary, fill allow-list (#215) Round 3 of the review of #215: - NEW-3 / M2: a TXT_MSG pair and single whose decoded JSON carries destPubKey and srcPubKey; the rename and the removal of the duplicate must find those keys too. - NEW-3 / M6, M7: after a survivor takes an earlier first_seen, s.packets, byPayloadType and the survivor's byNode lists are asserted in order. - NEW-4: fillableColumns is an explicit allow-list (a nullable column added later is not filled); the ingestor fills a NULL scope_name only, and "" is a value; the server's in-memory fill is pinned, with the one documented difference. - NEW-1: the in-memory pass logs the longest write-lock hold of a batch and the number of batches. Red at this commit: the allow-list test fails, and cmd/server does not build (lockHold does not exist yet). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… the scope fill and the reception gap (#215) - ingestor: fillableColumns is an explicit allow-list of the seven columns it filled on today's schema (route_type, payload_type, payload_version, decoded_json, from_pubkey, channel_hash, scope_name), intersected with the columns the table has. A nullable column added later is not filled unless it is listed. NULL and "" in scope_name stay different values (COALESCE). - server: the in-memory pass logs the longest write-lock hold of a batch, the total and the number of batches in its summary line, so a staging run can measure it. No behaviour change. - server: the file header documents why the scope fill differs from the ingestor's COALESCE for a "" survivor (StoreTx cannot tell NULL from "" and has no spare byte for a flag) and the known gap that a reception stored on a row the in-memory merge removed is picked up only after a server restart. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Rapport — CS-MacBook PR#222 runde 3 — head 741a9d6Status: Draft PR, NEW-1 (measurement only), NEW-2 (documented, no code), NEW-3 and NEW-4 are done; all local suites green and M2, M6 and M7 are now killed; CI for the new head was still running when I posted. Evidence: [T] = ran it, [A] = read in source or diff, [K] = known from the review or earlier work, not re-verified here. Review feedback addressed (commit
Tests (local, head 741a9d6) [T]
Rules [A][T]
CIRun Remainders
Proposed follow-up issue for NEW-1 (to be opened only if staging shows a problem)Title: Hash migration: shorten the in-memory write-lock hold of a merging batch Text:
|
…ove-post-packets Brings in #222 (content-hash migration moved to the ingestor), which also rewrote the knownServerWriteSQL comment. The map merged cleanly to backup.go 1, openapi.go 1, ping_score_history.go 15; only the comment conflicted. Both notes are kept: hash_migrate.go lost its 3 literals in #215 and routes.go lost its 3 with POST /api/packets (#223). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Relates to #215
Follow-up to #209. The content-hash migration wrote to the database from
cmd/server, on itsmode=rohandle. This moves the DB half to the ingestor and makes the server half in-memory only, including the merge of colliding transmissions.What was wrong (from the #209 review)
migrateContentHashesAsyncranUPDATE/DELETEontransmissionsandobservationsthrough the read-only handle. Every statement failed, each failure was logged ascollides — merging duplicate, and the same work was retried at every start.byHashpointed at one of them, andQueryPacketsreturned both.nodeHashesis keyed by hash. Keys made before a migration were never renamed, so a later observation was indexed a second time and eviction left the old key behind. After a merge, evicting one ghost also removed the key the other still needed.chargedBytes) was off by the hash length difference.perFallbackRelayByteswas not pinned by any test that runs by default.Investigation (rule 7): is the migration still needed for new rows?
No.
InsertTransmissionhashes withComputeContentHash(cmd/ingestor/db.go), and the two copies of the function are identical. Only rows written before the formula changed (Kpa-clawbot#786, Kpa-clawbot#787) can be stale, so this is a one-time conversion of old data, which is why it belongs in the ingestor and is recorded as done.Design
Ingestor (
cmd/ingestor/hash_migrate.go): async migrationcontent_hash_formula_v1, started frommainafter the ingest buffer is draining and cancelled on shutdown, like the route-mask backfill. It is deliberately not started fromOpenStore: that raced every test that opens a store and seeds rows with made-up hashes (it rehashed and merged them under the test).id, raw_hex, hashin id order, in batches; rewrites each batch in one transaction underwriterMu.ping_triggers).idx_observations_dedupwould reject) is dropped, and the survivor's copy wins, as it does when ingest meets a repeat reception. The earlier copy could win only by copying every observation column, the optional ones (resolved_path,raw_hex) included, and the in-memory observation too; the transmission keeps the earliestfirst_seeninstead. The survivor takes the earliestfirst_seen, the laterlast_seen, the union ofroute_mask(a grown mask is logged inroute_mask_changesfor running servers) and, for each column of an explicit allow-list (route_type,payload_type,payload_version,decoded_json,from_pubkey,channel_hash,scope_name; the ones the table has), the duplicate's value when its own is NULL; a value it has stays,""inscope_name(transport-scoped, region unknown) included. A nullable column added later is not filled unless it is added to the list, because its NULL may mean "pending". Aping_triggersrow follows the survivor, or goes if the survivor has one. The duplicate row is deleted._async_migrations, so it does not run again. If the formula changes again, bump the name.failed,context canceled) and resumes at the next start. It is logged ascancelled (will resume), notFAILED.Server (
cmd/server/hash_migrate.go): no SQL at all. It rehashes in memory and merges with the same rule (lowest id survives), so a server that has not restarted yet and the DB agree on which transmission is the survivor.s.packets,byTxID,byHash,byPayloadType,byNode,byPathHop, the subpath and resolved-pubkey indexes, the distance index andadvertPubkeys. Its charge is credited once; a dropped observation is credited when it is dropped.nodeHashesis renamed with the hash, finding the keys of a transmission through the transmission (its decoded pubkeys, its fallback relays, and its resolved relays through the resolved-pubkey index) instead of walking all ofnodeHashesper batch; the walk stays only as the fallback when the resolved-pubkey index is off. After a merge the lists that held the duplicate are rebuilt, so the shared hash cannot drop the survivor's key and a key only the duplicate held goes.first_seenand fills the scope and route type the survivor lacks (decoded_jsonandpayload_typeare not filled in memory: the content hash includes them, so the duplicates agree). The scope fill differs from the ingestor'sCOALESCEin one rare case, see Known remainders. When the survivor takes an earlierfirst_seen,s.packets,byPayloadTypeand the survivor'sbyNodelists are put back in order (eviction cuts from the head).mainstarts the in-memory pass afterStartupLoadDone()(LoadChunked and the background fill), not after the first chunk, so it sees the whole store.IngestNewObservationsis extracted asreindexTxPathand reused (behaviour unchanged).s.packets, not offsets into it: a merge removes elements and eviction trims the head. The batch is re-checked under the write lock (evicted or rehashed since is skipped).Not carried over: the relays of a merged observation are not re-derived for the survivor in memory (that needs
resolved_pathfrom the DB). They return at the next load, when the DB holds the merged row. Until then they are attributed to nothing, where before they were attributed to a ghost transmission that counted the same packet twice.Guards
TestServerHasNoPacketTableWrites: noUPDATE/DELETE/REPLACEon the packet tables anywhere incmd/server, and no statement orconncall at all inhash_migrate*.go.TestHashMigrate_IssuesNoWriteStatements_215: a probe on a recording driver connection, opened the way the server opens the DB (mode=ro), fails on any prepared statement that is not a read.hash_migrate.gogoes from 3 to 0.TestPerFallbackRelayBytes_IsPinned_215pins the constant from both sides in the default run (the= 0mutant is killed).Not changed
StoreTxstays 320 bytes (both layout tests pass unchanged), no newmap[string]interface{},cmd/serverstays read-only, workflows untouched.Performance
Interleaved test binaries (same fixtures, median of 5 runs per round, 4 rounds in round 2; the tests are in the PR, env-gated). "Before" is the previous head of this PR, "no rename" the same code with the
nodeHashesrename stubbed out, to read off its share:The rename is now about half of the migration at 50000 rows (21 of 43 ms) and about a third at 5000: it costs what it renames (150000 keys), not what the whole index holds, and it no longer grows with the store. Master's numbers include its DB writes on an in-memory SQLite.
Deploy plan
The migration deletes rows (merged duplicates) and a code rollback does not bring them back.
idx_observations_dedupexists on that database: whereobservationspredates it, nothing is dropped as a duplicate and both copies stay._async_migrationshasstatus = 'done'forcontent_hash_formula_v1and the log shows[hash-migrate] rehashed ...and[async-migration] "content_hash_formula_v1" done. Acancelled (will resume)line from a restart in between is expected; the run resumes. Record the ingestor'srehashed X of Y ... merged Zcounts and its duration.IngestNewObservations(the row is not in the store) and the poller's cursor moves past it, so it stays missing from memory until a reload. The common case is the "stale old row + current-hash new row" pair: the ingestor finds the new row by hash and stores new receptions on it. A restart after the ingestor isdoneloads the merged database and recovers them. Smaller reasons, all fixed by the same reload: the survivor's merged columns and observation ids come from the database, the relays of merged observations are re-derived at load, and a scope the server filled from a duplicate where the database kept""is reconciled (see Known remainders)./api/packetslatency to rise: a batch that merges holds the write lock for time that grows with the store (a reviewer measured ~12 ms at 100k and ~38 ms at 300k packets for a batch of 5000). The pass logs one summary line,[hash-migrate] Rehashed ... merged ...; max write-lock hold X ms over N batches (total Y ms): record it. After the restart the pass finds nothing stale and costs nothing.PRAGMA foreign_key_check;returns no rows.SELECT COUNT(*) FROM observations o LEFT JOIN transmissions t ON t.id = o.transmission_id WHERE t.id IS NULL;is 0, and the same forping_triggers.tx_id.[hash-migrate]log line; compare observation counts the same way.ComputeContentHashis unchanged, an older server's own pass finds nothing stale and attempts no write, an older ingestor ignores the unknown_async_migrationsrow). Data: only restoring the snapshot undoes the merges, and loses what was ingested since; the new code then runs the migration again, because the restored database has no record of it.Known remainders
IngestNewObservations, but it has to live until the ingestor finishes the row, which the read-only server cannot see, so a size cap would trade this gap for another. Not done.scope_nameand keeps"". In memory NULL and""are both""andStoreTxhas no spare byte for a flag (it stays 320 bytes), so the server fills a""survivor from the duplicate. They agree for the usual NULL survivor (a row from beforescope_nameexisted, merged with the same packet heard since). They differ for a""survivor with a duplicate that matched a region, which needs region keys to have changed between two receptions of one packet; the database keeps""and memory shows the region until the reload above. Pinned byTestHashMigrate_ScopeFillAgreesWithTheIngestorForNull_215andTestContentHashMigration_ScopeFillOnlyFillsNull_215.finishHashMergerunsslices.DeleteFuncover all ofs.packetsand over each affectedbyPayloadTypelist, and a survivor whosefirst_seenmoves triggers a stable re-sort ofs.packetsand of those lists. Follow-up issue to be opened if staging shows a problem.s.packetswithout the read lock survives-race: the window is a few instructions long. Not asserted.s.packets(first_seen) order; they could fill a NULL column from different copies when the copies disagree on a non-NULL value. Copies of one packet do not normally disagree. [analysis only]Tests
New and affected tests with
-race -count=3, both full suites once,go vet,gofmt -lon changed files, mutants per acceptance criterion: see the report comment on this PR.🤖 Generated with Claude Code