Repository navigation
feat(ingestor): opt-in retention for inactive_nodes, node_changes and observers (#329) - #350
Conversation
…ng_triggers and observers (#329) Four new retention knobs, each in days, 0/unset = keep forever (today's behaviour): inactiveNodeDays, nodeChangeDays, pingTriggerDays and observerPurgeDays. The ingestor applies them at startup and daily after the observer soft-delete, in bounded WriterTx batches, and logs the counts as [prune] lines. - inactive_nodes: by last_seen, skipping nodes that came back (their old row carries confirmed default_scope evidence and keeps them out of New Nodes) and observers still uploading (#199). - node_changes: by detected_at, over idx_node_changes_detected_at. - ping_triggers: by their own first_seen, not with their transmission. The server's Ping Scores history drops an entry once its trigger is gone, and data_pruned entries are exactly those whose transmission was pruned. No index covers first_seen, so the table is walked by tx_id ranges, starting below every key (ids can be 0 or negative). - observers: port of upstream#1886 (#34), plus the observer_neighbor_metrics guard and deletion of the current-only observer_neighbors snapshot. PruneOrphanRouteMaskChanges now uses the shared key-range walk. Co-authored-by: efiten <erwin.fiten@gmail.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ger is kept (#329) Pins the server half of why retention.pingTriggerDays has its own age: a data_pruned entry keeps rendering after its transmission is pruned, and leaves the history store (sender name included) once its ping_triggers row is deleted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Opt-in (CORESCOPE_RETENTION_PERF=1). Seeds ~5 GB of synthetic history and reports writer-lock hold, [db-slow-writer] lines and an ingest probe's lock wait for today's prunes and for the new table retention, on the first run, the next day and a run with nothing to delete. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rapport — CS-pve-agent3 PR#350 #329 — head 398af2bStatus: Draft ready for review: all four knobs implemented test-first, 24/24 mutants killed, CI green on every job that runs, and no Evidence: [T] = executed (test, command or CI output), [A] = analysis or argument, [K] = read in code or docs. Requirements
Always
Local runs
CI (run 37608978680,
|
| Job | Result |
|---|---|
| Go Build & Test | [T] pass (23m18s) |
| Playwright E2E Tests | [T] pass (25m22s) |
| Build & Publish Docker Image | [T] pass (57s) |
| Release Artifacts | skipped (tag-only) |
| Deploy Staging | skipped (push/fork-guarded) |
| Publish Badges & Summary | skipped (push-only) |
The known flaky test #271 did not appear; no re-run was needed.
Remaining
- [A] On first enable, the backlog delete can emit one or two
[db-slow-writer]lines atCOMMIT: 1 line (516 ms) in run 1 and 2 lines (615 ms and 1157 ms) in run 2. Today's prunes emit 85–128 in the same pass, so this is below today's level, but an operator may see them once. The daily steady state emitted none in either run. - [T] The two perf runs differ by disk noise: today's first-run
prune_metricshold was 2.1 s in run 1 and 14.0 s in run 2. The run-2 binary predates the last commit's start-key change (math.MinInt64instead of 0), which does not affect timing [A]. - [T] The local full Go suites ran before that last start-key change; CI ran both full suites on the final head.
- [K] Like the other retention values, the four knobs are read at ingestor startup, not hot-reloaded.
- [A] With
pingTriggerDaysset, the all-time Ping Scores records are bounded to that window. This is the cost of removing sender names, and it is documented. - No staging or production validation was done.
- port(upstream#1886): opt-in observerPurgeDays hard-delete for long-inactive observers #34 is left open; the PR description says this PR supersedes it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Operator decision on #329: ping_triggers keeps every row, because the all-time Ping Scores records join it with the history sidecar (#349, #241). TestTableRetentionLeavesPingTriggers sets every other knob (and a stale pingTriggerDays), runs the transmission prune first, and requires every ping_triggers row unchanged. The config test requires pingTriggerDays to be ignored, and the docs test requires neither config.example.json nor the user guide to offer it. Red on 259684f: the stale knob deletes the old triggers, TableRetention reads it, and both docs still list it. The tests of the removed knob go. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…329) Scope change on #329: ping_triggers must not be pruned (Ping Scores all-time, #349/#241). Remove the knob end to end: - RetentionConfig/TableRetention lose PingTriggerDays; a config that still sets it is ignored. The startup log lists the three knobs. - PrunePingTriggers and its wiring in runTableRetention go, and with them deleteInKeyRanges, which only existed to share the key-range walk with it: PruneOrphanRouteMaskChanges is back to master's code. - The server read-only invariant no longer names PrunePingTriggers. Of the two server tests, the data_pruned-renders one stays (the #329 acceptance criterion); the one pinning the knob's cost goes. - The perf test no longer prunes ping_triggers and fails if retention deletes one. - config.example.json and the user guide drop the knob and say that ping_triggers is kept forever. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rapport — CS-pve-agent3 PR#350 runde-2 — head 7228e08Review feedback addressed (commit Evidence: [T] = executed (test, command or CI output), [A] = analysis or argument, [K] = read in code or docs.
No finding was disputed. Mutants (round 2, one or more per finding, all killed)Run on a copy of the tree, one at a time, with the named test.
[T] Re-adding the knob itself (config field, Local runs (head
|
| Suite | Result |
|---|---|
cd cmd/ingestor && go test -count=1 -timeout 45m ./... |
[T] ok (1360 s) |
cd cmd/server && go test -race -count=1 -timeout 45m ./... |
[T] ok (1741 s) |
sh test-all.sh |
[T] 228 passed, 0 failed |
node test-frontend-helpers.js |
[T] 709 passed, 0 failed |
E2E on a local Go server, e2e-fixture.db prepared as in CI (freshen, seed SQL, corescope-migrate, three seed files) |
[T] test-e2e-playwright.js 132/135 (3 skipped); test-issue-199-inactive-observer-e2e.js 3/3; /api/ping-scores, /api/analytics/node-changes, /api/observers returned 200. Server stopped via pid |
Real ingestor binary on a copy of that fixture: all three knobs at 30 and a stale pingTriggerDays: 30, with a local stub MQTT broker so startup passes the connect step |
[T] logged table retention enabled (0 = off): inactiveNodeDays=30 nodeChangeDays=30 observerPurgeDays=30; deleted the 200-day-old inactive_nodes and node_changes rows (kept the #199 seeded node, last seen 10 days ago); ping_triggers dump md5 identical before and after, including a 200-day-old row |
TestTableRetentionTiming (opt-in, 5.34 GB), run 3 |
[T] PASS, 0 ping_triggers deleted. #329 first run: 2 [db-slow-writer] lines (both prune_node_changes COMMIT, 769 / 1016 ms) against 16 for today's prunes in the same pass; next day 0, nothing to delete 0. Full table in the PR description |
bash scripts/check-xss-sinks.sh --diff origin/master |
[T] no frontend changes to scan |
No new map[string]interface{} outside tests; workflows untouched |
[T] git diff origin/master...HEAD adds none; .github/ diff is empty |
[T] The first attempt at the two full Go suites hit Go's default 10-minute -timeout while both ran in parallel with the mutant runs, with no failing test before the timeout. They were rerun with -timeout 45m (CI uses 20m) and passed.
Branch
[T] origin/master 5cdeffd5 was merged in with a merge commit (259684f3, no conflicts) before the round-2 commits. Master has since moved to e8e99c4a (frontend-only changes), and git merge-tree shows that it still merges cleanly, so no second merge was made. No rebase, amend or force-push; the push was a fast-forward 398af2b6..7228e083.
CI (run 37622679705, pull_request, head 7228e08, attempt 1)
| Job | Result |
|---|---|
| Go Build & Test | [T] pass (20m3s) |
| Playwright E2E Tests | [T] pass (23m13s) |
| Build & Publish Docker Image | [T] pass (48s) |
| Release Artifacts | skipped (tag-only) |
| Deploy Staging | skipped (push/fork-guarded) |
| Publish Badges & Summary | skipped (push-only) |
The known flaky tests #271 and #301 did not fail; no job was re-run.
Remaining
- [K] The round-1 feature commit
7ebf4e7estill namesping_triggersin its subject. History was not rewritten (no amend or force-push); the round-2 commits record the removal. - [A] On first enable, the backlog delete can still emit one or two
[db-slow-writer]COMMITlines (run 3: two inprune_node_changes). Today's prunes emitted 16 in the same pass. The daily steady state emitted none in all three runs. - [K] Like the other retention values, the three knobs are read at ingestor startup, not hot-reloaded.
- No staging or production validation was done. port(upstream#1886): opt-in observerPurgeDays hard-delete for long-inactive observers #34 is left open; the description says this PR supersedes it. The PR stays a draft.
Review — CS-Minimax PR#350 — head 7228e08Dom: APPROVE med nits Independent, read-only round-2 review. Evidence: [T] = I executed it, [A] = analysis, [K] = read in code/docs. Everything below was run on the merged tree ( Findings
No blocking finding. No behaviour change outside the three opt-in knobs: with every knob unset the only added work is three calls that return immediately, which I also confirmed empirically on a real fixture DB (point 2 below). [T][A] The review points1.
2. Per knob: old rows removed, recent kept, unset changes nothing, batched inside
3. Perf on a large DB. Reproduced independently and the reported shape holds; my numbers are better than the author's, on a faster disk. [T]
Deleted on the first run: 4. Writes, docs, title/description. Confirmed.
Always-checks
Tests I ran (merged tree, against master
|
| Suite | Result |
|---|---|
cd cmd/ingestor && go test -count=1 -timeout 45m ./... |
[T] ok, 108.8s |
cd cmd/server && go test -race -count=1 -timeout 45m ./... |
[T] ok, 529.1s |
go vet ./... in both modules |
[T] clean |
sh test-all.sh |
[T] 230 passed, 0 failed (230 files) |
node test-frontend-helpers.js |
[T] 709 passed, 0 failed |
test-e2e-playwright.js against a local Go server on the CI-prepared e2e-fixture.db (freshen, the Kpa-clawbot#1486/Kpa-clawbot#1791 seed SQL, corescope-migrate, seeds 2073, 199 and 245) |
[T] 132/135 passed, 3 skipped |
test-issue-199-inactive-observer-e2e.js |
[T] 3/3 |
test-issue-2073-recent-adverts-e2e.js |
[T] 10/10 |
test-issue-1639-observers-sort-e2e.js, test-observer-iata-1188-e2e.js |
[T] pass |
/api/ping-scores, /api/analytics/node-changes, /api/observers |
[T] 200 |
TestTableRetentionTiming (opt-in, 5.34 GB) |
[T] PASS, table above |
Real-fixture retention run (three passes, ping_triggers md5 pinned) |
[T] PASS |
The server was started on a free high port and stopped by looking the listener up by port, never with $!. [T] One environment note: with a server built outside a git checkout, /api/health reports commit: unknown, so public/perf.js renders no Version card and test-e2e-playwright.js fails fast on "Version info lives on Perf dashboard, not in navbar". [T] It fails identically on master's tree under the same conditions, and CI avoids it because the server runs inside the checkout, where resolveCommit() falls back to git rev-parse. I reproduced CI by writing a .git-commit file; after that the suite ran to completion.
CI
[T] Run 37622679705, pull_request, head 7228e083, attempt 1, no re-runs. Per job: Go Build & Test pass (20m3s), Playwright E2E pass (23m13s), Build & Publish Docker Image pass (48s); Release Artifacts, Deploy Staging and Publish Badges skipped (tag-/push-/fork-gated). Neither known flake (#256 Hash Stats sort, #267 backfill write-hold) appeared.
My mutants (7, all killed)
Applied one at a time to a copy of the tree, each reverted from a pristine snapshot before the next.
| # | Mutant | Killed by |
|---|---|---|
| M1 | pruneInactiveNodesBatch loses the NOT EXISTS (… FROM nodes …) returned-node guard |
TestPruneInactiveNodesKeepsReturnedNode [T] |
| M2 | the observer purge loses the observer_neighbor_metrics guard |
TestPurgeStaleObserversKeepsObserverWithNeighborMetrics [T] |
| M3 | deleteInBatches returns after the first batch |
TestPruneInactiveNodesBatches, TestPruneNodeChangesBatches (deleted=2 left=4, 1 tx instead of 3) [T] |
| M4 | PruneNodeChanges guards on days < 0, so an unset knob prunes with cutoff = now |
TestRunTableRetentionUnsetChangesNothing, TestRunTableRetentionEachKnob [T] |
| M5 | PruneNodeChanges also deletes ping_triggers older than its cutoff |
TestTableRetentionLeavesPingTriggers [T] |
| M6 | vacuity / master behaviour: all three prunes return 0, nil |
9 tests, incl. every "deletes" test and TestTableRetentionLeavesPingTriggers on its non-vacuity assertions [T] |
| M7 | the observer purge leaves observer_neighbors behind |
TestPurgeStaleObserversDeletesNeighborSnapshot [T] |
Not verified
- No staging or production validation, and no server API key used — out of scope by instruction.
- The real ingestor binary end-to-end (it needs an MQTT source to pass the connect gate). I covered the same ground with the real-fixture run above, which exercises the production schema and the real
WriterTxpath but notmain()'s startup ordering; that ordering I only read. [K] CORESCOPE_RETENTION_PERFwas run once, not repeated, and on a different disk than the reported runs, so the absolute millisecond figures are not comparable to run 3's — only the relative claim is.scripts/check-xss-sinks.shwas run in a non-git export, so it reported "nothing to scan" rather than diffing; the conclusion rests on the file list of the diff. [A]- The Go suites were run without the ingestor
-raceflag (matching CI, which only racescmd/server). - The other CI jobs' content (channel lib, decrypt CLI, Docker image, css-vars lint) — taken from the green CI run, not re-run locally.
- Whether
#34is actually closed when this lands; I only read that the description says it supersedes it.
Relates to #329
Supersedes #34: its upstream#1886 port (
observerPurgeDays) is carried over here onto current master, with the two neighbour tables that port did not cover (see Observers below). #34 can be closed when this lands.Scope (operator decision on #329, 2026-10-07):
ping_triggersis not pruned. It keeps every row, because the all-time Ping Scores records join it with the history sidecar (#349 / #241). Round 2 removed thepingTriggerDaysknob from round 1 with all its code, tests and docs, and added a test that locksping_triggersas untouched.Plan
One logical change plus test commits:
feat(ingestor): three opt-inretentionknobs, the batched prunes, wiring at startup and in the daily observer-retention pass,config.example.json+ user-guide docs, and the read-only invariant extended to the three new method names.test(server): Ping Scores entries already markeddata_prunedstill render (the retention: opt-in pruning for inactive_nodes, node_changes, ping_triggers and soft-deleted observers #329 acceptance criterion).test(ingestor): an opt-in timing test on a multi-GB synthetic DB.test(ingestor)locksping_triggersout of the retention, andfix(ingestor)drops thepingTriggerDaysknob.Operator note: the target deployment will set all three knobs to 30 days. With
observerDays14,packetDays30 andmetricsDays30 that is consistent: an observer is purged once its packets and metrics have aged out on the same 30-day clock.No frontend change, so there are no customizer implications. The knobs are retention policy in
config.json, likepacketDays.Change
New knobs in
retention. Each is in days; 0 or unset keeps every row, exactly as today:inactiveNodeDaysinactive_nodeslast_seen(last advert) is older, except nodes that came back (anodesrow exists) and observers still uploading (#199, same predicate asMoveStaleNodes)nodeChangeDaysnode_changesdetected_atolder, oldest first overidx_node_changes_detected_atobserverPurgeDaysobserversinactive = 1), not seen for N days, and referenced by noobservations,observer_metrics,observer_neighbor_metricsordropped_packetsrow; theirobserver_neighborssnapshot goes in the same transactioncmd/ingestor(table_retention.go). They run at startup (after the metrics and transmission prunes, before the ingest gate opens) and in the existing daily observer-retention pass, right afterRemoveStaleObservers.WriterTxbatches of 1000 rows, likepruneBatches, so ingest waits for one batch at most. Each table logs a[prune] deleted N … older than D daysline; startup logs the enabled windows once.cmd/servergains no code.TestServerDBHasNoWriteMethodsnow also forbidsPruneInactiveNodes,PruneNodeChangesandPurgeStaleObserverson the server's*DB.ping_triggers: kept foreverNo knob, no delete. The transmission prune already leaves it alone. A config that still sets
pingTriggerDays(from round 1 of this draft) is ignored. The user guide andconfig.example.jsonsay thatping_triggersis kept and why.TestTableRetentionLeavesPingTriggerssets every other knob and that stale key, runs the transmission prune first, and requires everyping_triggersrow to be unchanged, including one older than all windows and one whose transmission was pruned.inactive_nodes: why returned nodes are keptMoveStaleNodesdoes not delete a node's oldinactive_nodesrow when it comes back. That row still carries confirmeddefault_scopeevidence thatUpdateNodeDefaultScopechecks across both tables, and it keeps the node out of New Nodes. Deleting it while the node is active would let inference downgrade confirmed scope and list a returning node as new. Once the node goes stale again,MoveStaleNodesreplaces the row with a freshlast_seen.A node purged after
inactiveNodeDaysthat adverts again later counts as new: it gets noresurrectedchange and shows in New Nodes. That is the intent of retention, and it is documented.Observers (from #34)
Taken from #34 / upstream#1886 (
PurgeStaleObservers, its tests, the docs), not the old branch as a whole. The review on #34 found twoobserver_idtables the port did not guard. I made these two decisions:observer_neighbor_metricsages out bymetricsDayslikeobserver_metrics, so it is a guard: the observer waits until those rows are gone.observer_neighborsis a current-only snapshot that only the observer itself replaces. As a guard it would keep a dead observer forever. Left behind, it would keep serving the dead observer's neighbours (GetObserverNeighborsdoes not joinobservers). So it is deleted with the observer in the same transaction.Each batch reads its candidate ids once and deletes exactly those. Rows with a NULL id, which SQLite allows in a
TEXT PRIMARY KEY, are skipped.Tests
Test-first: the new tests were run against a stub with master behaviour (no pruning), and every "deletes" test was red. They are green with the change. Per knob:
TestRunTableRetentionUnsetChangesNothing, plus each knob prunes only its own table);retentionBatchRowslowered in the test);inactive = 1, NULL id;ping_triggersuntouched:TestTableRetentionLeavesPingTriggers;pingTriggerDaysignored inTestTableRetentionConfig; neither doc offers it inTestConfigExampleDocumentsTableRetention. All three were red on the round-1 head;TestPingScoresDataPrunedEntryRendersWhileTriggerKept(server);config.example.jsondocuments the three knobs at 0.Mutation testing: round 1 had 24 targeted mutants, all killed. Round 2 adds 8, all killed. The tables are in the report comments.
Perf
TestTableRetentionTiming(opt-in,CORESCOPE_RETENTION_PERF=1) seeds a 5.34 GB synthetic DB:inactive_nodes, 500knode_changes, 50kping_triggers(no retention deletes them; the test fails if one goes);It runs today's prunes and then the table retention with the three windows at 30 days, in production order. A probe goroutine takes the writer lock every 5 ms, standing in for MQTT ingest. The default
[db-slow-writer]threshold of 500 ms applies.Run 3 is on the round-2 code (three knobs), with the fixture seeded fresh. Runs 1 and 2 (round 1, four knobs) are kept for comparison; their #329 numbers include the removed
ping_triggersprune (51 tx, hold max 24 ms in run 2).[db-slow-writer]lines run 1 / 2 / 3Run 3's first-run backlog deletes 18,000
inactive_nodes, 458,918node_changesand the 225 unreferenced soft-deleted observers, and noping_triggers. Its two slow lines are bothquery=COMMITinprune_node_changes(769 ms and 1016 ms, out of 459 batches; p99 230 ms). Those are commit/checkpoint stalls on a disk that, in the same pass, gave today's prunes 16 slow lines. No pass shows a spike above today's level, and the daily steady state adds no[db-slow-writer]line. Per-component first-run hold max, run 3:prune_inactive_nodes21.7 ms (19 tx),prune_node_changes1015.5 ms (459 tx),purge_observers39.8 ms (1 tx).Complexity: every batch is bounded (1000 rows deleted).
inactive_nodesandnode_changeswalk theirlast_seen/detected_atindex. The observer guards are four index seeks per candidate (EXPLAIN QUERY PLANchecked).Verification (round 2, head
7228e083)cd cmd/ingestor && go test -timeout 45m ./...: ok.cd cmd/server && go test -race -timeout 45m ./...: ok.sh test-all.sh: 228 files passed.node test-frontend-helpers.js: 709 passed.e2e-fixture.db, prepared as in CI (freshen, seed SQL,corescope-migrate, the three seed files):test-e2e-playwright.js132/135 passed (3 skipped);test-issue-199-inactive-observer-e2e.js3/3.pingTriggerDays: 30, logged the three knobs, deleted the oldinactive_nodesandnode_changesrows, and left a 200-day-oldping_triggersrow byte-identical.bash scripts/check-xss-sinks.sh --diff origin/master: clean (no frontend change). No newmap[string]interface{}outside tests. Workflows untouched.🤖 Generated with Claude Code