Repository navigation
port(upstream#1886): opt-in observerPurgeDays hard-delete for long-inactive observers - #34
adminopenclaw8-sketch wants to merge 2 commits into
Conversation
…observers (Kpa-clawbot#1886) ## Problem `RemoveStaleObservers` only soft-deletes — it sets `inactive = 1` and the row stays forever. On a long-running deployment those rows just accumulate: on a two-year-old instance roughly 25% of the `observers` table was rows nobody can ever see again. There is currently no way to reclaim them. ## Fix A second retention stage. `PurgeStaleObservers` hard-deletes rows that are: - already `inactive = 1` (so the soft-delete stage owns the decision of *when* an observer goes stale), **and** - older than `retention.observerPurgeDays`, **and** - referenced by nothing. New config field `retention.observerPurgeDays`, default `0` = disabled. Existing deployments are unaffected until they opt in. Set it above both `observerDays` and `packetDays` — below those the reference guards keep every candidate row anyway. ## Why the reference guards are the point `observations.observer_idx` is a bare rowid with no foreign key. Deleting a still-referenced observer silently orphans history — `packets_v` stops resolving the observer and those packets get mis-attributed. Nothing errors; the data just quietly goes wrong. So the statement guards on all three referencing tables: ```sql AND NOT EXISTS (SELECT 1 FROM observations o WHERE o.observer_idx = observers.rowid) AND NOT EXISTS (SELECT 1 FROM observer_metrics m WHERE m.observer_id = observers.id) AND NOT EXISTS (SELECT 1 FROM dropped_packets d WHERE d.observer_id = observers.id) ``` This is correctness, not defensive padding — it was found the hard way, by orphaning 280 observation rows during a manual purge that skipped one of these checks. Each guard has its own test. ## Performance Each `NOT EXISTS` is an index seek per candidate row (`idx_observations_observer_idx`, `idx_dropped_observer`, the `observer_metrics` PK), and `observers` is O(100). It runs on the existing daily retention tick alongside `RemoveStaleObservers`, never on the ingest path. ## Tests Eight tests in `cmd/ingestor/observer_purge_test.go`, written before the implementation: - deletes an unreferenced stale row - keeps a row referenced by `observations` — and asserts zero orphans afterwards - keeps a row referenced by `observer_metrics` - keeps a row referenced by `dropped_packets` - keeps a row that is old enough but still `inactive = 0` - keeps a row inside the retention window - no-ops when disabled (`0` and `-1`) - config accessor table test ## Invariant Writes stay in `cmd/ingestor` per Kpa-clawbot#1283. `cmd/server/readonly_invariant_test.go` now also forbids `PurgeStaleObservers` as a method on the server's `*DB`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit d821d9a) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Verificeret lokalt — og et hul i slettekontrakten er bekræftet ved eksekveringBranchen er synkroniseret med master Slettekontrakten, som den faktisk er
AND NOT EXISTS (SELECT 1 FROM observations o WHERE o.observer_idx = observers.rowid)
AND NOT EXISTS (SELECT 1 FROM observer_metrics m WHERE m.observer_id = observers.id)
AND NOT EXISTS (SELECT 1 FROM dropped_packets d WHERE d.observer_id = observers.id)Men der findes to tabeller mere på Konsekvensen er synlig i API'et, ikke kun i databasen:
Og rækkerne ryddes ikke op ad anden vej: Docstringen siger, at de tre garder "are what make the delete safe; they are correctness, not defensive padding". Det er rigtigt for de tre tabeller, de dækker — men sætningen læses let som at kontrakten er komplet, og det er den ikke. Dette kræver en produktbeslutning før merge, og det er derfor PR'en ikke bare venter på staging:
Jeg har ikke valgt for dig — det ændrer, hvad "hard-delete" betyder. Andre forhold, jeg har kontrolleret
Hvorfor PARKERETUd over produktbeslutningen ovenfor kræver opgavens regler staging før merge for retention-/sletteændringer, og udtrykkeligt med syntetiske data, aldrig produktionsdata, plus verificeret database-backup og isoleret mutationstest. Staging kan ikke nås herfra: docker-dæmonen kører ikke, Konkret testplan ligger i 🤖 Generated with Claude Code |
Split out of #25 (commit
61efd416there). This branch holds exactly one upstream change so it can be reviewed, tested and reverted on its own.Upstream
d821d9a390422179b6a5689cfdd33a77dd42b56egit cherry-pick -xonto masterfda24ca5; upstream authorship kept, and the commit message carries the(cherry picked from commit …)line.cmd/ingestor/{config,db,main}.go,config.example.json,cmd/server/readonly_invariant_test.go, newobserver_purge_test.go.Problem
RemoveStaleObserversonly soft-deletes (inactive = 1), so observer rows accumulate forever with no way to reclaim them.Change
New
retention.observerPurgeDays(default 0 = disabled). When set,PurgeStaleObservershard-deletes observers that are alreadyinactive = 1, not seen for N days, and not referenced by anyobservations,observer_metricsordropped_packetsrow. The reference guards matter becauseobservations.observer_idxhas no foreign key. It runs after the soft-delete at startup and on the observer retention ticker. The server's read-only invariant test now also asserts the server's*DBhas noPurgeStaleObservers.Adaptation to this fork
None. The cherry-pick applied without conflicts and the changed lines are identical to upstream.
Notes for review
Opt-in; nothing happens until an operator sets the value (it should be above both
observerDaysandpacketDays).Dependencies and merge order
fda24ca5and needs no other PR from this split.TestPruneOldNeighborMetricsdeterministic). If test(ingestor): make neighbor metrics pruning deterministic #33 lands first, the expected CI failure named below disappears; nothing in this PR depends on it.Verification
Local run of the same commands as CI's “Go Build & Test” job (server tests with
-race), on this branch and on masterfda24ca5under the same conditions (same machine, run one after another):fda24ca5go-ingestor-build-vetgo-ingestor-testTestBackfillTxLastSeen_ResolvesFromMaxObservationTimestamp,TestPruneOldNeighborMetricsgo-server-build-vetgo-server-test-racechannel-lib-testdecrypt-cli-build-testdockerfile-copy-invariantsdeclare -A), macOS has 3.2; identical on masterstaging-disk-monitorcss-vars-lintBaseline failures (fail identically on master; not introduced or changed here): see rows marked baseline failure, unchanged.
Browser validation (local, fixture DB, no staging/production): Not applicable (no frontend change).
Not run:
eslint(not installed locally; CI installs it on the fly).Expected GitHub CI: “Go Build & Test” is expected to fail on
TestPruneOldNeighborMetrics, which already fails on master (see #25's run). Downstream jobs (Playwright, image build) are therefore skipped. “Deploy Staging” and all GHCR publish steps only run onpushtomasterand cannot run for this PR.Two further ingestor tests have failed intermittently in this split's CI on branches whose
cmd/ingestortree is byte-identical to master (#27, #28), so they can also appear here without being caused by this change:TestBackfillTxLastSeen_ResolvesFromMaxObservationTimestamp: also reproduced locally on unmodified master.TestMQTTStallWatchdog_DisconnectedEscalationThrottled_1749: the suite flake that upstream test(ingestor): join the watchdog loop goroutine instead of only asking it to stop Kpa-clawbot/CoreScope#2003 (also split out of port(upstream): 26 clean upstream fixes — prune batching, /ws limits, observer liveness, watchdog race #25) addresses.GitHub CI result: run 34750395843 on
d33dd6fe. Go Build & Test: failure; all downstream jobs incl. Deploy Staging skipped. Failed tests:TestPruneOldNeighborMetrics: fails on master, documented baseline🤖 Generated with Claude Code