Repository navigation
port(upstream#2000): prune aged packets in bounded batches - #49
Conversation
…stalling ingest (Kpa-clawbot#2000) Reviewed at ecf0b37: query plans dumped and confirmed index-driven for all three statements, termination proven against concurrent ingest (first_seen is always time.Now()), FK child-first ordering required and correct, writer-stats assertions non-racy. Two low findings noted on the PR for follow-up: the dropped RowsAffected error now gates the loop, and ~0.53s batches will trip defaultSlowWriterMs=500. (cherry picked from commit fe37f10) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Uafhængig review + målt verifikation — VERDICT: ingen blockersGennemgået af en reviewer, der ikke skrev ændringen, med krav om at måle frem for at ræsonnere. Branchen er synkroniseret med master PR'ens kernepåstand holder. På en produktionsformet datamængde (16.000 forældede transmissions × 16 observationer = 256.000 observation-rækker — PR'ens egen "~260k observationer/dag"):
Miljø: darwin/arm64, Go 1.26.0, modernc.org/sqlite v1.34.5, Batchgrænsen er bevist, ikke antaget. På 2037 forældede rækker: Utilsigtet gevinst: chunking skærer peak Indeksplanen — påstanden VERIFICERET
Kommentarens "~10 µs til ~73 ms" er præcis. Revieweren rapporterede ærligt, at dens første forsøg modsagde kommentaren — det var dens eget harness: ved 200k rækker beholder SQLite det dækkende indeks for Differential mod implementeringen før PR'en14 datasæt: tom tabel, intet forældet, alt forældet på et eksakt multiplum af 250, 600 rækker med samme Desuden målt: de to statements i en batch materialiseres som SHOULD-FIX (dokumenteret, ikke rettet her — se begrundelse nederst)
Værste tilfælde: observation-sletningen er ikke bundet af 250
Lineært ved ~3,3–3,6 µs pr. slettet række, så hold ≈ Pre-eksisterende
|
Split out of #25 (commit
dcac010ethere). This branch holds exactly one upstream change so it can be reviewed, tested and reverted on its own.Upstream
fe37f1060cdf3f260943ad1cd15d3ed4ad176e48git cherry-pick -xonto masterfda24ca5; upstream authorship kept, and the commit message carries the(cherry picked from commit …)line.cmd/ingestor/maintenance.go, newprune_chunked_test.go.Problem
PruneOldPacketsdeleted a whole retention day inside a singleWriterTx.writerMuserialises every writer call, so MQTT ingest was blocked for the whole delete. Upstream measured about 35 s on an instance with ~260k observations/day.Change
Deletes in bounded batches of 250 transmissions (child observations first, same selection in both statements), committing and releasing the writer lock between batches. Batches are selected with
ORDER BY first_seen, id, whichidx_transmissions_first_seensatisfies, so the empty steady-state pass does not scan the table. A test pins that query plan. On error, the count already deleted by committed batches is returned.Adaptation to this fork
None. The cherry-pick applied without conflicts and the changed lines are identical to upstream.
Notes for review
idx_transmissions_first_seenexists in this fork's schema (internal/dbschema/dbschema.go,cmd/ingestor/db.go).Dependencies and merge order
fda24ca5and needs no other PR from this split.TestPruneOldNeighborMetricsdeterministic). If test(ingestor): make neighbor metrics pruning deterministic #33 lands first, the expected CI failure named below disappears; nothing in this PR depends on it.Verification
Local run of the same commands as CI's “Go Build & Test” job (server tests with
-race), on this branch and on masterfda24ca5under the same conditions (same machine, run one after another):fda24ca5go-ingestor-build-vetgo-ingestor-testTestPruneOldNeighborMetricschannel-lib-testdecrypt-cli-build-testdockerfile-copy-invariantsdeclare -A), macOS has 3.2; identical on masterstaging-disk-monitorcss-vars-lintBaseline failures (fail identically on master; not introduced or changed here): see rows marked baseline failure, unchanged.
Browser validation (local, fixture DB, no staging/production): Not applicable (no frontend change).
Not run:
eslint(not installed locally; CI installs it on the fly).Expected GitHub CI: “Go Build & Test” is expected to fail on
TestPruneOldNeighborMetrics, which already fails on master (see #25's run). Downstream jobs (Playwright, image build) are therefore skipped. “Deploy Staging” and all GHCR publish steps only run onpushtomasterand cannot run for this PR.Two further ingestor tests have failed intermittently in this split's CI on branches whose
cmd/ingestortree is byte-identical to master (#27, #28), so they can also appear here without being caused by this change:TestBackfillTxLastSeen_ResolvesFromMaxObservationTimestamp: also reproduced locally on unmodified master.TestMQTTStallWatchdog_DisconnectedEscalationThrottled_1749: the suite flake that upstream test(ingestor): join the watchdog loop goroutine instead of only asking it to stop Kpa-clawbot/CoreScope#2003 (also split out of port(upstream): 26 clean upstream fixes — prune batching, /ws limits, observer liveness, watchdog race #25) addresses.GitHub CI result: run 34752488840 on
76f42816. Go Build & Test: failure; all downstream jobs incl. Deploy Staging skipped. Failed tests:TestPruneOldNeighborMetrics: fails on master, documented baseline🤖 Generated with Claude Code