Repository navigation
port(upstream#1958): hash migration no longer reports success it never achieved - #42
adminopenclaw8-sketch wants to merge 2 commits into
Conversation
…ever achieved (Kpa-clawbot#1958) Part 2 of Kpa-clawbot#1856. **Part 1 is deliberately not fixed here** and the issue stays open for it; reasoning at the end. ## The bug `migrateContentHashesAsync` set `store.hashMigrationComplete` in a deferred func that ran unconditionally. Every DB failure inside the loop takes a `continue` (begin tx, prepare, commit), so the loop always reaches that defer, **including when not a single batch was written**. That is not hypothetical. The server has held a `mode=ro` handle since Kpa-clawbot#1283, so `Begin`, `Prepare` and `Commit` all fail, every batch is skipped, and `/api/stats` then answers `hashMigrationComplete: true` after migrating nothing. The migration is started unconditionally on every boot at `main.go:546`. ## The fix The three failure paths now count, and the defer only claims completion when the count is zero. When it is not, it logs once, naming the read-only handle as the expected cause and pointing at this issue, so an operator can tell "no work to do" apart from "could not do the work". **Nothing waits on the flag.** The only reader is `routes.go:828`, which reports it in `/api/stats`. Leaving it false on failure blocks nothing; it just stops the endpoint from lying. The in-memory index is untouched on failure. That was already true, because the index update runs only after a successful commit, and the test now asserts it so memory and disk cannot drift apart. ## Verification The regression test **fails on unmodified master**: ``` hash_migrate_test.go:115: hashMigrationComplete must stay false when no batch could be written; reporting true here is what Kpa-clawbot#1856 called self-reported success ``` It closes the DB handle to make writes fail. That is deterministic and exercises the identical path as a read-only handle (`Begin` errors, batch skipped); the in-memory test DB cannot be reopened read-only. The existing happy-path test still passes, so the flag still turns true on a real migration. `gofmt` clean, `go vet` clean, `cmd/server` suite ok in 59.7s. ## Why part 1 is not in here `handlePostPacket` writes to the same read-only handle and therefore always answers 500. I checked the error path before assuming it was misleading: it already returns `"transmission insert: attempt to write a readonly database"`, so the message is accurate. The endpoint is not confusing, it is simply dead. The issue asks maintainers directly: *"is this endpoint still wanted? If ingestion is MQTT-only now, deleting it is simpler than routing it through a handoff."* That is a product decision, not a fix, and inventing a middle answer would only add code without settling it. Worth noting the repository already has a precedent for the handoff shape: the server writes `request-<id>.json` and the ingestor consumes it (`cmd/ingestor/prune_geofilter.go`). Two things a decision should account for: the endpoint is documented in `openapi.go:69` and guarded by `requireAPIKey`, and `routes_test.go:4850` asserts it writes an observation row using the v3 schema, which passes only because the test DB is read-write. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit 56d6d4c) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Brings the branch up to master 834c8da so CI runs on current master. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Parked: this does not fix the case it targetsIndependent review found, and I then reproduced independently, that the premise in the code comment is not what SQLite does. The comment says that on the read-only handle the server has held since Kpa-clawbot#1283 "begin, prepare and commit all fail, every batch is skipped". Probing a real That is expected: the DSN sets no Consequence: all three new The test does not catch this because it closes the handle rather than opening it read-only. Two further things surfaced, both pre-existing but in the same function and relevant to any real fix:
For what it is worth: Why parked rather than patchedThe minimal correction (count
The branch is synced with current master and otherwise ready; it needs a decision on which of those to take. |
|
Closing as superseded: the content-hash migration now runs as a one-time ingestor migration (#222), so this port of upstream#1958 no longer applies. The branch is also ~690 commits behind and conflicts. |
Split out of #25 (commit
758aabebthere). This branch holds exactly one upstream change so it can be reviewed, tested and reverted on its own.Upstream
56d6d4c722da6bddf3b6b176f2207b309fd3a8e4git cherry-pick -xonto masterfda24ca5; upstream authorship kept, and the commit message carries the(cherry picked from commit …)line.cmd/server/hash_migrate.go,cmd/server/hash_migrate_test.go.Problem
migrateContentHashesAsyncsethashMigrationCompletein an unconditionaldefer. On the server's read-only DB handle (Kpa-clawbot#1283) every batch fails at begin/prepare/commit andcontinues, so the loop always reached the defer and/api/statsansweredhashMigrationComplete: truealthough nothing was migrated.Change
Counts failed batches (including a recovered panic). If any failed, the flag stays
falseand one[hash-migrate] INCOMPLETE …line is logged.Adaptation to this fork
None. The cherry-pick applied without conflicts and the changed lines are identical to upstream.
Notes for review
Visible change on this fork: our server opens SQLite
mode=ro, so after merge/api/statswill reporthashMigrationComplete: falseand the INCOMPLETE line appears at startup. That is the truthful value. In this tree the field is only exposed through/api/stats(routes.go); nothing underpublic/reads it.Dependencies and merge order
fda24ca5and needs no other PR from this split.TestPruneOldNeighborMetricsdeterministic). If test(ingestor): make neighbor metrics pruning deterministic #33 lands first, the expected CI failure named below disappears; nothing in this PR depends on it.Verification
Local run of the same commands as CI's “Go Build & Test” job (server tests with
-race), on this branch and on masterfda24ca5under the same conditions (same machine, run one after another):fda24ca5go-server-build-vetgo-server-test-racechannel-lib-testdecrypt-cli-build-testdockerfile-copy-invariantsdeclare -A), macOS has 3.2; identical on masterstaging-disk-monitorcss-vars-lintBaseline failures (fail identically on master; not introduced or changed here): see rows marked baseline failure, unchanged.
Browser validation (local, fixture DB, no staging/production): Not applicable (no frontend change).
Not run:
eslint(not installed locally; CI installs it on the fly).Expected GitHub CI: “Go Build & Test” is expected to fail on
TestPruneOldNeighborMetrics, which already fails on master (see #25's run). Downstream jobs (Playwright, image build) are therefore skipped. “Deploy Staging” and all GHCR publish steps only run onpushtomasterand cannot run for this PR.Two further ingestor tests have failed intermittently in this split's CI on branches whose
cmd/ingestortree is byte-identical to master (#27, #28), so they can also appear here without being caused by this change:TestBackfillTxLastSeen_ResolvesFromMaxObservationTimestamp: also reproduced locally on unmodified master.TestMQTTStallWatchdog_DisconnectedEscalationThrottled_1749: the suite flake that upstream test(ingestor): join the watchdog loop goroutine instead of only asking it to stop Kpa-clawbot/CoreScope#2003 (also split out of port(upstream): 26 clean upstream fixes — prune batching, /ws limits, observer liveness, watchdog race #25) addresses.GitHub CI result: run 34751729111 on
c272f6e3. Go Build & Test: failure; all downstream jobs incl. Deploy Staging skipped. Failed tests:TestPruneOldNeighborMetrics: fails on master, documented baseline🤖 Generated with Claude Code