Repository navigation
fix(server): bound resolved-path LRU entry age for /paths confirmation (#277) - #316
Conversation
…RU entry (#277) The ingestor's observation upsert can replace a stored resolved_path in place. /paths and /hop_analytics read the canonical path through apiResolvedPathLRU, which is never invalidated, so they keep attributing the tx to a node its stored path no longer contains. Adds a clock hook on the LRU (unused here) so the test can move time forward. Red on this commit; the next commit fixes it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
/paths, /hop_analytics and the packets page read the canonical resolved_path through apiResolvedPathLRU. The ingestor's upsert can replace a stored path in place (same row id), and the poll loop only reads new ids, so it cannot invalidate the entry; detecting the change would need a new ingestor change log and another query per poll tick. Bound the entry age instead: an entry older than resolvedPathLRUTTL (60 s) is a miss and is refreshed in place by the next read, keeping its FIFO slot. An entry stored after the lookup's clock (wall clock stepped back) is treated as expired. Nothing changes on the poll path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…th fetch (#277) Since #246 a hash-index candidate costs a canonical-path fetch and an LRU entry even when its stored path does not contain the target. Measured on the CI-prepared fixture and a 20x copy, no candidate is rejected after that fetch: prefix collisions, the usual non-member candidates, are dropped by the index check first. Rejected candidates therefore stay cached; this test pins the order that keeps them rare, for /paths and /hop_analytics. Numbers are in the PR description. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- TestPathLenFast_MatchesReference / _RandomisedAgainstReference become TestPathLen_EquivalenceGuard_Corpus / _Randomised. Comparing pathLen with the reference also passes without a fast path, so both now check pathLenFast directly and require it to serve the well-formed inputs. - TestNodePaths_StaleIndexEntryStillExcluded becomes ..._StaleIndexEntryExcludedWithoutSQLConfirm and also asserts the tx was decided from its canonical path (the pre-#246 pre-filter dropped it unread), so it fails on the old code without relying on the counter. - TestNodePaths_NoCanonicalPathStillConfirmedBySQL becomes ..._NoCanonicalPathConfirmedBySQLExactlyOnce; its comment says the query count is the only observable difference from the old code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…equest (#277) Giving LRU entries an age added a time.Now() to every cache hit. On a kvm-clock VM that is ~65 ns against a 15 ns hit, and /paths does one hit per candidate (up to ~4k per request on a 20x fixture copy). - The LRU clock is monotonic (time.Since of a package epoch): one clock read, and wall-clock steps cannot stretch or cut an entry's life, so the "stored in the future" guard is gone. An entry stored after a request's clock read is fresh. - fetchResolvedPathForTxBestAt / fetchResolvedPathForObsAt take the clock from the caller; /paths and /hop_analytics read it once per request. 1000 warm hits per request: 15.29 us on master, 15.27 us here (benchstat, n=10, p=0.74). A single-lookup call (packet detail, health) pays one clock read, ~40 ns. Tests pin one clock read per warm request and a minimum useful entry life. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rapport — CS-pve-agent2 PR#316 #277 — head 34e68c1Status: All three points addressed. Every requirement has a test and caught mutants. CI is green on all jobs with no reruns. The PR is still a draft; nothing merged, readied or closed. Evidence tags: [T] = test, benchmark or command run for this PR; [A] = code reading or analysis; [K] = taken from the issue, #246 or its review, not re-checked. Point 1 — stale resolved-path LRU
Choice and justification.
Poll-path cost [T] (benchstat, n=10, interleaved, 20× fixture copy). Code is unchanged; measured anyway:
Found while measuring. The first version read Point 2 — caching rejected candidates
Numbers per call:
Point 3 — test clarity
Perf:
|
| master | branch | ||
|---|---|---|---|
/paths |
152.9 ms | 154.5 ms | ~ (p=0.63) |
/hop_analytics |
72.2 ms | 73.7 ms | ~ (p=0.44) |
HTTP, sequential curl, all 204 nodes, two runs each, on the CI-prepared fixture and the 20× copy: cold /paths, warm /paths and /hop_analytics are equal within run-to-run noise. The full table is in the PR description [T].
The one intended cost is the first /paths per node after 62 s on the 20× copy: 1.79–2.06 ms mean on the branch vs 1.28–1.34 ms on master. That is a cold-equivalent re-read once per TTL [T].
Rules
- No new
map[string]interface{}(0 added outside tests) [T]. - No frontend files changed, so no hardcoded colours.
scripts/check-xss-sinks.sh --diff origin/masterreports nothing to scan [T]. - Fork guards: 9 in
deploy.ymland 1 inrelease-fast-path.yml. No.github/change [T]. - 5 commits, each authored and committed by
dborup <kontakt@meshview.dk>. No rebase, amend or force-push [T].
Local test runs
| Run | Result |
|---|---|
cmd/server go test -count=1 ./... |
ok (490 s) [T] |
Touched tests, -race -count=3 |
ok [T] |
cmd/ingestor go test ./... |
ok (1,161 s) with CI's -timeout 20m. The first run with Go's default 10 m timeout timed out; the package needs more than 10 min here [T] |
sh test-all.sh |
222 passed, 0 failed [T] |
node test-frontend-helpers.js |
709 passed, 0 failed [T] |
E2E against a local branch server on a CI-prepared e2e-fixture.db (stopped by pid) [T]:
| Suite | Result |
|---|---|
test-e2e-playwright.js |
132/135, 3 fixture skips |
test-issue-1146-path-link-contrast-e2e.js |
11/11 |
test-issue-1281-location-row-e2e.js |
6/6 |
test-issue-1206-resize-observer-leak-e2e.js |
28/28 |
test-path-inspector-e2e.js |
6/6 |
CI (run 37449209012, head 34e68c18)
| Job | Result |
|---|---|
| Go Build & Test | pass (23m36s) |
| Playwright E2E Tests | pass (23m42s) |
| Build & Publish Docker Image | pass (56s) |
| Release Artifacts / Deploy / Publish Badges | skipped (not a push to the main repo) |
No reruns. The known flaky test (#271) did not fail [T].
Remaining
- Pre-existing flake, not fixed here. Several
/pathshandler tests callstore.Load()withoutWaitIndexesReady, so they race the background path-hop index build and sometimes get503 index loading. Seen in 1 of 8 grouped local runs on the branch (…_HopName_CanonicalPathShowsTarget_1144) and in 2 of 25 onorigin/master(…_SortByRecency_1145,…_SortCountTiebreaker_1145) [T]. Worth its own issue. - Hash-index staleness is out of scope. A rewrite that adds a node to a stored path still only reaches the in-memory hash index on restart (documented in
resolved_path_backfill.go) [A]. - Dead fallback cache (follow-up). In
fetchResolvedPathForTxBest, the fallback caches the result under the sibling's obs id, but the next call looks up the longest obs id first. So the multi-row fallback query repeats on every call for those txs [A]. Not measured; separate from Follow-ups to #246: stale resolved-path LRU in /paths confirmation, rejected-candidate caching, test clarity #277. - Not measured: production-sized data, a multi-GB DB file, and concurrent
/pathsload. All timings are from the fixture and a synthetic 20× copy on one VM [A].
Review — CS-pve-agent1 PR#316 — head 34e68c1Dom: APPROVE med nits. F1 is a one-line test fix and I recommend applying it before merge. Evidence tags: [T] = test, benchmark or command I ran myself; [A] = my own code reading or analysis; [K] = taken from the PR or the author's report, not re-checked. Trees used: head Findings
Answers to the review points1. The choice (bounded entry age vs invalidation).
2. Repro test red on master, green on head.
With the F1 assertion added, B and B2 are caught and the unmutated code still passes [T]. 3. Clock and races.
4. Issue points 2 and 3.
5. Perf.
Rules
Tests I ran
E2E against a local Go server built from merged
The PR has no E2E of its own, since there is no frontend change. The four suites above are the ones that hit CI, run 37449209012 on Not verified
|
…#277) F1: the default branch of lruNow() was only checked for monotonicity, so a clock in ms or µs (TTL stretched to days/hours) passed every test. Assert it advances by nanoseconds across a 1 ms sleep. F2: the lruClock hook returns wall UnixNano while the default clock is a small offset from lruEpoch, so installing the hook while entries exist expires them all. Red until lruNow puts both in the same domain. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… clock (#277) The hook stands in for time.Now, but lruNow returned its UnixNano (about 1.8e18) while the default clock returns nanoseconds since lruEpoch. Installing the hook while entries existed therefore expired them all. lruNow now returns lruClock().Sub(lruEpoch), the same domain as time.Since(lruEpoch). The production branch is unchanged. The field comment no longer says nil = time.Now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rapport — CS-pve-agent3 PR#316 runde 2 — head 1fdff0fReview feedback addressed (commit
No other behaviour change. F3–F6 were out of scope for this round, as instructed. Evidence tags: [T] = test or command I ran for this round; [A] = my own code reading; [K] = taken from the review or the earlier report, not re-checked. Commits this round
Author and committer are Mutants (all run by me, on the merged tree, against the 12 touched tests)
The review showed that B and B2 survived the previous version of the test [K]. I ran B and B2 against the new test before writing the F2 fix: both fail [T]. The 12 touched tests pass with Local test runs (merged tree, head
|
| Run | Result |
|---|---|
cmd/server go test -count=1 ./... |
1st run: FAIL, one test, TestHandleAnalyticsSubpathsWithStore (expected 200, got 503). 2nd run: ok (608 s) [T] |
cmd/ingestor go test -count=1 -timeout 30m ./... |
ok (1,303 s) [T] |
sh test-all.sh |
225 passed, 0 failed [T] |
node test-frontend-helpers.js |
709 passed, 0 failed [T] |
About the one failure. It is the pre-existing 503 index loading race: the test calls store.Load() without WaitIndexesReady and races the background subpath index build. It is the same class as "Remaining 1" in the earlier report and F6 in the review. It is not #271 or #301. Neither the test nor the subpath code is touched by this PR. Repeated alone (-count=2000, run sequentially), it fails 4/2000 on the branch and 2/2000 on origin/master b0b9843c. A small sample beforehand (-count=300) gave 2/300 on the branch and 0/300 on master. So it reproduces on master and is not caused by this PR [T]. Still worth its own issue.
E2E against a local Go server built from 1fdff0f6, on a copy of e2e-fixture.db prepared as in CI: freshen-fixture.sh, the Kpa-clawbot#1486/Kpa-clawbot#1791 seed SQL taken from deploy.yml, corescope-migrate, then seeds 2073, 199 and 245. The server was stopped by pid [T]:
| Suite | Result |
|---|---|
test-e2e-playwright.js |
132/135, 3 fixture skips |
test-issue-1146-path-link-contrast-e2e.js |
11/11 |
test-issue-1281-location-row-e2e.js |
6/6 |
test-issue-1206-resize-observer-leak-e2e.js |
28/28 |
test-path-inspector-e2e.js |
6/6 |
Rules
- No new
map[string]interface{}(0 in the round-2 diff) [T]. - No frontend changes.
scripts/check-xss-sinks.sh --diff origin/masterreports nothing to scan (exit 0) [T]. - No
.github/change. Fork guards unchanged: 9 indeploy.yml, 1 inrelease-fast-path.yml[T]. gofmt: the three touched files are clean [T].
CI (run 37478466487, head 1fdff0f6, attempt 1)
| Job | Result |
|---|---|
| Go Build & Test | pass (23m50s) |
| Playwright E2E Tests | pass (24m23s) |
| Build & Publish Docker Image | pass (53s) |
| Release Artifacts / Deploy Staging / Publish Badges & Summary | skipped (not a push to the main repo) |
No reruns. The known flaky tests #271 and #301 did not fail [T].
Not done / not verified
- F3 (worst-case staleness wording), F4, F5 (probe not checked in) and F6 were out of scope for round 2 and are unchanged.
- The pre-existing
503 index loadingtest race is documented, not fixed. - The PR stays a draft. Nothing merged, readied or closed. No staging or production access was used.
Review — CS-pve-agent1 PR#316 — head 1fdff0fDom: APPROVE med nits. F1 and F2 from round 1 are fixed and verified. The remaining nits are documentation or report wording and don't block merge. Evidence tags: [T] = test, mutant or command I ran myself; [A] = my own code reading or analysis; [K] = taken from the PR or the author's report, not re-checked. Trees used:
The PR's own merge commit Findings
Answers to the review points1. F1: the production clock's unit is now pinned.
2. F2: the comment is correct, and late hook installation doesn't age entries wrongly.
3. Production behaviour: only what F1/F2 require.
4. CI per job (run 37478466487 on
No Rules
Tests I ran
E2E against a local Go server built from the merged tree, on
The server was stopped by port (
These are the suites that hit Not verified
|
Relates to #277
Follow-ups from the review of #246. All three points are addressed; only point 1 changes server behaviour.
Plan
/paths//hop_analytics#246tests pass trivially / only via query countsNo config or customizer implications. The TTL is an internal cache bound, not a user-facing value. No frontend change.
1. Stale LRU: bounded entry age
Why not invalidate in the poll loop. The ingestor's observation upsert (
ON CONFLICT … resolved_path = COALESCE(excluded.resolved_path, resolved_path)) rewrites the row in place and keeps its id. The server's poll loop only reads new ids (IngestNewObservations:WHERE o.id > ?), so no existing poll query ever sees the change. Invalidating there would need a new ingestor-side change log (asroute_mask_changesdoes for #89), a schema change, and at least one more query per poll tick. That is more than the issue's "no extra SQL per poll" bar allows.What this does instead.
apiResolvedPathLRUentry records when it was stored. An entry older thanresolvedPathLRUTTL(60 s) is a miss. The next read refreshes it in place and keeps its FIFO slot, solruOrdergets no duplicates.time.Sinceof a package epoch). Wall-clock steps can therefore neither extend nor cut an entry's life./pathsand/hop_analyticsread the clock once per request (fetchResolvedPathForTxBestAt). A cache hit with a per-hittime.Now()measured 80 ns against 15 ns on this kvm-clock VM, and/pathsdoes one hit per candidate.Trade-off. For up to 60 s after a rewrite, the old path can still be served. Each requested entry costs one primary-key read per 60 s (measured below as the "after 62 s" pass).
Not covered. A rewrite that adds the target to a path does not reach the in-memory hash index until restart. That is a pre-existing limitation, documented in
cmd/ingestor/resolved_path_backfill.go. This change only stops the endpoints from answering from an old path.2. Rejected candidates: measured, kept
Method: load the store, then for every node clear the LRU, call the endpoint, and classify every LRU entry the call created. An entry is either accepted, rejected by the canonical path, or rejected where the old SQL check would also have rejected it. That last class is the cost #246 added.
resolved_pathrows rewritten after load (synthetic stale index)/hop_analyticsgives identical counts.Prefix collisions are the common reason a candidate does not belong to the node. They are dropped by the hash-index check before any fetch. Only 64-bit FNV collisions and stale index entries reach the fetch. Even with an unrealistically high 5 % rewrite rate, that is at most 13 entries per call in a 10,000-entry LRU (0.13 %). Not caching them would make every repeat call re-read them. So caching stays, and
TestNodePaths_PrefixCollisionsExcludedBeforeCanonicalFetchpins the ordering that keeps them rare.3. Test clarity
TestPathLenFast_MatchesReference/_RandomisedAgainstReferenceare renamed toTestPathLen_EquivalenceGuard_Corpus/_Randomised. Both now comparepathLenFastdirectly and require it to serve inputs: the corpus's well-formed entries, and at least one non-empty random array. A build without a fast path now fails both.TestNodePaths_StaleIndexEntryStillExcludedis renamed to…_StaleIndexEntryExcludedWithoutSQLConfirm. It also asserts that the tx was decided from its canonical path. The pre-perf(server): remove per-candidate SQL and JSON parsing from /paths and /hop_analytics #246 SQL pre-filter dropped it unread, so the test now fails on the old code without relying on the query counter.TestNodePaths_NoCanonicalPathStillConfirmedBySQLis renamed to…_NoCanonicalPathConfirmedBySQLExactlyOnce. The comment states that the query count is the only observable difference from the old code (the response is identical by design).Commits
test(server): reproduce the stale LRU. Red on its own: both endpoints keep the tx an hour after its stored path stopped containing the target.fix(server): 60 s entry age.test(server): guard for point 2.test(server): point 3 renames and assertions.perf(server): one LRU clock read per request, and a monotonic clock.Perf proof
Master
c713e32avs this branch, same machine (4 vCPU VM, kvm-clock), same DB copies. Each server ran alone on its own copy.In-process, warm, all 204 nodes per op (benchstat, n=10, interleaved, 20× copy):
/paths/hop_analyticsfetchResolvedPathForObshit (packet detail, health)HTTP, sequential curl, all 204 nodes (mean / p95 ms; two runs each, run 2 in brackets):
/pathscold/pathswarm ×3/hop_analytics×3/pathscold/pathswarm ×3/hop_analytics×3/paths62 s laterThe last row is the price of the fix. After the TTL, the first call per node re-reads its entries and costs about as much as a cold call. Calls within the window are unchanged.
Complexity. O(1) extra per cache lookup (one subtraction). One clock read per
/paths//hop_analyticsrequest. No extra SQL on the poll path. The LRU stays bounded at 10,000 entries; each entry is 8 bytes larger.Checks
cmd/serverstays read-only. No new SQL;readonly_invariant_test.gopasses.map[string]interface{}.scripts/check-xss-sinks.sh --diff origin/masteris clean..github/change. Fork guards: 9 indeploy.yml, 1 inrelease-fast-path.yml.🤖 Generated with Claude Code