Skip to content

perf(nodes): cache region membership for /api/nodes?region= (#2101) - #2102

Closed
A13xB0 wants to merge 1 commit into
Kpa-clawbot:masterfrom
A13xB0:perf/nodes-region-cache-pr
Closed

A13xB0 wants to merge 1 commit into
Kpa-clawbot:masterfrom
A13xB0:perf/nodes-region-cache-pr

Conversation

@A13xB0

@A13xB0 A13xB0 commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Fixes #2101.

/api/nodes?region=X decided membership with an inline public_key IN (SELECT DISTINCT from_pubkey FROM transmissions ⋈ observations ⋈ observers …) subquery. It was uncached, unbounded in time, and evaluated twice per request (COUNT(*) and the page). fetchAllNodes() pages at 500, so a single Nodes view ran it several times, and one tab left open refreshing every minute was enough to occupy the reader pool.

On the ScotMesh instance (1.4 GB database, 1.47M observations, 2 vCPU, v3.12.0) each evaluation took about 10 s, a region-filtered page took 18–21 s, and corescope-server sat near 200% CPU. That matches what #2101 reports at larger scale.

What changes

Region membership (the set of public keys heard in a region set) is cached per canonical region set, together with an observations.id watermark. The canonical key is the sorted, de-duplicated codes from normalizeRegionCodes, so EDI,GLA, gla, edi and EDI,EDI,GLA share one entry. New file: cmd/server/nodes_region_cache.go.

  • Fresh for 30 s. After that the cached set is still served, and one background refresh (singleflight, the perf: /api/stats and /api/observers degrade to 10-17s under concurrent load (18k+ observers, ~4M observer_upserts/5min cycle) #1910 pattern) scans only observations past the watermark. That scan forces the join order with CROSS JOIN so it drives from the rowid range. Left to the planner, the sqlite_stat1 added in GetChannels/GetEncryptedChannels: cold query cost is driven by payload_type, not region (~10-13s solo) — EXPLAIN QUERY PLAN + numbers #2058 starts from the region's observers and walks their whole history.
  • Full rebuild every 30 min instead of a delta, so retention pruning, observer IATA changes and late from_pubkey backfills are not held forever. Full rebuilds are serialised across region sets, because each holds one of the four pooled connections for seconds and entries built together at startup fall due together.
  • Only the first request for a region set waits on a full scan, and concurrent first requests share it.
  • Failures and bounds. A failed refresh is logged and leaves the previous entry serving. Scans run under a context bounded by the rebuild interval. Entries are immutable once stored, and the cache holds at most 32 of them.
  • GetNodes binds the set through public_key IN (SELECT value FROM json_each(?)), so COUNT and SELECT no longer re-run the join.

Results are unchanged, except that a node newly heard in a region can take up to one refresh interval (about 30 s) to appear in that region's list. Region semantics are the same as before, so #1879 (attributing by home region) is neither addressed nor made harder: it would change what the cached set contains, not how it is cached.

Perf justification (rule 0).

  • A delta is a rowid range scan proportional to the observations ingested since the last refresh.
  • A full build costs the same as one evaluation of the old subquery, but runs at most once per region set per 30 min and off the request path, rather than twice per page request.
  • Per request, the work is now a json_each over at most a few thousand keys.
  • Memory per entry is one sorted key slice plus its JSON encoding: about 100 KB for the 1,390-node set seen on the soak instance.

Measurements

The same database, queried directly:

query time
old subquery, per evaluation (×2 per request) ~10 s
full build, EDI 9.4–14 s, once per 30 min
delta, last 300 observations 2 ms
delta, last 10k observations 22 ms

The ScotMesh production instance ran this algorithm for 28 h, probed every 5 minutes. That build was an earlier revision of this branch on a v3.12.0 base; see Validation for what changed after it.

v3.12.0 this branch
/api/nodes?limit=500&region=EDI 18–21 s p50 34 ms, max 215 ms (338 probes)
/api/nodes?limit=500 20–150 ms p50 26 ms, max 608 ms
/api/nodes, all traffic (/api/perf) p50 5.3 s over 24,372 requests p50 32 ms over 5,164 requests
corescope-server CPU, average ~31% of a core, pinned near 200% at peaks ~14% of a core
failed probes n/a 0 of 676

Memory was flat (1.00–1.03 GiB, under GOMEMLIMIT=1GiB).

Worth a decision

  • Full rebuild cost scales with the database. On the soak instance full rebuilds were the largest remaining cost of this path: 156 builds in 28 h, about 20 s each (three active region sets, two of them 35- and 55-code "everything ticked" selections). On a database the size of bug: Regional /api/nodes queries exhaust the SQLite connection pool #2101's that is proportionally more. nodeRegionRebuildInterval could reasonably be hours rather than 30 min, since pruned nodes lingering in a region's list matters little against multi-day retention. I left it at 30 min as the conservative choice and am happy to change it.
  • Cold start. Right after a restart, the first request for each region set waits on its full build, and those builds are serialised. On the soak instance the slowest first request after a deploy was 35 s, a 55-code set queued behind EDI. Every later request is served from the cache.
  • Customizer (rule 8). nodeRegionFreshTTL (30 s), nodeRegionRebuildInterval (30 min) and nodeRegionMaxEntries (32) are constants for now. The rebuild interval is the one worth exposing to operators later.
  • No request context on GetNodes. bug: Regional /api/nodes queries exhaust the SQLite connection pool #2101 also notes the queries ignore client disconnects. Membership is now off the request path and the remaining GetNodes queries are cheap, so I haven't threaded the request context through. The background scans have their own bounded context.

Validation

  • New tests: cmd/server/nodes_region_cache_test.go, 10 tests:
    • the old subquery is the oracle for result parity, on the v3 and v2 schemas, with lower-case, padded, duplicated and multi-region codes; each case that should match asserts the fixture actually produced matches, so an empty fixture cannot pass;
    • one scan per region set across COUNT, pages and equivalent spellings;
    • concurrent cold misses coalesce into one scan;
    • a delta picks up a newly heard node without a rebuild;
    • a delta with nothing new reuses the encoded keys and still advances the watermark;
    • a full rebuild drops pruned nodes;
    • a failed refresh keeps the previous entry serving;
    • the cache stays within its cap;
    • pooled connections are all returned after repeated refreshes.
  • Mutation checks. Each of these, applied to the implementation, makes the named test fail:
    • leaking the MAX(id) row: the connection test, with a clear message after 10 s;
    • never rebuilding: the pruning test;
    • removing the cap: the bounded test;
    • bypassing singleflight: 8 scans instead of 1.
  • The leak test guards a bug an earlier soak build had. That build issued both the ad hoc and the prepared MAX(id) query and scanned only one, leaking a pooled connection per refresh, and it hung every DB-backed endpoint within minutes. It was rolled back, fixed, and the 28 h soak above ran on the fixed build.
  • Server suite. go test -race ./... passes. One unrelated test is flaky under CPU load: TestPollerBroadcastsNewData failed 7 of 40 runs on unchanged master and 8 of 40 on this branch.
  • gofmt and go vet are clean.
  • Static analysis. I ran these on the changed code:
    • staticcheck -checks all: nothing;
    • SonarQube 26.9 (Go): 0 bugs, 0 vulnerabilities, 0 code smells, 0 security hotspots, 0% duplication, 88% coverage on the new file;
    • golangci-lint with a strict linter set: findings addressed except ones that conflict with this package's conventions (t.Parallel, varnamelen, noctx in tests, unchecked deferred Close) and a G201 false positive, since only ? placeholders and fixed fragments are formatted into the SQL.
  • Final revision on a live instance. This exact commit was redeployed to the ScotMesh instance:
    • /api/nodes?region= results compared against the old subquery on the live database: identical for EDI (320), EDI,GLA (321) and INV (235);
    • steady-state probes at 23–115 ms;
    • deltas logged in milliseconds (EDI: +1 nodes from obs 2424851..2425078 in 2ms).
  • Browser. Headless Chromium (Playwright) against that instance with the region filter set the way a visitor sets it:
    • the Nodes page renders 320 rows for EDI and 321 for EDI,GLA, matching the API, with no console or page errors;
    • warm page loads take about 5 s with or without a region (whole page), against 19 s for the first, cold EDI,GLA load;
    • the Map page does not send region to /api/nodes, so it is unaffected.
  • Rebase: on current master (093e320c), one commit.

🤖 Generated with Claude Code

…bot#2101)

`/api/nodes?region=X` filtered with an inline
`public_key IN (SELECT DISTINCT from_pubkey FROM transmissions ⋈ observations ⋈ observers ...)`
subquery. It was uncached and evaluated twice per request (COUNT(*) and the
page), and fetchAllNodes() pages at 500, so each Nodes or Map view ran it
several times. Measured on a 1.4 GB database (1.47M observations, 2 vCPU):
~10s per evaluation, 18-21s per region-filtered page on v3.12.0, and one open
tab refreshing every minute kept the reader pool and both cores busy.

Membership is cached per canonical region set (sorted, de-duplicated codes)
with an observations.id watermark:

- fresh for 30s; after that the cached set is served while one background
  refresh (singleflight, Kpa-clawbot#1910 pattern) scans only observations past the
  watermark, with the join order forced to drive from the rowid range;
- every 30 minutes the refresh is a full rebuild instead, so retention
  pruning, observer IATA changes and late from_pubkey backfills are not held
  forever. Full rebuilds are serialised across region sets, since each holds
  a pooled connection for seconds;
- only the first request for a region set waits on a full scan, and
  concurrent first requests share it;
- a failed refresh is logged and leaves the previous entry serving; scans
  run under a context bounded by the rebuild interval;
- entries are immutable once stored and the cache holds at most 32.

GetNodes binds the set through `public_key IN (SELECT value FROM json_each(?))`,
so COUNT and SELECT no longer re-run the join. Results match the previous
subquery (the tests use it as the oracle, on v2 and v3 schemas); the only
change is that a node newly heard in a region can take up to one refresh
interval to appear.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@A13xB0
A13xB0 marked this pull request as draft October 2, 2026 20:42
@A13xB0

A13xB0 commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

moved back to drafts due to it all rebuilding at the same time causing waves at 200% CPU for 4 minutes at a tie every 30 minutes

@A13xB0

A13xB0 commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

moving away from caching, leaving this here until I have a different solution

@mannkind

mannkind commented Oct 2, 2026

Copy link
Copy Markdown

I’ve been testing a smaller fix for this too.

Mine also stops doing work when the client disconnects. It might be good to cover that in your changes too so that requests nobody is waiting for don’t keep running? A shared refresh could still continue if other requests need it.

@A13xB0 A13xB0 closed this Oct 2, 2026
adminopenclaw8-sketch pushed a commit to dborup/CoreScope that referenced this pull request Oct 7, 2026
Review follow-ups on the Kpa-clawbot#2102 port.

1. Keep #38's from_pubkey contract inside the cache. The port had
   reintroduced COALESCE(t.from_pubkey, JSON_EXTRACT(decoded_json,
   '$.pubKey')), which #38 measured and rejected: it rescues no rows, and
   one corrupt ADVERT fails the whole query. Behind a cache that is worse
   than before, because the failed refresh leaves the entry stale for every
   later request too, not just the one that triggered it. #38's comment
   block moves to the query it documents.

2. Bound the cache by evicting one entry, not by clearing the map. The old
   setNodeRegionEntry replaced the whole map when a 33rd region set arrived,
   so a client cycling through region sets — ?region= is a query parameter —
   emptied the cache on every request and every following request paid a
   full observation scan, serialised behind nodeRegionFullMu. Entries now
   evict least-recently-used, and a region set too long to key is stored
   under its SHA-256 digest (channelListMaxKeyBytes's rule).

3. Make the documented freshness the delivered freshness. An entry older
   than nodeRegionMaxStale is no longer served blind: the caller waits for
   its refresh, which at that age is a full rebuild. Without the wait a
   refresh that keeps failing serves membership of unbounded age with only a
   log line to show it. A failing refresh still falls back to the stale
   entry rather than 500-ing /api/nodes?region=.

Tests: LRU eviction and the bound under 4x churn, the digest key, an
addition visible within nodeRegionFreshTTL, removal by retention and by an
observer IATA change gone within nodeRegionRebuildInterval and not by the
delta scan, the over-stale wait and its fallback, and the cached result set
against the uncached subquery for six region sets before and after new
ADVERTs. The port's legacy-backfill test is replaced by one that locks #38's
contract on the delta path. nodeRegionQueryHook now returns an error so a
test can fail a scan.

Perf (rule 0), nodes_region_cache_perf_176_test.go: 498MB / 1.5M
observations / 749k transmissions, ANALYZE run; one Nodes page load with a
region selected is 4 GetNodes pages of 500, 7 interleaved rounds of the same
test file compiled against master and against this branch. Warm page load
median 3729ms -> 31ms (121x; per GetNodes call 932ms -> 7.7ms, master's
932ms matching the 1-1.4s a region call costs on prod). Cold page load
median 4292ms -> 705ms (6.1x): the one full membership scan replaces eight
evaluations of the subquery.
dborup added a commit to dborup/CoreScope that referenced this pull request Oct 7, 2026
…ache

perf(nodes): port upstream region membership cache (Kpa-clawbot#2102)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug: Regional /api/nodes queries exhaust the SQLite connection pool

2 participants