Skip to content

perf(channels): coalesce concurrent GetChannels/GetEncryptedChannels cache misses per region - #138

Merged
dborup merged 5 commits into
masterfrom
codex/issue-109-channels-singleflight
Sep 30, 2026
Merged

dborup merged 5 commits into
masterfrom
codex/issue-109-channels-singleflight

Conversation

@dborup

@dborup dborup commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Relates to #109

Plan and design

The user asked for autonomous work, so the plan is written here instead of waiting for sign-off (AGENTS.md rule 5). Scope: the issue plus the extension in its comment (keyed caches + per-key singleflight). As instructed, this PR does not add upstream's message-page cache or its composite index.

Commits:

  1. 4abd9caf: tests plus two test hooks in db.go (channelsMissHook, channelsQueryHook), which count real query executions. Red on master.
  2. 048db3c9: the change.
  3. a39b2fd5 (review round): a channel list whose rows stop on an error is returned as an error, not as a truncated success.
  4. 7f32845b (review round): stored cache keys are bounded in size.

Claims in the issue, verified against master

  • GetChannels has a single-slot 60 s cache and no coalescing.
  • GetEncryptedChannels has no cache at all.
  • Every request that misses runs the GROUP BY scan itself. Measured: 16 concurrent cold callers ran 16 queries.

Change: new cmd/server/channels_list_cache.go

channelListCache backs both calls.

  • Key: the normalized region (channelsRegionKey: upper-cased, trimmed, de-duplicated, sorted codes; ""/all → all regions). So " aar " and "AAR,aar" share one entry and one flight, and different regions never share a result.
  • Coalescing: per-key singleflight (golang.org/x/sync, already a dependency), with a second cache check inside the flight. A caller that missed before another flight stored its result finds that result instead of querying again.
  • Errors: only successful results are stored. An error goes to that flight's callers only, and the next call queries again. Since a39b2fd5, both queryChannels and queryEncryptedChannels check rows.Err() after the row loop, so a SQLite step that fails part-way is an error too; a truncated list is never returned or cached as success.
  • Bounds: TTL 60 s. At most 64 keys, because the region comes from a query parameter. When full, expired entries are evicted first, then the entry closest to expiry. Since 7f32845b, a key longer than 256 bytes is stored, and used as the flight key, as sha256:<hex> (71 bytes), so a long region parameter (up to the 1 MiB header limit) is no longer kept byte for byte. Long regions are still cached, coalesced and kept apart. The 64-key bound is unchanged.

Behaviour change: the encrypted list is now cached

  • GetEncryptedChannels had no cache before. It now has the same 60 s TTL as GetChannels, as the issue comment asks.
  • So the encrypted part of /api/channels?includeEncrypted=true can be up to 60 s old.
  • The two lists expire independently, so for up to 60 s after a channel changes state it can appear in neither list or in both. A similar window already existed on master, where only the decrypted list was cached.

Preserved fork behaviour

How this differs from upstream Kpa-clawbot/CoreScope#2059 / Kpa-clawbot/CoreScope#1936

Upstream is read as a reference only; nothing was cherry-picked.

  • The fork's single-slot cache became a bounded keyed cache; upstream's map is unbounded.
  • The encrypted list gets the same cache and flight, which upstream's path did not have.
  • No message-page cache: fork handlers mutate returned maps (annotateMessageAreas, annotateBotReplyTouchedAreas), so that would need a copy-on-read contract first (separate decision per the issue comment).
  • No composite index: perf(db): pin GetChannels and CountFloodAdvertsForNode to their indexes (issue #100) #107 deliberately pins the existing index.

Acceptance criteria

Criterion Status Evidence
Same normalized region: at most one real GetChannels query Met TestGetChannelsCoalescesConcurrentMisses_109 (16 callers → 1), TestGetChannelsCoalescesOnTheNormalizedRegion_109 (4 spellings → 1)
Same normalized region: at most one real GetEncryptedChannels query Met TestGetEncryptedChannelsCoalescesConcurrentMisses_109 (16 → 1, then cached within the TTL)
Different region keys never coalesced or mixed Met TestDifferentRegionsAreNeverMixed_109, TestChannelListCacheKeyLengthIsBounded_109 (long keys too)
Errors shared only within the flight, never cached Met TestChannelErrorsAreSharedButNotCached_109, TestChannelRowsErrorIsNotCachedAsSuccess_109 (both kinds)
Second cache check inside the singleflight winner Met TestSecondCacheCheckInsideTheFlight_109 (both kinds)
#107 index pinning and fallback intact Met Query code unchanged; existing channel tests pass, including channel_proposals_test.go (adapted to the keyed cache)
#98 cache-slice protection intact Met channels_cache_append_test.go (adapted to the keyed cache) passes
Deterministic tests count real executions, not timing Met channelsQueryHook counts each real query; callers are held on a barrier so they really overlap
-race and cold concurrent measurement before/after Met See Tests

Tests

Passed Failed
Commit A (tests + hooks, no fix) 0 6
048db3c9 (_109 tests, including the cache-bound and key-normalization tests) 8 0
Review round, before each fix 8 2 (TestChannelRowsErrorIsNotCachedAsSuccess_109 both subtests; TestChannelListCacheKeyLengthIsBounded_109)
Review round, 7f32845b 10 0

Review-round details:

  • TestChannelRowsErrorIsNotCachedAsSuccess_109 fails the iteration after 1 row through a new test seam, channelsRowsHook (nil in production, like the existing hooks). It asserts the error, a re-query on the next call and the full list.
  • Before the fix, it got 1 of 7 (channels) and 1 of 43 (encrypted) rows back with no error.
  • Mutants: removing either rows.Err() check fails the matching subtest. Truncating long keys instead of hashing them mixes two 1 MiB regions and is caught. Skipping the store key is caught.
  • cd cmd/server && go test -count=1 ./... on 7f32845b: ok (60.9 s).

go test -race -run "_109|Channel|channel" on 048db3c9: ok (387.7 s on the loaded local machine).

BenchmarkColdConcurrentGetChannels_109: 16 concurrent cold callers (8 GetChannels + 8 GetEncryptedChannels, region AAR, 20k-row fixture).

Queries per round Time per round
master 16 about 440 ms
this branch 2 about 54 ms

Full server -race run on 048db3c9

It reported FAIL in two tests, neither in code this PR touches:

  • TestPollerBroadcastsNewData (WebSocket poller: expected data.packet.timestamp to exist). In isolation under -race it failed 1 of 40 runs on origin/master d264716c and 0 of 40 on this branch. That is a pre-existing intermittent failure, shown on master.
  • TestIssue1008_HandlerReturns503WhileSubpathIndexLoading (status 200 vs 503, timing-dependent). It passes 30/30 in isolation on both master and this branch. Reproduced on origin/master d264716c: go test -race -count=200 -cpu 1,2,8 -run '^TestIssue1008_HandlerReturns503WhileSubpathIndexLoading$' failed 2 of 600 runs, twice, with the identical message (status = 200, want 503). It is a scheduling race in the test: the background subpath build on a tiny DB can finish before the handler call. This PR does not touch that build or the handler. The master full-suite run that passed (ok, 1307 s) simply did not hit it.

Not verified

  • Behaviour with production traffic and region mix, and on a production-size database.
  • Cache hit rate under real concurrency.
  • A real mid-iteration SQLite error: the test injects it through the rows seam.
  • The review's P3 on TTL expiry (no test drives an entry past its expiry) and its nit on the expired-first eviction pass are not addressed in this round.

Overlap with other open PRs

  • cmd/server/db.go is changed only by this PR.
  • The branch was cut from 85bfee49 and merges cleanly with origin/master. A merge simulation of all 13 batch branches in issue order merges this one without conflicts.

🤖 Generated with Claude Code

https://claude.ai/code/session_011FcyXW5RdFzLZhuL1ntAsY

dborup and others added 2 commits September 29, 2026 08:23
…s misses (#109)

Adds two test seams to DB, nil in production: channelsMissHook (after a
cache miss) and channelsQueryHook (right before each real query), and
channels_singleflight_109_test.go, which counts real queries on a seeded
fixture (pinTestDB/pinTestSeed, the #107 fixture).

On master all 6 fail: 16 concurrent cold GetChannels("AAR") run 16
queries, normalized region forms are not shared, GetEncryptedChannels
has no coalescing and no cache, a failing flight runs 8 queries for 8
callers, and a caller held past its miss queries again after another
caller has filled the cache. Results per region are already correct;
the tests pin that too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019TcZHooUiiknVWbECVWzk8
…normalized region (#109)

GetChannels had a single-slot 60s cache and no request coalescing;
GetEncryptedChannels had neither. Every request that missed ran the
GROUP BY scan itself.

Both now go through channelListCache (channels_list_cache.go):
- keyed by the normalized region (sorted, de-duplicated codes), so
  " aar " and "AAR,aar" share one entry and one flight, and different
  regions never share a result;
- per-key singleflight with a second cache check inside the flight;
- only successful results are stored: an error goes to that flight's
  callers and the next call queries again;
- bounded to 64 keys (the region is a query parameter): expired entries
  go first, then the one closest to expiry.

The query bodies are unchanged: #107's pinned-then-unpinned fallback is
the same code, now in queryChannels(key); #98's full-slice append in
handleChannels is untouched and the cached slices are never modified.
No new map[string]interface{} (the cached slice moved from a DB field
to channelListEntry in db.go; count unchanged at 66).

Benchmark, 16 concurrent cold callers (8 GetChannels, 8
GetEncryptedChannels, region AAR, 20k-row fixture): master 16 queries
and about 440 ms per round, this branch 2 queries and about 54 ms.

TestDifferentRegionsAreNeverMixed_109 now spreads both kinds over every
region (the first version paired each region with one kind only).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019TcZHooUiiknVWbECVWzk8
@dborup

dborup commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

Independent review of 048db3c9

Verdict: APPROVE with nits. This is a recommendation only; merging is the owner's call.

Reviewed head: 048db3c9dcf08ac639cfe83013ab4d4b6229245d (unchanged before and after the review). Work done on git archive trees of the head, of commit A 4abd9caf, of origin/master ad011021, and of the merge-tree result 20cfb5ca (head merged onto master).

Labels: [F] freshly verified by me · [T] taken from the PR text · [A] assumption · [K] known limitation.

Findings

  1. P3. TTL expiry is untested. cmd/server/channels_list_cache.go:56. Mutant M12 makes get ignore expires. Entries then never expire, and /api/channels would serve the first result per region until the entry is evicted. All 8 new tests and the related existing tests still pass [F]. No test drives an entry past its expiry, because load reads time.Now() directly. A 40-line probe (reviewer probe, kept locally, not committed) sets the stored entry's expires into the past and asserts a re-query. The probe passes on the head and fails on M12 (channels: expired entry served from cache (1 queries), want a re-query) [F]. Suggest adding it or an equivalent test.

  2. P3. The cache bound counts keys, not bytes. cmd/server/channels_list_cache.go:37, :43, :68-85. The key is the normalized region query parameter. A single code has no length cap, so one placeholder and a successful (empty) query mean the entry gets cached. Reviewer probe (kept locally, not committed): 64 distinct 1 MiB region strings × both kinds kept 128.0 MiB of key bytes (64+64 entries) for up to 60 s [F]. Go's default MaxHeaderBytes is 1 MiB and main.go:546 does not lower it, so the path is reachable over HTTP [F]. A front proxy with a smaller URL limit would reduce this [A]. On master the single slot held at most one key. Suggested fix: skip caching, or reject, when len(key) exceeds a small limit (for example 256 bytes, or more than N codes).

  3. P3. A mid-iteration SQL error is cached as a successful, truncated list. cmd/server/db.go:3159-3195 (queryChannels) and db.go:3289-3307 (queryEncryptedChannels). Neither checks rows.Err() after for rows.Next(), although db.go does so 25 times elsewhere [F]. If stepping fails part-way (for example a busy or IO error on the read-only connection), the partial slice is returned as success. It is then stored for 60 s and handed to every waiter of the flight. The issue says errors must not be cached as successful data. This was already true for GetChannels on master; it is new for GetEncryptedChannels, which had no cache before. Found by reading the code; not reproduced [A]. Fix: if err := rows.Err(); err != nil { return nil, err } in both functions.

  4. nit. The expired-first eviction pass is redundant. cmd/server/channels_list_cache.go:69-73. Mutant M7 removes it and survives. It is an equivalent mutant: the soonest-expiry pass that follows always picks an expired entry first when one exists, and get never serves expired entries. So the code comment's claim ("a fresh key replaces an expired one, not a live one") holds without the loop. The only difference is that M7 frees at most one expired entry per put instead of all of them. You can keep it for memory hygiene or drop it for simplicity.

Metadata

  • CI on 048db3c9 is complete: Go Build & Test, Playwright E2E and Docker are SUCCESS; the rest are SKIPPED [F].

  • Commits origin/master..048db3c9 [F]:

  • Files changed [F]: cmd/server/channels_list_cache.go (new, 121 lines), cmd/server/db.go (+85/−), cmd/server/channels_singleflight_109_test.go (new, 331 lines), cmd/server/channel_proposals_test.go (4 lines), cmd/server/channels_cache_append_test.go (14 lines). Both existing test files were adapted only to the new cache fields.

  • Merge base 85bfee49, not ad011021. Master's 4 newer commits do not touch the channel caches [F]. git merge-tree --write-tree origin/master 048db3c9 → 20cfb5ca, no conflicts [F]. On the merged tree, go vet is clean and the targeted tests pass [F].

  • .github/workflows/deploy.yml: not changed [F].

  • PR body [F]:

  • The PR is a draft [F].

Acceptance criteria (issue #109, including the upstream Kpa-clawbot#1936 comment)

Criterion Result
Concurrent misses for the same normalized region → at most one real GetChannels query Met [F]. 16 callers → 1 query; 4 spellings (AAR, aar, aar,AAR, Aar) → 1 query. M1 (no singleflight) and M5 (raw, un-normalized key) are both caught.
Same for GetEncryptedChannels Met [F]. 16 → 1, then served from the cache within the TTL. M1 is caught.
Different region keys never coalesced or mixed Met [F]. 4 regions × 2 kinds are compared with uncoalesced results. M4 (one global flight key) and M8 (encrypted uses the channels cache) are caught. Keys are separate per kind because each cache has its own singleflight.Group (db.go:101-102).
Errors shared only within the flight, not cached Met at the cache layer [F]. M3 (cache the error as empty success) is caught. See finding 3 for iteration errors that never reach this layer.
Second cache check inside the singleflight winner Met [F]. channels_list_cache.go:105-108. M2 (remove it) is caught by TestSecondCacheCheckInsideTheFlight_109, both kinds.
#107 index pinning and fallback intact Met [F]. The diff only moves the query body into queryChannels(key) and adds the hook; the pinned-then-unpinned SQL is unchanged. TestGetChannels_PinnedMatchesUnpinned, TestGetChannels_MissingIndexFallsBack and TestChannelsSQL_* pass.
#98 cached-slice protection intact Met [F]. routes.go:3419 and :3431 are unchanged. M10 (plain append(channels, encrypted...)) is caught by TestChannelsListDoesNotMutateCaches and TestChannelsCacheAppend_DBPath. The encrypted slice is now shared too; its only caller (routes.go:3412-3419) reads it as the append source and never writes into it [F].
Deterministic tests count real executions Met [F]. channelsQueryHook runs once per real query. On the head, the counts do not depend on timing (the 300 ms grace only makes callers overlap on commit A).
-race plus cold concurrent measurement before/after Met [F]. See the next sections.
Kpa-clawbot#1936 comment: keyed caches plus per-key singleflight for both lists Met [F]. The cache is keyed by normalized region and bounded at 64 keys (M6, unbounded, is caught). Bytes are not bounded; see finding 2.
Kpa-clawbot#1936 comment: do not port the message-page cache or the composite index Met [F]. There is no message cache and no schema or index change. The only db.go changes are in the channel list functions and the DB struct.

Test-first and mutants

Commit A (tests plus hooks only): all 6 tests present are red, for the intended reason [F]:

  • ran 16 queries, want 1
  • ran 12 queries, want 1
  • channels=16 encrypted=16, want 4 each
  • failing flight ran 8 queries
  • 2 queries: the held caller queried again

On the head, all 8 _109 tests are green [F]. Between A and the head, TestDifferentRegionsAreNeverMixed_109 changed its region index from i%len to (i/2)%len. The change is justified: with the old index, only 2 regions were exercised per kind, so the "one per region" count could not hold. Two tests were also added: bound and key normalization. A also adds the two hook fields to db.go. They are test seams, nil in production.

The mutant runs used -run '_109|ChannelsCacheAppend|ChannelsListDoesNotMutateCaches|GetChannels_|ChannelsSQL|ChannelsListIncludesApproved'. All files were restored afterwards, and shasum matches git show 048db3c9:<path> for channels_list_cache.go, db.go and routes.go [F].

# Mutant Result
M1 No singleflight (query right after the miss, still cache) caught (5 tests)
M2 Remove the second cache check inside the flight caught (SecondCacheCheck, both kinds)
M3 Cache an error as an empty success caught (ErrorsAreSharedButNotCached, both kinds)
M4 One constant flight key for all regions caught (DifferentRegionsAreNeverMixed)
M5 Raw region as the key (no normalization) caught (CoalescesOnTheNormalizedRegion)
M6 No size bound caught (ChannelListCacheIsBounded)
M7 No expired-first eviction pass survived: equivalent mutant (finding 4)
M8 Encrypted list uses the channels cache caught (DifferentRegionsAreNeverMixed)
M9 TTL = 0 (never cached) caught (EncryptedCoalesces, SecondCacheCheck, ChannelsCacheAppend_DBPath)
M10 #98 regression: append into the cached slice caught (ChannelsListDoesNotMutateCaches, ChannelsCacheAppend_DBPath)
M11 Key not sorted caught (ChannelsRegionKey)
M12 get ignores expiry survived: test gap (finding 1). Killed by the reviewer probe.

Suites run locally

  • cmd/server on the head, go test -race -count=1 -timeout 60m ./...: ok github.com/corescope/server 910.562s, exit 0 [F]. This ran on a loaded machine alongside other reviewers' suites. The two flaky tests named in the PR (TestPollerBroadcastsNewData, TestIssue1008_...) did not fail in this run, so their flakiness is [T] and I did not check it.
  • Test and benchmark counts from go test -list [F]:
    • master ad011021: 2041
    • head: 2039 (older base)
    • merged tree: 2050, which is master + 8 tests + 1 benchmark.
  • Targeted run on the head: go test -race -count=5 -run '_109|ChannelsCacheAppend|ChannelsListDoesNotMutateCaches|GetChannels_': ok (431.8 s) [F].
  • readonly_invariant_test.go (all 4 tests) and TestServerDBConnIsReadOnly: PASS [F].
  • go vet: clean on the head and on the merged tree [F].
  • cmd/ingestor: not touched and not run.

Performance and security

  • Perf proof reproduced [F]. BenchmarkColdConcurrentGetChannels_109 (16 cold callers, 8 + 8, region AAR, 20k-row fixture, -benchtime=20x -count=3). For commit A, I used a reviewer copy of the benchmark, adapted to A's cache fields.
    • Commit A: 480 / 437 / 430 ms per round, 16.00 queries/round.
    • Head: 56 / 58 / 69 ms per round, 2.000 queries/round.
    • This matches the PR's claim (≈440 ms → ≈54 ms).
  • Warm path [A]. It adds channelsRegionKey (split, sort and join over a handful of codes) and a mutex per call. That is negligible next to JSON encoding of the response. I did not benchmark it.
  • Cancellation. DB calls take no context, and Do (not DoChan) is used. So no caller's cancellation can abort other callers' results, and a slow query blocks its waiters exactly as long as it would have blocked each of them separately [F].
  • Panics. x/sync v0.10.0 singleflight re-panics in all waiters, and net/http recovers per request [A].
  • Boundedness.
    • Entries are capped at 64 per kind, with eviction in O(64) under the mutex [F]. Bytes are not bounded (finding 2).
    • The singleflight map holds only in-flight keys and deletes each on completion [F].
    • No goroutines or timers are added [F].
  • Aliasing. Stored entries are never modified after put. e.expires is set on a fresh entry before it is stored, and callers only read [F].
  • New map[string]interface{}. None added: db.go has 66 before and 66 after, total non-test cmd/server is 710 before and 710 after, and the new file has 0 [F].
  • Writes. No DB writes; the read-only invariant tests pass [F].
  • Behaviour change [K]. includeEncrypted=true now shows the encrypted list up to 60 s stale, as the issue comment asks. The two lists expire independently, so for up to 60 s a channel can appear in neither or both lists after it changes state [A]. A similar window already existed on master, because only the decrypted list was cached.

Not verified

dborup and others added 2 commits September 30, 2026 08:21
…109)

queryChannels and queryEncryptedChannels did not check rows.Err() after
the row loop. A SQLite step that fails part-way ends rows.Next() like the
last row does, so the truncated list was returned as success and, since
#109, cached for the TTL and shared with every waiter of the flight. Both
now return the error, which the cache already never stores.

Test: TestChannelRowsErrorIsNotCachedAsSuccess_109 fails the iteration
after one row (both kinds) through a new test seam, channelsRowsHook,
and asserts the error, a re-query on the next call and the full list.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011FcyXW5RdFzLZhuL1ntAsY
The cache key is the normalized region query parameter, whose length is
caller-controlled (up to the 1 MiB header limit). The 64-entry bound
alone let 64 such keys per list stay in memory for the TTL. A key over
256 bytes is now stored, and used as the flight key, as
"sha256:<hex>" (71 bytes). Long regions are still cached, coalesced and
kept apart; the 64-entry bound is unchanged.

Test: TestChannelListCacheKeyLengthIsBounded_109 (300 codes, two 1 MiB
codes differing in the last byte, and a short code; both lists).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011FcyXW5RdFzLZhuL1ntAsY

dborup commented Sep 30, 2026

Copy link
Copy Markdown
Owner Author

Review feedback addressed (commit 7f32845b)

  1. rows.Err() after the row loop (a39b2fd5):
    • queryChannels and queryEncryptedChannels now return the step error. A list cut short part-way is no longer returned or cached as success.
    • New test TestChannelRowsErrorIsNotCachedAsSuccess_109 fails the iteration after 1 row (both kinds) through a test seam, channelsRowsHook (nil in production). It asserts the error, a re-query on the next call and the full list afterwards.
    • Red before the fix: 1 of 7 rows (channels) and 1 of 43 (encrypted) came back with err <nil>. Green after.
    • Removing either check fails its subtest.
  2. Cache key length (7f32845b):
    • A key over 256 bytes is stored, and used as the flight key, as sha256:<hex> (71 bytes). Long regions stay cached, coalesced and separate. The 64-entry bound is unchanged.
    • New test TestChannelListCacheKeyLengthIsBounded_109 covers 300 codes, two 1 MiB codes that differ in the last byte, and AAR, on both lists.
    • Red before: keys of 1389, 1048576 and 1048577 bytes were kept. Green after.
    • A mutant that truncates instead of hashing mixes the two 1 MiB regions and is caught.
  3. PR text:
    • It now says that GetEncryptedChannels has a 60 s TTL, so the includeEncrypted=true list can be up to 60 s old.
    • It also covers the independent expiry of the two lists and the new fixes.

Test counts:

  • _109 tests: 8 passed / 2 failed before the fixes, 10/0 after.
  • The related channel tests (ChannelsCacheAppend, ChannelsListDoesNotMutateCaches, GetChannels_*, ChannelsSQL_*, ChannelsListIncludesApproved) are all green.
  • cd cmd/server && go test -count=1 ./...: ok (60.9 s).

Not addressed in this round: the TTL-expiry test gap and the expired-first eviction nit.


Generated by Claude Code

Only the cache and the flight use the bounded (sha256) key; the query must
still get the full normalized region. Without this, passing the digest to
the query would look up the region "sha256:..." and cache an empty list for
60 s, and no test caught it (review mutant G3).

Also correct the comment on why a digest never equals a short key: region
codes are upper-case, so none starts with the lower-case "sha256:".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dborup

dborup commented Sep 30, 2026

Copy link
Copy Markdown
Owner Author

Review feedback addressed (commit 390af21f)

  1. Re-review finding P3 (mutant G3: the query receives the hashed cache key): new test TestChannelListLongKeyQueriesFullRegion_109 asserts that both GetChannels and GetEncryptedChannels query the full normalized region for a key longer than channelListMaxKeyBytes. Red with G3 (both kinds report sha256:…), green on the head.
  2. Re-review nit: the comment on channelListMaxKeyBytes now gives the actual reason a digest can't equal a short key (region codes are upper-case, the digest prefix is lower-case sha256:); normalized keys can contain :.
  3. go test -race -count=5 -run _109 ok; full cmd/server go test -race -count=1 ./... ok (325.7 s); gofmt/vet clean on the touched files. No production code change.

@dborup
dborup marked this pull request as ready for review September 30, 2026 13:57
@dborup
dborup merged commit e829d1e into master Sep 30, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant