Summary
The lazy distance-index build can crash the server with fatal error: sync: unlock of unlocked mutex. PacketStore.distLazyOnce (a sync.Once) is reassigned to a zero value while Do may be running on it. sync.Once.Do holds its internal mutex for the whole callback; overwriting the struct zeroes that mutex, so the deferred unlock at the end of the build fails fatally (not recoverable, process exits).
Found by the independent review of #145 (which re-runs analytics right after the startup load and so makes the window larger). A second review reproduced the sync.Once mechanism deterministically (fatal error: sync: unlock of unlocked mutex from sync.(*Once).doSlow) and found a lock-order deadlock in the same state handling (see below). Analysis against master 5f493f1d.
Where
cmd/server/store.go:
TriggerDistanceIndexBuild() (around lines 4658–4707): under distLazyMu it may reset s.distLazyOnce = sync.Once{} (around line 4680), then starts go s.distLazyOnce.Do(func() { ... }) (around line 4688). distLazyBuilding is only set inside the callback.
- The background chunk-load completion path (around lines 1677–1680) does
s.distLazyBuilt = false; s.distLazyOnce = sync.Once{} under distLazyMu, without checking distLazyBuilding.
Failure scenario
- A
/api/analytics/distance request calls TriggerDistanceIndexBuild(); the goroutine enters Do, which locks the Once's mutex and runs the build (buildDistanceIndex under s.mu, which can take a while on a large store).
- The background chunk load finishes and resets
s.distLazyOnce = sync.Once{} while step 1 is still inside Do.
- The build finishes;
Do's deferred m.Unlock() runs on the zeroed mutex → fatal error: sync: unlock of unlocked mutex, server exits.
Holding distLazyMu around the reset does not help: Do's own mutex is what gets overwritten. It is also a data race on the Once's fields (done, m), which -race should report.
Proposed fix
Stop resetting a sync.Once. The state already has distLazyBuilding/distLazyBuilt under distLazyMu, which is enough:
TriggerDistanceIndexBuild: under distLazyMu, return if distLazyBuilding; apply the existing debounce if distLazyBuilt; otherwise set distLazyBuilding = true before starting the goroutine, and run the build without a sync.Once.
- Chunk-load path: set
distLazyBuilt = false (so the next request rebuilds) and bump a data generation; never touch a Once. The build records the generation it actually read and rebuilds again only if the dataset changed after that snapshot (see the lock-order section).
- Remove the
distLazyOnce field.
Acceptance criteria
- Reproduce first: a test that holds the build open via the existing
distanceBuildHook seam, triggers the chunk-load reset path meanwhile, and shows the failure (a -race report or the fatal error in a subprocess) on master.
- See the additional criteria in the lock-order section below.
- With the fix, the same test passes under
-race -count=20; no sync.Once is reassigned anywhere in cmd/server (grep guard test is fine).
- Existing behaviour stays: at most one build at a time, debounce (Δobs > 5 % or 5 min), 202 while building, rebuild after the background load.
- Full
cmd/server suite under -race.
Lock-order deadlock (same code, separate bug)
Removing the sync.Once alone does not fix this:
TriggerDistanceIndexBuild() takes distLazyMu, then s.mu.RLock() to read totalObs (debounce path, when a build has completed before).
- The background-load completion path holds
s.mu.Lock() and then takes distLazyMu (around lines 1669–1681).
Opposite lock order: a distance request holding distLazyMu waits for s.mu.RLock() while the loader holds s.mu and waits for distLazyMu → both block forever.
Added requirements
- Never acquire
s.mu (Lock or RLock) while holding distLazyMu. Read totalObs before taking distLazyMu (or keep it atomic), or restructure so the locks are never nested in opposite orders.
- Prefer a generation/dirty counter over a plain
rebuildRequested flag: record the generation the build actually read; rebuild again only if the dataset changed after that snapshot. This avoids a redundant expensive build when the loader finishes before the queued build gets s.mu.
Added acceptance criteria
- Deterministic test of the
sync.Once crash (subprocess) or a real -race reproduction on master.
- Deterministic test of the lock-order deadlock on master, passing with the fix.
- A guard that no code path takes
s.mu while holding distLazyMu.
- Many concurrent triggers start exactly one build.
- Loader finishes before the build acquires
s.mu: no extra build.
- Loader changes the dataset after the build's snapshot: exactly one new build.
202 + Retry-After, debounce and post-load invalidation unchanged; full cmd/server suite under -race.
Order
Fix this issue first, then sync #145 with it (or land it as a clearly separate first commit in the same series).
Summary
The lazy distance-index build can crash the server with
fatal error: sync: unlock of unlocked mutex.PacketStore.distLazyOnce(async.Once) is reassigned to a zero value whileDomay be running on it.sync.Once.Doholds its internal mutex for the whole callback; overwriting the struct zeroes that mutex, so the deferred unlock at the end of the build fails fatally (not recoverable, process exits).Found by the independent review of #145 (which re-runs analytics right after the startup load and so makes the window larger). A second review reproduced the
sync.Oncemechanism deterministically (fatal error: sync: unlock of unlocked mutexfromsync.(*Once).doSlow) and found a lock-order deadlock in the same state handling (see below). Analysis against master5f493f1d.Where
cmd/server/store.go:TriggerDistanceIndexBuild()(around lines 4658–4707): underdistLazyMuit may resets.distLazyOnce = sync.Once{}(around line 4680), then startsgo s.distLazyOnce.Do(func() { ... })(around line 4688).distLazyBuildingis only set inside the callback.s.distLazyBuilt = false; s.distLazyOnce = sync.Once{}underdistLazyMu, without checkingdistLazyBuilding.Failure scenario
/api/analytics/distancerequest callsTriggerDistanceIndexBuild(); the goroutine entersDo, which locks the Once's mutex and runs the build (buildDistanceIndexunders.mu, which can take a while on a large store).s.distLazyOnce = sync.Once{}while step 1 is still insideDo.Do's deferredm.Unlock()runs on the zeroed mutex →fatal error: sync: unlock of unlocked mutex, server exits.Holding
distLazyMuaround the reset does not help:Do's own mutex is what gets overwritten. It is also a data race on the Once's fields (done,m), which-raceshould report.Proposed fix
Stop resetting a
sync.Once. The state already hasdistLazyBuilding/distLazyBuiltunderdistLazyMu, which is enough:TriggerDistanceIndexBuild: underdistLazyMu, return ifdistLazyBuilding; apply the existing debounce ifdistLazyBuilt; otherwise setdistLazyBuilding = truebefore starting the goroutine, and run the build without async.Once.distLazyBuilt = false(so the next request rebuilds) and bump a data generation; never touch a Once. The build records the generation it actually read and rebuilds again only if the dataset changed after that snapshot (see the lock-order section).distLazyOncefield.Acceptance criteria
distanceBuildHookseam, triggers the chunk-load reset path meanwhile, and shows the failure (a-racereport or the fatal error in a subprocess) on master.-race -count=20; nosync.Onceis reassigned anywhere incmd/server(grep guard test is fine).cmd/serversuite under-race.Lock-order deadlock (same code, separate bug)
Removing the
sync.Oncealone does not fix this:TriggerDistanceIndexBuild()takesdistLazyMu, thens.mu.RLock()to readtotalObs(debounce path, when a build has completed before).s.mu.Lock()and then takesdistLazyMu(around lines 1669–1681).Opposite lock order: a distance request holding
distLazyMuwaits fors.mu.RLock()while the loader holdss.muand waits fordistLazyMu→ both block forever.Added requirements
s.mu(Lock or RLock) while holdingdistLazyMu. ReadtotalObsbefore takingdistLazyMu(or keep it atomic), or restructure so the locks are never nested in opposite orders.rebuildRequestedflag: record the generation the build actually read; rebuild again only if the dataset changed after that snapshot. This avoids a redundant expensive build when the loader finishes before the queued build getss.mu.Added acceptance criteria
sync.Oncecrash (subprocess) or a real-racereproduction on master.s.muwhile holdingdistLazyMu.s.mu: no extra build.202 + Retry-After, debounce and post-load invalidation unchanged; fullcmd/serversuite under-race.Order
Fix this issue first, then sync #145 with it (or land it as a clearly separate first commit in the same series).