Skip to content

test(server): replace distance lock timing threshold with a deterministic check - #79

Merged
dborup merged 5 commits into
masterfrom
codex/port-upstream-2039-lock-threshold
Sep 23, 2026
Merged

dborup merged 5 commits into
masterfrom
codex/port-upstream-2039-lock-threshold

Conversation

@adminopenclaw8-sketch

@adminopenclaw8-sketch adminopenclaw8-sketch commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This replaces the timing threshold in cmd/server/distance_lock_contention_test.go with a deterministic check of the lock invariant from upstream issue 1239. The PR started as a port of upstream PR 2039 (Kpa-clawbot/CoreScope), which raised the limit from 150 µs to 5000 µs. That port turned out not to catch the regression it was meant to guard against.

Only the test file changes. There are no production, workflow or configuration changes.

Why the threshold had to go

The old test timed bare s.mu.Lock()/Unlock() cycles while 8 goroutines ran computeAnalyticsDistance, and compared the average with a fixed limit. How long a cycle takes depends on core count, scheduler, -race and machine speed, so no single number separates healthy code from a regression on every machine.

Healthy runs measured 156–402 µs on CI runners (upstream readings), which is why 150 µs was flaky.

Measured at ee008091 on a 10-core Mac (Go 1.27) against the 5000 µs limit, with the regression being the pre-fix s.mu.RLock(); defer s.mu.RUnlock() held across the whole compute:

Build GOMAXPROCS Runs min / median / max Passed 5000 µs
healthy 10 / 2 / 4 / 8 30 each 0–65 µs 120/120
healthy -race 10 / 2 / 4 / 8 12 each 0–542 µs 48/48
regression 10 (default) 30 4493 / 4649 / 5924 µs 26/30 (missed)
regression 8 30 4736 / 4875 / 6928 µs 21/30 (missed)
regression 4 30 6649 / 6892 / 8069 µs 0/30 (1.3× margin)
regression 2 30 9882 / 10132 / 10758 µs 0/30
regression -race 10 / 2 / 4 / 8 12 each 20.3–53.7 ms 0/48

The regression passed the 5 ms limit on the fast default configuration. Healthy CI readings reach 402 µs, while the regression drops toward 4.5 ms as cores are added. That gap keeps shrinking on faster hardware, so any fixed limit would eventually either flake or miss.

New test method

TestComputeAnalyticsDistanceReleasesLockBeforeCompute checks lock state directly instead of measuring time:

  1. The test enters tx.decodedOnce.Do for one transmission and blocks inside it. With a region and an area set, the compute's tx.ParsedDecoded() call in the area filter must wait. That filter is the first step after s.mu.RUnlock().
  2. The test waits until the compute goroutine's runtime.Stack shows it parked there. The goroutine must be one this test created, inside computeAnalyticsDistance, ParsedDecoded and sync.(*Once).doSlow. At that point the compute is provably past its snapshot. This follows the existing pattern in get_store_stats_cache_test.go.
  3. s.mu.TryLock() must succeed. NewPacketStore starts no goroutines, so a failure can only mean the compute still holds the read lock.
  4. Still holding the write lock, the test assigns fresh distHops/distPaths and releases the park. The compute must then either return, or show up parked in [sync.RWMutex.RLock], which means it took the lock again. The result must come from the snapshot.

Nothing in the pass condition depends on time. The 10 s deadlines only turn a broken setup into a failure instead of a hang. The old measurement is kept as BenchmarkComputeAnalyticsDistanceWriterCycle, which is never asserted and does not run in CI.

Scope limitation: only the region+area path can be parked. The default path (region="", area="") shares the lock acquisition, so a change to that shared code is caught, but a lock added only on the default path is not. The test documents this.

Verification

All numbers are from the merge result 27b404fd, on macOS arm64 with 10 cores and Go 1.27. The mutants were applied mechanically to fresh worktrees at that commit.

Build Normal (GOMAXPROCS 10/1/2/4/8) -race (GOMAXPROCS 10/1/2/4/8) Failure reason
healthy 500/500 pass (100 each) 150/150 pass (30 each) —
M1: RLock + defer RUnlock across the whole compute (the pre-fix shape) 250/250 fail 100/100 fail TryLock invariant, every run
M2: RUnlock moved to just after the area filter 250/250 fail 100/100 fail TryLock invariant, every run
M4: early RUnlock kept, RLock taken again for sort/dedupe 250/250 fail 100/100 fail re-lock invariant, every run
M3: area filter skipped, so the park point is never reached 6/6 fail 2/2 fail seam timeout (the test cannot pass vacuously)

No mutant passed a single run. There were no data races, panics or retries, and no failures came from an unexpected cause.

Other checks:

  • cd cmd/server && go test -timeout 20m -count=1 ./...: ok (22.8 s)
  • cd cmd/server && go test -timeout 20m -race -count=1 ./...: ok (183.6 s), 0 DATA RACE warnings
  • go vet ./...: clean
  • gofmt -l on the changed file: clean
  • git diff --check: clean

Independent review

A separate reviewer agent that did not write the change reviewed it twice.

  • First round: no blockers, four non-blocking findings.
    • Its own mutant M4 escaped the first version. This led to the step-4 check, and M4 is now caught every time.
    • The scope doc overclaimed coverage; fixed.
    • The benchmark lacked a readiness barrier; fixed.
    • It found two pre-existing production issues, listed below.
  • Second round: no blockers.
    • It found no hidden hardware dependence; runs under CPU load, at GOMAXPROCS=1 and in no-inline builds all passed.
    • Mutants that take s.mu some other way (write Lock(), or RLock in a spawned goroutine) fail on the timeout and never pass.
    • Its two low-severity suggestions were applied in 9d817e40: match only this test's goroutine, and name the stuck case.

Out of scope: pre-existing production issues

Both are tracked for separate follow-up and are not changed here:

  • In the area-only path (region="", area set), computeAnalyticsDistance reads s.distHops/s.distPaths after RUnlock, not the snapshot. This is a data race under -race, and it can index past the end if ingest shrinks the slice.
  • updateDistanceIndexForTxs and eviction compact distHops/distPaths in place in the backing array a snapshot may still be reading. The "append-only" safety comment in computeAnalyticsDistance is therefore inaccurate. Step 4 of this test assigns fresh slices and does not model this.

Commits

  • 9b122b5e test(server): prove the lock release instead of timing writers
  • c717d6fb reviewfix(test): catch a lock re-taken later in the distance compute
  • 9d817e40 reviewfix(test): match only this test's compute goroutine, name the stuck case
  • 27b404fd regular merge of master (6334c427); no conflicts, and nothing under cmd/server changed on master

🤖 Generated with Claude Code

Openclaw and others added 5 commits September 23, 2026 07:02
…ng writers

The distance lock test gated on average writer Lock/Unlock latency, first
at 150µs, then 5000µs. Neither limit holds across machines. Healthy code
measures 156-402µs on CI runners. With the RLock deliberately held across
the compute, the 5000µs limit let the regression pass 4/30 runs at the
default GOMAXPROCS and 9/30 at GOMAXPROCS=8 on a 10-core Mac (4.5-6.9ms).

The test now checks the lock state directly. It parks computeAnalyticsDistance
on a tx.decodedOnce the test holds, in the area filter that runs right after
the snapshot, and confirms the park from the goroutine's stack. s.mu.TryLock()
must then succeed. Holding the write lock, the test swaps distHops/distPaths
as ingest would, then checks that the result comes from the snapshot. Nothing
is timed; the deadlines only turn a broken setup into a failure instead of a
hang.

The writer-latency measurement stays as an unasserted benchmark. There are
no production code changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The independent review found an escape. A compute that releases s.mu at
the snapshot but takes RLock again around the sort/dedupe passed the
TryLock check at the park point. The test now keeps the write lock while
it releases the park. From then on the compute either returns without
touching s.mu or blocks in RLock, which its goroutine header shows as
[sync.RWMutex.RLock]. That mutant now fails 250/250 normal and 100/100
under -race.

The review also flagged three smaller points:
- Doc scope. The comment now says only the region+area path is driven,
  and that step 4 assigns fresh slices rather than modelling in-place
  compaction.
- The benchmark now waits until every reader has finished one compute
  before its timer starts.
- The benchmark comment no longer claims to be the exact old measurement.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tuck case

These are the re-review's two low-severity suggestions. The stack matcher
now also requires that this test created the goroutine, so a compute
goroutine leaked by another test cannot trigger a false failure. The
step-4 timeout message now names the likely cause: s.mu taken some other
way, such as Lock().

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@adminopenclaw8-sketch adminopenclaw8-sketch changed the title test(server): calibrate distance lock threshold (port of upstream #2039) test(server): replace distance lock timing threshold with a deterministic check Sep 23, 2026
@dborup
dborup merged commit 4f59b99 into master Sep 23, 2026
11 of 12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants