Skip to content

test(anomaly): guard fixture search work both ways; make -bench . finish - #127

Merged
dborup merged 2 commits into
masterfrom
codex/anomaly-bench-hygiene
Sep 29, 2026
Merged

dborup merged 2 commits into
masterfrom
codex/anomaly-bench-hygiene

Conversation

@dborup

@dborup dborup commented Sep 28, 2026

Copy link
Copy Markdown
Owner

Summary

Benchmark hygiene in internal/anomaly, following the independent review of PR #108. Only *_test.go files change, and there are no production code changes. The library is still not wired into any binary.

  1. The work guard now covers the fixture rule and checks both ways. TestPeriodicSearchWorkIsBounded did not see the +41 % rise in visited gaps that fix(anomaly): periodic refinement, mixed-burst labels and censored-episode confidence #108 caused on the fixture rule's alternating shape (72.96 → 103.0 gaps per pulse). It also could not see a drop in work, which would mean lost search coverage.
  2. BenchmarkBytesPerStream never finished. It repeated a whole heap measurement b.N times with the timer stopped, so go test -run '^$' -bench . ./... did not complete. This predates fix(anomaly): periodic refinement, mixed-burst labels and censored-episode confidence #108.

1. Work guard (periodic_bench_test.go)

Before:

  • The guard ran only maxPeriodicRule() and worstPeriodicRule(). The rule behind BenchmarkPeriodicSearch's fixture/... cases (experimentalPeriodic()) was not covered.
  • measuredSteps was keyed by shape/MinCoverage. The fixture and max rules both have MinCoverage 0.6, so that key could not tell them apart.
  • The check was one-sided: it failed only above +25 %.

Now:

  • Rules: the guard runs fixture, max and worst. measuredSteps is keyed by rule/shape and holds master's actual totals (18a13264, 4 × 256 pulses per rule and shape).
  • Two-sided margin: it fails above +25 % (more work), and also below −25 %. The error for the lower bound says the search visits markedly fewer gaps, which means lost search coverage, and that a deliberate optimisation should update measuredSteps.
  • Comment: it now states that the counters count only visited gaps. Work per gap, such as the periodRange memo, is not visible to them. TestRangeMemoIsNeverStale covers the memo, and only BenchmarkPeriodicSearch measures the CPU.
rule/shape previous value (key) new measuredSteps (master 18a13264)
fixture/periodic not guarded 33,289
fixture/alternating not guarded 104,271
fixture/jitter not guarded 54,979
fixture/missing not guarded 34,850
fixture/bursty not guarded 82,025
fixture/noise not guarded 67,588
max/periodic 236,631 (periodic/0.6) 236,631
max/alternating 1,466,800 (alternating/0.6) 1,466,835
max/jitter 247,574 (jitter/0.6) 247,953
max/missing 243,414 (missing/0.6) 243,414
max/bursty 199,730 (bursty/0.6) 199,882
max/noise 148,177 (noise/0.6) 147,811
worst/periodic 236,631 (periodic/1) 236,631
worst/alternating 1,466,800 (alternating/1) 1,466,835
worst/jitter 6,242,089 (jitter/1) 6,434,483
worst/missing 1,174,373 (missing/1) 1,174,373
worst/bursty 199,730 (bursty/1) 199,882
worst/noise 148,177 (noise/1) 147,811

Mutant evidence (run on scratch copies of this branch, full package suite):

Mutant in periodic.go Work guard Other tests red
Seed tolerance reintroduced: the refinement range uses tol(seed) instead of periodRange red: fixture/alternating: 73761 steps in total, less than measured 104271 -25%: … lost search coverage … TestPeriodicRelativeJitterRefinesWithTheCandidateTolerance, TestRefinedCandidateReplacesOnlyWhenItWouldSignal
periodRange memo removed (memoRange always recomputes) green, as designed: visited gaps are unchanged TestRangeMemoIsNeverStale: 0 memo hits of 200000 queries

With the old one-sided guard, the first mutant would have passed the guard: its steps fall from 104,271 to 73,761.

2. BenchmarkBytesPerStream (bench_test.go)

Before: the whole measurement sat inside for it := 0; it < b.N; it++, with the timer stopped for the full iteration. The measurement is a new Detector, 4 × n events and two forced GCs. Because the timed work was about 0 ns, the framework kept raising b.N towards 10⁹ and ran that many full measurements.

Now:

  • Each sub-benchmark takes one heap measurement (bytesPerStream(n, periodic)) outside the b.N loop, with the timer stopped, and reuses it when the framework calls the sub-benchmark again with a larger b.N.
  • The b.N loop is empty, and B/stream is reported with b.ReportMetric as before.
  • ns/op is meaningless for this benchmark, and the comment says so.
  • b.Loop() is not used, because the module's go.mod targets Go 1.22.
Command, on unchanged master 18a13264 vs this branch before (master) after (this branch)
go test -run '^$' -bench . -count=1 ./... never finished: stopped by hand after 13 min 35 s, still in the first sub-benchmark BytesPerStream/streams=1000/periodic=false (-timeout 10m does not cover benchmarks) 63.5 s, 43 benchmark lines, exit 0
go test -run '^$' -bench . -benchtime=1x -count=1 ./... 4.0 s (5 s wall) 3.9 s (4 s wall)

Load during the runs: load average 3.4–3.9 on 12 cores.

B/stream is the same measurement as before, now taken exactly once per sub-benchmark:

streams periodic master (-benchtime=1x) this branch (full run)
1,000 false 760.7 760.7
1,000 true 1345 1345
10,000 false 747.9 747.9
10,000 true 1332 1332
100,000 false 739.0 739.0
100,000 true 1323 1323

Test results (local, go1.26.0 darwin/arm64, in internal/anomaly)

  • gofmt -l . lists nothing, and go vet ./... is clean.
  • go test -race -count=1 ./...: ok (12.2 s).
  • go test -run TestPeriodicSearchWorkIsBounded -count=3 -v ./...: 3/3 pass, and all 18 rule/shape totals are identical in all three runs (the counters are deterministic).
  • go test -run '^$' -bench . -benchtime=1x -count=1 ./...: completes in 3.9 s.
  • go test -run '^$' -bench . -count=1 ./...: completes in 63.5 s.

Scope

  • Changes only internal/anomaly/periodic_bench_test.go and internal/anomaly/bench_test.go, in one commit each.
  • No production code, golden, config, CI or deploy changes.

🤖 Generated with Claude Code

dborup and others added 2 commits September 28, 2026 17:23
Follow-up to PR #108, which raised the fixture rule's alternating shape
from 72.96 to 103.0 gaps visited per pulse (+41%) without any guard
seeing it: TestPeriodicSearchWorkIsBounded ran only the max and worst
rules, and only checked the upper side.

- The guard now also runs the fixture rule that BenchmarkPeriodicSearch
  uses for its "fixture/..." cases; measuredSteps is keyed by rule name
  and shape (the fixture and max rules share MinCoverage 0.6, so the old
  shape/MinCoverage key could not tell them apart).
- measuredSteps holds master's actual totals (e.g. worst/jitter
  6,242,089 -> 6,434,483 after #108).
- The 25% margin is two-sided. More work fails as before; markedly fewer
  gaps visited fails too, with a message that it means lost search
  coverage and that a deliberate optimization updates measuredSteps.
  Reintroducing the seed's tolerance in chainRefined now fails here
  (fixture/alternating 73,761 < 104,271 - 25%).
- The comment states that the counters count visited gaps only: removing
  the periodRange memo leaves them unchanged, and TestRangeMemoIsNeverStale
  covers the memo.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…h . finishes

BenchmarkBytesPerStream ran its whole measurement (a new Detector, 4*n
events, two forced GCs) inside the b.N loop with the timer stopped. The
timed work was ~0, so the framework kept raising b.N towards 1e9 and
repeated the measurement b.N times: `go test -run '^$' -bench . ./...`
never finished (on 18a1326, aborted after 13.5 min, still in the first
sub-benchmark; -timeout does not cover benchmarks). This predates #108.

Each sub-benchmark now takes one heap measurement (bytesPerStream)
outside the b.N loop, reuses it when the framework calls it again with a
larger b.N, runs an empty b.N loop and reports B/stream with
b.ReportMetric as before (ns/op is meaningless here). b.Loop is not used:
the module targets Go 1.22.

B/stream is unchanged (760.7, 1345, 747.9, 1332, 739.0, 1323 for 1k, 10k
and 100k streams without and with periodicity, the same as master with
-benchtime=1x); the full `-bench . -count=1` now completes in 63.5s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dborup
dborup marked this pull request as ready for review September 29, 2026 05:59
@dborup
dborup merged commit 97cebd9 into master Sep 29, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant