Skip to content

fix(anomaly): periodic refinement, mixed-burst labels and censored-episode confidence - #108

Merged
dborup merged 5 commits into
masterfrom
codex/anomaly-periodic-followup
Sep 28, 2026
Merged

dborup merged 5 commits into
masterfrom
codex/anomaly-periodic-followup

Conversation

@dborup

@dborup dborup commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Summary

This PR fixes all four findings from the review of the anomaly library merged in PR #94. That library is not yet wired into production: no server or ingestor code imports it, and CI only vets and tests it as its own module. This PR adds no runtime integration, schema change or deploy, and every change is inside internal/anomaly/.

# Finding Cause Fix Reproduction (red on the commit before the fix)
1 P2, narrow range: a valid period is missed when [MinPeriod, MaxPeriod] is narrower than the jitter estimate() dropped every seed g/k outside the range before chainRefined could refine it Out-of-range seeds are kept while seedInReach holds, but used only for their refinement; in-range seeds are always measured TestPeriodicNarrowRangeFindsPeriodBehindOutOfRangeSeeds (absolute and relative)
2 P3, mixed bursts: a periodic burst with both expected and unclassified events is labelled unclassified The chain label counted only all-expected pulses A chain with any expected event, not all expected, is mixed; monitored still wins TestPeriodicBurstTrafficLabel
3 P2, relative jitter: with JitterRel > 0 a valid period is never found chainRefined built the common range with the seed's tolerance, not each candidate period's periodRange returns exactly {P : |g − k·P| ≤ tol(P)}, and the refined period is clamped into common range ∩ [MinPeriod, MaxPeriod] TestPeriodicRelativeJitterRefinesWithTheCandidateTolerance
4 P3, lost ConfidenceLow: a censored new-stream episode opens as low, but later transitions come out route_observed Confidence was lowered only from the current trigger's evidence; the episode did not carry it The episode records censored evidence, like its label and coverage, and every transition of that episode is low: quiet end, eviction end and reopen within cooldown TestCensoredNewStreamKeepsLowConfidenceOn{QuietEnd,EvictionEnd,Reopen}

Commits

  1. 827cf50e test: reproduces findings 1 and 2.
  2. 74f86e81 fix: findings 1 and 2.
  3. 9899b9f6 reviewfix: always measure in-range seeds, and prove seedInReach exhaustively.
  4. 0a30e8b4 test only: reproduces findings 3 and 4. All four new tests fail on 9899b9f6 at their intended assertions; the controls pass.
  5. cfbc7c84 fix: findings 3 and 4, the clamp, the memo, the golden update and the train-1240 expectation.

Files: periodic.go, detector.go and state.go, plus tests in periodic_followup_test.go, confidence_followup_test.go, review_fixes_test.go, chance_test.go and periodic_bench_test.go, and testdata/replay_golden.json.

Details of findings 1 and 2 (from the earlier commits)

Finding 1. seedInReach(g, k) keeps a seed while k·MinPeriod − tol(MinPeriod) ≤ g ≤ k·MaxPeriod + tol(MaxPeriod).

  • This is a necessary condition for some in-range P to explain g as k periods with its own tolerance, because k·P ± tol(P) never decreases in P.
  • It is proven exhaustively over small integers, and by sampling at nanosecond scale.
  • Seeds per pulse are still at most recentGaps·(MaxMissing+1), so searchCells (the Chance union bound) and periodicWorkBound still hold.

Finding 2. The chain is evaluated when its newest pulse starts, which is causal, so that pulse holds only its first event. The repro cases therefore do not rely on the newest pulse.

Details of finding 3

periodRange(g, k) returns the exact integer range of periods P that explain gap g as k periods, each with its own tolerance tol(P) = max(JitterAbs, ⌊JitterRel·P⌋).

  • JitterRel = 0: the two integer divisions of the old code.
    • TestPeriodRangeWithAbsoluteJitterIsTheOldRange compares them with the old formula as an oracle: exhaustive for g ≤ 2000, k ≤ 9 and JitterAbs ≤ 40, plus 200,000 samples at nanosecond scale.
    • Two off-by-one mutants of the absolute branch turn that test red.
  • JitterRel > 0: floating-point estimates of g/(k+JitterRel) and g/(k−JitterRel) are stepped to the exact ends.
    • This relies on k·P ± tol(P) never decreasing in P.
    • TestPeriodRangeIsExactlyThePeriodsThatExplainTheGap checks it by brute force over small integers for absolute, relative and combined jitter, and checks the defining ends on 200,000 samples at nanosecond scale.
  • Union bound: there is still at most one refined period per seed. The searchCells argument, that the highest lower end L_j of the common range explains every other gap, carries over with L_j as the lowest period explaining gap j. Chance and searchCells are unchanged.

Clamp into the search range (approved). With exact ranges, a common range can straddle MaxPeriod. The phase estimate (the mean gap) then lies beyond it and used to be rejected, although periods in range explain every gap. Without the clamp, 15 of 10,000 narrow-range edge trains that 9899b9f6 found were lost (McNemar p = 6e-5). With the clamp none are lost anywhere, and under absolute jitter it only adds detections at the range ends: narrow-abs6 edge trains go from 9,991 to 10,000. TestPeriodicRefinementIsClampedIntoTheSearchRange covers both jitter types, with fixture guards that the true period meets every criterion and that the mean gap exceeds MaxPeriod.

Memo. periodRange results are cached in the scratch that the Detector owns.

  • Slot: one per (pulse mod recentGaps, k), so at most 144 entries, allocated once.
  • Staleness: an entry is used only when g, JitterAbs and JitterRel all match, so it is never stale.
  • Tests: TestRangeMemoIsNeverStale makes 200,000 interleaved queries across rules, gaps, multiples and shared slots.
  • Equivalence: on 560,000 probe trains the outputs (signal counts, on-period signals and first pulse per train) are byte-identical to the version without the memo.

Details of finding 4

  • Where it is recorded: episode.note now also records whether the hot signal came from censored NewStreamEvidence. move() copies it onto each transition, and emit() lowers the confidence of every transition of such an episode, whether or not it has a trigger.
  • Scope: a new episode, after the cooldown, starts without it. Rate and Periodic rules are unaffected, because their confidence is a function of the key alone.
  • Tests: there is one test for each path, and each path has its own killing mutant. The controls check that an uncensored stream is never lowered and that low does not leak into the next episode.

Golden and changed expectations (approved)

  • testdata/replay_golden.json: exactly one field changes. The quiet end of censored episode 7771b168 goes from route_observed to low; that episode opens as low in the same golden. The new sha256 is 3d6fbf08….
    • TestGoldenChangedOnlyInTheDocumentedChanceValues keeps its pinned hash unchanged. It now asserts that this one quiet end is low and maps it back before hashing, so any other change to the golden still fails.
    • Both the original golden and a golden with one extra change fail the guard.
  • TestRefinedCandidateReplacesOnlyWhenItWouldSignal, train 1240: the expectation changes from 57/1/57 to 57/2/69, with the reason in the test comment.
    • The Monte Carlo train, broken by a 187.5 s gap at 58, is found again at gap 69 as 172.48 s over 12 pulses: residual 2.561 s ≤ 2.587 s, Chance 1.1e-11 ≤ 1e-9.
    • The seed-tolerance range only reached 171.75 s, whose residual of 2.84 s fails MaxJitterFraction.
  • measuredSteps: not updated. The total gaps visited change by +3.1 % at most (jitter/1: 6,242,089 → 6,434,483), well inside the 25 % guard. The peak per pulse is 114 chains and 11,765 steps, against a bound of 302 and 79,314.

Mutation evidence

All 22 mutants are killed, each on a scratch copy of the final tree with the full package test suite.

Mutant Killed by (among others)
N1: old seed tolerance in refinement …RelativeJitterRefinesWithTheCandidateTolerance, the train-1240 test
N2 and N3: absolute branch of periodRange off by one (lo, hi) TestPeriodRangeWithAbsoluteJitterIsTheOldRange, and 7 more
N4: relative branch dropped (only absolute fixed) relative-jitter repro, …IsExactlyThePeriodsThatExplainTheGap, clamp test, and 4 more
N5: no exact stepping (floating-point estimates only) TestPeriodRangeIsExactlyThePeriodsThatExplainTheGap
N6: ConfidenceLow lost on quiet end …OnQuietEnd, …OnReopen, golden replay
N7: ConfidenceLow lost on eviction end …OnEvictionEnd
N8: ConfidenceLow lost on reopen …OnReopen
N9: note never records censoring all three P3 tests and the golden replay
N10: emit ignores the episode flag (master behaviour) all three P3 tests and the golden replay
N11: low leaks into the next episode TestNewStreamConfidenceIsNotLoweredWithoutCensoring
N12: no clamp into the search range TestPeriodicRefinementIsClampedIntoTheSearchRange
N13–N15: memo key ignores JitterRel, JitterAbs, or the gap TestRangeMemoIsNeverStale
M1–M7: the mutants from the earlier review as before

False alarms and detection

Paired probe: identical trains fed to master b3e44761, 9899b9f6 and this PR, with 10,000 trains of 200 pulses per cell, each train using a fixed seed. A train counts as a false alarm if it signals at least once, and as a detection if it signals within tol of its true period.

Configs: fixture (the test rule: JitterAbs 5 s, JitterRel 2 %), rel2 (JitterRel 2 %), rel3abs1-tight (3 % + 1 s, MaxChance 1e-9), rel25 (JitterRel 0.25, the maximum), narrow-rel2 and narrow-abs6 (range [299 s, 301 s]), abs6, rel2-miss8 (MaxMissing 8) and rel2-loose (MaxChance 1).

False alarms (null trains; Poisson with a mean gap of 300 s, and uniform gaps in [30 s, 570 s])

config null master 9899b9f this PR this PR − 9899b9f (paired: +new/−gone, Δ [95% CI]) McNemar p
fixture poisson 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
fixture uniform 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
rel2 poisson 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
rel2 uniform 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
rel3abs1-tight poisson 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
rel3abs1-tight uniform 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
rel25 poisson 5 5 5 +0/−0, +0.00 pp [+0.00, +0.00] 1
rel25 uniform 255 255 267 +17/−5, +0.12 pp [+0.03, +0.21] 0.017
narrow-rel2 poisson 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
narrow-rel2 uniform 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
abs6 poisson 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
abs6 uniform 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
rel2-miss8 poisson 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
rel2-miss8 uniform 1 1 1 +0/−0, +0.00 pp [+0.00, +0.00] 1
rel2-loose poisson 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
rel2-loose uniform 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
narrow-abs6 poisson 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1
narrow-abs6 uniform 0 0 0 +0/−0, +0.00 pp [+0.00, +0.00] 1

Master and 9899b9f6 are identical in every null cell.

Confirmed on 100,000 further independent trains for rel25:

  • Uniform: 2.451 % → 2.582 %, +162/−31 paired trains, +0.131 pp [95 % CI +0.104, +0.158], McNemar p = 1.3e-22.
  • Attribution: the whole increase comes from periodRange. A variant without the clamp gives identical results.
  • Poisson: 55 vs 54 of 100,000, no change.
  • Realistic settings (fixture, rel2, rel2-loose, rel3abs1-tight), uniform null, 100,000 trains each: 1 vs 1 and 0 vs 0 (95 % upper bound 0.003 %).

This is accepted as a documented limitation: see "Known limitations".

Detection (no train detected before is lost)

trains config master 9899b9f this PR lost / gained vs master lost / gained vs 9899b9f
jittered fixture 10000 10000 10000 0 / 0 0 / 0
jittered rel2 10000 10000 10000 0 / 0 0 / 0
jittered rel3abs1-tight 10000 10000 10000 0 / 0 0 / 0
jittered rel25 10000 10000 10000 0 / 0 0 / 0
jittered narrow-rel2 10000 10000 10000 0 / 0 0 / 0
jittered abs6 10000 10000 10000 0 / 0 0 / 0
jittered rel2-miss8 10000 10000 10000 0 / 0 0 / 0
jittered rel2-loose 10000 10000 10000 0 / 0 0 / 0
jittered narrow-abs6 10000 10000 10000 0 / 0 0 / 0
edge fixture 9902 9902 9902 0 / 0 0 / 0
edge rel2 10000 10000 10000 0 / 0 0 / 0
edge rel3abs1-tight 6940 6940 6940 0 / 0 0 / 0
edge rel25 3850 3850 3851 0 / 1 0 / 1
edge narrow-rel2 0 10000 10000 0 / 10000 0 / 0
edge abs6 10000 10000 10000 0 / 0 0 / 0
edge rel2-miss8 10000 10000 10000 0 / 0 0 / 0
edge rel2-loose 10000 10000 10000 0 / 0 0 / 0
edge narrow-abs6 0 9991 10000 0 / 10000 0 / 9

"Jittered" trains have residuals up to ±0.95·tol, 10 % missing pulses and 2 % stray pulses. "Edge" trains have residuals of 0.90–0.99·tol, with two of three gaps long, as in the finding 3 reproduction. Master detects no narrow-range edge train at all; that is finding 1.

Performance

BenchmarkPeriodicSearch was run interleaved: prebuilt test binaries for master, 9899b9f6 and this PR, 3 rounds × 2 runs each (n = 6 per version), with the machine's load average between 3.0 and 4.4 on 12 cores. The table shows medians.

config/shape master ns/op 9899b9f ns/op this PR ns/op Δ vs 9899b9f Δ vs master steps/pulse master / 9899b9f / PR allocs/op master / 9899b9f / PR
fixture/periodic 289 293 295 +0.7% +2.1% 32.95 / 32.95 / 32.95 0 / 0 / 0
fixture/alternating 517 544 736 +35.4% +42.5% 72.96 / 72.96 / 103.0 0 / 0 / 0
fixture/jitter 754 759 771 +1.5% +2.2% 40.46 / 40.46 / 40.41 0 / 0 / 0
fixture/missing 303 310 314 +1.4% +3.5% 34.95 / 34.96 / 34.96 0 / 0 / 0
fixture/bursty 1272 1343 1583 +17.9% +24.4% 83.28 / 83.77 / 83.69 0 / 0 / 0
fixture/noise 1101 1180 1364 +15.5% +23.8% 69.39 / 70.35 / 69.80 0 / 0 / 0
max/periodic 1399 1422 1426 +0.3% +2.0% 255.2 / 255.2 / 255.2 0 / 0 / 0
max/alternating 5428 5604 6096 +8.8% +12.3% 1524 / 1555 / 1556 0 / 0 / 0
max/jitter 7460 7504 7515 +0.1% +0.7% 259.0 / 259.2 / 259.4 0 / 0 / 0
max/missing 1444 1468 1466 -0.2% +1.5% 260.8 / 260.8 / 260.8 1 / 1 / 1
max/bursty 3424 3626 4100 +13.1% +19.7% 202.6 / 205.9 / 206.2 0 / 0 / 0
max/noise 2590 2772 3152 +13.7% +21.7% 151.8 / 154.9 / 154.7 0 / 0 / 0
max-4rules/periodic 5596 5684 5730 +0.8% +2.4% 255.3 / 255.3 / 255.4 3 / 3 / 3
max-4rules/alternating 21828 22470 24398 +8.6% +11.8% 1525 / 1555 / 1557 0 / 0 / 0
max-4rules/jitter 18860 18943 18974 +0.2% +0.6% 259.0 / 259.2 / 259.4 3 / 3 / 3
max-4rules/missing 5758 5832 5861 +0.5% +1.8% 261.0 / 261.0 / 260.9 4 / 4 / 4
max-4rules/bursty 13033 13823 15698 +13.6% +20.5% 202.6 / 205.9 / 206.2 0 / 0 / 0
max-4rules/noise 9756 10495 11950 +13.9% +22.5% 151.8 / 154.9 / 154.7 0 / 0 / 0
worst/periodic 1402 1422 1426 +0.3% +1.8% 255.2 / 255.3 / 255.2 0 / 0 / 0
worst/alternating 5424 5616 6082 +8.3% +12.1% 1524 / 1555 / 1556 0 / 0 / 0
worst/jitter 103131 109180 115340 +5.6% +11.8% 7013 / 7580 / 7769 0 / 0 / 0
worst/missing 5540 5608 5828 +3.9% +5.2% 1266 / 1266 / 1266 0 / 0 / 0
worst/bursty 3432 3622 4120 +13.8% +20.1% 202.6 / 206.0 / 206.2 0 / 0 / 0
worst/noise 2584 2776 3168 +14.1% +22.6% 151.8 / 154.9 / 154.7 0 / 0 / 0

n = 6 per version (3 interleaved rounds x 2). Delta vs 9899b9f: -0.2% .. +35.4%; vs master: +0.6% .. +42.5%.

Costs:

  • Shapes with relative jitter that exercise refinement: fixture/alternating is +35 %, and 41 % more gaps are visited per pulse (73 → 103), because the exact ranges let the range walk go further. Bursty and noise shapes are about +13–18 %.
  • Periodic, jitter and missing shapes: within ±6 %.
  • Allocations: none new.
  • Memo effect: without it the overhead was up to +78 % (bursty/noise +40–72 %). The memo roughly halves that.
  • Absolute scale: the largest per-pulse cost for the fixture rule is about 1.6 µs.

Test results

All of the following were run locally on the final tree cfbc7c84:

  • In internal/anomaly: gofmt -l . is empty, go vet ./... is clean, and go test -count=1 ./... and go test -race -count=1 ./... pass.
  • TestIndependent* run with -count=10: 60/60 pass.
  • git diff --check is clean.
  • The golden sha256 is 3d6fbf08….
  • TestOfflineReplay skips, as on master: it needs explicit replay flags.

Review

  • Earlier commits: an internal reviewer agent approved 74f86e81 and 9899b9f6 in a pre-review, with no blockers.
  • Commits 4 and 5: these have not been reviewed yet. An independent review of the whole PR follows.

Known limitations

  • More chance periodicities in very regular traffic at the widest tolerance. At JitterRel = 0.25 with uniform, more-regular-than-Poisson gaps, false alarms rise by 0.131 pp (2.451 % → 2.582 %).
    • Chance is a union bound under a Poisson null. The searchCells comment already notes it is not a bound for traffic more regular than Poisson.
    • The Poisson rate is unchanged, and realistic JitterRel of 2–3 % shows 0 vs 0.
  • Clamp at the range ends. The refined period can now be exactly MinPeriod or MaxPeriod when the train's own estimate lies just beyond the range, as long as an in-range period explains every gap. A steady train whose gaps are all outside the range, such as constant 294 s gaps with range [299 s, 301 s], is still not reported, because the seed's own chain already explains it (TestPeriodicOutOfRangeSeedIsNeverReported).
  • seedInReach is necessary, not sufficient. Some admitted out-of-range seeds cannot be refined. They cost work but stay within the documented bound.
  • Label timing. The evaluation is causal: a periodic chain is labelled with the events that exist when its newest pulse starts.
  • Scope of the evidence. The numbers come from synthetic benchmarks and probes. No production traffic was replayed, since the library is not integrated.

🤖 Generated with Claude Code

dborup and others added 5 commits September 28, 2026 14:31
Two deterministic reproductions for the periodic detector in the merged
(not yet integrated) internal/anomaly library. Both fail on master:

- A 5m train with gaps alternating 294s/306s, MinPeriod=299s,
  MaxPeriod=301s, tolerance 6s: 300s explains every gap and meets every
  criterion (asserted as a fixture guard), but every seed g/k lies outside
  the range and is dropped before chainRefined can refine it into 300s.
- A periodic chain whose bursts merge an expected and an unclassified
  event is labelled unclassified instead of mixed when the newest pulse
  starts with an unclassified event, or when the only mixed burst sits
  among unclassified pulses.

Controls cover the label cases that are already correct.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Periodic detector of the merged, not yet integrated internal/anomaly
library.

Narrow period range: estimate() dropped every seed g/k outside
[MinPeriod, MaxPeriod] before chainRefined could refine it, so a range
narrower than the jitter (gaps 294s/306s, range 299s..301s, tol 6s) found
no period although 300s meets every criterion. A seed is now measured
while some period in range may explain its gap as k periods
(seedInReach: k*MinPeriod - tol(MinPeriod) <= g <= k*MaxPeriod +
tol(MaxPeriod), sound for absolute and relative jitter since k*P - tol(P)
never decreases in P). An out-of-range seed only feeds its refinement and
is never the estimate itself. The seed count per pulse is still at most
recentGaps*(MaxMissing+1), so searchCells (Chance) and periodicWorkBound
are unchanged; measuredSteps is refreshed (+0..6.8% gaps visited).

Mixed bursts: the chain label counted only all-expected pulses, so a
chain whose only expected events sat in bursts together with
unclassified events came out unclassified. Such a chain is now mixed;
monitored still wins, and all-expected chains stay expected.

Adds a property test for seedInReach (soundness under absolute,
relative and combined jitter, and tight bounds) and an invariant test
that an out-of-range seed is measured but never reported.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…exhaustively

From the independent review of 74f86e8 (verdict: approve, one
should-fix).

- With a tolerance below k-1 ns, seedInReach alone rejected a gap in
  (k*MaxPeriod+tol, k*MaxPeriod+k-1] whose seed floor(g/k) = MaxPeriod
  is in range, so the fix could drop a seed master measured and lower the
  estimate. estimate() now measures every in-range seed and, beyond them,
  the out-of-range seeds seedInReach keeps: the seed pool is a superset
  of master's and the in-range ranking is unchanged. Regression test:
  TestPeriodicInRangeSeedIsAlwaysMeasured (fails with seedInReach-only
  admission).
- seedInReach's comment states that the test is necessary, not
  sufficient, and gives the monotonicity argument including the 1 ns
  truncation step of tolNanos.
- TestSeedInReachIsSoundExhaustively checks every explained gap over small
  integer periods, with absolute, relative and combined jitter and every
  k; the sampled test no longer restates the formula as a tightness check.
- The narrow-range reproduction also runs with a relative tolerance.

Work counters of TestPeriodicSearchWorkIsBounded are unchanged by this
commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fidenceLow

The two remaining findings from the review of the merged (not yet
integrated) anomaly library. Every new test fails on 9899b9f at its
intended assertion:

- TestPeriodicRelativeJitterRefinesWithTheCandidateTolerance: JitterRel 2%,
  gaps repeating 305.95s, 305.95s, 294.05s. 300s explains every gap within
  its own 6s tolerance and meets every criterion (asserted as a fixture
  guard), but chainRefined builds the common range with the seed's
  tolerance, so no seed refines to a period that explains the train and it
  is never detected.
- TestCensoredNewStreamKeepsLowConfidenceOn{QuietEnd,EvictionEnd,Reopen}:
  a new-stream episode opened from censored evidence is reported low, but
  its quiet end, its eviction end and a reopen within the cooldown (by an
  uncensored reactivation) come out as route_observed.

Controls (pass before and after): an uncensored stream is never lowered,
and the low confidence does not carry into a new episode after the
cooldown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eLow per episode

Relative jitter (P2): chainRefined intersected the ranges
[(g-tol)/k, (g+tol)/k] with the seed's tolerance, although a period P
explains gap g with its own, tol(P) = max(JitterAbs, JitterRel*P). The
new periodRange returns exactly {P : |g - k*P| <= tol(P)}: the integer
ends as before when JitterRel is 0 (proven against the old formula as an
oracle), and otherwise float estimates of g/(k+JitterRel) and
g/(k-JitterRel) stepped to the exact ends (k*P -+ tol(P) never decrease
in P). Still at most one refined period per seed, so searchCells and the
Chance union bound are unchanged.

The refined period is now clamped into the common range intersected with
[MinPeriod, MaxPeriod]. With exact ranges a common range can straddle
MaxPeriod; unclamped, the phase estimate above it was rejected although
periods in range explain every gap (15 of 10,000 narrow-range edge
trains found by 9899b9f were lost without it). Under absolute jitter
this only adds detections at the range ends.

periodRange results are memoized in the Detector-owned scratch, one slot
per (pulse mod recentGaps, k), at most recentGaps*(maxMissingLimit+1)
entries; an entry is used only when g, JitterAbs and JitterRel match.

Lost ConfidenceLow (P3): the episode now records that one of its hot
signals came from censored new-stream evidence (like its traffic label
and coverage), every transition carries it, and emit lowers the
confidence for any transition of such an episode: quiet end, eviction end
and a reopen within the cooldown. A new episode starts without it.

Golden: exactly one field changes, the quiet end of censored episode
7771b168 from route_observed to low (sha256 3d6fbf08...). The golden
guard keeps its pinned hash and maps that one documented field back.
The pinned signals of Monte Carlo train 1240 become 57/2/69: the train,
broken at gap 58, is found again as 172.48s (residual 2.561s <= 2.587s,
Chance 1.1e-11), which the seed-tolerance range could not reach.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dborup

dborup commented Sep 28, 2026

Copy link
Copy Markdown
Owner Author

Review feedback addressed in commit 0a30e8b4 (tests only) and commit cfbc7c84 (fixes). These are the two remaining findings from the review of the merged anomaly library, so all four findings are now covered in this PR.

  1. Relative jitter (P2). Refinement now uses each candidate period's own tolerance, through the new periodRange. The repro TestPeriodicRelativeJitterRefinesWithTheCandidateTolerance fails on 9899b9f6 and passes now. With absolute jitter only, periodRange gives exactly the old integers; TestPeriodRangeWithAbsoluteJitterIsTheOldRange checks this against the old formula as an oracle.
  2. Lost ConfidenceLow (P3). A censored new-stream episode keeps low on its quiet end, its eviction end and a reopen within the cooldown. There is one repro per path, each red on 9899b9f6, and controls check that uncensored streams and later episodes are not lowered.
  3. Clamp. The refined period is clamped into the common range ∩ [MinPeriod, MaxPeriod]. Without it, 15 of 10,000 narrow-range edge trains that 9899b9f6 found were lost. With it, no detection is lost in any probe cell; under absolute jitter it only adds detections at the range ends. Test: TestPeriodicRefinementIsClampedIntoTheSearchRange.
  4. Golden. Exactly one field changes: the quiet end of episode 7771b168 goes from route_observed to low (sha256 3d6fbf08…). The golden guard keeps its pinned hash and maps that one field back before hashing.
  5. Train 1240 expectation. It changes from 57/1/57 to 57/2/69; the reason is in the test comment.
  6. Memo. periodRange results are cached in the Detector's scratch, with a bounded 144 entries and a full-input key. Outputs are byte-identical to the version without the memo on 560,000 probe trains, and TestRangeMemoIsNeverStale covers it.
  7. Evidence (all numbers are in the PR description):
    • Mutations: all 22 are red, 15 new and the 7 earlier ones.
    • False alarms: 10,000 paired trains per cell. The only change is at JitterRel 0.25 with a uniform null, +0.131 pp; it is accepted and documented as a known limitation. The Poisson rate is unchanged.
    • Detection: no train detected by master or 9899b9f6 is lost.
    • Performance: interleaved A/B; up to +35 % against 9899b9f6 on relative-jitter shapes, with no new allocations.
    • Checks: -race, and TestIndependent* run 10 times, all green.

@dborup dborup changed the title fix(anomaly): periodic detector misses narrow-range periods and mislabels mixed bursts fix(anomaly): periodic refinement, mixed-burst labels and censored-episode confidence Sep 28, 2026
@dborup
dborup marked this pull request as ready for review September 28, 2026 15:18
@dborup
dborup merged commit 18a1326 into master Sep 28, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant