Repository navigation
port(upstream#2003): join the watchdog loop goroutine in tests - #48
Conversation
…ng it to stop (Kpa-clawbot#2003) Verified before merging: on upstream/master `go test ./cmd/ingestor -run TestMQTTStallWatchdog -count=20` fails; on this branch the same command passes. The flake blocked CI on Kpa-clawbot#2000. (cherry picked from commit 02feb2a) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Why do I (and other contributors from the original repo) get a message every time you do something on your own repo? |
Add TestStartWatchdogTestLoop_StopJoinsLoop. It holds the watchdog loop inside a tick with a blocking emit callback and checks that stop does not return until the callback is released and the loop has exited. Ordering uses channels only. The test waits for done to be closed, so "stop has not returned" cannot just mean the stop goroutine was not scheduled yet. Every goroutine is joined on every path, including failures. With a close-only stop the test fails; with the join it passes. Factor the stop closure into joiningWatchdogStop so the test can hold done and exited. startWatchdogTestLoop behaves exactly as before. Correct the helper comment. The previous commit message and comment said two loops both read LastForceReconnectUnix as 0 before either wrote it. The failure traced locally was different: the loop from EscalateOnPersistentDisconnect_1749, whose clock was 420s ahead, was still running when the throttle test started. It read the non-zero stamp that the throttle test's own loop had just written, measured more than forceReconnectThrottle against its own clock, and forced a second reconnect. Test-only. No production code changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…pa-clawbot#2003) startWatchdogTestLoop's doc comment claims "It is safe to call more than once". joiningWatchdogStop (factored out by the StopJoinsLoop commit) guards close(done) with sync.Once, so a second call skips that close and falls straight to <-exited, which returns immediately once exited is closed. That claim was never exercised. Call stop() twice and assert the second call returns within 1s, matching this package's existing timeout/select idiom. This is a different guarantee than TestStartWatchdogTestLoop_ StopJoinsLoop (previous commit): that test proves a single stop() call waits for the loop to actually exit, not just for done to close. This test proves a second, concurrent-or-sequential stop() call does not hang. Independently authored, rebuilt on top of the StopJoinsLoop commit after that commit landed and changed the tail of this file; not cherry-picked, no unrelated changes carried over. Verified: both new tests together, watchdog/liveness/asyncemit group (count=5, 180/180 pass, 0 fail), same group under -race (count=5, 180/180 pass, 0 data races), gofmt clean, git diff --check clean, no production code touched. Full cmd/ingestor package has two pre-existing issues, independently reproduced on the clean base (3729e12) with this change stashed out: TestPruneOldNeighborMetrics fails 5/5 (dated-fixture time-bomb) and TestBackfillTxLastSeen_ ResolvesFromMaxObservationTimestamp is flaky, 2/5 fails with no watchdog code involved (async-migration timing flake). Neither is touched or caused by this commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
MACMINI — reconciliation handoff: both watchdog tests now on this branchDeveloper-authorized reconciliation of the two Commits, in order
These are genuinely different guarantees, not duplicates: one is about a single stop() actually joining; the other is about calling stop() twice being safe. Both exercise Provenance of commit 3Reauthored directly on top of Fresh validation (this push)
PushRefetched immediately before pushing; remote was still exactly Deviations from the original plan
PR #48 is otherwise unchanged: same base, same first two commits, no merge/rebase/force-push/amend. Handing off to whatever review/merge gate this PR normally goes through next. |
|
Correction to my handoff comment above: I said run |
The comment claimed that losing the sync.Once guard would make a second stop() call hang, so the test would catch it via the 1s timeout. Verified by mutation: removing sync.Once instead makes the second close(done) call panic with "close of closed channel", which fails the test immediately, not by timing out. The 1s timeout only catches a regression in <-exited's closed-channel read. Also states plainly that the two stop() calls are sequential, from the same goroutine, not concurrent. Comment-only. No change to the test's assertions, to StopJoinsLoop, or to production code. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Brings in: - #64: fix(nodes): separate advert timestamps from confirmed relay activity - #48: test(ingestor): reconcile watchdog Stop-test coverage (StopJoinsLoop + StopIsIdempotent), fixing the CI watchdog flake this PR previously hit No conflicts. PR #28's own change (public/rx-coverage.js) is untouched by this merge.
Brings in: - #64: fix(nodes): separate advert timestamps from confirmed relay activity - #48: test(ingestor): reconcile watchdog Stop-test coverage (StopJoinsLoop + StopIsIdempotent) No conflicts. PR #50's own change (channel-rainbow.json, internal/channel/channel_rainbow_test.go) is untouched by this merge. Does not address the row-height test Kpa-clawbot#1122/Kpa-clawbot#1124 failure — that is investigated separately, unresolved.
|
Please stop mentioning me in every pr you meege from upstream. Erwin |
Split out of #25 (commit
6e4929b2there), plus one follow-up commit from this fork. Scope:cmd/ingestortests only; no production code changes.Commits
b2ad17a7: port of upstream Kpa-clawbot/CoreScope#2003 by @efiten (merged 2026-09-11, upstream commit02feb2a88ef05284445f022eb87aca78273fbb4c). Applied withgit cherry-pick -xonto masterfda24ca5with no conflicts. Authorship is kept and the patch-id matches the upstream commit.3729e12a(new, this fork): a regression test for the helper, plus a correction to its comment. Details below.29b3573e(new, this fork, MACMINI):TestStartWatchdogTestLoop_StopIsIdempotent— a different guarantee than commit 2'sStopJoinsLoop(that one proves a singlestop()waits for the loop to exit; this one proves a secondstop()call does not hang). Rebuilt on top of commit 2 after it landed and changed this file's tail; not cherry-picked, no unrelated content carried over. Resolves the "Open point" below.Problem
TestMQTTStallWatchdog_DisconnectedEscalationThrottled_1749failed intermittently when run with the other watchdog tests. Those tests only asked their watchdog loop to stop (close(done)) and never waited for it to exit. A loop from an earlier test could therefore still be running when the next test restoredlivenessRegistryand registered its own sources.Observed mechanism (corrects commit 1)
Commit 1's message and helper comment say two loops both read
LastForceReconnectUnixas 0 before either wrote it. That is not what the local trace showed:TestMQTTStallWatchdog_EscalateOnPersistentDisconnect_1749received its last tick, with a fabricated clock +420 s ahead, and its test returned without waiting for it.LastForceReconnectUnix.forceReconnectThrottle(60 s) had elapsed, so it forced a second reconnect.The throttle itself behaved correctly: it was given two clocks for one source. The "both read 0" ordering can also be forced with instrumentation, but it is not the failure that was caught. Commit 1 is left as is (no history rewrite). Commit 2 corrects the comment, and this description replaces the commit message's explanation.
Change
startWatchdogTestLoopreturns astopthat closesdoneand waits for the loop goroutine to exit. The_1749and_1810watchdog tests now use it.joiningWatchdogStop(behaviour unchanged) so the new test can holddoneandexited.TestStartWatchdogTestLoop_StopJoinsLoop:emit, and the callback blocks until the test releases it. This holds the loop inside a tick.stopis called in a goroutine. The test waits untildoneis closed. That proves the stop goroutine has actually runstop, so "not returned yet" cannot just mean the goroutine was never scheduled.emitis still blocked, the test asserts thatstophas not returned. A joiningstopcannot return at this point at all. A 100 ms window bounds only how long a broken close-onlystopgets to reveal itself; it cannot fail the correct helper.stopreturns, and thatexitedwas already closed whenstopreturned (a second check).time.Afterbounds are safety nets that turn a hang into a failure. Deferred cleanup releasesemit, joins the stop goroutine, joins the loop and restores the registry on every path, includingt.Fatal.Evidence
Local (macOS arm64, go1.26.0)
Includes the causal diagnosis. It was instrumented only in scratch copies and was not committed.
Cause: an instrumented run of the watchdog family on master
693eb045caught one failure with the event order above. Deterministic scratch probes then forced that order and reproduced it: 2 reconnects on master's harness, exactly 1 with this PR's harness.Regression test, red/green. The mutation (removing
<-exitedfromjoiningWatchdogStop) was made only in a scratch copy.-count=20-race -count=20stop(mutation)stop(this PR)An independent reviewer repeated the mutation (50/50 FAIL; 20/20 FAIL with
-race -cpu 1) and ran 300 green runs with-race -cpu 1,2,8 -count=100.On
3729e12a:go vetclean.gofmt -lclean on the touched files.git diff --checkclean.-run 'Watchdog|Liveness|StallWatchdog'):-count=20PASS;-race -count=5PASS.cmd/ingestor(normal and-race): only the known baselinesTestPruneOldNeighborMetricsandTestBackfillTxLastSeen_ResolvesFromMaxObservationTimestampfail, and both are fixed on current master.-racereports a pre-existing data race inTestStatsFileWriter_SampledAtMatchesProcIOSampledAt. The stats writer goroutine is never stopped, and cleanup restoresreadProcSelfIOFnunder it. It reproduces 8/30 with-run TestStatsFileWriteralone. It is in files that neither this PR nor current master changes, and it is not addressed here.Scratch merge with current master
99d336e4(clean merge, no watchdog overlap in master's drift):go vetclean; watchdog family-count=10and-race -count=3PASS; fullcmd/ingestorPASS.Linux / GitHub CI
Not instrumented. CI shows only whether the tests pass on Linux. It does not show why the earlier CI failures happened. The same mechanism is likely but was proven only locally.
Dependencies and merge order
Open point (resolved)
Commit 3 (
29b3573e) reconciles this. MACMINI's localffe6b0c6(on top ofb2ad17a7) held the sameTestStartWatchdogTestLoop_StopIsIdempotenttest but no longer fast-forwarded onto commit 2's head; it was never pushed and stays local/unpushed (preserved under the local refmacmini/preserved-ffe6b0c6-stop-is-idempotent, not deleted). Commit 3 reauthors that same test directly on top of commit 2 as a fresh commit instead.e7f0635f(a diverged local prep-branch commit with ~1000 unrelated lines) was never a candidate for this PR and remains untouched, local, and unpushed. Verification for commit 3: both new tests together, the watchdog/liveness/asyncemit group at-count=5(180/180 pass, 0 fail) both normal and-race, 0 data races,gofmt/git diff --checkclean, no production code touched. The two known full-package baselines above were independently re-confirmed pre-existing on the clean commit-2 base with commit 3 stashed out (TestPruneOldNeighborMetrics5/5 fail;TestBackfillTxLastSeen_ResolvesFromMaxObservationTimestampflaky, 2/5 fail) before commit 3 was added.GitHub CI result
Run
35304301379for3729e12a(commit 2) was in progress at push time; the push then cancelled most of it, presumably via a concurrency group on the branch: final conclusioncancelled. ✅ Go Build & Test finishedsuccessbefore the cancellation; Playwright E2E, Docker Build & Publish, and Deploy Staging werecancelled; Release Artifacts wasskipped. A new run started for29b3573e(commit 3, current head); see that run for the current result.🤖 Generated with Claude Code