Skip to content

fix(ingestor): quiet watchdog retry noise and make force-reconnect shutdown-safe (#102, #103) - #210

Merged
dborup merged 4 commits into
masterfrom
codex/issue-102-103-mqtt-watchdog
Oct 4, 2026
Merged

dborup merged 4 commits into
masterfrom
codex/issue-102-103-mqtt-watchdog

Conversation

@adminopenclaw8-sketch

@adminopenclaw8-sketch adminopenclaw8-sketch commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Relates to #102, #103

Two fixes to the MQTT watchdog's force-reconnect in cmd/ingestor.

Plan

  1. Tests first (commit fc0c068f, red on master), in cmd/ingestor/mqtt_watchdog_102_103_test.go:
  2. Fix (commit 4db36ab8) in main.go (buildForceReconnectFn) and mqtt_watchdog.go (newAsyncEmit).
  3. Test sharpening (commit c9f9f693): the race test keeps the emitters running across stop.
  4. Review fixes (commit 29c2d4d9): the retry-pending line keeps paho's error and claims less (IsConnected() in disconnecting is a hint, not a guarantee); emit after stop is pinned as not counted in watchdogLogDropCount; the emit/stop race test runs 200 rounds.

#102: retry noise during the initial ConnectRetry loop

What paho v1.5.0 does (client.go Connect, status.go Connecting):

paho status Connect() returns paho keeps retrying? IsConnected()
reconnecting success token (no-op, AutoReconnect) yes true
connecting (initial ConnectRetry loop) errStatusMustBeDisconnected yes true (ConnectRetry)
disconnecting, willReconnect set errStatusMustBeDisconnected usually, not guaranteed (see below) true (willReconnect)
disconnecting, willReconnect clear errStatusMustBeDisconnected no false
disconnected starts a fresh attempt – false

buildForceReconnectFn still calls Connect() exactly as before. It now classifies the result:

  • errStatusMustBeDisconnected together with IsConnected() == true is logged as info: WATCHDOG force-reconnect: Connect() returned <err>; paho reports a retry pending (IsConnected=true), not starting a new attempt. paho's errors are unexported, so the error is matched on its text, kept in one constant. The real-paho test fails if a paho upgrade changes that text.
  • In connecting this is reliable: paho's ConnectRetry loop keeps going. In disconnecting, IsConnected() reflects paho's willReconnect flag, which is sticky: an earlier auto-reconnect sets it, a successful reconnect never clears it, and a Disconnect() from connected or reconnecting leaves it alone. So after a Disconnect() paho can report a retry pending that never comes. That makes IsConnected() a strong hint, not a guarantee, so the line keeps paho's error and claims no more than paho reports.
  • The same error with IsConnected() == false (nothing will retry), and any other error, are still logged as WATCHDOG force-reconnect Connect() failed: ….

The doc comment no longer says that Connect() is a no-op in every retrying state. It now lists the cases from the table.

#103: send on closed channel at shutdown

maybeForceReconnect runs ForceReconnectFn in a goroutine that is not joined, and then emits "reconnect attempt issued". On SIGTERM, runLivenessWatchdog's stop waits for the loop goroutine and then closes the emit queue. A force-reconnect still blocked in Connect()/Disconnect() then emits on the closed channel and panics.

Fix: newAsyncEmit's emit checks a stopped flag under a mutex, and stop sets the flag and closes the queue under the same mutex. A line emitted after stop is dropped. stop still does not wait for a blocked ForceReconnectFn, so shutdown is not delayed by a hanging connect. I chose this over joining the goroutines for that reason, and because any late emitter is now safe, not just this one.

No change in normal operation: before stop, emit queues the line, or counts a drop when the queue is full, exactly as before.

Tests

Test Master This PR
TestForceReconnect_RealPaho_InitialRetryLoopIsNotAConnectFailure_102 red: 5 × "Connect() failed", 0 × "retry already in progress" green
TestForceReconnect_RealPaho_ConnectErrorWhileDisconnectingIsLogged_102 green (pins the genuine case) green
TestForceReconnect_RealPaho_DisconnectWhileReconnectingKeepsConnectError_102 red: the line had no error (disconnecting with willReconnect set after Disconnect(); no retry follows) green
TestBuildForceReconnectFn_OtherConnectErrorIsLogged_102 green (pins other errors) green
TestNewAsyncEmit_EmitAfterStopDoesNotPanic_103 red: send on closed channel green
TestNewAsyncEmit_EmitAfterStopIsNotCountedAsDrop_103 green (pins that a drop after stop is not a "queue full" drop) green
TestNewAsyncEmit_ConcurrentEmitDuringStop_103 (200 rounds) red: panic and data race green
TestRunLivenessWatchdog_ForceReconnectInFlightAtShutdown_103 red: panic: send on closed channel (test binary crashes) green

fakeClient.IsConnected (mqtt_force_reconnect_race_test.go) now returns a field instead of panicking, because the fix calls it after a Connect() error.

Performance

The watchdog emit path is not a hot path: it runs about once per source per 60 s tick, plus one line per force-reconnect. The uncontended mutex adds about 15 ns per emit. In a scratch benchmark (not committed, n=6, darwin/arm64), the old emit took a median of 46 ns/op and the new one 62 ns/op, with the same 16 B and 1 alloc. The packet ingest path is untouched.

Rules

  • Changes stay in cmd/ingestor; cmd/server, internal/ and .github are untouched (fork guards in deploy.yml: 9).
  • No new map[string]interface{}.
  • No new configuration values.

🤖 Generated with Claude Code

dborup and others added 3 commits October 4, 2026 10:37
…102, #103)

Red on master:

- #102: a real paho client in its initial ConnectRetry loop (status
  connecting). Five force-reconnects log "Connect() failed" five times,
  though paho keeps retrying and connects once the broker is up.
- #103: newAsyncEmit's emit after stop panics with "send on closed
  channel", also when racing stop (and -race reports a data race). The
  issue's scenario: a force-reconnect whose ForceReconnectFn blocks, the
  watchdog stops, then the reconnect returns and its goroutine emits
  "reconnect attempt issued": the test binary panics.

Green on master, pinning what must stay an error: Connect()'s status
error while paho is disconnecting with no reconnect pending, and any
other Connect() error. fakeClient.IsConnected now returns a field.

Relates to #102, #103

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…utdown-safe (#102, #103)

#102: while paho is in its initial ConnectRetry loop (status connecting),
or disconnecting with a reconnect pending, Connect() returns paho's
errStatusMustBeDisconnected and paho keeps retrying on its own (paho
v1.5.0 client.go Connect, status.go Connecting). buildForceReconnectFn
now recognises that error together with IsConnected()==true and logs
"retry already in progress in paho" instead of "Connect() failed". The
same error with no retry pending (disconnecting after a Disconnect), and
any other error, are still logged as failures. The doc comment no longer
claims Connect() is a no-op in every retrying state.

#103: the force-reconnect goroutine is not joined at shutdown, so its
"reconnect attempt issued" emit could run after stop had closed the
queue and panic with "send on closed channel". newAsyncEmit's emit now
checks a stopped flag under a mutex and drops lines after stop. stop
still does not wait for a blocked ForceReconnectFn.

No change in normal operation: before stop, emit queues or counts a
drop exactly as before.

Relates to #102, #103

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The emitters now start before stop and keep emitting until after it has
returned, so -race sees unsynchronised reads both sides of the stopped
write. A mutant that reads the flag without the mutex is caught in 18 of
20 runs (was 2 of 10); master still fails 20 of 20.

Relates to #103

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@adminopenclaw8-sketch

Copy link
Copy Markdown
Collaborator Author

Rapport — CS-Minimax PR#210 #102+#103 — head c9f9f69

Status: ready for review (draft): both issues fixed in cmd/ingestor, tests red on master and green here, 5 mutants killed, CI green.

Base 6bdaa8f7. Commits:

  • fc0c068f: tests, red on master;
  • 4db36ab8: the fix;
  • c9f9f693: a stronger race test.

Paho behaviour comes from the paho.mqtt.golang v1.5.0 source in the module cache (client.go Connect/IsConnected/Disconnect, status.go Connecting/Disconnecting), not from comments.

Evidence tags: [T] test or CI, [A] analysis, [K] known, not re-run.

Acceptance criteria

Issue Criterion Met Evidence
#102 Recognise paho's connecting state and log it as info, not as a Connect() failure yes: errStatusMustBeDisconnected together with IsConnected()==true logs retry already in progress in paho, no new attempt needed [T][A]
#102 Genuine Connect() failures are still logged as errors yes: the same error with no retry pending (disconnecting after Disconnect) and any other error log Connect() failed [T]
#102 Correct the code comment yes: the buildForceReconnectFn doc lists Connect() per paho status (reconnecting, connecting, disconnecting, disconnected) [A]
#102 Test with a real paho client in connecting, reusing the in-test broker yes: TestForceReconnect_RealPaho_InitialRetryLoopIsNotAConnectFailure_102 reuses forceReconnectTestBroker; 5 triggers; paho's loop keeps sending CONNECT and connects after the broker is up [T]
#103 Emit path is shutdown-safe yes: newAsyncEmit checks a stopped flag under a mutex; emit after stop is dropped [T][A]
#103 Test: blocking ForceReconnectFn, then shutdown, then unblock; no panic; -race yes: TestRunLivenessWatchdog_ForceReconnectInFlightAtShutdown_103 uses the real runLivenessWatchdog; it also asserts that stop() does not wait for the blocked reconnect [T]
both No behaviour change in normal operation before stop, emit queues or counts a drop as before; Connect() is still called in the same cases; only the log line for the retrying case changes [A]

Tests

Test On master (fc0c068f) On head
TestForceReconnect_RealPaho_InitialRetryLoopIsNotAConnectFailure_102 red: 5 × "Connect() failed", 0 × "retry already in progress" [T] green [T]
TestForceReconnect_RealPaho_ConnectErrorWhileDisconnectingIsLogged_102 green, pins the genuine case [T] green [T]
TestBuildForceReconnectFn_OtherConnectErrorIsLogged_102 green, pins other errors [T] green [T]
TestNewAsyncEmit_EmitAfterStopDoesNotPanic_103 red: send on closed channel [T] green [T]
TestNewAsyncEmit_ConcurrentEmitDuringStop_103 red: panic and DATA RACE, 20/20 runs [T] green [T]
TestRunLivenessWatchdog_ForceReconnectInFlightAtShutdown_103 red: panic: send on closed channel (test binary crashes) [T] green [T]
Run Result
cd cmd/ingestor && go test -count=1 ./... ok (101 s) [T]
go test -race -count=5 -run '_10[23]$' . 6 tests × 5 passes, no race [T]
go test -race -run 'Watchdog|ForceReconnect|AsyncEmit|MQTT|Liveness' . ok, so the existing watchdog and paho tests still pass [T]
go vet ./... ok [T]
gofmt -l on the 4 changed files clean [T]
sh test-all.sh 214 passed, 0 failed [T]

Mutants

Each mutant was applied to the fix commit, run under -race against all 6 new tests, then reverted.

# Mutant Killed by
M1 connectRetryInProgress returns false (old behaviour) …InitialRetryLoopIsNotAConnectFailure_102 [T]
M2 drop && client.IsConnected() …ConnectErrorWhileDisconnectingIsLogged_102 [T]
M3 drop the error-text match (only IsConnected()) …OtherConnectErrorIsLogged_102 [T]
M4 if stopped { becomes if false && stopped { …EmitAfterStopDoesNotPanic_103, …ConcurrentEmitDuringStop_103, …ForceReconnectInFlightAtShutdown_103 [T]
M5 emit reads stopped without the mutex …ConcurrentEmitDuringStop_103 (data race): 18 of 20 runs after the sharpening in c9f9f693, 2 of 10 before [T]

All other tests stay green under each mutant, so each mutant is caught by its own test. [T]

Performance

The watchdog emit path is not a hot path (about one line per source per 60 s tick). [A] In a scratch benchmark (not committed, n=6), uncontended emit went from a median of 46 ns/op to 62 ns/op, with the same 16 B and 1 alloc. [T]

Rules

  • Changes are in cmd/ingestor only; cmd/server, internal/ and .github are unchanged against the base, and deploy.yml still has 9 fork guards. [T]
  • No new map[string]interface{} (0 added lines). [T]
  • No new config values. [A]

CI (run 37189990781, head c9f9f693)

Job Result
Go Build & Test success (ingestor package ok, 118 s with coverage) [T]
Playwright E2E Tests success [T]
Build & Publish Docker Image success [T]
Release Artifacts skipped (PR) [T]
Deploy Staging skipped (PR) [T]
Publish Badges & Summary skipped (PR) [T]

Remaining

  • The error is matched on the text of paho's unexported errStatusMustBeDisconnected. A paho upgrade that changes the text fails …InitialRetryLoopIsNotAConnectFailure_102; it is not silently missed. [A]
  • The "reconnect attempt issued" line of a force-reconnect still in flight at shutdown is now dropped, not logged. It is not counted in WatchdogLogDropCount, which keeps meaning "queue full". [A]
  • origin/master has moved to 8bafcdf2 (fix(ui): 48px navigation and channel buttons (port of upstream 2078) #201, UI only) since the base. GitHub reports the PR as mergeable. [T]
  • Not run: browser validation (backend-only change) and staging. [A]

Generated by Claude Code

@adminopenclaw8-sketch

Copy link
Copy Markdown
Collaborator Author

Review — CS-Macmini PR#210 mqtt-watchdog — head c9f9f69

Dom: APPROVE med nits

This is an independent, read-only review of head c9f9f693 and of the tree that merges it into origin/master 0f88865b (merged tree 7dbcb43e). Paho behaviour comes from the paho.mqtt.golang v1.5.0 source in the module cache (client.go Connect/IsConnected/Disconnect/internalConnLost, status.go), not from comments.

Evidence tags: [T] run here, [A] analysis of the source, [K] taken from the author's report or CI, not re-run.

Findings

# Severity Finding Evidence
F1 should-fix, non-blocking IsConnected()==true in status disconnecting does not always mean a reconnect will follow. Paho's willReconnect flag is sticky: ConnectionLost(true) sets it, a successful reconnect never clears it, and Disconnecting() from connected or reconnecting leaves it alone. So after a user Disconnect(), Connect() can return errStatusMustBeDisconnected while IsConnected() is true, and the PR logs "retry already in progress" although nothing will retry. The doc comment on connectRetryInProgress, the buildForceReconnectFn comment and the PR's acceptance table all state the opposite. [T][A]
F2 nit The emit-after-stop drop is not pinned. A mutant that counts it in watchdogLogDropCount survives, though the PR says the counter keeps meaning "queue full". [T]
F3 nit One shutdown-ordering mutant is caught only some of the time. It closes the queue and only then sets stopped, each under the lock. TestNewAsyncEmit_ConcurrentEmitDuringStop_103 catches it in about 1–2 of 20 runs under -race and 98 of 200 without. [T]
F4 info Matching on the error text is the only option here. Paho does not export errStatusMustBeDisconnected, so errors.Is cannot be used. Connect() stores the error unwrapped (t.setError(err)), so exact == on Error() is correct. A change of the text fails the real-paho test, so it would not be missed silently. [A]

F1 in detail

These are review-only probe tests, not committed. They use a real paho client against the in-test broker and read paho's internal status through reflection.

  • P2 [T]: on a fresh client, status=connected willReconnect=false. After one auto-reconnect (broker down, then up), it is status=connected willReconnect=true.
  • P1 [T]: the client is connected, the broker goes down, and paho starts reconnecting. Then go client.Disconnect(0) runs. Paho shows status=disconnecting willReconnect=true IsConnected=true, and buildForceReconnectFn logs WATCHDOG force-reconnect: retry already in progress in paho, no new attempt needed. The broker then comes back up. Three seconds later the client has still not reconnected: IsConnectionOpen=false IsConnected=false.
  • P3 [T]: the production path. After an auto-reconnect (flag now stale), the session is made half-open and buildForceReconnectFn runs: Disconnect(250), then Connect(). In 5 of 5 triggers, Disconnect(250) had already reached disconnected. Connect() started a fresh attempt (status afterwards connecting), so neither line was logged.

How far this reaches [A]:

  • The watchdog itself calls Disconnect only from connected. The mislabel therefore needs that Disconnect(250) to still be in disconnecting when Connect() runs, on a client that has auto-reconnected at least once. That is the fix: stop watchdog force-reconnect from racing paho's own retry loop Kpa-clawbot/CoreScope#1897 race, which the buildForceReconnectFn comment says normally finishes well within the quiesce. P3 agrees.
  • P1 can also happen when a shutdown Disconnect coincides with a force-reconnect that is still in flight. That is harmless.
  • No hang turns silent. Both lines are plain log.Printf; the ingestor has no log levels, so the frequency and visibility are unchanged. The stall WARN, the hourly heartbeat and the bug(mqtt): watchdog goroutine goes completely silent — 3 sources stalled 75 min, zero WATCHDOG log lines, no force-reconnect Kpa-clawbot/CoreScope#1749 ESCALATION path (paho disconnected … with no auto-reconnect — forcing reconnect) are untouched, and the ESCALATION path catches a client left disconnected.
  • What does change: a disconnecting state that never completes (willReconnect set, comms workers never stop) now prints an assertive "retry already in progress" every throttle window. Before, it printed the raw paho error.

Suggested fix, small and in this PR or a follow-up:

  • Keep the error in the info line and make it less assertive, for example Connect() returned %v; paho reports a retry pending (IsConnected=true), not starting a new attempt.
  • Correct the two comments: IsConnected() in disconnecting reflects paho's willReconnect, which is sticky after any earlier auto-reconnect, so it is a strong hint, not a guarantee.

1. #102: paho behaviour

  • The status table holds for connecting and reconnecting [A]:
    • Connecting() returns errAlreadyConnectedOrReconnecting for connected/reconnecting. Connect() turns that into a success token when AutoReconnect is set.
    • For connecting/disconnecting, Connecting() returns errStatusMustBeDisconnected. Connect() returns that as an error token.
    • For disconnected, a fresh attempt starts.
    • IsConnected() is true for connected, for connecting with ConnectRetry, for reconnecting, and for disconnecting && willReconnect.
  • errStatusMustBeDisconnected together with IsConnected()==true is reliable for the initial ConnectRetry loop. The time between the two reads is harmless: connecting can only move to connected (true is correct) or disconnecting after a Disconnect (false is correct). [A]
  • For disconnecting the signal is not reliable; see F1. [T][A]
  • The text match is robust enough; see F4. [A]

2. #103: shutdown

  • Can emit block stop, or the reverse? No. emit holds the mutex only around a non-blocking select with default, so it cannot block even when the queue is full. stop holds it only around stopped = true; close(queue). realEmit, the possibly blocking sink, runs in the drain goroutine without the mutex. stop waits on drained after unlocking. A sink that blocks stop through <-drained is pre-existing fix(#1749): decouple watchdog emit from blocking I/O (root cause) Kpa-clawbot/CoreScope#1853 behaviour and unchanged. [A]
  • Is the queue still closed correctly? Yes. It is closed exactly once, under the mutex, and after the watchdog loop has exited (<-loopExited comes before stopEmit() in runLivenessWatchdog). A second stop() would still panic on the double close, as before; nothing calls it twice. [A]
  • Do goroutines leak after stop? Not beyond the blocked call. The force-reconnect goroutine lives until ForceReconnectFn returns. Its emit is then a no-op and it exits. TestRunLivenessWatchdog_ForceReconnectInFlightAtShutdown_103 asserts that the goroutine is gone after release, and that stop() does not wait for it. [T] In production, ForceReconnectFn is Disconnect(250) plus a non-waiting Connect(), so it returns within about 250 ms. [A]
  • Is drop counting before stop unchanged? Yes. The select/default branch is byte-identical, and the existing TestNewAsyncEmit_NeverBlocksWhenWriterStuck still asserts a non-zero drop count. [T][A] Drops after stop are not pinned; see F2.

3. Normal operation

Before stop, emit queues or counts a drop exactly as before, and Connect()/Disconnect() are called under the same conditions. The only change is the text of the log line in the retry case. [A] The existing watchdog and real-paho tests pass on the merged tree. [T]

4. Tests and mutants

Run Result
cd cmd/ingestor && go test -race -count=1 ./... on the merged tree (origin/master 0f88865b + head), go1.27.0 darwin/arm64 ok (439 s) [T]
go test -race -count=10 -run '_10[23]$' . on the merged tree (the 6 new tests) ok (5.9 s) [T]
CI on head (Go Build & Test, Playwright, Docker) success [K]

My own mutants were each applied to a copy of head and run with go test -race -count=3 -run '_10[23]$|AsyncEmit|ForceReconnect':

# Mutant Result
M6 Shutdown race: emit reads stopped under the lock, unlocks, then sends (check-then-act) killed: ConcurrentEmitDuringStop_103 3/3, plus DATA RACE [T]
M7 Shutdown race: stop closes the queue under the lock, unlocks, then sets stopped under the lock killed only some of the time: survived -count=3; with -count=20 under -race, 1–2 failures in each of 5 batches; without -race, 98 of 200 runs fail (F3) [T]
M8 && client.IsConnected() becomes && !client.IsConnectionOpen() killed: ConnectErrorWhileDisconnectingIsLogged_102, TestBuildForceReconnectFn_LogsConnectError [T]
M9 Emit after stop increments watchdogLogDropCount survived (F2) [T]

Separately, the author reports mutants M1–M5 as killed [K].

5. Performance

Not a hot path. newAsyncEmit is used only by runLivenessWatchdog. emit is called from the watchdog loop (at most a few lines per source per tick) and from force-reconnect goroutines (throttled to 1 per forceReconnectThrottle per source). handleMessage still does only an atomic store, and the mutex is never taken on the message path. connectRetryInProgress runs only on a Connect() error. [A]

6. Rules

  • Only cmd/ingestor/ changed: 4 files against origin/master on the merged tree, and .github is unchanged. [T]
  • No new map[string]interface{}: 0 added lines. [T]
  • Fork guards: deploy.yml has 9, release-fast-path.yml has 1, both unchanged against master. [T]
  • No closing keywords in the title, body or the 3 commit messages ("Relates to" only). [T]
  • All 3 commits have author and committer dborup <kontakt@meshview.dk>. [T]

Not verified

  • Behaviour on staging or prod, and with any paho version other than v1.5.0.
  • Whether a slow message handler can stretch the watchdog's Disconnect(250) past the quiesce, which is the condition for F1 to reach production. Only the normal half-open path (P3) was probed.
  • Browser validation: not applicable, backend only.
  • The author's M1–M5 were not re-run.

The head was c9f9f693 both before and after this review (git ls-remote). The PR is still a draft and was not modified.


Generated by Claude Code

… emit-after-stop (#102, #103)

Review feedback on #210:

- F1: IsConnected() in status disconnecting reflects paho's sticky
  willReconnect flag, so it is a strong hint that a retry follows, not a
  guarantee. The info line now keeps paho's error and says only what paho
  reports: "Connect() returned <err>; paho reports a retry pending
  (IsConnected=true), not starting a new attempt". The comments on
  buildForceReconnectFn and connectRetryInProgress say so. A new real-paho
  test covers Disconnect() while paho is reconnecting: disconnecting,
  IsConnected()=true, no retry follows, and the line keeps the error.
- F2: a new test pins that emit after stop is not counted in
  watchdogLogDropCount.
- F3: the concurrent emit/stop test runs 200 rounds, so a stop that closes
  the queue before it marks itself stopped is caught reliably.

Relates to #102, #103

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@adminopenclaw8-sketch

Copy link
Copy Markdown
Collaborator Author

Rapport — CS-Minimax PR#210 runde 2 — head 29c2d4d

Status: review nits F1–F3 addressed in one commit in cmd/ingestor; local tests green under -race; M7 and M9 now killed in every run; CI pending when this was posted; PR stays draft.

Review feedback addressed (commit 29c2d4d9)

Evidence tags: [T] test or run here, [A] analysis of the source (paho.mqtt.golang v1.5.0 client.go, status.go), [K] known, not re-run.

  1. F1 (should-fix): the retry-pending line claimed too much.
    • New info line, which keeps paho's error: MQTT [<tag>] WATCHDOG force-reconnect: Connect() returned status can only transition to connecting from disconnected; paho reports a retry pending (IsConnected=true), not starting a new attempt. A Connect() error with IsConnected()==false, and any other error, is still logged as Connect() failed: …. [T]
    • Comments corrected in two places, the buildForceReconnectFn doc and the connectRetryInProgress doc. In connecting, IsConnected()==true is reliable, because the ConnectRetry loop keeps going. In disconnecting it reflects paho's willReconnect, and that flag is sticky: an auto-reconnect sets it, a successful reconnect never clears it, and Disconnect() from connected or reconnecting leaves it alone. So it is a strong hint, not a guarantee. [A]
    • The PR description now says the same. The status table has a disconnecting, willReconnect set row marked "usually, not guaranteed", and the explanation is below the table. Relates to #102, #103 is still at the top. [T]
    • Tests: …InitialRetryLoopIsNotAConnectFailure_102 now asserts 5 lines that each start with Connect() returned <paho error>; paho reports a retry pending. …ConnectErrorWhileDisconnectingIsLogged_102 and …OtherConnectErrorIsLogged_102 assert that the new marker is absent. [T]
    • New test TestForceReconnect_RealPaho_DisconnectWhileReconnectingKeepsConnectError_102 covers the reviewer's P1 with a real paho client and the in-test broker:
      • setup: connected, broker down, paho reconnecting (willReconnect=true), then Disconnect(0);
      • premise: status disconnecting with IsConnected()==true and IsConnectionOpen()==false;
      • the force-reconnect logs the line with paho's error exactly once;
      • with the broker back up, paho settles to IsConnected()==false, so the retry it reported never happens. [T]
      • It is red on c9f9f693 (0 lines with the error). [T]
      • It is deterministic: Connect() is used as a side-effect-free status probe (success token while reconnecting, the status error once disconnecting), and MaxReconnectInterval is 3 s, so paho's sleep of 1 s or more covers the 5 ms polling. [A]
      • It takes about 1 s per run. [T]
  2. F2 (nit): emit after stop was not pinned. New TestNewAsyncEmit_EmitAfterStopIsNotCountedAsDrop_103: after stop, asyncEmitQueueSize+10 emits leave WatchdogLogDropCount() unchanged. The package has no t.Parallel, so the global counter is not disturbed by other tests. [T] No production change.
  3. F3 (nit): ConcurrentEmitDuringStop_103 caught M7 only some of the time. I did the cheap fix: the test now runs 200 independent rounds (fresh newAsyncEmit, 8 emitters across stop) and reports the first panic with its round number. It takes 0.13 s under -race. [T] There is no test seam, and production code is unchanged. Under M7, in 100 runs under -race, the first failing round was min 0, median 11, max 68 of 200, so the margin is large. [T]

Tests

Run (cmd/ingestor, go1.27.0 darwin/arm64) Result
go test -race -count=1 ./... at 29c2d4d9 ok (409 s) [T]
go test -race -count=10 -run '_10[23]$|ForceReconnect|AsyncEmit' . (new and touched tests, plus the existing force-reconnect/asyncEmit tests) ok (64 s) [T]
go vet . ok [T]
gofmt -l on the 2 changed files clean [T]
Red first: the two #102 tests that assert the new line, run against the test change before the fix red: 0 matching lines each [T]

Mutants

Each mutant was applied to a copy of 29c2d4d9 (git archive), not to the worktree.

# Mutant Result
M7 stop closes the queue under the lock, unlocks, then sets stopped under the lock killed by ConcurrentEmitDuringStop_103: 100/100 runs under -race, 200/200 without (send on closed channel). Review round: 1–2 of 20 under -race, 98 of 200 without. [T]
M9 emit after stop increments watchdogLogDropCount killed by EmitAfterStopIsNotCountedAsDrop_103: 10/10 (counted 266 drops … want 0). No other test fails. [T]
M10 (extra) the retry-pending line drops paho's error killed by …InitialRetryLoopIsNotAConnectFailure_102 and …DisconnectWhileReconnectingKeepsConnectError_102, 3/3 each [T]

M1–M5 from round 1 and M6/M8 from the review were not re-run. [K] The code they target is unchanged except for the log text.

Scope

  • Changed: cmd/ingestor/main.go (one log line plus comments) and cmd/ingestor/mqtt_watchdog_102_103_test.go. [T]
  • buildForceReconnectFn keeps its signature (client mqtt.Client, tag string) func(), so the later conflict with fix(brokerurl): mask user, password, query values and decoded forms without residue (#159) #213 stays limited to the log calls. [T]
  • No other behaviour change: Connect()/Disconnect() are called in the same cases, the classification is unchanged, and newAsyncEmit is unchanged. [A]
  • No new map[string]interface{}. [T]
  • The commit author and committer are dborup <kontakt@meshview.dk>. [T]
  • Pushed as a normal fast-forward (c9f9f693..29c2d4d9), with no rebase or force-push. [T]

CI and remaining

  • CI on 29c2d4d9: pending when this was posted (1 check running). The last CI result, on c9f9f693, was green. [K]
  • Remaining:
    • the error is still matched on paho's unexported error text, as in round 1 (F4, info);
    • whether a slow message handler can stretch Disconnect(250) past the quiesce, which is what lets F1 reach production, was not probed;
    • browser validation and staging are not applicable or not run (backend only).

Generated by Claude Code

dborup-agent pushed a commit that referenced this pull request Oct 4, 2026
main() is not unit-testable, so a source-text guard (as for the route
mask backfill) pins its share of the wiring: prepareMQTTSource, then
setup.attachClient(client) before the first Connect(), and the initial
connect error logged with setup.secrets. Red at fe40aae. Kills the
review's mutant M5 and its variants, and guards the coming merge with
#210 in the same loop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dborup
dborup marked this pull request as ready for review October 4, 2026 11:14
@dborup
dborup merged commit a0086bd into master Oct 4, 2026
6 checks passed
dborup-agent pushed a commit that referenced this pull request Oct 4, 2026
Conflict in cmd/ingestor/main.go buildForceReconnectFn, resolved by
keeping both sides:
- from #210: the doc comment on paho's Connect() statuses, the switch on
  err with connectRetryInProgress, and the retry-pending info line;
- from #213: the secrets ...string parameter, and errForLog(err,
  secrets...) in both the retry-pending and the Connect() failed line.
connectRetryInProgress still reads the raw error, whose text it
compares with paho's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dborup-agent pushed a commit that referenced this pull request Oct 4, 2026
#159)

#210 added a retry-pending info line next to the Connect() failed line.
The test drives it with paho's status error and IsConnected()=true, and
a configured password that occurs in that text, and asserts the line is
masked and still classified as retry-pending (classification reads the
raw error).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants