Skip to content

fix(live): recover the shared WebSocket from silent half-open connections - #140

Merged
dborup merged 4 commits into
masterfrom
codex/issue-117-ws-heartbeat
Sep 30, 2026
Merged

dborup merged 4 commits into
masterfrom
codex/issue-117-ws-heartbeat

Conversation

@dborup

@dborup dborup commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Relates to #117

Plan and design

The user asked for autonomous work, so the plan is written here instead of waiting for sign-off (AGENTS.md rule 5).

Commits:

  1. 70fb5659: tests that reproduce the bug (red on master).
  2. 1d32bb7c: test refinement. The reconnect counter now counts only timers whose callback is connectWS itself, and one case was added for a pull during the reconnect delay (still red on master).
  3. 40c72e68: the fix.

Server (cmd/server/websocket.go)

  • writePump writes {"type":"heartbeat"} (20 bytes) right after the protocol ping, on the same ticker.
  • It uses the same single writer goroutine per client, so there is no competing writer and no hub lock.
  • The interval is now Hub.pingInterval (default 30 s) so tests can shorten it.
  • Pong-based dead-client handling (the 60 s read deadline and the pong refresh) is unchanged.

Client (public/app.js)

Liveness:

  • Every received frame refreshes wsLastFrameAt.
  • One watchdog timer checks for silence of WS_STALE_MS = 75 s, measured from socket creation. That is two heartbeat intervals plus slack, so one late or lost heartbeat is tolerated, and a handshake that never completes is caught too.
  • A stale socket is replaced once.

Heartbeat handling:

  • Heartbeats are matched by exact bytes first in onmessage, before the logo pulse, the /stats//nodes cache invalidation and every onWS listener. That includes the Packets pause buffer, which fills from a listener.
  • Packet messages are unchanged.

Reconnect races:

  • connectWS() cancels a pending reconnect and detaches the old socket's handlers before closing it, so a late close event from a replaced socket cannot schedule another connection.
  • As a result, onclose, the watchdog, resume checks and pull-to-reconnect cannot stack timers or sockets.

Resume:

  • visibilitychange (to visible) and online run the check at once, because a hidden or sleeping tab's timers run late.
  • The check does nothing while a reconnect is already scheduled, so an ordinary close keeps the configured WS_RECONNECT_MS delay (default 3 s).

Clock steps:

  • A backward wall-clock step gives a negative silence reading, which is treated as stale instead of re-arming for the size of the step.
  • A forward step (sleep) is measured against real elapsed time via Date.now().
  • Either direction costs at most one extra reconnect.

Pull-to-reconnect:

  • It replaces the socket at once, including when the socket is OPEN and possibly half-open.
  • On master, pull on a socket that was not OPEN left an extra live socket (the old onclose still scheduled a reconnect).
  • On master, pull on an OPEN socket waited for the close event.

Differences from upstream Kpa-clawbot/CoreScope#2020

Upstream is read as a reference only; nothing was cherry-picked.

  • The server side and the core client design (detach before close, watchdog from creation, heartbeat filter first, resume check, negative-silence rule) follow upstream.
  • The watchdog and resume check do nothing while a reconnect is pending. In upstream, a resume or watchdog tick after onclose could bypass the configured WS_RECONNECT_MS. The issue explicitly asks to keep that delay for ordinary close events.
  • onclose ignores a socket that is no longer current, as an extra guard next to the detach.
  • The fork's tests are written from scratch against the fork's app.js: a real DOMContentLoaded boot, and a fake WebSocket whose close event arrives late the way a real one does.
  • A Go test pins the relation between WS_STALE_MS, the heartbeat bytes and the server interval.

Old tabs during a rolling static-asset update

A tab that still runs an app.js from before this change receives the heartbeat every 30 s as an ordinary message:

  • The brand logo pulses once.
  • The /stats and /nodes API cache is invalidated 5 s later.
  • onWS listeners see type: "heartbeat". Every fork listener ignores that type (Live, Map, Nodes, Channels, Observers filter on type; nav stats just refresh).
  • On a paused Packets page the pause counter rises by 2 per minute. On replay, non-packet messages are filtered out.

Nothing breaks, and a reload ends it. Old tabs do not get the watchdog.

Acceptance criteria

Criterion Status Evidence
Small heartbeat from the existing single writer loop, with no competing writer and dead-client handling unchanged Met TestWritePumpSendsHeartbeatOnEveryPingTick_117, TestHeartbeatDoesNotChangeDeadClientHandling_117, TestBroadcastsUnchangedAlongsideHeartbeats_117
Liveness tracked from socket creation, any frame refreshes it, a silent socket or handshake beyond the documented threshold is replaced exactly once Met JS tests 1–4; threshold documented in app.js and here (75 s)
onclose, watchdog, online, visible resume and pull cannot stack timers or sockets Met JS tests 10–16 (one live socket, at most one pending reconnect)
Old handlers are detached before intentional replacement Met JS test 1 (onmessage === null on the replaced socket, and its late close event does nothing)
Heartbeats consumed before logo pulse, cache invalidation, pause buffers and onWS; packets unchanged Met JS tests 5–6
Configured reconnect delay kept for ordinary close; old-tab behaviour documented Met JS tests 11–12; section above

Tests

test-issue-117-ws-watchdog.js

Runs the real app.js in a vm with a fake clock, fake timers and a fake WebSocket, booted through DOMContentLoaded. Registered in test-all.sh and the deploy.yml unit step.

Passed Failed
master d264716c 5 11
this branch 16 0

The 5 that pass on master are guards: heartbeats and traffic keep a socket, packets are dispatched, the close delay is kept, and resume during a pending reconnect is covered.

Mutation check: 11 source mutations were each applied once and all 11 were caught:

  • heartbeat filter removed
  • handler detach removed
  • reconnect-pending guard removed
  • resume wiring removed
  • negative-silence rule removed
  • cancel of the pending reconnect removed
  • watchdog-from-creation removed
  • frame refresh removed
  • drop in connectWS removed
  • threshold ×10
  • the pull race

Go: cmd/server/ws_heartbeat_117_test.go

  • 5 tests: default interval 30 s, heartbeat on every tick next to the ping, frame shape, broadcasts unchanged, and one writer with the read deadline unchanged.
  • One more test checks that app.js's WS_STALE_MS exceeds two intervals and that WS_HEARTBEAT equals the server bytes.
  • On master the file does not compile (no Hub.pingInterval, no wsHeartbeat); on this branch all pass.
  • go test -race -run "_117|Hub|Broadcast|Poller|WS|WebSocket|CheckOrigin|Limit": ok (27.9 s).
  • Full server -race suite (cd cmd/server && go test -race -count=1 ./...): ok github.com/corescope/server 1271.906s.

Existing suites

Suite Result
test-pull-to-reconnect.js 6/6
test-pull-to-reconnect-1091.js 8/8
test-live.js 110/110
test-packet-filter.js 92/92
test-aging.js 19/19
test-frontend-helpers.js 705 passed / 2 failed, identical on master

eslint on public/app.js and scripts/check-xss-sinks.sh --diff origin/master are clean.

Browser check

Local Chromium against a server built from this branch on test-fixtures:

  • Real server: one {"type":"heartbeat"} frame arrived about 30.5 s after load, and an onWS listener did not see it.
  • Silent socket: the WebSocket was routed so it opened but never sent anything, with Playwright's fake clock. It was not replaced at 74 s, it was replaced by 76 s, and the new socket was OPEN.
  • No page errors in either case.

Perf

  • Client: one Date.now() and one string compare per WS message, and one timer per socket.
  • Server: one extra 20-byte write per client per 30 s on the goroutine that already writes the ping.

Not verified

  • A real proxy or NAT half-open drop, laptop sleep, and mobile tab freezing. Only the fake-clock tests and the routed-socket browser check cover these; a frozen but healthy tab may do one unnecessary reconnect on resume.
  • Firefox and Safari.
  • A rolling deploy with old tabs open (the behaviour above is derived from the code).

Overlap with other open PRs

🤖 Generated with Claude Code

https://claude.ai/code/session_019TcZHooUiiknVWbECVWzk8


Generated by Claude Code

dborup and others added 3 commits September 29, 2026 09:04
test-issue-117-ws-watchdog.js runs the real app.js with a fake clock,
timers and WebSocket: 10 of 15 fail on master (no watchdog for a silent
OPEN socket or a stuck handshake, heartbeats dispatched to listeners,
no resume/online check, pull-to-reconnect leaving extra sockets).

ws_heartbeat_117_test.go: heartbeat on every ping tick next to the
unchanged protocol ping, broadcasts unchanged, one writer goroutine,
and app.js's threshold/heartbeat bytes in step with the server. Does
not compile on master (no Hub.pingInterval, no wsHeartbeat).

Relates to #117

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019TcZHooUiiknVWbECVWzk8
…ect delay (#117)

pendingReconnects() now counts only timers whose callback is connectWS
itself (the watchdog timer mentions connectWS in its body), and a new
case checks that pull-to-reconnect during the configured reconnect
delay cancels the pending reconnect. 5 of 16 pass on master.

Relates to #117

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019TcZHooUiiknVWbECVWzk8
…117)

Server: writePump writes {"type":"heartbeat"} right after the protocol
ping on the same tick (Hub.pingInterval, default 30s), from the one
writer goroutine per client. Pong-based dead-client handling and the
60s read deadline are unchanged.

Client (app.js):
- any frame refreshes the liveness clock, measured from socket
  creation; a socket silent for WS_STALE_MS (75s) is replaced once;
- heartbeats are consumed before the logo pulse, cache invalidation and
  every onWS listener (so before pause buffers); packets are unchanged;
- connectWS() cancels a pending reconnect and detaches + closes the old
  socket, so onclose, watchdog, resume and pull cannot stack sockets;
- visible-tab resume and `online` check at once; the check is inert
  while a reconnect is scheduled, so ordinary close keeps its
  configured WS_RECONNECT_MS delay;
- a backward wall-clock step is treated as stale;
- pull-to-reconnect replaces the socket at once also when it is OPEN.

Relates to #117

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019TcZHooUiiknVWbECVWzk8
@dborup

dborup commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

Independent review of 40c72e68

Verdict: APPROVE with nits. This is a recommendation only; merging is the owner's call.

Reviewed head: 40c72e687a3a01f7ff584f7aae8e947d2b512719 (unchanged before and after the review). Work done on a git archive of the head, of commit A 70fb5659 (test-first), of 1d32bb7c (test refinement) and of origin/master ad011021, plus a separate head copy for mutants.

Labels: [F] freshly verified by me · [T] taken from the PR text · [A] assumption · [K] known limitation.

Findings

  1. P3. The Go test for heartbeat cadence does not enforce "every tick". cmd/server/ws_heartbeat_117_test.go:52 (assertion at :70). The test reads 4 heartbeats within 5 s and then checks pings >= heartbeats. A mutant that writes the heartbeat on every other tick passes all six _117 tests. That matters because with a heartbeat every 60 s, the client's 75 s WS_STALE_MS no longer tolerates a single lost heartbeat, which is the stated design margin. A reviewer probe that asserts pings <= heartbeats + 1 is green on the head (3/3 runs) and red on that mutant ("11 pings but only 6 heartbeats"). The probe is kept locally by the reviewer (not committed). [F]
  2. P3. The public WebSocket API spec is not updated for the new message type. docs/api-spec.md:2282-2287 says "All WebSocket messages use this envelope" with type being "packet" or "message" and a data object. The server now pushes {"type":"heartbeat"} every 30 s to every client, with no data field. The endpoint is unauthenticated and documented for third parties. A consumer that reads msg.data.* without filtering on type would now throw. Old-tab behaviour is documented only in the PR body and in a code comment (cmd/server/websocket.go:78-85), not in repo docs. [F]
  3. nit. The new Hub field breaks gofmt alignment in the struct. cmd/server/websocket.go:22. gofmt -d now wants to realign allowedOrigins and limits (lines 20-21). The Client struct drift (lines 89-90) predates this PR, so gofmt -l already lists the file on master. [F]
  4. nit. Two guards are redundant. public/app.js:873 (if (ws !== sock) return;) and the clearTimeout(wsWatchdogTimer) in onclose (:874-875). Deleting either one leaves all 16 JS tests green (mutants J2 and J11). Both are equivalent mutants: dropWS() nulls every handler before ws is reassigned, and a leftover watchdog returns early because wsReconnectTimer is set. The PR calls the first one "an extra guard", so this is deliberate defence in depth, not a defect. [F]

No P1 or P2 findings.

Metadata

  • CI on 40c72e68 is complete: Go Build & Test, Playwright E2E and Docker are SUCCESS; the rest are SKIPPED. [F]
  • Commits in origin/master..40c72e68: 70fb5659, 1d32bb7c, 40c72e68. Author and committer are dborup <kontakt@meshview.dk> on all three. [F]
  • Files: .github/workflows/deploy.yml (+1), cmd/server/websocket.go (+19/−4), cmd/server/ws_heartbeat_117_test.go (new, 154), public/app.js (+87/−16), test-all.sh (+1), test-issue-117-ws-watchdog.js (new, 352). [F]
  • Merge base is d264716c, not ad011021, so the branch is behind master. git merge-tree --write-tree origin/master 40c72e68 is clean (tree 7581213d). [F]
  • PR body: "Relates to fix(live): recover the shared WebSocket from silent half-open connections #117", with no closing keyword. closingIssuesReferences is []. No @mentions. Upstream appears only as `Kpa-clawbot/CoreScope#2020` in code format, with no github.com/Kpa-clawbot URL. The PR is a draft. [F]
  • Workflow: the only change to deploy.yml is + node test-issue-117-ws-watchdog.js in the unit-test list (line 162). No github.repository == 'Kpa-clawbot/CoreScope' line, trigger, permission or job changed. [F]

Acceptance criteria (issue #117)

Criterion Result
Small heartbeat from the existing single writer loop, with no competing writer and dead-client handling unchanged Met [F]. The 20-byte text frame is written in writePump right after the ping on the same tick (websocket.go:279-287), under the same 10 s write deadline. The only other conn write is the pre-existing WriteControl in Hub.Close (:166), which gorilla allows concurrently. readPump (60 s deadline plus pong refresh) is unchanged. A mutant with a second heartbeat goroutine was caught (6 DATA RACE reports plus a test failure).
Liveness tracked from socket creation, any frame refreshes it, a silent socket or handshake beyond the documented threshold is replaced exactly once Met [F]. wsLastFrameAt and the watchdog are set in connectWS (app.js:869-870) and refreshed in onmessage (:885). The 75 s threshold is documented at :669-677. Browser: a blackholed OPEN socket was replaced 75.1 s after its last frame, and a hung handshake (server SIGSTOPped) was replaced at 74.8 s after creation. Both were replaced exactly once.
onclose, watchdog, online, visible resume and pull cannot stack timers or sockets Met [F]. There is one module-level wsWatchdogTimer and one wsReconnectTimer. connectWS clears both and calls dropWS() first. checkWSLiveness does nothing while a reconnect is pending (:844). JS tests 10–16 pass; mutants J3, J7 and J12 are caught. In the browser, SPA navigation across #/live, #/nodes, #/map and #/packets after a replacement added no socket.
Old handlers detached before intentional replacement Met [F]. dropWS (:829-838) nulls all four handlers, then calls close(). Mutant J6 is caught.
Heartbeats consumed before logo pulse, cache invalidation, pause buffers and onWS; packets unchanged Met [F]. Exact-byte match at :888, before Logo.pulse. JS test 5 covers it and mutant J1 is caught. Browser: 1 heartbeat frame received, and an onWS listener saw 0 heartbeats.
Configured reconnect delay kept for ordinary close; old-tab behaviour documented Met [F]/[T]. onclose still uses window.WS_RECONNECT_MS || 3000 (:881); JS test 11 passes with 5000. Old-tab behaviour is documented in the PR body [T]. The public API doc is not updated (finding 2).

Test-first and mutants

  • test-issue-117-ws-watchdog.js (head version): 5 passed / 11 failed on master, 16/0 on the head. Commit A's own 15-test version gives 5/10 on A. [F]
  • The test changed only between A and 1d32bb7c (+15/−3): the reconnect counter now counts only connectWS timers, and one pull-during-delay case was added. There were no test changes in the fix commit. The change is justified. [F]
  • ws_heartbeat_117_test.go does not compile on master (hub.pingInterval undefined). Behavioural red is shown by mutant G1. All 6 pass on the head. [F]
Mutant (head copy) Result
G1 heartbeat write removed caught (TestWritePumpSendsHeartbeatOnEveryPingTick_117)
G2 protocol ping removed caught (same test)
G3 default interval 40 s caught (TestHubDefaultPingInterval_117, TestClientStaleThreshold…_117)
G4 heartbeat from a second goroutine (competing writer) caught (test failure + 6× DATA RACE under -race)
G5 heartbeat only every other tick survived → finding 1 (probe catches it)
J1 heartbeat filter removed caught (1 test)
J2 ws !== sock guard in onclose removed survived, equivalent (finding 4)
J3 reconnect-pending guard removed caught
J4 frame refresh in onmessage removed caught (3)
J5 negative-silence rule removed caught
J6 handler detach removed caught
J7 cancel of pending reconnect in connectWS removed caught
J8 setupWSResumeCheck() wiring removed caught (3)
J9 watchdog-from-creation removed caught (3)
J10 WS_STALE_MS ×10 caught (JS resume test); ×0.8 is also pinned by the Go test (G3)
J11 onclose keeps the watchdog survived, equivalent (finding 4)
J12 dropWS() in connectWS removed caught (8)
J13 old pull behaviour (close only when OPEN) caught

test-pull-to-reconnect.js (6/6) and test-pull-to-reconnect-1091.js (8/8) stayed green on every JS mutant. shasum of app.js and websocket.go in the mutant copy matched git show 40c72e68:<path> after restore. [F]

Suites run locally

  • Go, head: cd cmd/server && go test -race -count=1 -timeout 60m ./... → ok github.com/corescope/server 910.989s, exit 0. [F]
  • readonly_invariant_test.go (4 tests) passes on the head. go vet . is clean. [F]
  • JS: all 38 non-E2E test-*.js files that load public/app.js, plus test-packet-filter.js and test-aging.js, were run on master and the head, and the per-file results are identical. Pre-existing failures on both: test-frontend-helpers.js 705/2, test-packets.js 115/13, test-issue-1470-card-bg-contrast.js 9/1, test-issue-1648-m3-emoji-scan.js rc=1, test-rx-coverage-escape.js rc=1. Everything else passes, including test-live.js 110/0, test-packet-filter.js 92/0, test-aging.js 19/0 and both pull-to-reconnect suites. [F]
  • test-e2e-playwright.js fail-fasts on "Version info lives on Perf dashboard" on both master and the head (known noise). [F]

Browser

Local Chromium (Playwright) against servers built from the head (port 13740) and master (13741) on the freshened e2e fixture. Scripts: r140/halfopen.js, r140/sigstop.js. [F]

  • True half-open, head. A TCP proxy in front of the server silently drops all bytes of the established WS connection, forwarding no FIN, while new connections pass through. A heartbeat arrived at 31.4 s, and the onWS listener saw 0 heartbeats. The proxy blackholed the socket at 34.7 s. A new socket was created and OPEN at 106.5 s (75 s after the last frame). Browser sockets afterwards were [CLOSING, OPEN]: the old one is detached and waiting for its close handshake. After SPA navigation plus 33 s, a heartbeat arrived on the new socket, no extra socket appeared, there were 0 page errors, and server /api/health websocket.clients was 1.
  • True half-open, master. No socket was replaced within 110 s of the blackhole. The page kept one "OPEN" socket while the server reported clients: 0, which is the fix(live): recover the shared WebSocket from silent half-open connections #117 bug reproduced.
  • Server SIGSTOP, head. This gives a silent socket and then a hung handshake. The replacement was created at 75.9 s, 74.8 s after the first socket. It stayed CONNECTING while the server was stopped. After SIGCONT it went OPEN within 8 s, browser sockets were [CLOSED, OPEN], server clients: 1, and there were 0 page errors.

Performance and security

  • Server [F]: one extra 20-byte write per client per pingInterval (30 s), on the goroutine that already writes the ping, with no hub lock. At 2K clients that is about 67 small writes per second, spread by connect time. Broadcast and the poller are untouched, and the heartbeat bypasses the 256-slot send channel, so a busy client cannot drop it. No benchmark is given; none is needed for a claim this structural [A].
  • Client [F]: one Date.now() and one string compare per message, and at most one watchdog timer plus one reconnect timer globally. No per-message allocation is added.
  • Reconnect storms [F]/[K]: stale replacement is at most one attempt per 75 s per tab. The ordinary-close path keeps the pre-existing fixed WS_RECONNECT_MS with no backoff (unchanged, and the issue asks to preserve it). A half-open client's server-side slot is normally already freed by the 60 s pong deadline before the 75 s client watchdog fires, so the WebSocket /ws: per-IP rate limit, conn cap, and source-IP deny list (follow-up to #1793) Kpa-clawbot/CoreScope#1794 per-IP cap is not charged twice. By default the per-IP cap is off and the upgrade rate is 30/min.
  • Leaks [F]: no new goroutines. The ticker is stopped in the writePump defer as before, and no Go-side timers were added. The client clears the watchdog on replace and on close. A replaced half-open socket lingers in CLOSING until the browser's close timeout, bounded by the browser.
  • No DOM sinks touched, no new map[string]interface{} (websocket.go count 1 on both master and the head), no DB writes, and readonly_invariant_test.go is green. [F]

Not verified

  • Firefox and Safari; mobile tab freeze, bfcache and a real laptop sleep. The PR's "Not verified" section lists the same, and it is honest. [K]
  • The PR's claim that "a frozen but healthy tab may do one unnecessary reconnect on resume" follows from the code (visibilitychange may run before queued onmessage tasks). I did not reproduce it. [A]
  • A rolling deploy with old tabs open. I did not run it; the PR's description matches the master onmessage code path. [A]
  • The full cmd/server race suite was not run on master for comparison, because the head run was green.

…heartbeat

# Conflicts:
#	.github/workflows/deploy.yml
#	test-all.sh

dborup commented Sep 30, 2026

Copy link
Copy Markdown
Owner Author

Review feedback addressed (commit 8a3dbb1d)

  1. Merged origin/master (4ce26563) into the branch with a merge commit to resolve the conflict in .github/workflows/deploy.yml and test-all.sh. The fix(live): recover the shared WebSocket from silent half-open connections #117 line (test-issue-117-ws-watchdog.js) and the ui: show the running CoreScope version in the navigation drawer #111 line from feat(nav): show the running CoreScope version in the navigation drawer #142 (test-issue-111-drawer-version.js) are both kept, side by side. Nothing else changed: git diff origin/master..8a3dbb1d touches the same 6 files as before, with +614/−20.

  2. The 9 lines containing github.repository == 'Kpa-clawbot/CoreScope' in deploy.yml are byte-identical to master and to the previous head 40c72e68.

  3. Tests run locally on 8a3dbb1d:

    • test-issue-117-ws-watchdog.js 16/0
    • go test -run _117 ok
    • test-issue-111-drawer-version.js 14/0
    • test-issue-111-drawer-version-e2e.js 7/0
    • test-issue-125-live-toggles-wiring.js 4/0
    • test-issue-125-live-toggles-early-e2e.js 14/0
    • test-pull-to-reconnect.js 6/0
    • test-pull-to-reconnect-1091.js 8/0
    • test-live.js 110/0

    The E2E suites ran in local Chromium against a server built from this head on the migrated e2e fixture.

The review nits (the Go cadence assertion, the docs/api-spec.md heartbeat type, gofmt alignment) are not part of this round, which was scoped to the conflict only.


Generated by Claude Code

@dborup
dborup marked this pull request as ready for review September 30, 2026 12:32
@dborup
dborup merged commit 91d0ac5 into master Sep 30, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant