fix(etcp): count progress ahead of a probe as liveness and keep the catchup until the server speaks - #6
Conversation
…atchup until the server speaks Liveness: the watcher took only inbound frames as proof of life, and the probe queues behind any unsent backlog. On a slow uplink with the server silent, a healthy link was declared dead; each reconnect replayed the backlog and the probe answer queued behind it again, so the upload could loop without finishing. At 1 KiB/s, 10 of 96 KiB arrived in an hour of fake time. The writer now sends each batch in 4 KiB pieces and signals liveness for a piece written while a probe still waits behind it: the echo cannot come before the probe is sent, and once the send buffer is full a write completes only as the peer acknowledges data. Writes with no probe behind them, the probe itself included, never count, so neither keystrokes nor large packets written into a dead link keep it alive. Two limits remain and are documented on Dialer.KeepAlive: an uplink under about 400 B/s, and data already in the kernel send buffer, which drains out of sight. Replay trim: recover trimmed the ring to ReplayLimit right after writing our catchup, but the server counts a catchup only once it has decoded the whole message. A link lost in between made the next recover fail with ErrReplayExceeded. The ring now holds everything from the sequence the peer acknowledged until the server's first frame on the new link, which upstream sends only after decoding our catchup (BackedWriter::write waits on the recover mutex, src/base/BackedWriter.cpp:17-18 and src/base/Connection.cpp:109,134-142 at et-v7.0.0). The hold is capped at twice ReplayLimit of written bytes, beyond which the next recover fails with ErrReplayExceeded as before. BenchmarkDrainBacklog shows no significant change against main (benchstat p=0.57). Verified: go test -race ./... (etcp also -count=20), golangci-lint (linux and windows), go fix -diff, wasm vet, and the e2e suite against etserver on localhost. Each new test was checked by removing the production line it pins and watching it fail.
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Warning Review limit reached
Reviews can continue after your included limit without a manual trigger. An admin must approve usage-based billing. Next included review available in 51 minutes. View limit detailsLimit details: You’ve used the included review currently available. Your 61 included PR review attempts over the past 7 days set your current allowance at 1 review per hour. Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. Review configuration: ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Essentials Run ID: 📒 Files selected for processing (2)
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Essentials Run ID: 📒 Files selected for processing (11)
Included review availability: This review used your included allowance. 0 included reviews remain after this review. Your included PR review attempts over the past 7 days set your current allowance at 1 review per hour. WalkthroughThe link writer now reports progress in bounded pieces while a keepalive probe waits behind queued data. Recovery now retains unacknowledged catchup entries until the server sends a frame on the new link, then trims the ring. ChangesETCP Link Liveness and Recovery
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Bug fix Merge Risk: ⚪ Minimal · up to The reviewed reconnect changes have no established issue requiring resolution before merge. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Comment |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The only noted issue is a non-blocking comment wording nit.
Review effort: Lite
Findings: 1
Open (1)
What changed in this PR
Fixes reconnect reliability by recognizing upload progress as liveness and retaining replay data until the server responds.
Changes:
- Adds chunked, probe-aware liveness signaling.
- Retains catchup data with bounded ring growth.
- Adds recovery, liveness, throttling, and benchmark coverage.
- Updates keep-alive and replay-limit documentation.
| File | Description |
|---|---|
internal/etcp/throttle_test.go |
Adds throttling test support; includes a minor comment-wording nit. |
internal/etcp/ring.go |
Implements bounded hold-aware trimming. |
internal/etcp/ring_test.go |
Tests ring retention limits. |
internal/etcp/recover.go |
Preserves catchup data until the server speaks. |
internal/etcp/recover_test.go |
Tests recovery and replay retention. |
internal/etcp/liveness_test.go |
Tests liveness during slow and dead-link uploads. |
internal/etcp/link.go |
Implements chunked writes and probe-aware liveness. |
internal/etcp/link_internal_test.go |
Adds internal link behavior coverage. |
internal/etcp/helpers_test.go |
Updates test helpers. |
internal/etcp/drain_bench_test.go |
Updates backlog-draining benchmarks. |
internal/etcp/dialer.go |
Documents updated liveness and replay bounds. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Summary
Two reconnect-path fixes in
internal/etcp.Liveness under a slow upload. The watcher took only inbound frames as proof of life, and the KEEP_ALIVE probe queues behind any unsent backlog. On a slow uplink with a silent server, a healthy link was declared dead; each reconnect replayed the backlog and the probe's echo queued behind it again, so an upload could loop without finishing (at 1 KiB/s, 10 of 96 KiB arrived in an hour of fake time). The writer now sends each batch in 4 KiB pieces and signals liveness for a piece written while a probe still waits behind it. The echo cannot come before the probe is sent, and once the send buffer is full a write completes only as the peer acknowledges data. Writes with no probe behind them, the probe itself included, never count, so neither keystrokes nor large packets written into a dead link keep it alive. Two limits remain and are documented on
Dialer.KeepAlive: an uplink under about 400 B/s, and data already in the kernel send buffer, which drains out of sight.Replay trim after recover.
recovertrimmed the ring to ReplayLimit right after writing our catchup, but the server counts a catchup only once it has decoded the whole message. A link lost in between made the next recover fail withErrReplayExceeded. The ring now holds everything from the sequence the peer acknowledged until the server's first frame on the new link, which upstream sends only after decoding our catchup (BackedWriter::writewaits on the recover mutex thatConnection::recoverholds until then; src/base/BackedWriter.cpp:17-18 and src/base/Connection.cpp:109,134-142 at et-v7.0.0). The hold is capped at twice ReplayLimit of written bytes, beyond which a later recover fails withErrReplayExceededas before. TheReplayLimitdoc now states the resulting bound (about three times ReplayLimit plus one packet).Test plan
TestCatchupKeptUntilServerSpeaks,TestServerPacketReleasesCatchup,TestRingTrimHoldCeiling,TestLivenessSlowUploadKeepsLink,TestLivenessTypingDoesNotHideDeadLink,TestLivenessLargeWritesDoNotHideDeadLink. Each was seen failing before the fix or with the production line it pins removed, including each branch ofprobeBehind.go test -race ./..., andgo test ./internal/etcp -race -count=20with no flakes.golangci-lint runon linux and windows,go fix -diff, js/wasm and wasip1 vet, e2e vet.BenchmarkDrainBacklogagainst main: no significant change (benchstat p=0.57).Summary by CodeRabbit