fix(daemon-client): a restart probe cut short by the RPC deadline reports the timeout - #3056
Conversation
…orts the timeout Node timers start from the event loop's cached clock, so the restart health probe's timer can expire while performance.now() is still short of the RPC deadline. The retry then threw 'Remote daemon is unavailable' instead of the request timeout, which flaked 'a delayed restart health probe stops at the RPC deadline without retrying' on CI. The health probe now marks a result that ran out of time, and the restart retry reports a probe the deadline capped that timed out as the RPC timing out.
There was a problem hiding this comment.
All reported issues were addressed across 3 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
|
[claude-opus-5-5] responding on behalf of @okwasniewski iOS Smoke Tests failure on the test-pin commit is unrelated: |
|
I reviewed 83cc5a7. A restart probe cut short by the RPC deadline now reports the timeout, and I found no problems in the change. Not blocking: the I did not run the new test on main or on this head, so the claim that it fails on main rests on reading the old code with a frozen CI is still pending, because Smoke Tests was running when I checked. The change only touches the remote instance-mismatch restart retry, and smoke runs drive a local daemon, so I expect no overlap. There are no conflicts. Once Smoke Tests finishes green, this is ready for human review. |
|
Smoke Tests has now failed on 83cc5a7, in the live iOS simulator step "wait for Automation lab" (job). That run drives a local daemon and does not reach the remote instance-mismatch retry this PR changes, so it looks unrelated. The code verdict is unchanged; a rerun of Smoke Tests should confirm it. |
Summary
daemon-client-transport.test.ts> "a delayed restart health probe stops at the RPC deadline without retrying" flakes on CI. It failed the Coverage job on #3055 withRemote daemon is unavailablewhere it expectsdaemon_transport_timeout, and it fails the same way locally under load.This is a real race in the client, not just in the test. After a
remote_instance_mismatch,retryAfterRemoteInstanceMismatchprobes/healthwith exactly the RPC's remaining budget, then checksperformance.now()against the deadline. Node timers start from the event loop's cached clock, which can lag behind real time. So the probe's timer can expire whileperformance.now()is still short of the deadline. The remaining budget reads positive, and the client reports the daemon as unavailable when the RPC actually timed out.The fix:
readDaemonHttpHealthnow marks a result that ran out of time (timedOut: true), from either its socket timeout or its abort signal.Remote daemon is unavailable.Wire ledger:
RemoteDaemonHealthandreadDaemonHttpHealthgot new digests and new acks.timedOutis client-local: it is never read from a/healthpayload, and the request and the accepted payload fields are unchanged. The two acks it replaces had expired with the digests.Validation
performance.now(), so the probe's timeout lands before the deadline. Onmainit fails with the exact CI error (Remote daemon is unavailable); with this change it passes.daemon-client-transport.test.ts20x in a row with oneyesper core: 20/20 green.pnpm check:affected --run: passed.pnpm test:coverage:ci: passed. Changed-line gate: passed (100%, 10/10).