Skip to content

fix(daemon-client): report the caller deadline when the restart health probe runs to it - #3053

Closed
thymikee wants to merge 3 commits into
mainfrom
fix/deflake-delayed-health-probe
Closed

thymikee wants to merge 3 commits into
mainfrom
fix/deflake-delayed-health-probe

Conversation

@thymikee

@thymikee thymikee commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Summary

Root cause: after an instance mismatch, retryAfterRemoteInstanceMismatch probes /health with the remaining RPC budget. The probe's timer can fire a few ms before performance.now() reaches the deadline. The remaining-time check then saw a sliver of budget, so the failed probe was classified as Remote daemon is unavailable instead of daemon_transport_timeout. A loaded CI runner hits this and a fast local machine does not.

Fix: the health probe now reports which event ended it. Its own budget timer sets a typed timedOut outcome (internal to daemon-client). The retry throws the request timeout only when the probe timed out and its budget was the caller's remaining time. Refusals, resets, aborts and 5xx still report Remote daemon is unavailable. The public return shapes of readRemoteDaemonHealth and RemoteDaemonHealth are unchanged; two wire-ledger digests were re-pinned (health request and payload are unchanged).

Found on unrelated PR CI (#3014, #3025, #3029, #3033, #3036).

Validation

  • Commit 3c5d77b
  • daemon-client-transport.test.ts passed 20/20 in a loop.
  • pnpm check:affected --run: all runnable checks passed (including daemon-wire-compat).
  • Tests: the early-timer test skews performance.now by 50 ms and asserts /health was reached; it fails on the previous code. A new test has /health return 503 near the deadline and expects Remote daemon is unavailable; it fails if any failed probe near the deadline is treated as a timeout.

Review in cubic

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Size Report

Metric Base Current Diff
Installed (including dependencies) 4.88 MB 4.88 MB +375 B
Package (unpacked) 4.88 MB 4.88 MB +375 B
Package (download) 1.46 MB 1.46 MB +116 B

Startup median (7 runs, lower is better):

Scenario Base Current Diff
CLI --version 26.1 ms 26.6 ms +0.5 ms
CLI --help 80.8 ms 81.1 ms +0.3 ms

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 issues found across 2 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="src/daemon-client/daemon-client-transport.ts">

<violation number="1" location="src/daemon-client/daemon-client-transport.ts:293">
P2: For a remaining budget below 5 ms, this threshold is negative, so an immediately failed health probe is reported as `daemon_transport_timeout` instead of `Remote daemon is unavailable`. Do not apply the timer slop when the probe budget is at most the slop window.</violation>
</file>

<file name="src/daemon-client/__tests__/daemon-client-transport.test.ts">

<violation number="1" location="src/daemon-client/__tests__/daemon-client-transport.test.ts:240">
P3: This regression test only distinguishes fixed code from unfixed code while the health-probe timer fires within ~2 ms of its nominal deadline. Without the fix the misclassification depends on `remainingMs = deadline - performance.now()`, which at timer-fire time equals `2 - δ` where δ is the timer's real fire latency; any δ ≥ 2 ms (a loaded CI, exactly where this bug was observed) makes the old code throw `daemon_transport_timeout` and the test passes against the buggy implementation. Test-assert `probing === true` after the rejects so the skew demonstrably engaged, and widen the simulated clock skew to ~4 ms (still safely under the 5 ms slop, so the fix path keeps triggering via `elapsed_seen ≥ probeTimeoutMs - 5`) to give the discrimination real margin.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

if (probeTimeoutMs === undefined || probeTimeoutMs > REMOTE_DAEMON_HEALTHCHECK_TIMEOUT_MS) {
return false;
}
return performance.now() - probeStartedAt >= probeTimeoutMs - PROBE_TIMER_SLOP_MS;

@cubic-dev-ai cubic-dev-ai Bot Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: For a remaining budget below 5 ms, this threshold is negative, so an immediately failed health probe is reported as daemon_transport_timeout instead of Remote daemon is unavailable. Do not apply the timer slop when the probe budget is at most the slop window.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At src/daemon-client/daemon-client-transport.ts, line 293:

<comment>For a remaining budget below 5 ms, this threshold is negative, so an immediately failed health probe is reported as `daemon_transport_timeout` instead of `Remote daemon is unavailable`. Do not apply the timer slop when the probe budget is at most the slop window.</comment>

<file context>
@@ -272,6 +280,19 @@ async function retryAfterRemoteInstanceMismatch(
+  if (probeTimeoutMs === undefined || probeTimeoutMs > REMOTE_DAEMON_HEALTHCHECK_TIMEOUT_MS) {
+    return false;
+  }
+  return performance.now() - probeStartedAt >= probeTimeoutMs - PROBE_TIMER_SLOP_MS;
+}
+
</file context>
Suggested change
return performance.now() - probeStartedAt >= probeTimeoutMs - PROBE_TIMER_SLOP_MS;
return (
probeTimeoutMs > PROBE_TIMER_SLOP_MS &&
performance.now() - probeStartedAt >= probeTimeoutMs - PROBE_TIMER_SLOP_MS
);
Fix with cubic

});
const clock = vi
.spyOn(performance, 'now')
.mockImplementation(() => realNow() - (probing ? 2 : 0));

@cubic-dev-ai cubic-dev-ai Bot Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: This regression test only distinguishes fixed code from unfixed code while the health-probe timer fires within ~2 ms of its nominal deadline. Without the fix the misclassification depends on remainingMs = deadline - performance.now(), which at timer-fire time equals 2 - δ where δ is the timer's real fire latency; any δ ≥ 2 ms (a loaded CI, exactly where this bug was observed) makes the old code throw daemon_transport_timeout and the test passes against the buggy implementation. Test-assert probing === true after the rejects so the skew demonstrably engaged, and widen the simulated clock skew to ~4 ms (still safely under the 5 ms slop, so the fix path keeps triggering via elapsed_seen ≥ probeTimeoutMs - 5) to give the discrimination real margin.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At src/daemon-client/__tests__/daemon-client-transport.test.ts, line 240:

<comment>This regression test only distinguishes fixed code from unfixed code while the health-probe timer fires within ~2 ms of its nominal deadline. Without the fix the misclassification depends on `remainingMs = deadline - performance.now()`, which at timer-fire time equals `2 - δ` where δ is the timer's real fire latency; any δ ≥ 2 ms (a loaded CI, exactly where this bug was observed) makes the old code throw `daemon_transport_timeout` and the test passes against the buggy implementation. Test-assert `probing === true` after the rejects so the skew demonstrably engaged, and widen the simulated clock skew to ~4 ms (still safely under the 5 ms slop, so the fix path keeps triggering via `elapsed_seen ≥ probeTimeoutMs - 5`) to give the discrimination real margin.</comment>

<file context>
@@ -222,6 +222,35 @@ test('a delayed restart health probe stops at the RPC deadline without retrying'
+  });
+  const clock = vi
+    .spyOn(performance, 'now')
+    .mockImplementation(() => realNow() - (probing ? 2 : 0));
+  try {
+    const port = await listenOnLoopback(server);
</file context>
Fix with cubic

@thymikee

Copy link
Copy Markdown
Member Author

Thanks for the fix. At 7566714 the restart probe still guesses "timeout" from elapsed time, so the flake is narrowed but not removed, and a real outage can now be reported as a timeout. There are no conflicts.

The check at https://github.com/callstack/agent-device/blob/7566714/src/daemon-client/daemon-client-transport.ts#L293 never asks which event ended the probe. readDaemonHttpHealth (lines 168-177) returns the same {reachable:false} for a timeout, ECONNREFUSED, a reset, an abort and a 5xx. So with a budget under 5 ms, or a refusal or 503 in the last 5 ms, a down daemon gets a retriable daemon_transport_timeout instead of "Remote daemon is unavailable" with daemonBaseUrl. In the other direction, a timer that fires more than 5 ms early still gives "unavailable", so the CI flake can come back under load. The rule to satisfy: a failed restart probe is a caller-deadline timeout if and only if the probe's own budget timer ended it and that budget was the caller's remaining time. The owner of that fact is readDaemonHttpHealth. Its timeout handler and its AbortSignal.timeout abort (check signal.aborted in the error handler, not the error text) should set a typed field on RemoteDaemonHealth, for example timedOut: true. retryAfterRemoteInstanceMismatch then throws handleRequestTimeout when health.timedOut && probeTimeoutMs !== undefined && probeTimeoutMs <= REMOTE_DAEMON_HEALTHCHECK_TIMEOUT_MS. That lets you delete PROBE_TIMER_SLOP_MS, the probeStartedAt read and probeRanToCallerDeadline.

The new test at https://github.com/callstack/agent-device/blob/7566714/src/daemon-client/__tests__/daemon-client-transport.test.ts#L240 does not reliably separate old code from new code. The 2 ms skew only helps when the timer fires less than 2 ms late; at 2 ms or more, the old code already throws daemon_transport_timeout, so the test can pass without the fix. It also never asserts that /health was reached. Once the typed timedOut signal exists, please use a large skew (for example 50 ms) so the old code deterministically throws "unavailable". Please also add a negative test where /health returns 503, or the server is closed, with a small remaining budget. It should assert the message "Remote daemon is unavailable" and no daemon_transport_timeout reason. Assert that /health was reached in both tests.

Not blocking: timeoutMs ?? 0 at https://github.com/callstack/agent-device/blob/7566714/src/daemon-client/daemon-client-transport.ts#L263 is a dead fallback, and the new throw duplicates the handleRequestTimeout block in remainingRemoteRequestTimeoutMs (lines 306-310), so a small shared helper or dropping the fallback would tidy both, but take or leave it.

The failing Smoke Tests job looks unrelated to this diff. It is the iOS simulator run against a local daemon with no base URL, so it never reaches the changed branch, and it failed with wait_readiness_exhausted in target discovery. I did not run the tests locally or open the smoke artifacts. The timer-lateness reasoning above is from reading the code, not from measurement. Once the probe is classified from the typed timeout signal and both tests are in, this is ready for another look.

@thymikee
thymikee force-pushed the fix/deflake-delayed-health-probe branch from 7566714 to 3c5d77b Compare September 29, 2026 10:16
@thymikee

Copy link
Copy Markdown
Member Author

Thanks for the review. Pushed 3c5d77b.

  • Probe guessed a timeout from elapsed time: the health probe now records whether its own timer ended it (typed outcome inside daemon-client, not on the wire type). The retry throws the request timeout only when the probe timed out and its budget was the caller's remaining time. Refusals, resets, aborts and 5xx report 'Remote daemon is unavailable'. The slop constant and elapsed-time check are gone, which also covers the sub-5 ms budget case. (0809263)
  • Wire ledger: the two pinned functions changed, so I re-pinned only their digests, noting the health request and payload are unchanged. (3c5d77b)
  • Early-timer test: now skews by 50 ms and asserts /health was reached. It fails on the previous code.
  • Negative test: /health returns 503 near the deadline and the error stays 'Remote daemon is unavailable' with no timeout reason; /health reached is asserted.
  • Nits: dropped the dead timeoutMs fallback and shared one timeout-error helper.
  • Smoke Tests: checking now.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 3 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="src/daemon-client/daemon-client-transport.ts">

<violation number="1" location="src/daemon-client/daemon-client-transport.ts:141">
P2: This refactor moves the wire-relevant health request tokens out of the ledger-covered `readDaemonHttpHealth` declaration. Previously `buildDaemonHttpUrl`, `transport.request({method: 'GET', ...})`, the `statusCode < 500` threshold, and body-read/abort handling all lived inside `readDaemonHttpHealth`'s digested body; now they live in `probeDaemonHttpHealth`/`collectHealthResponse`, neither of which is listed in `test/wire-compat/surface.ts`'s /health consumer group. The re-pinned digest for `readDaemonHttpHealth` now covers only a delegation, so a future wire-breaking change to the health request shape inside these helpers will no longer move the ledger digest — coverage the /health boundary previously had is silently lost. Add the new helpers to the surface (with a `wire-mutations` proof if they introduce a new break class) or note the moved tokens as reviewer-owned `uncovered`, per the wire-compat README's 'never claim coverage you don't provide'.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

return (await probeDaemonHttpHealth(info, probeTimeoutMs)).health;
}

async function probeDaemonHttpHealth(

@cubic-dev-ai cubic-dev-ai Bot Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This refactor moves the wire-relevant health request tokens out of the ledger-covered readDaemonHttpHealth declaration. Previously buildDaemonHttpUrl, transport.request({method: 'GET', ...}), the statusCode < 500 threshold, and body-read/abort handling all lived inside readDaemonHttpHealth's digested body; now they live in probeDaemonHttpHealth/collectHealthResponse, neither of which is listed in test/wire-compat/surface.ts's /health consumer group. The re-pinned digest for readDaemonHttpHealth now covers only a delegation, so a future wire-breaking change to the health request shape inside these helpers will no longer move the ledger digest — coverage the /health boundary previously had is silently lost. Add the new helpers to the surface (with a wire-mutations proof if they introduce a new break class) or note the moved tokens as reviewer-owned uncovered, per the wire-compat README's 'never claim coverage you don't provide'.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At src/daemon-client/daemon-client-transport.ts, line 141:

<comment>This refactor moves the wire-relevant health request tokens out of the ledger-covered `readDaemonHttpHealth` declaration. Previously `buildDaemonHttpUrl`, `transport.request({method: 'GET', ...})`, the `statusCode < 500` threshold, and body-read/abort handling all lived inside `readDaemonHttpHealth`'s digested body; now they live in `probeDaemonHttpHealth`/`collectHealthResponse`, neither of which is listed in `test/wire-compat/surface.ts`'s /health consumer group. The re-pinned digest for `readDaemonHttpHealth` now covers only a delegation, so a future wire-breaking change to the health request shape inside these helpers will no longer move the ledger digest — coverage the /health boundary previously had is silently lost. Add the new helpers to the surface (with a `wire-mutations` proof if they introduce a new break class) or note the moved tokens as reviewer-owned `uncovered`, per the wire-compat README's 'never claim coverage you don't provide'.</comment>

<file context>
@@ -125,21 +135,30 @@ async function readDaemonHttpHealth(
+  return (await probeDaemonHttpHealth(info, probeTimeoutMs)).health;
+}
+
+async function probeDaemonHttpHealth(
+  info: DaemonInfo,
+  probeTimeoutMs?: number,
</file context>
Fix with cubic

@thymikee

Copy link
Copy Markdown
Member Author

Thanks, this resolves the earlier finding at 3c5d77b. readDaemonHttpHealth now reports a typed timedOut from its own timer and abort signal, and the restart path throws the caller-deadline timeout only when that probe ran on the caller's remaining time. PROBE_TIMER_SLOP_MS and the elapsed-time guess are gone. The wire-compat ledger change only renames the recorded probe outcome and matches the new field. Smoke Tests was still running when I checked. There are no conflicts.

@thymikee thymikee added the ready-for-human Valid work that needs human implementation, judgment, or maintainer merge label Sep 29, 2026
@thymikee

Copy link
Copy Markdown
Member Author

Small correction to my previous comment on 3c5d77b, with the full notes. The verdict does not change.

This PR is ready at 3c5d77b. The fixes from the earlier review at 7566714 are in, and I found no remaining problems in the code.

Not blocking, and you can take or leave these: isCallerDeadlineProbeBudget re-derives which bound won min(REMOTE_DAEMON_HEALTHCHECK_TIMEOUT_MS, probeTimeoutMs), a choice probeDaemonHttpHealth already makes at line 154, and the result leaves readRemoteDaemonHealth through the onProbeTimedOut callback. The probe could report which budget ended it, for example timedOut: 'caller' | 'cap' | false, from an internal variant that returns it. Also, both amended rationales in ledger.json now say the health request and payload are unchanged twice, so keeping only the appended clause about the internal timer outcome would read better.

I did not run the tests or the wire-compat gate. The digests and the loopback test behavior come from reading the code only. I also did not check that Node emits 'error' or 'aborted' with signal.aborted set on every platform when a body read stalls. The req 'timeout' path covers the idle case on its own.

Smoke Tests is still in progress with no failure so far. The change reaches the remote-daemon instance-mismatch retry, which needs a baseUrl. canConnectHttp only goes through the refactored probe and its result does not change. Smoke runs a local daemon, so it does not reach the retry path. There are no conflicts. Smoke Tests must finish green before merge.

@thymikee

Copy link
Copy Markdown
Member Author

Closing: #3056 landed the same fix on main (typed timeout on the restart health probe, with the ledger re-pin), so this PR is superseded.

@thymikee thymikee closed this Sep 29, 2026
@thymikee
thymikee deleted the fix/deflake-delayed-health-probe branch September 29, 2026 12:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-for-human Valid work that needs human implementation, judgment, or maintainer merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant