Skip to content

fix(android): wait out a device-offline refusal and retry once - #3055

Merged
thymikee merged 7 commits into
callstack:mainfrom
okwasniewski:oskar/adb-waits-out-device-offline
Sep 30, 2026
Merged

thymikee merged 7 commits into
callstack:mainfrom
okwasniewski:oskar/adb-waits-out-device-offline

Conversation

@okwasniewski

@okwasniewski okwasniewski commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

An Android emulator can drop back to offline for a few seconds after sys.boot_completed while adbd restarts. In that window, the host adb refuses every command with adb: device offline before anything reaches the device. adb-failure.ts already classified this as device_offline and marked it retriable, but no caller ever acted on that.

In the e2e mobile benchmark, the first test after boot hit that window during the snapshot helper install (job):

APP_UNREACHABLE: snapshot failed: Android snapshot helper failed: Failed to install Android snapshot helper: device offline: adb: device offline

The device-scoped executor (createSerialAdbExecutor) now reacts to a device offline refusal, whether it comes back as a result or as a thrown error. It runs adb -s <serial> wait-for-device against the same adb server and environment, then retries the command once. Only the bare host refusal counts: stderr exactly adb: device offline or error: device offline, empty stdout, and no timeout. Output from a command that ran on the device never matches. The refusal comes from the host adb, so the command never ran and the retry cannot repeat a side effect. If the device is still offline, the second refusal surfaces with the existing classification and hint.

The retry is bounded so short probes keep their budgets:

  • A caller's timeoutMs covers the wait and the retry together. The wait takes at most half of what is left, and 15 s when the caller sets no timeout. With no budget left, the refusal stands. An abort during the wait skips the retry.
  • A caller without a timeout gets an unbounded retry, as its first attempt was.
  • A device that stays offline through its wait fails fast for 30 s instead of making every call wait again. The mark is keyed by adb server and serial, and the device's next answer clears it (a timeout or a cancel does not). Without this, a wedged device would cost 15 s per call to doctor, the boot uptime probe, and snapshot-helper retirement.

Touched: 2 source files (adb-provider-scope.ts, adb-failure.ts) and 2 test files.

Validation

  • Live, on a Pixel 7 API 36 emulator: adb -s emulator-5554 reconnect device followed immediately by a guarded shell true through createDeviceAdbExecutor. On main, 5/5 attempts fail with adb: device offline (11-13 ms). With this change, 6/6 succeed (440-560 ms) with and without a caller timeout, and so does the thrown-error path. This also confirms the real host serializes the waitFor-only invocation.

  • adb-provider-scope.test.ts covers:

    • an offline result retries after the wait (exact host calls asserted);
    • a thrown offline error gets the same single retry;
    • the wait and the retry share the caller's timeout;
    • an abort during the wait skips the retry;
    • a non-offline failure is not retried;
    • device output that mentions device offline is not retried;
    • a refusal with no budget left stands;
    • the wait uses the command's adb server and env;
    • a device that stays offline stops waiting until it answers again, and a timeout doesn't clear that.

    Each new assertion fails without the fix.

  • adb-executor.test.ts: the retriable-classification test now models a device that stays offline (the command, the wait, the retry).

  • pnpm check:affected --run: passed. pnpm test:coverage:ci: passed. Changed-line gate: passed (100%, 51/51).

Copilot AI balanced review requested due to automatic review settings September 29, 2026 11:14

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 4 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="packages/platform-android/src/adb-provider-scope.ts">

<violation number="1" location="packages/platform-android/src/adb-provider-scope.ts:129">
P2: When the caller passes no `timeoutMs`, the retry runs with `timeoutMs: undefined` — unbounded. The wait gets a 15s default (`ANDROID_DEVICE_OFFLINE_WAIT_MS`), so a device that answers the wait but then wedges (e.g. adbd still restarting, or an install held on an OEM confirmation dialog) holds the caller for 15s plus an unlimited second attempt — the kind of long hang the PR's fast-fail goal is meant to avoid. Give the no-timeout retry the same 15s default as the wait.</violation>
</file>

Tip: instead of fixing issues one by one fix them all with cubic

Re-trigger cubic

Comment thread packages/platform-android/src/adb-provider-scope.ts Outdated
Comment thread packages/platform-android/src/adb-provider-scope.ts Outdated
Comment thread packages/platform-android/src/adb-provider-scope.ts Outdated
Comment thread packages/platform-android/src/adb-provider-scope.ts
options?.signal?.throwIfAborted();
const retryOptions =
budgetMs === undefined
? options

@cubic-dev-ai cubic-dev-ai Bot Sep 29, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: When the caller passes no timeoutMs, the retry runs with timeoutMs: undefined — unbounded. The wait gets a 15s default (ANDROID_DEVICE_OFFLINE_WAIT_MS), so a device that answers the wait but then wedges (e.g. adbd still restarting, or an install held on an OEM confirmation dialog) holds the caller for 15s plus an unlimited second attempt — the kind of long hang the PR's fast-fail goal is meant to avoid. Give the no-timeout retry the same 15s default as the wait.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At packages/platform-android/src/adb-provider-scope.ts, line 129:

<comment>When the caller passes no `timeoutMs`, the retry runs with `timeoutMs: undefined` — unbounded. The wait gets a 15s default (`ANDROID_DEVICE_OFFLINE_WAIT_MS`), so a device that answers the wait but then wedges (e.g. adbd still restarting, or an install held on an OEM confirmation dialog) holds the caller for 15s plus an unlimited second attempt — the kind of long hang the PR's fast-fail goal is meant to avoid. Give the no-timeout retry the same 15s default as the wait.</comment>

<file context>
@@ -60,15 +60,105 @@ export function createDeviceAdbExecutor(
+  options?.signal?.throwIfAborted();
+  const retryOptions =
+    budgetMs === undefined
+      ? options
+      : { ...options, timeoutMs: Math.max(1, budgetMs - (Date.now() - startedAt)) };
+  const retry = await attemptAdb(serial, async () => await run(retryOptions));
</file context>
Suggested change
? options
? { ...options, timeoutMs: ANDROID_DEVICE_OFFLINE_WAIT_MS }
Fix with cubic

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[claude-opus-5-5] responding on behalf of @okwasniewski

Declining. A caller with no timeoutMs already runs its first attempt unbounded; the retry keeps that same contract, and bounding it would change behavior for callers that deliberately set none (long installs). The added cost is at most the 15 s wait, and only once per stayed-offline window. Leaving this open for the maintainer.

Comment thread packages/platform-android/src/adb-provider-scope.ts Outdated
Copilot AI review requested due to automatic review settings September 29, 2026 11:30

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@okwasniewski

Copy link
Copy Markdown
Contributor Author

[claude-opus-5-5] responding on behalf of @okwasniewski

The Coverage failure is unrelated to this diff: daemon-client-transport.test.ts > "a delayed restart health probe stops at the RPC deadline without retrying" (Remote daemon is unavailable instead of daemon_transport_timeout). Root cause is a client race (Node timers vs performance.now() at the RPC deadline), fixed in #3056. A maintainer rerun of Coverage should go green.

@thymikee

Copy link
Copy Markdown
Member

The reconnect fix in 290b28b looks right, but one test does not prove the main rule of this PR. The retry must share the caller's timeout with the wait, and no test checks the retry half of that.

In adb-provider-scope.test.ts, the wait-for-device stub resolves at once. So the retry timeout is 4000 or 3999 ms. If { ...options, timeoutMs: retryMs } at adb-provider-scope.ts:134 became plain options, the retry would get the full 4000 ms and toBeLessThanOrEqual(4_000) would still pass. A change like that could ship green, and total time could reach 1.5x the caller's timeout. Please make the wait stub use up a known amount of time, for example with fake timers or a Date.now stub that moves 1500 ms in the wait branch. Then assert that the retry's timeout is 4000 minus 1500 (or at most 2500), so that mutation fails the test.

Not blocking, and you can take or leave these. The offline-key at adb-provider-scope.ts:77 puts the per-call port first, while deviceServerPort uses installed ?? requested, so the key could be derived from the resolved route instead. The abort test at line 211 and the sticky-timeout test at line 311 use a bare .rejects.toThrow(), so asserting the error code and reason would make them stricter. devicesStayingOffline at adb-provider-scope.ts:103 is module-global and never drops expired entries, so deleting them on read and adding a test reset would help.

The live Pixel 7 API 36 reconnect run appears only in the PR body, with no log attached. It also called createDeviceAdbExecutor directly, not the snapshot-helper install route from the original incident. The managed-lease path (-P port -s serial wait-for-device) was checked by reading the code, not by running it. I did not run the tests, and the missed mutation comes from reading the stub's timing.

The only failing check is the daemon-client transport test 'a delayed restart health probe stops at the RPC deadline without retrying'. This PR changes only packages/platform-android, and that test does not touch those files, so it looks unrelated. There are no conflicts. Once the timeout test is fixed, the next step before merge is a maintainer rerun of the Coverage job.

An emulator drops to offline for a few seconds after boot while adbd
restarts. adb refuses commands on the host side then, so a helper
install could fail with 'adb: device offline'. The device-scoped
executor now waits for the device and runs the command once more.

The wait and retry stay inside the caller's timeoutMs (wait gets at
most half, 15s cap). A device that stays offline through its wait
fails fast for 30s instead of making every call wait.
…budget

- match only stderr that is exactly adb's device-offline refusal, empty
  stdout, no timeout, so device output never triggers a rerun
- wait takes half of the remaining budget; no budget left, refusal stands
- wait uses the command's adb server and env
- timeouts and cancels keep the stayed-offline mark
- key the mark by adb server and serial
Copilot AI review requested due to automatic review settings September 29, 2026 16:28
@okwasniewski
okwasniewski force-pushed the oskar/adb-waits-out-device-offline branch from 290b28b to 0be2e35 Compare September 29, 2026 16:28

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

…left

The wait stub now moves the clock 1500 ms, so the retry must get 4000 - 1500 ms;
a retry given the caller's full timeout fails the test. The abort and timeout
cases assert the error they reject with. The stayed-offline key takes the
route's installed adb server first, as the route itself does.
Copilot AI balanced review requested due to automatic review settings September 30, 2026 10:47

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

…fusal

The adb failure classifier now records hostRefusal (adbHostRefusal on thrown
errors) when stderr is exactly the host adb's device-offline refusal, stdout is
empty, and the command did not time out. The retry keys only on that detail, so
the separate stderr predicate is gone. The broad device_offline family is
unchanged. Stayed-offline marks move to their own module, which drops lapsed
marks when it records a new one.
Copilot AI balanced review requested due to automatic review settings September 30, 2026 11:13

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

* origin/main:
  0.21.17
  feat(daemon): report the host CPU architecture in /health (callstack#3048)
  feat: add daemon policy to confine devices, commands, and device shutdown (callstack#3064)
  test(web): wait for the killed fake daemon to be reaped before asserting it is gone (callstack#3066)
  fix(ios): write the simulator clipboard from the runner (callstack#3065)
  test(daemon-client): a restart probe that fails outright near the RPC deadline reports the daemon unavailable (callstack#3058)
  fix(daemon-client): a client whose daemon lost the start race adopts the winner (callstack#3057)
  fix(android): honor boot --timeout as the emulator boot deadline (callstack#3059)
  test(ios-smoke): wait once more when the runner is still starting behind a deep link (callstack#3063)

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 issues found across 5 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="packages/platform-android/src/adb-failure.test.ts">

<violation number="1" location="packages/platform-android/src/adb-failure.test.ts:30">
P3: The expected `hint` is derived from `classifyAndroidAdbFailure('device offline')` — the same function under test — so the hint assertion compares the implementation with itself and can never fail on wrong hint text. Both code paths also share the same `ANDROID_ADB_DEVICE_OFFLINE_FAILURE` constant, making the comparison structurally guaranteed. Inline the literal hint text (or drop the hint field) so the expected value is independent of the code it verifies.</violation>
</file>

<file name="packages/platform-android/src/adb-provider-scope.ts">

<violation number="1" location="packages/platform-android/src/adb-provider-scope.ts:123">
P2: The stayed-offline cooldown is a module-level map with no in-flight guard between the `has()` check and the `mark()` at the end of `retryOnceAfterDeviceOffline`. When two adb commands for the same device run concurrently while it is offline, both pass `devicesStayingOffline.has(device) === false` before either one marks, so both execute the full up-to-15s `wait-for-device` and both retry; the second `mark()` then re-arms the cooldown to a fresh 30s window on top of the first. The "fails fast" benefit only applies to commands arriving after both retries finish, so parallel e2e adb calls each pay the full wait during a boot-offline window. Consider keying an in-flight wait promise so concurrent refusals for the same server+serial join the running wait/retry instead of duplicating it.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

Comment thread packages/platform-android/src/adb-stayed-offline-devices.ts Outdated
): Promise<AndroidAdbExecutorResult> {
const startedAt = Date.now();
const first = await attemptAdb(device, async () => await run(options));
if (!first.offline || devicesStayingOffline.has(device)) {

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: The stayed-offline cooldown is a module-level map with no in-flight guard between the has() check and the mark() at the end of retryOnceAfterDeviceOffline. When two adb commands for the same device run concurrently while it is offline, both pass devicesStayingOffline.has(device) === false before either one marks, so both execute the full up-to-15s wait-for-device and both retry; the second mark() then re-arms the cooldown to a fresh 30s window on top of the first. The "fails fast" benefit only applies to commands arriving after both retries finish, so parallel e2e adb calls each pay the full wait during a boot-offline window. Consider keying an in-flight wait promise so concurrent refusals for the same server+serial join the running wait/retry instead of duplicating it.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/platform-android/src/adb-provider-scope.ts, line 123:

<comment>The stayed-offline cooldown is a module-level map with no in-flight guard between the `has()` check and the `mark()` at the end of `retryOnceAfterDeviceOffline`. When two adb commands for the same device run concurrently while it is offline, both pass `devicesStayingOffline.has(device) === false` before either one marks, so both execute the full up-to-15s `wait-for-device` and both retry; the second `mark()` then re-arms the cooldown to a fresh 30s window on top of the first. The "fails fast" benefit only applies to commands arriving after both retries finish, so parallel e2e adb calls each pay the full wait during a boot-offline window. Consider keying an in-flight wait promise so concurrent refusals for the same server+serial join the running wait/retry instead of duplicating it.</comment>

<file context>
@@ -118,7 +120,7 @@ async function retryOnceAfterDeviceOffline(
   const startedAt = Date.now();
   const first = await attemptAdb(device, async () => await run(options));
-  if (!first.offline || (devicesStayingOffline.get(device) ?? 0) > Date.now()) {
+  if (!first.offline || devicesStayingOffline.has(device)) {
     return first.outcome();
   }
</file context>
Fix with cubic

for (const stderr of ['adb: device offline\n', "error: device 'emulator-5554' offline"]) {
assert.deepEqual(classifyAndroidAdbFailure(stderr), {
reason: 'device_offline',
hint: classifyAndroidAdbFailure('device offline')?.hint,

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The expected hint is derived from classifyAndroidAdbFailure('device offline') — the same function under test — so the hint assertion compares the implementation with itself and can never fail on wrong hint text. Both code paths also share the same ANDROID_ADB_DEVICE_OFFLINE_FAILURE constant, making the comparison structurally guaranteed. Inline the literal hint text (or drop the hint field) so the expected value is independent of the code it verifies.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/platform-android/src/adb-failure.test.ts, line 30:

<comment>The expected `hint` is derived from `classifyAndroidAdbFailure('device offline')` — the same function under test — so the hint assertion compares the implementation with itself and can never fail on wrong hint text. Both code paths also share the same `ANDROID_ADB_DEVICE_OFFLINE_FAILURE` constant, making the comparison structurally guaranteed. Inline the literal hint text (or drop the hint field) so the expected value is independent of the code it verifies.</comment>

<file context>
@@ -23,6 +23,38 @@ test('ADB failure classification keeps transport on stderr and install verdicts
+  for (const stderr of ['adb: device offline\n', "error: device 'emulator-5554' offline"]) {
+    assert.deepEqual(classifyAndroidAdbFailure(stderr), {
+      reason: 'device_offline',
+      hint: classifyAndroidAdbFailure('device offline')?.hint,
+      retriable: true,
+      hostRefusal: true,
</file context>
Fix with cubic

Comment thread packages/platform-android/src/adb-stayed-offline-devices.test.ts Outdated
attemptAdb asks isOfflineRefusalResult or isOfflineRefusalError instead of
reading the classification inline, which keeps it under the fallow
complexity threshold.
Copilot AI balanced review requested due to automatic review settings September 30, 2026 11:21

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@thymikee

Copy link
Copy Markdown
Member

I found no code blockers at dd05dd9. The earlier findings from #3055 (comment) are fixed, and I found no new problems in the code.

Not blocking, and you can take or leave these: the expected hint in https://github.com/callstack/agent-device/blob/dd05dd9/packages/platform-android/src/adb-failure.test.ts#L30 comes from classifyAndroidAdbFailure('device offline')?.hint, so the test compares the hint with itself and a literal string would be stronger; and the PR body still says "Touched: 2 source files", while the latest change also added adb-stayed-offline-devices.ts and the adbHostRefusal error detail.

I did not run the tests, and the live emulator run exists only in the PR body with no log, so it is not independently confirmed. That run also called createDeviceAdbExecutor directly, not the snapshot-helper install route from the incident.

Smoke Tests passed on dd05dd9. Coverage fails, and this PR causes it: the eager-closure budget test reports that packages/platform-android/src/mechanics.ts now evaluates 177 modules on import, one more than the merge-base (log). The new adb-stayed-offline-devices.ts module is the likely extra static edge. Please load it lazily or fold it into an existing module so the budget holds. There are no conflicts.

@thymikee thymikee added the ready-for-human Valid work that needs human implementation, judgment, or maintainer merge label Sep 30, 2026
… scope

A separate module added one module to the eager import closure of
platform-android mechanics.ts, which failed the eager-closure budget.
Copilot AI balanced review requested due to automatic review settings September 30, 2026 11:38

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@thymikee

Copy link
Copy Markdown
Member

The PR is ready. The fixes since the earlier review (dd05dd9) look good at 192f759, and no code changes are needed. The Coverage failure was the eager-closure budget (177 vs 176 modules for mechanics.ts), and this change fixes it by removing the separate module and its static edge. All 13 checks are passing now.

Not blocking: createStayedOfflineDevices and StayedOfflineDevices are now exported from the provider-scope module and only adb-provider-scope.test.ts uses them, so they are test-only exports of a production module; you can test the window through the retry route or leave it as it is.

I did not run the tests or the eager-closure budget test locally. The budget fix is inferred from the removed import edge and the green checks. This change does not touch adb-failure.ts, which the earlier review already covered. The emulator run is still described only in the PR body, and since this change is a pure in-package move, it needs no new device evidence. I know of no conflicts, and the next step is a maintainer merge decision.

@thymikee
thymikee merged commit eb51a88 into callstack:main Sep 30, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-for-human Valid work that needs human implementation, judgment, or maintainer merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants