Skip to content

test(ios-smoke): wait once more when the runner is still starting behind a deep link - #3063

Merged
thymikee merged 1 commit into
callstack:mainfrom
okwasniewski:oskar/ios-smoke-waits-out-runner-start
Sep 29, 2026
Merged

thymikee merged 1 commit into
callstack:mainfrom
okwasniewski:oskar/ios-smoke-waits-out-runner-start

Conversation

@okwasniewski

@okwasniewski okwasniewski commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

The iOS smoke test live iOS simulator fixture E2E keeps failing at wait for Automation lab, right after open --relaunch --launch-url agent-device-test-app:///automation?.... The final error is app 'com.callstack.agentdevicelab' is not running (wait_capture_stalled, APP_NOT_RUNNING). The failure screenshot shows iOS's "Open in "Agent Device Tester"?" confirmation, unanswered.

It hit #3056 once and #3057 twice today. In #3057's run the uploaded runner.log shows the cause:

  1. The deep-link open succeeds at 14:51:37. The relaunch restarts the runner, and xcodebuild launches it at 14:51:33.
  2. The helper's first 15 s destination wait (14:51:37 to 14:51:52) ends before the runner is up: Running tests at 14:51:54, LISTENER_READY at 14:51:55. That is runner-start readiness exhaustion.
  3. answerDeepLinkConfirmation returns on runner-start exhaustion without probing (test(ios-smoke): answer deep-link prompt after readiness exhaustion #3044 gated the probe to target-discovery), so the prompt is never answered.
  4. The caller's next wait reaches the now-running runner, which answers READ_TARGET_NOT_RUNNING on every read (14:51:55 to 14:52:14) until the wait times out.

The helper now waits once more after a runner-start timeout, without an alert probe. The next wait is the first that can see the pending launch: it ends on APP_NOT_RUNNING, and the existing path answers the confirmation. A second runner-start timeout still returns without probing. That keeps #3044's point: a runner that really failed costs one extra 15 s wait, not the whole 5-wait budget with probes. The miss classification moves into classifyDestinationMiss (fallow complexity); its outcomes are unchanged for every other reason.

Two test-harness files changed; production behavior is unchanged.

Validation

  • ios-simulator-e2e-deep-link-confirmation.test.ts (13 pass):

    • a runner-start timeout, then a pending launch, gives wait, wait, alert get, alert accept, wait;
    • two runner-start timeouts return after the second wait with no probe.

    Both fail on main.

  • pnpm check:affected --run: passed, including fallow.

Review in cubic

…ind a deep link

A relaunch can restart the runner, and on a loaded CI Mac its start
outlasts the first 15 s destination wait. That wait ends in runner-start
readiness exhaustion, which returned without probing, so iOS's 'Open in
Agent Device Tester?' confirmation stayed unanswered and the caller's
wait failed on APP_NOT_RUNNING. One more wait lets the started runner
see the pending launch and the helper answer it; a second runner-start
timeout still returns. The miss classification moves into its own
function.
Copilot AI balanced review requested due to automatic review settings September 29, 2026 16:28

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 2 files

Re-trigger cubic

@thymikee

Copy link
Copy Markdown
Member

The code looks right at 9b3995e, and all 14 checks pass. I found no conflicts, so this is ready for human review.

Not blocking: live-deep-link-confirmation.ts adds a local DestinationMissDetails type and a readinessExhaustedPhase helper, and it compares against the raw string 'runner-start'. @agent-device/contracts/wait already exports readinessPhaseOf and ReadinessPhase, so you could use readinessPhaseOf(details) behind a reason === WAIT_REASONS.readinessExhausted check. The new inline comment could also be shorter, with the one-wait rule moved into the doc block of answerDeepLinkConfirmation. Take or leave both.

I did not run the unit tests locally, so their result on the PR head and on main comes from reading the code. I could not read the #3057 runner log, so the timeline of the runner starting after the first 15 s wait comes from the PR body. CI does exercise this helper, but it does not show that this change removes the flake. Whether one extra 15 s wait is enough on a loaded CI host is best judged from the next few iOS smoke runs.

@thymikee thymikee added the ready-for-human Valid work that needs human implementation, judgment, or maintainer merge label Sep 29, 2026
@thymikee
thymikee merged commit 5bb104d into callstack:main Sep 29, 2026
14 checks passed
thymikee added a commit to okwasniewski/agent-device that referenced this pull request Sep 30, 2026
* origin/main:
  0.21.17
  feat(daemon): report the host CPU architecture in /health (callstack#3048)
  feat: add daemon policy to confine devices, commands, and device shutdown (callstack#3064)
  test(web): wait for the killed fake daemon to be reaped before asserting it is gone (callstack#3066)
  fix(ios): write the simulator clipboard from the runner (callstack#3065)
  test(daemon-client): a restart probe that fails outright near the RPC deadline reports the daemon unavailable (callstack#3058)
  fix(daemon-client): a client whose daemon lost the start race adopts the winner (callstack#3057)
  fix(android): honor boot --timeout as the emulator boot deadline (callstack#3059)
  test(ios-smoke): wait once more when the runner is still starting behind a deep link (callstack#3063)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-for-human Valid work that needs human implementation, judgment, or maintainer merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants