You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
iOS smoke lane: two untracked failure signatures from the #2491 attribution pass (MAIN_THREAD_TIMEOUT on alert dismiss; replay timeout_cleanup_pending) #3337
Gesture replay TIMEOUT with timeout_cleanup_pending — run 36115336398, failing step "Run gesture pan-duration smoke replay". examples/test-app/replays/gesture-pan-duration.ad, attempt 1/3: a wait command hit TIMEOUT after 180000ms (timeoutMode: cooperative, timeoutCleanupPending: true, junit wall 182s), and the attempt's runner log ends in ** BUILD INTERRUPTED **. The settings replay in the same run passed. A command timeout leaving cleanup pending is its own class — the retry policy in fix(test): make the iOS simulator smoke lane's red/green signal meaningful (#2491) #3336 deliberately does not re-issue hard timeouts.
Evidence
Both runs upload a single artifact named ios-artifacts:
gh run download 36115336398 --repo callstack/agent-device -n ios-artifacts → work/agent-device/agent-device/test/artifacts/replays-ios-gesture-pan-duration/3be9273812dd9d58/examples__test-app__replays__gesture-pan-duration.ad/attempt-1/failure.txt (plus replay-timing.ndjson; runner log under .tmp/agent-device-state/sessions/…gesture-pan-duration_attempt-1/runner.log)
Scope
Triage each signature to an owning defect (runner main-thread stall vs alert-dialog path for #1; timeout-cleanup lifecycle / cooperative-timeout teardown for #2), then make each red step assert its real contract. No retry-layer change is implied — #3336's policy intentionally leaves hard timeouts red.
Adding a third scenario-level signature per the #3336 coordinator pass (coordinating from #3343, which this supersedes for this signature).
3. smoke:webview-remote-content: wait text "Jump to form" 20000 → wait_deadline_exceeded with high readableCaptures — run 37843633972 (PR #3336 lane, 98321a279, 2026-10-08T21:16:46Z). Typed shape: timeoutMs: 20000, captures: 25, readableCaptures: 24, captureTruncated: true, diagnosticId: mv01do8x-5064e742. Every poll readable at 250-900 ms — the route answered the whole budget; the WKWebView page simply was not in the tree. The post-failure snapshot seconds later (test/artifacts/ios-simulator/smoke/1791493983644-15998/failed-step-102-snapshot.json) already containsJump to form x2, Account form, Email address under a healthy/tree verdict — the cold-WebContent load finished just past the deadline.
Lane-corroboration (not mine): the scenario predates this PR (#3084/#2902/#2493 on main); main was green on iOS the same day (run 37815426076); #3330 and #3334 pass Smoke on the same simulator UDID (5FEB61C5-…) and runner class with no cross-ref concurrency group, so host sharing is in play. The retry policy added in #3336 declined to re-issue this by design (wait_deadline_exceeded is named non-retriable: a readable miss is the product answering).
Decision rule for future occurrences (per the coordinator pass): this is a cold-host timing flake UNLESS it reproduces on main, or readableCaptures stays high across N runs with the page link never appearing in any capture — that second shape would make it a page-load-path product defect. The owning fix for a real latency class is the PAGE_LOAD_WAIT_MS budget or the page-load path, never widening the lane's retriable-reason set.
Signature 3 non-reproduction: re-run of 37843633972 @ 98321a2 (identical head, same simulator UDID) passed smoke:webview-remote-content — all checks green. Per the recorded decision rule this occurrence is the cold-host timing-flake shape: no main reproduction, and the original run's post-failure snapshot proved the content arrives seconds late. Keeping the signature filed here so a recurrence is matched rather than re-triaged.
Why
While attributing recent
ios.ymlmain failures for #2491 (see #2491 (comment)), two signatures had no owning issue:MAIN_THREAD_TIMEOUTonalert dismiss— run 37660079219,smoke:automation-inputscenario, step "dismiss native alert". Typed details:reason: runner_main_thread_timeout,dispatched: "unknown",runnerFailureReason: runner_main_thread_execution_timeout, hint says the runner abandoned the command's main-thread work past its execution watchdog. The finished-at-boundary misreport fix(ios-runner): runMainThreadWork reports MAIN_THREAD_TIMEOUT for work that finished at the timeout boundary #2782 is closed and does not cover this shape; looks like a genuine main-thread stall during alert dismissal. Adjacent prior art: iOS alert activation can tap after the command deadline #2956, iOS runner smoke: alert-observation XCTests are red or flaky on main #2546.TIMEOUTwithtimeout_cleanup_pending— run 36115336398, failing step "Run gesture pan-duration smoke replay".examples/test-app/replays/gesture-pan-duration.ad, attempt 1/3: awaitcommand hitTIMEOUT after 180000ms(timeoutMode: cooperative,timeoutCleanupPending: true, junit wall 182s), and the attempt's runner log ends in** BUILD INTERRUPTED **. The settings replay in the same run passed. A command timeout leaving cleanup pending is its own class — the retry policy in fix(test): make the iOS simulator smoke lane's red/green signal meaningful (#2491) #3336 deliberately does not re-issue hard timeouts.Evidence
Both runs upload a single artifact named
ios-artifacts:gh run download 37660079219 --repo callstack/agent-device -n ios-artifacts→work/agent-device/agent-device/test/artifacts/ios-simulator/smoke/1791396322253-59034/failed-step.txt(plusstep-history.json,failed-step-55.png, snapshot)gh run download 36115336398 --repo callstack/agent-device -n ios-artifacts→work/agent-device/agent-device/test/artifacts/replays-ios-gesture-pan-duration/3be9273812dd9d58/examples__test-app__replays__gesture-pan-duration.ad/attempt-1/failure.txt(plusreplay-timing.ndjson; runner log under.tmp/agent-device-state/sessions/…gesture-pan-duration_attempt-1/runner.log)Scope
Triage each signature to an owning defect (runner main-thread stall vs alert-dialog path for #1; timeout-cleanup lifecycle / cooperative-timeout teardown for #2), then make each red step assert its real contract. No retry-layer change is implied — #3336's policy intentionally leaves hard timeouts red.