You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Test runs on top of agent-device failed in ways nobody could act on. An action timed out, the framework retried it, and the app got two taps. A replayed press fired a render before its button existed and failed, while the same step passed by hand. The error said Element not found or timed out, and the only way to decide "retry or not" was to match message text.
Consumer data (tester-army e2e, mobile.yml, 723 runs, one week): 303 failures from snapshot quality, 111 from session lifecycle, 88 from gestures, 57 from input. The dominant single flake (Login Form, 42 of 68 first-pass failures) is fixed by #3051 and released in v0.21.17. This issue is about the class that remains: actions that run twice, and taps that run too early.
What we get
One answer to "is a resend safe?" A failed interaction carries error.details.dispatched: no (nothing reached the device) or unknown (it may have landed). Consumers retry on no, look first on unknown, and stop matching message text. We stop guessing inside the daemon too. There is no yes: no backend can prove execution on its failure path (the runner's own XCTEST_RECORDED_FAILURE says "may not have been performed"), so a third value would only ever be a false claim.
No double actions, on any backend. Every internal retry keys on that same proof. Twelve WebDriver routes and the Maestro direct iOS tap resent blindly; they no longer do. The remaining blind resends (daemon no-change retry, runner restart) are tracked below.
Scripted taps absorb a render race. A replayed press, click, or longpress waits up to 2 seconds for a target that is not on screen yet, then taps once. Agents and the interactive CLI keep one-shot behavior.
A table that fails CI when a backend lies. A row is one way a backend can fail plus the dispatched value that failure must carry (ios-runner.reply.runner-busy → no; a read-only command such as get is always no, because a read has nothing to repeat). For every row, a test makes the real code produce that failure and asserts the value. A row without a test, or a test without a row, fails the gate. When a backend changes what it can prove, the gate says so before a user does.
Every platform and provider inherits this. The three behaviors live above the backends: the dispatched seam, the read-only rule, and the readiness wait run in the daemon; the polling engine lives in capture-kit. A new backend gets unknown on every unclassified failure, no on every read, and the readiness wait, with zero backend code. It adds rows only for the refusals it can prove.
Cross-platform: what is inherited, what is declared
flowchart TB
subgraph consumers [Consumers]
E2E[e2e frameworks]
AG[agents via MCP / CLI]
SDK[Node client]
end
subgraph daemon [Daemon: inherited by every backend]
RD[Readiness wait<br/>replay + Node client, 2 s cap]
DS[dispatched seam<br/>fills unknown once, never overwrites<br/>9 daemon rows]
RO[Read-only rule<br/>registry trait → dispatched: no]
EN[observeUntil engine<br/>capture-kit, one deadline rule]
end
subgraph backends [Backends: declare rows, prove them]
IOS[Apple runner<br/>iOS, tvOS, macOS, visionOS<br/>25 rows]
AND[Android<br/>adb + helper<br/>9 rows]
WEB[WebDriver<br/>3 rows]
NEW[HarmonyOS, Vega, Linux,<br/>next provider<br/>0 rows = all unknown, still safe]
end
consumers --> daemon --> backends
G[(dispatch-disclosure.json<br/>one driver test per producer, gate)]
G -.enforces.- IOS
G -.enforces.- AND
G -.enforces.- WEB
Loading
A backend declares a row when it can prove the action never reached the device (no: a selector miss, a covered target, a runner RUNNER_BUSY, an adb host refusal). Everything else stays unknown, which is safe by construction: a wrong no makes a consumer resend an action that ran; a wrong unknown costs one snapshot. A multi-step action (press --count 25, chunked typing) says no only when its first step was refused; once any step reached the device, a later refusal is unknown.
The readiness wait and the engine need nothing from a backend. They consume the same snapshot a backend already serves.
Performance
Hit path: one capture, same as today. With the readiness option set, a target present on the first capture is tapped after that capture; measured on a private iPhone 17 Pro simulator, every one of 46 clicks resolved in one poll (waitedMs median 245 ms, which is the capture itself). Per-run wall time was 34.5 s against 36.5 s p50 on a loaded host over 12 pairs, inside the load noise; the soak issue below measures it properly and has a 3% keep rule.
Miss path: bounded at 2 s, and only where it is set. MCP and the interactive CLI never wait; an agent that misses on a wrong selector gets the fast error it gets today.
Failure path: cheaper. A WebDriver mutation that timed out used to pay a second full request timeout before failing; it fails after one. dispatched is a field on an error, no round trip.
Engine: no extra polls. The budget bounds when a poll starts, and a capture in flight is never abandoned mid-way; cancellation joins it. The two loops moved onto the engine (post-gesture stability, scroll movement) keep their cadence and verdicts.
Request envelope grows only by the budget, on client and daemon, so a replayed step never times out on the client while the daemon is still inside its wait.
Before and after: what a failure tells us
sequenceDiagram
participant T as Test framework
participant D as Daemon
participant B as Backend (runner, adb, WebDriver)
T->>D: press "Continue"
D->>B: tap
B--xD: reply lost
Note over D: before: error text "timed out"
D-->>T: error (text only)
T->>T: regex on message → retry
T->>D: press "Continue"
D->>B: tap
Note over B: the app got two taps
Loading
sequenceDiagram
participant T as Test framework
participant D as Daemon
participant B as Backend (runner, adb, WebDriver)
T->>D: press "Continue"
D->>B: tap
B--xD: reply lost
Note over D: after: details.dispatched = "unknown"
D-->>T: error + dispatched
T->>D: snapshot
D-->>T: the next screen is up
Note over T: no retry, the step passed
Loading
error.details.dispatched
Meaning
What a consumer does
"no"
Nothing reached the device
Retry as is
"unknown"
It may have landed
Snapshot before any retry
absent
Not classified
Treat as unknown
Before and after: where retries live
flowchart LR
subgraph before [Before: five retry sites, five rules]
W[WebDriver transport<br/>resends 12 mutating routes on timeout]
M[Maestro direct iOS tap<br/>re-taps on 4 message strings]
N[Daemon no-change retry<br/>resends when the tree looks the same]
R[Runner restart<br/>resends after a lost reply]
A[adb<br/>marks offline retriable, nobody acts]
end
Loading
flowchart LR
subgraph after [After: one rule, one table]
P[Producer classifies its own failure<br/>no / unknown]
G[(dispatch-disclosure.json<br/>46 rows, one driver test per producer)]
C{Retry only when<br/>dispatched = no}
P --> G --> C
C -->|WebDriver| W2[zero retries on mutations]
C -->|Maestro direct tap| M2[falls back only on no]
C -->|adb host refusal| A2[retries once, offline proven pre-dispatch]
C -->|daemon| N2[no-change retry deleted]
C -->|runner| R2[restart resend needs proof]
end
Loading
Before and after: a replayed tap
stateDiagram-v2
[*] --> Capture
Capture --> Tap: target present
Capture --> Fail: target absent
Tap --> [*]
Fail --> [*]: selector_not_found
note right of Fail
before: one look, one chance
end note
Loading
stateDiagram-v2
[*] --> Capture
Capture --> Tap: target present
Capture --> FailNow: covered, off-screen, ambiguous
Capture --> Wait: target absent
Wait --> Capture: every 200 ms, up to 2 s
Wait --> Expired: 2 s passed
Capture --> Sparse: tree went empty
Tap --> [*]
FailNow --> [*]: same error as live
Expired --> [*]: selector_not_found + details.readiness
Sparse --> [*]: reason capture_sparse
Loading
The wait covers one case only: the target is not there yet. Everything else fails at once, as it does today, on every platform.
Before and after: the polling loops
Three loops (post-gesture stability, scroll movement, readiness) each had their own cadence, deadline, and idea of what "timed out" means. They share one engine, observeUntil: the budget bounds when a poll starts, a capture may overrun by one interval, cancellation joins the running capture, and every poll is on the timeline. Each loop keeps its own verdict (what counts as "done"). A loop a new backend needs (a helper's fill verification, a provider's stability probe) is a verdict plus a schedule, not a new deadline rule.
What does not change
MCP tools and the interactive CLI: one look, one tap, byte-identical responses.
No new flag for agents. readinessTimeoutMs is set by replay, or by the Node client.
dispatched says what happened up to the device. It never claims what the app did next.
A wait never proves the target stays; a target that vanishes after the poll fails like today.
Why
Test runs on top of agent-device failed in ways nobody could act on. An action timed out, the framework retried it, and the app got two taps. A replayed
pressfired a render before its button existed and failed, while the same step passed by hand. The error saidElement not foundortimed out, and the only way to decide "retry or not" was to match message text.Consumer data (tester-army e2e,
mobile.yml, 723 runs, one week): 303 failures from snapshot quality, 111 from session lifecycle, 88 from gestures, 57 from input. The dominant single flake (Login Form, 42 of 68 first-pass failures) is fixed by #3051 and released in v0.21.17. This issue is about the class that remains: actions that run twice, and taps that run too early.What we get
error.details.dispatched:no(nothing reached the device) orunknown(it may have landed). Consumers retry onno, look first onunknown, and stop matching message text. We stop guessing inside the daemon too. There is noyes: no backend can prove execution on its failure path (the runner's ownXCTEST_RECORDED_FAILUREsays "may not have been performed"), so a third value would only ever be a false claim.press,click, orlongpresswaits up to 2 seconds for a target that is not on screen yet, then taps once. Agents and the interactive CLI keep one-shot behavior.dispatchedvalue that failure must carry (ios-runner.reply.runner-busy→no; a read-only command such asgetis alwaysno, because a read has nothing to repeat). For every row, a test makes the real code produce that failure and asserts the value. A row without a test, or a test without a row, fails the gate. When a backend changes what it can prove, the gate says so before a user does.dispatchedseam, the read-only rule, and the readiness wait run in the daemon; the polling engine lives in capture-kit. A new backend getsunknownon every unclassified failure,noon every read, and the readiness wait, with zero backend code. It adds rows only for the refusals it can prove.Cross-platform: what is inherited, what is declared
flowchart TB subgraph consumers [Consumers] E2E[e2e frameworks] AG[agents via MCP / CLI] SDK[Node client] end subgraph daemon [Daemon: inherited by every backend] RD[Readiness wait<br/>replay + Node client, 2 s cap] DS[dispatched seam<br/>fills unknown once, never overwrites<br/>9 daemon rows] RO[Read-only rule<br/>registry trait → dispatched: no] EN[observeUntil engine<br/>capture-kit, one deadline rule] end subgraph backends [Backends: declare rows, prove them] IOS[Apple runner<br/>iOS, tvOS, macOS, visionOS<br/>25 rows] AND[Android<br/>adb + helper<br/>9 rows] WEB[WebDriver<br/>3 rows] NEW[HarmonyOS, Vega, Linux,<br/>next provider<br/>0 rows = all unknown, still safe] end consumers --> daemon --> backends G[(dispatch-disclosure.json<br/>one driver test per producer, gate)] G -.enforces.- IOS G -.enforces.- AND G -.enforces.- WEBA backend declares a row when it can prove the action never reached the device (
no: a selector miss, a covered target, a runnerRUNNER_BUSY, an adb host refusal). Everything else staysunknown, which is safe by construction: a wrongnomakes a consumer resend an action that ran; a wrongunknowncosts one snapshot. A multi-step action (press --count 25, chunked typing) saysnoonly when its first step was refused; once any step reached the device, a later refusal isunknown.The readiness wait and the engine need nothing from a backend. They consume the same snapshot a backend already serves.
Performance
waitedMsmedian 245 ms, which is the capture itself). Per-run wall time was 34.5 s against 36.5 s p50 on a loaded host over 12 pairs, inside the load noise; the soak issue below measures it properly and has a 3% keep rule.dispatchedis a field on an error, no round trip.Before and after: what a failure tells us
sequenceDiagram participant T as Test framework participant D as Daemon participant B as Backend (runner, adb, WebDriver) T->>D: press "Continue" D->>B: tap B--xD: reply lost Note over D: before: error text "timed out" D-->>T: error (text only) T->>T: regex on message → retry T->>D: press "Continue" D->>B: tap Note over B: the app got two tapssequenceDiagram participant T as Test framework participant D as Daemon participant B as Backend (runner, adb, WebDriver) T->>D: press "Continue" D->>B: tap B--xD: reply lost Note over D: after: details.dispatched = "unknown" D-->>T: error + dispatched T->>D: snapshot D-->>T: the next screen is up Note over T: no retry, the step passederror.details.dispatched"no""unknown"unknownBefore and after: where retries live
flowchart LR subgraph before [Before: five retry sites, five rules] W[WebDriver transport<br/>resends 12 mutating routes on timeout] M[Maestro direct iOS tap<br/>re-taps on 4 message strings] N[Daemon no-change retry<br/>resends when the tree looks the same] R[Runner restart<br/>resends after a lost reply] A[adb<br/>marks offline retriable, nobody acts] endflowchart LR subgraph after [After: one rule, one table] P[Producer classifies its own failure<br/>no / unknown] G[(dispatch-disclosure.json<br/>46 rows, one driver test per producer)] C{Retry only when<br/>dispatched = no} P --> G --> C C -->|WebDriver| W2[zero retries on mutations] C -->|Maestro direct tap| M2[falls back only on no] C -->|adb host refusal| A2[retries once, offline proven pre-dispatch] C -->|daemon| N2[no-change retry deleted] C -->|runner| R2[restart resend needs proof] endBefore and after: a replayed tap
stateDiagram-v2 [*] --> Capture Capture --> Tap: target present Capture --> Fail: target absent Tap --> [*] Fail --> [*]: selector_not_found note right of Fail before: one look, one chance end notestateDiagram-v2 [*] --> Capture Capture --> Tap: target present Capture --> FailNow: covered, off-screen, ambiguous Capture --> Wait: target absent Wait --> Capture: every 200 ms, up to 2 s Wait --> Expired: 2 s passed Capture --> Sparse: tree went empty Tap --> [*] FailNow --> [*]: same error as live Expired --> [*]: selector_not_found + details.readiness Sparse --> [*]: reason capture_sparseThe wait covers one case only: the target is not there yet. Everything else fails at once, as it does today, on every platform.
Before and after: the polling loops
Three loops (post-gesture stability, scroll movement, readiness) each had their own cadence, deadline, and idea of what "timed out" means. They share one engine,
observeUntil: the budget bounds when a poll starts, a capture may overrun by one interval, cancellation joins the running capture, and every poll is on the timeline. Each loop keeps its own verdict (what counts as "done"). A loop a new backend needs (a helper's fill verification, a provider's stability probe) is a verdict plus a schedule, not a new deadline rule.What does not change
readinessTimeoutMsis set by replay, or by the Node client.dispatchedsays what happened up to the device. It never claims what the app did next.Delivery
In order:
dispatchedon failed interactions, table and gate (merged)retryOnNoChange,pendingInteractionOutcome) #3073 Delete the daemon no-change retry path → PR fix(daemon): delete the no-change interaction retry path #3083 (on main with feat(errors): disclose whether a failed interaction reached the device #3071)hostRefusaland derivedispatched: "no"from it #3075 adb:hostRefusalon every host refusal,dispatched: nofrom it → PR fix(android): report dispatched no for adb host refusals #3079 (merged)open --launch-urlshould answer the iOS "Open in <app>?" confirmation #3076open --launch-urlanswers the iOS "Open in app?" sheet → PR feat(open): answer the iOS launch confirmation for --launch-url #3084 (merged)open --launch-url: a post-open target discovery that fails on its deadline must not skip the confirmation read (CI logs of the feat(open): answer the iOS launch confirmation for --launch-url #3084 reds; the first version of the issue blamed the read's budget, which was wrong)simctl openurlhas no timeout (found while reproducing open --launch-url: a post-open target discovery that fails on its deadline must not skip the launch confirmation read #3096)pnpm check:production-exportsorientationrefused by every WebDriver route reportsdispatched: no(merged)Design record: ADR 0011 gains the
targetReadinessandoutcomeObservationcells; the golden table lives atcontracts/fixtures/dispatch-disclosure.json.