Skip to content

Reliability contract: every failed action says whether it ran, scripted taps wait for their target, every platform inherits both #3069

Description

@thymikee

Why

Test runs on top of agent-device failed in ways nobody could act on. An action timed out, the framework retried it, and the app got two taps. A replayed press fired a render before its button existed and failed, while the same step passed by hand. The error said Element not found or timed out, and the only way to decide "retry or not" was to match message text.

Consumer data (tester-army e2e, mobile.yml, 723 runs, one week): 303 failures from snapshot quality, 111 from session lifecycle, 88 from gestures, 57 from input. The dominant single flake (Login Form, 42 of 68 first-pass failures) is fixed by #3051 and released in v0.21.17. This issue is about the class that remains: actions that run twice, and taps that run too early.

What we get

  1. One answer to "is a resend safe?" A failed interaction carries error.details.dispatched: no (nothing reached the device) or unknown (it may have landed). Consumers retry on no, look first on unknown, and stop matching message text. We stop guessing inside the daemon too. There is no yes: no backend can prove execution on its failure path (the runner's own XCTEST_RECORDED_FAILURE says "may not have been performed"), so a third value would only ever be a false claim.
  2. No double actions, on any backend. Every internal retry keys on that same proof. Twelve WebDriver routes and the Maestro direct iOS tap resent blindly; they no longer do. The remaining blind resends (daemon no-change retry, runner restart) are tracked below.
  3. Scripted taps absorb a render race. A replayed press, click, or longpress waits up to 2 seconds for a target that is not on screen yet, then taps once. Agents and the interactive CLI keep one-shot behavior.
  4. A table that fails CI when a backend lies. A row is one way a backend can fail plus the dispatched value that failure must carry (ios-runner.reply.runner-busy → no; a read-only command such as get is always no, because a read has nothing to repeat). For every row, a test makes the real code produce that failure and asserts the value. A row without a test, or a test without a row, fails the gate. When a backend changes what it can prove, the gate says so before a user does.
  5. Every platform and provider inherits this. The three behaviors live above the backends: the dispatched seam, the read-only rule, and the readiness wait run in the daemon; the polling engine lives in capture-kit. A new backend gets unknown on every unclassified failure, no on every read, and the readiness wait, with zero backend code. It adds rows only for the refusals it can prove.

Cross-platform: what is inherited, what is declared

flowchart TB
    subgraph consumers [Consumers]
        E2E[e2e frameworks]
        AG[agents via MCP / CLI]
        SDK[Node client]
    end
    subgraph daemon [Daemon: inherited by every backend]
        RD[Readiness wait<br/>replay + Node client, 2 s cap]
        DS[dispatched seam<br/>fills unknown once, never overwrites<br/>9 daemon rows]
        RO[Read-only rule<br/>registry trait → dispatched: no]
        EN[observeUntil engine<br/>capture-kit, one deadline rule]
    end
    subgraph backends [Backends: declare rows, prove them]
        IOS[Apple runner<br/>iOS, tvOS, macOS, visionOS<br/>25 rows]
        AND[Android<br/>adb + helper<br/>9 rows]
        WEB[WebDriver<br/>3 rows]
        NEW[HarmonyOS, Vega, Linux,<br/>next provider<br/>0 rows = all unknown, still safe]
    end
    consumers --> daemon --> backends
    G[(dispatch-disclosure.json<br/>one driver test per producer, gate)]
    G -.enforces.- IOS
    G -.enforces.- AND
    G -.enforces.- WEB
Loading

A backend declares a row when it can prove the action never reached the device (no: a selector miss, a covered target, a runner RUNNER_BUSY, an adb host refusal). Everything else stays unknown, which is safe by construction: a wrong no makes a consumer resend an action that ran; a wrong unknown costs one snapshot. A multi-step action (press --count 25, chunked typing) says no only when its first step was refused; once any step reached the device, a later refusal is unknown.

The readiness wait and the engine need nothing from a backend. They consume the same snapshot a backend already serves.

Performance

  • Hit path: one capture, same as today. With the readiness option set, a target present on the first capture is tapped after that capture; measured on a private iPhone 17 Pro simulator, every one of 46 clicks resolved in one poll (waitedMs median 245 ms, which is the capture itself). Per-run wall time was 34.5 s against 36.5 s p50 on a loaded host over 12 pairs, inside the load noise; the soak issue below measures it properly and has a 3% keep rule.
  • Miss path: bounded at 2 s, and only where it is set. MCP and the interactive CLI never wait; an agent that misses on a wrong selector gets the fast error it gets today.
  • Failure path: cheaper. A WebDriver mutation that timed out used to pay a second full request timeout before failing; it fails after one. dispatched is a field on an error, no round trip.
  • Engine: no extra polls. The budget bounds when a poll starts, and a capture in flight is never abandoned mid-way; cancellation joins it. The two loops moved onto the engine (post-gesture stability, scroll movement) keep their cadence and verdicts.
  • Request envelope grows only by the budget, on client and daemon, so a replayed step never times out on the client while the daemon is still inside its wait.

Before and after: what a failure tells us

sequenceDiagram
    participant T as Test framework
    participant D as Daemon
    participant B as Backend (runner, adb, WebDriver)
    T->>D: press "Continue"
    D->>B: tap
    B--xD: reply lost
    Note over D: before: error text "timed out"
    D-->>T: error (text only)
    T->>T: regex on message → retry
    T->>D: press "Continue"
    D->>B: tap
    Note over B: the app got two taps
Loading
sequenceDiagram
    participant T as Test framework
    participant D as Daemon
    participant B as Backend (runner, adb, WebDriver)
    T->>D: press "Continue"
    D->>B: tap
    B--xD: reply lost
    Note over D: after: details.dispatched = "unknown"
    D-->>T: error + dispatched
    T->>D: snapshot
    D-->>T: the next screen is up
    Note over T: no retry, the step passed
Loading
error.details.dispatched Meaning What a consumer does
"no" Nothing reached the device Retry as is
"unknown" It may have landed Snapshot before any retry
absent Not classified Treat as unknown

Before and after: where retries live

flowchart LR
    subgraph before [Before: five retry sites, five rules]
        W[WebDriver transport<br/>resends 12 mutating routes on timeout]
        M[Maestro direct iOS tap<br/>re-taps on 4 message strings]
        N[Daemon no-change retry<br/>resends when the tree looks the same]
        R[Runner restart<br/>resends after a lost reply]
        A[adb<br/>marks offline retriable, nobody acts]
    end
Loading
flowchart LR
    subgraph after [After: one rule, one table]
        P[Producer classifies its own failure<br/>no / unknown]
        G[(dispatch-disclosure.json<br/>46 rows, one driver test per producer)]
        C{Retry only when<br/>dispatched = no}
        P --> G --> C
        C -->|WebDriver| W2[zero retries on mutations]
        C -->|Maestro direct tap| M2[falls back only on no]
        C -->|adb host refusal| A2[retries once, offline proven pre-dispatch]
        C -->|daemon| N2[no-change retry deleted]
        C -->|runner| R2[restart resend needs proof]
    end
Loading

Before and after: a replayed tap

stateDiagram-v2
    [*] --> Capture
    Capture --> Tap: target present
    Capture --> Fail: target absent
    Tap --> [*]
    Fail --> [*]: selector_not_found
    note right of Fail
        before: one look, one chance
    end note
Loading
stateDiagram-v2
    [*] --> Capture
    Capture --> Tap: target present
    Capture --> FailNow: covered, off-screen, ambiguous
    Capture --> Wait: target absent
    Wait --> Capture: every 200 ms, up to 2 s
    Wait --> Expired: 2 s passed
    Capture --> Sparse: tree went empty
    Tap --> [*]
    FailNow --> [*]: same error as live
    Expired --> [*]: selector_not_found + details.readiness
    Sparse --> [*]: reason capture_sparse
Loading

The wait covers one case only: the target is not there yet. Everything else fails at once, as it does today, on every platform.

Before and after: the polling loops

Three loops (post-gesture stability, scroll movement, readiness) each had their own cadence, deadline, and idea of what "timed out" means. They share one engine, observeUntil: the budget bounds when a poll starts, a capture may overrun by one interval, cancellation joins the running capture, and every poll is on the timeline. Each loop keeps its own verdict (what counts as "done"). A loop a new backend needs (a helper's fill verification, a provider's stability probe) is a verdict plus a schedule, not a new deadline rule.

What does not change

  • MCP tools and the interactive CLI: one look, one tap, byte-identical responses.
  • No new flag for agents. readinessTimeoutMs is set by replay, or by the Node client.
  • dispatched says what happened up to the device. It never claims what the app did next.
  • A wait never proves the target stays; a target that vanishes after the poll fails like today.

Delivery

In order:

Design record: ADR 0011 gains the targetReadiness and outcomeObservation cells; the golden table lives at contracts/fixtures/dispatch-disclosure.json.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions