Skip to content

fix(apple-runner): restart without resending a command written then lost - #3082

Merged
thymikee merged 18 commits into
feat/dispatch-disclosurefrom
fix/runner-restart-resend-needs-proof
Oct 1, 2026
Merged

thymikee merged 18 commits into
feat/dispatch-disclosurefrom
fix/runner-restart-resend-needs-proof

Conversation

@thymikee

@thymikee thymikee commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Summary

A mutating command whose reply stays lost now fails with details.reason: runner_reply_lost and dispatched: "unknown". This covers a runner that dies mid-command, and every status-recovery verdict that finds neither a retained result nor a runner answer: completed without a retained reply, still accepted or started, notAccepted, an unknown or missing lifecycle state, a failed status probe, and a command with no id to probe. The command is sent once and never replayed, and the next command starts a new runner. One builder makes this error. The transport's own reason stays under details.transportReason. A read keeps its transport error, so the read-only resend loop resends it, and the read never carries the reason.

The restart-and-resend rule after a connect failure is owned by shouldRestartRunnerBeforeCommandSend(error, command). The rule restarts only when no attempt could have written the command (resolveFirstAttemptDispatch(error) === 'no') or when the command is read-only. A simctl curl exit-28 POST carries dispatched: unknown, so it never restarts a mutation. The restart carries the first attempt's DispatchDisclosure, not a boolean.

agent-device longpress @e3 5000 --json   # runner killed mid-press: runner_reply_lost, sent once

Closes #3074. Part of #3069. Stacked on #3071. 14 files touched.

Validation

Tested commit: 4e88908.

pnpm check:affected --run (fail-open full set): every runnable check passed. In vitest-related, 1333 of 1334 files passed. The one failure was a 5 s timeout under host load in a file this PR does not touch (screenshot-density.test.ts, and press-target-readiness.test.ts on the run before). Each passes alone. The checks after vitest-related were then run on the same commit, and all passed. Also run: pnpm typecheck, pnpm lint, apple runner vitest (167 files, 1704 tests), check:layering, check:fallow, format:check, the eager-closure budget test, and the runner read-only parity test.

Live on f52f120 (iPhone 17 Pro simulator, iOS 26.2, Release test app): longpress @e3 5000 --debug, with the runner process killed 1.3 s after the runner accepted the command. The request log has one ios_runner_command_send for longPress, then a failed status probe, invalidation_decision status_probe_failed, and ios_runner_session_invalidated transport_error_after_command_send. The error has reason: runner_reply_lost, recovery: status_probe_failed, dispatched: unknown, and no runnerRestarted. snapshot -i then succeeded, and the next runner command started a new runner session (port 61516, was 61032). The structural changes after that commit do not change this path.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Size Report

Metric Base Current Diff
Installed (including dependencies) 4.92 MB 4.92 MB +561 B
Package (unpacked) 4.92 MB 4.92 MB +561 B
Package (download) 1.47 MB 1.47 MB +155 B

Startup median (7 runs, lower is better):

Scenario Base Current Diff
CLI --version 27.6 ms 26.3 ms -1.3 ms
CLI --help 81.1 ms 81.1 ms -0.0 ms

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 9 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="packages/platform-apple/src/runner/runner-lifecycle.ts">

<violation number="1" location="packages/platform-apple/src/runner/runner-lifecycle.ts:542">
P3: `discloseRestartDispatch` now reports `dispatched: "no"` for read-only commands when the restart's resend fails with an already-disclosed `"unknown"` verdict (the `ios-runner.transport.read-only-written-then-lost` fixture asserts this). The parallel status-recovery path in `runner-command-recovery.ts` still discloses the same user-visible case as `"unknown"` (`handleCompletedRunnerStatus`'s `read_only_completed_without_retained_response` branch, and the fallthrough after a `notAccepted` probe for `lifecycle_state_not_recoverable`), so an agent deciding whether to retry on `details.dispatched` sees different values for the same class of failure depending on which recovery flow fired. Decide one semantics for read-only lost replies and pin both paths (and the fixture) to it.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

): AppError {
if (!restartedSession) return discloseDispatch(error, firstAttemptUnwritten ? 'no' : 'unknown');
if (!firstAttemptUnwritten) return discloseDispatch(error, 'unknown');
if (isReadOnlyRunnerCommand(evidence.command)) return discloseDispatch(error, 'no');

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: discloseRestartDispatch now reports dispatched: "no" for read-only commands when the restart's resend fails with an already-disclosed "unknown" verdict (the ios-runner.transport.read-only-written-then-lost fixture asserts this). The parallel status-recovery path in runner-command-recovery.ts still discloses the same user-visible case as "unknown" (handleCompletedRunnerStatus's read_only_completed_without_retained_response branch, and the fallthrough after a notAccepted probe for lifecycle_state_not_recoverable), so an agent deciding whether to retry on details.dispatched sees different values for the same class of failure depending on which recovery flow fired. Decide one semantics for read-only lost replies and pin both paths (and the fixture) to it.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/platform-apple/src/runner/runner-lifecycle.ts, line 542:

<comment>`discloseRestartDispatch` now reports `dispatched: "no"` for read-only commands when the restart's resend fails with an already-disclosed `"unknown"` verdict (the `ios-runner.transport.read-only-written-then-lost` fixture asserts this). The parallel status-recovery path in `runner-command-recovery.ts` still discloses the same user-visible case as `"unknown"` (`handleCompletedRunnerStatus`'s `read_only_completed_without_retained_response` branch, and the fallthrough after a `notAccepted` probe for `lifecycle_state_not_recoverable`), so an agent deciding whether to retry on `details.dispatched` sees different values for the same class of failure depending on which recovery flow fired. Decide one semantics for read-only lost replies and pin both paths (and the fixture) to it.</comment>

<file context>
@@ -500,21 +525,23 @@ function markRunnerRestartError(
 ): AppError {
-  if (!restartedSession) return discloseDispatch(error, firstAttemptUnwritten ? 'no' : 'unknown');
-  if (!firstAttemptUnwritten) return discloseDispatch(error, 'unknown');
+  if (isReadOnlyRunnerCommand(evidence.command)) return discloseDispatch(error, 'no');
+  if (!evidence.firstAttemptUnwritten) return discloseDispatch(error, 'unknown');
+  if (!restartedSession) return discloseDispatch(error, 'no');
</file context>
Fix with cubic

Comment thread packages/platform-apple/src/runner/__tests__/runner-recovery-wiring.test.ts Outdated
Comment thread packages/platform-apple/src/runner/__tests__/runner-recovery-wiring.test.ts Outdated
@thymikee

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

Head is now 4758d2786d.

1. Read-only restart discloses no but status recovery says unknown (runner-lifecycle.ts ~L542): fixed in be1b89f.
Both runner paths now report the transport fact. If the first attempt may have written the command, the value is unknown, for a read or a mutation. The read-only no now has one owner: the router seam src/daemon/request-dispatch-disclosure.ts, which keys it on the registry's recordingEffect: 'observes-app' (row daemon.read-only-command). The runner branch keyed on its own read-only trait was a second copy of that rule from a different source, so I removed it. The recycle-budget refusal had the same problem, so it now uses only firstAttemptUnwritten. The runner still resends a read after the restart. Only the disclosure changes.

  • ios-runner.transport.read-only-written-then-lost is now unknown.
  • New row ios-runner.status.read-only-completed-without-retained-reply is unknown. Its driver goes through the real connect loop and the recovery module, so the two paths are pinned to the same value.
  • Mutations: putting the read-only no branch back in discloseRestartDispatch makes the restart row fail. Setting dispatched: 'no' on read_only_completed_without_retained_response makes the new status row fail.

2. Unconsumed second ok for the mutation (runner-recovery-wiring.test.ts ~L170): fixed in 4758d27. I dropped it.

3. Premise says "restarted", but the path is status recovery (~L164): fixed in 4758d27, with one pushback.

  • I changed the comment on the existing test. It models status answering notAccepted, the way a runner whose journal did not survive a restart would. The path is status recovery.
  • Harness gap: the fake server could only hang up on one request and stayed up, so it could not model a runner that dies. I added the smallest hook: an { kind: 'exit' } reply that hangs up and stops listening. A second fake server acts as the restarted runner.
  • New test for press/fill (tap/type): the runner dies on the command. Across both servers the command is sent once, it fails with dispatched: unknown, the dead session is invalidated, and the next readText is served by the restarted runner. Mutation: declaring tap/type with READ_ONLY_TRAITS makes both rows fail.
  • New test for get (readText): the runner dies mid-read. The daemon restarts it (runner_connect_failed_before_command_send), and the read is resent once and succeeds. Mutation: reducing canResendAfterRestart to firstAttemptUnwritten makes it fail.
  • Pushback: with the real transport, a mutating command cannot reach restartSessionAndRunCommand with a possibly-written first attempt. Mutations are sent by sendRunnerCommandOnce (runner-exchange.ts, the non-read-only branch of sendRunnerCommandAfterPreflight). Only the connect loop waitForRunner stamps runner_connect_refused (runner-startup-transport.ts, runnerConnectFailureDetails call sites), and that loop sends only reads and the readiness preflight. The preflight restart always passes firstAttemptUnwritten: true. So the mutating arm of the new gate can only be driven through the mocked lifecycle driver (ios-runner.transport.written-then-lost). No fake-server test can reach it.
  • Also, when a runner dies mid-mutation, the failure goes to status recovery, and the status probe fails there. That path rethrows the bare transport error (fetch failed, dispatched: unknown) with no runner_reply_lost reason. So the new test does not assert that reason. runner-command-recovery.test.ts ("a failing status probe … rethrows the transport error") pins this as the current contract. Adding the reason there is a separate change.

Checks run: pnpm typecheck; vitest on packages/platform-apple/src/runner (68 files, 768 tests) and on packages/contracts; pnpm check:layering; pnpm check:fallow; pnpm run format:check. All pass.

@thymikee

thymikee commented Sep 30, 2026 •

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

Follow-up in 50fc6fc:

  • Typed reason on the path that occurs: a runner dies mid-mutation and the status probe then fails. The error now has reason: runner_reply_lost (the name the restart arm uses), recovery: status_probe_failed and dispatched: unknown. It keeps the transport message and details, and the transport error as cause. The recovery unit test that pinned the bare fetch failed rethrow now asserts the reason. So do the real-transport press/fill tests. Mutation: dropping reason: RUNNER_REPLY_LOST_REASON from buildStatusProbeFailedError makes all three fail.
  • Mutating arm of canResendAfterRestart: kept. The ADR 0011 gap list now says that ios-runner.transport.written-then-lost proves the rule, not a production route. Over the real transport that arm cannot be reached, because mutations go through sendRunnerCommandOnce and only the connect loop raises the restart trigger. Mutation (from c6188ef): making canResendAfterRestart return true fails the mocked row, because the tap is resent.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 8 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

Comment thread packages/platform-apple/src/runner/__tests__/runner-dispatch-disclosure.test.ts Outdated
@thymikee

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

Hang-up count tied to the connect-loop cadence (runner-dispatch-disclosure.test.ts ~L81): fixed in f247110, with a complexity follow-up in the next commit. The fake runner server has a new hangUpAlways reply that is never consumed. The read-only driver uses it in place of 20 hang-ups, so the read is refused until the connect loop gives up, however often the loop retries. The 400 ms timeout now only bounds how long the loop runs. Mutation: setting RUNNER_CONNECT_ATTEMPT_INTERVAL_MS to 10 or 50 (with a 1 ms retry base delay) does not change the outcome, and the row still passes.

@thymikee

Copy link
Copy Markdown
Member Author

The change looks correct in code, but the one reachable behavior it adds has no live proof yet. I reviewed 1edcb78. CI is green: 14 checks, none failing. No conflicts.

The reachable change is the typed reason runner_reply_lost on a mutation whose runner dies mid-command (runner-command-recovery.ts#L207). The route is: status probe fails, then retainInvalidation, then the reason with recovery status_probe_failed. The only live run was on 4708d4f, and it killed the runner before the press. That run hit the existing refused-connect restart-and-resend arm, not this route, and the PR's Risks line says so. So we have no proof that a real runner death mid-mutation gives this error and that the next command boots a new runner. Please run this on head on an iOS simulator with the Release test app: start a long mutation such as longpress @eN 5000 --debug (or a gesture or long swipe) and kill the XCTest runner while it is in flight. The request log must show exactly one ios_runner_command_send for it. The error must show details.reason=runner_reply_lost, recovery status_probe_failed (or lifecycle_state_not_recoverable) and dispatched=unknown. There must be no runnerRestarted resend, and an ios_runner_session invalidation with transport_error_after_command_send. Then snapshot -i must succeed on a new runner sessionId. A run that kills the runner before the send does not count.

Could the change be smaller? The connect loop in waitForRunner only carries read-only commands, so the new no-resend arm in restartSessionAndRunCommand sees a read, or a mutation with unwritten proof from the readiness preflight. Could the call site pass the resend permission (isRunnerPreSendRefusal || isReadOnlyRunnerCommand) as the entry condition, or drop the mutating arm and its error builder? Then the reachable delta is the runner_reply_lost reason on the two status-recovery errors, about 30 lines. The ADR 0011 sentence already records the rule, so nothing needs deciding first. Closing #3074 would need a note that its premise (restart resends a written mutation) has no production route.

Not blocking, take or leave: the no-resend arm at runner-lifecycle.ts#L437 is only driven by a mocked transport row, because mutations never enter the connect loop; the notAccepted and runner-dies-mid-command tests in runner-recovery-wiring.test.ts#L173 pass on the merge-base except for the runner_reply_lost assertion, so they could be retitled to that reason; buildStatusProbeFailedError spreads transportError.details and then overwrites reason and hint, so a transport reason would be lost (none of today's producers set one, but a separate key would keep it); and the PR body is stale, since it says reads disclose dispatched: "no" when the runner reports unknown and the router says no, it says 9 files when the diff has 12, it names 4708d4f instead of head, and it records check:affected exiting 1 with no green rerun.

I did not run the tests or the mutations. I judged "fails without the fix" from the merge-base code, and "next command gets a fresh runner" is only shown with a mocked ensureRunnerSession.

Before merge, we need the live run on 1edcb78 described above, and a refreshed PR body with the head sha, a green check:affected result and the router-owned "no".

@thymikee

thymikee commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

Head is now f52f120b1d.

Simplification (6739a58): I took it.

  • The connect-failure call site in executeRunnerCommandAttempt now restarts only when isRunnerPreSendRefusal(appErr) || isReadOnlyRunnerCommand(command). Any other connect failure falls through to status recovery, as before.
  • Deleted: the no-resend arm in restartSessionAndRunCommand, canResendAfterRestart, the firstAttemptError param, buildRunnerRestartedWithoutResendError, and the comment edits in runner-error-classification.ts (now identical to the base). I also reverted a reorder in discloseRestartDispatch that did not change behaviour.
  • Production delta in this round: +26 / −72 lines. Against the merge-base, production is now +47 / −3 in runner-command-recovery.ts and runner-lifecycle.ts.
  • The ADR 0011 rule sentence stays, as its own paragraph. I moved it out of the "Remaining gaps" paragraph that feat(errors): disclose whether a failed interaction reached the device #3071 just rewrote, so the PR merges cleanly with the new base.
  • The row ios-runner.transport.written-then-lost no longer has a mocked lifecycle driver. It now runs tap over the real transport against the fake runner's { kind: 'exit' } hook. It asserts one tap send, reason: runner_reply_lost, recovery: status_probe_failed, no runnerRestarted, and invalidation with transport_error_after_command_send. Mutation: removing the call-site gate (always restart) fails 8 runner tests.

Reason key: buildStatusProbeFailedError keeps a transport details.reason as details.transportReason before it sets reason: runner_reply_lost. The probe-failed recovery test pins this.

Retitles (f52f120): a $acceptanceCommand whose status answers notAccepted fails as runner_reply_lost, sent once and a $acceptanceCommand whose runner dies mid-command fails as runner_reply_lost, not resent.

Live on f52f120: new iPhone 17 Pro sim, iOS 26.2, Release test app. I ran longpress @e3 5000 --debug --json and killed the sim's runner process during the press. The runner accepted the command at 23:32:47.630, the kill was at 23:32:48.922, and the send failed at 23:32:49.030. This was the first try.

47.612 ios_runner_session_reuse          sessionId …:61032:1790811100968
49.030 ios_runner_command_send  1402ms   {"command":"longPress","error":"fetch failed"}
52.035 ios_runner_command_send  3003ms   {"command":"status","error":"Runner did not accept connection"}
52.035 ios_runner_command_invalidation_decision {"decision":"retained","reason":"status_probe_failed","runnerReachable":false}
52.035 ios_runner_session_invalidated    {"reason":"transport_error_after_command_send"}
  • Error: reason: runner_reply_lost, recovery: status_probe_failed, dispatched: unknown, with no runnerRestarted. The log has exactly one longPress send.
  • snapshot -i --json then succeeded. On this sim it is served by the AX bridge and does not touch the runner. The next runner command (longpress @e3 100) started a new runner, session …:61516:1790811200162.

Gate: pnpm check:affected --run is green on f52f120. The first run had one 5 s timeout under load (fallow-fixture-policy.test.ts), which passes alone. The PR body is refreshed.

@thymikee thymikee left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code-quality review. The shipped behavior is right: a mutation whose runner dies mid-command is sent once and fails with runner_reply_lost, through the status probe. The structure has three problems:

  • The new restart gate at the call site never changes behavior on any input that can reach it. The defect is in the classifier that labels an exit-28 POST as runner_connect_refused.
  • The written-then-lost proof is a boolean that later code converts back into 'no' | 'unknown', instead of the existing DispatchDisclosure type.
  • runner_reply_lost is set on some lost-reply outcomes and not on others, by four hand-built copies of the same error.

Inline comments cover each point. No file crosses 1k lines.

if (
shouldRestartRunnerBeforeCommandSend(appErr) &&
session &&
(firstAttemptUnwritten || isReadOnlyRunnerCommand(command))

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This conjunct guards a state that cannot happen. Mutations go through sendRunnerCommandOnce and never enter the connect loop (waitForRunner / postCommandViaSimulator), and that loop is the only producer of runner_connect_refused. The PR description says the same. For a mutation, the only restartBeforeSend error that can arrive is a marked readiness-preflight failure, and isRunnerPreSendRefusal is already true for that. For a read, the disjunct is true. So (firstAttemptUnwritten || isReadOnlyRunnerCommand(command)) is true for every input that can reach this line. No test exercises the side where it refuses either: the written-then-lost row moved to the status-probe path, and every remaining mutation driver uses unwrittenConnectRefusal().

The real defect is in the code that owns the classification. restartBeforeSend is documented as "connect-shaped failure before the command was sent". But postCommandViaSimulator (runner-startup-transport.ts ~L517) labels a curl exit-28 POST that has dispatched: 'unknown' as runner_connect_refused. Fix it there. Either stop labelling a request that was posted and then timed out as a connect refusal, or make shouldRestartRunnerBeforeCommandSend(error, command) in runner-error-classification.ts own the write-proof rule. Then delete this call-site guard. AGENTS.md says to repair at the owning type and not to add guards that rebuild another source of truth.

signal,
restartReason: 'runner_connect_failed_before_command_send',
firstAttemptUnwritten: isRunnerPreSendRefusal(appErr),
firstAttemptUnwritten,

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

firstAttemptUnwritten is a boolean derived from a typed fact that already exists: the error's details.dispatched, which is 'no' for curl exit 7 or ECONNREFUSED. Downstream code then turns it back into 'no' | 'unknown' with hand-written ternaries: once for the recycle-budget throw in restartSessionAndRunCommand, and twice in discloseRestartDispatch. A reader has to decode the boolean every time it is used.

Carry the first attempt's disclosure instead:

  • Add firstAttemptDispatched: DispatchDisclosure, computed once by one classifier, for example resolveFirstAttemptDispatch(error) next to isRunnerCommandProvablyUnwritten.
  • Replace the ternaries with discloseDispatch(err, params.firstAttemptDispatched).
  • Have the readiness-preflight call site pass 'no' instead of true.

This models the proof as a value of the domain type and removes branches instead of adding them. Together with the L359 comment, the restart decision becomes one answer from one classifier.

* `details.reason` of a failure whose command was written, lost its reply, and has no proof it did
* not run. The command is not resent; the caller observes the screen before acting again.
*/
export const RUNNER_REPLY_LOST_REASON = 'runner_reply_lost';

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

runner_reply_lost is documented as: the command was written, its reply was lost, there is no proof it did not run, and it is not resent. But only some outcomes that match this carry it: status_probe_failed and the unknown/missing-lifecycle case.

  • completed_without_retained_response (handleCompletedRunnerStatus) and command_still_in_flight (runnerStatusInFlightError) are also written, reply lost, dispatched: 'unknown', and not resent. They do not set the reason.
  • In the other direction, buildStatusProbeFailedError sets it on read-only commands too. That error keeps the code, message and connect details, so isRetryableRunnerError still holds, and the read-only resend loop in runAppleRunnerCommand resends the read up to TRANSPORT_RESEND_ATTEMPTS times. A read that runs out of attempts therefore ends with a reason that says "not resent".

A consumer that keys on details.reason gets a signal that is incomplete and sometimes false. The reason belongs to the recovery verdict, not to each error builder. Set it in one place, for example in applyRunnerTransportRecovery for every non-recovered mutation verdict with dispatched: 'unknown'. Add one table-driven test that every lost-reply row in dispatch-disclosure.json carries it.

}

/** The lost reply could not be placed because the status probe itself failed. */
function buildStatusProbeFailedError(

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the fourth hand-built copy of the same lost-reply AppError: command, commandId, recovery, hint, logPath, the transportError message, and transportError as the cause. The other three are the inline unknown-lifecycle literal (~L289–305), handleCompletedRunnerStatus and runnerStatusInFlightError.

This copy also has a different shape. The other three drop the transport details. This one spreads ...transportError.details, overwrites reason, and moves the original reason into a new transportReason key that only this path has. To find the transport's reason, a consumer now has to know which path built the error.

Extract one buildLostReplyError(command, transportError, options, { recovery, lifecycleState?, message?, hint }). Make it own the detail shape, the reply-lost reason (see the L54 comment), and one rule for keeping the transport's own reason. Then rewrite this function and the inline literal as calls to it.

@thymikee

thymikee commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

This PR is ready. The earlier findings at 1edcb78 are fixed in f52f120, and I found nothing new to block on. Not blocking, and you can take or leave it: the guard (firstAttemptUnwritten || isReadOnlyRunnerCommand(command)) at https://github.com/callstack/agent-device/blob/f52f120/packages/platform-apple/src/runner/runner-lifecycle.ts#L359 looks true for every input that reaches it, and no test covers its refuse side, so you could let shouldRestartRunnerBeforeCommandSend own the write-proof rule, or stop labelling a posted-then-timed-out curl exit 28 as runner_connect_refused in runner-startup-transport.ts, and then drop the guard.

The code review is done. I read the quoted request-log lines from your live run, but I saw no raw artifact, so that evidence is author-reported. I did not run the runner vitest suite or your gate-removal mutation. From reading the code, the refuse side looks untested, which does not match the "8 tests fail" claim.

CI is still pending. Smoke Tests is in progress and has not failed, and the other 13 checks pass. If its iOS lane runs runner mutations, it exercises the changed route, which is the restart gating in executeRunnerCommandAttempt and the status-probe recovery. No conflicts. Smoke Tests must finish green before merge, and nothing in this delta stops that.

@thymikee thymikee added the ready-for-human Valid work that needs human implementation, judgment, or maintainer merge label Oct 1, 2026
@thymikee
thymikee force-pushed the feat/dispatch-disclosure branch from b2027c2 to d7b442d Compare October 1, 2026 06:42
@thymikee
thymikee force-pushed the fix/runner-restart-resend-needs-proof branch from f52f120 to 9393095 Compare October 1, 2026 06:52
@thymikee

thymikee commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

Replayed onto #3071's d7b442d040. Conflicts: docs/adr/0011-interaction-guarantee-contract.md only, in two commits. I kept the base's runBatch paragraph and dropped this branch's older sentence that said a batch outside the daemon keeps the failing step's value, because the base now contradicts it. The written-then-lost row sentence stays only in its final form (the paragraph above "Remaining gaps").

Added test(apple-runner): the runner read-only set matches the registry's observes-app commands (src/__tests__/runner-read-only-registry-parity.test.ts). The runner has no mapping to registry commands, so the test declares one: each runner wire command is paired with the registry request that issues it, and readOnly must equal observes-app for it; alert actions are checked the same way. Runner-internal commands (status, uptime, appState, gestureViewport, sequence, shutdown, activate, terminate, targetReset) are listed with their expected value. recordStart and recordStop are the one divergence: record observes the app, but the runner must not resend them, so they are declared stricter than the registry. The test loads the package-private traits module by computed file URL, because the layering gate rejects a relative or type import across the package boundary and the module has no subpath.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 1 file (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

Comment thread src/__tests__/runner-read-only-registry-parity.test.ts Outdated
…t command

A command whose reply was lost and whose status probe answers notAccepted
(a restarted runner's empty journal) or an unnamed state now fails with
details.reason runner_reply_lost beside dispatched unknown, so a caller
keys on the typed reason instead of the message before it observes the
screen.

Tests (fake runner server through runAppleRunnerCommand and the recovery
module):
- runner-recovery-wiring: press and fill (runner tap and type) accept the
  command, drop the connection, status answers notAccepted: one dispatched
  command, error dispatched unknown with reason runner_reply_lost, the
  session is invalidated, and the next command gets a runner. Mutation:
  delete `reason: RUNNER_REPLY_LOST_REASON` from the unknown-state error ->
  both rows fail.
- runner-recovery-wiring: a snapshot whose reply is dropped is resent and
  succeeds. Mutation: declare snapshot with DEFAULT_TRAITS (not read-only)
  -> fails (the lost read goes to status recovery and is not resent).
- runner-command-recovery: notAccepted yields reason runner_reply_lost and
  dispatched unknown. Mutation: same deletion as above -> fails.
A connect-shaped failure keeps its restart, but the command is sent again
on the restarted runner only with proof the first send could not have run
it: the attempt provably wrote nothing (a connect refused before the write,
dispatched no), the readiness preflight gave up before the write, or the
command is read-only by its runner trait. A mutating command whose first
attempt may have written it now fails with reason runner_reply_lost and
dispatched unknown after the restart; a restarted runner's journal is empty,
so no status answer can prove the first send did not run. A read-only
command's restart failure discloses no: a read has no side effect to repeat.

dispatch-disclosure.json: ios-runner.transport.written-then-lost keeps
unknown with a trigger that says the command is not resent;
ios-runner.transport.read-only-written-then-lost is new and discloses no.

Tests (runner-lifecycle-dispatch-disclosure, real connect loop over a
stubbed fetch and simctl curl exit 28, restart succeeds):
- written-then-lost: tap restarts the runner, is not sent on the restarted
  runner, fails unknown with reason runner_reply_lost. Mutation: make
  canResendAfterRestart return true -> fails (the tap is replayed and the
  request succeeds).
- read-only-written-then-lost: snapshot is sent again on the restarted
  runner and its failure discloses no. Mutations: drop the read-only term
  from canResendAfterRestart -> fails (no resend); drop the read-only
  branch of discloseRestartDispatch -> fails (unknown).
runner-command-retry: four mutating-restart fixtures now carry the proof a
real refused connect stamps (dispatched no); without it they would test the
removed replay.
The restart path disclosed `no` for a read-only command whose first send
may have run, while the status-recovery path discloses `unknown` for the
same read (read_only_completed_without_retained_response, and the
notAccepted/unknown-state fallthrough). Both runner paths now report the
transport fact: a first attempt that may have written the command is
`unknown`, read or mutation. The read-only `no` has one owner, the daemon
router seam (request-dispatch-disclosure.ts), which keys it on the
registry's recordingEffect `observes-app` for every producer
(daemon.read-only-command). A second copy keyed on the runner's own
read-only trait reconstructed that rule from a different source of truth,
and the two runner paths disagreed. The recycle-budget refusal on the
restart path also stops deriving `no` from the read-only trait and uses
only whether the first attempt was provably unwritten.

The runner still resends a read on the restarted runner
(canResendAfterRestart); only the disclosure changes.

dispatch-disclosure.json:
- ios-runner.transport.read-only-written-then-lost: now `unknown`; the
  trigger names the router rule that turns it into `no` for the caller.
- ios-runner.status.read-only-completed-without-retained-reply: new,
  `unknown`, driven through the real connect loop and recovery module.
  It pins the status-recovery read to the same value as the restart read.

Mutations:
- put back `if (isReadOnlyRunnerCommand(evidence.command)) return
  discloseDispatch(error, 'no')` in discloseRestartDispatch ->
  read-only-written-then-lost fails.
- set `dispatched: 'no'` on read_only_completed_without_retained_response
  -> status.read-only-completed-without-retained-reply fails.
runner-recovery-wiring: the notAccepted mutation test no longer scripts a
second `ok` for the mutation. Nothing consumed it, and it suggested that
a resend was expected. Its comment now says what the test models: status
answers notAccepted, the way a runner whose journal did not survive a
restart would. The path is status recovery, not the daemon's restart.

The fake runner server gains an `exit` reply: it hangs up and stops
listening, like a runner process that died mid-command. Before this, a
test could only hang up on one request while the same server stayed up.
Two new tests use a second server as the restarted runner:

- press/fill (runner tap/type): the runner dies on the command. The
  mutation is sent once, is not sent to the restarted runner, fails with
  dispatched unknown, and the dead session is invalidated. A readText on
  the next request is served by the restarted runner. Mutation: declare
  tap and type with READ_ONLY_TRAITS (the resend trait) -> both rows fail.
- get (runner readText): the runner dies mid-read. The connect loop gives
  up with a possibly-written POST (simctl curl exit 28), the daemon
  restarts the runner (runner_connect_failed_before_command_send), and the
  read is resent once and succeeds. Mutation: reduce canResendAfterRestart
  to `evidence.firstAttemptUnwritten` -> fails (no resend).

With the real transport, a mutating command never reaches
restartSessionAndRunCommand with a possibly-written first attempt.
Mutations are sent by sendRunnerCommandOnce, and only the connect loop
(waitForRunner, used for reads and the readiness preflight) stamps
runner_connect_refused. So the mutating arm of canResendAfterRestart is
covered only by the mocked lifecycle driver
(ios-runner.transport.written-then-lost). A dying runner sends a mutation
to status recovery, and a failed status probe rethrows the transport
error without a runner_reply_lost reason.
A runner that dies mid-mutation drops the reply, and the status probe that
follows fails. The error that reaches the caller was then a bare transport
error (`fetch failed`, dispatched unknown), with no typed reason. #3074
requires a reason that names the lost reply, so this path now builds an
error with reason runner_reply_lost (the name the restart arm already uses)
and recovery status_probe_failed. The error keeps the transport message
and details, and the transport error as cause. The session is still
invalidated, and the command is still not resent.

ADR 0011 gap list: the row ios-runner.transport.written-then-lost proves
the rule behind canResendAfterRestart's mutating arm, not a production
route. Over the real transport the arm cannot be reached, because
mutations go through sendRunnerCommandOnce and only the connect loop
raises the restart trigger.

Tests:
- runner-command-recovery: a failing status probe now asserts reason
  runner_reply_lost, recovery status_probe_failed, dispatched unknown, and
  the transport error as cause.
- runner-recovery-wiring: press/fill whose runner dies mid-command (real
  fake-runner transport) now also asserts reason runner_reply_lost.
Mutation: drop `reason: RUNNER_REPLY_LOST_REASON` from
buildStatusProbeFailedError -> all three fail.
…e connect loop

The read-only lost-response driver scripted 20 hang-ups and a 400 ms
timeout. That count depended on the connect loop's retry cadence
(RUNNER_CONNECT_ATTEMPT_INTERVAL_MS): with a faster cadence the queue runs
out, and the fake answers the next attempt with `ok`. The fake runner
server now has a `hangUpAlways` reply that is never consumed, so the read
is refused until the loop gives up, however often it retries. The timeout
now only bounds how long the loop runs.

Mutation: set RUNNER_CONNECT_ATTEMPT_INTERVAL_MS to 10 or 50 (with a 1 ms
retry base delay) -> ios-runner.status.read-only-completed-without-retained-reply
still passes.
The request handler of the fake runner server went over the fallow complexity threshold after hangUpAlways was added. The choice of the next scripted reply is now in takeScriptedResponse. Behavior is unchanged: the apple-runner suite passes (768 tests).
…e call site

A mutation never reaches the connect loop, so the no-resend arm after a
restart had no production route. The call site now restarts only when the
first attempt provably wrote nothing or the command is read-only, and the
arm and its error builder are gone. The written-then-lost row is driven
through a fake runner that dies mid-command. A failed status probe keeps a
transport reason under transportReason.
…lassifier

shouldRestartRunnerBeforeCommandSend now takes the command and owns the
write-proof rule: a connect-loop failure restarts and resends only when no
attempt could have written the command, or the command is read-only. An
exit-28 POST that carries dispatched unknown no longer restarts a mutation,
whatever call site asks. The call-site guard in executeRunnerCommandAttempt
is gone.
… through the restart

restartSessionAndRunCommand takes firstAttemptDispatched: DispatchDisclosure
instead of a boolean that three ternaries decoded back into no or unknown.
resolveFirstAttemptDispatch computes it once beside
isRunnerCommandProvablyUnwritten, and the restart rule reads the same answer.
The readiness-preflight call site passes no.
…ith one reason

Status recovery verdicts now carry either the runner's own answer or the
facts of a lost reply. applyRunnerTransportRecovery turns a lost reply into
one buildLostReplyError for a mutation, so runner_reply_lost is set on every
lost-reply verdict (completed without a retained reply, still in flight,
unknown or missing lifecycle, status probe failed, status unavailable) and
the transport's reason stays under transportReason. A read keeps its
transport error, which the read-only resend loop resends, so it never
carries the reason. The four hand-built copies are gone.
@thymikee
thymikee force-pushed the fix/runner-restart-resend-needs-proof branch from 9393095 to 4e88908 Compare October 1, 2026 07:31
@thymikee

thymikee commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

Head is now 4e88908fc3, rebased on feat/dispatch-disclosure a5ac453380.

1. The call-site guard never changed behavior (runner-lifecycle.ts L359): fixed in 2ffd18a.
The write-proof rule moved into its owner. shouldRestartRunnerBeforeCommandSend(error, command) restarts only when the connect failure provably wrote nothing, or when the command is read-only. An exit-28 POST (runner_connect_refused with dispatched: unknown) on a mutation now gets false from the classifier.

  • Deleted: the call-site conjunct (firstAttemptUnwritten || isReadOnlyRunnerCommand(command)) and its import (runner-lifecycle.ts +2/−8).
  • Tests: a classifier test where the same exit-28 error gives false for tap and true for snapshot. A lifecycle test sends tap through the real connect loop with curl exit 28. It asserts no restart, one tap send, and dispatched: unknown.
  • Mutation: making the classifier ignore the proof (return true) fails both tests.
  • Parity test: kept. The runner's read-only list still makes decisions: the restart rule above, the read/mutation split in recovery (item 3), and the read-only resend loop in runAppleRunnerCommand.

2. A boolean instead of the domain type (runner-lifecycle.ts L369): fixed in 66ef675.
restartSessionAndRunCommand now takes firstAttemptDispatched: DispatchDisclosure. resolveFirstAttemptDispatch(error) computes it once, next to isRunnerCommandProvablyUnwritten, and the restart rule reads the same answer. The readiness-preflight call site passes 'no'.

  • Deleted: the three ternaries. The recycle-budget throw and the no-restarted-session branch are now discloseDispatch(err, firstAttemptDispatched).
  • Test: a read whose first POST may have been written and whose restart fails discloses unknown.
  • Mutations: passing a constant 'no' at the connect call site fails that test. Passing 'unknown' at the readiness-preflight site fails the ios-runner.pre-send.readiness-preflight-after-restart row.

3. One lost-reply error, one reason (runner-command-recovery.ts L54, L489): fixed in 4e88908.
Each non-recovered verdict now carries one of two failures: runnerAnswer (the runner's own failure, read back from the journal) or lostReply facts (recovery, lifecycle state, message, hint). applyRunnerTransportRecovery resolves the failure in one place:

  • A mutation gets buildLostReplyError. This one builder owns the detail shape, reason: runner_reply_lost, and one rule for the transport's reason, which is always kept as transportReason.
  • A read gets its transport error, so the read-only resend loop resends it.
  • The reason now covers completed_without_retained_response, command_still_in_flight, and status_recovery_unavailable as well.

What changed:

  • Deleted: the inline unknown-lifecycle literal, the error built in handleCompletedRunnerStatus, runnerStatusInFlightError, buildStatusProbeFailedError, and the | undefined recovery result (runner-command-recovery.ts +133/−140).
  • Not covered: runner_reported_failure. That is the runner's own answer and keeps its own reason, for example runner_main_thread_timeout.
  • Behavior change for reads: a read whose status is unknown or notAccepted now rethrows its transport error and is resent, the same as every other lost read. Before, it failed with a non-retryable error.

Tests:

  • In runner-dispatch-disclosure.test.ts, every row's test asserts that reason === runner_reply_lost is true exactly for the lost-reply rows (completed-without-retained-reply, accepted, started, notAccepted, probe-failed, unavailable). The row ios-runner.transport.written-then-lost already asserts it in its own driver.
  • A test.each over four read outcomes (probe failed, notAccepted, started, completed without a reply) asserts no reason and a retryable error.

Mutations:

  • Removing the reason from the builder fails all 6 lost-reply rows.
  • Removing the read branch fails the read-only row and all 4 read cases.

ADR 0011 now names the ios-runner.status.* lost-reply rows.

Gate on 4e88908fc3: check:affected --run (fail-open full set) passed every runnable check. In vitest-related, 1333 of 1334 files passed. The one failure was a 5 s timeout under host load in an untouched file (screenshot-density.test.ts), which passes alone.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 8 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

Comment thread docs/adr/0011-interaction-guarantee-contract.md Outdated
Comment thread packages/platform-apple/src/runner/runner-command-recovery.ts Outdated
@thymikee

thymikee commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

Addressed cubic's latest four findings:

  • P2 runner-command-recovery.ts: a lost reply re-classified as retryable by its own text. 75a418f6b0. LostReply.message is now required, and status_probe_failed and status_recovery_unavailable get their own fixed message. The transport text stays in details.transportError and cause. RUNNER_REPLY_LOST_REASON moved into runner-error-classification.ts. A new first row in RUNNER_ERROR_RULES keys on details.reason === runner_reply_lost and sets retryable, drainResend, connectRetry, restartBeforeSend and restartAfterReadinessPreflight to false, so text a lost reply carries cannot resend it. New tests: (1) isRetryableRunnerError, shouldRetryRunnerConnectError and shouldRestartRunnerBeforeCommandSend are false for a lost reply that carries fetch failed, ECONNREFUSED or socket hang up text. (2) An end-to-end runAppleRunnerCommand tap whose send fails with fetch failed and whose status probe fails with ECONNREFUSED fails as runner_reply_lost, is not retryable, and sends the tap once. The runner-command-retry assertions that pinned the old message ('fetch failed') now assert the reason and details.transportError.
  • P2 ADR 0011: ios-runner.status.* is not only the lost-reply rows. 72f27a68d1. The ADR now names the six lost-reply rows and ios-runner.transport.written-then-lost. It also says that the ios-runner.status.failed* rows carry the runner's own answer.
  • P3 parity map: keyboardReturn. 28e492fa76. Changed to positionals: ['enter'] to match keyboardEnter in interactor.ts and the daemon's return → enter normalization.
  • P3 connect-loop mutation test does not tell a lost reply from a plain transport failure. 25ce56e7d5. It now asserts reason: runner_reply_lost and recovery: status_probe_failed.

Checks: typecheck, lint, format:check, layering and fallow audit pass, and the apple-runner project passes 767/767. A 5 s timeout failed under load and passed when run alone. In unit-core, 2 timeouts fail when run alone, and neither relates to this change: screenshot-density fails the same way at the base, and scripts/fuzz/harness is outside this diff.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 10 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="packages/platform-apple/src/runner/runner-command-recovery.ts">

<violation number="1" location="packages/platform-apple/src/runner/runner-command-recovery.ts:506">
P3: `lostReplyWithoutStatusMessage` rebuilds the same sentence skeleton as the inline message in `handleRunnerCommandStatusRecovery` (`lifecycle_state_not_recoverable`): both are `Runner command "…" lost its transport response and … so agent-device invalidated the runner session instead of replaying the command.`. Three lost-reply messages now share the identical structure, so the shared "invalidated instead of replaying" clause can drift between the two paths. Feed the lifecycle branch through the same builder, e.g. `lostReplyWithoutStatusMessage(command.command, lifecycleState ? `lifecycle status was "${lifecycleState}"` : 'lifecycle status was missing')`.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

return details;
}

function lostReplyWithoutStatusMessage(command: string, statusOutcome: string): string {

@cubic-dev-ai cubic-dev-ai Bot Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: lostReplyWithoutStatusMessage rebuilds the same sentence skeleton as the inline message in handleRunnerCommandStatusRecovery (lifecycle_state_not_recoverable): both are Runner command "…" lost its transport response and … so agent-device invalidated the runner session instead of replaying the command.. Three lost-reply messages now share the identical structure, so the shared "invalidated instead of replaying" clause can drift between the two paths. Feed the lifecycle branch through the same builder, e.g. lostReplyWithoutStatusMessage(command.command, lifecycleState ? lifecycle status was "${lifecycleState}" : 'lifecycle status was missing').

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/platform-apple/src/runner/runner-command-recovery.ts, line 506:

<comment>`lostReplyWithoutStatusMessage` rebuilds the same sentence skeleton as the inline message in `handleRunnerCommandStatusRecovery` (`lifecycle_state_not_recoverable`): both are `Runner command "…" lost its transport response and … so agent-device invalidated the runner session instead of replaying the command.`. Three lost-reply messages now share the identical structure, so the shared "invalidated instead of replaying" clause can drift between the two paths. Feed the lifecycle branch through the same builder, e.g. `lostReplyWithoutStatusMessage(command.command, lifecycleState ? `lifecycle status was "${lifecycleState}"` : 'lifecycle status was missing')`.</comment>

<file context>
@@ -504,6 +503,10 @@ function readReadinessPreflightRecoveryDetails(
   return details;
 }
 
+function lostReplyWithoutStatusMessage(command: string, statusOutcome: string): string {
+  return `Runner command "${command}" lost its transport response and ${statusOutcome}, so agent-device invalidated the runner session instead of replaying the command.`;
+}
</file context>
Fix with cubic

@thymikee

thymikee commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

Reviewed 25ce56e: no blocking issues. The new commits keep the restart-resend decision in one place, the connect-failure classifier. A lost reply now keys on its typed reason, not on transport text. A mutation whose request may have been written is never resent, and the first attempt's dispatched value survives the restart.

Not blocking, take or leave: the parity test in src/__tests__/runner-read-only-registry-parity.test.ts copies the readOnly values of nine runner-internal commands, so those rows compare the constant with itself; typing the counterpart rows as the registry's request type would also remove the as never cast. Also, buildLostReplyError replaces a boot-failure hint with the generic lost-reply hint for a mutation on a physical device; is that intended?

CI is still running (Smoke Tests, Repo Guards, Coverage), and iOS Smoke Tests run the runner route this changes. There are no conflicts. I read the code and did not run the tests; the lost-reply path is covered by the fake runner server only, with no live device run.

@thymikee
thymikee merged commit d37a9f4 into feat/dispatch-disclosure Oct 1, 2026
14 checks passed
@thymikee
thymikee deleted the fix/runner-restart-resend-needs-proof branch October 1, 2026 09:05
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-01 09:06 UTC

@thymikee

thymikee commented Oct 1, 2026 •

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

This PR merged into feat/dispatch-disclosure at 25ce56e7d5 before the three take-or-leave notes landed, so they are on #3071 instead (7f2d6d1c01, c679e60385, e6869a0827): the parity rows are typed from the registry's DispatchedCommand with both as never casts gone (the nine internal rows already asserted against a literal table); buildLostReplyError keeps the transport error's own hint and falls back to the lost-reply hint only when there is none (the replacement was not intended); the lost-reply sentence is built once from { command, cause }. Runner suite 823/823 on #3071's c6734bbf04.

thymikee added a commit that referenced this pull request Oct 1, 2026
…ost (#3082)

* fix(apple-runner): name the lost reply when status cannot place a lost command

A command whose reply was lost and whose status probe answers notAccepted
(a restarted runner's empty journal) or an unnamed state now fails with
details.reason runner_reply_lost beside dispatched unknown, so a caller
keys on the typed reason instead of the message before it observes the
screen.

Tests (fake runner server through runAppleRunnerCommand and the recovery
module):
- runner-recovery-wiring: press and fill (runner tap and type) accept the
  command, drop the connection, status answers notAccepted: one dispatched
  command, error dispatched unknown with reason runner_reply_lost, the
  session is invalidated, and the next command gets a runner. Mutation:
  delete `reason: RUNNER_REPLY_LOST_REASON` from the unknown-state error ->
  both rows fail.
- runner-recovery-wiring: a snapshot whose reply is dropped is resent and
  succeeds. Mutation: declare snapshot with DEFAULT_TRAITS (not read-only)
  -> fails (the lost read goes to status recovery and is not resent).
- runner-command-recovery: notAccepted yields reason runner_reply_lost and
  dispatched unknown. Mutation: same deletion as above -> fails.

* fix(apple-runner): restart without resending a command written then lost

A connect-shaped failure keeps its restart, but the command is sent again
on the restarted runner only with proof the first send could not have run
it: the attempt provably wrote nothing (a connect refused before the write,
dispatched no), the readiness preflight gave up before the write, or the
command is read-only by its runner trait. A mutating command whose first
attempt may have written it now fails with reason runner_reply_lost and
dispatched unknown after the restart; a restarted runner's journal is empty,
so no status answer can prove the first send did not run. A read-only
command's restart failure discloses no: a read has no side effect to repeat.

dispatch-disclosure.json: ios-runner.transport.written-then-lost keeps
unknown with a trigger that says the command is not resent;
ios-runner.transport.read-only-written-then-lost is new and discloses no.

Tests (runner-lifecycle-dispatch-disclosure, real connect loop over a
stubbed fetch and simctl curl exit 28, restart succeeds):
- written-then-lost: tap restarts the runner, is not sent on the restarted
  runner, fails unknown with reason runner_reply_lost. Mutation: make
  canResendAfterRestart return true -> fails (the tap is replayed and the
  request succeeds).
- read-only-written-then-lost: snapshot is sent again on the restarted
  runner and its failure discloses no. Mutations: drop the read-only term
  from canResendAfterRestart -> fails (no resend); drop the read-only
  branch of discloseRestartDispatch -> fails (unknown).
runner-command-retry: four mutating-restart fixtures now carry the proof a
real refused connect stamps (dispatched no); without it they would test the
removed replay.

* test(apple): share the unwritten connect refusal fixture to keep the retry test file within its size ratchet

* fix(apple-runner): report the restart transport fact for a read, not no

The restart path disclosed `no` for a read-only command whose first send
may have run, while the status-recovery path discloses `unknown` for the
same read (read_only_completed_without_retained_response, and the
notAccepted/unknown-state fallthrough). Both runner paths now report the
transport fact: a first attempt that may have written the command is
`unknown`, read or mutation. The read-only `no` has one owner, the daemon
router seam (request-dispatch-disclosure.ts), which keys it on the
registry's recordingEffect `observes-app` for every producer
(daemon.read-only-command). A second copy keyed on the runner's own
read-only trait reconstructed that rule from a different source of truth,
and the two runner paths disagreed. The recycle-budget refusal on the
restart path also stops deriving `no` from the read-only trait and uses
only whether the first attempt was provably unwritten.

The runner still resends a read on the restarted runner
(canResendAfterRestart); only the disclosure changes.

dispatch-disclosure.json:
- ios-runner.transport.read-only-written-then-lost: now `unknown`; the
  trigger names the router rule that turns it into `no` for the caller.
- ios-runner.status.read-only-completed-without-retained-reply: new,
  `unknown`, driven through the real connect loop and recovery module.
  It pins the status-recovery read to the same value as the restart read.

Mutations:
- put back `if (isReadOnlyRunnerCommand(evidence.command)) return
  discloseDispatch(error, 'no')` in discloseRestartDispatch ->
  read-only-written-then-lost fails.
- set `dispatched: 'no'` on read_only_completed_without_retained_response
  -> status.read-only-completed-without-retained-reply fails.

* test(apple-runner): drive a real runner death through the fake runner

runner-recovery-wiring: the notAccepted mutation test no longer scripts a
second `ok` for the mutation. Nothing consumed it, and it suggested that
a resend was expected. Its comment now says what the test models: status
answers notAccepted, the way a runner whose journal did not survive a
restart would. The path is status recovery, not the daemon's restart.

The fake runner server gains an `exit` reply: it hangs up and stops
listening, like a runner process that died mid-command. Before this, a
test could only hang up on one request while the same server stayed up.
Two new tests use a second server as the restarted runner:

- press/fill (runner tap/type): the runner dies on the command. The
  mutation is sent once, is not sent to the restarted runner, fails with
  dispatched unknown, and the dead session is invalidated. A readText on
  the next request is served by the restarted runner. Mutation: declare
  tap and type with READ_ONLY_TRAITS (the resend trait) -> both rows fail.
- get (runner readText): the runner dies mid-read. The connect loop gives
  up with a possibly-written POST (simctl curl exit 28), the daemon
  restarts the runner (runner_connect_failed_before_command_send), and the
  read is resent once and succeeds. Mutation: reduce canResendAfterRestart
  to `evidence.firstAttemptUnwritten` -> fails (no resend).

With the real transport, a mutating command never reaches
restartSessionAndRunCommand with a possibly-written first attempt.
Mutations are sent by sendRunnerCommandOnce, and only the connect loop
(waitForRunner, used for reads and the readiness preflight) stamps
runner_connect_refused. So the mutating arm of canResendAfterRestart is
covered only by the mocked lifecycle driver
(ios-runner.transport.written-then-lost). A dying runner sends a mutation
to status recovery, and a failed status probe rethrows the transport
error without a runner_reply_lost reason.

* fix(apple-runner): name the lost reply when the status probe fails

A runner that dies mid-mutation drops the reply, and the status probe that
follows fails. The error that reaches the caller was then a bare transport
error (`fetch failed`, dispatched unknown), with no typed reason. #3074
requires a reason that names the lost reply, so this path now builds an
error with reason runner_reply_lost (the name the restart arm already uses)
and recovery status_probe_failed. The error keeps the transport message
and details, and the transport error as cause. The session is still
invalidated, and the command is still not resent.

ADR 0011 gap list: the row ios-runner.transport.written-then-lost proves
the rule behind canResendAfterRestart's mutating arm, not a production
route. Over the real transport the arm cannot be reached, because
mutations go through sendRunnerCommandOnce and only the connect loop
raises the restart trigger.

Tests:
- runner-command-recovery: a failing status probe now asserts reason
  runner_reply_lost, recovery status_probe_failed, dispatched unknown, and
  the transport error as cause.
- runner-recovery-wiring: press/fill whose runner dies mid-command (real
  fake-runner transport) now also asserts reason runner_reply_lost.
Mutation: drop `reason: RUNNER_REPLY_LOST_REASON` from
buildStatusProbeFailedError -> all three fail.

* test(apple-runner): hang up on every read attempt without counting the connect loop

The read-only lost-response driver scripted 20 hang-ups and a 400 ms
timeout. That count depended on the connect loop's retry cadence
(RUNNER_CONNECT_ATTEMPT_INTERVAL_MS): with a faster cadence the queue runs
out, and the fake answers the next attempt with `ok`. The fake runner
server now has a `hangUpAlways` reply that is never consumed, so the read
is refused until the loop gives up, however often it retries. The timeout
now only bounds how long the loop runs.

Mutation: set RUNNER_CONNECT_ATTEMPT_INTERVAL_MS to 10 or 50 (with a 1 ms
retry base delay) -> ios-runner.status.read-only-completed-without-retained-reply
still passes.

* test(apple-runner): pick the fake runner's next reply in one helper

The request handler of the fake runner server went over the fallow complexity threshold after hangUpAlways was added. The choice of the next scripted reply is now in takeScriptedResponse. Behavior is unchanged: the apple-runner suite passes (768 tests).

* refactor(apple-runner): gate the restart resend at the connect-failure call site

A mutation never reaches the connect loop, so the no-resend arm after a
restart had no production route. The call site now restarts only when the
first attempt provably wrote nothing or the command is read-only, and the
arm and its error builder are gone. The written-then-lost row is driven
through a fake runner that dies mid-command. A failed status probe keeps a
transport reason under transportReason.

* test(apple-runner): name runner_reply_lost in the lost-reply test titles

* test(apple-runner): the runner read-only set matches the registry's observes-app commands

* fix(apple-runner): decide the restart resend in the connect-failure classifier

shouldRestartRunnerBeforeCommandSend now takes the command and owns the
write-proof rule: a connect-loop failure restarts and resends only when no
attempt could have written the command, or the command is read-only. An
exit-28 POST that carries dispatched unknown no longer restarts a mutation,
whatever call site asks. The call-site guard in executeRunnerCommandAttempt
is gone.

* refactor(apple-runner): carry the first attempt's dispatch disclosure through the restart

restartSessionAndRunCommand takes firstAttemptDispatched: DispatchDisclosure
instead of a boolean that three ternaries decoded back into no or unknown.
resolveFirstAttemptDispatch computes it once beside
isRunnerCommandProvablyUnwritten, and the restart rule reads the same answer.
The readiness-preflight call site passes no.

* refactor(apple-runner): build every lost-reply failure in one place with one reason

Status recovery verdicts now carry either the runner's own answer or the
facts of a lost reply. applyRunnerTransportRecovery turns a lost reply into
one buildLostReplyError for a mutation, so runner_reply_lost is set on every
lost-reply verdict (completed without a retained reply, still in flight,
unknown or missing lifecycle, status probe failed, status unavailable) and
the transport's reason stays under transportReason. A read keeps its
transport error, which the read-only resend loop resends, so it never
carries the reason. The four hand-built copies are gone.

* fix(apple-runner): key a lost reply on its typed reason, never on transport text

* docs(adr): name the lost-reply dispatch-disclosure rows in ADR 0011

* test(apple-runner): pair keyboardReturn with the keyboard enter request that issues it

* test(apple-runner): pin the lost-reply branch in the connect-loop mutation test
thymikee added a commit that referenced this pull request Oct 1, 2026
#3071)

* feat(errors): disclose whether a failed interaction reached the device

details.dispatched carries no, yes, or unknown on interaction failures,
anchored on the ADR 0014 side-effect seam: a failure before the seam is
no, one after it is unknown unless a producer proves yes. Producers: the
iOS runner reply and status-probe verdict tables, Android adb input and
helper gestures, the Android post-tap in-app guard. The rows live in
contracts/fixtures/dispatch-disclosure.json and every row is driven
through its real producer by an owning test, gated for coverage.

* fix(errors): only a pre-dispatch refusal discloses dispatched no

The daemon fallback now fills unknown for any failure no producer
classified; the session runtime revision was not a sound witness (it does
not advance for sessionless requests, a replaced session resets it, and
several side effects never advance it). Target resolution, keyboard
occlusion, and targeted-touch admission stamp no where they refuse, and
the covered-target refusal gains details.reason target_covered.

* feat(apple-runner): publish runner_busy and runner_main_thread_timeout as details.reason

RUNNER_BUSY and MAIN_THREAD_TIMEOUT reached the wire only as
details.runnerErrorCode. The one runner-code classifier now also sets
details.reason, so a consumer reads a single field; runnerErrorCode stays.

* fix(ios): fall back from a direct selector tap only when it provably did not dispatch

isDirectIosSelectorFallbackError matched message text (fetch failed, timed
out, runner did not accept connection, invalid runner response), several
of which arrive after the tap already ran; the tree path then tapped again.
It now keys on details.dispatched === 'no'. The runner discloses no for a
failure before the exchange, a pre-send recovery verdict (connect refused,
readiness preflight, RUNNER_BUSY), a failed restart, a spent recycle
budget, and its selector refusals (ELEMENT_NOT_FOUND, ELEMENT_OFFSCREEN,
AMBIGUOUS_MATCH).

* fix(errors): disclose dispatched yes where a producer ran the action on purpose

scroll_no_progress is raised after the scroll gesture ran, and an Android
fill whose verification fails after the last clear-and-retype pass typed
the text; both now say yes instead of falling to the daemon's unknown.

* fix(errors): claim no only for a connect that provably wrote nothing, and for read-only commands

A runner connect failure classified for replay-after-restart said no even
when an attempt had already posted the command (a fetch deadline, a simctl
curl that exited after the POST). The connect loop now records whether
every attempt was refused before writing (ECONNREFUSED, a usbmux socket
that never opened, curl exit 7, a pre-attempt deadline), and only that
failure is a pre-send refusal; the restart path reports unknown when the
first attempt may have written. The daemon fallback stamps no for a
command the registry declares read-only (recordingEffect observes-app).

* docs(interaction): explain dispatched and the runner refusal reasons

* refactor(apple-runner): the transport discloses dispatch itself instead of a write-evidence flag

The runner transport is the only code that knows whether request bytes
left, so it now stamps the public `details.dispatched` where it throws,
replacing the private `runnerCommandUnwritten` details flag:

- usbmux socket refusal, per-attempt connection deadline, simctl curl
  exit 7: `dispatched: no`; simctl curl with any other exit: `unknown`.
- the connect loop's aggregate: `unknown` (overwriting) when any attempt
  may have written the command, else fill-if-absent `no`.
- isRunnerCommandProvablyUnwritten reads `dispatched === 'no'` or a
  refused connection; markRunnerCommandUnwritten is deleted.

Mutation checks: removing both `no` stamps (aggregate fill and curl
exit 7) fails the golden row
ios-runner.pre-send.connect-refused-before-write; stamping curl exit 7
as `unknown` fails it too; dropping the aggregate `unknown` overwrite
fails the new waitForRunner test for a written attempt followed by a
refused curl fallback.

* refactor(contracts): drop the unread phase column from the dispatch-disclosure table

No consumer read the before-seam/after-seam phase beyond a vocabulary membership check; the daemon seam does not infer a verdict from it. Removes the column from all 46 rows, the row type, DISPATCH_DISCLOSURE_PHASES, and the check.

* fix(daemon): a read-only command discloses dispatched no over any producer verdict

`dispatched` answers whether a resend is safe. A command the registry
declares read-only (`recordingEffect: 'observes-app'`) has no side
effect, so the interaction seam now sets `no` for it over a producer's
`unknown` or `yes`, on both the thrown and the response path. Mutations
keep the fill-if-absent `unknown`. The seam is renamed
discloseInteractionDispatch since it no longer only fills.

The golden row daemon.read-only-command now states it applies even when
a producer classified the failure. ADR 0011 and the commands docs name
the rule.

Mutation checks: reverting the response path to fill-if-absent fails
"a read-only command discloses no over a producer verdict" (get text
whose capture throws dispatched unknown/yes); reverting the thrown path
fails "a read-only command discloses no over a producer verdict it
throws".

Consumers unaffected: direct-ios-selector's fallback check and the
runner lifecycle read `dispatched` inside the producer, before the seam,
and only for tap/interaction commands.

* feat(errors)!: dispatched is two-valued, no or unknown

No consumer acts differently on `yes` and `unknown`: the docs said
snapshot on both, and the only branching consumer (the direct iOS selector
fallback) keys on `no`. No producer can prove execution on its failure
path either: the runner documents XCTEST_RECORDED_FAILURE as "the action
may not have been performed", the Android helper's ok=false comes from an
outer catch that also covers pre-injection failures, and WebDriver has no
execution receipt. `DispatchDisclosure` is now `'no' | 'unknown'`.

Every former `yes` producer:
- runner reply table and status recovery: an unlisted or missing code,
  `completed` without a retained reply, and a journal `failed` without a
  refusal code are `unknown`.
- post-action guard: stamps nothing; the daemon seam fills `unknown`.
- Android helper ok=false (one-shot and session) and fill verification:
  stamp `unknown` themselves, so their platform-level drivers pin the
  verdict without the daemon seam. The helper result carries no count of
  events injected before the failure, so it cannot say `no`.
- scroll_no_progress stamps `unknown` itself: `scroll` routes through the
  generic runtime, not the interaction seam, so nothing would fill it.

Golden table: every row keeps its scenario; each `yes` verdict is now
`unknown`, and triggers that claimed execution say what is proven.
`ios-runner.reply.executed-failure` is renamed `unlisted-code`.

Mutations: runner default `?? 'unknown'` -> `?? 'no'` fails
ios-runner.reply.{MAIN_THREAD_TIMEOUT,unlisted-code} and status.failed;
dropping the seam's `discloseUnclassifiedDispatch(error, 'unknown')` fails
post-action-guard.android-press-left-app; dropping the scroll_no_progress
stamp fails daemon.scroll-no-progress.

* fix(errors): a later step's refusal never claims no for the whole series

`dispatched: no` describes the whole requested operation. A series that
issues more than one device-reaching step (press --count split into runner
sequence chunks, swipe --count, adb text chunks) let a later step's `no`
(RUNNER_BUSY on chunk 2, a refused connect after the runner crashed) stand
for the whole interaction, so a consumer would resend the series and the
app would get extra taps.

One kernel helper owns the rule: `discloseDispatchAfterSteps(error,
dispatchedSteps)` keeps `no` only while no step was dispatched, and
otherwise stamps `unknown` with `details.dispatchedSteps` (added to any
count the failing step carries, so nested series compose). The local
Android copy is gone; `discloseAdbInputDispatch` now classifies one input.

Series loops, enumerated by
  git grep -n -A22 -E '^[[:space:]]*(for|while) \(' -- packages src \
    | LC_ALL=C grep -a -E '^(packages/[^/]+/src|src)/' \
    | LC_ALL=C grep -a -v -E '\.test\.ts|__tests__' \
    | LC_ALL=C grep -a -E 'await (.*[^a-zA-Z])?(runCommand|interactor\.(tap|longPress)|doubleTap\.call|interactions\.gesture|runAndroidShell|runAdbShell|sendAndroidImeHelperText|clearAndroidImeHelperText|typeAndroidShellChunk|focusAndroid|scroll|runStep|runJson|typeText|runAlert)\('
and each hit that is an interaction series now applies the helper:
- platform-apple runner-sequence.ts runApplePressSeries (sequence chunks)
- contracts touch-runtime.ts executeGenericPress (press --count on
  Android, Linux, and every generic interactor)
- daemon interaction-gesture.ts runSwipeRepetitions (swipe --count)
- platform-android text-input.ts: fillAndroid (focus tap, then a pass),
  fillAndroidImeHelper, typeAndroidImeHelper, typeAndroidShell,
  clearFocusedText
- platform-android device-input-state.ts keyboard dismiss keyevents
- capture-kit scroll-edge-state.ts and daemon scroll-until.ts scroll passes
- maestro daemon-runtime-port-observation.ts scrollUntilTypedMaestroTarget
- platform-linux runPacedScrollSteps, platform-web runPacedScroll,
  provider-limrun per-character typeText
Left as is: platform-apple alert.ts actOnAppleAlert retries only after an
ALERT_NOT_FOUND refusal (`no`), so no earlier attempt dispatched; alert.ts
84 is a read poll. The other hits (app lifecycle, settings, permissions,
notifications, perf, macOS host provider, ime-helper broadcast extras) are
not interactions or are calls after the loop. WebDriver actions land with
their own package branch.

Table rows and drivers: ios-runner.series.later-chunk-refused (press
--count 25 over the real send stack; chunk 2 answers RUNNER_BUSY ->
unknown, dispatchedSteps 1) and daemon.series.swipe-later-repetition-
refused (swipe --count 2 through the daemon handler; repetition 2 fails
`no` -> unknown, dispatchedSteps 1).

Mutations: `throw error` instead of the helper in runApplePressSeries
fails ios-runner.series.later-chunk-refused; the same in
runSwipeRepetitions fails daemon.series.swipe-later-repetition-refused;
dropping the `dispatchedSteps === 0` guard in the helper fails its kernel
test.

* fix(apple-runner): unlisted runner codes are unknown; pre-gesture refusals are no

An unlisted or missing runner code already reads `unknown` (the default
became `unknown` with the two-valued type), in the reply table and in the
status-recovery journal path, which share classifyRunnerReportedError.
This adds the one code the runner provably emits before any gesture:

INVALID_ARGS -> no. Every emission in the runner sources
(apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/):
- RunnerTests+Transport.swift:23, :106, :169: request body too large,
  not UTF-8, or undecodable; nothing is dispatched.
- RunnerTests+CommandDispatch.swift:33, :59, :78: status, appState,
  pasteboardWrite argument guards, before the read or the pasteboard write.
- RunnerTests+CommandExecution.swift:256, :296, :316, :346 (scroll and
  desktopScroll direction and duration guards), :290 and :336 (no gesture
  or wheel plan), :593 and :599 (missing or invalid gesturePlan): each
  returns before executeScrollDragGesture, desktopScrollAt, or
  plannedGestureExecution.
- RunnerTests+ScrollDragExecution.swift:86 (non-finite drag coordinates)
  and :159 (no synthesized coordinate frame), before synthesis or dragAt.
- RunnerTests+SequenceExecution.swift:25, :28, :130, :135 (via :219):
  sequence validation, before assembleSequenceExecution at :46.

Not listed, with the evidence:
- UNSUPPORTED_OPERATION stays `unknown`. It is also the reply after a
  gesture ran: RunnerSynthesizedGesture.m:350-353 returns "private XCTest
  event synthesis failed" after synthesizeWithError: ran, and
  RunnerTests+TvRemote.swift:119-126 returns "element tap failed" after
  element.tap() raised. The row ios-runner.reply.UNSUPPORTED_OPERATION pins it.
- "Active app interaction viewport is unavailable"
  (RunnerTests+CommandExecution.swift:628) is a pre-gesture refusal, but
  its code is COMMAND_FAILED; mapping it would key on message text. It
  stays `unknown` until the runner gives it its own code.

Mutations: removing the INVALID_ARGS entry fails
ios-runner.reply.INVALID_ARGS (driven through the real reply path against
the fake runner); adding UNSUPPORTED_OPERATION as `no` fails
ios-runner.reply.UNSUPPORTED_OPERATION.

* fix(daemon): every failed interaction carries dispatched, plain errors included

The interaction seam and the runner lifecycle both rethrew a failure that
was not an AppError untouched, so a plain Error from a backend reached the
wire without `details.dispatched`. Both now normalize at the boundary with
asAppError (fallback code UNKNOWN, the same code the request router's
normalizeError gives a plain Error, so the wire code does not change)
before disclosing: the seam fills `unknown` (or `no` for a read-only
command), and the runner lifecycle keeps its `no` for a failure before the
exchange started.

Tests: a plain Error thrown by the bound runtime reaches the caller of
handleInteractionCommands as an AppError whose normalized details say
`unknown`; a plain Error from ensureRunnerSession reaches the caller of
runAppleRunnerCommand with `no` and no command sent.

Mutations: restoring the raw rethrow for plain Errors in the seam fails the
first; restoring `if (!(error instanceof AppError)) throw error` in
executeRunnerCommand fails the second.

* fix(errors): close the smaller dispatch-disclosure gaps from review

- adb text: only a non-empty chunk is a dispatched step, so `type "\nx"`
  with adb missing fails its leading ENTER with `no` instead of `unknown`.
  Row android-adb.input-text.tool-missing-before-first-input.
- The daemon.unclassified.session-replaced row is dropped: its driver
  replaced the session inside the touch, after the runtime had read it,
  and the seam reads only the failure (interaction-dispatch-disclosure.ts
  has no reference to the session store), so no placement of the
  replacement can change the verdict. daemon.unclassified covers the case.
- waitForRunner: the early-exit error after the loop gets the same write
  evidence as the connect error, so a runner that exited before any
  attempt could write says `no`.
- fill: admitFill wraps the whole preamble (surface policy, target parse,
  bind, recorded-parameter and settle-flag checks, @ref preamble); a
  refusal it returns or throws is `no`. There is no single admission seam
  across interaction commands (press has admitTargetedTouch, fill had none),
  so the fill preamble is wrapped as a whole, and prepareTouchDispatch, the
  bind every touch route shares, marks its own refusal. Row
  daemon.refusal.fill-admission.
- Driver ownership compares slash-separated paths on every platform.
- fill: the second-pass input failure is its own row,
  android-adb.fill.second-pass-input-failed (unknown, dispatchedSteps 2),
  driven through fillAndroid; android-adb.fill.unverified keeps the
  verification-exhausted case.
- Table checks: fallsBack exactly on maestro-direct rows, implementedBy
  only on webdriver rows, values only no or unknown, and every driver file
  carries the per-row loop DISPATCH_DISCLOSURE_ROW_LOOP.
- Tests: the overwrite test overwrites a seam-filled value; the direct
  iOS fallback pins that each legacy message falls back with `no` and not
  without it; the direct-iOS suite resets its capture mock per test; the
  refused-fallback transport test asserts the fallback ran once.

Mutations: counting an empty chunk as a step fails
tool-missing-before-first-input; dropping the focus-tap count in
fillAndroid fails second-pass-input-failed; returning admitFill's refusal
unwrapped fails daemon.refusal.fill-admission; throwing the bare early-exit
error fails the new waitForRunner test; a fallsBack or implementedBy on a
daemon row fails the vocabulary test; skipping a driver's per-row loop
fails the driver-file test; a message sniff returning false for legacy
texts fails the direct iOS fallback test. The path-separator fix has no
POSIX mutation.

* refactor(interaction): split the series and fill admission below the complexity gate

fallow flagged typeAndroidImeHelper, executeGenericPress, and
readFillAdmission after the series and admission changes. The two Android
text loops, which had the same shape, now share one plan
(planAndroidTextSteps) and one sender (sendAndroidTextSteps) that owns the
dispatched-step count; the ENTER between lines goes through
pressAndroidEnterKey, so a missing adb before any input is `no` on the IME
path too. executeGenericPress delegates one press to pressOnce, and fill
admission reads its session and target in readFillTarget.

Behavior is unchanged otherwise: pauses fall after every chunk except the
last of the text, and an empty line still sends nothing. The existing
Android dispatch-disclosure, text-input, and fill suites pass unchanged;
counting an empty chunk as a step in sendAndroidTextSteps still fails
android-adb.input-text.tool-missing-before-first-input.

* chore(gates): re-derive the DispatchDisclosure wire digest for the two-valued type

`DispatchDisclosure` is new on this branch (unreleased, so no ack), and
its declaration changed to `'no' | 'unknown'`; the digest is the one
test/wire-compat/wire-compat.test.ts prints from declaration-digest.ts.
The DaemonError and NormalizedError acks keep their digests; their
rationale now names the two values.

* fix(errors): a failure after the mutation was sent never keeps a producer's no

`dispatched: no` may describe a failure only while nothing reached the device.
Two owners now record each mutation send that returned and downgrade any later
failure of the same command through the kernel rule (`no` -> `unknown`, sent
count added to `details.dispatchedSteps`):

- Daemon: a request-scoped `RequestDispatchLedger`. The sends are the ADR 0014
  side-effect seams of the interaction route (every `expireRefFrame` there):
  the runtime backend touch members and `performGesture`
  (interaction-runtime.ts), the direct iOS selector tap, and `type`.
  `discloseRequestDispatch` reads the ledger. The swipe series counter is
  superseded by the ledger and removed.
- Maestro port: every `MaestroPublicOperation` kind is classified in
  `MAESTRO_OPERATION_MUTATES` (exhaustive record); `invokeMaestroPublicOperation`
  counts a mutating send that returned ok, and each port operation (one Maestro
  command) discloses the sends it made. This replaces the scrollUntilVisible
  counter, which counted only after the settle, and also covers inputText,
  eraseText, and tapOn retries.

Post-dispatch phases enumerated: `grep -rn expireRefFrame src/daemon/interaction`
(4 send seams); after each: afterRun escape guard, Android dialog readiness
after-phase, response payload build, finalize. The runtime `--settle`/`--verify`
observation (post-action-observation.ts, settle.ts) is best-effort and never
throws, so it cannot carry a `no` out; the escape guard's foreground read can.

Rows: daemon.post-dispatch.press-then-foreground-read-refused,
maestro-port.scroll-until.capture-after-scroll-refused,
maestro-port.input-text.settle-capture-refused,
maestro-port.scroll-until.first-capture-refused (control: no send, no stands).
Mutations: dropping the ledger increment in `sendRecordedMutation` fails the
press and swipe rows; dropping the increment in `invokeMaestroPublicOperation`
fails both maestro-port unknown rows.

* fix(errors): composite and helper-IME steps disclose through the same helper

A `no` from a later sub-step of one interactor operation no longer describes
the whole operation.

Composite enumeration: interactor operations that issue more than one
device-reaching send whose producer can say `no`. The `no` producers are
(grep "discloseDispatch(.*'no'"): adb input (`discloseAdbInputDispatch`), the
Apple runner, and daemon refusals. Their multi-send operations:
- Android `doubleTap` (src/core/interactors/android.ts): two `input tap` sends,
  uncounted. Now `doubleTapAndroid` in input-actions.ts discloses a failure of
  the second tap through `discloseDispatchAfterSteps(error, 1)`.
- Android type/fill/clear, keyboard dismiss, Apple press series and sequence
  chunks: already counted on this branch.
- Apple `doubleTap`, limrun and WebDriver `doubleTap`, HarmonyOS doubleClick:
  one send each. Linux doubleClick sends several xdotool/ydotool calls, but no
  Linux producer says `no`, so nothing can leak.

Helper IME: every broadcast goes through `sendAndroidImeHelperBroadcast`, which
now classifies its failure with `discloseAdbInputDispatch` (TOOL_MISSING before
start is `no`, anything after start `unknown`), the same classifier the shell
path uses. The classifier moves to adb-failure.ts so ime-helper.ts need not
import the input-action module.

Rows: android-adb.input-tap.double-tap-second-refused,
android-helper.ime.broadcast-tool-missing, android-helper.ime.broadcast-failed.
Mutations: rethrowing the raw error in `doubleTapAndroid` fails the double-tap
row (`no`); removing the classification in the broadcast sender fails both IME
rows (unset).

* fix(daemon): every mutating command route discloses dispatch

The dispatch seam was wired only into the interaction route, so a generic
`scroll` (and every other non-interaction command) could fail with
`dispatched` unset. The request router now owns the seam: `executeLockedRequest`
wraps the handler chain and the generic dispatcher in `discloseRequestDispatch`,
keyed by the registry `recordingEffect` (resolveCommandRecordingEffect):
`observes-app` is `no`; `mutates-app` keeps a producer verdict, downgrades a
`no` after a sent mutation, and fills `unknown` otherwise; a command without a
declared effect passes through. `handleInteractionCommands` no longer wraps
itself, so the seam runs once per request (a second pass would double the
ledger count).

The router also owns the ledger: it passes it through the handler chain to the
interaction route and records the generic route's bound `execute` as its
mutation send. A generic admission refusal is `no`, like the interaction
admission refusals; the pre-dispatch helpers move next to the seam.

`scroll --until` now proves its own pre-gesture refusals: a failure before the
first gesture, other than the gesture send's own, is `no`.

Residual: failures before the router seam (session lookup, lock, lease
admission in prepareLockedRequestScope) stay unset. Commands outside the
interaction and generic routes (open, close, install, settings, alert,
react-native) get the fill but record no ledger sends; their post-mutation
steps can still carry a producer `no` if one appears there.

Rows (src/daemon/__tests__/request-dispatch-disclosure.test.ts, real router):
daemon.route.scroll-transport-failure, daemon.route.scroll-until-before-first-gesture,
daemon.route.scroll-then-dialog-read-refused, daemon.route.read-only-get.
Mutations: routing without `discloseRequestDispatch` fails the scroll, Android
scroll, and get rows; dropping the ledger wrap on `execute` fails the Android
scroll row; rethrowing in scroll-until fails the --until row.

* fix(apple-runner): pre-send validation and restart replay disclose their verdict

Sequence validation ran inside the transport catch of runApplePressSeries, so
an invalid step (a non-finite point) reached the wire undisclosed and the daemon
filled `unknown`. Every validation refusal now comes from one constructor,
`invalidSequence`, which carries `reason: runner_sequence_invalid` and
`dispatched: no`; the series builds and validates every chunk before it sends
the first, so a refusal can never follow a sent chunk.

A restart replay after an unwritten first attempt returned the replay failure
as is, so a replay failure without producer evidence escaped undisclosed. It
now keeps an existing verdict, is `no` for a pre-send refusal
(isRunnerPreSendRefusal), and is `unknown` otherwise.

Rows: ios-runner.series.invalid-step-before-send,
ios-runner.transport.replay-failed-after-unwritten-first-attempt.
Mutations: `invalidSequence` disclosing `unknown` fails the invalid-step row;
passing the replay failure through fails the replay row.

* test(daemon): type the route driver's dialog-read stub as the adapter declares it

* refactor(daemon): the request scope owns the dispatch ledger; every bound mutation records itself

The ledger was threaded route by route (InteractionRouteInput,
InteractionRuntimeInput, RequestHandlerChainParams, executePlatformCommand), so
`find <q> type`, session, snapshot/alert, react-native, and record/trace routes
never recorded their sends, and the Maestro port kept a second ledger plus
MAESTRO_OPERATION_MUTATES, which restated the registry.

The ledger now lives on the request execution scope and is handed to
createRequestRuntimeBindings, the one place bindDevice and bindExactDevice
narrow a binding. Every narrowed projection wraps each operation that
RUNTIME_OPERATION_EFFECTS (src/daemon/runtime-operation-effects.ts) declares
`mutates`, and records it when its send returns. The table is a Record over the
runtime operation keys, so an operation with no declared effect fails typecheck
and the completeness test. A replay action runs on a ledger of its own
(req.internal.dispatchLedger) and moves its sends into the parent's.

The router discloses around the scope's ledger: once it holds a sent mutation,
a failure is `unknown` with the count added to details.dispatchedSteps, for any
declared effect or none; a read with an empty ledger stays `no`.

Deleted: the route dispatchLedger fields, sendRecordedMutation,
MaestroMutationLedger, MAESTRO_OPERATION_MUTATES, and the per-command Maestro
wrapper. The three maestro-port rows become one router row.

Rows: daemon.route.find-type-refused-after-focus,
daemon.route.session-close-then-finalize-refused,
daemon.route.maestro-deferred-settle-capture-refused.
Mutations: dropping recordBoundMutations from bindDevice fails the find-type,
close, Maestro, and scroll-then-dialog-read rows; dropping the nested ledger
from the replay action invoker fails the Maestro row; deleting one entry from
RUNTIME_OPERATION_EFFECTS fails typecheck and the completeness test.

* fix(daemon): a batch reports no only when no executed step was a mutation

Each batch step ran as a nested request on a fresh ledger, and the batch
descriptor declares no recording effect, so the router passed the batch
response through with the failing step's `dispatched: no`. `[press A (lands),
press 'label=Missing']` reported `no` with `executed: 1`, and a retry pressed A
twice.

The router now hands every nested request it invokes (batch steps, find's
delegated click and fill) a ledger of its own through recordNestedRequests and
moves its sends into the parent's ledger, as it already did for replay actions.
The batch failure then reports `unknown` with the executed steps' sends in
details.dispatchedSteps. A failed nested response keeps only the steps its
producer counted inside the failing send, so the parent counts each send once.

Rows: daemon.route.batch-mutation-then-refused-step,
daemon.route.batch-read-then-refused-step.
Mutation: passing handleRequest as the chain's invoke fails the
mutation-then-refused row.

* refactor(provider-webdriver): disclose through the kernel's DispatchDisclosure type

The transport declared its own 'no' | 'unknown' union because the kernel type
had not landed when #3070 merged. After the rebase both exist; the transport now
uses the kernel's DispatchDisclosure, so the vocabulary has one declaration.

* test(provider-webdriver): drive the three WebDriver disclosure rows through the real transport

The webdriver rows waited on #3070 behind `implementedBy`. With it merged, the
transport test owns them (`webdriver.` prefix): each sends a mutating POST
through WebDriverTransport and the real fetch to a local socket. A port nothing
listens on is refused before send (`no`); a driver that reads the body and never
answers hits the transport deadline after send (`unknown`); a driver that
answers 503 received the request (`unknown`). The row markers are gone and the
triggers name what each driver does.

The contracts fixtures subpath is a test-only import of the provider package;
the eager-closure budgets are unchanged.

Rows: webdriver.connect-refused, webdriver.timeout-after-send,
webdriver.http-5xx-after-send.
Mutations: the connect-refused error disclosing `unknown` fails the refused
row; the timeout error disclosing `no` fails the timeout row; dropping the
>= 500 disclosure fails the 5xx row.

* docs(adr): state the read-only rule and the remaining disclosure gaps

ADR 0011 now describes the one request ledger: operation effects declared once
in RUNTIME_OPERATION_EFFECTS, bound operations recording their own sends, and
nested requests moving theirs into the parent's. `no` for a read is defined as
"a resend repeats no app-visible action": the registry has no trait that
separates a pure read from a device control, so record, trace, and perf stay
`observes-app`, and their recorder and profiler controls are declared
`repeatable` at the operation level. The section lists the remaining gaps:
failures before the router's locked scope (session, lock, lease, and policy
admission) and the public runBatch with a caller-supplied invoke.

The rewrite also removes the code spans that were hard-wrapped mid-token
(`'mutates- app'`, `'observes- app'`) and the Maestro-port and
`implementedBy` sentences, which no longer apply. commands.md states the
read rule and the batch/replay rule.

* fix(apple-runner): a press series counts completed taps, not sequence chunks

details.dispatchedSteps of runApplePressSeries counted sequence chunks while
every other series counts device inputs. It now counts the taps the runner
reported completed: across finished chunks when a later chunk is refused, plus
the failing chunk's own completed steps when a step inside it fails. The
ios-runner.series.later-chunk-refused row now asserts 20, and its trigger says
so. The runner disclosure test header names its third category: pre-send
validation that refuses before anything is sent.

Mutations: counting chunks fails the later-chunk row (1, not 20); dropping the
failing chunk's completed steps fails the new runner-sequence test (20, not 22).

* chore(gates): drop the implementedBy row marker and its check

The webdriver rows were the only rows naming another branch; they now have a
driver, so the field, its ownership filter, and the gate assertion that only
webdriver rows could carry it have no remaining use. Every row must now have a
driver file.

* fix(batch): report unknown after an executed mutating step

* fix(interaction): readiness refusals disclose dispatched no

The readiness wait now refuses before any touch on main: a sparse capture
(capture_sparse) stamps no where it is thrown, and an exhausted budget
inherits no from the selector failure it decorates. A daemon table row
drives the exhausted wait through the real readiness path.

* fix(daemon): scroll loops leave counting bound sends to the request ledger

Each scroll --until and edge pass runs the bound scrollDirection, which the
request ledger already records, and the router adds the ledger to any
producer count. The two loops counted the same sends again, so two passes
reported four. Only a producer counting sub-steps inside one bound
operation keeps its own counter. Two router rows pin exactly 2 after a
refused third pass.

* refactor(errors): one dispatched-steps rule, applied to details

discloseDispatchAfterSteps now wraps detailsAfterDispatchedSteps instead
of restating its rule. withoutLedgerSteps stays: a nested request's own
router pass adds its ledger to the failure, which the parent then adds
again through its own ledger.

* refactor(daemon): unexport two types only their own modules read

* fix(batch): a batch reports no only when every executed step is a declared read

* docs(adr): state the batch dispatched rule as observes-app reads only, pin undeclared step

* fix(daemon): delete the no-change interaction retry path (#3083)

* fix(daemon)!: delete the no-change interaction retry path

The daemon re-sent a coordinate tap when the next capture found the
surface unchanged, if the request carried the interactionOutcome
retry-on-no-change opt-in. Nothing in the repository set it, and the
resend could not know whether the first tap was dispatched, so a tap that
landed on a screen that had not re-rendered yet ran twice.

Removed:
- `interactionOutcome` from `CommandFlags`. No CLI token, env var, config
  key, Node option or MCP field ever carried it; the daemon now has no
  reader, and the session recorder keeps only declared flags, so a raw
  key sent by a peer is dropped.
- `SessionState.pendingInteractionOutcome` and its type (the R7 owner
  entry leaves in the chore(gates) commit).
- The pending-outcome branch of `resolveDeferredInteractionOutcome`, the
  `scheduleOutcomeRetry` mark, and every retry helper in
  interaction-outcome-policy.ts (`markPendingInteractionOutcome`,
  `retryPendingInteractionOutcome`, `InteractionRetryTap`,
  `classifyInteractionSurfaceChange`, the settle diagnostics).
- interaction-retry-tap.ts and the `inspectFacts`/`bindDevice` params the
  snapshot capture carried only to build the retry seam.
- `retryPositionals`/`scheduleInteractionOutcomeRetry` on
  `finalizeTouchInteraction`, and `pointPositionals`.

Tests:
- interaction-ios-tap-outcome.test.ts "a coordinate click the next
  snapshot finds unchanged is not re-sent": a click on 104,222, then a
  `snapshot` whose request carries the same runtime bindings the click
  used and returns the same tree, asserts one `press`. Mutation: restoring
  the base retry branch with its opt-in gate forced on (base
  `shouldRetryTouchOnNoChange` returning true) makes it see three presses
  (verified).
- interaction-touch-runtime.test.ts "press @ref taps the resolved point
  once and records the ref". Mutation: recording the resolved point
  positionals instead of the request's `@e1` in finalizeTouchInteraction
  fails it.
- interaction-outcome-policy.test.ts: the classifyInteractionSurfaceChange
  cases now pin areInteractionSurfaceSignaturesStable, which the
  stabilization loop and scroll movement still use. Mutation: dropping
  `stateMarkers(node)` from interactionSurfaceSemanticKey fails the
  checked-only flip case.
- snapshot-handler-freshness.test.ts: the annotation-survival case now
  rides the post-gesture deferred capture. Mutation: returning only
  `{ snapshot }` from resolvedPostGestureCapture fails it (verified).

* refactor(daemon): move the deferred outcome interface into post-gesture-stabilization.ts

With the pending-outcome retry gone, deferred-interaction-outcome.ts only
marks and resolves post-gesture stabilization (its R7 field) and Android
freshness. It moves to src/daemon/post-gesture-stabilization.ts; the
exported names stay, so callers change only their import path. No shim is
left at the old path.

The surviving cases of deferred-interaction-outcome.test.ts move unchanged
into post-gesture-stabilization.test.ts, which already tested this
module's stabilization loop. The capture-kit and fixture comments name the
new owner.

No new tests: the move carries its tests unchanged.

* docs: drop the no-change retry from the deferred outcome docs

CONTEXT.md's "Deferred interaction outcome" no longer lists outcome
retry, docs/agents/selector-capture.md states that the daemon never
re-sends an interaction to settle its outcome, and ADR 0004/0023 name
post-gesture-stabilization.ts where they named the deleted module.

* refactor(daemon): drop the unused logPath from the deferred outcome capture

`resolveDeferredInteractionOutcome` read `logPath` only to build the
runner context of a re-fired tap. With the retry gone nothing reads it, so
the parameter leaves the capture interface and its one production caller
(`captureSnapshot`).

No new tests: removing an unread parameter changes no behavior; the
typecheck is the gate that a caller still passing it fails.

* chore(gates): retire the pendingInteractionOutcome R7 owner entry

SessionState no longer has pendingInteractionOutcome, so its row leaves
SESSION_STATE_FIELD_OWNERS, and the postGestureStabilization owner moves
to src/daemon/post-gesture-stabilization.ts. R10 measures the merge-base,
so the shrink (16 -> 15 writer-owned fields) banks without a baseline edit.

* docs(agents): scope the no-resend rule to the daemon's deferred-outcome path

* docs(agents): keep the deferred-outcome rule to one line inside the docs budget

* refactor(move): keep the deferred outcome interface in deferred-interaction-outcome.ts

Reverts 401869a47b. The module still exports markDeferredInteractionOutcome,
resolveDeferredInteractionOutcome and DeferredInteractionOutcomeMark, owns
Android freshness as well as post-gesture stabilization, and CONTEXT.md
keeps "Deferred interaction outcome" as its glossary term. The surviving
cases move back into deferred-interaction-outcome.test.ts, ADR 0004/0023,
the capture-kit comment and the R7 owner entry name the original path.

git diff -M90% --stat:
 docs/adr/0004-ios-snapshot-backend-strategy.md     |   2 +-
 docs/adr/0023-end-state-hop-trace.md               |   4 +-
 packages/capture-kit/src/post-gesture-stability.ts |   2 +-
 scripts/layering/session-state.ts                  |   2 +-
 .../__tests__/deferred-interaction-outcome.test.ts | 197 +++++++++++++++++++++
 src/daemon/__tests__/is-runtime.test.ts            |   2 +-
 .../__tests__/post-gesture-no-effect-claim.test.ts |   4 +-
 .../post-gesture-stabilization-fixtures.ts         |   4 +-
 .../__tests__/post-gesture-stabilization.test.ts   | 192 +-------------------
 .../__tests__/snapshot-runtime-disclosure.test.ts  |   2 +-
 ...lization.ts => deferred-interaction-outcome.ts} |   0
 src/daemon/direct-ios-selector.ts                  |   2 +-
 src/daemon/interaction-outcome-policy.ts           |   2 +-
 .../find-target-activation-disclosure.test.ts      |   2 +-
 .../interaction-capture-disclosure.test.ts         |   2 +-
 .../interaction/internal/interaction-runtime.ts    |   2 +-
 src/daemon/interaction/internal/types.ts           |   2 +-
 src/daemon/request-generic-dispatch.ts             |   2 +-
 src/daemon/scroll-movement.ts                      |   2 +-
 src/daemon/selector-capture-runtime.ts             |   2 +-
 .../internal/session-open-execution.ts             |   2 +-
 src/daemon/snapshot-capture.ts                     |   2 +-
 22 files changed, 223 insertions(+), 210 deletions(-)

* refactor(move): name the surface comparators interaction-surface-signature.ts

With the outcome retry gone, interaction-outcome-policy.ts only builds and
compares interaction surface signatures. It moves to
interaction-surface-signature.ts with its test. stripInternalInteractionFlags
moves to deferred-interaction-outcome.ts, which already reads
flags.postGestureStabilization, and its test moves with it. Neither path
has a fallow baseline entry, so no baseline changes.

git diff -M90% --stat:
 docs/adr/0023-end-state-hop-trace.md                        |  2 +-
 packages/capture-kit/src/post-gesture-stability.ts          |  2 +-
 packages/capture-kit/src/snapshot-chrome.ts                 |  2 +-
 src/daemon/__tests__/deferred-interaction-outcome.test.ts   | 11 +++++++++++
 .../__tests__/interaction-surface-baseline-evidence.test.ts |  2 +-
 ...policy.test.ts => interaction-surface-signature.test.ts} | 13 +------------
 src/daemon/__tests__/post-gesture-no-effect-claim.test.ts   |  2 +-
 .../__tests__/post-gesture-stabilization-verdict.test.ts    |  2 +-
 src/daemon/__tests__/post-gesture-stabilization.test.ts     |  2 +-
 src/daemon/deferred-interaction-outcome.ts                  | 12 ++++++++++--
 ...n-outcome-policy.ts => interaction-surface-signature.ts} |  9 ---------
 src/daemon/interaction/internal/find.ts                     |  2 +-
 src/daemon/interaction/internal/interaction-common.ts       |  2 +-
 src/daemon/scroll-movement.ts                               |  4 ++--
 src/daemon/session-state.ts                                 |  2 +-
 15 files changed, 34 insertions(+), 35 deletions(-)

* docs: state the current capture ordering without the retired retry

Drop the no-resend sentence from the deferred-outcome mark comment, the
history line from docs/agents/selector-capture.md, and the retry framing
from the interaction-touch-runtime test header.

* test(daemon): drop tests that pass with or without the retry

Neither coordinate-tap test could set the retired opt-in, so one press
held on the base branch too. Delete the unchanged-click test; keep the
corroborated coordinate tap only for its corroboration assertions.

The one captureSnapshot case left in snapshot-handler-capture-retry.test.ts
moves to snapshot-capture.test.ts, which mirrors snapshot-capture.ts,
without the Apple runner and iOS hint mocks only the retry cases used.

* docs(adr): drop the automatic no-change retry from ADR 0014

No automatic no-change retry exists. Maestro retryTapIfNoChange sends an
ordinary press that crosses the leaf seam like any other action.

* test: drop no-op retry canary, fake timers in stabilization test, widen CommandFlags waiver

* fix(apple-runner): restart without resending a command written then lost (#3082)

* fix(apple-runner): name the lost reply when status cannot place a lost command

A command whose reply was lost and whose status probe answers notAccepted
(a restarted runner's empty journal) or an unnamed state now fails with
details.reason runner_reply_lost beside dispatched unknown, so a caller
keys on the typed reason instead of the message before it observes the
screen.

Tests (fake runner server through runAppleRunnerCommand and the recovery
module):
- runner-recovery-wiring: press and fill (runner tap and type) accept the
  command, drop the connection, status answers notAccepted: one dispatched
  command, error dispatched unknown with reason runner_reply_lost, the
  session is invalidated, and the next command gets a runner. Mutation:
  delete `reason: RUNNER_REPLY_LOST_REASON` from the unknown-state error ->
  both rows fail.
- runner-recovery-wiring: a snapshot whose reply is dropped is resent and
  succeeds. Mutation: declare snapshot with DEFAULT_TRAITS (not read-only)
  -> fails (the lost read goes to status recovery and is not resent).
- runner-command-recovery: notAccepted yields reason runner_reply_lost and
  dispatched unknown. Mutation: same deletion as above -> fails.

* fix(apple-runner): restart without resending a command written then lost

A connect-shaped failure keeps its restart, but the command is sent again
on the restarted runner only with proof the first send could not have run
it: the attempt provably wrote nothing (a connect refused before the write,
dispatched no), the readiness preflight gave up before the write, or the
command is read-only by its runner trait. A mutating command whose first
attempt may have written it now fails with reason runner_reply_lost and
dispatched unknown after the restart; a restarted runner's journal is empty,
so no status answer can prove the first send did not run. A read-only
command's restart failure discloses no: a read has no side effect to repeat.

dispatch-disclosure.json: ios-runner.transport.written-then-lost keeps
unknown with a trigger that says the command is not resent;
ios-runner.transport.read-only-written-then-lost is new and discloses no.

Tests (runner-lifecycle-dispatch-disclosure, real connect loop over a
stubbed fetch and simctl curl exit 28, restart succeeds):
- written-then-lost: tap restarts the runner, is not sent on the restarted
  runner, fails unknown with reason runner_reply_lost. Mutation: make
  canResendAfterRestart return true -> fails (the tap is replayed and the
  request succeeds).
- read-only-written-then-lost: snapshot is sent again on the restarted
  runner and its failure discloses no. Mutations: drop the read-only term
  from canResendAfterRestart -> fails (no resend); drop the read-only
  branch of discloseRestartDispatch -> fails (unknown).
runner-command-retry: four mutating-restart fixtures now carry the proof a
real refused connect stamps (dispatched no); without it they would test the
removed replay.

* test(apple): share the unwritten connect refusal fixture to keep the retry test file within its size ratchet

* fix(apple-runner): report the restart transport fact for a read, not no

The restart path disclosed `no` for a read-only command whose first send
may have run, while the status-recovery path discloses `unknown` for the
same read (read_only_completed_without_retained_response, and the
notAccepted/unknown-state fallthrough). Both runner paths now report the
transport fact: a first attempt that may have written the command is
`unknown`, read or mutation. The read-only `no` has one owner, the daemon
router seam (request-dispatch-disclosure.ts), which keys it on the
registry's recordingEffect `observes-app` for every producer
(daemon.read-only-command). A second copy keyed on the runner's own
read-only trait reconstructed that rule from a different source of truth,
and the two runner paths disagreed. The recycle-budget refusal on the
restart path also stops deriving `no` from the read-only trait and uses
only whether the first attempt was provably unwritten.

The runner still resends a read on the restarted runner
(canResendAfterRestart); only the disclosure changes.

dispatch-disclosure.json:
- ios-runner.transport.read-only-written-then-lost: now `unknown`; the
  trigger names the router rule that turns it into `no` for the caller.
- ios-runner.status.read-only-completed-without-retained-reply: new,
  `unknown`, driven through the real connect loop and recovery module.
  It pins the status-recovery read to the same value as the restart read.

Mutations:
- put back `if (isReadOnlyRunnerCommand(evidence.command)) return
  discloseDispatch(error, 'no')` in discloseRestartDispatch ->
  read-only-written-then-lost fails.
- set `dispatched: 'no'` on read_only_completed_without_retained_response
  -> status.read-only-completed-without-retained-reply fails.

* test(apple-runner): drive a real runner death through the fake runner

runner-recovery-wiring: the notAccepted mutation test no longer scripts a
second `ok` for the mutation. Nothing consumed it, and it suggested that
a resend was expected. Its comment now says what the test models: status
answers notAccepted, the way a runner whose journal did not survive a
restart would. The path is status recovery, not the daemon's restart.

The fake runner server gains an `exit` reply: it hangs up and stops
listening, like a runner process that died mid-command. Before this, a
test could only hang up on one request while the same server stayed up.
Two new tests use a second server as the restarted runner:

- press/fill (runner tap/type): the runner dies on the command. The
  mutation is sent once, is not sent to the restarted runner, fails with
  dispatched unknown, and the dead session is invalidated. A readText on
  the next request is served by the restarted runner. Mutation: declare
  tap and type with READ_ONLY_TRAITS (the resend trait) -> both rows fail.
- get (runner readText): the runner dies mid-read. The connect loop gives
  up with a possibly-written POST (simctl curl exit 28), the daemon
  restarts the runner (runner_connect_failed_before_command_send), and the
  read is resent once and succeeds. Mutation: reduce canResendAfterRestart
  to `evidence.firstAttemptUnwritten` -> fails (no resend).

With the real transport, a mutating command never reaches
restartSessionAndRunCommand with a possibly-written first attempt.
Mutations are sent by sendRunnerCommandOnce, and only the connect loop
(waitForRunner, used for reads and the readiness preflight) stamps
runner_connect_refused. So the mutating arm of canResendAfterRestart is
covered only by the mocked lifecycle driver
(ios-runner.transport.written-then-lost). A dying runner sends a mutation
to status recovery, and a failed status probe rethrows the transport
error without a runner_reply_lost reason.

* fix(apple-runner): name the lost reply when the status probe fails

A runner that dies mid-mutation drops the reply, and the status probe that
follows fails. The error that reaches the caller was then a bare transport
error (`fetch failed`, dispatched unknown), with no typed reason. #3074
requires a reason that names the lost reply, so this path now builds an
error with reason runner_reply_lost (the name the restart arm already uses)
and recovery status_probe_failed. The error keeps the transport message
and details, and the transport error as cause. The session is still
invalidated, and the command is still not resent.

ADR 0011 gap list: the row ios-runner.transport.written-then-lost proves
the rule behind canResendAfterRestart's mutating arm, not a production
route. Over the real transport the arm cannot be reached, because
mutations go through sendRunnerCommandOnce and only the connect loop
raises the restart trigger.

Tests:
- runner-command-recovery: a failing status probe now asserts reason
  runner_reply_lost, recovery status_probe_failed, dispatched unknown, and
  the transport error as cause.
- runner-recovery-wiring: press/fill whose runner dies mid-command (real
  fake-runner transport) now also asserts reason runner_reply_lost.
Mutation: drop `reason: RUNNER_REPLY_LOST_REASON` from
buildStatusProbeFailedError -> all three fail.

* test(apple-runner): hang up on every read attempt without counting the connect loop

The read-only lost-response driver scripted 20 hang-ups and a 400 ms
timeout. That count depended on the connect loop's retry cadence
(RUNNER_CONNECT_ATTEMPT_INTERVAL_MS): with a faster cadence the queue runs
out, and the fake answers the next attempt with `ok`. The fake runner
server now has a `hangUpAlways` reply that is never consumed, so the read
is refused until the loop gives up, however often it retries. The timeout
now only bounds how long the loop runs.

Mutation: set RUNNER_CONNECT_ATTEMPT_INTERVAL_MS to 10 or 50 (with a 1 ms
retry base delay) -> ios-runner.status.read-only-completed-without-retained-reply
still passes.

* test(apple-runner): pick the fake runner's next reply in one helper

The request handler of the fake runner server went over the fallow complexity threshold after hangUpAlways was added. The choice of the next scripted reply is now in takeScriptedResponse. Behavior is unchanged: the apple-runner suite passes (768 tests).

* refactor(apple-runner): gate the restart resend at the connect-failure call site

A mutation never reaches the connect loop, so the no-resend arm after a
restart had no production route. The call site now restarts only when the
first attempt provably wrote nothing or the command is read-only, and the
arm and its error builder are gone. The written-then-lost row is driven
through a fake runner that dies mid-command. A failed status probe keeps a
transport reason under transportReason.

* test(apple-runner): name runner_reply_lost in the lost-reply test titles

* test(apple-runner): the runner read-only set matches the registry's observes-app commands

* fix(apple-runner): decide the restart resend in the connect-failure classifier

shouldRestartRunnerBeforeCommandSend now takes the command and owns the
write-proof rule: a connect-loop failure restarts and resends only when no
attempt could have written the command, or the command is read-only. An
exit-28 POST that carries dispatched unknown no longer restarts a mutation,
whatever call site asks. The call-site guard in executeRunnerCommandAttempt
is gone.

* refactor(apple-runner): carry the first attempt's dispatch disclosure through the restart

restartSessionAndRunCommand takes firstAttemptDispatched: DispatchDisclosure
instead of a boolean that three ternaries decoded back into no or unknown.
resolveFirstAttemptDispatch computes it once beside
isRunnerCommandProvablyUnwritten, and the restart rule reads the same answer.
The readiness-preflight call site passes no.

* refactor(apple-runner): build every lost-reply failure in one place with one reason

Status recovery verdicts now carry either the runner's own answer or the
facts of a lost reply. applyRunnerTransportRecovery turns a lost reply into
one buildLostReplyError for a mutation, so runner_reply_lost is set on every
lost-reply verdict (completed without a retained reply, still in flight,
unknown or missing lifecycle, status probe failed, status unavailable) and
the transport's reason stays under transportReason. A read keeps its
transport error, which the read-only resend loop resends, so it never
carries the reason. The four hand-built copies are gone.

* fix(apple-runner): key a lost reply on its typed reason, never on transport text

* docs(adr): name the lost-reply dispatch-disclosure rows in ADR 0011

* test(apple-runner): pair keyboardReturn with the keyboard enter request that issues it

* test(apple-runner): pin the lost-reply branch in the connect-loop mutation test

* test(apple-runner): type the read-only parity rows as the registry request and drop the cast

* fix(apple-runner): keep the transport error's own hint on a mutation lost reply

* refactor(apple-runner): build the lost-reply message once from the command and cause

* docs(wire-compat): say an older peer ignores or rejects an unknown flag
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-for-human Valid work that needs human implementation, judgment, or maintainer merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant