fix(apple-runner): restart without resending a command written then lost - #3082
Conversation
Size Report
Startup median (7 runs, lower is better):
|
There was a problem hiding this comment.
1 issue found across 9 files
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="packages/platform-apple/src/runner/runner-lifecycle.ts">
<violation number="1" location="packages/platform-apple/src/runner/runner-lifecycle.ts:542">
P3: `discloseRestartDispatch` now reports `dispatched: "no"` for read-only commands when the restart's resend fails with an already-disclosed `"unknown"` verdict (the `ios-runner.transport.read-only-written-then-lost` fixture asserts this). The parallel status-recovery path in `runner-command-recovery.ts` still discloses the same user-visible case as `"unknown"` (`handleCompletedRunnerStatus`'s `read_only_completed_without_retained_response` branch, and the fallthrough after a `notAccepted` probe for `lifecycle_state_not_recoverable`), so an agent deciding whether to retry on `details.dispatched` sees different values for the same class of failure depending on which recovery flow fired. Decide one semantics for read-only lost replies and pin both paths (and the fixture) to it.</violation>
</file>
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
| ): AppError { | ||
| if (!restartedSession) return discloseDispatch(error, firstAttemptUnwritten ? 'no' : 'unknown'); | ||
| if (!firstAttemptUnwritten) return discloseDispatch(error, 'unknown'); | ||
| if (isReadOnlyRunnerCommand(evidence.command)) return discloseDispatch(error, 'no'); |
There was a problem hiding this comment.
P3: discloseRestartDispatch now reports dispatched: "no" for read-only commands when the restart's resend fails with an already-disclosed "unknown" verdict (the ios-runner.transport.read-only-written-then-lost fixture asserts this). The parallel status-recovery path in runner-command-recovery.ts still discloses the same user-visible case as "unknown" (handleCompletedRunnerStatus's read_only_completed_without_retained_response branch, and the fallthrough after a notAccepted probe for lifecycle_state_not_recoverable), so an agent deciding whether to retry on details.dispatched sees different values for the same class of failure depending on which recovery flow fired. Decide one semantics for read-only lost replies and pin both paths (and the fixture) to it.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/platform-apple/src/runner/runner-lifecycle.ts, line 542:
<comment>`discloseRestartDispatch` now reports `dispatched: "no"` for read-only commands when the restart's resend fails with an already-disclosed `"unknown"` verdict (the `ios-runner.transport.read-only-written-then-lost` fixture asserts this). The parallel status-recovery path in `runner-command-recovery.ts` still discloses the same user-visible case as `"unknown"` (`handleCompletedRunnerStatus`'s `read_only_completed_without_retained_response` branch, and the fallthrough after a `notAccepted` probe for `lifecycle_state_not_recoverable`), so an agent deciding whether to retry on `details.dispatched` sees different values for the same class of failure depending on which recovery flow fired. Decide one semantics for read-only lost replies and pin both paths (and the fixture) to it.</comment>
<file context>
@@ -500,21 +525,23 @@ function markRunnerRestartError(
): AppError {
- if (!restartedSession) return discloseDispatch(error, firstAttemptUnwritten ? 'no' : 'unknown');
- if (!firstAttemptUnwritten) return discloseDispatch(error, 'unknown');
+ if (isReadOnlyRunnerCommand(evidence.command)) return discloseDispatch(error, 'no');
+ if (!evidence.firstAttemptUnwritten) return discloseDispatch(error, 'unknown');
+ if (!restartedSession) return discloseDispatch(error, 'no');
</file context>
|
[claude-fable-5-1] responding on behalf of @thymikee Head is now 1. Read-only restart discloses
2. Unconsumed second 3. Premise says "restarted", but the path is status recovery (~L164): fixed in 4758d27, with one pushback.
Checks run: |
|
[claude-fable-5-1] responding on behalf of @thymikee Follow-up in 50fc6fc:
|
There was a problem hiding this comment.
All reported issues were addressed across 8 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Fix all with cubic | Re-trigger cubic
|
[claude-fable-5-1] responding on behalf of @thymikee Hang-up count tied to the connect-loop cadence (runner-dispatch-disclosure.test.ts ~L81): fixed in f247110, with a complexity follow-up in the next commit. The fake runner server has a new |
|
The change looks correct in code, but the one reachable behavior it adds has no live proof yet. I reviewed 1edcb78. CI is green: 14 checks, none failing. No conflicts. The reachable change is the typed reason Could the change be smaller? The connect loop in Not blocking, take or leave: the no-resend arm at runner-lifecycle.ts#L437 is only driven by a mocked transport row, because mutations never enter the connect loop; the I did not run the tests or the mutations. I judged "fails without the fix" from the merge-base code, and "next command gets a fresh runner" is only shown with a mocked Before merge, we need the live run on 1edcb78 described above, and a refreshed PR body with the head sha, a green |
|
[claude-fable-5-1] responding on behalf of @thymikee Head is now Simplification (6739a58): I took it.
Reason key: Retitles (f52f120): Live on f52f120: new iPhone 17 Pro sim, iOS 26.2, Release test app. I ran
Gate: |
thymikee
left a comment
There was a problem hiding this comment.
Code-quality review. The shipped behavior is right: a mutation whose runner dies mid-command is sent once and fails with runner_reply_lost, through the status probe. The structure has three problems:
- The new restart gate at the call site never changes behavior on any input that can reach it. The defect is in the classifier that labels an exit-28 POST as
runner_connect_refused. - The written-then-lost proof is a boolean that later code converts back into
'no' | 'unknown', instead of the existingDispatchDisclosuretype. runner_reply_lostis set on some lost-reply outcomes and not on others, by four hand-built copies of the same error.
Inline comments cover each point. No file crosses 1k lines.
| if ( | ||
| shouldRestartRunnerBeforeCommandSend(appErr) && | ||
| session && | ||
| (firstAttemptUnwritten || isReadOnlyRunnerCommand(command)) |
There was a problem hiding this comment.
This conjunct guards a state that cannot happen. Mutations go through sendRunnerCommandOnce and never enter the connect loop (waitForRunner / postCommandViaSimulator), and that loop is the only producer of runner_connect_refused. The PR description says the same. For a mutation, the only restartBeforeSend error that can arrive is a marked readiness-preflight failure, and isRunnerPreSendRefusal is already true for that. For a read, the disjunct is true. So (firstAttemptUnwritten || isReadOnlyRunnerCommand(command)) is true for every input that can reach this line. No test exercises the side where it refuses either: the written-then-lost row moved to the status-probe path, and every remaining mutation driver uses unwrittenConnectRefusal().
The real defect is in the code that owns the classification. restartBeforeSend is documented as "connect-shaped failure before the command was sent". But postCommandViaSimulator (runner-startup-transport.ts ~L517) labels a curl exit-28 POST that has dispatched: 'unknown' as runner_connect_refused. Fix it there. Either stop labelling a request that was posted and then timed out as a connect refusal, or make shouldRestartRunnerBeforeCommandSend(error, command) in runner-error-classification.ts own the write-proof rule. Then delete this call-site guard. AGENTS.md says to repair at the owning type and not to add guards that rebuild another source of truth.
| signal, | ||
| restartReason: 'runner_connect_failed_before_command_send', | ||
| firstAttemptUnwritten: isRunnerPreSendRefusal(appErr), | ||
| firstAttemptUnwritten, |
There was a problem hiding this comment.
firstAttemptUnwritten is a boolean derived from a typed fact that already exists: the error's details.dispatched, which is 'no' for curl exit 7 or ECONNREFUSED. Downstream code then turns it back into 'no' | 'unknown' with hand-written ternaries: once for the recycle-budget throw in restartSessionAndRunCommand, and twice in discloseRestartDispatch. A reader has to decode the boolean every time it is used.
Carry the first attempt's disclosure instead:
- Add
firstAttemptDispatched: DispatchDisclosure, computed once by one classifier, for exampleresolveFirstAttemptDispatch(error)next toisRunnerCommandProvablyUnwritten. - Replace the ternaries with
discloseDispatch(err, params.firstAttemptDispatched). - Have the readiness-preflight call site pass
'no'instead oftrue.
This models the proof as a value of the domain type and removes branches instead of adding them. Together with the L359 comment, the restart decision becomes one answer from one classifier.
| * `details.reason` of a failure whose command was written, lost its reply, and has no proof it did | ||
| * not run. The command is not resent; the caller observes the screen before acting again. | ||
| */ | ||
| export const RUNNER_REPLY_LOST_REASON = 'runner_reply_lost'; |
There was a problem hiding this comment.
runner_reply_lost is documented as: the command was written, its reply was lost, there is no proof it did not run, and it is not resent. But only some outcomes that match this carry it: status_probe_failed and the unknown/missing-lifecycle case.
completed_without_retained_response(handleCompletedRunnerStatus) andcommand_still_in_flight(runnerStatusInFlightError) are also written, reply lost,dispatched: 'unknown', and not resent. They do not set the reason.- In the other direction,
buildStatusProbeFailedErrorsets it on read-only commands too. That error keeps the code, message and connect details, soisRetryableRunnerErrorstill holds, and the read-only resend loop inrunAppleRunnerCommandresends the read up toTRANSPORT_RESEND_ATTEMPTStimes. A read that runs out of attempts therefore ends with a reason that says "not resent".
A consumer that keys on details.reason gets a signal that is incomplete and sometimes false. The reason belongs to the recovery verdict, not to each error builder. Set it in one place, for example in applyRunnerTransportRecovery for every non-recovered mutation verdict with dispatched: 'unknown'. Add one table-driven test that every lost-reply row in dispatch-disclosure.json carries it.
| } | ||
|
|
||
| /** The lost reply could not be placed because the status probe itself failed. */ | ||
| function buildStatusProbeFailedError( |
There was a problem hiding this comment.
This is the fourth hand-built copy of the same lost-reply AppError: command, commandId, recovery, hint, logPath, the transportError message, and transportError as the cause. The other three are the inline unknown-lifecycle literal (~L289–305), handleCompletedRunnerStatus and runnerStatusInFlightError.
This copy also has a different shape. The other three drop the transport details. This one spreads ...transportError.details, overwrites reason, and moves the original reason into a new transportReason key that only this path has. To find the transport's reason, a consumer now has to know which path built the error.
Extract one buildLostReplyError(command, transportError, options, { recovery, lifecycleState?, message?, hint }). Make it own the detail shape, the reply-lost reason (see the L54 comment), and one rule for keeping the transport's own reason. Then rewrite this function and the inline literal as calls to it.
|
This PR is ready. The earlier findings at 1edcb78 are fixed in f52f120, and I found nothing new to block on. Not blocking, and you can take or leave it: the guard The code review is done. I read the quoted request-log lines from your live run, but I saw no raw artifact, so that evidence is author-reported. I did not run the runner vitest suite or your gate-removal mutation. From reading the code, the refuse side looks untested, which does not match the "8 tests fail" claim. CI is still pending. Smoke Tests is in progress and has not failed, and the other 13 checks pass. If its iOS lane runs runner mutations, it exercises the changed route, which is the restart gating in |
b2027c2 to
d7b442d
Compare
f52f120 to
9393095
Compare
|
[claude-fable-5-1] responding on behalf of @thymikee Replayed onto #3071's Added |
There was a problem hiding this comment.
All reported issues were addressed across 1 file (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Fix all with cubic | Re-trigger cubic
…t command A command whose reply was lost and whose status probe answers notAccepted (a restarted runner's empty journal) or an unnamed state now fails with details.reason runner_reply_lost beside dispatched unknown, so a caller keys on the typed reason instead of the message before it observes the screen. Tests (fake runner server through runAppleRunnerCommand and the recovery module): - runner-recovery-wiring: press and fill (runner tap and type) accept the command, drop the connection, status answers notAccepted: one dispatched command, error dispatched unknown with reason runner_reply_lost, the session is invalidated, and the next command gets a runner. Mutation: delete `reason: RUNNER_REPLY_LOST_REASON` from the unknown-state error -> both rows fail. - runner-recovery-wiring: a snapshot whose reply is dropped is resent and succeeds. Mutation: declare snapshot with DEFAULT_TRAITS (not read-only) -> fails (the lost read goes to status recovery and is not resent). - runner-command-recovery: notAccepted yields reason runner_reply_lost and dispatched unknown. Mutation: same deletion as above -> fails.
A connect-shaped failure keeps its restart, but the command is sent again on the restarted runner only with proof the first send could not have run it: the attempt provably wrote nothing (a connect refused before the write, dispatched no), the readiness preflight gave up before the write, or the command is read-only by its runner trait. A mutating command whose first attempt may have written it now fails with reason runner_reply_lost and dispatched unknown after the restart; a restarted runner's journal is empty, so no status answer can prove the first send did not run. A read-only command's restart failure discloses no: a read has no side effect to repeat. dispatch-disclosure.json: ios-runner.transport.written-then-lost keeps unknown with a trigger that says the command is not resent; ios-runner.transport.read-only-written-then-lost is new and discloses no. Tests (runner-lifecycle-dispatch-disclosure, real connect loop over a stubbed fetch and simctl curl exit 28, restart succeeds): - written-then-lost: tap restarts the runner, is not sent on the restarted runner, fails unknown with reason runner_reply_lost. Mutation: make canResendAfterRestart return true -> fails (the tap is replayed and the request succeeds). - read-only-written-then-lost: snapshot is sent again on the restarted runner and its failure discloses no. Mutations: drop the read-only term from canResendAfterRestart -> fails (no resend); drop the read-only branch of discloseRestartDispatch -> fails (unknown). runner-command-retry: four mutating-restart fixtures now carry the proof a real refused connect stamps (dispatched no); without it they would test the removed replay.
…retry test file within its size ratchet
The restart path disclosed `no` for a read-only command whose first send may have run, while the status-recovery path discloses `unknown` for the same read (read_only_completed_without_retained_response, and the notAccepted/unknown-state fallthrough). Both runner paths now report the transport fact: a first attempt that may have written the command is `unknown`, read or mutation. The read-only `no` has one owner, the daemon router seam (request-dispatch-disclosure.ts), which keys it on the registry's recordingEffect `observes-app` for every producer (daemon.read-only-command). A second copy keyed on the runner's own read-only trait reconstructed that rule from a different source of truth, and the two runner paths disagreed. The recycle-budget refusal on the restart path also stops deriving `no` from the read-only trait and uses only whether the first attempt was provably unwritten. The runner still resends a read on the restarted runner (canResendAfterRestart); only the disclosure changes. dispatch-disclosure.json: - ios-runner.transport.read-only-written-then-lost: now `unknown`; the trigger names the router rule that turns it into `no` for the caller. - ios-runner.status.read-only-completed-without-retained-reply: new, `unknown`, driven through the real connect loop and recovery module. It pins the status-recovery read to the same value as the restart read. Mutations: - put back `if (isReadOnlyRunnerCommand(evidence.command)) return discloseDispatch(error, 'no')` in discloseRestartDispatch -> read-only-written-then-lost fails. - set `dispatched: 'no'` on read_only_completed_without_retained_response -> status.read-only-completed-without-retained-reply fails.
runner-recovery-wiring: the notAccepted mutation test no longer scripts a second `ok` for the mutation. Nothing consumed it, and it suggested that a resend was expected. Its comment now says what the test models: status answers notAccepted, the way a runner whose journal did not survive a restart would. The path is status recovery, not the daemon's restart. The fake runner server gains an `exit` reply: it hangs up and stops listening, like a runner process that died mid-command. Before this, a test could only hang up on one request while the same server stayed up. Two new tests use a second server as the restarted runner: - press/fill (runner tap/type): the runner dies on the command. The mutation is sent once, is not sent to the restarted runner, fails with dispatched unknown, and the dead session is invalidated. A readText on the next request is served by the restarted runner. Mutation: declare tap and type with READ_ONLY_TRAITS (the resend trait) -> both rows fail. - get (runner readText): the runner dies mid-read. The connect loop gives up with a possibly-written POST (simctl curl exit 28), the daemon restarts the runner (runner_connect_failed_before_command_send), and the read is resent once and succeeds. Mutation: reduce canResendAfterRestart to `evidence.firstAttemptUnwritten` -> fails (no resend). With the real transport, a mutating command never reaches restartSessionAndRunCommand with a possibly-written first attempt. Mutations are sent by sendRunnerCommandOnce, and only the connect loop (waitForRunner, used for reads and the readiness preflight) stamps runner_connect_refused. So the mutating arm of canResendAfterRestart is covered only by the mocked lifecycle driver (ios-runner.transport.written-then-lost). A dying runner sends a mutation to status recovery, and a failed status probe rethrows the transport error without a runner_reply_lost reason.
A runner that dies mid-mutation drops the reply, and the status probe that follows fails. The error that reaches the caller was then a bare transport error (`fetch failed`, dispatched unknown), with no typed reason. #3074 requires a reason that names the lost reply, so this path now builds an error with reason runner_reply_lost (the name the restart arm already uses) and recovery status_probe_failed. The error keeps the transport message and details, and the transport error as cause. The session is still invalidated, and the command is still not resent. ADR 0011 gap list: the row ios-runner.transport.written-then-lost proves the rule behind canResendAfterRestart's mutating arm, not a production route. Over the real transport the arm cannot be reached, because mutations go through sendRunnerCommandOnce and only the connect loop raises the restart trigger. Tests: - runner-command-recovery: a failing status probe now asserts reason runner_reply_lost, recovery status_probe_failed, dispatched unknown, and the transport error as cause. - runner-recovery-wiring: press/fill whose runner dies mid-command (real fake-runner transport) now also asserts reason runner_reply_lost. Mutation: drop `reason: RUNNER_REPLY_LOST_REASON` from buildStatusProbeFailedError -> all three fail.
…e connect loop The read-only lost-response driver scripted 20 hang-ups and a 400 ms timeout. That count depended on the connect loop's retry cadence (RUNNER_CONNECT_ATTEMPT_INTERVAL_MS): with a faster cadence the queue runs out, and the fake answers the next attempt with `ok`. The fake runner server now has a `hangUpAlways` reply that is never consumed, so the read is refused until the loop gives up, however often it retries. The timeout now only bounds how long the loop runs. Mutation: set RUNNER_CONNECT_ATTEMPT_INTERVAL_MS to 10 or 50 (with a 1 ms retry base delay) -> ios-runner.status.read-only-completed-without-retained-reply still passes.
The request handler of the fake runner server went over the fallow complexity threshold after hangUpAlways was added. The choice of the next scripted reply is now in takeScriptedResponse. Behavior is unchanged: the apple-runner suite passes (768 tests).
…e call site A mutation never reaches the connect loop, so the no-resend arm after a restart had no production route. The call site now restarts only when the first attempt provably wrote nothing or the command is read-only, and the arm and its error builder are gone. The written-then-lost row is driven through a fake runner that dies mid-command. A failed status probe keeps a transport reason under transportReason.
…bserves-app commands
…lassifier shouldRestartRunnerBeforeCommandSend now takes the command and owns the write-proof rule: a connect-loop failure restarts and resends only when no attempt could have written the command, or the command is read-only. An exit-28 POST that carries dispatched unknown no longer restarts a mutation, whatever call site asks. The call-site guard in executeRunnerCommandAttempt is gone.
… through the restart restartSessionAndRunCommand takes firstAttemptDispatched: DispatchDisclosure instead of a boolean that three ternaries decoded back into no or unknown. resolveFirstAttemptDispatch computes it once beside isRunnerCommandProvablyUnwritten, and the restart rule reads the same answer. The readiness-preflight call site passes no.
…ith one reason Status recovery verdicts now carry either the runner's own answer or the facts of a lost reply. applyRunnerTransportRecovery turns a lost reply into one buildLostReplyError for a mutation, so runner_reply_lost is set on every lost-reply verdict (completed without a retained reply, still in flight, unknown or missing lifecycle, status probe failed, status unavailable) and the transport's reason stays under transportReason. A read keeps its transport error, which the read-only resend loop resends, so it never carries the reason. The four hand-built copies are gone.
9393095 to
4e88908
Compare
|
[claude-fable-5-1] responding on behalf of @thymikee Head is now 1. The call-site guard never changed behavior (runner-lifecycle.ts L359): fixed in 2ffd18a.
2. A boolean instead of the domain type (runner-lifecycle.ts L369): fixed in 66ef675.
3. One lost-reply error, one reason (runner-command-recovery.ts L54, L489): fixed in 4e88908.
What changed:
Tests:
Mutations:
ADR 0011 now names the Gate on |
There was a problem hiding this comment.
All reported issues were addressed across 8 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Fix all with cubic | Re-trigger cubic
…st that issues it
|
[claude-fable-5-1] responding on behalf of @thymikee Addressed cubic's latest four findings:
Checks: typecheck, lint, format:check, layering and fallow audit pass, and the apple-runner project passes 767/767. A 5 s timeout failed under load and passed when run alone. In unit-core, 2 timeouts fail when run alone, and neither relates to this change: |
There was a problem hiding this comment.
1 issue found across 10 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="packages/platform-apple/src/runner/runner-command-recovery.ts">
<violation number="1" location="packages/platform-apple/src/runner/runner-command-recovery.ts:506">
P3: `lostReplyWithoutStatusMessage` rebuilds the same sentence skeleton as the inline message in `handleRunnerCommandStatusRecovery` (`lifecycle_state_not_recoverable`): both are `Runner command "…" lost its transport response and … so agent-device invalidated the runner session instead of replaying the command.`. Three lost-reply messages now share the identical structure, so the shared "invalidated instead of replaying" clause can drift between the two paths. Feed the lifecycle branch through the same builder, e.g. `lostReplyWithoutStatusMessage(command.command, lifecycleState ? `lifecycle status was "${lifecycleState}"` : 'lifecycle status was missing')`.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Fix all with cubic | Re-trigger cubic
| return details; | ||
| } | ||
|
|
||
| function lostReplyWithoutStatusMessage(command: string, statusOutcome: string): string { |
There was a problem hiding this comment.
P3: lostReplyWithoutStatusMessage rebuilds the same sentence skeleton as the inline message in handleRunnerCommandStatusRecovery (lifecycle_state_not_recoverable): both are Runner command "…" lost its transport response and … so agent-device invalidated the runner session instead of replaying the command.. Three lost-reply messages now share the identical structure, so the shared "invalidated instead of replaying" clause can drift between the two paths. Feed the lifecycle branch through the same builder, e.g. lostReplyWithoutStatusMessage(command.command, lifecycleState ? lifecycle status was "${lifecycleState}" : 'lifecycle status was missing').
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/platform-apple/src/runner/runner-command-recovery.ts, line 506:
<comment>`lostReplyWithoutStatusMessage` rebuilds the same sentence skeleton as the inline message in `handleRunnerCommandStatusRecovery` (`lifecycle_state_not_recoverable`): both are `Runner command "…" lost its transport response and … so agent-device invalidated the runner session instead of replaying the command.`. Three lost-reply messages now share the identical structure, so the shared "invalidated instead of replaying" clause can drift between the two paths. Feed the lifecycle branch through the same builder, e.g. `lostReplyWithoutStatusMessage(command.command, lifecycleState ? `lifecycle status was "${lifecycleState}"` : 'lifecycle status was missing')`.</comment>
<file context>
@@ -504,6 +503,10 @@ function readReadinessPreflightRecoveryDetails(
return details;
}
+function lostReplyWithoutStatusMessage(command: string, statusOutcome: string): string {
+ return `Runner command "${command}" lost its transport response and ${statusOutcome}, so agent-device invalidated the runner session instead of replaying the command.`;
+}
</file context>
|
Reviewed 25ce56e: no blocking issues. The new commits keep the restart-resend decision in one place, the connect-failure classifier. A lost reply now keys on its typed reason, not on transport text. A mutation whose request may have been written is never resent, and the first attempt's Not blocking, take or leave: the parity test in CI is still running (Smoke Tests, Repo Guards, Coverage), and iOS Smoke Tests run the runner route this changes. There are no conflicts. I read the code and did not run the tests; the lost-reply path is covered by the fake runner server only, with no live device run. |
|
|
[claude-fable-5-1] responding on behalf of @thymikee This PR merged into |
…ost (#3082) * fix(apple-runner): name the lost reply when status cannot place a lost command A command whose reply was lost and whose status probe answers notAccepted (a restarted runner's empty journal) or an unnamed state now fails with details.reason runner_reply_lost beside dispatched unknown, so a caller keys on the typed reason instead of the message before it observes the screen. Tests (fake runner server through runAppleRunnerCommand and the recovery module): - runner-recovery-wiring: press and fill (runner tap and type) accept the command, drop the connection, status answers notAccepted: one dispatched command, error dispatched unknown with reason runner_reply_lost, the session is invalidated, and the next command gets a runner. Mutation: delete `reason: RUNNER_REPLY_LOST_REASON` from the unknown-state error -> both rows fail. - runner-recovery-wiring: a snapshot whose reply is dropped is resent and succeeds. Mutation: declare snapshot with DEFAULT_TRAITS (not read-only) -> fails (the lost read goes to status recovery and is not resent). - runner-command-recovery: notAccepted yields reason runner_reply_lost and dispatched unknown. Mutation: same deletion as above -> fails. * fix(apple-runner): restart without resending a command written then lost A connect-shaped failure keeps its restart, but the command is sent again on the restarted runner only with proof the first send could not have run it: the attempt provably wrote nothing (a connect refused before the write, dispatched no), the readiness preflight gave up before the write, or the command is read-only by its runner trait. A mutating command whose first attempt may have written it now fails with reason runner_reply_lost and dispatched unknown after the restart; a restarted runner's journal is empty, so no status answer can prove the first send did not run. A read-only command's restart failure discloses no: a read has no side effect to repeat. dispatch-disclosure.json: ios-runner.transport.written-then-lost keeps unknown with a trigger that says the command is not resent; ios-runner.transport.read-only-written-then-lost is new and discloses no. Tests (runner-lifecycle-dispatch-disclosure, real connect loop over a stubbed fetch and simctl curl exit 28, restart succeeds): - written-then-lost: tap restarts the runner, is not sent on the restarted runner, fails unknown with reason runner_reply_lost. Mutation: make canResendAfterRestart return true -> fails (the tap is replayed and the request succeeds). - read-only-written-then-lost: snapshot is sent again on the restarted runner and its failure discloses no. Mutations: drop the read-only term from canResendAfterRestart -> fails (no resend); drop the read-only branch of discloseRestartDispatch -> fails (unknown). runner-command-retry: four mutating-restart fixtures now carry the proof a real refused connect stamps (dispatched no); without it they would test the removed replay. * test(apple): share the unwritten connect refusal fixture to keep the retry test file within its size ratchet * fix(apple-runner): report the restart transport fact for a read, not no The restart path disclosed `no` for a read-only command whose first send may have run, while the status-recovery path discloses `unknown` for the same read (read_only_completed_without_retained_response, and the notAccepted/unknown-state fallthrough). Both runner paths now report the transport fact: a first attempt that may have written the command is `unknown`, read or mutation. The read-only `no` has one owner, the daemon router seam (request-dispatch-disclosure.ts), which keys it on the registry's recordingEffect `observes-app` for every producer (daemon.read-only-command). A second copy keyed on the runner's own read-only trait reconstructed that rule from a different source of truth, and the two runner paths disagreed. The recycle-budget refusal on the restart path also stops deriving `no` from the read-only trait and uses only whether the first attempt was provably unwritten. The runner still resends a read on the restarted runner (canResendAfterRestart); only the disclosure changes. dispatch-disclosure.json: - ios-runner.transport.read-only-written-then-lost: now `unknown`; the trigger names the router rule that turns it into `no` for the caller. - ios-runner.status.read-only-completed-without-retained-reply: new, `unknown`, driven through the real connect loop and recovery module. It pins the status-recovery read to the same value as the restart read. Mutations: - put back `if (isReadOnlyRunnerCommand(evidence.command)) return discloseDispatch(error, 'no')` in discloseRestartDispatch -> read-only-written-then-lost fails. - set `dispatched: 'no'` on read_only_completed_without_retained_response -> status.read-only-completed-without-retained-reply fails. * test(apple-runner): drive a real runner death through the fake runner runner-recovery-wiring: the notAccepted mutation test no longer scripts a second `ok` for the mutation. Nothing consumed it, and it suggested that a resend was expected. Its comment now says what the test models: status answers notAccepted, the way a runner whose journal did not survive a restart would. The path is status recovery, not the daemon's restart. The fake runner server gains an `exit` reply: it hangs up and stops listening, like a runner process that died mid-command. Before this, a test could only hang up on one request while the same server stayed up. Two new tests use a second server as the restarted runner: - press/fill (runner tap/type): the runner dies on the command. The mutation is sent once, is not sent to the restarted runner, fails with dispatched unknown, and the dead session is invalidated. A readText on the next request is served by the restarted runner. Mutation: declare tap and type with READ_ONLY_TRAITS (the resend trait) -> both rows fail. - get (runner readText): the runner dies mid-read. The connect loop gives up with a possibly-written POST (simctl curl exit 28), the daemon restarts the runner (runner_connect_failed_before_command_send), and the read is resent once and succeeds. Mutation: reduce canResendAfterRestart to `evidence.firstAttemptUnwritten` -> fails (no resend). With the real transport, a mutating command never reaches restartSessionAndRunCommand with a possibly-written first attempt. Mutations are sent by sendRunnerCommandOnce, and only the connect loop (waitForRunner, used for reads and the readiness preflight) stamps runner_connect_refused. So the mutating arm of canResendAfterRestart is covered only by the mocked lifecycle driver (ios-runner.transport.written-then-lost). A dying runner sends a mutation to status recovery, and a failed status probe rethrows the transport error without a runner_reply_lost reason. * fix(apple-runner): name the lost reply when the status probe fails A runner that dies mid-mutation drops the reply, and the status probe that follows fails. The error that reaches the caller was then a bare transport error (`fetch failed`, dispatched unknown), with no typed reason. #3074 requires a reason that names the lost reply, so this path now builds an error with reason runner_reply_lost (the name the restart arm already uses) and recovery status_probe_failed. The error keeps the transport message and details, and the transport error as cause. The session is still invalidated, and the command is still not resent. ADR 0011 gap list: the row ios-runner.transport.written-then-lost proves the rule behind canResendAfterRestart's mutating arm, not a production route. Over the real transport the arm cannot be reached, because mutations go through sendRunnerCommandOnce and only the connect loop raises the restart trigger. Tests: - runner-command-recovery: a failing status probe now asserts reason runner_reply_lost, recovery status_probe_failed, dispatched unknown, and the transport error as cause. - runner-recovery-wiring: press/fill whose runner dies mid-command (real fake-runner transport) now also asserts reason runner_reply_lost. Mutation: drop `reason: RUNNER_REPLY_LOST_REASON` from buildStatusProbeFailedError -> all three fail. * test(apple-runner): hang up on every read attempt without counting the connect loop The read-only lost-response driver scripted 20 hang-ups and a 400 ms timeout. That count depended on the connect loop's retry cadence (RUNNER_CONNECT_ATTEMPT_INTERVAL_MS): with a faster cadence the queue runs out, and the fake answers the next attempt with `ok`. The fake runner server now has a `hangUpAlways` reply that is never consumed, so the read is refused until the loop gives up, however often it retries. The timeout now only bounds how long the loop runs. Mutation: set RUNNER_CONNECT_ATTEMPT_INTERVAL_MS to 10 or 50 (with a 1 ms retry base delay) -> ios-runner.status.read-only-completed-without-retained-reply still passes. * test(apple-runner): pick the fake runner's next reply in one helper The request handler of the fake runner server went over the fallow complexity threshold after hangUpAlways was added. The choice of the next scripted reply is now in takeScriptedResponse. Behavior is unchanged: the apple-runner suite passes (768 tests). * refactor(apple-runner): gate the restart resend at the connect-failure call site A mutation never reaches the connect loop, so the no-resend arm after a restart had no production route. The call site now restarts only when the first attempt provably wrote nothing or the command is read-only, and the arm and its error builder are gone. The written-then-lost row is driven through a fake runner that dies mid-command. A failed status probe keeps a transport reason under transportReason. * test(apple-runner): name runner_reply_lost in the lost-reply test titles * test(apple-runner): the runner read-only set matches the registry's observes-app commands * fix(apple-runner): decide the restart resend in the connect-failure classifier shouldRestartRunnerBeforeCommandSend now takes the command and owns the write-proof rule: a connect-loop failure restarts and resends only when no attempt could have written the command, or the command is read-only. An exit-28 POST that carries dispatched unknown no longer restarts a mutation, whatever call site asks. The call-site guard in executeRunnerCommandAttempt is gone. * refactor(apple-runner): carry the first attempt's dispatch disclosure through the restart restartSessionAndRunCommand takes firstAttemptDispatched: DispatchDisclosure instead of a boolean that three ternaries decoded back into no or unknown. resolveFirstAttemptDispatch computes it once beside isRunnerCommandProvablyUnwritten, and the restart rule reads the same answer. The readiness-preflight call site passes no. * refactor(apple-runner): build every lost-reply failure in one place with one reason Status recovery verdicts now carry either the runner's own answer or the facts of a lost reply. applyRunnerTransportRecovery turns a lost reply into one buildLostReplyError for a mutation, so runner_reply_lost is set on every lost-reply verdict (completed without a retained reply, still in flight, unknown or missing lifecycle, status probe failed, status unavailable) and the transport's reason stays under transportReason. A read keeps its transport error, which the read-only resend loop resends, so it never carries the reason. The four hand-built copies are gone. * fix(apple-runner): key a lost reply on its typed reason, never on transport text * docs(adr): name the lost-reply dispatch-disclosure rows in ADR 0011 * test(apple-runner): pair keyboardReturn with the keyboard enter request that issues it * test(apple-runner): pin the lost-reply branch in the connect-loop mutation test
#3071) * feat(errors): disclose whether a failed interaction reached the device details.dispatched carries no, yes, or unknown on interaction failures, anchored on the ADR 0014 side-effect seam: a failure before the seam is no, one after it is unknown unless a producer proves yes. Producers: the iOS runner reply and status-probe verdict tables, Android adb input and helper gestures, the Android post-tap in-app guard. The rows live in contracts/fixtures/dispatch-disclosure.json and every row is driven through its real producer by an owning test, gated for coverage. * fix(errors): only a pre-dispatch refusal discloses dispatched no The daemon fallback now fills unknown for any failure no producer classified; the session runtime revision was not a sound witness (it does not advance for sessionless requests, a replaced session resets it, and several side effects never advance it). Target resolution, keyboard occlusion, and targeted-touch admission stamp no where they refuse, and the covered-target refusal gains details.reason target_covered. * feat(apple-runner): publish runner_busy and runner_main_thread_timeout as details.reason RUNNER_BUSY and MAIN_THREAD_TIMEOUT reached the wire only as details.runnerErrorCode. The one runner-code classifier now also sets details.reason, so a consumer reads a single field; runnerErrorCode stays. * fix(ios): fall back from a direct selector tap only when it provably did not dispatch isDirectIosSelectorFallbackError matched message text (fetch failed, timed out, runner did not accept connection, invalid runner response), several of which arrive after the tap already ran; the tree path then tapped again. It now keys on details.dispatched === 'no'. The runner discloses no for a failure before the exchange, a pre-send recovery verdict (connect refused, readiness preflight, RUNNER_BUSY), a failed restart, a spent recycle budget, and its selector refusals (ELEMENT_NOT_FOUND, ELEMENT_OFFSCREEN, AMBIGUOUS_MATCH). * fix(errors): disclose dispatched yes where a producer ran the action on purpose scroll_no_progress is raised after the scroll gesture ran, and an Android fill whose verification fails after the last clear-and-retype pass typed the text; both now say yes instead of falling to the daemon's unknown. * fix(errors): claim no only for a connect that provably wrote nothing, and for read-only commands A runner connect failure classified for replay-after-restart said no even when an attempt had already posted the command (a fetch deadline, a simctl curl that exited after the POST). The connect loop now records whether every attempt was refused before writing (ECONNREFUSED, a usbmux socket that never opened, curl exit 7, a pre-attempt deadline), and only that failure is a pre-send refusal; the restart path reports unknown when the first attempt may have written. The daemon fallback stamps no for a command the registry declares read-only (recordingEffect observes-app). * docs(interaction): explain dispatched and the runner refusal reasons * refactor(apple-runner): the transport discloses dispatch itself instead of a write-evidence flag The runner transport is the only code that knows whether request bytes left, so it now stamps the public `details.dispatched` where it throws, replacing the private `runnerCommandUnwritten` details flag: - usbmux socket refusal, per-attempt connection deadline, simctl curl exit 7: `dispatched: no`; simctl curl with any other exit: `unknown`. - the connect loop's aggregate: `unknown` (overwriting) when any attempt may have written the command, else fill-if-absent `no`. - isRunnerCommandProvablyUnwritten reads `dispatched === 'no'` or a refused connection; markRunnerCommandUnwritten is deleted. Mutation checks: removing both `no` stamps (aggregate fill and curl exit 7) fails the golden row ios-runner.pre-send.connect-refused-before-write; stamping curl exit 7 as `unknown` fails it too; dropping the aggregate `unknown` overwrite fails the new waitForRunner test for a written attempt followed by a refused curl fallback. * refactor(contracts): drop the unread phase column from the dispatch-disclosure table No consumer read the before-seam/after-seam phase beyond a vocabulary membership check; the daemon seam does not infer a verdict from it. Removes the column from all 46 rows, the row type, DISPATCH_DISCLOSURE_PHASES, and the check. * fix(daemon): a read-only command discloses dispatched no over any producer verdict `dispatched` answers whether a resend is safe. A command the registry declares read-only (`recordingEffect: 'observes-app'`) has no side effect, so the interaction seam now sets `no` for it over a producer's `unknown` or `yes`, on both the thrown and the response path. Mutations keep the fill-if-absent `unknown`. The seam is renamed discloseInteractionDispatch since it no longer only fills. The golden row daemon.read-only-command now states it applies even when a producer classified the failure. ADR 0011 and the commands docs name the rule. Mutation checks: reverting the response path to fill-if-absent fails "a read-only command discloses no over a producer verdict" (get text whose capture throws dispatched unknown/yes); reverting the thrown path fails "a read-only command discloses no over a producer verdict it throws". Consumers unaffected: direct-ios-selector's fallback check and the runner lifecycle read `dispatched` inside the producer, before the seam, and only for tap/interaction commands. * feat(errors)!: dispatched is two-valued, no or unknown No consumer acts differently on `yes` and `unknown`: the docs said snapshot on both, and the only branching consumer (the direct iOS selector fallback) keys on `no`. No producer can prove execution on its failure path either: the runner documents XCTEST_RECORDED_FAILURE as "the action may not have been performed", the Android helper's ok=false comes from an outer catch that also covers pre-injection failures, and WebDriver has no execution receipt. `DispatchDisclosure` is now `'no' | 'unknown'`. Every former `yes` producer: - runner reply table and status recovery: an unlisted or missing code, `completed` without a retained reply, and a journal `failed` without a refusal code are `unknown`. - post-action guard: stamps nothing; the daemon seam fills `unknown`. - Android helper ok=false (one-shot and session) and fill verification: stamp `unknown` themselves, so their platform-level drivers pin the verdict without the daemon seam. The helper result carries no count of events injected before the failure, so it cannot say `no`. - scroll_no_progress stamps `unknown` itself: `scroll` routes through the generic runtime, not the interaction seam, so nothing would fill it. Golden table: every row keeps its scenario; each `yes` verdict is now `unknown`, and triggers that claimed execution say what is proven. `ios-runner.reply.executed-failure` is renamed `unlisted-code`. Mutations: runner default `?? 'unknown'` -> `?? 'no'` fails ios-runner.reply.{MAIN_THREAD_TIMEOUT,unlisted-code} and status.failed; dropping the seam's `discloseUnclassifiedDispatch(error, 'unknown')` fails post-action-guard.android-press-left-app; dropping the scroll_no_progress stamp fails daemon.scroll-no-progress. * fix(errors): a later step's refusal never claims no for the whole series `dispatched: no` describes the whole requested operation. A series that issues more than one device-reaching step (press --count split into runner sequence chunks, swipe --count, adb text chunks) let a later step's `no` (RUNNER_BUSY on chunk 2, a refused connect after the runner crashed) stand for the whole interaction, so a consumer would resend the series and the app would get extra taps. One kernel helper owns the rule: `discloseDispatchAfterSteps(error, dispatchedSteps)` keeps `no` only while no step was dispatched, and otherwise stamps `unknown` with `details.dispatchedSteps` (added to any count the failing step carries, so nested series compose). The local Android copy is gone; `discloseAdbInputDispatch` now classifies one input. Series loops, enumerated by git grep -n -A22 -E '^[[:space:]]*(for|while) \(' -- packages src \ | LC_ALL=C grep -a -E '^(packages/[^/]+/src|src)/' \ | LC_ALL=C grep -a -v -E '\.test\.ts|__tests__' \ | LC_ALL=C grep -a -E 'await (.*[^a-zA-Z])?(runCommand|interactor\.(tap|longPress)|doubleTap\.call|interactions\.gesture|runAndroidShell|runAdbShell|sendAndroidImeHelperText|clearAndroidImeHelperText|typeAndroidShellChunk|focusAndroid|scroll|runStep|runJson|typeText|runAlert)\(' and each hit that is an interaction series now applies the helper: - platform-apple runner-sequence.ts runApplePressSeries (sequence chunks) - contracts touch-runtime.ts executeGenericPress (press --count on Android, Linux, and every generic interactor) - daemon interaction-gesture.ts runSwipeRepetitions (swipe --count) - platform-android text-input.ts: fillAndroid (focus tap, then a pass), fillAndroidImeHelper, typeAndroidImeHelper, typeAndroidShell, clearFocusedText - platform-android device-input-state.ts keyboard dismiss keyevents - capture-kit scroll-edge-state.ts and daemon scroll-until.ts scroll passes - maestro daemon-runtime-port-observation.ts scrollUntilTypedMaestroTarget - platform-linux runPacedScrollSteps, platform-web runPacedScroll, provider-limrun per-character typeText Left as is: platform-apple alert.ts actOnAppleAlert retries only after an ALERT_NOT_FOUND refusal (`no`), so no earlier attempt dispatched; alert.ts 84 is a read poll. The other hits (app lifecycle, settings, permissions, notifications, perf, macOS host provider, ime-helper broadcast extras) are not interactions or are calls after the loop. WebDriver actions land with their own package branch. Table rows and drivers: ios-runner.series.later-chunk-refused (press --count 25 over the real send stack; chunk 2 answers RUNNER_BUSY -> unknown, dispatchedSteps 1) and daemon.series.swipe-later-repetition- refused (swipe --count 2 through the daemon handler; repetition 2 fails `no` -> unknown, dispatchedSteps 1). Mutations: `throw error` instead of the helper in runApplePressSeries fails ios-runner.series.later-chunk-refused; the same in runSwipeRepetitions fails daemon.series.swipe-later-repetition-refused; dropping the `dispatchedSteps === 0` guard in the helper fails its kernel test. * fix(apple-runner): unlisted runner codes are unknown; pre-gesture refusals are no An unlisted or missing runner code already reads `unknown` (the default became `unknown` with the two-valued type), in the reply table and in the status-recovery journal path, which share classifyRunnerReportedError. This adds the one code the runner provably emits before any gesture: INVALID_ARGS -> no. Every emission in the runner sources (apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/): - RunnerTests+Transport.swift:23, :106, :169: request body too large, not UTF-8, or undecodable; nothing is dispatched. - RunnerTests+CommandDispatch.swift:33, :59, :78: status, appState, pasteboardWrite argument guards, before the read or the pasteboard write. - RunnerTests+CommandExecution.swift:256, :296, :316, :346 (scroll and desktopScroll direction and duration guards), :290 and :336 (no gesture or wheel plan), :593 and :599 (missing or invalid gesturePlan): each returns before executeScrollDragGesture, desktopScrollAt, or plannedGestureExecution. - RunnerTests+ScrollDragExecution.swift:86 (non-finite drag coordinates) and :159 (no synthesized coordinate frame), before synthesis or dragAt. - RunnerTests+SequenceExecution.swift:25, :28, :130, :135 (via :219): sequence validation, before assembleSequenceExecution at :46. Not listed, with the evidence: - UNSUPPORTED_OPERATION stays `unknown`. It is also the reply after a gesture ran: RunnerSynthesizedGesture.m:350-353 returns "private XCTest event synthesis failed" after synthesizeWithError: ran, and RunnerTests+TvRemote.swift:119-126 returns "element tap failed" after element.tap() raised. The row ios-runner.reply.UNSUPPORTED_OPERATION pins it. - "Active app interaction viewport is unavailable" (RunnerTests+CommandExecution.swift:628) is a pre-gesture refusal, but its code is COMMAND_FAILED; mapping it would key on message text. It stays `unknown` until the runner gives it its own code. Mutations: removing the INVALID_ARGS entry fails ios-runner.reply.INVALID_ARGS (driven through the real reply path against the fake runner); adding UNSUPPORTED_OPERATION as `no` fails ios-runner.reply.UNSUPPORTED_OPERATION. * fix(daemon): every failed interaction carries dispatched, plain errors included The interaction seam and the runner lifecycle both rethrew a failure that was not an AppError untouched, so a plain Error from a backend reached the wire without `details.dispatched`. Both now normalize at the boundary with asAppError (fallback code UNKNOWN, the same code the request router's normalizeError gives a plain Error, so the wire code does not change) before disclosing: the seam fills `unknown` (or `no` for a read-only command), and the runner lifecycle keeps its `no` for a failure before the exchange started. Tests: a plain Error thrown by the bound runtime reaches the caller of handleInteractionCommands as an AppError whose normalized details say `unknown`; a plain Error from ensureRunnerSession reaches the caller of runAppleRunnerCommand with `no` and no command sent. Mutations: restoring the raw rethrow for plain Errors in the seam fails the first; restoring `if (!(error instanceof AppError)) throw error` in executeRunnerCommand fails the second. * fix(errors): close the smaller dispatch-disclosure gaps from review - adb text: only a non-empty chunk is a dispatched step, so `type "\nx"` with adb missing fails its leading ENTER with `no` instead of `unknown`. Row android-adb.input-text.tool-missing-before-first-input. - The daemon.unclassified.session-replaced row is dropped: its driver replaced the session inside the touch, after the runtime had read it, and the seam reads only the failure (interaction-dispatch-disclosure.ts has no reference to the session store), so no placement of the replacement can change the verdict. daemon.unclassified covers the case. - waitForRunner: the early-exit error after the loop gets the same write evidence as the connect error, so a runner that exited before any attempt could write says `no`. - fill: admitFill wraps the whole preamble (surface policy, target parse, bind, recorded-parameter and settle-flag checks, @ref preamble); a refusal it returns or throws is `no`. There is no single admission seam across interaction commands (press has admitTargetedTouch, fill had none), so the fill preamble is wrapped as a whole, and prepareTouchDispatch, the bind every touch route shares, marks its own refusal. Row daemon.refusal.fill-admission. - Driver ownership compares slash-separated paths on every platform. - fill: the second-pass input failure is its own row, android-adb.fill.second-pass-input-failed (unknown, dispatchedSteps 2), driven through fillAndroid; android-adb.fill.unverified keeps the verification-exhausted case. - Table checks: fallsBack exactly on maestro-direct rows, implementedBy only on webdriver rows, values only no or unknown, and every driver file carries the per-row loop DISPATCH_DISCLOSURE_ROW_LOOP. - Tests: the overwrite test overwrites a seam-filled value; the direct iOS fallback pins that each legacy message falls back with `no` and not without it; the direct-iOS suite resets its capture mock per test; the refused-fallback transport test asserts the fallback ran once. Mutations: counting an empty chunk as a step fails tool-missing-before-first-input; dropping the focus-tap count in fillAndroid fails second-pass-input-failed; returning admitFill's refusal unwrapped fails daemon.refusal.fill-admission; throwing the bare early-exit error fails the new waitForRunner test; a fallsBack or implementedBy on a daemon row fails the vocabulary test; skipping a driver's per-row loop fails the driver-file test; a message sniff returning false for legacy texts fails the direct iOS fallback test. The path-separator fix has no POSIX mutation. * refactor(interaction): split the series and fill admission below the complexity gate fallow flagged typeAndroidImeHelper, executeGenericPress, and readFillAdmission after the series and admission changes. The two Android text loops, which had the same shape, now share one plan (planAndroidTextSteps) and one sender (sendAndroidTextSteps) that owns the dispatched-step count; the ENTER between lines goes through pressAndroidEnterKey, so a missing adb before any input is `no` on the IME path too. executeGenericPress delegates one press to pressOnce, and fill admission reads its session and target in readFillTarget. Behavior is unchanged otherwise: pauses fall after every chunk except the last of the text, and an empty line still sends nothing. The existing Android dispatch-disclosure, text-input, and fill suites pass unchanged; counting an empty chunk as a step in sendAndroidTextSteps still fails android-adb.input-text.tool-missing-before-first-input. * chore(gates): re-derive the DispatchDisclosure wire digest for the two-valued type `DispatchDisclosure` is new on this branch (unreleased, so no ack), and its declaration changed to `'no' | 'unknown'`; the digest is the one test/wire-compat/wire-compat.test.ts prints from declaration-digest.ts. The DaemonError and NormalizedError acks keep their digests; their rationale now names the two values. * fix(errors): a failure after the mutation was sent never keeps a producer's no `dispatched: no` may describe a failure only while nothing reached the device. Two owners now record each mutation send that returned and downgrade any later failure of the same command through the kernel rule (`no` -> `unknown`, sent count added to `details.dispatchedSteps`): - Daemon: a request-scoped `RequestDispatchLedger`. The sends are the ADR 0014 side-effect seams of the interaction route (every `expireRefFrame` there): the runtime backend touch members and `performGesture` (interaction-runtime.ts), the direct iOS selector tap, and `type`. `discloseRequestDispatch` reads the ledger. The swipe series counter is superseded by the ledger and removed. - Maestro port: every `MaestroPublicOperation` kind is classified in `MAESTRO_OPERATION_MUTATES` (exhaustive record); `invokeMaestroPublicOperation` counts a mutating send that returned ok, and each port operation (one Maestro command) discloses the sends it made. This replaces the scrollUntilVisible counter, which counted only after the settle, and also covers inputText, eraseText, and tapOn retries. Post-dispatch phases enumerated: `grep -rn expireRefFrame src/daemon/interaction` (4 send seams); after each: afterRun escape guard, Android dialog readiness after-phase, response payload build, finalize. The runtime `--settle`/`--verify` observation (post-action-observation.ts, settle.ts) is best-effort and never throws, so it cannot carry a `no` out; the escape guard's foreground read can. Rows: daemon.post-dispatch.press-then-foreground-read-refused, maestro-port.scroll-until.capture-after-scroll-refused, maestro-port.input-text.settle-capture-refused, maestro-port.scroll-until.first-capture-refused (control: no send, no stands). Mutations: dropping the ledger increment in `sendRecordedMutation` fails the press and swipe rows; dropping the increment in `invokeMaestroPublicOperation` fails both maestro-port unknown rows. * fix(errors): composite and helper-IME steps disclose through the same helper A `no` from a later sub-step of one interactor operation no longer describes the whole operation. Composite enumeration: interactor operations that issue more than one device-reaching send whose producer can say `no`. The `no` producers are (grep "discloseDispatch(.*'no'"): adb input (`discloseAdbInputDispatch`), the Apple runner, and daemon refusals. Their multi-send operations: - Android `doubleTap` (src/core/interactors/android.ts): two `input tap` sends, uncounted. Now `doubleTapAndroid` in input-actions.ts discloses a failure of the second tap through `discloseDispatchAfterSteps(error, 1)`. - Android type/fill/clear, keyboard dismiss, Apple press series and sequence chunks: already counted on this branch. - Apple `doubleTap`, limrun and WebDriver `doubleTap`, HarmonyOS doubleClick: one send each. Linux doubleClick sends several xdotool/ydotool calls, but no Linux producer says `no`, so nothing can leak. Helper IME: every broadcast goes through `sendAndroidImeHelperBroadcast`, which now classifies its failure with `discloseAdbInputDispatch` (TOOL_MISSING before start is `no`, anything after start `unknown`), the same classifier the shell path uses. The classifier moves to adb-failure.ts so ime-helper.ts need not import the input-action module. Rows: android-adb.input-tap.double-tap-second-refused, android-helper.ime.broadcast-tool-missing, android-helper.ime.broadcast-failed. Mutations: rethrowing the raw error in `doubleTapAndroid` fails the double-tap row (`no`); removing the classification in the broadcast sender fails both IME rows (unset). * fix(daemon): every mutating command route discloses dispatch The dispatch seam was wired only into the interaction route, so a generic `scroll` (and every other non-interaction command) could fail with `dispatched` unset. The request router now owns the seam: `executeLockedRequest` wraps the handler chain and the generic dispatcher in `discloseRequestDispatch`, keyed by the registry `recordingEffect` (resolveCommandRecordingEffect): `observes-app` is `no`; `mutates-app` keeps a producer verdict, downgrades a `no` after a sent mutation, and fills `unknown` otherwise; a command without a declared effect passes through. `handleInteractionCommands` no longer wraps itself, so the seam runs once per request (a second pass would double the ledger count). The router also owns the ledger: it passes it through the handler chain to the interaction route and records the generic route's bound `execute` as its mutation send. A generic admission refusal is `no`, like the interaction admission refusals; the pre-dispatch helpers move next to the seam. `scroll --until` now proves its own pre-gesture refusals: a failure before the first gesture, other than the gesture send's own, is `no`. Residual: failures before the router seam (session lookup, lock, lease admission in prepareLockedRequestScope) stay unset. Commands outside the interaction and generic routes (open, close, install, settings, alert, react-native) get the fill but record no ledger sends; their post-mutation steps can still carry a producer `no` if one appears there. Rows (src/daemon/__tests__/request-dispatch-disclosure.test.ts, real router): daemon.route.scroll-transport-failure, daemon.route.scroll-until-before-first-gesture, daemon.route.scroll-then-dialog-read-refused, daemon.route.read-only-get. Mutations: routing without `discloseRequestDispatch` fails the scroll, Android scroll, and get rows; dropping the ledger wrap on `execute` fails the Android scroll row; rethrowing in scroll-until fails the --until row. * fix(apple-runner): pre-send validation and restart replay disclose their verdict Sequence validation ran inside the transport catch of runApplePressSeries, so an invalid step (a non-finite point) reached the wire undisclosed and the daemon filled `unknown`. Every validation refusal now comes from one constructor, `invalidSequence`, which carries `reason: runner_sequence_invalid` and `dispatched: no`; the series builds and validates every chunk before it sends the first, so a refusal can never follow a sent chunk. A restart replay after an unwritten first attempt returned the replay failure as is, so a replay failure without producer evidence escaped undisclosed. It now keeps an existing verdict, is `no` for a pre-send refusal (isRunnerPreSendRefusal), and is `unknown` otherwise. Rows: ios-runner.series.invalid-step-before-send, ios-runner.transport.replay-failed-after-unwritten-first-attempt. Mutations: `invalidSequence` disclosing `unknown` fails the invalid-step row; passing the replay failure through fails the replay row. * test(daemon): type the route driver's dialog-read stub as the adapter declares it * refactor(daemon): the request scope owns the dispatch ledger; every bound mutation records itself The ledger was threaded route by route (InteractionRouteInput, InteractionRuntimeInput, RequestHandlerChainParams, executePlatformCommand), so `find <q> type`, session, snapshot/alert, react-native, and record/trace routes never recorded their sends, and the Maestro port kept a second ledger plus MAESTRO_OPERATION_MUTATES, which restated the registry. The ledger now lives on the request execution scope and is handed to createRequestRuntimeBindings, the one place bindDevice and bindExactDevice narrow a binding. Every narrowed projection wraps each operation that RUNTIME_OPERATION_EFFECTS (src/daemon/runtime-operation-effects.ts) declares `mutates`, and records it when its send returns. The table is a Record over the runtime operation keys, so an operation with no declared effect fails typecheck and the completeness test. A replay action runs on a ledger of its own (req.internal.dispatchLedger) and moves its sends into the parent's. The router discloses around the scope's ledger: once it holds a sent mutation, a failure is `unknown` with the count added to details.dispatchedSteps, for any declared effect or none; a read with an empty ledger stays `no`. Deleted: the route dispatchLedger fields, sendRecordedMutation, MaestroMutationLedger, MAESTRO_OPERATION_MUTATES, and the per-command Maestro wrapper. The three maestro-port rows become one router row. Rows: daemon.route.find-type-refused-after-focus, daemon.route.session-close-then-finalize-refused, daemon.route.maestro-deferred-settle-capture-refused. Mutations: dropping recordBoundMutations from bindDevice fails the find-type, close, Maestro, and scroll-then-dialog-read rows; dropping the nested ledger from the replay action invoker fails the Maestro row; deleting one entry from RUNTIME_OPERATION_EFFECTS fails typecheck and the completeness test. * fix(daemon): a batch reports no only when no executed step was a mutation Each batch step ran as a nested request on a fresh ledger, and the batch descriptor declares no recording effect, so the router passed the batch response through with the failing step's `dispatched: no`. `[press A (lands), press 'label=Missing']` reported `no` with `executed: 1`, and a retry pressed A twice. The router now hands every nested request it invokes (batch steps, find's delegated click and fill) a ledger of its own through recordNestedRequests and moves its sends into the parent's ledger, as it already did for replay actions. The batch failure then reports `unknown` with the executed steps' sends in details.dispatchedSteps. A failed nested response keeps only the steps its producer counted inside the failing send, so the parent counts each send once. Rows: daemon.route.batch-mutation-then-refused-step, daemon.route.batch-read-then-refused-step. Mutation: passing handleRequest as the chain's invoke fails the mutation-then-refused row. * refactor(provider-webdriver): disclose through the kernel's DispatchDisclosure type The transport declared its own 'no' | 'unknown' union because the kernel type had not landed when #3070 merged. After the rebase both exist; the transport now uses the kernel's DispatchDisclosure, so the vocabulary has one declaration. * test(provider-webdriver): drive the three WebDriver disclosure rows through the real transport The webdriver rows waited on #3070 behind `implementedBy`. With it merged, the transport test owns them (`webdriver.` prefix): each sends a mutating POST through WebDriverTransport and the real fetch to a local socket. A port nothing listens on is refused before send (`no`); a driver that reads the body and never answers hits the transport deadline after send (`unknown`); a driver that answers 503 received the request (`unknown`). The row markers are gone and the triggers name what each driver does. The contracts fixtures subpath is a test-only import of the provider package; the eager-closure budgets are unchanged. Rows: webdriver.connect-refused, webdriver.timeout-after-send, webdriver.http-5xx-after-send. Mutations: the connect-refused error disclosing `unknown` fails the refused row; the timeout error disclosing `no` fails the timeout row; dropping the >= 500 disclosure fails the 5xx row. * docs(adr): state the read-only rule and the remaining disclosure gaps ADR 0011 now describes the one request ledger: operation effects declared once in RUNTIME_OPERATION_EFFECTS, bound operations recording their own sends, and nested requests moving theirs into the parent's. `no` for a read is defined as "a resend repeats no app-visible action": the registry has no trait that separates a pure read from a device control, so record, trace, and perf stay `observes-app`, and their recorder and profiler controls are declared `repeatable` at the operation level. The section lists the remaining gaps: failures before the router's locked scope (session, lock, lease, and policy admission) and the public runBatch with a caller-supplied invoke. The rewrite also removes the code spans that were hard-wrapped mid-token (`'mutates- app'`, `'observes- app'`) and the Maestro-port and `implementedBy` sentences, which no longer apply. commands.md states the read rule and the batch/replay rule. * fix(apple-runner): a press series counts completed taps, not sequence chunks details.dispatchedSteps of runApplePressSeries counted sequence chunks while every other series counts device inputs. It now counts the taps the runner reported completed: across finished chunks when a later chunk is refused, plus the failing chunk's own completed steps when a step inside it fails. The ios-runner.series.later-chunk-refused row now asserts 20, and its trigger says so. The runner disclosure test header names its third category: pre-send validation that refuses before anything is sent. Mutations: counting chunks fails the later-chunk row (1, not 20); dropping the failing chunk's completed steps fails the new runner-sequence test (20, not 22). * chore(gates): drop the implementedBy row marker and its check The webdriver rows were the only rows naming another branch; they now have a driver, so the field, its ownership filter, and the gate assertion that only webdriver rows could carry it have no remaining use. Every row must now have a driver file. * fix(batch): report unknown after an executed mutating step * fix(interaction): readiness refusals disclose dispatched no The readiness wait now refuses before any touch on main: a sparse capture (capture_sparse) stamps no where it is thrown, and an exhausted budget inherits no from the selector failure it decorates. A daemon table row drives the exhausted wait through the real readiness path. * fix(daemon): scroll loops leave counting bound sends to the request ledger Each scroll --until and edge pass runs the bound scrollDirection, which the request ledger already records, and the router adds the ledger to any producer count. The two loops counted the same sends again, so two passes reported four. Only a producer counting sub-steps inside one bound operation keeps its own counter. Two router rows pin exactly 2 after a refused third pass. * refactor(errors): one dispatched-steps rule, applied to details discloseDispatchAfterSteps now wraps detailsAfterDispatchedSteps instead of restating its rule. withoutLedgerSteps stays: a nested request's own router pass adds its ledger to the failure, which the parent then adds again through its own ledger. * refactor(daemon): unexport two types only their own modules read * fix(batch): a batch reports no only when every executed step is a declared read * docs(adr): state the batch dispatched rule as observes-app reads only, pin undeclared step * fix(daemon): delete the no-change interaction retry path (#3083) * fix(daemon)!: delete the no-change interaction retry path The daemon re-sent a coordinate tap when the next capture found the surface unchanged, if the request carried the interactionOutcome retry-on-no-change opt-in. Nothing in the repository set it, and the resend could not know whether the first tap was dispatched, so a tap that landed on a screen that had not re-rendered yet ran twice. Removed: - `interactionOutcome` from `CommandFlags`. No CLI token, env var, config key, Node option or MCP field ever carried it; the daemon now has no reader, and the session recorder keeps only declared flags, so a raw key sent by a peer is dropped. - `SessionState.pendingInteractionOutcome` and its type (the R7 owner entry leaves in the chore(gates) commit). - The pending-outcome branch of `resolveDeferredInteractionOutcome`, the `scheduleOutcomeRetry` mark, and every retry helper in interaction-outcome-policy.ts (`markPendingInteractionOutcome`, `retryPendingInteractionOutcome`, `InteractionRetryTap`, `classifyInteractionSurfaceChange`, the settle diagnostics). - interaction-retry-tap.ts and the `inspectFacts`/`bindDevice` params the snapshot capture carried only to build the retry seam. - `retryPositionals`/`scheduleInteractionOutcomeRetry` on `finalizeTouchInteraction`, and `pointPositionals`. Tests: - interaction-ios-tap-outcome.test.ts "a coordinate click the next snapshot finds unchanged is not re-sent": a click on 104,222, then a `snapshot` whose request carries the same runtime bindings the click used and returns the same tree, asserts one `press`. Mutation: restoring the base retry branch with its opt-in gate forced on (base `shouldRetryTouchOnNoChange` returning true) makes it see three presses (verified). - interaction-touch-runtime.test.ts "press @ref taps the resolved point once and records the ref". Mutation: recording the resolved point positionals instead of the request's `@e1` in finalizeTouchInteraction fails it. - interaction-outcome-policy.test.ts: the classifyInteractionSurfaceChange cases now pin areInteractionSurfaceSignaturesStable, which the stabilization loop and scroll movement still use. Mutation: dropping `stateMarkers(node)` from interactionSurfaceSemanticKey fails the checked-only flip case. - snapshot-handler-freshness.test.ts: the annotation-survival case now rides the post-gesture deferred capture. Mutation: returning only `{ snapshot }` from resolvedPostGestureCapture fails it (verified). * refactor(daemon): move the deferred outcome interface into post-gesture-stabilization.ts With the pending-outcome retry gone, deferred-interaction-outcome.ts only marks and resolves post-gesture stabilization (its R7 field) and Android freshness. It moves to src/daemon/post-gesture-stabilization.ts; the exported names stay, so callers change only their import path. No shim is left at the old path. The surviving cases of deferred-interaction-outcome.test.ts move unchanged into post-gesture-stabilization.test.ts, which already tested this module's stabilization loop. The capture-kit and fixture comments name the new owner. No new tests: the move carries its tests unchanged. * docs: drop the no-change retry from the deferred outcome docs CONTEXT.md's "Deferred interaction outcome" no longer lists outcome retry, docs/agents/selector-capture.md states that the daemon never re-sends an interaction to settle its outcome, and ADR 0004/0023 name post-gesture-stabilization.ts where they named the deleted module. * refactor(daemon): drop the unused logPath from the deferred outcome capture `resolveDeferredInteractionOutcome` read `logPath` only to build the runner context of a re-fired tap. With the retry gone nothing reads it, so the parameter leaves the capture interface and its one production caller (`captureSnapshot`). No new tests: removing an unread parameter changes no behavior; the typecheck is the gate that a caller still passing it fails. * chore(gates): retire the pendingInteractionOutcome R7 owner entry SessionState no longer has pendingInteractionOutcome, so its row leaves SESSION_STATE_FIELD_OWNERS, and the postGestureStabilization owner moves to src/daemon/post-gesture-stabilization.ts. R10 measures the merge-base, so the shrink (16 -> 15 writer-owned fields) banks without a baseline edit. * docs(agents): scope the no-resend rule to the daemon's deferred-outcome path * docs(agents): keep the deferred-outcome rule to one line inside the docs budget * refactor(move): keep the deferred outcome interface in deferred-interaction-outcome.ts Reverts 401869a47b. The module still exports markDeferredInteractionOutcome, resolveDeferredInteractionOutcome and DeferredInteractionOutcomeMark, owns Android freshness as well as post-gesture stabilization, and CONTEXT.md keeps "Deferred interaction outcome" as its glossary term. The surviving cases move back into deferred-interaction-outcome.test.ts, ADR 0004/0023, the capture-kit comment and the R7 owner entry name the original path. git diff -M90% --stat: docs/adr/0004-ios-snapshot-backend-strategy.md | 2 +- docs/adr/0023-end-state-hop-trace.md | 4 +- packages/capture-kit/src/post-gesture-stability.ts | 2 +- scripts/layering/session-state.ts | 2 +- .../__tests__/deferred-interaction-outcome.test.ts | 197 +++++++++++++++++++++ src/daemon/__tests__/is-runtime.test.ts | 2 +- .../__tests__/post-gesture-no-effect-claim.test.ts | 4 +- .../post-gesture-stabilization-fixtures.ts | 4 +- .../__tests__/post-gesture-stabilization.test.ts | 192 +------------------- .../__tests__/snapshot-runtime-disclosure.test.ts | 2 +- ...lization.ts => deferred-interaction-outcome.ts} | 0 src/daemon/direct-ios-selector.ts | 2 +- src/daemon/interaction-outcome-policy.ts | 2 +- .../find-target-activation-disclosure.test.ts | 2 +- .../interaction-capture-disclosure.test.ts | 2 +- .../interaction/internal/interaction-runtime.ts | 2 +- src/daemon/interaction/internal/types.ts | 2 +- src/daemon/request-generic-dispatch.ts | 2 +- src/daemon/scroll-movement.ts | 2 +- src/daemon/selector-capture-runtime.ts | 2 +- .../internal/session-open-execution.ts | 2 +- src/daemon/snapshot-capture.ts | 2 +- 22 files changed, 223 insertions(+), 210 deletions(-) * refactor(move): name the surface comparators interaction-surface-signature.ts With the outcome retry gone, interaction-outcome-policy.ts only builds and compares interaction surface signatures. It moves to interaction-surface-signature.ts with its test. stripInternalInteractionFlags moves to deferred-interaction-outcome.ts, which already reads flags.postGestureStabilization, and its test moves with it. Neither path has a fallow baseline entry, so no baseline changes. git diff -M90% --stat: docs/adr/0023-end-state-hop-trace.md | 2 +- packages/capture-kit/src/post-gesture-stability.ts | 2 +- packages/capture-kit/src/snapshot-chrome.ts | 2 +- src/daemon/__tests__/deferred-interaction-outcome.test.ts | 11 +++++++++++ .../__tests__/interaction-surface-baseline-evidence.test.ts | 2 +- ...policy.test.ts => interaction-surface-signature.test.ts} | 13 +------------ src/daemon/__tests__/post-gesture-no-effect-claim.test.ts | 2 +- .../__tests__/post-gesture-stabilization-verdict.test.ts | 2 +- src/daemon/__tests__/post-gesture-stabilization.test.ts | 2 +- src/daemon/deferred-interaction-outcome.ts | 12 ++++++++++-- ...n-outcome-policy.ts => interaction-surface-signature.ts} | 9 --------- src/daemon/interaction/internal/find.ts | 2 +- src/daemon/interaction/internal/interaction-common.ts | 2 +- src/daemon/scroll-movement.ts | 4 ++-- src/daemon/session-state.ts | 2 +- 15 files changed, 34 insertions(+), 35 deletions(-) * docs: state the current capture ordering without the retired retry Drop the no-resend sentence from the deferred-outcome mark comment, the history line from docs/agents/selector-capture.md, and the retry framing from the interaction-touch-runtime test header. * test(daemon): drop tests that pass with or without the retry Neither coordinate-tap test could set the retired opt-in, so one press held on the base branch too. Delete the unchanged-click test; keep the corroborated coordinate tap only for its corroboration assertions. The one captureSnapshot case left in snapshot-handler-capture-retry.test.ts moves to snapshot-capture.test.ts, which mirrors snapshot-capture.ts, without the Apple runner and iOS hint mocks only the retry cases used. * docs(adr): drop the automatic no-change retry from ADR 0014 No automatic no-change retry exists. Maestro retryTapIfNoChange sends an ordinary press that crosses the leaf seam like any other action. * test: drop no-op retry canary, fake timers in stabilization test, widen CommandFlags waiver * fix(apple-runner): restart without resending a command written then lost (#3082) * fix(apple-runner): name the lost reply when status cannot place a lost command A command whose reply was lost and whose status probe answers notAccepted (a restarted runner's empty journal) or an unnamed state now fails with details.reason runner_reply_lost beside dispatched unknown, so a caller keys on the typed reason instead of the message before it observes the screen. Tests (fake runner server through runAppleRunnerCommand and the recovery module): - runner-recovery-wiring: press and fill (runner tap and type) accept the command, drop the connection, status answers notAccepted: one dispatched command, error dispatched unknown with reason runner_reply_lost, the session is invalidated, and the next command gets a runner. Mutation: delete `reason: RUNNER_REPLY_LOST_REASON` from the unknown-state error -> both rows fail. - runner-recovery-wiring: a snapshot whose reply is dropped is resent and succeeds. Mutation: declare snapshot with DEFAULT_TRAITS (not read-only) -> fails (the lost read goes to status recovery and is not resent). - runner-command-recovery: notAccepted yields reason runner_reply_lost and dispatched unknown. Mutation: same deletion as above -> fails. * fix(apple-runner): restart without resending a command written then lost A connect-shaped failure keeps its restart, but the command is sent again on the restarted runner only with proof the first send could not have run it: the attempt provably wrote nothing (a connect refused before the write, dispatched no), the readiness preflight gave up before the write, or the command is read-only by its runner trait. A mutating command whose first attempt may have written it now fails with reason runner_reply_lost and dispatched unknown after the restart; a restarted runner's journal is empty, so no status answer can prove the first send did not run. A read-only command's restart failure discloses no: a read has no side effect to repeat. dispatch-disclosure.json: ios-runner.transport.written-then-lost keeps unknown with a trigger that says the command is not resent; ios-runner.transport.read-only-written-then-lost is new and discloses no. Tests (runner-lifecycle-dispatch-disclosure, real connect loop over a stubbed fetch and simctl curl exit 28, restart succeeds): - written-then-lost: tap restarts the runner, is not sent on the restarted runner, fails unknown with reason runner_reply_lost. Mutation: make canResendAfterRestart return true -> fails (the tap is replayed and the request succeeds). - read-only-written-then-lost: snapshot is sent again on the restarted runner and its failure discloses no. Mutations: drop the read-only term from canResendAfterRestart -> fails (no resend); drop the read-only branch of discloseRestartDispatch -> fails (unknown). runner-command-retry: four mutating-restart fixtures now carry the proof a real refused connect stamps (dispatched no); without it they would test the removed replay. * test(apple): share the unwritten connect refusal fixture to keep the retry test file within its size ratchet * fix(apple-runner): report the restart transport fact for a read, not no The restart path disclosed `no` for a read-only command whose first send may have run, while the status-recovery path discloses `unknown` for the same read (read_only_completed_without_retained_response, and the notAccepted/unknown-state fallthrough). Both runner paths now report the transport fact: a first attempt that may have written the command is `unknown`, read or mutation. The read-only `no` has one owner, the daemon router seam (request-dispatch-disclosure.ts), which keys it on the registry's recordingEffect `observes-app` for every producer (daemon.read-only-command). A second copy keyed on the runner's own read-only trait reconstructed that rule from a different source of truth, and the two runner paths disagreed. The recycle-budget refusal on the restart path also stops deriving `no` from the read-only trait and uses only whether the first attempt was provably unwritten. The runner still resends a read on the restarted runner (canResendAfterRestart); only the disclosure changes. dispatch-disclosure.json: - ios-runner.transport.read-only-written-then-lost: now `unknown`; the trigger names the router rule that turns it into `no` for the caller. - ios-runner.status.read-only-completed-without-retained-reply: new, `unknown`, driven through the real connect loop and recovery module. It pins the status-recovery read to the same value as the restart read. Mutations: - put back `if (isReadOnlyRunnerCommand(evidence.command)) return discloseDispatch(error, 'no')` in discloseRestartDispatch -> read-only-written-then-lost fails. - set `dispatched: 'no'` on read_only_completed_without_retained_response -> status.read-only-completed-without-retained-reply fails. * test(apple-runner): drive a real runner death through the fake runner runner-recovery-wiring: the notAccepted mutation test no longer scripts a second `ok` for the mutation. Nothing consumed it, and it suggested that a resend was expected. Its comment now says what the test models: status answers notAccepted, the way a runner whose journal did not survive a restart would. The path is status recovery, not the daemon's restart. The fake runner server gains an `exit` reply: it hangs up and stops listening, like a runner process that died mid-command. Before this, a test could only hang up on one request while the same server stayed up. Two new tests use a second server as the restarted runner: - press/fill (runner tap/type): the runner dies on the command. The mutation is sent once, is not sent to the restarted runner, fails with dispatched unknown, and the dead session is invalidated. A readText on the next request is served by the restarted runner. Mutation: declare tap and type with READ_ONLY_TRAITS (the resend trait) -> both rows fail. - get (runner readText): the runner dies mid-read. The connect loop gives up with a possibly-written POST (simctl curl exit 28), the daemon restarts the runner (runner_connect_failed_before_command_send), and the read is resent once and succeeds. Mutation: reduce canResendAfterRestart to `evidence.firstAttemptUnwritten` -> fails (no resend). With the real transport, a mutating command never reaches restartSessionAndRunCommand with a possibly-written first attempt. Mutations are sent by sendRunnerCommandOnce, and only the connect loop (waitForRunner, used for reads and the readiness preflight) stamps runner_connect_refused. So the mutating arm of canResendAfterRestart is covered only by the mocked lifecycle driver (ios-runner.transport.written-then-lost). A dying runner sends a mutation to status recovery, and a failed status probe rethrows the transport error without a runner_reply_lost reason. * fix(apple-runner): name the lost reply when the status probe fails A runner that dies mid-mutation drops the reply, and the status probe that follows fails. The error that reaches the caller was then a bare transport error (`fetch failed`, dispatched unknown), with no typed reason. #3074 requires a reason that names the lost reply, so this path now builds an error with reason runner_reply_lost (the name the restart arm already uses) and recovery status_probe_failed. The error keeps the transport message and details, and the transport error as cause. The session is still invalidated, and the command is still not resent. ADR 0011 gap list: the row ios-runner.transport.written-then-lost proves the rule behind canResendAfterRestart's mutating arm, not a production route. Over the real transport the arm cannot be reached, because mutations go through sendRunnerCommandOnce and only the connect loop raises the restart trigger. Tests: - runner-command-recovery: a failing status probe now asserts reason runner_reply_lost, recovery status_probe_failed, dispatched unknown, and the transport error as cause. - runner-recovery-wiring: press/fill whose runner dies mid-command (real fake-runner transport) now also asserts reason runner_reply_lost. Mutation: drop `reason: RUNNER_REPLY_LOST_REASON` from buildStatusProbeFailedError -> all three fail. * test(apple-runner): hang up on every read attempt without counting the connect loop The read-only lost-response driver scripted 20 hang-ups and a 400 ms timeout. That count depended on the connect loop's retry cadence (RUNNER_CONNECT_ATTEMPT_INTERVAL_MS): with a faster cadence the queue runs out, and the fake answers the next attempt with `ok`. The fake runner server now has a `hangUpAlways` reply that is never consumed, so the read is refused until the loop gives up, however often it retries. The timeout now only bounds how long the loop runs. Mutation: set RUNNER_CONNECT_ATTEMPT_INTERVAL_MS to 10 or 50 (with a 1 ms retry base delay) -> ios-runner.status.read-only-completed-without-retained-reply still passes. * test(apple-runner): pick the fake runner's next reply in one helper The request handler of the fake runner server went over the fallow complexity threshold after hangUpAlways was added. The choice of the next scripted reply is now in takeScriptedResponse. Behavior is unchanged: the apple-runner suite passes (768 tests). * refactor(apple-runner): gate the restart resend at the connect-failure call site A mutation never reaches the connect loop, so the no-resend arm after a restart had no production route. The call site now restarts only when the first attempt provably wrote nothing or the command is read-only, and the arm and its error builder are gone. The written-then-lost row is driven through a fake runner that dies mid-command. A failed status probe keeps a transport reason under transportReason. * test(apple-runner): name runner_reply_lost in the lost-reply test titles * test(apple-runner): the runner read-only set matches the registry's observes-app commands * fix(apple-runner): decide the restart resend in the connect-failure classifier shouldRestartRunnerBeforeCommandSend now takes the command and owns the write-proof rule: a connect-loop failure restarts and resends only when no attempt could have written the command, or the command is read-only. An exit-28 POST that carries dispatched unknown no longer restarts a mutation, whatever call site asks. The call-site guard in executeRunnerCommandAttempt is gone. * refactor(apple-runner): carry the first attempt's dispatch disclosure through the restart restartSessionAndRunCommand takes firstAttemptDispatched: DispatchDisclosure instead of a boolean that three ternaries decoded back into no or unknown. resolveFirstAttemptDispatch computes it once beside isRunnerCommandProvablyUnwritten, and the restart rule reads the same answer. The readiness-preflight call site passes no. * refactor(apple-runner): build every lost-reply failure in one place with one reason Status recovery verdicts now carry either the runner's own answer or the facts of a lost reply. applyRunnerTransportRecovery turns a lost reply into one buildLostReplyError for a mutation, so runner_reply_lost is set on every lost-reply verdict (completed without a retained reply, still in flight, unknown or missing lifecycle, status probe failed, status unavailable) and the transport's reason stays under transportReason. A read keeps its transport error, which the read-only resend loop resends, so it never carries the reason. The four hand-built copies are gone. * fix(apple-runner): key a lost reply on its typed reason, never on transport text * docs(adr): name the lost-reply dispatch-disclosure rows in ADR 0011 * test(apple-runner): pair keyboardReturn with the keyboard enter request that issues it * test(apple-runner): pin the lost-reply branch in the connect-loop mutation test * test(apple-runner): type the read-only parity rows as the registry request and drop the cast * fix(apple-runner): keep the transport error's own hint on a mutation lost reply * refactor(apple-runner): build the lost-reply message once from the command and cause * docs(wire-compat): say an older peer ignores or rejects an unknown flag
Summary
A mutating command whose reply stays lost now fails with
details.reason: runner_reply_lostanddispatched: "unknown". This covers a runner that dies mid-command, and every status-recovery verdict that finds neither a retained result nor a runner answer: completed without a retained reply, stillacceptedorstarted,notAccepted, an unknown or missing lifecycle state, a failed status probe, and a command with no id to probe. The command is sent once and never replayed, and the next command starts a new runner. One builder makes this error. The transport's own reason stays underdetails.transportReason. A read keeps its transport error, so the read-only resend loop resends it, and the read never carries the reason.The restart-and-resend rule after a connect failure is owned by
shouldRestartRunnerBeforeCommandSend(error, command). The rule restarts only when no attempt could have written the command (resolveFirstAttemptDispatch(error) === 'no') or when the command is read-only. A simctl curl exit-28 POST carriesdispatched: unknown, so it never restarts a mutation. The restart carries the first attempt'sDispatchDisclosure, not a boolean.agent-device longpress @e3 5000 --json # runner killed mid-press: runner_reply_lost, sent onceCloses #3074. Part of #3069. Stacked on #3071. 14 files touched.
Validation
Tested commit: 4e88908.
pnpm check:affected --run(fail-open full set): every runnable check passed. In vitest-related, 1333 of 1334 files passed. The one failure was a 5 s timeout under host load in a file this PR does not touch (screenshot-density.test.ts, andpress-target-readiness.test.tson the run before). Each passes alone. The checks after vitest-related were then run on the same commit, and all passed. Also run:pnpm typecheck,pnpm lint, apple runner vitest (167 files, 1704 tests),check:layering,check:fallow,format:check, the eager-closure budget test, and the runner read-only parity test.Live on f52f120 (iPhone 17 Pro simulator, iOS 26.2, Release test app):
longpress @e3 5000 --debug, with the runner process killed 1.3 s after the runner accepted the command. The request log has oneios_runner_command_sendforlongPress, then a failedstatusprobe,invalidation_decisionstatus_probe_failed, andios_runner_session_invalidatedtransport_error_after_command_send. The error hasreason: runner_reply_lost,recovery: status_probe_failed,dispatched: unknown, and norunnerRestarted.snapshot -ithen succeeded, and the next runner command started a new runner session (port 61516, was 61032). The structural changes after that commit do not change this path.