Skip to content

fix(record): keep the daemon alive when a record request times out - #3199

Merged
thymikee merged 3 commits into
callstack:mainfrom
pvedula7:fix/recording-terminal
Oct 4, 2026
Merged

thymikee merged 3 commits into
callstack:mainfrom
pvedula7:fix/recording-terminal

Conversation

@pvedula7

@pvedula7 pvedula7 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

On iOS simulators, record stop on a long recording can outlast the 90s client timeout (we saw 40–122s). record used the default reset-daemon timeout policy, so the client killed the daemon mid-export. That left the session's screen-recording manifest in a non-terminal state. A record-only session holds no device claim, so nothing reconciles it, and every later record start on that device fails with cleanup-unconfirmed ("has not reached a confirmed terminal state") until someone runs record stop on that exact session.

In practice this breaks retries in tester-army/e2e: attempt 0 fails, its video stop times out, and attempt 1 can't start recording.

record now uses preserve-daemon, like snapshot and press: the export finishes, a retried record stop returns it (#2534), and the next start is admitted. The local timeout hint now says to retry record stop, as the remote hint already did. 5 files.

Validation

  • f2f48a8a4: pnpm check:affected --run passes. New route test: a local record stop timeout never signals the daemon (fails without the fix).
  • iOS 26.5 simulator: replayed the e2e retry sequence with a stalled export. main: the retry's record start fails with cleanup-unconfirmed. This branch: the daemon survives, the retried stop returns the video, and the next record start succeeds.
  • Risk: a genuinely hung record request no longer resets the daemon; it's cancelled like other preserve-daemon commands.

Review in cubic

A record stop export routinely outlasts the 90s client envelope on long
recordings. The default reset-daemon policy SIGKILLed the daemon
mid-export, leaving the session's screen-recording manifest open with no
owner. A record-only session holds no device claim, so nothing reconciled
it, and every later record start on the device refused with
cleanup-unconfirmed until that exact session ran record stop.

record now preserves the daemon on timeout: the export finishes, a retried
record stop serves it, and the next record start is admitted. The local
timeout hint now names that retry, as the remote one already did.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 5 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread src/daemon-client/daemon-client-timeout.ts Outdated
@thymikee

thymikee commented Oct 3, 2026

Copy link
Copy Markdown
Member

Reviewed f2f48a7... correction: the reviewed commit is f2f48a8. The code looks good, and I found nothing that blocks merge. CI is green, with one check reported and passing, and there are no conflicts.

Not blocking: the repeat-safe record stop paragraph in https://github.com/callstack/agent-device/blob/f2f48a8/website/docs/docs/commands.md#L1095 still calls the in-progress export typical for a remote daemon, but a local daemon now survives the timeout too, so please drop that qualifier or name both. You can take or leave this.

The open inline thread from Cubic on the "still exporting" hint wording still applies (#3199 (comment)). A stop that was still queued is dropped at lock entry, so no export ran. The retry instruction is right, and only the wording overstates.

The PR body claims live simulator validation, but I could not see a transcript, so I can't confirm it. I did not trace physical iOS or macOS runner-backed recordings. cleanupTimedOutIosRunnerBuilds still kills the runner xcodebuild on each local timeout (daemon-client-timeout.ts:70). That predates this PR, and I did not confirm it. If it hits a runner-backed stop, "still exporting" may be false and recovery would rely on the stop-sequence journal. I also did not check whether runRunner picks up an ambient request id, or how Android record stop uses the request signal beyond start. If you have the transcript, please link it, and ideally include a runner-backed stop that times out and a record stop retry that returns the finished file.

A record stop that times out while still queued for the device lock is
dropped at lock entry, so no export ran. The timeout hint now says the
daemon may still be exporting and keeps the retry instruction. The docs
no longer call the in-progress export typical only for a remote daemon,
since a local daemon now survives the timeout too.
The client's timeout cleanup pkills every Apple runner xcodebuild on the
host. On a physical iOS device or macOS the runner is the recorder, so a
timed-out record stop could kill the export the preserved daemon was
still finishing. record now skips that sweep. The daemon still cancels
its own runner work for a timed-out request through the request signal,
and the exclusion is keyed on the command, never on the declared
platform.
@pvedula7

pvedula7 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor Author

Two commits on top of f2f48a8:

  • ab1fc6d: the hint says the daemon "may still be exporting", and the docs paragraph now says "on a local or remote daemon".
  • ba61c63: cleanupTimedOutIosRunnerBuilds was an issue. Simulators record with simctl, but physical iOS (CoreDevice) and macOS record through the runner, and that cleanup's pattern matches every runner xcodebuild on the host. The repro at f2f48a8 logged timedOutRunnerPidsTerminated: 1 on the record stop timeout. record timeouts now skip the cleanup (decided by command, not declared platform), and the route test asserts no pkill.

Live run on ba61c63 below. I didn't do a runner-backed stop: simulators always use simctl and I didn't have a cabled device. The simulator run shows the cleanup no longer firing (0 terminated, runner pid unchanged) and the retried stop returning the file. pnpm check:affected --run passes on ba61c63.

Not changed: other commands' timeouts still sweep every runner on the host. Happy to open an issue for that.

Live run transcript

Live run: record stop times out, retry returns the file (ba61c63)

iPhone Air simulator, iOS 26.5, isolated --state-dir, fresh pnpm build. Simulator recording is simctl recordVideo; the overlay burn-in helper was SIGSTOPped so the export runs past the 90s client window.

$ agent-device boot --session rec-timeout-3199 --platform ios --udid <sim>
$ agent-device open com.apple.Preferences --session rec-timeout-3199
$ agent-device record start take1.mp4 --quality medium --scope device --session rec-timeout-3199 --json
  success: true, recording: started, recordingBackend: "simctl recordVideo"
$ for i in 1..5: agent-device press 'label="General"' && agent-device back     # starts the session's runner
  daemon pid 95421 | this sim's runner xcodebuild pid 95884 | other sim's runner pid 10559

[17:59:09] $ agent-device record stop --session rec-timeout-3199 --json
[17:59:10]   (overlay export helper SIGSTOPped)
[18:00:39]   error: COMMAND_FAILED "Daemon request timed out"
             details: { reason: daemon_transport_timeout, timeoutMs: 90000 }
             hint: The daemon may still be exporting the recording. Run agent-device record stop
                   --session rec-timeout-3199 again to wait for that export and receive the completed recording.
             diagnostic daemon_request_timeout: { command: record, timedOutRunnerPidsTerminated: 0,
                                                  daemonPreservedAfterTimeout: true }
           after the timeout: daemon 95421 alive | this sim's runner 95884 still running | other sim's runner 10559 still running

[18:00:39]   (overlay export helper resumed)
[18:00:43] $ agent-device record stop --session rec-timeout-3199 --json
             success: true, recording: stopped, recorder: confirmed, durationMs: 106648, outPath: take1.mp4
           ffprobe take1.mp4: mp4, 13.15s, 14,174,294 bytes

$ agent-device record stop --session rec-timeout-3199 --json        # repeat
  success: true, recording: stopped
$ agent-device record start take2.mp4 --scope device --hide-touches --session rec-timeout-3199 --json
  success: true, recording: started
$ agent-device record stop --session rec-timeout-3199 --json
  success: true, recording: stopped
$ agent-device close --session rec-timeout-3199
  success: true

For comparison, the same repro at f2f48a8 logged timedOutRunnerPidsTerminated: 1 on the record stop timeout: the client's pkill sweep killed runner xcodebuild. On a simulator the recorder is simctl, so the export survived anyway. On a physical iOS device or macOS the runner is the recorder.

Not covered: a runner-backed record stop (physical iOS over CoreDevice, or macOS). Simulators always record with simctl (startAppleRecording routes device.kind === 'simulator' to simctl), and I had no cabled device for this run. That path is covered by the unit test, which asserts no pkill on a record timeout.

@thymikee

thymikee commented Oct 4, 2026

Copy link
Copy Markdown
Member

This PR is ready. The cubic-dev-ai P2 thread on the retry hint (daemon-client-timeout.ts:141-148) is fixed at ba61c63, so you can resolve it: #3199 (comment). The earlier findings from f2f48a8 are addressed, and the one check reported on this head passes.

Not blocking, and you can take or leave these: a timed-out local record start --platform ios no longer runs the sweep, but the timeout hint at https://github.com/callstack/agent-device/blob/ba61c63/src/daemon-client/daemon-client-timeout.ts#L158 still says Apple runner work was aborted, because the declared platform alone sets the evidence flag. Gating that note on a real termination count, or on a sweep-eligibility trait next to timeoutPolicy, would keep the hint and the sweep on one source. A batch that contains record stop times out as batch, so the command carve-out at line 72 does not cover it. Open PR #3193 deletes this sweep and rewrites the same handler and route test, so whichever PR lands second needs a rebase, and the carve-out here becomes dead code if #3193 lands first.

The runner-backed recording path (physical iOS over CoreDevice, and macOS) was not run live, and the PR says so. Coverage is the unit and route tests plus a simulator run in the transcript. I did not run the tests locally or reproduce the transcript.

No conflicts. Before merge, sequence this against #3193 or rebase onto it.

@thymikee thymikee added the ready-for-human Valid work that needs human implementation, judgment, or maintainer merge label Oct 4, 2026
@thymikee
thymikee merged commit 294dc7d into callstack:main Oct 4, 2026
16 checks passed
thymikee added a commit that referenced this pull request Oct 4, 2026
Rebase resolution against main@294dc7d (#3199), which rewrote the same hint
formatter this PR restructures. Both changes are semantic and both survive
here:

- #3199 made `record` a preserve-daemon policy and generalized the
  keep-exporting retry to LOCAL timeouts, because a preserved daemon may
  still be exporting (and a stop queued for the device lock is dropped
  before any export starts, hence "may"). This PR's shape decides
  "preserved" from the probe verdict + declared policy instead of the old
  sweep/reset booleans, so the retry branch keys on `!resetDaemon` and
  `action === 'stop'` BEFORE the remote split, with the remote wording
  merely adding "remote ". The old remote-only "is still exporting" claim
  becomes the shared "may still be exporting" one at #3199's second commit.
- The reset branch stays this PR's: a daemon the probe proved unresponsive
  was SIGKILLed, is no longer exporting, and gets the probe-verdict wording
  with no keep-exporting promise — the exact case #3199's note about a
  mid-export reset predates, where the probe now makes the reset rarer.

Route coverage gains a local `record stop` row: the positional rides the
real transport context into the hint, the preserve policy must skip the
probe entirely (connections: 1), and no sweep call may fire — #3199's
runner-survival requirement, which this PR satisfies by removing the sweep
for every command instead of excluding `record` by name.
thymikee added a commit that referenced this pull request Oct 4, 2026
Rebase resolution against main@294dc7d (#3199), which rewrote the same hint
formatter this PR restructures. Both changes are semantic and both survive
here:

- #3199 made `record` a preserve-daemon policy and generalized the
  keep-exporting retry to LOCAL timeouts, because a preserved daemon may
  still be exporting (and a stop queued for the device lock is dropped
  before any export starts, hence "may"). This PR's shape decides
  "preserved" from the probe verdict + declared policy instead of the old
  sweep/reset booleans, so the retry branch keys on `!resetDaemon` and
  `action === 'stop'` BEFORE the remote split, with the remote wording
  merely adding "remote ". The old remote-only "is still exporting" claim
  becomes the shared "may still be exporting" one at #3199's second commit.
- The reset branch stays this PR's: a daemon the probe proved unresponsive
  was SIGKILLed, is no longer exporting, and gets the probe-verdict wording
  with no keep-exporting promise — the exact case #3199's note about a
  mid-export reset predates, where the probe now makes the reset rarer.

Route coverage gains a local `record stop` row: the positional rides the
real transport context into the hint, the preserve policy must skip the
probe entirely (connections: 1), and no sweep call may fire — #3199's
runner-survival requirement, which this PR satisfies by removing the sweep
for every command instead of excluding `record` by name.
harrisrobin added a commit to harrisrobin/agent-device that referenced this pull request Oct 5, 2026
…ts budget against record's envelope

- `exportProcessedVideo` uses `Deadline.fromTimeoutMs(...).remainingMs()` instead of a hand-rolled
  `Date.now() + budgetMs`.
- The budget's comment no longer says the client resets the daemon: since callstack#3199 `record` keeps it
  on timeout, and the caller gets "Daemon request timed out" while the daemon finishes the stop.
- `HELPER_EXIT_GRACE_MS` names what it covers: start-up, then verifying the output or cancelling the
  export, and exiting.
- The check that the budget fits `record stop`'s envelope moves from a 90_000 literal in
  `overlay.test.ts` to the root timeout-policy test, against `record`'s resolved envelope.
thymikee pushed a commit that referenced this pull request Oct 5, 2026
…e record request (#3219)

* fix(recording): render the touch overlay at most 30 fps and inside the record request

* refactor(recording): take the overlay deadline from host-kit; check its budget against record's envelope

- `exportProcessedVideo` uses `Deadline.fromTimeoutMs(...).remainingMs()` instead of a hand-rolled
  `Date.now() + budgetMs`.
- The budget's comment no longer says the client resets the daemon: since #3199 `record` keeps it
  on timeout, and the caller gets "Daemon request timed out" while the daemon finishes the stop.
- `HELPER_EXIT_GRACE_MS` names what it covers: start-up, then verifying the output or cancelling the
  export, and exiting.
- The check that the budget fits `record stop`'s envelope moves from a 90_000 literal in
  `overlay.test.ts` to the root timeout-policy test, against `record`'s resolved envelope.

* test(recording): the overlay budget leaves record stop a 20 s reserve inside its envelope

The check let the budget take all but a millisecond of the envelope; the recorder stop, the copies
and the playability checks around the overlay need room beyond it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-for-human Valid work that needs human implementation, judgment, or maintainer merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants