Skip to content

fix(apple-runner): fence prep spawns behind a start-owned admission - #3239

Merged
thymikee merged 2 commits into
mainfrom
t3code/fence-apple-prep-spawns
Oct 6, 2026
Merged

thymikee merged 2 commits into
mainfrom
t3code/fence-apple-prep-spawns

Conversation

@thymikee

@thymikee thymikee commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Problem (#3220)

A runner start builds, launches, and health-checks without holding the device session lock for the whole sequence. A close racing that window kills the in-flight xcodebuild build-for-testing, and the start then answers the kill by retrying — spawning a second concurrent build on the same DerivedData directory while close waits for the lock the retry holds.

Design

Every start carries an admission token scoped to that start (RunnerStartAdmission in runner-artifact.ts, exported through the runner-xctestrun.ts barrel). Per-device state is only a pending-teardown count and the set of in-flight tokens — the device map holds in-flight tokens and nothing else.

  • Teardown owns the fence, not the device. A teardown that can strand a build (non-retained close, daemon-wide abort/stop) closes the in-flight tokens before killing prep processes, holds the fence across its body, and settles it in finally. There is no device-global fence lifetime and no clear-after-teardown step.
  • Refusal is scoped to the retiring start. A closed token refuses its own start's retries with a typed runnerStartRetiredError (reason device_teardown); a fresh open that only queued behind a settled fence is readmitted and proceeds normally. Idle-stop and speculative-release do not fence — the reviewer's requested semantics.
  • The prepare loop owns one token. prepareLocalIosRunner mints a loop admission and hands it to each ensureRunnerSession attempt through options.startAdmission, so the loop's replacement build is the authorized continuation of the same start (and is never reopened by a settled fence).
  • Cancellation owns a scoped kill. A caller leaving on its own deadline marks the start retry-pending (A wait timeout during runner start stops the runner, so the retry pays the start again #2894 keeps the build for the retry); stopRunnerPrepProcessesWithoutLiveOwner stops only builds whose token has no interested waiters and no pending retry. Every caller that can still cancel counts as an interested waiter, including requestId-only daemon owners (reserveRunnerStartOwnerInterest).
  • abortAll/stopAll fence all device starts for their duration instead of awaiting per-device session locks (review P2: mid-phase lock holders no longer block the sweep).

Layout

The machinery lives in runner-artifact.ts beside the prep ledger it gates. A separate static module was tried and empirically fails the eager-closure budget (a new static edge grows façade closures); co-location adds zero modules to any closure, removes both dynamic loaders, and drops the .fallowrc.json host allowlist entry entirely.

Evidence

  • Deterministic control of the incident at the exec seam: runner-session-close-prep-fence.test.ts (close parked mid-prep-kill → retry refused with typed reason, exactly one build ever spawned; queued caller readmitted after settle; negative control that a plain post-teardown open still builds).
  • runner-start-admission.test.ts (token/fence/readmit/loop semantics), runner-start-budget.test.ts (last-waiter kill, deadline ownership, teardown-then-late-cancel with a registered later build, per-device scoping, owner interest), runner-artifact-start-admission.test.ts (spawn-gate enforcement).
  • Live iPhone 17 simulator: cold open → snapshot → close --shutdown → cached reopen → close, all clean; at most one build-for-testing in flight at every sample; daemon and session logs show no admission-token noise or errors.
  • pnpm check:affected --run green (incl. eager-closure budgets, layering, fallow, format, lint).

Closes #3220

@thymikee thymikee added the bug Something isn't working label Oct 5, 2026
@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Size Report

Metric Base Current Diff
Installed (including dependencies) 5.10 MB 5.10 MB +3.0 kB
Package (unpacked) 5.10 MB 5.10 MB +3.0 kB
Package (download) 1.53 MB 1.53 MB +994 B

Startup median (7 runs, lower is better):

Scenario Base Current Diff
CLI --version 27.6 ms 27.5 ms -0.1 ms
CLI --help 84.1 ms 81.9 ms -2.2 ms

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 11 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread packages/platform-apple/src/runner/runner-session.ts Outdated
Comment thread packages/platform-apple/src/runner/runner-start-budget.ts
Comment thread packages/platform-apple/src/runner/runner-start-admission.ts Outdated
Comment thread packages/platform-apple/src/runner/runner-session.ts Outdated
@thymikee

thymikee commented Oct 5, 2026

Copy link
Copy Markdown
Member Author

The fence design in d8644c0 is not ready, and the live evidence for #3220 is missing. Smoke Tests is still running, and the diff touches ensureRunnerSession, the non-retained close and the iOS runner build spawn, so a failure there is not presumed unrelated. I did not run any tests; I judged the regression tests by reading the pre-change route.

The live run in the PR settled the same way on origin/main, so no retry ever reached the spawn fence (runner-session.ts#L759). It proves close during build, not the refused replacement build, so the "no respawn" done-when in #3220 is not shown. Please run one live iOS simulator run: start a cold open/prepare (prewarm with health check), then send a non-retained close during build-for-testing. The request or daemon log should show ios_runner_start_admission_retired with reason device_teardown, a retry rejected with details.runnerStartRetired, exactly one xcodebuild build-for-testing pid, and close finishing with no build child left. If the window cannot be forced on your host, please say so and rely on the deterministic test.

The fence is a device-global registry that refuses every newcomer while a teardown runs. The retry in the prepare loop (runner-lifecycle.ts lines 209 and 428) calls a fresh ensureRunnerSession, which looks like an independent open, so the registry must refuse both. That is what causes the idle-stop, speculative-release and overlapping-teardown windows. Would a smaller design work? The prepare attempt loop (prepareLocalIosRunner/runPrepareAttempt) would own one admission token and pass it into each ensureRunnerSession retry. Teardown would close the tokens of starts in flight, and independent opens would queue on the session lock as before, with no clear-after-teardown step and no per-device fence lifetime. The start, not the device, would own the token, so ensureRunnerSession needs an optional admission input from the retry loop, and the device map holds only in-flight tokens.

Two open inline threads still stand: concurrent callers share one admission before registering (P1), and fence cleared by any teardown. Also raw signal ignores requestId-only owners, idle-stop open gets closed admission, and zero-signal test registers no prep child still apply. The env isolation thread does not apply: the forks pool keeps isolation on, and both tests reassign the env in beforeEach, like sibling suites. Please resolve it.

Before merge, count every caller that can still cancel, including requestId-only daemon owners, as an interested waiter, registered before the first await. Make the fence refuse only the retiring start's own retries, not idle-stop or concurrent-teardown newcomers. Then add the live refused-retry evidence above.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 6 files (changes from recent commits).

Requires human review: Auto-approval blocked because this review re-detected 1 unresolved issue already reported by Cubic.

Re-trigger cubic

Comment thread packages/platform-apple/src/runner/runner-start-budget.ts

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 2 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment thread packages/platform-apple/src/runner/runner-session.ts Outdated
@thymikee
thymikee force-pushed the t3code/fence-apple-prep-spawns branch from 127db94 to 5d74815 Compare October 5, 2026 20:59

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 8 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment thread packages/capture-kit/src/recording/overlay.ts Outdated
Comment thread apple/runner/AgentDeviceRunner/RecordingScripts/recording-overlay.swift Outdated
Comment thread packages/platform-apple/src/runner/__tests__/runner-start-budget.test.ts Outdated
@thymikee
thymikee force-pushed the t3code/fence-apple-prep-spawns branch from 5d74815 to 1efebe8 Compare October 5, 2026 21:21
@thymikee

thymikee commented Oct 6, 2026

Copy link
Copy Markdown
Member Author

Thanks for the update. At 1efebe8 all 19 checks pass and there are no conflicts. The first-round fixes are in: close no longer waits on cold starts that have no session, and the lock-holding owner now registers its resolved request signal. Two defects remain, so this is not ready to merge.

The fence still decides by lock-queue order, not by which start it retires. rerouteSettledRunnerStart (runner-start-admission.ts#L120) lets a retry that queues behind close wake after the teardown clears the fence, and it then spawns a fresh build-for-testing. The prepare and recovery retries in runner-lifecycle.ts (lines 105, 209, 428) each call a fresh ensureRunnerSession, and close enqueues right after its prep sweep, so this is the usual order. Your own test "a start queued behind a close starts fresh once that close has settled" runs the #3220 scenario (a retry after the first child is stopped) and asserts two builds. So a closed device can still get a runner that holds its lease, which is the respawn #3220 is meant to remove. The rule should be: a start retired by a teardown is refused for every retry its own caller issues, and callers that did not issue that start never inherit the fence. The retry owners are prepareLocalIosRunner/runPrepareAttempt, recoverBadCachedRunnerArtifact and the restart path at runner-lifecycle.ts:428. Each should hold one admission token and pass it into ensureRunnerSession. A call without a token takes the device's current or fresh admission. Please rewrite that test so a retry issued after close has enqueued is refused, while an independent open after close settles builds.

A caller with only a requestId is still not counted as a waiter until it holds the lock. reserveRunnerStartOwnerInterest runs inside the lock task (runner-session.ts#L127), and raceRunnerStartAgainstCaller returns early when there is no signal (runner-start-budget.ts#L116). If R1 holds a cold build and R2 (requestId-only) is queued on the same admission, R1's cancel sees no other waiter and closes the admission. R2 then wakes and fails with a cancellation it never issued, where main would have built its own runner. This breaks the #3220 rule that canceling one waiter must preserve work another still needs. The rule should be: every caller whose cancellation can be observed, through resolveRunnerStartupSignal, registers as a waiter synchronously when it captures the admission, before the lock and before any await, and stays registered until its own wait ends. That means one registration site, in the race, so reserveRunnerStartOwnerInterest can go. Please add a control with two requestId-only callers, one on a hanging build and one queued: abort the first, then assert the second is not refused and builds.

Could a smaller design cover both? The retry loop would own one admission token, created in prepareLocalIosRunner and the restart path and passed to every ensureRunnerSession retry. Teardown would close the tokens of starts in flight, and calls without a token would queue on the session lock as before. That would remove clearRunnerStartAdmissionAfterTeardown, rerouteSettledRunnerStart, redirectedTo and the device-fence lifetime, and the first issue would be fixed by construction. Waiter interest would be registered once, at capture. The earlier question about the size of this mechanism has no reply yet. The first change needed is for ensureRunnerSession to accept an optional admission from its retry owners, with the device map holding only in-flight start tokens.

The open inline threads still apply for the vacuous zero-signal assertion in runner-start-budget.test.ts (r4187333221), the queued requestId-only starts (r4187333211) and the shared fence deleted by overlapping teardowns (r4187333242). These can be resolved: r4187333227 (the forks pool isolates workers and beforeEach reassigns both env vars), r4187333249 (idle-stop and speculative newcomers wake after the clear and re-route), r4188785937 and r4188785949 (rebase artifacts, the diff since 14e1bd7 touches no capture-kit or Swift file), r4187333233 (the lock-holding owner now registers its signal, and the queued remainder is the second issue above), r4188422578 (an uncounted fire-and-forget prewarm is intended), r4188482183 (the per-device settle now covers only registered sessions) and r4188785962 (stray asterisk gone).

I read the code and your tests but did not run them. I did not list every daemon surface that reaches ensureRunnerSession with a requestId and no signal, and I checked the pure-move claim by line-set comparison. The live close-during-build run never reached the fence on either tree, so the absence of a respawn rests on the PR body. Before merge, the fence must follow start identity, so a retry of a start that close retired is refused even after close has queued, and every requestId-only caller must count as a waiter before its first await.

@thymikee
thymikee force-pushed the t3code/fence-apple-prep-spawns branch 2 times, most recently from 4ec271a to 1ee5cf6 Compare October 6, 2026 09:42
…3220)

A runner start builds, launches, and health-checks without holding the
device session lock, so a teardown racing that window could kill the
build and still let the start retry into a second concurrent
`xcodebuild build-for-testing` on the same DerivedData directory.

Every start now carries an admission token. Teardown of an in-flight
start closes the token before killing prep processes, so later calls
see an explicit retired verdict instead of retrying into the race; an
open that only queued behind a settled fence is readmitted and
proceeds normally. The prepare loop owns one token across its retry so
the replacement build stays authorized, and a caller leaving on its own
deadline marks the start retry-pending so teardown does not stop a
build another owner still waits on. abortAll/stopAll fence all device
starts for their duration instead of awaiting per-device session locks.

The machinery lives in runner-artifact.ts beside the prep ledger it
gates; no new static module edges, no host allowlist.
@thymikee
thymikee force-pushed the t3code/fence-apple-prep-spawns branch from 1ee5cf6 to 7ca8147 Compare October 6, 2026 10:00
@thymikee

thymikee commented Oct 6, 2026

Copy link
Copy Markdown
Member Author

Rewritten at 7ca8147 on the smaller design you sketched — both rules are now structural, and the mechanism is smaller than the first round's.

The shape. The retry loop owns one admission token: prepareLocalIosRunner mints it (openRunnerStartLoopAdmission) and passes it to every ensureRunnerSession attempt through options.startAdmission; recoverBadCachedRunnerArtifact inherits it inside the options it already receives (pinned by the prepare-recovery assertion in runner-command-retry.test.ts that the rebuild call carries the same token as the first). A call without a token takes a fresh self-minted token and otherwise queues on the session lock exactly as before. Teardown closes the tokens of starts in flight before its prep sweep, holds while it runs, and settles in finally. clearRunnerStartAdmissionAfterTeardown, rerouteSettledRunnerStart, redirectedTo and the device-fence lifetime are gone — the device map holds only pendingTeardowns and the in-flight token set.

Fence follows start identity. A token minted by ensureRunnerSession itself may be readmitted only when it closed behind a fence that has since settled and nothing ever ran under it — so the queued independent open proceeds (your requested test: a start queued behind a close starts fresh once that close has settled / a start after a settled teardown opens fresh and builds), while a retry of a start close retired is refused even after close has queued (close during a cold build leaves the start no way to spawn a replacement build, asserted with the typed reason). A loop token is never readmitted — the loop's health retry is the replacement build the fence exists for (a loop-supplied token is never reopened by a settled fence). The restart path at runner-lifecycle.ts keeps fresh tokens per ensureRunnerSession call: it has no build-retry loop after the dead session is invalidated, and its restart attempt registers synchronously like any open.

Waiter interest. Every caller whose cancellation can be observed counts before its first await: the token itself registers synchronously at capture, raceRunnerStartAgainstCaller registers its signal synchronously, and the requestId-only owner registers through reserveRunnerStartOwnerInterest with the resolved startup signal. I did try collapsing that to the single registration site in the race you proposed — hoisting the owner registration to synchronous capture requires a static import of runner-start-budget.ts from runner-session.ts, which the eager-closure budget gate rejects (measured twice, and the alternative of a new module failed it too, see below). The lock task is not a correctness gap: the queued requestId-only case you described is pinned as requested — a queued requestId-only start survives the starting caller disconnecting in runner-session-close-prep-fence.test.ts: R1 on a hanging build with requestId and no signal, R2 requestId-only queued behind it; aborting R1 stops R1's build and R2 still builds (two spawns, second never refused).

Size. The mechanism is ~280 added lines inside runner-artifact.ts (tokens, device state, gates, waiter counting, documentation) plus ~70 in the race, with no new module and no new static edge: the separate runner-start-admission.ts is gone — a new static module edge grows four to seven façade closures and fails scripts/__tests__/eager-closure-budgets.test.ts, so co-locating beside the prep ledger it gates is what let both dynamic loaders and the .fallowrc.json host allowlist entry go away. Production diff: +719/−199 across 8 files, of which runner-adoption.ts (+106) is a pure move that keeps runner-session.ts under the 1k-line rule.

Live evidence. The iPhone 17 run did not force the fence window (the cold build landed while open still held the only caller, so close never raced a second start) — as you allowed, saying so plainly: the refused-retry evidence rests on the deterministic control at the exec seam. The live run did verify normal open/close/reopen on a real simulator with at most one build-for-testing in flight at every sample and no admission noise in the daemon or session logs. If you can force the window on your host (two callers, close enqueued mid-build), the diagnostic to expect is a canceled-request error with details.runnerStartRetired: true and runnerStartRetirementReason: "device_teardown" on the retry, and one build pid total.

All inline threads are answered and resolved; the vacuous zero-signal assertion now registers the later start's build (4951) and asserts it stays running; env overrides are restored in afterEach; idle-stop and speculative release no longer fence at all. pnpm check:affected --run green at 7ca8147.

… file (#3220)

The fence work grew runner-command-retry.test.ts past its merge-base length,
which the test-file size ratchet refuses. The two bad-cache recovery tests
mirror the prepare artifact decision, so they move to
runner-lifecycle-prepare-artifact.test.ts, which already carries that harness.
@thymikee

thymikee commented Oct 6, 2026

Copy link
Copy Markdown
Member Author

At 38527ee this is ready for human review. Both findings from the earlier review (#3239 (comment)) are fixed: each retry owner now holds its own start admission and passes it through options.startAdmission, and the separate admission module and the reroute and fence-lifetime code are gone. All 19 checks pass at this commit, and I know of no conflicts. Nothing else blocks merge.

Not blocking, and you can take or leave these: (1) when a prepare caller's own deadline fires, the loop's finally in runner-lifecycle.ts#L96 drops the admission from the in-flight set while the detached build still runs, so a cancel from another request on the same device can stop the build the #2894 retry meant to join; the rule is that an admission leaves the in-flight set only when no start holding it is still running, so could the loop finish it after the last ensureRunnerSession promise settles? (2) The production-route control in runner-session-close-prep-fence.test.ts#L187 issues its retry before close enqueues, so the refusal comes from pendingTeardowns and not from start identity; a test that drives prepareLocalIosRunner through a real close and asserts one build plus a runnerStartRetired refusal would pin the case from the earlier review. (3) Each call now creates its own admission and the two owner paths are mutually exclusive, so an admission never holds more than one waiter; could the waiter set collapse to one owner signal, with the four two-waiter tests rewritten as two admissions on one device?

On the simplicity question: the one-admission-per-owner design answers it, and the collapse in (3) is the only simplification left.

I ran no tests; I judged the regression coverage by reading the pre-change route. Live refused-retry evidence is absent, and the deterministic exec-seam control stands in for it, as the earlier review allowed since the window could not be forced on an iPhone 17 simulator. I did not measure the eager-closure budget claim, which no longer matters with per-call admissions. Note (1) assumes the prepare request's signal is deregistered once its caller deadline answers, and I did not trace the daemon request scope to confirm that.

On the earlier review threads, these still apply: none at P1 or P2, and none at lower priority. These are fixed at this head and can be resolved: #3239 (comment), #3239 (comment), #3239 (comment), #3239 (comment), #3239 (comment), #3239 (comment), #3239 (comment), #3239 (comment). These do not apply: #3239 (comment) (a prewarm with no signal or requestId has no caller who could cancel it, same as main's #3193 behavior), #3239 (comment) and #3239 (comment) (rebase artifacts; the diff against 37fa3e7 has no capture-kit or Swift file).

@thymikee thymikee added the ready-for-human Valid work that needs human implementation, judgment, or maintainer merge label Oct 6, 2026
@thymikee
thymikee merged commit 1163b7a into main Oct 6, 2026
19 checks passed
@thymikee
thymikee deleted the t3code/fence-apple-prep-spawns branch October 6, 2026 11:51
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-06 11:51 UTC

thymikee added a commit to okwasniewski/agent-device that referenced this pull request Oct 6, 2026
* origin/main: (77 commits)
  fix(apple-runner): fence prep spawns behind a start-owned admission (callstack#3239)
  0.21.22
  test(apple): own the simctl settings plan tests in simctl-settings.test.ts (callstack#3244)
  fix(limrun): report the session device id in iOS settings refusals (callstack#3243)
  0.21.21
  feat(remote): add a host-allocated macos-app lease backend (callstack#3236)
  test(android): bound the screenshot write wait by wall time, not event-loop turns (callstack#3250)
  feat(recording): cap the touch overlay frame rate at the caller's --fps (callstack#3241)
  fix(ad-script): let .ad scripts carry scroll --until and wait capture flags (callstack#3197) (callstack#3234)
  feat(provider-webdriver): keyboard enter, dismiss, and status over WebDriver (callstack#3233)
  feat(selectors): match role= against snapshot kind with a node-scoped alias window (callstack#3232)
  fix(provider-webdriver): read field values, placeholders, secure fields, and checked state from page source (callstack#3231)
  feat(replay): accept --test-ime on test and replay so flow-owned Android opens opt into the test IME (callstack#3235)
  refactor(daemon): route daemon-level diagnostics through one scope helper (callstack#3242)
  docs(adr): correct ADR 0031 pointer event delivery evidence (callstack#3245)
  fix(ios): stop a tap's post-gesture lookup from recording an XCTest failure (callstack#3060) (callstack#3237)
  fix(recording): render the touch overlay at most 30 fps and inside the record request (callstack#3219)
  fix(daemon): keep an idle daemon alive only for retained leases (callstack#3227)
  fix(provider-webdriver): send an empty JSON object on bodyless POSTs (callstack#3230)
  Feat/maestro repeat while (callstack#3214)
  ...
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working ready-for-human Valid work that needs human implementation, judgment, or maintainer merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(apple-runner): fence prep spawns after teardown or last-waiter cancellation

1 participant