Skip to content

feat(interaction): readiness wait for replay and the Node client on a shared observation engine - #3072

Merged
thymikee merged 23 commits into
mainfrom
proto/reliability-contract
Oct 1, 2026
Merged

thymikee merged 23 commits into
mainfrom
proto/reliability-contract

Conversation

@thymikee

@thymikee thymikee commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Summary

Part of #3069. Merge after #3070 and #3071. The post-gesture and scroll migrations moved to a stacked PR, so this PR is the engine plus its one consumer, the readiness wait, and can be dropped whole if the #3077 soak says so.

press, click, and longpress take an operator-only readinessTimeoutMs. With it, a target that is not on screen yet is looked for every 200 ms, up to the value (capped at 2 s), then the requested interaction runs once. A covered, off-screen, or ambiguous target still fails at once; only an unreadable capture is waited out (and if every capture is unreadable, that typed error surfaces, not selector_not_found); a sparse capture ends the wait with reason: capture_sparse. Which commands take the budget is one registry trait, targetReadiness: 'budgeted'; a command without it refuses the key; the envelope widens by the capped budget only for those commands; replay reads the trait.

Replay: recorded steps carry targetEvidence and are verified before dispatch, so the pre-dispatch gate now polls under the same schedule, then hands the remaining budget to the dispatch (one 2 s budget per step, not two). A gate that waited emits the same interaction_target_readiness diagnostic as the daemon, and a selector-miss divergence carries error.details.readiness.

await client.interactions.press({ selector: 'label="Continue"', readinessTimeoutMs: 2_000 });

The wait runs on observeUntil in capture-kit: one cadence, one deadline rule (the budget bounds when a poll starts), a declared per-capture deadline mode (captureDeadline: 'cancel' | 'none'; a capture that cannot honor the signal is judged late, never stalled), ride-out with lastError, and a per-poll timeline with a failed outcome. The verdict stays with the caller. ADR 0011 gains the targetReadiness cell. resolution.ts (1411 lines on main) is split by domain question into ten modules, the largest 323 (refactor(move), bodies unchanged). captureDeadline defaults to 'none'; readiness opts into 'cancel'; the schedule has one owner in the selectors policy; the replay gate polls under it and hands the remainder to the dispatch.

Over the 1,000-line budget by maintainer decision; the split is most of the churn. Docs: replay-e2e.md, client-api.md.

Validation

Tested commit d476097d60 (rebased on main after #3070): all 19 CI checks pass; local pnpm check:affected --run: 1591 of 1592 files green, scripts/fuzz/harness.test.ts (untouched here) timed out once and passes alone.

  • Engine: 14 tests (deadline modes, ride-out lastError, minPolls past the budget, failed outcome). Readiness: all-unreadable surfaces the typed reason; scenarios for target on the third capture, never appears, first capture hits, option absent, sparse.
  • Replay: target renders on the second capture → dispatched with its guard and one diagnostic; renders at 1.6 s → dispatch gets 400 ms; never → selector-miss with readiness; fill gets no wait. .ad round-trip of annotations at target-annotation-serde.ts:169/249.
  • Registry: the trait set equals the budgeted targetReadiness cells; fill refuses readinessTimeoutMs; 999_000 widens by 2000.
  • Live: 46 interleaved A/B runs on a private iPhone 17 Pro simulator under host load, 46/46 both ways, the wait never engaged; cost parity on the hit path only. Keep-or-drop follows the nightly soak (Readiness wait: nightly soak and keep-or-drop rule #3077). The annotated-replay live run with an engaged wait is still open.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Size Report

Metric Base Current Diff
Installed (including dependencies) 4.90 MB 4.91 MB +9.1 kB
Package (unpacked) 4.90 MB 4.91 MB +9.1 kB
Package (download) 1.47 MB 1.47 MB +3.0 kB

Startup median (7 runs, lower is better):

Scenario Base Current Diff
CLI --version 18.3 ms 20.1 ms +1.8 ms
CLI --help 56.4 ms 54.9 ms -1.4 ms

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-01 06:14 UTC

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

6 issues found across 48 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="packages/capture-kit/src/post-gesture-stability.ts">

<violation number="1" location="packages/capture-kit/src/post-gesture-stability.ts:158">
P2: This adapter discards `observeUntil`’s deadline signal, so a snapshot capture that runs past the 1.5s/3.5s budget cannot be cancelled and joined by the shared engine. Pass the observer signal through `PostGestureStabilityHooks.capture` and the snapshot capture path; otherwise the loop can remain in flight with the pending stabilization record uncleared until the backend eventually resolves.</violation>

<violation number="2" location="packages/capture-kit/src/post-gesture-stability.ts:206">
P2: The migration changes the quiet-window verdict on the stakes path the old loop never hit: `observeUntil` now aborts an in-flight capture at the per-poll deadline (each subsequent capture is capped at `Math.max(remainingMs, intervalMs)`, e.g. 200 ms once the 1.5 s budget is nearly spent) and reports `stalled`, which this line converts into a hard throw. The previous loop passed no deadline into the capture, waited for it, and still returned the `unsettled`/value outcome (with the `post_gesture_snapshot_stabilization_timeout` warning) after the deadline passed. A capture slower than its remaining budget slice therefore escalates a previously graceful `unsettled` settlement into a thrown error, contradicting the PR's "without changing their verdicts" claim and skipping the timeout diagnostic.</violation>
</file>

<file name="packages/contracts/src/interaction-guarantees.ts">

<violation number="1" location="packages/contracts/src/interaction-guarantees.ts:167">
P2: The shared waiver misclassifies the runtime tree paths as having no always-on outcome observation. Runtime dispatch already performs post-action Android escape checks and iOS failure corroboration, so mark this cell as runtime or narrow `outcomeObservation` to successful-return verification.</violation>

<violation number="2" location="packages/contracts/src/interaction-guarantees.ts:545">
P2: The coordinate path is not inapplicable to outcome observation; it uses the shared post-dispatch observation lifecycle even without an element. Classify this cell as runtime, or redefine the guarantee explicitly around element identity rather than action outcome.</violation>
</file>

<file name="src/daemon/scroll-movement.ts">

<violation number="1" location="src/daemon/scroll-movement.ts:274">
P1: This callback drops `observeUntil`’s deadline signal, so a hung snapshot can keep `observeScrollMovement` awaiting indefinitely instead of returning `surface-unsettled` after 1.5 seconds. Thread the signal through `readOneCapture` and the scroll capture input to `captureSnapshot`.</violation>
</file>

<file name="src/commands/interaction/runtime/interaction-snapshot-capture.ts">

<violation number="1" location="src/commands/interaction/runtime/interaction-snapshot-capture.ts:38">
P1: The readiness poll's `forceFresh` flag is dropped by the daemon interaction backend, so later polls can keep reading the cached tree instead of observing a target that appears during the wait. Thread this flag through the interaction capture options and bound capture before relying on it for readiness retries.</violation>
</file>

Tip: instead of fixing issues one by one fix them all with cubic

Re-trigger cubic

Comment thread packages/capture-kit/src/observe-until.ts Outdated
Comment thread src/daemon/scroll-movement.ts Outdated
await sleep(params.pollMs ?? MOVEMENT_POLL_MS);
}
const observed = await observeUntil<CaptureReading, SurfaceJudgement>({
capture: () => readOneCapture(baseline, params.capture),

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: This callback drops observeUntil’s deadline signal, so a hung snapshot can keep observeScrollMovement awaiting indefinitely instead of returning surface-unsettled after 1.5 seconds. Thread the signal through readOneCapture and the scroll capture input to captureSnapshot.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/daemon/scroll-movement.ts, line 274:

<comment>This callback drops `observeUntil`’s deadline signal, so a hung snapshot can keep `observeScrollMovement` awaiting indefinitely instead of returning `surface-unsettled` after 1.5 seconds. Thread the signal through `readOneCapture` and the scroll capture input to `captureSnapshot`.</comment>

<file context>
@@ -257,30 +265,32 @@ async function pollForSurfaceVerdict(
-    await sleep(params.pollMs ?? MOVEMENT_POLL_MS);
-  }
+  const observed = await observeUntil<CaptureReading, SurfaceJudgement>({
+    capture: () => readOneCapture(baseline, params.capture),
+    schedule: {
+      intervalMs: params.pollMs ?? SCROLL_MOVEMENT_SCHEDULE.intervalMs,
</file context>
Fix with cubic

const result = await runtime.backend.captureSnapshot(toBackendContext(runtime, options), {
interactiveOnly,
includeRects: true,
...(forceFresh ? { forceFresh: true } : {}),

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: The readiness poll's forceFresh flag is dropped by the daemon interaction backend, so later polls can keep reading the cached tree instead of observing a target that appears during the wait. Thread this flag through the interaction capture options and bound capture before relying on it for readiness retries.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/commands/interaction/runtime/interaction-snapshot-capture.ts, line 38:

<comment>The readiness poll's `forceFresh` flag is dropped by the daemon interaction backend, so later polls can keep reading the cached tree instead of observing a target that appears during the wait. Thread this flag through the interaction capture options and bound capture before relying on it for readiness retries.</comment>

<file context>
@@ -0,0 +1,50 @@
+  const result = await runtime.backend.captureSnapshot(toBackendContext(runtime, options), {
+    interactiveOnly,
+    includeRects: true,
+    ...(forceFresh ? { forceFresh: true } : {}),
+  });
+  const snapshot =
</file context>
Fix with cubic


const observed = await observeUntil<T, PostGestureStabilityOutcome<T>>({
...(params.initial !== undefined ? { initial: params.initial } : {}),
capture: () => hooks.capture(),

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This adapter discards observeUntil’s deadline signal, so a snapshot capture that runs past the 1.5s/3.5s budget cannot be cancelled and joined by the shared engine. Pass the observer signal through PostGestureStabilityHooks.capture and the snapshot capture path; otherwise the loop can remain in flight with the pending stabilization record uncleared until the backend eventually resolves.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/capture-kit/src/post-gesture-stability.ts, line 158:

<comment>This adapter discards `observeUntil`’s deadline signal, so a snapshot capture that runs past the 1.5s/3.5s budget cannot be cancelled and joined by the shared engine. Pass the observer signal through `PostGestureStabilityHooks.capture` and the snapshot capture path; otherwise the loop can remain in flight with the pending stabilization record uncleared until the backend eventually resolves.</comment>

<file context>
@@ -129,24 +138,34 @@ export async function runPostGestureStabilityLoop<T, S extends readonly unknown[
+
+  const observed = await observeUntil<T, PostGestureStabilityOutcome<T>>({
+    ...(params.initial !== undefined ? { initial: params.initial } : {}),
+    capture: () => hooks.capture(),
+    schedule: POST_GESTURE_STABILITY_SCHEDULE,
+    verdict: (latest, previousValue) => {
</file context>
Fix with cubic

// after a tap is an ambiguous outcome, per docs/agents/selector-capture.md) is scoped to
// navigation-sensitive Android actions and target-authored gestures, not the tap-shaped runtime
// paths; --settle/--verify are the opt-in ways to observe one.
const TAP_OUTCOME_NOT_OBSERVED_GAP: GuaranteeEnforcement = {

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: The shared waiver misclassifies the runtime tree paths as having no always-on outcome observation. Runtime dispatch already performs post-action Android escape checks and iOS failure corroboration, so mark this cell as runtime or narrow outcomeObservation to successful-return verification.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/contracts/src/interaction-guarantees.ts, line 167:

<comment>The shared waiver misclassifies the runtime tree paths as having no always-on outcome observation. Runtime dispatch already performs post-action Android escape checks and iOS failure corroboration, so mark this cell as runtime or narrow `outcomeObservation` to successful-return verification.</comment>

<file context>
@@ -149,6 +159,42 @@ const SHARED_RESPONSE_CONSTRUCTION: GuaranteeEnforcement = {
+// after a tap is an ambiguous outcome, per docs/agents/selector-capture.md) is scoped to
+// navigation-sensitive Android actions and target-authored gestures, not the tap-shaped runtime
+// paths; --settle/--verify are the opt-in ways to observe one.
+const TAP_OUTCOME_NOT_OBSERVED_GAP: GuaranteeEnforcement = {
+  kind: 'waived',
+  reason:
</file context>
Fix with cubic

Comment thread src/commands/common-input-fields.ts
Comment thread packages/capture-kit/src/observe-until.test.ts Outdated
Comment thread test/integration/provider-scenarios/scroll-movement-observation.test.ts Outdated
Comment thread packages/capture-kit/src/observe-until.ts Outdated
@thymikee

Copy link
Copy Markdown
Member Author

Thanks for the PR. At c19112a I found four defects that need fixing before merge, and I have a question on scope at the end.

The post-gesture adapter passes capture: () => hooks.capture() (line 158), so the capture ignores the engine's deadline signal. When a capture outlives the remaining budget, which is normal for 300-800 ms iOS captures near the end of the 1.5 s budget, captureWithin returns {kind:'stalled', error: undefined} and line 206 runs throw undefined. Before this change the loop judged the late capture and returned it with postGestureOutcome: 'unsettled' and the post_gesture_snapshot_stabilization_timeout warning. Now, on a surface that never settles (spinner, animation), the next snapshot or interaction fails with an unknown error. The rule: a capture that does not honor the signal must never be classified as stalled, and a stalled or expired end must give the same outcome as the old deadline expiry. Please make this hold for every observeUntil caller (readiness, post-gesture, scroll), for example with a per-capture deadline mode of none versus cancel. Please add a test where a capture advances the clock past the remaining budget and the result is unsettled, not a throw.

Scroll has the same problem. readOneCapture is wrapped without the signal (line 274), and the first capture is now bounded by the budget. A slow post-scroll capture that showed movement ends as stalled, and line 290 maps it to budgetExpiredVerdict. The result is a blind surface-unsettled and a scroll_movement_budget_expired warning, where it used to claim movement. This also contradicts the statement that scroll verdicts are unchanged. The same rule applies: judge a late capture from a signal-blind capture, and judge the value a stalled capture returned before falling back to the expired verdict. Please add a scenario test with a slow capture.

In observe-until.ts the ride-out branch returns without setting state.lastError, and nothing else assigns it. If readiness is on and every capture is ridden out as unreadable content, the loop ends expired with no last and no lastError. Then selector-readiness.ts:227 cannot fire, and readinessExhaustedFailure reports selector_not_found against an empty node list. Every replay step and Node caller sets readinessTimeoutMs, so users would be told to fix a selector that is correct. Please record poll.error in that branch. Please add an observe-until test and a readiness test where every capture throws an unreadable-content error, and assert that the typed unreadable reason surfaces.

I think the replay default is applied too late (session-replay-action-runtime.ts:148). Recorded press, click, and longpress steps carry targetEvidence, and the verify step in verify-dispatch.ts (lines 127-160) captures once and classifies the target before dispatchStep. A target that has not rendered yet gives a selector-miss REPLAY_DIVERGENCE, and the step is never dispatched. So the main goal of the linked issues, replaying a step recorded against a loading screen, would only hold for unannotated steps. The #3077 soak would also count no engaged waits for recorded scripts. The rule: the readiness budget must cover every gate that can refuse a press, click, or longpress step for a missing target. Please either poll the pre-dispatch target classification under the same observeUntil schedule, or defer the selector-miss result to the dispatched wait. Please add a step-loop test with targetEvidence where the target appears on the second capture. I read the recorder path but not the .ad serializer round-trip, so please confirm that recorded scripts carry these annotations.

Not blocking: the docs (replay-e2e.md:66) point replay users at error.details.readiness and selector_not_found, but replay wraps failures as REPLAY_DIVERGENCE and SAFE_CAUSE_DETAIL_KEYS omits readiness, and client-api.md says every failure after the wait carries readiness, which covered, ambiguous, and non-ridden-out failures do not; the timeout envelope at timeout-policy.ts:86 adds the raw readinessTimeoutMs although the daemon caps the wait at the 2 s promoted-target max, and it applies to commands that ignore the field; isRideOutCaptureError in observation.ts:21 is called only by a test while production uses isUnreadableCaptureContentError, so it adds a subpath export and a second ride-out rule; and baselineEndsInDirection at scroll-movement.ts:268 now runs before every scroll and emits scroll_movement_edge_rest_required even when the surface never changes, where it used to run lazily. You can take or leave all of these.

Would it help to split this into two PRs? The first would hold observeUntil and the readiness wait, which is the one new consumer and is gated by #3077. The second would migrate post-gesture stability and scroll movement once the engine can express "no per-capture deadline, judge late captures", with slow-capture scenario tests. The first two defects come from that second half, and the split would bring the first PR under about 700 net production lines. If #3077 decides to drop the wait, the first PR goes away whole and leaves no migrated loops behind. For that, observeUntil needs a declared per-capture deadline mode, and the replay pre-dispatch gate needs to accept the readiness schedule.

I ran no tests and no live device runs, so the first two findings come from reading the code. The reported live A/B runs covered only the hit path. Before merge I would like three live runs. First, an iOS simulator replay with --debug of a recorded, targetEvidence-annotated .ad script whose press target renders after the first capture. It should show an interaction_target_readiness diagnostic with polls > 1 and the step passing, not a pre-dispatch selector-miss REPLAY_DIVERGENCE. Second, a gesture followed by a snapshot on a screen that never settles. It should show the post_gesture_snapshot_stabilization_timeout warning and a returned snapshot, not a thrown error. Third, a scroll on a large tree where the post-scroll capture takes longer than 1.5 s. It should report movement moved, not surface-unsettled.

On CI, the checks do not clear the change yet. Both CodeQL Analyze logs end at "Uploading results" with no error, and the java-kotlin job shows only a Maven fetch warning. This PR changes no Kotlin or Java, so those look like upload problems. Coverage and Compatibility & Provenance have no failed-step log that I could read, and Coverage runs the suites this change touches, so it cannot be cleared without one. Smoke Tests is still running and exercises the interaction and scroll routes this PR rewrites.

Next, please fix the first two defects by keeping the judge-late-capture behavior in the post-gesture and scroll adapters. Then record lastError and apply readiness at the replay pre-dispatch gate. Then share the live miss-path run through replay of an annotated script.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

5 issues found across 30 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="src/commands/interaction/runtime/resolution-disclosure.test.ts">

<violation number="1" location="src/commands/interaction/runtime/resolution-disclosure.test.ts:9">
P3: This test re-asserts the two disclosures that `tryResolveRefNode discloses exact for a resolved ref and label-fallback for label recovery` in ref-target-resolution.test.ts (lines 80-101) already verifies end-to-end, and that test was added in the same commit. `buildRefResolution` is only reachable through `tryResolveRefNode`/`adoptPreresolvedRefTarget`, so this file adds no coverage; the fields it uniquely contributes — the `ref` and `node` passthrough — are never asserted. Drop the file, or give it distinct value by asserting `buildRefResolution('e1', node, kind).ref` and that the returned node is the passed node.</violation>
</file>

<file name="src/commands/interaction/runtime/native-ref-interaction.ts">

<violation number="1" location="src/commands/interaction/runtime/native-ref-interaction.ts:64">
P1: `preflightNativeRefInteraction` validates `@ref` against the latest observation instead of the authorized ref frame. After a read-only capture advances `session.snapshot`, native fast paths can guard and report a different node while the backend acts on the authorized ref; use the same frame selection as normal ref resolution.</violation>

<violation number="2" location="src/commands/interaction/runtime/native-ref-interaction.ts:131">
P2: Native dispatch reports `resolution.kind: 'exact'` even when `fallbackLabel` recovered the target. Preserve the preflight resolution so clients receive the required `label-fallback` disclosure, matching the runtime path.</violation>
</file>

<file name="src/commands/interaction/runtime/target-visibility-stages.ts">

<violation number="1" location="src/commands/interaction/runtime/target-visibility-stages.ts:163">
P3: The generated recovery command does not escape apostrophes in selectors. Copying a hint for a selector such as `label=O'Reilly` produces invalid shell quoting, so the documented `scroll --until` recovery cannot run; escape single quotes before wrapping the selector.</violation>
</file>

<file name="packages/command-registry/src/types.ts">

<violation number="1" location="packages/command-registry/src/types.ts:83">
P3: Adding `targetReadiness` to `CommandTimeoutPolicy` makes it legal to declare the trait inside a descriptor's own `timeoutPolicy`, which the doc comment right above explicitly forbids ("Descriptors declare it as `targetReadiness`, never on their `timeoutPolicy`"). If a future descriptor did attach it there, `registry.ts`'s merge (`descriptor.targetReadiness ? {...} : descriptor.timeoutPolicy`) stores the policy object as-is, so `readinessBudgetMs` would honor the trait while `commandAcceptsReadinessBudget` (which reads `descriptor.targetReadiness`) would not — envelope widening and replay default-budget gating diverge silently. The sealing invariant is currently prose-only.</violation>
</file>

Requires human review: Auto-approval blocked because this review re-detected 1 unresolved issue already reported by Cubic.
Tip: instead of fixing issues one by one fix them all with cubic

Re-trigger cubic

preAction?: SurfaceScopedNodes;
}> {
const session = await runtime.sessions.get(options.session ?? 'default');
const storedSnapshot = session?.snapshot;

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: preflightNativeRefInteraction validates @ref against the latest observation instead of the authorized ref frame. After a read-only capture advances session.snapshot, native fast paths can guard and report a different node while the backend acts on the authorized ref; use the same frame selection as normal ref resolution.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/commands/interaction/runtime/native-ref-interaction.ts, line 64:

<comment>`preflightNativeRefInteraction` validates `@ref` against the latest observation instead of the authorized ref frame. After a read-only capture advances `session.snapshot`, native fast paths can guard and report a different node while the backend acts on the authorized ref; use the same frame selection as normal ref resolution.</comment>

<file context>
@@ -0,0 +1,135 @@
+  preAction?: SurfaceScopedNodes;
+}> {
+  const session = await runtime.sessions.get(options.session ?? 'default');
+  const storedSnapshot = session?.snapshot;
+  const nodes = storedSnapshot?.nodes;
+  if (!storedSnapshot || !nodes || normalizeRef(target.ref) === null) return {};
</file context>
Fix with cubic

return {
kind: 'ref',
target: { kind: 'ref', ref: target.ref },
resolution: EXACT_REF_RESOLUTION,

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Native dispatch reports resolution.kind: 'exact' even when fallbackLabel recovered the target. Preserve the preflight resolution so clients receive the required label-fallback disclosure, matching the runtime path.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/commands/interaction/runtime/native-ref-interaction.ts, line 131:

<comment>Native dispatch reports `resolution.kind: 'exact'` even when `fallbackLabel` recovered the target. Preserve the preflight resolution so clients receive the required `label-fallback` disclosure, matching the runtime path.</comment>

<file context>
@@ -0,0 +1,135 @@
+  return {
+    kind: 'ref',
+    target: { kind: 'ref', ref: target.ref },
+    resolution: EXACT_REF_RESOLUTION,
+    ...preflight,
+    ...(formattedBackendResult ? { backendResult: formattedBackendResult } : {}),
</file context>
Fix with cubic

test('buildRefResolution is the shared exact and label-fallback disclosure constructor', () => {
const node = selectorSnapshot().nodes[0]!;

assert.deepEqual(buildRefResolution('e1', node, 'exact').resolution, {

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: This test re-asserts the two disclosures that tryResolveRefNode discloses exact for a resolved ref and label-fallback for label recovery in ref-target-resolution.test.ts (lines 80-101) already verifies end-to-end, and that test was added in the same commit. buildRefResolution is only reachable through tryResolveRefNode/adoptPreresolvedRefTarget, so this file adds no coverage; the fields it uniquely contributes — the ref and node passthrough — are never asserted. Drop the file, or give it distinct value by asserting buildRefResolution('e1', node, kind).ref and that the returned node is the passed node.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/commands/interaction/runtime/resolution-disclosure.test.ts, line 9:

<comment>This test re-asserts the two disclosures that `tryResolveRefNode discloses exact for a resolved ref and label-fallback for label recovery` in ref-target-resolution.test.ts (lines 80-101) already verifies end-to-end, and that test was added in the same commit. `buildRefResolution` is only reachable through `tryResolveRefNode`/`adoptPreresolvedRefTarget`, so this file adds no coverage; the fields it uniquely contributes — the `ref` and `node` passthrough — are never asserted. Drop the file, or give it distinct value by asserting `buildRefResolution('e1', node, kind).ref` and that the returned node is the passed node.</comment>

<file context>
@@ -0,0 +1,19 @@
+test('buildRefResolution is the shared exact and label-fallback disclosure constructor', () => {
+  const node = selectorSnapshot().nodes[0]!;
+
+  assert.deepEqual(buildRefResolution('e1', node, 'exact').resolution, {
+    source: 'ref',
+    phase: 'pre-action',
</file context>
Fix with cubic

): string {
if (!direction) return 'Scroll toward it,';
if (!selector) return `Scroll ${direction} toward it,`;
return `Run scroll ${direction} --until '${selector}' to bring it on screen,`;

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The generated recovery command does not escape apostrophes in selectors. Copying a hint for a selector such as label=O'Reilly produces invalid shell quoting, so the documented scroll --until recovery cannot run; escape single quotes before wrapping the selector.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/commands/interaction/runtime/target-visibility-stages.ts, line 163:

<comment>The generated recovery command does not escape apostrophes in selectors. Copying a hint for a selector such as `label=O'Reilly` produces invalid shell quoting, so the documented `scroll --until` recovery cannot run; escape single quotes before wrapping the selector.</comment>

<file context>
@@ -0,0 +1,219 @@
+): string {
+  if (!direction) return 'Scroll toward it,';
+  if (!selector) return `Scroll ${direction} toward it,`;
+  return `Run scroll ${direction} --until '${selector}' to bring it on screen,`;
+}
+
</file context>
Suggested change
return `Run scroll ${direction} --until '${selector}' to bring it on screen,`;
return `Run scroll ${direction} --until '${selector.replaceAll("'", "'\\''")}' to bring it on screen,`;
Fix with cubic

* so the envelope can widen by the readiness budget. Descriptors declare it as `targetReadiness`,
* never on their `timeoutPolicy`.
*/
targetReadiness?: CommandTargetReadiness;

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: Adding targetReadiness to CommandTimeoutPolicy makes it legal to declare the trait inside a descriptor's own timeoutPolicy, which the doc comment right above explicitly forbids ("Descriptors declare it as targetReadiness, never on their timeoutPolicy"). If a future descriptor did attach it there, registry.ts's merge (descriptor.targetReadiness ? {...} : descriptor.timeoutPolicy) stores the policy object as-is, so readinessBudgetMs would honor the trait while commandAcceptsReadinessBudget (which reads descriptor.targetReadiness) would not — envelope widening and replay default-budget gating diverge silently. The sealing invariant is currently prose-only.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/command-registry/src/types.ts, line 83:

<comment>Adding `targetReadiness` to `CommandTimeoutPolicy` makes it legal to declare the trait inside a descriptor's own `timeoutPolicy`, which the doc comment right above explicitly forbids ("Descriptors declare it as `targetReadiness`, never on their `timeoutPolicy`"). If a future descriptor did attach it there, `registry.ts`'s merge (`descriptor.targetReadiness ? {...} : descriptor.timeoutPolicy`) stores the policy object as-is, so `readinessBudgetMs` would honor the trait while `commandAcceptsReadinessBudget` (which reads `descriptor.targetReadiness`) would not — envelope widening and replay default-budget gating diverge silently. The sealing invariant is currently prose-only.</comment>

<file context>
@@ -75,8 +75,21 @@ export type CommandTimeoutPolicy = {
+   * so the envelope can widen by the readiness budget. Descriptors declare it as `targetReadiness`,
+   * never on their `timeoutPolicy`.
+   */
+  targetReadiness?: CommandTargetReadiness;
 };
 
</file context>
Fix with cubic

@thymikee

Copy link
Copy Markdown
Member Author

Thanks for the update. I reviewed bd0aff2. The delta mostly moves target-resolution code out of resolution.ts, but the four defects from the first review are still open, and there is no live replay run yet.

First, a capture that ignores the engine's signal can end as stalled. In post-gesture-stability.ts, the capture adapter at line 158 ignores the deadline signal. If a 300-800 ms iOS capture outlives the rest of the 1.5 s budget, observeUntil returns {kind:'stalled', error: undefined} and line 206 runs throw undefined. Before this PR, the loop judged the late capture and returned unsettled with the post_gesture_snapshot_stabilization_timeout warning. Now a surface that never settles (spinner, animation) fails the next snapshot or interaction with an unknown error. The rule should be that a capture that ignores the signal is never classified as stalled, and a stalled or expired end gives the same outcome as the old deadline expiry. Please enforce it in observeUntil (for example a per-capture deadline mode of none vs cancel), so readiness, post-gesture and scroll all inherit it. Please add a test where the capture advances the clock past the remaining budget and the result is unsettled, not a throw.

Second, the same engine rule breaks scroll in scroll-movement.ts. Line 274 wraps readOneCapture without the signal, so a slow post-scroll capture that showed movement ends as stalled. Line 290 then maps every non-done result to budgetExpiredVerdict. A scroll that moved now reports a surface-unsettled result and a scroll_movement_budget_expired warning, where it used to report movement. This contradicts the claim that verdicts are unchanged. Once the rule above is in place, please judge the value a stalled or late capture returned before falling back to the expired result, and add a scroll-movement-observation scenario with a slow capture that moved.

Third, the ride-out branch loses the unreadable reason. In observe-until.ts, the branch returns undefined without recording poll.error in state.lastError, which is only initialised at line 116. If readiness rides out every capture as unreadable content, the loop ends expired with last and lastError both undefined. Then selector-readiness cannot surface the reason, and users get selector_not_found against an empty node list, which tells them to fix a selector that is correct. Please set state.lastError = poll.error in that branch. Please add an observe-until test and a selector-readiness test where every capture throws the unreadable-content error, and assert the typed unreadable reason.

Fourth, the replay readiness wait skips the verify gate. In session-replay-action-runtime.ts, the new change swaps the hardcoded set for commandAcceptsReadinessBudget, but the default budget is still added only at dispatch in buildReplayActionFlags. Recorded press, click and longpress steps carry targetEvidence, and session-replay-target-verification.ts checks them before dispatch. That file is unchanged, so an absent or unreadable capture there returns REPLAY_DIVERGENCE with no readiness wait. For recorded steps, which is the case #3069 targets, the wait never runs when the target is still loading at the verify gate. The rule should be that every replay gate that reads the target before dispatch applies the same readiness budget as dispatch. Please apply the budget in the pre-dispatch verification capture for any command with targetReadiness of budgeted, and add a replay test with targetEvidence where the target appears after the first capture.

No live replay run is attached yet. Please replay a recorded .ad script with targetEvidence on a simulator, with a press target that appears after about 1 s. Attach the passing step output and the --debug ndjson poll timeline. Please also attach a miss run where the target never appears, showing error.details.readiness and a typed reason, not selector_not_found on an empty tree.

I read the code and the new registry consistency test but did not run any suite, and I could not observe device behavior. Smoke Tests and Coverage are still running. This change moves the press, click and longpress target-resolution code that smoke tests exercise, so a failure there could relate to it. There is no failure to point at yet. No conflicts. Before merge, please fix the four defects at their owners, add a regression test for each, and attach the live pass and miss runs.

@thymikee

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

Split as you suggested, and the four defects are fixed at c1280738ee.

Split. This PR is now the engine plus its one consumer, the readiness wait. The post-gesture and scroll migrations, the outcomeObservation cell, and the scroll scenario moved to feat/observation-engine-loops (stacked on this branch; PR to follow once the migrations use the new deadline mode and carry slow-capture scenarios). If #3077 drops the wait, this PR goes away whole and leaves no migrated loops.

Defects 1 and 2 (late captures) (1ea9c29cb7): the engine now declares a per-capture deadline mode, captureDeadline: 'cancel' | 'none'. Under 'none' the capture gets no signal, a late capture is judged, and the loop ends done or expired; under 'cancel' an aborted capture ends stalled. Readiness keeps 'cancel'; the trace that its signal reaches the platform capture is in the commit body (toBackendContext → interaction-runtime.ts:137 → snapshot-capture → captureSnapshotSignal, the same cancellation wait uses). The migrations were removed from this PR rather than patched; they come back on the stacked branch with 'none' and the slow-capture tests you asked for.

Defect 3 (lastError) (1ea9c29cb7): the ride-out branch records poll.error; readinessExhaustedFailure raises it for both expired and stalled when no tree was observed. Engine test: all polls ridden out → expired with lastError. Readiness test: every capture throws an unreadable-content error → that typed error surfaces, not selector_not_found.

Defect 4 (replay gate) (da77619439, 1ec434350b): the pre-dispatch target classification now polls under the same observeUntil schedule (in replay-port, which already depends on capture-kit; ad-replay stays free of it), only for a budgeted command with a selector token. Deferring the miss to dispatch was rejected: it would lose the identity guard and turn selector-miss into action-failure. Then the step has one budget: the dispatch gets max(0, budget − gate waitedMs), so a target rendering at 1.6 s leaves the tap 400 ms. The gate emits interaction_target_readiness (same phase and data as the daemon) when it polled more than once, and an annotated selector-miss divergence carries error.details.readiness. Recorded scripts do carry the annotations: written at ad-script/src/internal/script-formatting.ts:43-48 → target-annotation-serde.ts:169, parsed at script.ts:100 → :249, round-trip tests at target-annotation-serde.test.ts:68/74. Step-loop tests: target on the second capture passes with one diagnostic; never appears diverges as before.

Non-blocking, all taken (3b25e815a5, 333c31c9ab): SAFE_CAUSE_DETAIL_KEYS includes readiness; docs say where readiness appears (only expired/stalled/sparse; covered, off-screen, ambiguous and all-unreadable carry their own error; "performs the requested interaction"); the envelope widens by min(budget, 2000) and only for commands with the targetReadiness: 'budgeted' registry trait, which also replaces the hardcoded replay set (a parity test ties the trait to the guarantee cells; fill refuses the key; the 2000 is pinned to the selector row's maxTimeoutMs by test rather than imported, so the command-registry closure does not grow); isRideOutCaptureError and the observation subpath are deleted; the eager baselineEndsInDirection went with the scroll migration.

One cubic finding was false and led to a deletion: press captures never reach the 750 ms selector cache (that cache lives only in the selector backend serving wait/find/is/get), so the forceFresh plumbing this branch had added was unreachable and is removed (04b0c6930a, trace in the commit body).

Live runs: the two from the migration half move with it. The annotated replay with an engaged wait is still open: 46 local runs never engaged the wait on this host, which is the reason #3077 exists.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

3 issues found across 44 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="packages/replay-port/src/daemon-port/session-replay-runtime-engine-adapter.ts">

<violation number="1" location="packages/replay-port/src/daemon-port/session-replay-runtime-engine-adapter.ts:210">
P2: This branch collapses sparse and exhausted unreadable captures into generic target-unavailable failures. Preserve the typed capture outcome through `AdReplayTargetObservation` and map sparse to `capture_sparse` and exhausted unreadable captures to their original typed error before building the divergence.</violation>
</file>

<file name="src/commands/interaction/runtime/selector-readiness.test.ts">

<violation number="1" location="src/commands/interaction/runtime/selector-readiness.test.ts:118">
P3: `notEqual(error.details?.reason, 'selector_not_found')` asserts absence loosely: it passes for any other `reason` value (e.g. a wrapping failure that somehow preserves the `androidSnapshotHelperFailureReason` detail), and duplicates the magic string instead of the contract. The declared behavior is that the ridden-out unreadable error is surfaced unchanged with no selector-miss conversion, so pin that exactly: `assert.equal(error.details?.reason, INTERACTION_ERROR_REASONS.selectorNotFound)` would also be wrong — assert the selector-miss reason is *absent*: `assert.equal(error.details?.reason, undefined)` (import `INTERACTION_ERROR_REASONS` from `@agent-device/selectors/interaction-error` if you want to assert against the constant instead of the literal).</violation>
</file>

<file name="src/daemon/__tests__/replay-target/session-replay-target-verification-runtime.test.ts">

<violation number="1" location="src/daemon/__tests__/replay-target/session-replay-target-verification-runtime.test.ts:249">
P3: The test name promises `fill` refuses a selector miss on one capture, but nothing asserts `fill` captured only once. A regression that made `fill` wait for its target — then time out with the target never rendering — would still end in a selector-miss and pass this test unchanged (only the wait-then-hit variant is caught, indirectly, by the mocked second capture). Mirror the click gate test's `toBeGreaterThan(1)` assertion and pin the single capture.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

const observation = await captureTargetObservation(action);
lastObservation = observation;
if (observation.state !== 'available') {
return { state: 'unavailable', reason: observation.reason, hint: observation.hint };

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This branch collapses sparse and exhausted unreadable captures into generic target-unavailable failures. Preserve the typed capture outcome through AdReplayTargetObservation and map sparse to capture_sparse and exhausted unreadable captures to their original typed error before building the divergence.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/replay-port/src/daemon-port/session-replay-runtime-engine-adapter.ts, line 210:

<comment>This branch collapses sparse and exhausted unreadable captures into generic target-unavailable failures. Preserve the typed capture outcome through `AdReplayTargetObservation` and map sparse to `capture_sparse` and exhausted unreadable captures to their original typed error before building the divergence.</comment>

<file context>
@@ -167,46 +202,62 @@ export function createAdReplayStepRuntime(params: {
+        const observation = await captureTargetObservation(action);
+        lastObservation = observation;
+        if (observation.state !== 'available') {
+          return { state: 'unavailable', reason: observation.reason, hint: observation.hint };
+        }
+        const session = ctx.sessionStore.get();
</file context>
Fix with cubic

(error: unknown) => {
assert.ok(error instanceof AppError);
assert.equal(error.details?.androidSnapshotHelperFailureReason, 'system-window-only');
assert.notEqual(error.details?.reason, 'selector_not_found');

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: notEqual(error.details?.reason, 'selector_not_found') asserts absence loosely: it passes for any other reason value (e.g. a wrapping failure that somehow preserves the androidSnapshotHelperFailureReason detail), and duplicates the magic string instead of the contract. The declared behavior is that the ridden-out unreadable error is surfaced unchanged with no selector-miss conversion, so pin that exactly: assert.equal(error.details?.reason, INTERACTION_ERROR_REASONS.selectorNotFound) would also be wrong — assert the selector-miss reason is absent: assert.equal(error.details?.reason, undefined) (import INTERACTION_ERROR_REASONS from @agent-device/selectors/interaction-error if you want to assert against the constant instead of the literal).

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/commands/interaction/runtime/selector-readiness.test.ts, line 118:

<comment>`notEqual(error.details?.reason, 'selector_not_found')` asserts absence loosely: it passes for any other `reason` value (e.g. a wrapping failure that somehow preserves the `androidSnapshotHelperFailureReason` detail), and duplicates the magic string instead of the contract. The declared behavior is that the ridden-out unreadable error is surfaced unchanged with no selector-miss conversion, so pin that exactly: `assert.equal(error.details?.reason, INTERACTION_ERROR_REASONS.selectorNotFound)` would also be wrong — assert the selector-miss reason is *absent*: `assert.equal(error.details?.reason, undefined)` (import `INTERACTION_ERROR_REASONS` from `@agent-device/selectors/interaction-error` if you want to assert against the constant instead of the literal).</comment>

<file context>
@@ -93,3 +93,31 @@ test('runtime press caps a readinessTimeoutMs larger than the row maxTimeoutMs a
+    (error: unknown) => {
+      assert.ok(error instanceof AppError);
+      assert.equal(error.details?.androidSnapshotHelperFailureReason, 'system-window-only');
+      assert.notEqual(error.details?.reason, 'selector_not_found');
+      return true;
+    },
</file context>
Suggested change
assert.notEqual(error.details?.reason, 'selector_not_found');
assert.equal(error.details?.reason, undefined);
Fix with cubic


const response = await scene.replay();

expect(scene.invoked.length).toBe(0);

@cubic-dev-ai cubic-dev-ai Bot Sep 30, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The test name promises fill refuses a selector miss on one capture, but nothing asserts fill captured only once. A regression that made fill wait for its target — then time out with the target never rendering — would still end in a selector-miss and pass this test unchanged (only the wait-then-hit variant is caught, indirectly, by the mocked second capture). Mirror the click gate test's toBeGreaterThan(1) assertion and pin the single capture.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/daemon/__tests__/replay-target/session-replay-target-verification-runtime.test.ts, line 249:

<comment>The test name promises `fill` refuses a selector miss on one capture, but nothing asserts `fill` captured only once. A regression that made `fill` wait for its target — then time out with the target never rendering — would still end in a selector-miss and pass this test unchanged (only the wait-then-hit variant is caught, indirectly, by the mocked second capture). Mirror the click gate test's `toBeGreaterThan(1)` assertion and pin the single capture.</comment>

<file context>
@@ -177,6 +194,63 @@ test('a selector-miss divergence blocks dispatch and never sends the action', as
+
+  const response = await scene.replay();
+
+  expect(scene.invoked.length).toBe(0);
+  expect(response.ok).toBe(false);
+  if (response.ok) return;
</file context>
Suggested change
expect(scene.invoked.length).toBe(0);
expect(scene.invoked.length).toBe(0);
expect(mockDispatchCommand).toHaveBeenCalledTimes(1);
Fix with cubic

@thymikee

Copy link
Copy Markdown
Member Author

The earlier findings at bd0aff2 are mostly addressed in c128073, but two things remain: a duplicated readiness schedule and missing live evidence.

The replay gate and the dispatch still build the same readiness schedule in two packages. replayStepReadinessSchedule rebuilds { intervalMs: poll.intervalMs, budgetMs: Math.min(readinessTimeoutMs, poll.maxTimeoutMs) }. It reads SELECTOR_PIPELINE_POLICIES.promotedTarget.poll directly. readinessScheduleFor builds the same value from params.pipeline.poll. The PR says the gate and the dispatch poll under the same schedule, so neither refuses a target the other would still wait for. Two copies enforce that today. If the dispatch changes its cap, cadence, or budget origin, the gate can refuse a target with a selector-miss divergence that the dispatch would still wait for, and the step is never sent. The rule should be: one function derives the schedule from the pipeline poll row and the caller budget, and both sides call it. Could readinessScheduleFor move next to SELECTOR_PIPELINE_POLICIES in @agent-device/selectors/selector-pipeline-policy, which owns the poll row? Then resolution.ts and replayStepReadinessSchedule (through the dynamic import already there) would both call it, and the local copy goes away.

The gate wait changes a device-facing route, and no live run shows it engaging. Replay of an annotated press, click, or longpress now re-captures the device before dispatch, at this call. The 46 local runs you report never entered the wait. Unit tests use a fake capture and an instant clock, so the capture cadence, the launch-race retry inside each poll, and the hand-off of the remaining budget to the dispatch are not proven on a device. Please make the route deterministic, for example with a test-app screen whose target button renders about 1 s after navigation. Then attach two simulator runs with --debug. In the pass run, replay a .ad with the # agent-device:target-v1 annotation above the press. The ndjson must show an interaction_target_readiness line from the gate with polls greater than 1 and end done, and the step must pass. In the miss run, the target never appears. It must show REPLAY_DIVERGENCE with details.divergence.kind selector-miss, details.readiness.end expired, waitedMs near 2000, and no dispatched press.

Not blocking, and you can take or leave it: ReplayDaemonDependencies.clock is set only by the test fixture, and the replay gate counts its budget from the start of the wait while the dispatch counts from its first capture, which sits a little against the "same schedule" comment there.

Is there a smaller shape than this? The split answers the earlier challenge, and the one new seam (observeTarget merging capture and classify, plus the dynamic import of capture-kit/observe-until) is reasonable. The only smaller change I would ask for is the first finding: derive the schedule in the selectors policy owner instead of keeping two local copies.

I read the code at c128073 and ran no tests. I did not trace whether press captures ever reach the 750 ms selector snapshot reuse window, so the reason for deleting forceFresh is unverified. I took the cancellation trace (toBackendContext to interaction-runtime.ts:137) from your description without reading it. The gate honors cancel only between polls, because the capture and the interval sleep ignore the signal. That seems acceptable to me, but no test covers cancelling the gate.

The Smoke Tests job failed in "Install Linux desktop dependencies", where the apt step timed out after 6 minutes, before any test ran. The diff touches no CI config, apt, or Linux desktop setup, so this looks unrelated to the change. Please rerun it, since a real Smoke result is needed for the press and click resolution route this PR changes. Before merge, the schedule needs one owner, the annotated-replay pass and miss runs need to be attached, and Smoke Tests needs a green rerun.

…get, two loops migrated

Prototype only. Shared observeUntil loop in capture-kit with one ride-out
classifier in contracts; promotedTarget declares a 2 s / 200 ms readiness
budget consumed by selector resolution with the first capture unbounded;
post-gesture stability and scroll movement run as verdict predicates over
the engine; targetReadiness and outcomeObservation guarantee cells.
…gpress

promotedTarget.poll becomes {intervalMs, maxTimeoutMs} (no row default). A
new operator-only readinessTimeoutMs common input field (no CLI flag, no MCP
exposure) gates resolveSelectorInteractionTarget's poll: a positive integer
caps at the row's maxTimeoutMs and runs the readiness loop; absent (MCP,
CLI, and every caller today) takes the one-attempt path, so a wrong-selector
miss fails on the first capture instead of paying a 2s wait. Replay defaults
readinessTimeoutMs: 2_000 onto dispatched press/click/longpress steps unless
already set.

Split resolution.ts (1,411 lines) three ways: selector-readiness.ts owns the
readiness poll and its capture-attempt/failure machinery,
interaction-snapshot-capture.ts and covered-interaction-error.ts are new
leaves shared with resolution.ts without a value-import cycle (verified via
the layering scan). interaction-guarantees.ts's targetReadiness via now
points at selector-readiness.ts#pollForSelectorReadiness.
…t.ts

client.test.ts is already over the test-file-size-ratchet tripwire and may
not grow; the new press/click/longpress readinessTimeoutMs test moves to its
own small file instead.
resolution.ts (1,131 lines) keeps the entry and the point/selector target
kinds (249 lines). The rest moves, function bodies unchanged:

- ref-target-resolution.ts (262): how an @ref becomes a node, and when it
  is refused (adopt/read, frame reconciliation, tryResolveRefNode,
  refMissRefusal).
- resolution-disclosure.ts (179): what a resolution records and reports
  (EXACT_REF_RESOLUTION, buildRefResolution and its ResolvedRefNode type,
  selector disclosure, non-hittable hint, press recording retarget).
- target-visibility-stages.ts (219): is the resolved target reachable on
  screen: the node-stage runner, the covered refusal (covered-interaction-
  error.ts folded in and deleted) and the off-screen stage.
- native-ref-interaction.ts (135): the native-ref preflight and dispatch.
- resolution-touch-point.ts (48): node to touch point, shared by the ref and
  selector kinds, so neither kind module owns the other.
- replay-target-guard.ts (65): assertExpectedResolvedTarget and
  assertReplayTargetResolution, used by the ref and selector kinds and by
  selector-is/selector-read.
- interaction-resolution-request.ts (73): InteractionAction,
  ResolveInteractionTargetParams and ExpectedResolvedTarget. The layering
  scan (R9/R10) refused the split while sibling modules type-imported
  resolution.ts: the largest type cycle grew from 6 to 8 files. The request
  types now sit below every module that reads them.

Not a pure move: `export` added to symbols now crossing a module boundary;
two comments that pointed at "above" now name the file; the stale
fallow-ignore on EXACT_REF_RESOLUTION (now statically imported) and on
pollForSelectorReadiness (already statically imported by resolution.ts)
are removed. interaction-guarantees.ts `via` paths and three cross-package
comments follow the symbols.

resolution.test.ts (706) splits along the same seams: resolution.test.ts
350, ref-target-resolution.test.ts 126, resolution-disclosure.test.ts 19,
target-visibility-stages.test.ts 227. Test bodies unchanged.

git diff -M90% --stat: 24 files changed, 1399 insertions(+), 1312
deletions(-). No file passes the 90% rename threshold (it is a split);
resolution.ts +18/-900, resolution.test.ts +1/-357. A line-multiset
comparison of old against new bodies (imports and blank lines excluded)
differs only in the export keywords and comments listed above. No fallow
baseline entry was keyed on the moved paths.
…laration

Replay hardcoded press/click/longpress as the commands that get a default
readinessTimeoutMs. Nothing declared that set per command:
readinessTimeoutMs is a common input field every command reads.

press, click and longpress now declare `targetReadiness: 'budgeted'` in the
command registry. Replay reads it through commandAcceptsReadinessBudget.
resolveCommandTimeoutPolicy attaches the trait to the resolved timeout
policy, and resolveCommandRequestTimeoutMs widens the envelope by the
readiness budget only when that trait is present. Before, it widened for
any command whose flags carried readinessTimeoutMs; only these three
commands forward the field to request flags, so no request changes.

A test asserts that the commands with the trait equal the commands the
`runtime` targetReadiness guarantee cells in interaction-guarantees.ts
apply to. The replay default test iterates the declared set.
…igrations move to a stacked PR

The post-gesture stability and scroll-movement loops return to their main
versions, together with the scroll-movement observation provider scenario and
the outcomeObservation guarantee cell that described them. They move to a
stacked PR: both loops run captures that may not honor a per-capture deadline
signal, so the engine needs a declared per-capture deadline mode and
slow-capture scenarios before they can migrate without turning a late capture
into a stalled end.

The shared ride-out classifier had no production reader (readiness rides out
isUnreadableCaptureContentError), so it is deleted with the contracts
observation subpath. ObservationEnd has one reader, the engine, and now lives
in observe-until.ts.
…den-out errors

observeUntil schedules now declare captureDeadline. 'cancel' keeps the
existing contract: each capture is armed with the remaining budget and a
capture that ends at or past its deadline ends the loop stalled. 'none' hands
the capture no deadline, so a capture that cannot be cancelled is still judged
when it finishes late and the loop ends done or expired, never stalled.

A ridden-out poll now records its error, so a loop whose every capture was
ridden out ends expired with lastError set; the readiness wait raises that
typed unreadable-content error instead of a selector miss against an empty
tree. A terminal capture failure is labelled 'failed' in the poll timeline
instead of 'observed'.
…ap its envelope

readinessTimeoutMs was a common input every structured command accepted and
ignored. Only commands whose descriptor declares targetReadiness: 'budgeted'
now carry it, through targetReadinessFields; the common input reader refuses
the key with INVALID_ARGS for every command whose fields do not declare it.

The request envelope widened by the raw readinessTimeoutMs, so a caller value
of 999000 stretched it by 999 s while the poll itself stops at the
promotedTarget row's ceiling. The widening now uses the same capped budget.

The press readiness provider scenario claimed to exercise forceFresh through
the selector capture cache; the press route captures through the interaction
backend, which keeps no such cache, so the comments now say so.
…rt readiness

A recorded press/click/longpress step carries target evidence, and replay
classified that target from one capture before dispatch. A target that had not
rendered yet diverged as a selector miss before the dispatch's own readiness
wait could run, so the replay default budget never applied to annotated steps.

The engine now asks the port to observeTarget, which captures and classifies
in one capability. For a step whose dispatch would wait (a readiness-budgeted
command resolving a selector), the port re-captures under the same schedule
while the recorded target is a selector miss; any other classification or an
unavailable capture ends the wait. A target that never appears diverges as
before. The divergence capture takes no per-capture signal, so the wait
declares captureDeadline 'none'. ReplayDaemonDependencies gains an optional
clock the wait paces itself by.

REPLAY_DIVERGENCE now carries the cause's readiness detail at
error.details.readiness.
The client API page says the command performs the requested interaction after
the wait, and names the failures that carry readiness evidence and those that
do not. The replay page says how a missing target surfaces for annotated and
unannotated steps.
…ess capture route has no cache

The readiness poll asked for a fresh capture from its second poll on, but no
press, click, or longpress capture can be served from the selector capture
cache, so the flag and the backend option that carried it are removed and
selector-runtime-backend.ts, backend.ts, and backend-snapshot-options.ts
return to main.

Trace:
- selector-readiness.ts pollSelectorReadinessOnce captured through
  attemptSelectorResolution, which calls captureInteractionSnapshot
  (interaction-snapshot-capture.ts:35) and so runtime.backend.captureSnapshot.
- The daemon press route builds that runtime with createInteractionRuntimeForRoute
  (src/daemon/interaction/internal/interaction-touch-press.ts:10, used for
  longPress/click/press at :177-187).
- Its backend is createInteractionBackend
  (src/daemon/interaction/internal/interaction-runtime.ts:126), whose
  captureSnapshot (:132-139) forwards interactiveOnly, preferredBackend,
  includeRects, and the signal, and nothing cache-related.
- That capture calls captureSnapshotForSession (src/daemon/interaction/index.ts:21),
  which calls the daemon captureSnapshot (src/daemon/snapshot-capture.ts:77).
- captureSnapshot runs captureSnapshotAttempt, which runs captureSnapshotData
  (snapshot-capture.ts:99) and captures live through the bound capture or
  captureSnapshotWithInteractor on every call.
- The 750 ms cache (SELECTOR_CAPTURE_CACHE_TTL_MS,
  src/daemon/selector-capture-runtime.ts:24, read at :253 and :280) exists only
  inside createSelectorBackend (src/daemon/selector-runtime-backend.ts:142).
- createSelectorRuntimeForDevice is reached from the selector read and wait
  routes (selector-runtime.ts, wait-runtime.ts:158), never from
  interaction-touch-press.ts.

The runtime-selector targetReadiness contract scenario keyed its late target on
the removed option; it now keys on the capture count.
…he dispatch

A replayed step has one readiness budget. The pre-dispatch target gate and the
dispatch each spent the full budget, so a target that rendered late could
cost a step twice the budget. When the gate polled more than once, the
dispatch now receives what is left, max(0, budget - gate waitedMs), as its
readinessTimeoutMs; 0 makes the dispatch resolve its target once. A gate that
matched on its first capture leaves the dispatch the whole budget.

A gate that polled more than once emits the interaction_target_readiness
diagnostic the dispatched wait emits, with polls, waitedMs, end, and the
step's command. A selector-miss divergence after such a wait carries the same
evidence at error.details.readiness, where an unannotated step's
target-not-found failure already carries it.
…lector policy

The envelope widening imported the promotedTarget selector row to read its
2 s poll ceiling, which made every command-registry entry and src/cli.ts
evaluate the selector policy module. The cap is now READINESS_BUDGET_MAX_MS in
timeout-policy.ts, pinned by test to the row's maxTimeoutMs; selectors cannot
import command-registry, so the test is the tie.

The same eager-closure gate showed two more new edges from this branch:
- replay's native and test command entries evaluated observe-until and the
  selector policy through the target gate. Both now load inside the gate's
  readiness path, the only code that uses them.
- src/cli.ts evaluated the new target-readiness-grammar module. Its one
  function moves into interaction/metadata.ts, its one caller, and its tests
  into metadata.test.ts.
The replay target gate and the dispatch each built the promotedTarget
readiness schedule from the row's poll budget. readinessScheduleFor now lives
next to SELECTOR_PIPELINE_POLICIES and both callers use it: the dispatch
directly, the replay gate through its existing dynamic import.

The two waits also counted the budget from different origins: the gate from
the start of its wait, the dispatch from the end of its first capture. The
schedule now carries budgetFrom: 'first-capture' for both, and the budget the
gate leaves the dispatch is the step budget minus the gate's time after its
first capture ended. The first capture is the one-attempt lookup a step pays
without a wait, in the gate as in the dispatch.
…ne; cancel is an opt-in

captureDeadline is now optional and defaults to 'none'. A capture only ends the
loop stalled when its caller declares 'cancel', so an adapter whose capture
ignores the signal cannot claim cancellation by omission. The selector
readiness poll keeps its explicit 'cancel': its poll signal reaches the
platform as CaptureSnapshotInput.signal. The replay target gate drops its
explicit 'none'.
An error thrown out of the replay run reached the wire as code and message
only; the catch rebuilt it with errorResponse and dropped details, hint, and
diagnostic references. A request canceled during the pre-dispatch target gate
therefore lost its request_canceled reason. The catch now normalizes the error
with normalizeError and appends the run's artifactPaths to its details.
The target gate's divergence capture takes no signal, so the gate honors a
canceled request only at a poll boundary. A request canceled during the gate's
second capture ends the wait when that capture returns: two captures, the
request_canceled reason on the response, and no dispatch.
observeUntil checked the caller's signal only at the top of each iteration,
before the sleep. A cancel that arrived during the sleep still paid one more
capture before the loop noticed. The loop now checks the signal after the
sleep and ends with the typed request-canceled error before starting a poll.
The host-kit sleep takes no signal, so the check after the sleep is the
boundary.
@thymikee

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

Both blocking items are addressed at d476097d60 (rebased on 0aeaba7bfc after #3070 merged).

One owner for the schedule (1f8f627030). readinessScheduleFor(poll, readinessTimeoutMs) now lives beside the poll row it derives from, in packages/selectors/src/selector-pipeline-policy.ts:249, and both the dispatch (resolution.ts) and the replay gate (through its existing dynamic import) call it; both local copies are deleted. A selectors test asserts the gate and the dispatch produce the same schedule for every readiness-budgeted command. The budget origin is one as well: both waits count from the end of their first capture (the one-attempt lookup a step pays anyway), and the gate hands the dispatch max(0, budget − time after the gate's first capture ended); a 500 ms first capture plus seven misses leaves the dispatch 400 ms. The "same schedule" comment is true now.

Engine default (1b93af6adc): captureDeadline defaults to 'none'; 'cancel' is an explicit opt-in that readiness makes with its signal trace. A signal-blind adapter can no longer declare 'cancel' by accident, which is the shape #3078's two regressions had.

Cancel (9c76e9a3d8, d476097d60): cancelling the request during the gate wait ends at the next poll boundary with reason: request_canceled and no dispatch (test: exactly two captures); and observeUntil now checks the signal after the inter-poll sleep, so a cancel that lands during the sleep costs no further capture (the host-kit sleep takes no signal, so the post-sleep check is the boundary). One fix this needed: native-command.ts's outer catch rebuilt thrown errors as code and message only, dropping typed details such as request_canceled; it now normalizes (b0921253b4).

Clock (fa7f36af7d): kept, following AgentDeviceRuntime.clock (optional, wall clock when absent, never set in production); one JSDoc sentence says so.

Live runs: on this head (d476097d60, built), a fresh iPhone 17 Pro simulator (iOS 26.2), the Release fixture app, isolated state dir. The deterministic target is the app's own Settings screen: "Load diagnostics" renders "Retry diagnostics" about 1.1 s later, so no app change was needed. The script was recorded with open --relaunch --save-script, and the recorder annotated every press with # agent-device:target-v1 {...}, so both replays cover the annotated path; the wait line was removed so the Retry press follows the Load press directly.

Pass, replay pass.ad --session live --debug --json: exit 0, replayed: 8. The gate's line for the Retry press:

{"phase":"interaction_target_readiness","command":"replay","data":{"polls":3,"waitedMs":1475,"end":"done","command":"press"}}

The other three presses show polls: 1 (118–144 ms).

Miss, same script with the last press changed to press "label=\"Retry diagnostics never\"" (annotation kept): exit 1, error.code: REPLAY_DIVERGENCE, error.details.divergence.kind: "selector-miss" (cause.code: SELECTOR_MISS, step 7), error.details.readiness: {"polls":5,"waitedMs":2576,"end":"expired"}, and no tap after the Load press in the request log (only snapshot sends until the last readiness line). Two notes on the numbers: readiness sits at error.details.readiness (the same place the unannotated path puts it), and waitedMs is 2576 rather than ~2000 because the budget counts from the end of the first capture and a capture in flight at the deadline completes (one-interval overrun), which is the engine's stated rule.

Smoke Tests: the earlier red was the Install Linux desktop dependencies apt timeout, before any repository code ran; the rerun is green.

@thymikee

Copy link
Copy Markdown
Member Author

This PR is ready. Both blocking findings from the earlier review at c128073 are fixed in d476097, and the delta no longer has any blocking problem.

Not blocking, and you can take or leave these: the adapter rebuilds the engine's budget start from the poll timeline in budgetSpentMs, so could ObservationEvidence report the remaining or spent budget and let you delete that helper; and the "one schedule" test in session-replay-action-runtime.test.ts calls readinessScheduleFor with promotedTarget directly, so it proves the flag hand-off but not that the dispatch route picks that row.

CI shows 21 checks with none failing. The earlier Smoke failure was an apt-install timeout that ran before any repository code, and the rerun is reported green. There are no conflicts.

For evidence, I took the live pass and miss runs from your quoted output and did not inspect the raw ndjson or request-log artifacts. I ran no tests, and I confirmed the regression is real by reading the pre-change code in 1481f36. The miss run's 2576 ms wait against the 2000 ms cap matches the engine's rule that the budget counts from the end of the first capture, so one capture can overrun. I did not re-check the request envelope's headroom, since that code is outside this delta. I also did not re-review the resolution.ts domain split from c5bd3fd, which the earlier review covered. I only checked that removing the local readinessScheduleFor there is equivalent.

Nothing else stands in the way, so this is ready for a maintainer to merge.

@thymikee thymikee added the ready-for-human Valid work that needs human implementation, judgment, or maintainer merge label Sep 30, 2026
@thymikee
thymikee added this pull request to stack #3086 October 1, 2026 06:13
@thymikee
thymikee merged commit 06d3e81 into main Oct 1, 2026
21 of 22 checks passed
@thymikee
thymikee deleted the proto/reliability-contract branch October 1, 2026 06:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-for-human Valid work that needs human implementation, judgment, or maintainer merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant