Skip to content

refactor(interaction): move post-gesture stability and scroll movement onto observeUntil - #3078

Merged
thymikee merged 7 commits into
mainfrom
feat/observation-engine-loops
Oct 1, 2026
Merged

thymikee merged 7 commits into
mainfrom
feat/observation-engine-loops

Conversation

@thymikee

@thymikee thymikee commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Summary

Part of #3069. Stacked on #3072 (base branch proto/reliability-contract); do not merge before it.

Both loops now judge a late capture instead of ending stalled (the engine defaults to captureDeadline: 'none' on the base; readiness is the one opt-in). Post-gesture: every non-done end returns main's outcome (last value, unsettled unless the last pair agreed, plus the stabilization-timeout warning); a capture error is rethrown as itself. Scroll: only an expired end maps to budgetExpiredVerdict; the edge-rest analysis is lazy again. Clock-driven adapter tests cover a capture that overruns the budget, a 2 s first capture, and a late poll that shows movement. The scroll provider scenario mocks the Simulator AX bridge probe that was costing 1.5 s per test. The outcomeObservation cell's waiver reasons are corrected from the traced marks; the coordinate cell joins the bounded gap list.

Moves the two existing polling loops onto observeUntil: post-gesture stability (packages/capture-kit/src/post-gesture-stability.ts) and scroll movement (src/daemon/scroll-movement.ts), plus the outcomeObservation guarantee cell and the scroll provider scenario. Verdicts are meant to stay exactly as on main (#2984's edge-rest rule, the distrust deadline reset, the no-effect bar).

Still a draft only because it must merge after #3072.

Validation

Tested commit 938a38c260. Local pnpm check:affected --run: 1593 of 1594 files green; daemon-session-idle-expiry.test.ts (untouched here) hit a tmp-dir cleanup race once and passes alone. Live, private iPhone 17 Pro simulator, this head: a gesture then snapshot on a page whose counter ticks every 100 ms returned the snapshot with the post_gesture_snapshot_stabilization_timeout warning (attempts: 4, durationMs: 1843), no UNKNOWN; scroll down on a 3000-node page reported movement: "moved" with a 3757 ms post-scroll capture. Both used local Safari pages because the fixture app has no never-settling screen and no tree over 1.5 s on this host. Details in the review thread.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Size Report

Metric Base Current Diff
Installed (including dependencies) 4.91 MB 4.91 MB +513 B
Package (unpacked) 4.91 MB 4.91 MB +513 B
Package (download) 1.47 MB 1.47 MB +162 B

Startup median (7 runs, lower is better):

Scenario Base Current Diff
CLI --version 21.7 ms 22.4 ms +0.8 ms
CLI --help 64.5 ms 65.0 ms +0.5 ms

@thymikee

Copy link
Copy Markdown
Member Author

Thanks for the migration. At 66eaec5 the code has two regressions, and the iOS and Android smoke jobs fail with what looks like the first one. I did not check that the base branch passes the same jobs, and I found no conflicts.

The post-gesture schedule at https://github.com/callstack/agent-device/blob/66eaec5/packages/capture-kit/src/post-gesture-stability.ts#L227 uses captureDeadline 'cancel', but the capture ignores the signal. A capture that finishes at or past max(remaining budget, 200 ms) ends the loop as 'stalled', even though it succeeded. Two ordinary cases hit this: a 1.4 s poll on an unsettled screen, and a first capture slower than 1.5 s. Then throw observed.error throws undefined, and toAppError turns it into UNKNOWN 'Unknown error'. On main the loop judged that late capture and returned the value, or 'unsettled' with post_gesture_snapshot_stabilization_timeout. So the first snapshot-backed read after a gesture (snapshot, get text, is, wait) fails instead of returning a snapshot with a warning. Both smoke failures match this: iOS get text and Android scrollToVisibleSelector both give UNKNOWN 'Unknown error'. I did not open the request ndjson, so this rests on the code path and not on a captured stack. The rule: no capture that the old loop would have judged may end the loop without an outcome. Every observeUntil adapter whose capture ignores the signal must declare captureDeadline 'none', and every non-done end must map to main's outcome. Please switch this schedule to 'none' and drop the stalled rethrow, or throw only a typed error. Please add a clock-driven test where a capture runs past the remaining budget and the loop returns unsettled plus the warning.

The scroll schedule at https://github.com/callstack/agent-device/blob/66eaec5/src/daemon/scroll-movement.ts#L68 has the same shape. readOneCapture ignores the signal, and a post-scroll capture of 1.5 s or more, or a later poll that outlives its tail deadline, ends 'stalled'. Then observed.kind !== 'done' routes to budgetExpiredVerdict, which returns 'surface-unsettled'. On main that capture was judged, and a changed surface returned 'moved'. On large iOS trees a scroll that moved content would lose its movement claim and emit a spurious scroll_movement_budget_expired warning. Please use 'none' here too, and add a clock-driven test where a capture slower than 1.5 s that shows movement reports 'moved'.

The new scenarios in https://github.com/callstack/agent-device/blob/66eaec5/test/integration/provider-scenarios/scroll-movement-observation.test.ts#L1 use instant captures only. They would pass the same on main and on this head, so they do not guard the timing changes above. No post-gesture test covers a slow capture either. The scenario also times out at the 5 s default in your local run. Please add injected-clock slow-capture cases for both loops as adapter unit tests, and fix or raise the scenario timeout so it runs green.

Both device-facing polling paths change, and the description only reports typecheck. Please attach two live runs. First, on an iOS simulator, do a gesture and then snapshot --json on a screen that never settles. The response should return the snapshot with the post_gesture_snapshot_stabilization_timeout warning and no UNKNOWN error. Second, run scroll down --json on a large tree whose post-scroll capture exceeds 1.5 s, with the durationMs visible in the --debug ndjson. The response should carry movement 'moved'. The iOS and Android smoke jobs must also be green.

Not blocking, take or leave: baselineEndsInDirection in src/daemon/scroll-movement.ts is now eager, so a dynamic import and analysis run on every scroll and scroll_movement_edge_rest_required can fire without any 'changed' reading, so restoring the lazy changeNeedsRest ??= would fit your own list; and the waiver reasons at https://github.com/callstack/agent-device/blob/66eaec5/packages/contracts/src/interaction-guarantees.ts#L167 say the deferred-outcome mark applies only to navigation-sensitive Android actions and target-authored gestures, while markDeferredInteractionOutcome also calls markPostGestureStabilization and markPendingInteractionOutcome, and the coordinate cell reads 'inapplicable' although that path can carry --verify or --settle. I did not trace those dispatch paths, so please check which of them reach the two marks and correct the reasons or narrow the cell.

Could this PR be split? The two loop migrations are the refactor: set captureDeadline 'none' on both, keep changeNeedsRest lazy, and map every non-done end back to main's outcome. The outcomeObservation guarantee column is a separate contract change and might fit better in its own PR. With 'none', the post-gesture adapter would need no stalled or failed branch beyond rethrowing a real capture error. #3072 (observeUntil) would need to merge first. Would it also help if observeUntil defaulted captureDeadline to 'none' and required an explicit opt-in for captures that honor the signal? That would make a signal-blind 'cancel' adapter impossible to declare by accident.

Before merge, both schedules need 'none' with the slow-capture tests, the smoke jobs need to pass, and the two live runs need to be attached.

@thymikee
thymikee force-pushed the proto/reliability-contract branch from c128073 to d476097 Compare September 30, 2026 19:01
@thymikee
thymikee force-pushed the feat/observation-engine-loops branch from 66eaec5 to 578bebb Compare September 30, 2026 19:01
@thymikee

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

Both regressions are fixed at 938a38c260 (rebased on #3072's head, where observeUntil now defaults to captureDeadline: 'none' and 'cancel' is an explicit opt-in, so a signal-blind adapter cannot declare it by accident).

  • Post-gesture (fix(capture-kit): judge a late post-gesture capture instead of ending stalled): the schedule uses the default; the stalled rethrow is gone; every non-done end returns main's outcome (the last value, unsettled unless its last pair agreed, plus post_gesture_snapshot_stabilization_timeout); a real capture error is rethrown as itself. Clock-driven tests: a capture that overruns the budget returns the late value, unsettled, one warning; a 2 s first capture still forms the quiet pair. Mutation: switching the schedule to 'cancel' fails both.
  • Scroll (fix(daemon): judge a late post-scroll capture; keep the edge-rest analysis lazy): default schedule; only an expired end maps to budgetExpiredVerdict; failed rethrows; changeNeedsRest ??= runs inside the capture after a changed reading, so the edge analysis and its dynamic import are lazy again (a test asserts an untouched surface never runs readScrollEdgeState). Clock-driven tests: a 2 s first capture that shows movement reports moved; a later 2 s poll is judged and reports moved.
  • Scroll scenario timeout: it was not the loop's timers. The scenario stubs launchctl/ps so the route resolves a stable target, which made every capture eligible for the Simulator AX bridge; the bridge then ran against the real host (xcrun, bridge prep, simctl spawn, 250 ms socket retries to its deadline) before falling back to the scripted runner, about 1.5 s per test. The scenario now mocks @agent-device/platform-apple/snapshot-source to report unsupported at once; the file runs in 1.5–1.8 s and the full provider-integration project passes 224/224.
  • outcomeObservation cell: the old waiver reason was false; every tap path finalizes through finalizeTouchInteraction → markDeferredInteractionOutcome, whose three marks are Android snapshot freshness (press/click/back/open on Android), markPendingInteractionOutcome (coordinate press/click on mobile, only under interactionOutcome.retryOnNoChange, which nothing in the repo sets), and markPostGestureStabilization (swipe/scroll/gesture swipe on mobile, or any action with postGestureStabilization: true, also unset). The shared waiver now says exactly that, and the coordinate cell moved from inapplicable to the same gap waiver (it carries --verify/--settle), joining the bounded gap list under Interaction guarantee gaps (ADR 0011 umbrella) #1081.

Live runs, private iPhone 17 Pro simulator, iOS 26.2, this branch built. The fixture app has no never-settling screen and no tree over 1.5 s on this host, so both used local Safari pages:

  1. A page with a counter ticking every 100 ms: scroll down 0.3 then snapshot --json --debug → success: true, the snapshot returned, data.warnings carries "The surface was still changing after scroll down 0.3 when this tree was read…", and the request ndjson shows post_gesture_snapshot_stabilization_timeout with attempts: 4, durationMs: 1843, lastPairAgreed: false, then request_success. No UNKNOWN. Negative control on the Home tab: no warning.
  2. A page of 3000 links: scroll down 0.3 --json --debug → data.movement: "moved", ndjson snapshot_capture durationMs: 3757 (backend xctest) for the post-scroll capture, scroll_movement_observed {movement: moved, attempts: 1, durationMs: 3894}. On the Catalog list the post-scroll capture is 140–170 ms (bridge) even under CPU load, so the app alone cannot exercise the late-capture path.

The iOS and Android smoke jobs rerun on this head; the Lint & Format red was two unformatted test files, fixed in the last commit.

@thymikee

Copy link
Copy Markdown
Member Author

This PR is ready at 938a38c. The stalled-capture route that broke the iOS and Android smoke jobs at 66eaec5 is gone, and all 13 checks now pass, including both smoke jobs.

Not blocking, and you can take or leave it: in post-gesture-stability.ts, surfaceOf keeps a one-slot cache keyed on the last value. The verdict calls surfaceOf(latest) first (https://github.com/callstack/agent-device/blob/938a38c/packages/capture-kit/src/post-gesture-stability.ts#L167), and that overwrites the slot. The call for previousValue on the next line then misses, so readSurface runs 2n-1 times for n captures where main ran it n times (a plain simulation gives 7 calls for 4 captures). Keeping a previous-surface slot, or calling surfaceOf(previousValue) before surfaceOf(latest), would fix it.

The live-run numbers come from the PR body. I did not open the raw --debug ndjson for the 3757 ms scroll capture or the 1843 ms stabilization timeout. The engine from #3072 shortens the last sleep to the remaining budget, so the final poll can land up to 200 ms earlier than on main. Capture counts and outcome kinds match main. In scroll-movement.ts, startedAt is rebuilt as Date.now() - observed.waitedMs, which mixes the real clock with the injected one under tests. I did not check which tests read it.

There are no conflicts. Please merge #3072 (proto/reliability-contract) first, then take this PR out of draft.

@thymikee thymikee added the ready-for-human Valid work that needs human implementation, judgment, or maintainer merge label Sep 30, 2026
Base automatically changed from proto/reliability-contract to main October 1, 2026 06:14
…ment onto observeUntil

Restores the two loop migrations, their scroll-movement provider scenario, and
the outcomeObservation guarantee cell on top of the engine and readiness wait.
Both schedules declare captureDeadline 'cancel', the engine's behavior before
the mode existed.
… stalled

The post-gesture capture hook takes no signal, so declaring 'cancel' let the
loop end stalled on a capture that finished past its deadline and rethrow it,
where main judged the late capture. The schedule now uses the default 'none'.
Every non-done end maps to main's outcome: the stabilization-timeout warning
and the last value, marked unsettled unless its pair agreed. A capture error
is rethrown as itself.

The loop takes an optional clock for its elapsed-time reads and the engine.
Tests with an injected clock pin a capture that overruns the remaining budget
(judged, unsettled, one warning) and a first capture slower than the budget
(still forms the quiet pair); both fail with the schedule switched to 'cancel'.
…lysis lazy

The scroll capture takes no signal, so the movement schedule now uses the
default 'none': a capture that finishes past the budget is judged, and only a
budget that ends while the verdict still continues maps to
budgetExpiredVerdict. A capture error is rethrown.

The migration asked baselineEndsInDirection before the first capture. It is
lazy again: the edge analysis and its dynamic import run only after a changed
reading, as on main.

Tests with an injected clock pin a first capture slower than the budget that
shows movement (moved) and a later poll that outlives the budget (judged,
moved); both fail with the schedule switched to 'cancel'. A third pins that an
untouched surface never asks the edge question of the baseline.
…n the scroll scenario

The scroll scenario stubs launchctl and ps so the snapshot route resolves one
stable app target. That also makes every capture eligible for the Simulator AX
bridge, which no scenario provider scripts: the bridge source runs xcrun on the
host, builds or locates its binary, spawns it with simctl, and retries its
socket until the bridge deadline before the route falls back to the scripted
runner. On a host with Xcode that cost more than a second per test, and under
the full provider-integration run both tests timed out at 5 s; the scroll loop
itself took 8 ms and 204 ms.

The scenario now mocks the bridge source to report unsupported at once, as it
does on a host without Xcode, so every capture takes the scripted runner. The
file runs in 1.5 to 1.8 s in the full project run.
The outcomeObservation waiver said the deferred-outcome mark applied only to
navigation-sensitive Android actions and target-authored gestures. The tap
paths finalize through finalizeTouchInteraction, which calls
markDeferredInteractionOutcome with the tap point: press/click get the Android
snapshot-freshness mark, the no-change tap retry when the request sets
interactionOutcome.retryOnNoChange, and post-gesture stabilization when the
request sets postGestureStabilization. Target drag (gesture drag) gets none.
The waiver now names these marks and says the next capture judges them, never
the tap's own response.

The coordinate cell said inapplicable, but coordinate press/click reach the
same finalize call with the same marks and accept --verify/--settle. It now
shares the runtime paths' gap waiver, and the bounded gap list gains
coordinate/outcomeObservation.
@thymikee
thymikee force-pushed the feat/observation-engine-loops branch from 938a38c to dfe9629 Compare October 1, 2026 06:15
@thymikee
thymikee marked this pull request as ready for review October 1, 2026 06:15

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

5 issues found across 7 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="packages/contracts/src/interaction-guarantees.ts">

<violation number="1" location="packages/contracts/src/interaction-guarantees.ts:192">
P3: This waiver claims the direct route "observes no outcome beyond the runner's own report", but the success path finalizes through finalizeTouchInteraction with scheduleInteractionOutcomeRetry defaulting to true, so it sets the same deferred markers (pending-outcome retry, post-gesture stabilization) that TAP_OUTCOME_NOT_OBSERVED_GAP documents as "judged by the next capture, never in this response". Scope the claim to the response ("never in its own response") and name the deferred markers as the sibling does, so the two waivers that the comment says share the same reasoning stay equally precise.</violation>

<violation number="2" location="packages/contracts/src/interaction-guarantees.ts:604">
P2: This shared waiver is inaccurate for the `maestro-non-hittable-fallback` path: `fill` can execute the fallback while carrying `--verify` or `--settle`, and its response can include that observation. Scope the direct-selector waiver separately or update this path’s contract to model the runtime fill case.</violation>
</file>

<file name="packages/capture-kit/src/post-gesture-stability.ts">

<violation number="1" location="packages/capture-kit/src/post-gesture-stability.ts:152">
P2: The loop caches surface metadata by object identity even though capture values are not required to be immutable or unique. If a provider reuses and updates a capture object, later polls compare stale signature/backend data and can report a false settle or no-effect result; read the surface for each observation instead of caching it.</violation>
</file>

<file name="packages/capture-kit/src/post-gesture-stability.test.ts">

<violation number="1" location="packages/capture-kit/src/post-gesture-stability.test.ts:72">
P3: The loop's other timeout shape is untested: every case here passes `needsBaselineDistrust: false` and never triggers a rebase, so nothing exercises the branch where the budget expires after a quiet pair already agreed — the loop then returns a bare `{ value }` with no `unsettled` outcome (post-gesture-stability.ts: the `if (lastPairAgreed) return { value }` path, which comments say is reached when a rebase or distrust verdict keeps polling on an at-rest surface). This is exactly the branch the PR claims to preserve ('budget can expire on a surface that is already at rest'); a regression would silently change when the agent sees the stabilization warning. Add a test with `needsBaselineDistrust: true` where captures keep agreeing on the unchanged baseline past the 1.5s cap, and assert `postGestureOutcome` stays undefined while the stabilization_timeout warning fires.</violation>
</file>

<file name="test/integration/provider-scenarios/scroll-movement-observation.test.ts">

<violation number="1" location="test/integration/provider-scenarios/scroll-movement-observation.test.ts:106">
P3: `scrollEntry()` returns `{ x, y, x2, y2 }`, but `honoredScrollSwipeMidpoint` (packages/contracts/src/scroll-command.ts:122-133) reads `x1`/`y1`/`x2`/`y2`, so the swipe midpoint is always `undefined` here. The comment's claim that midpoint (201, 437) sits inside CONTAINER never materializes, and in the no-progress test the `containerHoldsSwipe` gate in decideEdgeVerdict is silently skipped — the refusal passes for the wrong reason (no container-position verification). Use `x1`/`y1` to match the leaf contract.</violation>
</file>

Tip: instead of fixing issues one by one fix them all with cubic

Re-trigger cubic

via: 'src/daemon/interaction/internal/interaction-touch-response.ts#buildInteractionResponseData',
},
targetReadiness: DIRECT_IOS_SINGLE_QUERY_READINESS,
outcomeObservation: DIRECT_IOS_OUTCOME_NOT_OBSERVED_GAP,

@cubic-dev-ai cubic-dev-ai Bot Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This shared waiver is inaccurate for the maestro-non-hittable-fallback path: fill can execute the fallback while carrying --verify or --settle, and its response can include that observation. Scope the direct-selector waiver separately or update this path’s contract to model the runtime fill case.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/contracts/src/interaction-guarantees.ts, line 604:

<comment>This shared waiver is inaccurate for the `maestro-non-hittable-fallback` path: `fill` can execute the fallback while carrying `--verify` or `--settle`, and its response can include that observation. Scope the direct-selector waiver separately or update this path’s contract to model the runtime fill case.</comment>

<file context>
@@ -558,6 +601,7 @@ export const INTERACTION_DISPATCH_PATHS: Record<InteractionPathId, InteractionPa
         via: 'src/daemon/interaction/internal/interaction-touch-response.ts#buildInteractionResponseData',
       },
       targetReadiness: DIRECT_IOS_SINGLE_QUERY_READINESS,
+      outcomeObservation: DIRECT_IOS_OUTCOME_NOT_OBSERVED_GAP,
     },
   },
</file context>
Fix with cubic

Comment on lines +152 to +157
let surfaceCache: CapturedSurface<T, S> | undefined;

const surfaceOf = (value: T): CapturedSurface<T, S> => {
if (surfaceCache?.value === value) return surfaceCache;
surfaceCache = { value, ...hooks.readSurface(value) };
return surfaceCache;

@cubic-dev-ai cubic-dev-ai Bot Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: The loop caches surface metadata by object identity even though capture values are not required to be immutable or unique. If a provider reuses and updates a capture object, later polls compare stale signature/backend data and can report a false settle or no-effect result; read the surface for each observation instead of caching it.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/capture-kit/src/post-gesture-stability.ts, line 152:

<comment>The loop caches surface metadata by object identity even though capture values are not required to be immutable or unique. If a provider reuses and updates a capture object, later polls compare stale signature/backend data and can report a false settle or no-effect result; read the surface for each observation instead of caching it.</comment>

<file context>
@@ -116,38 +126,52 @@ export function decidePostGestureStabilityVerdict<S extends readonly unknown[]>(
-  // the deadline can expire on a surface that is already at rest.
+  // the budget can expire on a surface that is already at rest.
   let lastPairAgreed = false;
+  let surfaceCache: CapturedSurface<T, S> | undefined;
+
+  const surfaceOf = (value: T): CapturedSurface<T, S> => {
</file context>
Suggested change
let surfaceCache: CapturedSurface<T, S> | undefined;
const surfaceOf = (value: T): CapturedSurface<T, S> => {
if (surfaceCache?.value === value) return surfaceCache;
surfaceCache = { value, ...hooks.readSurface(value) };
return surfaceCache;
const surfaceOf = (value: T): CapturedSurface<T, S> => ({
value,
...hooks.readSurface(value),
});
Fix with cubic


const outcome = await runPostGestureStabilityLoop({
pending: PENDING,
needsBaselineDistrust: false,

@cubic-dev-ai cubic-dev-ai Bot Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The loop's other timeout shape is untested: every case here passes needsBaselineDistrust: false and never triggers a rebase, so nothing exercises the branch where the budget expires after a quiet pair already agreed — the loop then returns a bare { value } with no unsettled outcome (post-gesture-stability.ts: the if (lastPairAgreed) return { value } path, which comments say is reached when a rebase or distrust verdict keeps polling on an at-rest surface). This is exactly the branch the PR claims to preserve ('budget can expire on a surface that is already at rest'); a regression would silently change when the agent sees the stabilization warning. Add a test with needsBaselineDistrust: true where captures keep agreeing on the unchanged baseline past the 1.5s cap, and assert postGestureOutcome stays undefined while the stabilization_timeout warning fires.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/capture-kit/src/post-gesture-stability.test.ts, line 72:

<comment>The loop's other timeout shape is untested: every case here passes `needsBaselineDistrust: false` and never triggers a rebase, so nothing exercises the branch where the budget expires after a quiet pair already agreed — the loop then returns a bare `{ value }` with no `unsettled` outcome (post-gesture-stability.ts: the `if (lastPairAgreed) return { value }` path, which comments say is reached when a rebase or distrust verdict keeps polling on an at-rest surface). This is exactly the branch the PR claims to preserve ('budget can expire on a surface that is already at rest'); a regression would silently change when the agent sees the stabilization warning. Add a test with `needsBaselineDistrust: true` where captures keep agreeing on the unchanged baseline past the 1.5s cap, and assert `postGestureOutcome` stays undefined while the stabilization_timeout warning fires.</comment>

<file context>
@@ -0,0 +1,120 @@
+
+  const outcome = await runPostGestureStabilityLoop({
+    pending: PENDING,
+    needsBaselineDistrust: false,
+    hooks: hooksFor(clock, [{ signature: ['a'], costMs: 0 }, late]),
+    clock,
</file context>
Fix with cubic

command: 'ios.runner.scroll',
deviceId: DEVICE_ID,
platform: 'apple',
result: { x: 201, y: 487, x2: 201, y2: 387 },

@cubic-dev-ai cubic-dev-ai Bot Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: scrollEntry() returns { x, y, x2, y2 }, but honoredScrollSwipeMidpoint (packages/contracts/src/scroll-command.ts:122-133) reads x1/y1/x2/y2, so the swipe midpoint is always undefined here. The comment's claim that midpoint (201, 437) sits inside CONTAINER never materializes, and in the no-progress test the containerHoldsSwipe gate in decideEdgeVerdict is silently skipped — the refusal passes for the wrong reason (no container-position verification). Use x1/y1 to match the leaf contract.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At test/integration/provider-scenarios/scroll-movement-observation.test.ts, line 106:

<comment>`scrollEntry()` returns `{ x, y, x2, y2 }`, but `honoredScrollSwipeMidpoint` (packages/contracts/src/scroll-command.ts:122-133) reads `x1`/`y1`/`x2`/`y2`, so the swipe midpoint is always `undefined` here. The comment's claim that midpoint (201, 437) sits inside CONTAINER never materializes, and in the no-progress test the `containerHoldsSwipe` gate in decideEdgeVerdict is silently skipped — the refusal passes for the wrong reason (no container-position verification). Use `x1`/`y1` to match the leaf contract.</comment>

<file context>
@@ -0,0 +1,233 @@
+    command: 'ios.runner.scroll',
+    deviceId: DEVICE_ID,
+    platform: 'apple',
+    result: { x: 201, y: 487, x2: 201, y2: 387 },
+  };
+}
</file context>
Suggested change
result: { x: 201, y: 487, x2: 201, y2: 387 },
result: { x1: 201, y1: 487, x2: 201, y2: 387 },
Fix with cubic

const DIRECT_IOS_OUTCOME_NOT_OBSERVED_GAP: GuaranteeEnforcement = {
kind: 'waived',
reason:
"gap: this replay-only route observes no outcome beyond the runner's own report; --verify/--settle do not apply here (see the inapplicable cells on this row), and the shared ambiguous-failure corroboration only reconsiders a thrown error, never a successful dispatch.",

@cubic-dev-ai cubic-dev-ai Bot Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: This waiver claims the direct route "observes no outcome beyond the runner's own report", but the success path finalizes through finalizeTouchInteraction with scheduleInteractionOutcomeRetry defaulting to true, so it sets the same deferred markers (pending-outcome retry, post-gesture stabilization) that TAP_OUTCOME_NOT_OBSERVED_GAP documents as "judged by the next capture, never in this response". Scope the claim to the response ("never in its own response") and name the deferred markers as the sibling does, so the two waivers that the comment says share the same reasoning stay equally precise.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At packages/contracts/src/interaction-guarantees.ts, line 192:

<comment>This waiver claims the direct route "observes no outcome beyond the runner's own report", but the success path finalizes through finalizeTouchInteraction with scheduleInteractionOutcomeRetry defaulting to true, so it sets the same deferred markers (pending-outcome retry, post-gesture stabilization) that TAP_OUTCOME_NOT_OBSERVED_GAP documents as "judged by the next capture, never in this response". Scope the claim to the response ("never in its own response") and name the deferred markers as the sibling does, so the two waivers that the comment says share the same reasoning stay equally precise.</comment>

<file context>
@@ -165,6 +180,19 @@ const DIRECT_IOS_SINGLE_QUERY_READINESS: GuaranteeEnforcement = {
+const DIRECT_IOS_OUTCOME_NOT_OBSERVED_GAP: GuaranteeEnforcement = {
+  kind: 'waived',
+  reason:
+    "gap: this replay-only route observes no outcome beyond the runner's own report; --verify/--settle do not apply here (see the inapplicable cells on this row), and the shared ambiguous-failure corroboration only reconsiders a thrown error, never a successful dispatch.",
+  trackingIssue: GAPS_UMBRELLA_ISSUE,
+};
</file context>
Suggested change
"gap: this replay-only route observes no outcome beyond the runner's own report; --verify/--settle do not apply here (see the inapplicable cells on this row), and the shared ambiguous-failure corroboration only reconsiders a thrown error, never a successful dispatch.",
"gap: this replay-only route observes no outcome in its own response beyond the runner's report — the deferred markers set after dispatch (pending outcome retry, post-gesture stabilization) are judged by the next capture, never here; --verify/--settle do not apply (see the inapplicable cells on this row), and the shared ambiguous-failure corroboration only reconsiders a thrown error, never a successful dispatch.",
Fix with cubic

@thymikee

thymikee commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

This PR is ready. The earlier review at 938a38c found it clean, and dfe9629 has the same patch, so that result still holds. #3081 does not touch these files, so I carried the review over without a fresh run against b9b410a.

Not blocking: the verdict in post-gesture-stability.ts calls surfaceOf(latest) before surfaceOf(previousValue), and the one-slot cache evicts the previous entry, so hooks.readSurface runs 2n-1 times for n captures instead of n, which means each post-gesture poll builds the full-tree signature twice. Keeping the previous surface in a let previousSurface updated at the end of each verdict, or reading previousValue first, would fix it. You can take or leave this.

Smoke Tests on dfe9629 is still running and has not failed. The diff overlaps its route, since iOS and Android smoke run gestures and scrolls through post-gesture-stability.ts and scroll-movement.ts. This route broke smoke at 66eaec5, and both smoke jobs passed at 938a38c. A regression is not expected, but a red result would need a look before anyone blames a flake.

I did not run the tests locally. Merge waits on Smoke Tests finishing green. I know of no conflicts.

@thymikee

thymikee commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

iOS Smoke Tests red on dfe9629383 (run 36823829712): not caused by this branch. I reran the failed job once (attempt 2).

What happened

The E2E daemon was healthy. A second daemon took its daemon.json away. That daemon was an orphan from the Preflight iOS runner step's prepare. When it idle-reaped, its shutdown deleted the live daemon's daemon.json. After that, every client found daemon.lock held by a live daemon (pid 45692) with no daemon.json. That gives daemon_startup_failed with retainedLockProcess: true and hasInfo: false.

Evidence (from the ios-artifacts artifact)

  1. The step before get text succeeded. replay clear-state-launch-url.yaml --maestro finished ok at 06:28:11.502Z (events.ndjson, request 18c0c87075cef86c, 7549 ms). Every snapshot inside it was ok. The get text client started about 150 ms later and found no daemon.json.
  2. The only shutdown record in daemon.log comes from a different daemon:
    {"ts":"2026-10-01T06:28:10.109Z","phase":"ios_runner_session_detach_skipped","session":"daemon","command":"daemon","data":{"sessionId":"5FEB61C5-…:50565:1790835745178","reason":"runner_process_dead",…}}
    
    • Runner session :50565:1790835745178 was created at 06:22:25. It is the runner from the preflight prepare (sessions/cwd_a87d3c22dc05c4d7_ios/runner.log, port 50565, xcodebuild 06:22:32).
    • The E2E daemon's runner session is …:50971:1790836017209 (06:26:57). All 34 sessionId mentions in the E2E request logs are this one. Its open logged lease_absent and started its own xcodebuild at 06:27:02.
    • So the daemon that shut down never served the E2E. It also never served the Settings replay step, which started its own runner at 06:23:41.
  3. The timing matches idle reap. The prepare request finished at 06:23:10.031Z. DAEMON_IDLE_REAP_DEFAULT_MS is 300 s, and ios.yml does not override it. 06:23:10.031 + 300 s = 06:28:10.031. The detach line is 78 ms later, after closeDaemonServers returned on a server with no connections.
  4. That shutdown path deletes metadata it does not own. In src/daemon/server/daemon-runtime.ts, shutdown() calls removeInfo(infoPath). removeInfo (server-lifecycle.ts:65) does an unconditional unlinkSync. releaseDaemonLock does check existing.pid !== process.pid. So the orphan removed pid 45692's daemon.json and left its lock, which is exactly hasInfo: false, hasLock: true.
  5. pid 45692 was still alive and working after the failure. It appended to request log a73a5e63ef2b3506.ndjson at 06:30:11.587Z (AX-bridge simctl spawn … snapshot-bridge serve exit, durationMs: 167783). The runner it owns kept logging AGENT_DEVICE_RUNNER_IDLE_KEEPALIVE until 06:31:26. There was no Daemon error:, no unhandled rejection, no fatal diagnostic and no stall. Each new spawn only printed Daemon lock is held by another process; exiting.
  6. The branch cannot reach this. The diff touches post-gesture-stability.ts, scroll-movement.ts, interaction-guarantees.ts and their tests. None of them touch daemon lifecycle, metadata or idle reap. The scenario up to the failure (open, wait, alert, get, settings clear-app-state, maestro launchApp/assertVisible) has no post-gesture stability or scroll-movement call.
  7. Main and the previous head did not hit this. Main 36823707501 and 938a38c260 36770653287 are green, and their daemon.log has no foreign detach/shutdown record. In this run, the slower preflight (about 34 s from clean:daemon to prepare request start) and the slower E2E (30.6 s runner health check on open, 15 s wait miss before the deep-link alert) put the orphan's reap exactly in the E2E window.

Not proven from the artifacts

The artifacts do not show how the prepare daemon lost its daemon.json/lock ownership while staying alive. The E2E daemon's publication truncates daemon.log and removes that history. The defect that turns an orphan into a red job exists on main independent of this PR: an exiting daemon's removeInfo must check that daemon.json names its own pid, the same way releaseDaemonLock checks the lock. I recommend a separate main-side fix for that, not a change in this PR.

The post-gesture-stability.ts cache note (2n−1 readSurface calls) is still open as a non-blocking item. This comment does not change code.

@thymikee

thymikee commented Oct 1, 2026 •

Copy link
Copy Markdown
Member Author

[claude-fable-5-1] responding on behalf of @thymikee

iOS Smoke red on dfe9629383 (run 36823829712, attempt 2): not caused by this branch. The Simulator AX bridge never served a capture in this CI session. The depth-frontier assertion is a bridge-health check, so it fails when the bridge is down. Evidence is from the run's ios-artifacts.

Mechanism

  • assertSimulatorBridgeSnapshot (test/integration/ios-simulator-e2e/live-snapshot-depth-frontier.ts:128) requires snapshotQuality.backend === undefined, so the snapshot must come from the AX bridge. This snapshot came from XCTest ("backend":"tree", warning Simulator AX snapshot unavailable (circuit-disabled); used XCTest for this app generation.).
  • The scenario runs only open --relaunch → wait → wait → snapshot --depth 1. No gesture or scroll runs, so post-gesture-stability.ts and scroll-movement.ts (the only production files this PR changes) are not on the path.
  • The circuit opened on the fresh app generation (pid 34231, after the relaunch) because the bridge failed to connect. Nothing from an earlier gesture can carry over:
    22a355a3 (open)  07:18:50.083 ios.snapshot-source.acquire 79   error: transport-failure: application-server-unavailable
    22a355a3 (open)  07:18:55.328 ios.snapshot-source.acquire 5002 error: timeout: bridge-connect-deadline
    d1e8b775 (wait)  07:19:01.317 ios_snapshot_route_fallback reason: bridge-connect-deadline  generation: 34231:...[7577]...07:18:47
    b7fa2953 (wait)  07:19:02.116 ios_snapshot_route_fallback reason: circuit-disabled
    83481fc0 (snap)  07:19:02.637 ios_snapshot_route_fallback reason: circuit-disabled -> snapshot_capture backend: xctest
    
  • The bridge was already failing at the session's first request, before any gesture ran:
    4f8304ce (open)  07:15:33.820 ios.snapshot-source.acquire 6019 error: timeout: bridge-request-deadline
    9d19dfd2 (wait)  07:15:40.339 ios.snapshot-source.acquire 5514 error: timeout: bridge-connect-deadline
    9d19dfd2         07:16:42.532 exec_command xcrun simctl spawn … snapshot-bridge serve  exit after 67706 ms
    
  • The failing response's snapshotDiagnostics.backends is {"xctest":59}: zero bridge captures in the whole session.

Comparison

  • Previous head 938a38c260 (run 36770653287): green, and this assertion passed.
  • Main b9b410a6f7 (run 36823707501): green.
  • Same head, attempt 1: it failed on the daemon-metadata issue (Daemon shutdown deletes another daemon's daemon.json #3087) and did not reach this scenario. A grep of its artifact finds 1 circuit-disabled, 3 application-server-unavailable and 1 bridge-preparation-pending.

Action: I ran gh run rerun 36823829712 --failed once. No code change, so I did not take the optional previousSurface cache note in this round.

@thymikee
thymikee merged commit 9e54064 into main Oct 1, 2026
19 of 21 checks passed
@thymikee
thymikee deleted the feat/observation-engine-loops branch October 1, 2026 09:05
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-01 09:06 UTC

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-for-human Valid work that needs human implementation, judgment, or maintainer merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant