Skip to content

feat: add typed point inspection - #2999

Open
csark0812 wants to merge 22 commits into
callstack:mainfrom
csark0812:chris/agent/inspect-point
Open

csark0812 wants to merge 22 commits into
callstack:mainfrom
csark0812:chris/agent/inspect-point

Conversation

@csark0812

@csark0812 csark0812 commented Sep 27, 2026 •

Copy link
Copy Markdown

Summary

  • add capture.inspectPoint to the typed client and inspect-point --x/--y to the CLI and MCP command surfaces
  • return bounded accessibility descriptors ordered smallest-to-largest, with an explicit no-element-at-point result for honest misses
  • advertise the capability only for iOS Simulator runtimes and propagate runner, target, and transport failures
  • expand the existing XCTest readText point implementation without weakening existing readText behavior

Validation

  • pnpm check:affected --run (all runnable checks passed)
  • provider-backed iOS scenario covers the public command, coordinate flags, transport, and descriptor response
  • package tarball, runtime conformance, docs coverage, Swift parse, daemon wire compatibility, and related tests passed

Live verification boundary

The repository's GitHub-authoritative iOS runner lanes and a live iOS Simulator proof are still required. This change does not claim live Messages, share-sheet, action-extension, or sticker-surface verification from the local run.

Review in cubic

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 52 files

Tip: instead of fixing issues one by one fix them all with cubic

Re-trigger cubic

Comment thread packages/command-registry/src/registry.ts Outdated
Comment thread packages/contracts/src/client-capture.ts Outdated
Comment thread packages/platform-android/src/runtime.ts Outdated
Comment thread packages/command-registry/src/flag-definitions-action.ts
Comment thread packages/platform-apple/src/runtime.test.ts Outdated
Comment thread packages/platform-apple/src/interactor.ts Outdated
Comment thread src/client/client-types.ts Outdated

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 8 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

Comment thread src/__tests__/cli-client-commands.test.ts Outdated
@thymikee

Copy link
Copy Markdown
Member

Reviewed at 471df4a. This needs code changes before merge, mainly to stop the new inspection route from changing existing get text behavior.

readText now always returns ok:true, even on a miss (https://github.com/callstack/agent-device/blob/471df4a/apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+CommandExecution.swift#L425). Before this PR, a miss returned ok:false with 'readText did not resolve text', and the TS side threw on it, so get text failed. Now a miss reaches readTextForNode as no-text-at-point, and it silently falls back to the snapshot's fallbackText, which is often empty for text fields, with only a warn diagnostic. So on iOS, get text on a text field or an unlabelled node whose live read misses now succeeds with an empty or stale string, where it used to fail loudly. This rides along with the new command, with no test and no CHANGELOG entry, and it contradicts the PR body. Can point inspection get its own runner command (say inspectPoint), or an opt-in request field that only interactor.inspectRunnerPoint sets, so the readText wire contract and error semantics stay the same for every existing caller? A runner-requests fixture for the new shape would help confirm the boundary.

readPointAt builds up to 24 descriptors before it resolves text, and each descriptor queries label, identifier, value, readableText (which reads label, identifier and value again), elementType, frame and isHittable (https://github.com/callstack/agent-device/blob/471df4a/apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+Interaction.swift#L337). The old readTextAt stopped at the first readable text, but it now delegates to readPointAt, so every existing get text live-read call pays the full walk, including the expensive isHittable query, across up to 24 elements. That adds latency to a path that never uses the descriptors, and it raises the timeout risk on deep trees. Keeping readTextAt's original early-return loop, and building descriptors only when inspection is actually requested, would keep the existing route at its old cost.

Not blocking, and can be taken or left: inspect-point's timeout policy differs from the sibling runner AX reads (registry.ts:1092), bindPointInspectionRuntime duplicates bindElementTextRuntime almost verbatim (selector-observation-runtime.ts:295), the daemon's point-inspection coordinate parsing diverges from the shared readPointPositionals helper (inspect-point-runtime.ts:9), a test proxy's spread loses its throwing stubs (cli-client-commands.test.ts:1252), an unrelated hunk drops six fallow-ignore suppressions whose comment still describes them (client-types.ts:36), inspect-point.ts is a pure re-export barrel over snapshot.ts (contradicting the barrel-only-at-package-boundary rule), and Interactor.inspectPoint carries an unforwarded surface option plus a role field that duplicates type, a typeof check where Number.isFinite fits better, and reused elementText reason codes instead of a platform-specific one (interactor-types.ts:378).

Is the layering here proportional to what this needs? The change touches the runner, platform-apple, contracts, daemon, and CLI/MCP, at 486 net production lines. A single dedicated runner command that reuses the Swift candidate filter and ordering, with its binding placed next to or inside the element-text runtime binding, would remove the readText overload, the copied bind helper, and the pass-through subpath and re-export file in one move. Positional x y CLI args, like press and focus already use, would also reuse readPointPositionals without new global flags. No ADR looks necessary here, but splitting the runner route out from readText first would make the rest of this land cleaner.

The one reported check is green, but it does not exercise the Swift runner, so it says nothing about the runner-side behavior change or the new query cost. Two live iOS Simulator runs with the rebuilt runner are still needed: agent-device inspect-point --x <x> --y <y> over a labelled control should return status inspected with non-empty elements ordered smallest-frame-first, and the same command over empty space should return no-element-at-point; separately, get text @ref on a text field whose value is longer than its label should still return the live value, and a runner miss should behave as it did on main, with daemon log timings for get text from before and after this change to show the added cost. None of the Swift runner or simulator behavior was run here, so the query-count latency estimate above is inferred from the XCUIElement call pattern, not measured, and the fallow unused-type gate and the MCP output schema against PointInspectionResult were not checked either.

Move point inspection off the readText runner command so existing get text semantics and cost are unchanged, then bring the live iOS Simulator evidence for both routes.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 2 files (changes from recent commits).

Requires human review: Auto-approval blocked because this review re-detected 2 unresolved issues already reported by Cubic.
Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

@thymikee

Copy link
Copy Markdown
Member

Reviewed at a55892e. This still has findings, and the delta did not touch the issue raised in the earlier review at 471df4a (#2999 (comment)).

The readText case in RunnerTests+CommandExecution.swift:425 still returns ok:true with text nil on a miss, so readRunnerTextAtPoint in packages/platform-apple/src/interactor.ts returns undefined instead of throwing. Before this PR, a miss answered ok:false 'readText did not resolve text' and the call threw. On iOS, get text on a node whose live read misses now succeeds silently with the snapshot fallback text, which can be empty or stale, and every existing caller sees this as an undocumented behavior change. The readText wire contract needs to stay ok:false on a miss for every caller it already has; give point inspection its own runner command, or an opt-in request field that only inspectRunnerPoint sets, and add a runner-requests fixture for the new shape.

In RunnerTests+Interaction.swift:334, readPointAt now tries privateAXPointInspection first on every simulator call, and that route takes the smallest-area element that has any text. That drops the old readTextAt rule where text inputs (prefersExpandedTextRead, including textInputCandidatesAt) win before smaller elements, and it never falls back to XCTest when the capture succeeds. pointInspectionText also casts with as? String, so a non-String AX value such as an NSNumber switch or slider value becomes nil, where privateAXPresentationString would describe it. On a simulator, get text over a text field with a smaller labelled child (placeholder, clear button, icon) can now return the child's label instead of the field's value, and nothing covers this with a test. The existing get text route needs to resolve text exactly as the pre-PR readTextAt did, with text inputs first; keep the private AX path out of readTextAt by gating it to the inspection command only, or apply the text-input-first ordering over the private tree, and reuse privateAXFields and privateAXPresentationString instead of the new string helpers.

In RunnerTests+Interaction.swift:391, privateAXPointInspection ignores response["truncated"] and the deep-extension pending/missed counts that privateAXSnapshotAcquisition uses (AXSnapshotFallback.swift:236-256). A capture capped at 5,000 nodes, or limited in depth, that omits the subtree under the point returns an empty element list, and that gets treated as a successful honest miss that skips the XCTest query. The 8 s deadline also does not bound the initial requestSnapshotFromClient call (RunnerAXSnapshotBridge.m:132), so the 'deadline-bounded' claim in the comment at line 330 does not hold. On a large tree, inspect-point and get text can report no element at a point that has one, and a wedged first AX call can still hold the main thread past the watchdog, which is the problem this PR exists to fix. An empty result should count as a miss only when the capture is complete: return nil and fall through when the capture is truncated or has pending or missed frontiers, reusing privateAXDepthLimited, and either correct the comment or run the call under the existing watchdog containment.

The comment change at RunnerTests+Interaction.swift:329 touches a device-facing simulator route, but the only evidence is Swift unit tests calling privateAXPointInspection(root:point:) directly with hand-built dictionaries. The one CI check is green, but it does not build or run the Swift runner, so it does not exercise either delta commit. Can this be validated with live iOS Simulator runs from the PR head: inspect-point <x> <y> over a system share sheet showing a bounded runtime and the sheet control rather than the app name, get text on a point with no text showing a failure rather than ok with empty text, get text on a filled text field showing the field's value, and a runner log line proving the private AX path was taken?

Isn't point inspection reusing privateAXSnapshotAcquisition's own completeness checks (truncated, privateAXDepthLimited) and privateAXFields the simpler path here, rather than parsing the raw tree a second time? Splitting point inspection into its own runner command would leave the readText route untouched and remove most of the findings above at once.

Get text keeping ok:false on a miss with its text-input-first resolution, followed by the live simulator runs above, is what needs to happen before this is ready to merge.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 30 files (changes from recent commits).

Requires human review: Auto-approval blocked because this review re-detected 1 unresolved issue already reported by Cubic.
Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 5 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

Comment thread scripts/integration-progress-model.ts Outdated

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 4 files (changes from recent commits).

Requires human review: Auto-approval blocked because this review re-detected 1 unresolved issue already reported by Cubic.
Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

Comment thread scripts/integration-progress-model.ts
@csark0812

Copy link
Copy Markdown
Author

I pushed eb2bef0 addressing the code-side concerns raised in this review. Point inspection remains an explicit inspectPoint runner mode; legacy readText requests still use the original readTextAt path and preserve its miss behavior. The inspection path consumes the contained AX snapshot route, rejects truncated/incomplete frontier captures, reuses the shared private AX field normalization (including non-string values), and prioritizes text-input values when choosing its text result. Added a regression for a smaller nested clear button plus a longer text-field value and numeric AX value normalization.\n\nValidation on this commit: pnpm test:unit passed (1,413 files / 11,482 tests, 1 skipped); pnpm typecheck, pnpm lint, pnpm format:check, pnpm build, and pnpm check:packaged-runner-swift passed. iOS Simulator XCTest build-for-testing succeeded, so the changed runner and XCTest source compile. I still cannot provide the requested live simulator trace here because CoreSimulatorService has no available runtimes; I am not claiming the live behavior/performance proof.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 3 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

@csark0812

Copy link
Copy Markdown
Author

Live follow-up on eb2bef0: on the booted iPhone 18 Pro (iOS 27.0), the public inspect-point --x 140 --y 80 returned inspected and the Agent Device Tester heading with frame {x:60,y:70,width:159.33,height:20.33} first. inspect-point --x 403 --y 875 returned no-element-at-point; note this coordinate is outside the 402×874 screen bounds, so it verifies the miss path but is not an in-bounds blank-surface proof.\n\nThe fixture loaded its real JavaScript surface from Metro, and a full-name field was populated with Alexandria Alexandra (snapshot confirms the value). Three get text @ref attempts timed out with typed MAIN_THREAD_TIMEOUT, including after closing/reopening the session and after blurring the field. Screenshot confirmed the simulator remained responsive. I cannot claim the requested live text-value oracle passed; this exposes an XCTest text-read timeout on the current simulator/fixture, and still needs a baseline comparison plus a successful field read. The session was closed and Metro stopped after the attempt.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 issues found across 1 file (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+Interaction.swift">

<violation number="1" location="apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+Interaction.swift:327">
P3: Every legacy `get text` call on an iOS Simulator now starts by capturing the full bounded AX tree (`privateAXPointInspection(app:x:y:)` calls `RunnerAXSnapshotBridge.snapshotTree` with maxNodes 5,000 and an 8-second deadline) even when the point resolves nothing. For misses and truncated captures — exactly the cases this guard falls through on — the code then also runs the legacy `app.descendants(matching: .any).allElementsBoundByIndex` query, so those calls acquire the tree twice and add the bridge capture's full latency on top of the path it was meant to replace. Consider only paying for the bridge capture when it can actually avoid the legacy query (e.g., skip the bounded read when the bridge is known-unavailable, or gate it on a cheap first check), or document the doubled cost for the common read-miss path.</violation>

<violation number="2" location="apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+Interaction.swift:327">
P3: `get text` output for the same coordinate can now differ between iOS Simulator and every other target: on simulators this guard resolves text from the private AX bridge (raw bridge frames and `pointInspectionReadableText`), while macOS/tvOS/physical iOS still use `readableText(for:)` over XCTest elements. The two backends can disagree about which element underlies a point (frame/containment), and the private path caps candidates at the first 24. This is the bounded-read trade-off the PR is deliberately making, but it changes the readText result surface for simulators; a provider-backed parity check asserting the same text on both paths (when both resolve) would confirm the claimed "preserve readText behavior" property rather than leaving it to live runs.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

// text fields, this uses the same text-input-first policy as that fallback.
// An incomplete capture never answers the request; it falls through to the
// legacy XCTest path so a capped tree cannot turn a real value into a miss.
if let inspection = privateAXPointInspection(app: app, x: x, y: y),

@cubic-dev-ai cubic-dev-ai Bot Sep 29, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: get text output for the same coordinate can now differ between iOS Simulator and every other target: on simulators this guard resolves text from the private AX bridge (raw bridge frames and pointInspectionReadableText), while macOS/tvOS/physical iOS still use readableText(for:) over XCTest elements. The two backends can disagree about which element underlies a point (frame/containment), and the private path caps candidates at the first 24. This is the bounded-read trade-off the PR is deliberately making, but it changes the readText result surface for simulators; a provider-backed parity check asserting the same text on both paths (when both resolve) would confirm the claimed "preserve readText behavior" property rather than leaving it to live runs.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+Interaction.swift, line 327:

<comment>`get text` output for the same coordinate can now differ between iOS Simulator and every other target: on simulators this guard resolves text from the private AX bridge (raw bridge frames and `pointInspectionReadableText`), while macOS/tvOS/physical iOS still use `readableText(for:)` over XCTest elements. The two backends can disagree about which element underlies a point (frame/containment), and the private path caps candidates at the first 24. This is the bounded-read trade-off the PR is deliberately making, but it changes the readText result surface for simulators; a provider-backed parity check asserting the same text on both paths (when both resolve) would confirm the claimed "preserve readText behavior" property rather than leaving it to live runs.</comment>

<file context>
@@ -318,6 +318,19 @@ extension RunnerTests {
+    // text fields, this uses the same text-input-first policy as that fallback.
+    // An incomplete capture never answers the request; it falls through to the
+    // legacy XCTest path so a capped tree cannot turn a real value into a miss.
+    if let inspection = privateAXPointInspection(app: app, x: x, y: y),
+      inspection.complete,
+      let text = inspection.text
</file context>
Fix with cubic

// text fields, this uses the same text-input-first policy as that fallback.
// An incomplete capture never answers the request; it falls through to the
// legacy XCTest path so a capped tree cannot turn a real value into a miss.
if let inspection = privateAXPointInspection(app: app, x: x, y: y),

@cubic-dev-ai cubic-dev-ai Bot Sep 29, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: Every legacy get text call on an iOS Simulator now starts by capturing the full bounded AX tree (privateAXPointInspection(app:x:y:) calls RunnerAXSnapshotBridge.snapshotTree with maxNodes 5,000 and an 8-second deadline) even when the point resolves nothing. For misses and truncated captures — exactly the cases this guard falls through on — the code then also runs the legacy app.descendants(matching: .any).allElementsBoundByIndex query, so those calls acquire the tree twice and add the bridge capture's full latency on top of the path it was meant to replace. Consider only paying for the bridge capture when it can actually avoid the legacy query (e.g., skip the bounded read when the bridge is known-unavailable, or gate it on a cheap first check), or document the doubled cost for the common read-miss path.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+Interaction.swift, line 327:

<comment>Every legacy `get text` call on an iOS Simulator now starts by capturing the full bounded AX tree (`privateAXPointInspection(app:x:y:)` calls `RunnerAXSnapshotBridge.snapshotTree` with maxNodes 5,000 and an 8-second deadline) even when the point resolves nothing. For misses and truncated captures — exactly the cases this guard falls through on — the code then also runs the legacy `app.descendants(matching: .any).allElementsBoundByIndex` query, so those calls acquire the tree twice and add the bridge capture's full latency on top of the path it was meant to replace. Consider only paying for the bridge capture when it can actually avoid the legacy query (e.g., skip the bounded read when the bridge is known-unavailable, or gate it on a cheap first check), or document the doubled cost for the common read-miss path.</comment>

<file context>
@@ -318,6 +318,19 @@ extension RunnerTests {
+    // text fields, this uses the same text-input-first policy as that fallback.
+    // An incomplete capture never answers the request; it falls through to the
+    // legacy XCTest path so a capped tree cannot turn a real value into a miss.
+    if let inspection = privateAXPointInspection(app: app, x: x, y: y),
+      inspection.complete,
+      let text = inspection.text
</file context>
Fix with cubic

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 2 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+Interaction.swift">

<violation number="1" location="apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+Interaction.swift:327">
P2: `readTextAt` now lets a complete bound AX capture answer the legacy `get text` request even when it resolves no text, skipping the live XCTest walk entirely. The comment directly above still promises the opposite — "a capped tree cannot turn a real value into a miss" — and this block now turns a bounded, textless capture into an authoritative miss. `complete` only means the 5_000-node/56-depth caps and the deep-extension bookkeeping were not hit; it says nothing about the snapshot matching the live hierarchy. Elements with missing/empty frames are dropped by the `!frame.isEmpty` guard, and this repo's own docs (`docs/adr/0004-ios-snapshot-backend-strategy.md`, `RunnerTests+AXSnapshotFallback.swift`) treat private AX as a fallback/recovery backend that can return sparse or divergent views. So `get text` can now return `ok:false "readText did not resolve text"` for a point that genuinely holds text — the exact miss-semantics change the earlier reviews asked to avoid. Prefer keeping the fallback when `inspection.text` is nil (only short-circuit with a proven value), or prove the miss with a lightweight live probe before treating it as authoritative.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Fix all with cubic | Re-trigger cubic

Comment on lines +327 to +331
if let inspection = privateAXPointInspection(app: app, x: x, y: y), inspection.complete {
// A complete miss is authoritative too. Falling through would repeat a
// full XCTest descendant walk after the bounded AX capture already
// established that no readable element contains the point.
return inspection.text

@cubic-dev-ai cubic-dev-ai Bot Sep 29, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: readTextAt now lets a complete bound AX capture answer the legacy get text request even when it resolves no text, skipping the live XCTest walk entirely. The comment directly above still promises the opposite — "a capped tree cannot turn a real value into a miss" — and this block now turns a bounded, textless capture into an authoritative miss. complete only means the 5_000-node/56-depth caps and the deep-extension bookkeeping were not hit; it says nothing about the snapshot matching the live hierarchy. Elements with missing/empty frames are dropped by the !frame.isEmpty guard, and this repo's own docs (docs/adr/0004-ios-snapshot-backend-strategy.md, RunnerTests+AXSnapshotFallback.swift) treat private AX as a fallback/recovery backend that can return sparse or divergent views. So get text can now return ok:false "readText did not resolve text" for a point that genuinely holds text — the exact miss-semantics change the earlier reviews asked to avoid. Prefer keeping the fallback when inspection.text is nil (only short-circuit with a proven value), or prove the miss with a lightweight live probe before treating it as authoritative.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+Interaction.swift, line 327:

<comment>`readTextAt` now lets a complete bound AX capture answer the legacy `get text` request even when it resolves no text, skipping the live XCTest walk entirely. The comment directly above still promises the opposite — "a capped tree cannot turn a real value into a miss" — and this block now turns a bounded, textless capture into an authoritative miss. `complete` only means the 5_000-node/56-depth caps and the deep-extension bookkeeping were not hit; it says nothing about the snapshot matching the live hierarchy. Elements with missing/empty frames are dropped by the `!frame.isEmpty` guard, and this repo's own docs (`docs/adr/0004-ios-snapshot-backend-strategy.md`, `RunnerTests+AXSnapshotFallback.swift`) treat private AX as a fallback/recovery backend that can return sparse or divergent views. So `get text` can now return `ok:false "readText did not resolve text"` for a point that genuinely holds text — the exact miss-semantics change the earlier reviews asked to avoid. Prefer keeping the fallback when `inspection.text` is nil (only short-circuit with a proven value), or prove the miss with a lightweight live probe before treating it as authoritative.</comment>

<file context>
@@ -324,11 +324,11 @@ extension RunnerTests {
-      let text = inspection.text
-    {
-      return text
+    if let inspection = privateAXPointInspection(app: app, x: x, y: y), inspection.complete {
+      // A complete miss is authoritative too. Falling through would repeat a
+      // full XCTest descendant walk after the bounded AX capture already
</file context>
Suggested change
if let inspection = privateAXPointInspection(app: app, x: x, y: y), inspection.complete {
// A complete miss is authoritative too. Falling through would repeat a
// full XCTest descendant walk after the bounded AX capture already
// established that no readable element contains the point.
return inspection.text
if let inspection = privateAXPointInspection(app: app, x: x, y: y),
inspection.complete,
let text = inspection.text
{
return text
}
Fix with cubic

@thymikee

Copy link
Copy Markdown
Member

Thanks for the update. Commit 612f054 fixes part of the earlier findings, but the runner still has two defects and no live evidence on this head. The single CI check is green, but it does not build the Swift runner: check:packaged-runner-swift only runs swiftc -parse, which cannot catch the compile failure below. No conflicts.

readableText(for:) now calls pointReadableText at RunnerTests+Interaction.swift#L523. That function is declared at line 505, inside the #if os(iOS) && targetEnvironment(simulator) block that closes at line 516. In any non-simulator build (physical iPhone, Apple TV, macOS) the symbol does not exist, so the runner will not compile and every Apple non-simulator session breaks at runner build. I read this from the #if placement and did not run a Swift build. Any helper that shared, unconditional code calls must be declared outside platform #if blocks. Please move pointReadableText, and the text-input type set it shares with privateAXPointInspection, above the #if. Then build the runner for a generic iOS device destination and for macOS.

readTextAt at line 323 now binds let candidates = ... before the ?? chain. Every legacy get text therefore pays the full app.descendants(matching: .any).allElementsBoundByIndex walk, including the text-field case that used to return early. The readPointAt XCTest fallback at line 356 has the same order. Could this explain the MAIN_THREAD_TIMEOUT seen in your live run on eb2bef0? I inferred that link and have no baseline run on the same fixture. The legacy readTextAt should keep its pre-PR order and cost. Evaluate the text-input candidates first, and build the descendants list only when that returns nil. Please apply the same order in the readPointAt fallback.

The changed runner routes are device-facing: the inspectPoint private-AX path, the contained bridge requests, and legacy readTextAt (line 345). The only live evidence is on eb2bef0, which is older than the Swift changes in c5f0dce, 27d95de and 612f054. It also shows a failing get text, and it has no share-sheet or in-bounds-miss run. On 612f054 against an iOS Simulator, please show these runs. (a) inspect-point over a system share sheet returns the sheet control within the command budget, with the AGENT_DEVICE_RUNNER_PRIVATE_AX line from runner.log. (b) get text on a filled text field returns its value with no MAIN_THREAD_TIMEOUT. (c) get text on an in-bounds point with no text returns the typed failure "readText did not resolve text". (d) inspect-point on an in-bounds blank area returns no-element-at-point. (e) A successful runner build for a physical-device destination.

Not blocking, take or leave: the new 2 s cap (RunnerAXSnapshotReadTimeout) and the IOS_SNAPSHOT_AX_CONTAINMENT_FAILED throw in RunnerAXSnapshotBridge.m#L470 also change the shared snapshot-recovery route, so a deep-extension request slower than 2 s now fails the whole snapshot. That has no CHANGELOG entry or live evidence, so document it or scope the cap to the inspection call. The truncated-capture error at CommandExecution.swift:431 says it "could not prove a miss" even when elements were found under the point, so a message about an incomplete capture would be more accurate.

Before merge, please move pointReadableText out of the simulator-only #if, restore the text-input-first order in readTextAt, and post the live simulator runs and device build on 612f054.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants