Skip to content

fix(ios): fill succeeds when app removes input after last character - #3161

Merged
thymikee merged 4 commits into
callstack:mainfrom
okwasniewski:oskar/ios-fill-input-removed-after-submit
Oct 4, 2026
Merged

thymikee merged 4 commits into
callstack:mainfrom
okwasniewski:oskar/ios-fill-input-removed-after-submit

Conversation

@okwasniewski

@okwasniewski okwasniewski commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

fill into an auto-submitting code field failed with XCTEST_RECORDED_FAILURE (reading frame on the removed input) and invalidated the runner.

The runner (Swift only) now:

  • binds the input's type and identifier at the first resolve. resolveTextEntryElement refuses any other input at every post, warmup read-back, verify poll and repair.
  • before the last post, a refused resolve returns TEXT_INPUT_NOT_FOCUSED. After full delivery, the fill is unverified ok "typed" only on XCTest's no-match (com.apple.dt.xctest.ui-testing.error 10008); any other error returns TEXT_INPUT_COMMIT_NOT_OBSERVED.
  • refuses repair when the bound input has no identifier.
  • composes with fix(ios): report fill into a non-echoing field as unconfirmed #3171: its unconfirmed check only reads inputs the bound resolver accepted, and the removal verdict runs first.

Limits:

  • A successor with the same type and identifier is indistinguishable, including an empty one: an unnamed field replaced mid-fill by an unnamed focused field gets the remaining posts. That fill fails TEXT_ENTRY_MISMATCH without repair, but the successor keeps the text.
  • The penalized-channel route (runSynthesizedReplacementRoute) is not covered.
  • Proving no-match costs ~2 s.

5 files.

Validation

Tested commit: 0eef3bb (rebased on upstream/main 294dc7d, after #3171).

  • pnpm check:affected --run --base upstream/main passed.
  • 61 text-entry runner XCTests (this PR's and fix(ios): report fill into a non-echoing field as unconfirmed #3171's) pass on an iOS 27.0 sim.
  • Live CLI, iOS 27.0 sim, one runner process: OTP digit-count fill ok unconfirmed ("0 of 6 digits" to "6 of 6 digits"); auto-submit fill ok with INPUT_REMOVED_AFTER_DELIVERY, no repair; then a tap.
  • CI: all green, including iOS Smoke (87 runner XCTests) and macOS Smoke.

Review in cubic

Copilot AI balanced review requested due to automatic review settings October 3, 2026 12:24

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 5 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Identity collisions and unchecked target re-resolution can still redirect typing or repair into another input.

Review effort: Balanced
Findings: 3 High severity

Open (3)
What changed in this PR

Updates the Apple runner so filling an auto-submitting input can succeed without invalidating the session when the input disappears after delivery.

Changes:

  • Uses snapshots for frame reads and captures input identity.
  • Skips verification and repair when the delivery input is detected as removed.
  • Adds auto-submit fixtures and regression tests.
File Description
apple/​runner/​AgentDeviceRunner/​AgentDeviceRunnerUITests/​UnitTests/​RunnerTests+TextTypingTests.swift Tests completed delivery and early removal.
apple/​runner/​AgentDeviceRunner/​AgentDeviceRunnerUITests/​UnitTests/​RunnerTests+TextEntryPolicyTests.swift Tests identity-based removal detection.
apple/​runner/​AgentDeviceRunner/​AgentDeviceRunnerUITests/​RunnerTests+TextTyping.swift Adds the post-delivery identity check.
apple/​runner/​AgentDeviceRunner/​AgentDeviceRunnerUITests/​RunnerTests+TextEntry.swift Captures snapshot-based input identity and frame.
apple/​runner/​AgentDeviceRunner/​AgentDeviceRunner/​AgentDeviceRunnerApp.m Adds an auto-submit replacement-input fixture.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@thymikee

thymikee commented Oct 3, 2026

Copy link
Copy Markdown
Member

I reviewed fe0461a. The fix is not ready to merge yet: the identity check does not cover every place the runner re-resolves the input, and the changed device path has no live run on this head.

The only live CLI run was on the pre-rebase head. The rebase re-applied deliveredTo and the post-plan withElement onto main's synthesized type plan, which is a different route (https://github.com/callstack/agent-device/blob/fe0461a/apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/RunnerTests+TextTyping.swift#L400). The reported run does not show that it reached the new branch, because runner.log has no AGENT_DEVICE_RUNNER_TEXT_ENTRY_INPUT_REMOVED_AFTER_DELIVERY. The fixture XCTests call executeTypeCommand directly and skip the daemon fill route, so nothing shows the changed path working on the head being merged. After the thread fixes, please run a live fill from the CLI through the daemon on the final head, against an auto-submit code field on an iOS simulator. The run should show the response ok with message "typed", runner.log containing AGENT_DEVICE_RUNNER_TEXT_ENTRY_INPUT_REMOVED_AFTER_DELIVERY and no REPAIR_TEXT_ENTRY line, and a second command in the same runner process with no restart. Please also run a mid-delivery replacement case that returns a typed failure.

Could this be simpler? The try? snapshot() change alone seems to fix the reported XCTEST_RECORDED_FAILURE: with no frame read, a vanished target already falls through verify as unverified. The identity type seems to exist only to stop index-bound queries from re-binding to a successor input. Could that check live in one place, resolveTextEntryElement? Every post, the warmup read-back, the verify polling and the repair all call it, so it would cover every site the open threads name. The rule would be that a resolution whose snapshot identity differs from the bound identity is refused. Before the last post, refused means .notFocused. After full delivery, it means unverified success, but only when the snapshot error is a proven no-match and nothing else is at that point. When the bound identity has an empty identifier, uniqueness cannot be proven, so the runner should refuse to repair instead of guessing. What would have to change first is that XCUIElement exposes no per-instance id, so element type plus identifier is unique only when the identifier is non-empty, and an unidentified successor must fail closed instead of comparing equal.

Not blocking, and fine to take or leave: the removed-after-delivery outcome returns plain ok "typed" with no response-level marker, so callers cannot tell it from a verified fill (an unverified marker or warning through the shared response builder could help); the reported state (input removed after the last character, no successor) is covered only by the predicate table, so a fixture variant that removes the field at length 6 without a successor would close that gap; and when the XCTest channel is penalized, fill takes runSynthesizedReplacementRoute, which never reaches the new guard, so either narrow the PR body claim or cover that route too.

The open review threads on this PR still apply: #3161 (comment) (replacement after warmup), #3161 (comment) (unnamed field identity), #3161 (comment) (transient error as removed), and #3161 (comment) (verify and repair identity). The three Copilot threads (r4173162982, r4173163019, r4173163055) repeat those same mechanisms, and they should resolve once the fix lands.

CI is still pending: Smoke Tests, Repo Guards and Coverage were queued or running with no failure logs. The iOS Smoke jobs run the runner text-entry code this diff changes, so a later failure there cannot be cleared as unrelated. There are no conflicts.

Before merge, please make the identity check hold at every resolve inside resolveTextEntryElement, count an input as removed only on a proven no-match, and then run the live CLI fill on the final head as described. I ran no XCTests and had no range-diff against the pre-rebase patch. I also did not confirm which XCTest error domain or code snapshot() throws for no-match versus a transient AX failure, or trace what the penalized-channel route does when the field vanishes.

Copilot AI balanced review requested due to automatic review settings October 3, 2026 16:09
@okwasniewski
okwasniewski force-pushed the oskar/ios-fill-input-removed-after-submit branch from fe0461a to e030d42 Compare October 3, 2026 16:09

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@okwasniewski

okwasniewski commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor Author

[claude-opus-5-5] responding on behalf of Oskar

Reworked as suggested and rebased on main. Head: e030d42.

Changes

  • resolveTextEntryElement owns the identity check. The first resolve binds type plus identifier; after that, any other input is refused at every post, warmup read-back, verify poll and repair. A bound target no longer falls back to the focused-keyboard route.
  • Refused before the last post: TEXT_INPUT_NOT_FOCUSED. After full delivery: unverified ok only when XCTest proves no-match, meaning snapshot() throws com.apple.dt.xctest.ui-testing.error 10008 and nothing at the point carries the identity. Otherwise TEXT_INPUT_COMMIT_NOT_OBSERVED.
  • Error codes measured on an iOS 27.0 sim: removed input 10008, multiple matches 10006, app not running 10001. A no-match snapshot() takes about 2 s, so candidates are gated by exists and the 2 s is paid once, in the removal check.
  • Empty bound identifier: repair is refused (REPAIR_REFUSED reason=unidentified-input), and the mismatch is reported.

Can the frame change alone fix it? It fixes the reported failure, but not the successor case. Fixture XCTests, iOS 27.0:

  • main: removal at length 6 with no successor records Failed to get matching snapshot: No matches found at the frame read. With a successor, main repairs into it: REPAIR_TEXT_ENTRY expectedLength=6 observedLength=0, typed after repair, next field value 123456.
  • frame fix only: no successor gives ok typed. With a successor it still repairs into the next field. Early replacement types 23456 into the successor and then repairs it.
  • this head: both success cases are ok typed with the next field empty. Early replacement and unnamed successor return typed failures.

Validation on e030d42

  • pnpm check:affected --run --base upstream/main: passed. macOS runner build passed.
  • 50 text-entry runner XCTests (built with AGENT_DEVICE_XCUITEST_INCLUDE_UNIT_TESTS=1): 50 passed, 0 failed.
  • Live CLI through the daemon, iOS 27.0 sim, SwiftUI app whose code field navigates to "Signed in" at 6 digits:
    fill 'id=code-field' 123456 -> success true ("Filled 6 chars")
    AGENT_DEVICE_RUNNER_TEXT_ENTRY_INPUT_REMOVED_AFTER_DELIVERY
    ...phase=verify durationMs=2066.4 ... phase=total durationMs=2755.8
    AGENT_DEVICE_RUNNER_COMMAND_COMPLETED command=type ... ok=1
    AGENT_DEVICE_RUNNER_COMMAND_COMPLETED command=tap ... ok=1   (Sign out, same runner)
    
    The early-navigation toggle sends the app to a focused "Name" field after the first digit:
    fill 'id=code-field' 123456                -> TEXT_INPUT_NOT_FOCUSED, screen shows name=[]
    fill 'id=code-field' 123456 --delay-ms 150 -> TEXT_INPUT_NOT_FOCUSED, screen shows name=[]
    
    A final normal fill also succeeded. Across all of it: one runner pid, one testCommand start (no restart), 0 REPAIR_TEXT_ENTRY lines. The CLI shows the daemon's fill message, not the runner's "typed"; the XCTests assert data.message == "typed".

CI on the old head

  • Coverage: session-device-claims.test.ts failed with ENOTEMPTY in its afterEach rmSync cleanup. That is a teardown race, not this diff. It passes on main CI (9a245d0, be3c810) and 3/3 locally.
  • iOS Smoke: all runner XCTests passed. The scenario then failed at open --relaunch --launch-url ... with a daemon_request_timeout (90 s), before any fill. fix/daemon-review-hardening (run 37130487744) failed at the same step the same way.

Left open

  • No response-level unverified marker. It would cross the Swift payload, runner contract, apple interactor and daemon builder, and secure fields need the same marker. Better as a follow-up.
  • The penalized-channel route (runSynthesizedReplacementRoute) never resolves an element, so the guard does not reach it. By reading, its commit wait polls by point and would end in TEXT_INPUT_COMMIT_NOT_OBSERVED. That is unverified, and the PR body no longer claims it.
  • A successor with the same type and non-empty identifier, or text in flight inside one burst post, cannot be told apart.

CI on e030d42: all 12 checks green. The iOS Smoke targeted lane ran 83 runner XCTests, all 5 new fill tests among them, with 0 failures. The macOS host lane ran 281 with 0 failures.

@thymikee

thymikee commented Oct 3, 2026

Copy link
Copy Markdown
Member

The earlier findings from fe0461a are fixed at e030d42, and I found no new blocking problems in the code. Posts now resolve through the bound resolver, repair is refused for an empty identifier, and a failed probe no longer reads as a removed input.

The iOS Smoke Tests are still queued with no logs. They run the runner text-entry code this PR changes (RunnerTests+TextEntry and TextTyping), so they must pass before merge. I know of no conflicts.

One open point remains: isBoundInput compares type and identifier. If an unnamed code field is replaced by an unnamed, focused successor mid-delivery, the remaining posts type into the successor and verify reports TEXT_ENTRY_MISMATCH. No test covers this case. Please state this limit next to the same-identifier limit in the PR body. You could also refuse focus and point re-resolution after the first post when the bound identity has no identifier.

On the open threads, the unnamed successor in posts and the same point from the other reviewer still apply, and both match the note above. These threads are fixed at e030d42, so you can resolve them: bound resolution at every post, probe error mapping, deliveredTo comparison removed, posts and warmup read-back, and verify polls and repair.

The live CLI run and the 50 passing runner XCTests are author-reported, and I ran no XCTests. I did not confirm that snapshot() throws code 10008 for a removed input, and I did not measure the added per-post cost on paced per-character plans. The penalized replacement route stays outside the guard, as you note.

Copilot AI balanced review requested due to automatic review settings October 3, 2026 21:17

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@okwasniewski

Copy link
Copy Markdown
Contributor Author

[claude-opus-5-5] responding on behalf of Oskar

Follow-up on the e030d42 review. Head: 3dbb8d9, a test-only commit.

Unnamed successor mid-delivery: documented as a limit, not fixed. I tried the suggested refusal (after the first post, refuse point and focus re-resolution when the identifier is empty) on an iOS 27.0 sim. It never fires. A coordinate fill binds an index-bound element (app.textFields.allElementsBoundByIndex), and after the replacement TextField (Element at index 0) itself resolves to the successor:

UNNAMEDPROBE resolved-by=element query=Attributes: TextField, ..., value: 23456, Keyboard Focused
AGENT_DEVICE_RUNNER_TEXT_ENTRY_REPAIR_REFUSED reason=unidentified-input
UNNAMEDPROBE ok=false code=TEXT_ENTRY_MISMATCH message=... expected "123456", observed "23456"

The outcome is the same with and without the refusal, so I reverted it. The refusal would only add a new NOT_FOCUSED failure when an ordinary unnamed field's index shifts. Closing the gap needs a per-instance element identity, which public XCTest does not expose. The new testFillFailsTypedWhenAnIndistinguishableInputReplacesItMidDelivery pins the current outcome: a typed TEXT_ENTRY_MISMATCH, no repair, and the successor holding 23456. The limit is in the PR body next to the same-identifier one.

Validation on 3dbb8d9

  • pnpm check:affected --run --base upstream/main: passed.
  • 51 text-entry runner XCTests (built with AGENT_DEVICE_XCUITEST_INCLUDE_UNIT_TESTS=1), iOS 27.0 sim: 51 passed.
  • Live CLI, iOS 27.0 sim, with an unnamed TextField("Note") added to the SwiftUI test app:
    • Default fill into the unnamed field (warmup plan): 30/30 ok, value read back.
    • --delay-ms 80 into the unnamed field: 26/31 ok. All 5 failures returned no text input found at the provided coordinates to clear in the unbound focus route, before any text-entry code ran. The runner tapped (201, 470), the field's position with the keyboard down, while the keyboard was up and the field sat at y=323. That points to the snapshot catching a layout mid-move, not to the bound resolver.
    • The same sequence on upstream main: unnamed fills 20/20 and 20/20 ok. The auto-submit fill 'id=code-field' 123456 failed 10/10 with XCTEST_RECORDED_FAILURE.
    • On this head the auto-submit fill passed 12/13. The one failure was TEXT_ENTRY_MISMATCH with the field still empty, right after the unnamed field's text keyboard had been up. It did not recur on retry.
  • CI: iOS Smoke passed, and its targeted runner lane ran 84 XCTests (the new test among them) with 0 failures. CI (Coverage, Integration, Repo Guards), Linux and Android passed. macOS Smoke is still queued.

Threads: replied on r4173162842 and r4173162982 with the above. All 7 threads are resolved.

@okwasniewski

Copy link
Copy Markdown
Contributor Author

[claude-opus-5-5] responding on behalf of Oskar

I checked the 5/31 paced unnamed-field failures from my last comment. It is not a regression in this diff: upstream main fails the same way, and the failing command never reaches code this PR changes.

Setup: one fresh iOS 27.0 sim. The upstream main runner and the 3dbb8d9 runner alternated in batches of 10 (main, PR, main, PR, ...), 40 fills per side. Each iteration ran the same sequence: open --relaunch, snapshot -i, fill @e4 'hello note', snapshot -i, fill @e4 'paced unnamed' --delay-ms 80. @e4 is an unnamed SwiftUI TextField("Note"). Both builds carried the same logging patch on the runner's focus route (not part of the PR).

Results, paced fill ok:

batch main PR
1 9/10 5/10
2 10/10 10/10
3 10/10 10/10
4 10/10 10/10
total 39/40 35/40

Fisher's exact test p=0.20. All 6 failures fell in the first two batches after the sim booted; the later 60 iterations had none.

Where it fails: every failure on both sides is no text input found at the provided coordinates to clear from focusTextInputForTextEntry. That function is unchanged here, and it runs before typeTextReliably. The runner log has only phase=focus for those commands, with no initial-resolve, so nothing is bound and the bound resolver never runs. Main and PR log the same line:

ABPROBE focus-entry point=(201,470) keyboardVisible=true textInputAt=nil fields={{16, 240}, {370, 34}} {{16, 306}, {370, 34}}
ABPROBE focus-exit textInputAt=nil stabilized=nil focusConfirmed=false readiness=nil element=nil

The daemon's snapshot gave the field's position with the keyboard down (y=470). By the time of the paced fill, the keyboard was up and the layout had moved the field to y=306-340, so nothing was at the requested point.

Focus-fallback hypothesis: ruled out. The bound-target change only affects posts inside typeTextReliably, and a refused post there returns TEXT_INPUT_NOT_FOCUSED. These failures are a different error from an earlier step, with no bound target. The one timing difference: the preceding fill is about 130 ms slower on this head (total 1190 ms vs 1061 ms; verify 328 ms vs 271 ms) from the identity snapshots. That moves where the next snapshot -i lands in the keyboard animation, but does not change any decision.

My earlier "main 20/20" was a smaller sample; with 40 per side, main also hits it (1/40).

@thymikee

thymikee commented Oct 3, 2026

Copy link
Copy Markdown
Member

The earlier findings from e030d42 are fixed at 3dbb8d9: posts, warmup read-back, and verify and repair now all resolve through the bound text-entry lookup, and the probe maps failures to unavailable instead of "gone". I found no remaining code problems. Not blocking, and you can take or leave it: the new test in https://github.com/callstack/agent-device/blob/3dbb8d9/apple/runner/AgentDeviceRunner/AgentDeviceRunnerUITests/UnitTests/RunnerTests+TextTypingTests.swift#L281 asserts that the successor holds "23456", which pins the known wrong-target side effect as expected output, so you could drop that assertion or say in the test name that it pins the current limit. I did not run the XCTests or the live CLI runs, so the passing runner tests and the auto-submit and paced numbers on the PR are your reports, and I have not confirmed that snapshot() on a removed element returns the no-match error code the probe maps. No conflicts are known. The Smoke Tests job is still in progress, and it covers the runner text-entry route this PR changes, so it needs to finish green before merge. On the other threads, the P1 on indistinguishable inputs (#3161 (comment)) and the Copilot thread on the same limit (#3161 (comment)) are resolved by the PR body note and the new test, and the P1 on bound resolution (#3161 (comment)), the P2 on the unavailable probe (#3161 (comment)), the P3 on the deliveredTo comparison (#3161 (comment)), and the two Copilot threads on bound resolution (#3161 (comment) and #3161 (comment)) are fixed at this commit, so none of the threads still apply and you can resolve them all.

Reading frame of an input the app removed after delivery recorded an XCTest failure, which the runner converts to XCTEST_RECORDED_FAILURE and the daemon treats as session-fatal. Bind targets through a throwing snapshot and skip verify/repair once the delivered input is gone, so it never verifies or repairs into the next screen's input.
Bind the input's type and identifier at the first resolve and check it in
resolveTextEntryElement, which every post, warmup read-back, verify poll and
repair call. Count the input as removed only on XCTest's no-match (10008).
Refuse repair when the bound input has no identifier.
Copilot AI balanced review requested due to automatic review settings October 4, 2026 07:36
@okwasniewski
okwasniewski force-pushed the oskar/ios-fill-input-removed-after-submit branch from 3dbb8d9 to 0eef3bb Compare October 4, 2026 07:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@thymikee

thymikee commented Oct 4, 2026

Copy link
Copy Markdown
Member

Thanks for the update. The code looks good at 0eef3bb, but the PR is missing live proof on the rebased route.

The earlier findings from 3dbb8d9 are fixed. The text-entry element is now resolved only through the bound identity on every post, read-back, verify and repair step. A removed input now maps to "no match" only for ui-testing error 10008, and any other error stays "unavailable". The app-wide focused-typing fallback and the repair path both refuse an unidentified identity. The two cubic threads on the indistinguishable-input limit and the Copilot thread on the same limit are resolved as benign: public XCTest cannot close that limit, a test pins it, and the PR body documents it. These threads are fixed at this head and can be resolved: #3161 (comment), #3161 (comment), #3161 (comment), #3161 (comment), #3161 (comment). These three are benign by design and can be resolved too: #3161 (comment) and #3161 (comment).

The live CLI runs (auto-submit fill id=code-field 123456 and the paced unnamed fills) ran on 3dbb8d9, before the rebase onto #3171. That rebase merged this PR's early return into the verify function that #3171 changed, so those runs do not cover the code at 0eef3bb. I did not run the runner XCTests or any device fill myself, and I did not confirm on device that snapshot() on a removed element throws error 10008. Please show a green Smoke Tests iOS runner XCTest lane on 0eef3bb with 0 failures. It must include testFillIntoAutoSubmittingDigitCountFieldReportsNoEvidenceFromItsSuccessor, testFillSucceedsWhenAppReplacesInputAfterLastCharacter and testFillSucceedsWhenAppRemovesInputAfterLastCharacter. Even better, one live fill into an auto-submitting code field on 0eef3bb that returns ok with message "typed", no verification block, and no XCTEST_RECORDED_FAILURE.

Smoke Tests is still running, and its iOS lane covers this exact route (executeTypeCommand, typeTextReliably, verifyTextEntryWithRepairIfNeeded), so a failure there would count against this PR. I know of no conflicts. Before merge, Smoke Tests must finish green on 0eef3bb with the iOS runner text-entry XCTests passing on the rebased route.

@okwasniewski

Copy link
Copy Markdown
Contributor Author

[claude-opus-5-5] responding on behalf of Oskar

Rebased onto upstream/main 294dc7d (after #3171). Head: 0eef3bb.

Conflicts. The only textual conflict was in the fixture app. The fixture now keeps #3171's --agent-device-text-entry-digit-count-value field class together with this PR's --agent-device-text-entry-unnamed-input. RunnerTests+TextTyping.swift merged cleanly. verifyTextEntryWithRepairIfNeeded now runs:

  1. This PR's removed-after-delivery verdict, first.
  2. fix(ios): report fill into a non-echoing field as unconfirmed #3171's settledUnconfirmedTextEntry, before repair and after a repair.
  3. The unnamed-input repair refusal.

#3171 and the bound identity. I kept both identities and did not merge them.

  • They compare different things on purpose. The binding (TextEntryInputIdentity, type plus identifier) must hold across every post while the field may still be moving with the keyboard: the A/B run saw it move from y=453 to y=306. So it leaves the frame out.
  • fix(ios): report fill into a non-echoing field as unconfirmed #3171's isSameElement uses the frame for unnamed fields, but only within its short settle window.
  • They cannot disagree in the success direction. settledUnconfirmedTextEntry reads through resolveTextEntryElement, so every observation it compares already passed the bound check. isSameElement can only add a stricter frame check on top, which fails closed.
  • The relationship is noted in the doc comment on TextEntryInputIdentity.

Can the unconfirmed path report success for a replaced input?

  • An input with a different identifier is never observed, so it cannot produce unconfirmed evidence. Removal takes the existing verdict: unverified ok "typed" without a verification field. New test: testFillIntoAutoSubmittingDigitCountFieldReportsNoEvidenceFromItsSuccessor.
  • An unnamed successor in the same place stays the documented limit. Its value echoes the delivered digits or equals the baseline, so the existing tests still get a typed failure.

Maintainer nit. Taken. The test is renamed to testFillPinsKnownLimitWhenAnIndistinguishableInputReplacesItMidDelivery. Its comment now says the 23456 assertion records the wrong-target side effect and should change once a per-instance identity exists.

Validation on 0eef3bb

  • pnpm check:affected --run --base upstream/main: passed.
  • 61 text-entry runner XCTests, this PR's and fix(ios): report fill into a non-echoing field as unconfirmed #3171's, built with AGENT_DEVICE_XCUITEST_INCLUDE_UNIT_TESTS=1 on an iOS 27.0 sim: 61 passed.
  • Live CLI, iOS 27.0 sim, one runner process (pid 98252, a single testCommand start):
    fill 'id=otp-field' 123456  -> success, verification "unconfirmed", before "0 of 6 digits", after "6 of 6 digits"
      AGENT_DEVICE_RUNNER_TEXT_ENTRY_UNCONFIRMED expectedLength=6 observedLength=13 repaired=0
    fill 'id=code-field' 123456 -> success, "Signed in" screen
      AGENT_DEVICE_RUNNER_TEXT_ENTRY_INPUT_REMOVED_AFTER_DELIVERY
    click 'label="Sign out"'    -> ok, same runner
    
    No REPAIR_TEXT_ENTRY lines.
  • CI: CI, Linux, Android, macOS Smoke and iOS Smoke all passed. The iOS targeted lane ran 87 runner XCTests with 0 failures, including both PRs' fill tests.
  • No new review threads. All 7 are resolved, and the PR is mergeable.

@thymikee

thymikee commented Oct 4, 2026

Copy link
Copy Markdown
Member

Thanks, this closes the evidence gap at 0eef3bb. The live CLI runs on the rebased head cover both paths: fill 'id=otp-field' 123456 returns the unconfirmed verdict, and fill 'id=code-field' 123456 reaches the "Signed in" screen through the removed-after-delivery path, with no repair lines and one runner process. CI is green, including iOS Smoke with the runner fill tests. The code verdict from my last review stands, so this is ready for a human review.

@thymikee thymikee added the ready-for-human Valid work that needs human implementation, judgment, or maintainer merge label Oct 4, 2026
@thymikee
thymikee merged commit ce62bee into callstack:main Oct 4, 2026
14 checks passed
pvedula7 added a commit to pvedula7/agent-device that referenced this pull request Oct 4, 2026
When the XCTest channel is penalized, iOS `fill` taps the daemon's point and
types through private synthesis. That route named the field only by the
point, but focusing can move the layout (keyboard avoidance, a bottom sheet
extending above the keyboard). The commit wait then re-read whatever sat at
the point: a fill that landed failed with TEXT_INPUT_COMMIT_NOT_OBSERVED, and
`fill ""` cleared the neighbouring field and reported success.

The input under the point is now found before the tap, and the target is
bound to its identity, the way callstack#3161 binds an element-route target. Reads and
clears take that input's handle while it still carries the identity, and
otherwise only the focused input or the app-wide identifier match that
carries it; a target bound before its tap never re-reads the point. The
lookup and the handle check run inside the text-input probe's issue
containment, so a read that cannot answer fails the fill closed instead of
ending the runner.

Finding the input took up to 2 s more on a React Native bottom sheet, so the
replacement's focus allowance rises from 2 s to 4 s.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-for-human Valid work that needs human implementation, judgment, or maintainer merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants