Repository navigation
fix(mobile): keep a message the host recorded but did not deliver in the phone chat, labelled not sent - #26140
brennanb2025 wants to merge 32 commits into
Conversation
Move the rejection words and the start-failure fact match from the desktop notices and the outbox's send disposition into src/shared, so the phone can word a not-sent row exactly as the desktop does. The desktop's output is unchanged.
… projection The shared projection listed rows the host recorded and then rejected after the whole conversation; only the desktop re-sorted them. Merge them into the conversation through the existing journal-order sort, so a reader that draws the list as it comes puts them where the host recorded them.
Record at send time which rows are already shown as not sent. Only those can never be the send's echo: a not-sent row that appears later is the send's own, so it settles the send and retires its bubble. The send's boundary is the newest row past any shown as not sent, and they never count as earlier copies of its text, so a same-text resend is not left as a phantom bubble or a 'Delivery unconfirmed' banner.
…top does The phone hid a message the host recorded and then did not deliver, and handed its text back to the composer with a banner. The text then lived only on the phone that sent it: other viewers saw nothing, and the sender's text was gone once the composer changed. The phone now projects these rows in place, with the host's reason in a muted line under the message (the same words the desktop uses). Its own send answers such a rejection like a held send: no banner and no hand-back, because the row holds the text. A retained id that replays its own recorded rejection is sent again once under a fresh id, as a replay a Stop took back already is. A send a Stop withdrew and one kept as a card behave as before.
…rded rejection The phone matched a structured send's bubble, photo preview and unconfirmed hold to the transcript by text and by the row that was newest when it was sent. Once rows the host recorded and rejected stay in the chat, that broke: the host moves a rejected row to where it was rejected, so a later image send's bubble never retired (R1P-1), and another send's not-sent row of the same text cleared this send's bubble early (R1P-3). The host records a send, and draws its row, under the id the phone sent it with. The structured send now returns that id with its outcome, and the bubble, the preview and the hold are settled by that row alone, delivered or not sent, wherever it sits; a hold is also settled by its own queued card. The text guards added for not-sent rows are gone, and the terminal lane's text matching is main's again. A recorded, rejected send now answers its own outcome, 'recorded-unsent', instead of reusing 'queued' (a card above the composer). The bridge echoes it under its id, so a not-sent photo keeps its local preview until its row arrives.
An attributed user message (an orchestration task, for one) is laid out as the agent's, and the not-sent line was drawn only for the person's own bubble, so a task the host recorded and rejected read as delivered on the phone. Draw the line for any user-role row shown as not sent, as the desktop does.
…re-sort Merging not-sent rows into the conversation before the stop rows were placed let one split a run of stopped sends, which gave the desktop a second 'stopped' row. Place the stop rows first, then merge the few not-sent rows in at their journal places in one pass; a row with no journal place keeps its neighbour. The desktop's order is base's again (0 differences in 20,000 random journals), and a chat with a not-sent row no longer sorts the whole list on every update.
Both chat clients now draw a recorded, rejected message in place; only the outline's ticks leave it out.
Round 1 settled a structured send's bubble and its held send only when a row with the send's id was drawn. Two cases never draw one: - a send the host took as a queued card that then went out under a fresh id; - a send whose own row the projection hides, because a same-text copy was sent after it was rejected, or because it was kept as a card. In the first the phone said "Delivery unconfirmed" for a message in the chat; in the second its bubble stayed for good. A send made under id I is now settled when the loaded journal holds a submission recorded under I, in any state, or one whose queuedMessageId is I (a held send also settles on its own card, as before). This is the rule the phone already uses to release an ack-lost id, now one shared index beside it (mobileStructuredSendsInJournal); the release keeps its narrower settled-only, same-body reading through its parameters. A photo binds to the send's row when that row is drawn, and otherwise retires unbound. The drafts hook takes the structured journal's submissions from the controller, as it takes the queued cards.
A photo send's answer can land after the frame carrying its record renders but before that frame's effects run. With another bubble pending, the retirement then saw the record, not yet the bound preview, and retired the photo unbound, so its delivered row showed the host path instead of the thumbnail. A photo bubble whose own row is drawn now waits to bind to it; only one whose row the chat hides retires unbound on the journal's record.
… the drafts The type pinned only that some array is passed; the structured-session mock now returns submissions and the drafts arguments must carry that array.
A local the harness assigned was narrowed to null at the reads, so the test file failed the mobile tests typecheck; read the hook through an object.
The controller check no longer rewrites the mock's existing cast line (the gate counts it as new), and the photo-race harness reads the hook through a plain object instead of writing a ref during render.
…reads The anti-slop naming rule rejects "shape"; the helper lists the rows in the order they are drawn.
|
The phone hid the host’s copy of a message that had been recorded but could not be delivered. Depending on rejection timing, it could disappear, return to the message box with a banner, or leave an unlabelled temporary copy looking sent. This PR keeps the recorded message in the chat with one muted not-sent reason and leaves the message box clear. The fix settles temporary copies by each message’s identity in the host’s record, including queued cards later sent to the agent, rather than matching text. It covers attributed messages, preserves desktop stopped-message grouping, and binds each photo preview to its own row. Deferred work keeps the earlier failed copy after a same-text resend and covers separate desktop/command changes after #25959; low-impact row-object recreation remains deferred. The final sync below resolves the earlier host-record lookup naming/module issue. On exact head Earlier phone merge head: The merge touched the phone message renderer and shared projection. Main also added phone visuals, subagent groups and image fingerprint changes; the rejection-label properties, shared reason rules, pending-copy identity matcher and photo-preview association were preserved. The screenshots embedded in the PR body prove QA head Earlier phone merge validation passed 86 root tests in 11 explicit files and 365 mobile tests in 28 explicit files, including message rendering, captioned-photo races, identity settlement, queued cards, stopped messages and the desktop sender. Queued web, node and CLI typechecks passed. Mobile typecheck completed with only the eight established missing generated modules, whose importers match main; no final-head native app, live provider, Android or SSH UI run was performed. All 37 checks completed on that earlier phone head That earlier phone head’s pre-push checks finished 21/22: the sole local failure is five native import-cycle warnings reproduced at the same locations against the merged main source. Both full and changed-code type-aware scans, React Doctor and localization checks passed. There is no new CI failure to classify against main. The PR remains draft. No ready, PR merge or auto-merge action was taken. A subsequent copy update and conflict-resolution sync are reported below. The isolated rig used private actual host/provider configuration paths with hooks off. A real-settings hash change paused the driver; metadata aligned it with the live Orca restart before our rig restart, and the coordinator explicitly authorized a new baseline and resume after isolation verification. Its writer remains unknown; no real configuration was copied, edited or restored. Original native evidence is retained; the invalid reconstructed high-resolution file is excluded. Final review of the merge (fresh reviewer, read-only, on Shared-copy update on final head The fresh review fixes also give automatic retry and failed-turn diagnostics the correct technical audience, preserve the refusal verdict about the requested response for Stop, and reuse the recorded cause for command refusal advice. Cleared chats point to the current conversation, unsupported commands have no pointless retry, and an active response asks the reader to wait or stop it. Failed Stops name the agent and ask the reader to check the chat; a refusal of the requested response does not establish that the whole agent is idle. Returned cards carry the chat’s agent name. The six locale catalogs and renderer translations use the same meanings. Both fresh copy reviews identified issues now fixed here: the complete historical marker sentence is withheld at clause boundaries, final transport errors after retries get safe named copy, older internal “provider” sentences are normalized, failed Stops do not infer global idle from a requested-response refusal, and refused commands reuse their recorded cause and next step. The final non-retrying error says “Codex ran into a problem. Check the chat before trying again.” without claiming the message was unsent. A typed final server-capacity refusal retains its readable explanation. The new final-error kind uses the existing row fact field; older clients retain the complete host sentence for an unfamiliar kind, with no new stream message or request shape. At the coordinator’s request, the branch then merged main Final validation passed 560 root tests in 34 files and 335 phone tests in 25 files. Queued node, renderer and CLI typechecks passed; the phone test-file gate covered 955 files with four intentional exclusions and main’s unchanged 122-file baseline. CI exposed stale copy expectations and one missing test-runtime registration, all corrected with 148 affected root tests and one phone retry test passing. Correction head The desktop proof embedded in the PR body comes from actual typing and Send clicks on The copy rig used private home and agent configuration paths, hooks off and a mock keychain. Hashes for Codex hooks, Claude settings and the real Orca tree were unchanged; Codex configuration and Claude state hashes changed, classified by the coordinator as “writer unknown, laptop-wide churn.” No real configuration was edited or restored. The PR remains draft; no ready, PR merge or auto-merge action was taken. Final round (10-09): main sync and the last wording fixes, at
Deferred: about a dozen generic lines still say "the agent" instead of its name (for example "Too many messages were waiting for the agent…"); they are readable and carry no codes. A bare code without an underscore ( Verified vs not: the five native phone rows from Round 5 were checked again in code after the sync and behave the same (the code paths they use are unchanged); they were not re-run on the final head. The new plain lines for Grok and the other agents are covered by tests (built from constructed error text; for Codex also a recorded real disconnected-stream run), not by a live run or screenshot. Readiness checklist (final diff at |
📝 WalkthroughWalkthroughThe pull request adds client message IDs to structured sends and uses journal receipts to reconcile mobile drafts, image previews, queued cards, and unconfirmed sends. Rejected submissions remain in transcript order and render as unsent messages with agent-specific notices. Failure handling now filters technical diagnostics, distinguishes provider errors from write failures, and uses agent-specific localized wording across mobile, renderer, ACP, Claude, and Codex flows. Priority: ➖ Normal Merge Risk: 🔵 Low · up to The change keeps rejected messages in the phone chat as intended. The remaining issues are small wording problems: a capitalized "The agent" in a mid-sentence caption, and a Korean particle that is inconsistent with the rest of the file. Both can be fixed before or after merge. 🚥 Pre-merge checks | ✅ 4 | ❓ 1❌ Failed checks (1 inconclusive)✅ Passed checks (4 passed)Full details: Docstring CoverageExplanation Docstring coverage is 33.78% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 74 functions across 79 files. (66 skipped: 7 unsupported, 59 over the file limit.) ✨ Finishing Touches 💡 1📝 Generate docstrings 💡
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
- A subagent's sign-in row no longer hides the chat's own sign-in steps: only rows the conversation renders count as already explained (phone and desktop). - A failed ACP turn (Grok, OpenCode, OMP) names the agent and quotes its reason only when a person can read it; raw error text stays on the failure fact. - The readable-text check itself withholds explanations that say "provider", so quoted refusals, retry causes and request errors follow the same rule. - Main's sign-in tests use safe example detail and assert technical detail is withheld.
The startup row test from main still expected 'Provider diagnostic' text on the row. This PR withholds that kind of detail, so the test now checks plain detail is kept literally, technical detail is withheld, and the agent's sign-in step and resend advice stay.
An ACP turn failure hid the agent's raw reason from the line but attached it nowhere, so 'Check the chat' pointed at nothing, and a later copy of a rate-limit reason replaced 'Grok usage limit reached.' with a generic line. The row now carries the reason as its Details frame and logs it once, and a usage-limit end keeps its lead whichever copy arrives first. The readable-text check also withholds 'providers'.
Main's new prompt-failure test now expects the named line with the agent's words in Details. Readable text also withholds upper-case snake codes with a digit (CAPTURE_PROVIDER_400, ERR_42), while digit-free setting names such as ANTHROPIC_API_KEY stay. The failure-row identity moves into AcpTurnFailureSource so the translator stays under its line limit.
Only upper-case snake codes ending in a number (CAPTURE_PROVIDER_400, ERR_42) are withheld, so next steps such as AWS_S3_BUCKET, R2_ACCESS_KEY and GPT_4O stay visible.
There was a problem hiding this comment.
Actionable comments posted: 2
ℹ️ Review info
⚙️ Run configuration
- Configuration used: Repository UI
- Review profile: CHILL
- Plan: Advanced
- Run ID:
b352e30c-263a-48de-aef6-32a0983f8937
📒 Files selected for processing (116)
config/scripts/structured-agent-session-start-failure-row-render.test.mjsconfig/scripts/vitest-sqlite-runtime-files.mjsmobile/src/session/MobileNativeChatMessage.test.tsmobile/src/session/MobileNativeChatMessage.tsxmobile/src/session/MobileNativeChatQueuedMessages.test.tsxmobile/src/session/MobileNativeChatView.unsent-notice.test.tsmobile/src/session/mobile-native-chat-merged-snapshot-parity-hooks.test.tsxmobile/src/session/mobile-native-chat-message-styles.tsmobile/src/session/mobile-native-chat-own-row-echo.test.tsmobile/src/session/mobile-native-chat-unsent-notices.test.tsmobile/src/session/mobile-structured-agent-session-contract.tsmobile/src/session/mobile-structured-agent-session-rpc.tsmobile/src/session/mobile-structured-agent-session-send.tsmobile/src/session/mobile-structured-composer-command.test.tsmobile/src/session/mobile-structured-composer-command.tsmobile/src/session/mobile-structured-queued-message-cards.test.tsmobile/src/session/mobile-structured-queued-message-cards.tsmobile/src/session/mobile-structured-send-delivery.test.tsmobile/src/session/mobile-structured-send-delivery.tsmobile/src/session/mobile-structured-send-new-action.test.tsmobile/src/session/mobile-structured-send-receipts.tsmobile/src/session/use-mobile-structured-agent-session-queued.test.tsxmobile/src/session/use-mobile-structured-agent-session-retry-rows.test.tsxmobile/src/session/use-mobile-structured-agent-session-send.test.tsxmobile/src/session/use-mobile-structured-agent-session.test.tsxmobile/src/session/use-mobile-structured-agent-session.tsmobile/src/session/use-mobile-structured-native-chat-send-bridge.tsmobile/src/session/use-mobile-structured-send-reached-host.tsmobile/src/session/use-mobile-structured-send-with-outcome.tssrc/main/acp/acp-structured-lane.tssrc/main/acp/acp-structured-prompt-failure.test.tssrc/main/acp/acp-structured-sign-in.test.tssrc/main/acp/acp-timeline-fixture.test-support.tssrc/main/acp/acp-timeline-generic.test.tssrc/main/acp/acp-timeline-translator.tssrc/main/acp/acp-timeline-turn-failures.test.tssrc/main/acp/acp-turn-failures.tssrc/main/claude/claude-api-retry-row.test.tssrc/main/claude/claude-command-turn.test.tssrc/main/codex/codex-app-server-connection.test.tssrc/main/codex/codex-app-server-request-error.tssrc/main/codex/codex-command-turn-claim.test.tssrc/main/codex/codex-journal-command-turn.tssrc/main/codex/codex-provider-diagnostic-copy.test.tssrc/main/codex/codex-provider-retry-row.test.tssrc/main/codex/codex-provider-retry-row.tssrc/main/codex/codex-structured-compaction-refusal.test.tssrc/main/codex/codex-structured-conversation-stop.test.tssrc/main/codex/codex-structured-dispatch-admission.test.tssrc/main/codex/codex-structured-journal-translation-streams.test.tssrc/main/codex/codex-structured-journal-translation.test.tssrc/main/codex/codex-structured-journal-translation.tssrc/main/codex/codex-structured-session-adapter.test.tssrc/main/codex/codex-structured-session-adapter.tssrc/main/codex/codex-structured-turn-end-settlement.test.tssrc/main/codex/codex-structured-turn-end-settlement.tssrc/main/codex/codex-structured-turn-open-wait.test.tssrc/main/codex/codex-structured-turn-open-wait.tssrc/main/codex/codex-structured-turn-start.tssrc/main/codex/codex-structured-turn-steer.test.tssrc/main/native-chat/agent-session-journal/journal-reducer-rejection-fact.test.tssrc/main/native-chat/agent-session-wire/structured-agent-session-codex-stop-row.test.tssrc/main/native-chat/agent-session-wire/structured-agent-session-command-turn.tssrc/main/native-chat/agent-session-wire/structured-agent-session-conversation-stop.test.tssrc/main/native-chat/agent-session-wire/structured-agent-session-queued-compact.test.tssrc/main/native-chat/agent-session-wire/structured-agent-session-send-preparation.test.tssrc/main/native-chat/agent-session-wire/structured-agent-session-stop-failure-copy.test.tssrc/main/native-chat/agent-session-wire/structured-agent-session-turns-cancel.tssrc/main/native-chat/agent-session-wire/structured-conversation-compaction.test.tssrc/main/pi/rpc-capture-replay.test.tssrc/main/runtime/structured-agent-session-codex-turn-end-settlement.test.tssrc/renderer/src/components/native-chat/NativeChatMessageList.provider-retry-runs.test.tsxsrc/renderer/src/components/native-chat/NativeChatNoticeRow.test.tsxsrc/renderer/src/components/native-chat/NativeChatQueuedMessageCard.failure-copy.test.tsxsrc/renderer/src/components/native-chat/NativeChatQueuedMessageCard.tsxsrc/renderer/src/components/native-chat/NativeChatQueuedMessageList.test.tsxsrc/renderer/src/components/native-chat/NativeChatStructuredSession.start-failure-notice.test.tsxsrc/renderer/src/components/native-chat/agent-session-command-refusal-words-text.tssrc/renderer/src/components/native-chat/agent-session-failure-words-text.test.tssrc/renderer/src/components/native-chat/agent-session-failure-words-text.tssrc/renderer/src/components/native-chat/agent-session-write-notice-text.tssrc/renderer/src/components/native-chat/structured-agent-session-delivery-notices.test.tssrc/renderer/src/components/native-chat/structured-agent-session-delivery-notices.tssrc/renderer/src/components/native-chat/structured-conversation-command-send.test.tssrc/renderer/src/components/native-chat/structured-conversation-command-send.tssrc/renderer/src/components/native-chat/use-structured-agent-session-command-send-gate.test.tsxsrc/renderer/src/components/native-chat/use-structured-agent-session-write-refusal.test.tsxsrc/renderer/src/components/native-chat/use-structured-agent-session.queued-gating.test.tsxsrc/renderer/src/i18n/en-runtime-required.jsonsrc/renderer/src/i18n/locales/en.jsonsrc/renderer/src/i18n/locales/es.jsonsrc/renderer/src/i18n/locales/fr.jsonsrc/renderer/src/i18n/locales/ja.jsonsrc/renderer/src/i18n/locales/ko.jsonsrc/renderer/src/i18n/locales/zh.jsonsrc/shared/agent-session-attachment-failure-words.tssrc/shared/agent-session-availability-sentences.tssrc/shared/agent-session-command-refusal-words.test.tssrc/shared/agent-session-command-refusal-words.tssrc/shared/agent-session-failure-copy.tssrc/shared/agent-session-failure-words.test.tssrc/shared/agent-session-failure-words.tssrc/shared/agent-session-failure.tssrc/shared/agent-session-provider-retry-words.tssrc/shared/agent-session-refusal-notice.test.tssrc/shared/agent-session-refusal-notice.tssrc/shared/agent-session-sign-in.test.tssrc/shared/agent-session-visible-failures.tssrc/shared/provider-diagnostic-person-text.test.tssrc/shared/provider-diagnostic-person-text.tssrc/shared/structured-agent-session-dispatch-rejection.test.tssrc/shared/structured-agent-session-dispatch-rejection.tssrc/shared/structured-agent-session-message-projection.tssrc/shared/structured-agent-session-rejection-words.test.tssrc/shared/structured-agent-session-rejection-words.tssrc/shared/structured-agent-session-start-failure-facts.ts
Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 6 remain after this review.
| const values = { agent: agentName ?? sayAgentSessionFailureTranslated('theAgent') } | ||
| return [ | ||
| sayAgentSessionFailureTranslated('writeFailed', values), | ||
| sayAgentSessionFailureTranslated( | ||
| card.waitsForAgent ? 'queueSendRetryWhenDone' : 'queueSendRetry', | ||
| values | ||
| ) | ||
| ].join(' ') |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Fix the capitalized fallback agent name in the queued-card retry caption.
Without an agent name, both captions put "The agent" from theAgent into the middle of the queueSendRetryWhenDone sentence. The rendered text is "Use Send to try again once The agent finishes."
src/renderer/src/components/native-chat/NativeChatQueuedMessageCard.tsx#L77-L84: use a fallback that suits the middle of a sentence, or a separate copy entry with no agent name. Then update the assertion atNativeChatQueuedMessageList.test.tsxLine 273.mobile/src/session/mobile-structured-queued-message-cards.ts#L79-L86: apply the same fix. Then update the assertion atmobile-structured-queued-message-cards.test.tsLine 145.
📍 Affects 2 files
src/renderer/src/components/native-chat/NativeChatQueuedMessageCard.tsx#L77-L84(this comment)mobile/src/session/mobile-structured-queued-message-cards.ts#L79-L86
| "queueSendRetryWhenDone": "{{agent}}가 끝나면 보내기로 다시 시도하세요.", | ||
| "commandStillWorking": "{{agent}}가 아직 작업 중입니다." |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Use the 이/가 particle form here as well.
The other Korean entries write {{agent}}이(가). These two entries write {{agent}}가. For agent names that end in a consonant, {{agent}}가 is ungrammatical. One example is the fallback "에이전트" if it is ever localized. Use the same 이(가) form as the rest of the file.
Proposed fix
- "queueSendRetryWhenDone": "{{agent}}가 끝나면 보내기로 다시 시도하세요.",
- "commandStillWorking": "{{agent}}가 아직 작업 중입니다."
+ "queueSendRetryWhenDone": "{{agent}}이(가) 끝나면 보내기로 다시 시도하세요.",
+ "commandStillWorking": "{{agent}}이(가) 아직 작업 중입니다."📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| "queueSendRetryWhenDone": "{{agent}}가 끝나면 보내기로 다시 시도하세요.", | |
| "commandStillWorking": "{{agent}}가 아직 작업 중입니다." | |
| "queueSendRetryWhenDone": "{{agent}}이(가) 끝나면 보내기로 다시 시도하세요.", | |
| "commandStillWorking": "{{agent}}이(가) 아직 작업 중입니다." |
ELI5
When you send a message from the Orca phone app and the computer running the agent records it but then can't hand it to the agent (the agent failed to start, Orca restarted before passing it on, the provider refused it), the phone hid the host's rejected row. Depending on when the rejection arrived, your message could disappear, return to the message box with a banner, or remain as a temporary copy that looked sent and gave no reason. The desktop already keeps such a message in the chat. Now the phone does too: the message stays where it was, with a small grey line under it saying it was not sent and why, and the text is not pushed back into the message box.
What Changed
The problem. Orca records every message you send on the machine running the agent (the host) before handing it over. If the hand-over then fails, the host marks the message rejected in that record. Since #24710 the desktop draws such a message in place, labelled not sent. The phone did not:
rejectedInPlace: false);Before / after on the phone.
The copy problem and change. The not-sent line could say that a “provider” refused the message and then quote an internal error code. The phone and desktop now name the chat's agent: “Codex couldn't receive this message. Send it again.” A content refusal retains the agent's readable explanation, with no blanket instruction to resend an unsupported message. Marker-shaped strings, protocol records and stack traces stay out of the visible explanation even when an older host labelled them for a person.
Codex previously treated every explicit rejection from its protocol as a content refusal. The host now recognises an explicit unwritten-send marker as a write failure, while ambiguous transport failures still remain unconfirmed. The same guard covers shortening-history failures, refused Stops, and retry explanations; automatic retries and their final error no longer show protocol codes or transport URLs. The final error says “Codex ran into a problem. Check the chat before trying again.” without claiming the message was unsent; a typed server-capacity refusal keeps its readable explanation. Older reason-only rows containing internal “provider” prose use the known agent’s neutral sentence, and embedded codes are rejected at clause boundaries while filenames remain readable. Unconfirmed shortening-history, Stop and answer lines name the agent and tell the reader to check the chat before repeating the action. A refused Stop reports that it did not stop, without claiming the agent is idle: the refusal may concern one response while another is running. Returned cards name the agent too. Refused commands use their recorded cause and next step: a cleared chat points to the current conversation, an unsupported command has no pointless retry advice, and a running response asks the reader to wait or stop it. Existing start-failure messages already derive their advice from the chat's state and keep raw diagnostics out.
Failure wording after the October 8 main sync. Every failure line follows one rule: plain words, the agent's name, and never "provider", an error code or a raw error. A next step appears only when it works. Four gaps found by the post-sync review are now closed:
codex login." The short caption stays only when the chat itself shows the sign-in row. This applies to the phone and the desktop: the desktop shows subagent rows only inside their own collapsible sections.provider_write_failed: write EPIPE. It now reads "OpenCode ran into a problem. Check the chat before trying again.", the same wording as Codex. A readable explanation is still quoted ("OpenCode ran into a problem: Upstream failed. Check the chat before trying again."). As on Codex's final error row, the agent's own text is kept under the row's Details disclosure on desktop and written to the app log, so "check the chat" points at something. When the agent ends the turn on a usage limit, the line keeps saying so ("Grok usage limit reached.") with the agent's text in Details, even if Grok's later copies of the reason arrive as plain errors or arrive before the limit end.CAPTURE_PROVIDER_400orERR_42: OpenCode's recorded "Internal error: CAPTURE_PROVIDER_400" now reads "OpenCode ran into a problem. Check the chat before trying again.", with the code under Details. Setting names without digits stay readable, because they are often the next step ("Set ANTHROPIC_API_KEY, then try again.").The mechanism.
use-mobile-structured-agent-session.ts): rejected messages are drawn in place, from the host's record, as the desktop does.MobileNativeChatMessage.tsx,mobile-native-chat-unsent-notices.ts): a muted line under any user message shown as not sent, worded from the host's typed failure by the same words the desktop uses.src/shared/structured-agent-session-rejection-words.ts,src/shared/structured-agent-session-start-failure-facts.ts): the rejection words and the "the start row already says why" rule moved out of desktop-only code so both clients say the same thing. The copy update also makes these shared failures name the agent and keeps technical details out of the inline sentence.src/shared/structured-agent-session-message-projection.ts): not-sent rows are placed at their position in the host's record in one pass, so the phone (which draws the list as given) shows them in order; the desktop's drawn output is byte-identical (checked on 30,000 random chats).mobile-structured-send-delivery.ts): a recorded, non-withdrawn rejection gets its own outcome, so nothing is handed back and no banner shows; the row holds the message. The final sync preserves main’s fresh action for every Send and its removal of the saved phone journal. Each action returns its fresh message ID to settle temporary text and photo rows; no saved ID or automatic resend is restored.use-mobile-native-chat-drafts.tsand the reconcile helpers): the phone shows a "pending" copy of what you sent until the host's copy arrives. That copy is now settled by the message's own id, once the host's record holds it in any state (delivered or not sent), instead of by matching text. Matching by text could leave a stray copy or a false "Delivery unconfirmed" once a not-sent copy of the same text was in the chat. The terminal-backed chat keeps its existing text matching.No new stream message or request shape. A final agent error uses a new row-only failure kind so it cannot be mistaken for a retry still running, an exited process or an unsent message; older readers retain the host’s complete sentence for an unfamiliar kind. A Stop refusal retains the optional verdict about the requested response; it does not describe the whole agent as idle. The copy update corrects the host's classification of an explicitly unwritten Codex send and changes the shared English failure sentences it publishes; older clients can still read the same failure kinds and legacy markers.
Why
The common pattern keeps a message the host recorded in the conversation and never hands its text back; only a message the host never recorded returns to the message box. Doing the same on the phone removes the last place such a message vanished or lived only in the composer.
Alternatives considered:
Differences from the common pattern, each labelled:
Linked Issue
N/A (maintainer). Rebuilds the phone half of #24918 on current main; #24918 is closed in favour of this PR and a follow-up.
Carried from #24918, re-reviewed here: the phone in-place rows and grey line, the shared words, ordering by the host's record, no hand-back for a recorded message, and the phone fixes found in #24918's review (a stray copy after resending the same text, a false "Delivery unconfirmed" for a send that landed, an image placeholder that never cleared). New here: settling pending copies by message id (replacing #24918's text-plus-snapshot matching), the separate send outcome, the line on another agent's rejected message, and keeping the desktop's stop rows unchanged.
Follow-up (PR B), after #25959 lands: a failed
/compactshown once instead of hidden, the muted desktop line, no rail tick for a not-sent message, and keeping the failed original when the same text is sent again.Visual Proof
Shared-copy proof, desktop on
8ed888640ef. Actual typing and Send clicks through an isolated hidden desktop renderer produced “Codex couldn't receive this message. Send it again.” beneath the retained message. A second Send retained the readable image refusal without telling the reader to resend it. The desktop and phone use this same shared English sentence. These two pictured sentences are unchanged by the later review fixes; the newer Stop, retry and command branches are covered by tests. This is desktop proof, not a new native phone run.The rig used private home and agent configuration paths, hooks off, an absolute stand-in executable and a mock keychain. Hashes for Codex hooks, Claude settings and the Orca directory were unchanged; Codex configuration and Claude state hashes changed, classified by the coordinator as “writer unknown, laptop-wide churn.” No real configuration was edited or restored. All recorded rig and driver processes exited.
The older native phone evidence below establishes retention, photo delivery state and layout on its stated head. Its three rejected-message images showed the superseded raw error line and have been removed; the desktop captures above demonstrate the shared replacement wording. Final-head native Round 5 proof follows; the older material remains historical.
Native iOS Round 5 on final head
c15c79bf109669676541ca63779630b0da167edf, October 8, 2026. The phone previously hid or misrepresented a message the host recorded but could not deliver. Actual native Sends now show one retained text or captioned-photo bubble with “Codex couldn't receive this message. Send it again.” beneath it, an empty message box and no banner or duplicate. The delivered control remains a normal bubble and reply. An owned headless iPhone 17 Pro simulator on iOS 26.5 ran the exact-head native Release app, paired to an isolated exact-head host with an absolute provider stand-in; these are full, unmodified native PNGs.Round 5 a: Recorded text rejection.
Round 5 b: Delivered control.
Round 5 c: Recorded captioned-photo rejection.
Round 5 d: Attributed long-reason stored layout fixture.
Round 5 e: Fresh same-text Send; earlier failure hidden (declared limitation).
The installed embedded bundle matched the built bundle by SHA-256. All UI input used the scoped Orca emulator; no desktop control, activation or visible window was used. The private host and provider environments, hooks-off settings and absolute stand-ins were verified, and the real Codex hooks and Claude settings hashes stayed unchanged. Owned simulator, host/profile/build resources and the clean child worktree were removed.
Coverage limits: the prepared folder’s agent chooser could not load its repository context, so these native cases ran in Floating Workspace. Reopening for the attributed fixture again showed the already documented stored-path photo fallback; no fresh-main reopen comparison was made to classify it. Photo proof in c is the actual Send-time observation. Actual other-agent delivery, real providers, Android, SSH UI and other failure timings remain untested; historical main-before screenshots below are retained.
Native iOS Round 4 on
c4328cd870d: all four requested cases verified, with the fixture and provenance limits below. The PR stays draft. The check used an owned iPhone 17 Pro simulator running iOS 26.5, paired to an isolated exact-head desktop host with a provider stand-in. All tested Sends and photo-picker actions were native Orca emulator input; these are native phone screenshots.80d45095d2fBefore: main, rejected text.
Before: main, rejected captioned photo.
Delivered control: actual native Send, accepted in the host record.
Loaded native chat.
The earlier control's lingering Working status came from an incomplete provider fixture and is excluded as accepted-delivery proof. The fresh corrected control reached accepted and completed, and the native capture shows its settled normal row without a reopen. Native baseline identity was checked from the installed embedded bundle and source map; all 18 included changed mobile/shared files match main, with no PR-only file included. Original full native PNGs were retained; uploads are full-screen JPEG derivatives. Headless screenshots of only the owned device were explicitly authorized; no desktop input or focus was used.
A hash-only audit briefly paused QA when the machine's real Claude settings changed. Metadata places that change just after the live Orca app restarted and before the isolated rig restart; the writer remains unknown. The coordinator authorized a new baseline and resume after every rig host/provider environment was verified private with hooks off. No real configuration was edited or restored.
The earlier attributed-row reopen capture also showed the photo using its stored-path fallback. No fresh-main reopen comparison was made to attribute that behavior to this PR; the native photo Send-time observation is recorded in the table above. A reconstructed higher-resolution transfer was invalid and is excluded from evidence.
Testing
Final head:
e94cf2d6b4d. This is the final sync with main (df922d37091, which includes main's typecheck fix #26781). The merge had no text conflicts. It adds three things: main's #26666 test updated to the plain named wording (see What Changed), the readable-text check now withholds upper-case codes with a digit (CAPTURE_PROVIDER_400,ERR_42) and keeps digit-free setting names (ANTHROPIC_API_KEY), and the failure row's names moved into one small type so main's Agent Client Protocol translator stays under its 300-line limit. One test of this PR's own (NativeChatQueuedMessageCard.failure-copy.test.tsx) was updated for main's new queued-card props. Local checks on the merged tree: 211 root test files (every test file this PR or main's 56 new commits changed in the agent-chat areas, plus all Agent Client Protocol suites) passed, 2,579 tests; phone tests for every touched phone file passed (25 files, 343 tests); queued node, web and CLI typechecks passed;orca-ci-checks --no-typecheckpassed 21/22, the only failure being the same five phone import-cycle warnings seen only with local phone packages installed. Lockfiles equal main's. CI on this head (run) passed every check that ran: static analysis and typecheck, all five test shards, relay integration, cross-version wire compatibility, both package jobs, the phone web bundle and both verify jobs. The skipped jobs (end-to-end, persistence, Git compatibility and similar) are path-gated and did not run for this change.Previous head
ff35ce35b9e: It fixes the two findings from the review ofe416d27c87f: a failed turn from an Agent Client Protocol agent hid its reason everywhere (the agent's text is now under Details and in the log, and a usage-limit line stays a usage-limit line), and the "provider" check missed the plural. The last commit only moves the new option two lines away from main's neighbouring edits so the branch stays mergeable without a main sync. Local checks: the Agent Client Protocol suites and the shared readable-text test (51 files, 443 tests) and the failure-words, sign-in, notice-row, delivery-notice, Pi replay and Codex retry-row suites (12 files, 218 tests) passed. The queued node typecheck passed.orca-ci-checks --no-typecheckpassed 21/22; the only failure is the same five phone import-cycle warnings as before. No phone file changed. CI on this head (run): the static-analysis job failed only on a type error that is on main itself.command-receipt-transaction.test.ts(#26653) passes one row to a writer callback that #26664 changed to take a list. The test shards were skipped behind it, andverifyfailed because that job failed. Known merge work for the final main sync: main's newacp-structured-prompt-failure.test.ts(#26666) expects an Agent Client Protocol agent's raw error as the whole line. This PR words it " ran into a problem…", so that test's expected lines change at the sync. The previous head merged with main fails it the same way.Previous head
e416d27c87f: It updates one test that main added in #26544: the startup sign-in row test still expected technical "Provider diagnostic" text on the row, which this PR withholds. The test now checks that plain detail stays literal, that technical detail is withheld, and that the agent's sign-in step and resend advice remain; no product code changed. The previous head1c3ed1f2b03fixed the four wording findings from the review of the October 8 main sync (d9ac767e852, which merged main58a58f4e4e9). Local checks on this head: every root test that imports a source file this PR changes (3,958 files, 37,480 passed). The one failure was the cross-version harness timing out after 45 seconds while extracting this head under machine load; CI's cross-version job passed. Phone tests passed (281 files, 3,621 tests). The queued node typecheck passed.orca-ci-checks --no-typecheckpassed 21/22; the only failure is the same five phone import-cycle warnings that main has. CI on this head (run): static analysis/typecheck, both package jobs, mobile bundle, cross-version wire compatibility, relay integration and test shards 1/5, 2/5, 3/5 and 5/5 passed. Shard 4/5 failed one test that only main has,preflight-command-exec.test.ts("gives an explicit binary the full timeout", from #25183). CI runs the PR merged with current main. That test computesDate.now() + 5000and later subtractsDate.now(), so it got 4999 ms when a millisecond passed in between. It is a timing flake in main's test, not ours. The earlier shard 2/5 failure,transcript-watch-inplace-drain-race.test.ts, passed on CI this time and 5 of 5 times locally; nothing it imports at runtime is changed by this PR.Earlier head:
c15c79bf109669676541ca63779630b0da167edf. Main5252884a9c7a34ab556ee8e3e1d8d4decae4982bwas merged once at the coordinator’s request after the fresh review. The resolved source passed 560 root tests in 34 files and 335 phone tests in 25 files. Queued node, renderer and CLI typechecks passed. The phone test-file gate included 955 files, four intentional exclusions and main’s unchanged 122-file baseline.CI exposed incomplete test updates for the changed copy and one missing Node-runtime registration; all were corrected, with 148 affected root tests and one phone retry test passing. The correction head
3c06915737cfinished CI without failure. The later Stop wording correction avoids inferring global idle from a refusal about one requested response. All 181 tests in six affected files and fresh queued node, renderer and CLI typechecks passed. A test uses the production Stop handler with a simulated Codex connection; this is not live-agent evidence. The last commit updates one runtime assertion and its name; all 28 tests in that file pass, and application code is unchanged from99ecd04607e.Required pre-push checks completed at 21/22. Only five native import-cycle warnings failed, independently reproduced at the exact same files and positions on the merged main with full dependency resolution; all localization, type-aware lint, changed-code quality and React Doctor checks passed. All 37 final-head checks completed: 21 passed, 16 skipped, no failures. Root CI passed all ten test shards, static analysis/typechecks, both package jobs, relay integration and aggregate verification; Mobile Checks passed. Skipped checks are not passes. Earlier CI results below belong to their stated commits.
Earlier phone-layout head:
c4328cd870d63fcd67f00dc47d9534babf59cfc8. All 32 final-head CI checks completed: 15 passed and 17 skipped after the infrastructure rerun. Every check that ran passed; skipped checks are not counted as passes.Passed: static analysis/typecheck, mobile verification, mobile web bundle, both package jobs, relay integration, and test shards 1/5, 2/5, 3/5 and 5/5.
Resolved infrastructure failure: test shard 4/5 rerun and the final verify job passed. The earlier nine failures were PowerShell aborts while loading Newtonsoft.Json; their relevant command/test files match merged main. No red final-head check remains.
Local final-head checks: targeted root tests passed (4 files, 67 tests), targeted mobile tests passed (7 files, 111 tests), and queued web/node typechecks passed. A fresh narrow review of the main merge found no new actionable issue.
Pre-push checks:
orca-ci-checks --no-typecheckcompleted with 20/22 checks passing, including changed-code quality and its type-aware scan, React Doctor and localization checks. The native audit's five import-cycle warnings reproduce at identical locations on the merged main source. The separate full type-aware audit never acquired a machine queue slot and exited after 1,801 seconds; it did not run and is not a pass. No queue bypass or process kill was used.Native phone QA: Round 4 verified text/photo rejection and actual main-before comparisons as detailed above; accepted-state control and attributed/long-reason fixture layout also passed. No tests or typechecks were rerun in this QA round.
Phone rows:
MobileNativeChatMessage.test.ts(the muted line under a not-sent message, including another agent's),MobileNativeChatView.unsent-notice.test.ts(view wiring),mobile-native-chat-unsent-notices.test.ts(host's words; "Your message was not sent." when the start row says why).Send answer:
mobile-structured-send-delivery.test.ts(a recorded rejection: no banner, nothing handed back; Stop, card and never-recorded cases unchanged),mobile-structured-send-new-action.test.ts(each Send is a fresh action, and its returned ID matches the outgoing message even after an unconfirmed answer),use-mobile-structured-agent-session-send.test.tsx(the row drawn in place through the real hook).Pending copies settled by message id:
mobile-native-chat-own-row-echo.test.ts(same-text resend, unacknowledged send that landed, own not-sent row with no "Delivery unconfirmed", photos and captioned photos, a row the host moves, two quick same-text sends, a card sent out after a lost answer, a hidden row),use-mobile-native-chat-drafts-own-row.test.tsanduse-mobile-native-chat-drafts-photo-race.test.ts(through the real hook), the controller test (the host's record reaches the drafts).Order: projection tests, including a not-sent row between two stopped sends and a stopped send with no position; a fuzz check over 30,000 random chats found the desktop's drawn output identical to main and the phone's order equal to the desktop's.
Words: the earlier phone-layout tests preserved desktop wording; the later shared-copy change updates both clients together.
Ablations: each behaviour and each wiring step reverted fails its own test.
Earlier wider runs (before the final main merge): the related root tests (80 files, 733 tests) and mobile tests (49 files; one file hit a 5-second load timeout at machine load ~70 and passes alone).
Platforms: shared and phone logic with no OS-specific code. SSH: every row comes from the host's published record. Folder workspaces: no git assumption.
I manually tested these changes locally (native iOS text/photo cases; fixture and untested variants are explicitly qualified above)
Automated tests added/updated, or explained why not below
Review
The earlier phone-merge review found no new actionable finding on its stated head. Two fresh shared-copy reviews then found marker leakage, incorrect diagnostic audiences, misleading Stop/command advice and missing names; those findings are fixed in the final head above, including the embedded older marker sentence and final error after retries. One low-impact review finding remains deferred: not-sent row objects are rebuilt as a reply streams. The final sync resolves the earlier host-record lookup naming/module issue by deriving receipt IDs directly from host submissions in a dedicated phone receipt module. The saved send queue follow-up also remains responsible for the documented history-page limit. Native attributed placement and long-reason wrapping passed using a labelled persisted fixture; the corrected accepted-state control passed using an actual native Send. Actual other-agent delivery, the main restored-text/banner variant, and photo fallback after a fresh-main reopen remain unverified.
Deferred follow-ups.
Agent skill upstream boundary
docs/reference/agent-skill-sharing-upstream-boundary.mdand copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.Notes
Author: @BrennanKB5
Checklist
N/Awith reasonpnpm lint,pnpm typecheck,pnpm test, andpnpm buildpass (or CI will cover; local preferred)