Repository navigation
fix(native-chat): the phone keeps a message the host did not deliver, and a failed /compact is said once - #24918
brennanb2025 wants to merge 40 commits into
Conversation
…it is not delivered A send the host recorded and then settled as not delivered (its agent failed to start, Orca restarted before handing it over) was hidden by the transcript projection unless the sender's outbox entry still carried it. Another window, the phone, or the sender after its entry was dropped saw the message vanish. The projection now shows every recorded, non-withdrawn rejection from the journal as a not-sent row at its journal position, and the delivery notices word it from the host's fact (no Retry where this client holds nothing to send). The phone row says "Message not sent". Not-sent rows take no rail tick, so the host's outline is unchanged. One predicate now reads a recovered unknown, including an older host's restart row without the marker.
…n the phone too The phone's banner repeated what the recorded message's row now says. A recorded rejection's words move to one shared function: the desktop row translates them, the phone row shows them in English, and a start-failure row that already says why leaves only "not sent". The phone's banner keeps only a Stop's withdrawal, which has no row.
…t original in place
… of this change It is unrelated to recorded messages staying in the chat; it moves to its own follow-up.
…tbox entry for The pane read the journal's rejection rows only while its own outbox held a rejected entry, so another window, or the sender after its entry was gone, drew the host's not-sent message as a plain bubble with no line under it. It now reads them whenever a row is shown as not sent.
…re it already is A rejected /compact drew as a not-sent "/compact" message while its reply or its turn's "Compaction failed" row already said it failed, so the failure showed twice. A rejected hand-off of a queued draft drew as a not-sent bubble while the draft's card kept the same text. Both stay hidden as before: the command entry is read from the journal's own item, the hand-off from its queued-draft link.
…t message A message shown as not sent has no rail tick, but the active-tick rule still lit its own id, so the rail lit nothing while it was the row being read, for example at the bottom of a chat ending in one. The tick before it stays lit.
…en sending, too Hosts from before the recovered marker publish only the restart reason. The working-state rule already read that reason as recovered; the send answer now uses the same predicate, so both agree on those hosts.
…ected The phone put a rejected message's text back in the composer even when the host had recorded it, so the same text showed in the chat as not sent and in the composer. The host's row now holds it, as on the desktop: the send reports that the host holds the text, so the composer and attachments are not restored. A message a Stop withdrew still goes back with its notice.
The line under a message that was not sent, on desktop and phone, used the
error color, as if the user had something to fix. It is a plain label now,
in the muted text color. A line that is still in doubt ("Message delivery is
unconfirmed.", or an expired id Orca can't confirm) keeps the error color,
as does the terminal chat's notice.
A /compact the host recorded and then rejected was hidden, so on several paths (blocked after the reply stopped waiting, rejected on restart, or seen from another window or the phone) it vanished with nothing saying it failed. It now stays as a not-sent row for every viewer, like any message the host recorded. Its line says only "Your message was not sent." when a loaded host row already says why: its turn's result row (found by the command's id) or the failed start's row. The sender's command reply no longer repeats it when the loaded journal shows the host recorded the command under the operation id, and the text stays the row's rather than coming back to the composer. The desktop pane's delivery-notice wiring moves into its own hook.
…rows The scan parsed every journal key on each update while a row showed as not sent, about 3-17 ms per call at 3000 items. Every result row is a status row, so the scan now skips the rest before parsing (about 0.02-0.04 ms).
…as rejected After a lost answer, the phone keeps that message's id for its text. If the host since recorded and rejected it, sending the same text again replayed the id, which answered with the same rejection: no banner, no bubble, and an empty composer, so the press did nothing. Such a replay now goes out under a fresh id, as a withdrawn one already does. "Sent, but this phone couldn't update its record" is no longer said for a resend the host recorded and rejected; its row says it was not sent.
…orded-outcomes-from-journal # Conflicts: # mobile/src/session/MobileNativeChatMessage.test.ts
Review summary (updated after #24710 and #24606 landed)The problem. When Orca recorded a native-chat message on the machine running the agent and then could not deliver it (the agent failed to start, Orca restarted mid-send, the provider refused it), the message used to vanish. #24710 has since fixed that on the desktop. What was left:
What changes for you.
What the review fixed.
Deferred (P3, none block this change).
Verified.
Not verified.
All four are to be captured before this is marked ready. It lands together with the follow-up that removes Retry. |
Takes main's side for the files #24710 also changed; this branch's remaining changes are re-applied on top in focused commits.
The start-failure facts, the same-failure match and the host-rejection words lived in the desktop renderer, so the phone could not share them. They live in one shared module now, which the desktop notices read.
…said once Design case 1 and "no silent dead end": a hidden /compact vanished with no word when its reply was gone, so it is shown like any recorded message, and its line says only "not sent" where its result row or the start's row says why.
…ace, as on the desktop Design case 1: the phone no longer hands such a message back or banners it (this branch's earlier change), so its row is where it lives; the phone now draws it as not sent, like the desktop, and leaves one a queued card holds to that card.
The not-sent line drew in the error color as if the user had something to fix. Main's notices now mark every plain "not sent" (the host's row, its outbox copy until the row loads, and a held send) so the row draws it muted, as the phone does; a line still in doubt keeps the error color. A batch that changes only that mark re-renders its row.
… outline Design D6: a not-sent message is no prompt the agent saw and the host's outline leaves it out, so the loaded rail does too (the rail skip and the active-tick walk-back already on this branch); the parity test pins that again.
…phone draws it there The projection listed rows shown as not sent after the whole conversation and left ordering to the reader. The desktop sorts; the phone draws the list as it comes, so an old not-sent message sat under every later message. The projection now merges them in at their journal positions: the desktop's sorted output is unchanged, and the phone draws them where the host recorded them.
The phone counted a row shown as not sent when it matched a send's echo, and took it as the newest row the echo had to land after. Once the resend lands, the host hides that row, so the echo expected one copy too many and never retired, leaving a plain duplicate bubble; an answer-lost resend read as unconfirmed. Not-sent rows no longer count toward either.
… it was not sent A /compact reply was silenced as soon as the journal held its submission, even while pending. If the reply beat the stream and a Stop then withdrew the command, the row was hidden too, so the command vanished with nothing said and its text gone. The reply is now silent only once the loaded journal draws the command as not sent.
A resend the host answered as unknown reads "Orca couldn't confirm what happened" and may have landed, but its line drew muted like a plain "not sent". It keeps the error color now; its words are unchanged.
…the rail does Drops a check that could never match: the outline projects no rows as not sent.
Echo matching skipped every row shown as not sent, so a send whose own row first appeared already rejected was never settled: a "Delivery unconfirmed" banner followed 20 s later, or its bubble stayed beside the row. A send now records the not-sent rows already present when it went out, as it records the queued cards; only those claim nothing, and a later one is its own.
…orded-outcomes-from-journal # Conflicts: # src/renderer/src/components/native-chat/NativeChatMessageRow.test.tsx # src/renderer/src/components/native-chat/NativeChatMessageRow.tsx # src/renderer/src/components/native-chat/structured-agent-session-delivery-notices.ts # src/renderer/src/components/native-chat/use-structured-agent-session-delivery-notices.ts
The merged notices test passed the line limit; its two tone tests move out.
Since "Sending…" landed, an unconfirmed entry the probe still resends reads as sending, so the case no longer tested the doubt tone. It is now one the user already retried, as main's notices test does.
React Doctor flagged the array lookups inside the echo loops and an Array<…> type; the snapshot is read as a Set and the type uses T[].
A rejected message's row disappeared once a later copy of the same text was recorded, so a failure the chat had shown could vanish from the record. The common pattern keeps the original: a rejected request stays where it was, matched by its own id, never hidden by text. Each rejected message now stays at its place with its muted not-sent line, and a resend is an ordinary new message that says nothing more. The phone's echo matching already leaves the original alone, so a resend ends as the original plus the delivered copy.
…ected message Main's #24710 draws a message Orca recorded and then rejected where the host placed it, as "Not sent", from the host's own history. The host side (the rejection moves the message to its place; a failed start writes its rejections and its row in one transaction) is taken as written. On the client, C2's settlement stays the only place an outbox entry ends: - #24710's reconcile module is folded into C2's: an entry whose rejected row is not loaded yet stays, never sent again, and draws the message until that row loads (or the host's window for its id ends). Nothing new is stored; the held submission says it was rejected. A batch that settles nothing writes nothing. - One "not sent" surface: main's rule hides a rejected message a live card holds or a later copy of the same text superseded, in the rows and the notices alike, on the desktop and the phone. The notice stays C2's muted one, and notices are kept as the same objects across unchanged batches. - No Retry: a message refused before it was recorded still goes back to the conversation's draft with the reason, as before. - A rejected /compact stays in the chat (#24918's rule), since the command's own reply says nothing once the journal shows it. - The phone and the desktop both draw these rows; only the host's outline, which older clients read, leaves them out.
#24918 and C2 already agreed on everything this head adds over main except wording sources, so the resolution keeps C2's mechanism: - A rejected /compact is case 1, shown muted and said once. The command's reply stays silent only while the journal draws it: a withdrawn, card-held or superseded copy still speaks (structuredAgentSessionJournalShowsSubmission now reads the shown-in-place set). - The phone draws recorded rejections in place with its live card ids. - The row's words come from the host's reason and fact, through C2's send-failure words module (send-disposition stays deleted). - The not-sent line keeps C2's `muted` flag; #24918's `notSent` field is not taken, and its Retry-era outbox notice tests stay out. - The list test for a rejected /compact is taken, on C2's notice signature.
Brings main through d3afb5c, which includes #24606 as it landed: in the native chat that is the version C2 already merged, so those files keep C2's resolution, and the locales keep C2's messageNotConfirmed line beside main's new keys. From #24918: - Not-sent rows come back at their journal places, so the phone draws them where the host recorded them. - A command's reply stays silent only once the loaded journal draws it as not sent (structuredAgentSessionJournalShowsRejection), so a pending command a Stop then withdraws still gets its reply. - The phone's echo matching no longer double-counts not-sent rows. - The delivery-line tone test is taken in C2's terms: a recorded rejection reads muted from its row or from this client's copy, and a message on its way reads only as sending. #24918's `notSent`/Retry notice shape stays out.
…orded-outcomes-from-journal # Conflicts: # mobile/src/session/MobileNativeChatMessage.tsx # mobile/src/session/use-mobile-structured-agent-session.ts
…wn module Keeps MobileNativeChatView within its line limit after the merge with main.
Drafts take #24905's merged store. A hand-back's text follows the shared returned-text rule (#25149) through returnNativeChatDraftText, which trims only the end and now trusts the append's durable answer; a composer is told of an append only when it changed the draft. The withdrawn-message restore test stays deleted: its one change was the rule's dedupe, which C2's returning tests already state. messagesUnsettled keeps C2's 'wait' in the moved reason-words table.
… own not-sent row The image matcher read rows through a helper that dropped every not-sent row, so a captioned image send the host accepted and then rejected kept its sent-looking placeholder beside its own not-sent row. It now applies the send-time rule the text matchers use: only a not-sent row already on screen when the send went out can't be its echo. The helper is gone; each matcher states its not-sent rule itself.
…orded-outcomes-from-journal
With a not-sent message followed by its failed start's row at the bottom, the rail lit no tick: the walk-back skipped only rows that were themselves not sent, and the start's row resolved to no tick or to the not-sent message's id. The lit tick now walks back to the nearest row whose tick exists, skipping rows in no turn and rows whose turn a not-sent message names.
Three delivered messages with replies, a not-sent message, then the failed start's row last, built through the renderer's own journal-to-slots path: the third message's tick stays lit when the chat fits and is pinned to the bottom, and with either row at the fold. With host-stated turns the start row is in no turn, which is the layout that lit nothing before the walk-back fix.
…orded-outcomes-from-journal
…jected-command test Main's chat appearance settings (#25657) replaced the message list's fontScale prop.
Brings main eaaae01 and #24918's rail-tick fix. Conflicts: the rejected-in-place and delivery tests keep C2's cases on main's fake probe clock; the rejection-cause test stays deleted (main changed only its timers); zh.json takes main's 智能体 wording beside C2's keys. A C2-only test's font mock follows main's font-size rename.
|
Closing in favour of a rebuild on current main. Since this PR was opened, main landed most of its desktop half (#24710), keeps a message Stop took back on screen (#25051), keeps a cut-off send as a queued card (#24660), and the desktop send layer is being reworked (#25959). Merging 126 commits of main into this branch would have meant resolving 12 conflicting files against changed premises, so the remaining work is split into two focused PRs:
The branch is kept for reference. |
ELI5
#24710 made the desktop chat keep a message Orca recorded but could not deliver, labelled "not sent", instead of letting it vanish. This change brings the same to the phone, makes a failed
/compactshow up and say its failure once instead of disappearing, draws the "not sent" line quietly (muted grey) instead of in red, and keeps the conversation rail from counting messages the agent never got.What Changed
The problem. When you send a message in a native chat, Orca first records it on the machine running the agent (the host), then hands it to the agent. If that hand-over fails (the agent failed to start, Orca restarted mid-send, the provider refused it), the host marks the message rejected in its record. #24710 now shows such a message on the desktop, in place, for every window. What it left:
/compactwas hidden everywhere. Its failure was said by the command's reply, which only the sending window sees and only while it waits for the reply. A/compactblocked after that wait, rejected after a restart, or viewed from another window or the phone vanished with nothing saying it failed.Before / after, as you see it.
/compactthe host recordedOne #24710 behaviour this removes: hiding a failed message when you send the same text again.
What #24710 already covers, and is reused here: the desktop row from the host's record for every window, its placement where it was rejected, no Retry on a recorded rejection, the outbox copy shown until the host's row loads, the queued-card handling, and the "the start row already says why" rule. One #24710 rule is removed: it hid a not-sent message once the same text was sent again, matching by text. Here a message is kept or hidden by its own id, never by its text, so history stays truthful.
The mechanism.
src/shared/structured-agent-session-message-projection.ts: a rejected message is hidden only when Stop withdrew it, its queued card holds it, or it is a queued hand-off (no longer when a later copy has the same text); rejected commands are no longer excluded; not-sent rows are merged into the conversation at their journal position, so every reader (desktop and phone) gets them in order without its own sort.src/shared/structured-agent-session-recorded-rejection-words.ts: the start-failure facts and the recorded-rejection wording fix(native-chat): show a message Orca accepted and then failed to deliver as "Not sent" in the chat #24710 added on the desktop move here, so the phone words the row the same way; a loaded command result row now counts as "already says why".use-mobile-structured-agent-session.tsdraws rejected rows in place;mobile-native-chat-unsent-notices.tsandMobileNativeChatMessage.tsxput the reason under the row;mobile-structured-send-delivery.tskeeps its banner and composer hand-back only for a Stop's withdrawal or a send the host never recorded. The phone's pending-send matching ignores not-sent rows that were already on screen when you sent, so retyping the same text never leaves a stray bubble or a false "Delivery unconfirmed".structured-conversation-command-send.ts(desktop) andmobile-structured-composer-command.ts(phone) stay quiet only when the loaded history already shows the command rejected in place; a withdrawn command, or one from a host too old to record commands, still answers in the composer.notSentflag;NativeChatMessageRow.tsxdraws it withtext-muted-foreground, the phone with its muted text colour. Doubt notices keep the error colour.native-chat-message-rail-items.tsgives a not-sent row no tick, matching the host's outline.native-chat-active-rail-item.tslights the nearest tick above whatever row you are reading, so a not-sent message, or the "couldn't start" row under it, keeps the message before it lit.structured-agent-session-unanswered-dispatch.ts: one check answers "did the host lose this send for good", including an older host's restart reason without the marker; the send-answer handling uses it.No wire change and nothing the host publishes changes.
Why
The common pattern keeps a message the host recorded in the conversation, for every viewer, and states its failure once from the host's own record; the text is not handed back. #24710 did that on the desktop. Doing the same on the phone removes the last place the message vanished or showed up twice (row and composer). A command is a message the host recorded too, so it follows the same rule rather than relying on a reply only one window sees.
Alternatives considered:
Differences from the common pattern, each labelled:
Linked Issue
N/A (maintainer). Part of the native-chat failed-send follow-up, after #24710.
Visual Proof
Captured in an isolated, hidden test app on a Windows machine, with a stand-in Claude (no real agent, no login), driven over the browser debugging protocol. Before = main dd39dcb, after = this branch at 9adf800. Later commits only add the phone fix for image messages, a merge of main, and the rail fix below.
Sending the same text again after a not-sent message. Before: the failed original disappears. After: it stays, labelled not sent, and the resend is a new message with a reply.
Agent fails to start. Before: red "Your message was not sent." After: the same line in muted grey; the start row says why.
A failed
/compact. Before: the command is hidden, "Run /compact again." appears twice, and "/compact" is left in the composer. After:/compactstays as not sent, the reason is said once, and the composer is clear.Three delivered messages and one not sent. Before: four rail ticks. After: three; the not-sent message takes none.
That run also showed the rail lighting no tick at all when the "couldn't start" row under a not-sent message was the row being read. That is fixed in 0fd5735, with a test that rebuilds the same chat (
native-chat-rail-not-sent-start.test.ts); the fix has not been re-checked in the running app.Not captured:
mobile-native-chat-unsent-notices.test.ts,mobile-native-chat-not-sent-resend.test.ts,mobile-structured-send-delivery.test.ts,mobile-structured-composer-command.test.ts,MobileNativeChatMessage.test.tsand the session send tests.Testing
Projection: rejected commands drawn as not sent; not-sent rows at their journal position; the desktop's sorted output unchanged.
Commands, said once: the real message list with a refused, blocked and start-failed
/compact(NativeChatMessageList.rejected-command.test.tsx); the sender's reply quiet when the row shows it and speaking when the command was withdrawn or the reply came first (structured-conversation-command-send.test.ts, the command start-failure hook test,mobile-structured-composer-command.test.ts).Phone: rows through the real session hook; words and the start-row rule (
mobile-native-chat-unsent-notices.test.ts); no banner or hand-back for a recorded rejection; same-text resend ends as one bubble, an unacknowledged send that landed is seen as landed, its own not-sent row settles it with no banner (mobile-native-chat-not-sent-resend.test.ts).Phone image messages: a captioned image whose own message shows as not sent retires its placeholder bubble (
mobile-native-chat-not-sent-resend.test.ts).Rail: the lit tick when reading a not-sent message or the "couldn't start" row under it, through the real chat-history-to-rail path (
native-chat-rail-not-sent-start.test.ts).Muted vs doubt colour, including an "Orca couldn't confirm" send (
structured-agent-session-delivery-notices.test.ts); rail ticks and lit tick (native-chat-active-rail-item.test.ts, rail/outline parity).Ablations: each change reverted makes its own test fail.
Wider: 98 root test files and 56 mobile test files around the changed modules pass (the one root file that needs main's new JSON-stream package could not load it in this checkout; CI installs it).
Typecheck: web, node and mobile clean apart from that missing package and the mobile generated modules this checkout does not build. Lint and the changed-code quality gate pass.
Platforms: shared, renderer and phone logic, no OS-specific code. SSH: every row comes from the host's published record. Folder workspaces: no git assumption.
I manually tested these changes locally
Automated tests added/updated, or explained why not below
Review
Agent skill upstream boundary
docs/reference/agent-skill-sharing-upstream-boundary.mdand copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.Notes
Author: @BrennanKB5
Checklist
N/Awith reasonpnpm lint,pnpm typecheck,pnpm test, andpnpm buildpass (or CI will cover; local preferred)