Repository navigation
fix(native-chat): a failed agent start fails only its own message - #26119
Draft
brennanb2025 wants to merge 29 commits into
Draft
brennanb2025 wants to merge 29 commits into
brennanb2025 wants to merge 29 commits into
Conversation
…start no longer holds the cards behind it
…for; the messages behind it each get their own start
… batch, which nothing writes now
…n its message, with no second note
…essages its own start rejected
…2025/start-failure-fails-its-own-message
…rejection, so clients that hide a rejected message still see why
…a later separate failure is announced again
…ow its own start wrote, in either write order
…that message; an older host's batch row matches by its failure
…ge is charged with that start
…eed announces once when nothing is owed, as on main
… exit's start row; the failed-start check is file-local
…, until a turn is delivered
… no row, in the tests that counted one per attempt
…owever often the chat is sent to
… on the journal's lane, as its batch is planned
…failure, so a CLI's timestamped stderr keeps one row per run
… starts, a /compact with no item of its own too
…, and that row never speaks for a message's failure
… for, so a row worded for /compact is that command's The exit wrote its row keyed by the start even when it rejected a handed /compact and worded the row "Run /compact again". That row then started a run, so a later message failing alike was read under the /compact's words. Keying the row by the message its words are for makes it a command's row, which never starts a run and is never dropped. A /compact's failure now always writes its own row, so the composer no longer needs to ignore rows loaded before the send.
Main now hands a message to a child still starting (no wait on its start) and holds the next one until the turn ahead opens. Resolved so a failed start still fails only its own message: the loop keeps its per-message writer and catch; the exit charges a start only to a still-queued message it was started for, so the delivery loop's start-wait tracking (waitingOn / awaits / deliveryAwaits) is gone with the wait. Tests that drove the removed start wait are rebuilt on refusals, exits before and after handover, and a dispatch that faults while starting.
… again, and the awaited command reads an empty id as none
…own, and the exit words its row only for a message it was handed - The delivery loop charges a gone child's failure only to the message that child was started for; a message whose pass merely joined that start goes on to its own. - The exit reads the message its words are for among sends it was handed, so a /compact still queued never words, or keys, the exit's row. - A send a child past its start refuses because it ended writes no start-failure row; the exit's own row says why. - The exit checks for itself whether the message its start was for is still queued. - Restart continuation is left as on main (its note beside a failed start is main's).
…arting child before that child exits
Contributor
Author
|
Status at The problem. When a native chat's agent failed to start, Orca rejected every message waiting behind it with the same error. Messages typed while the agent was starting never got their own try. So a one-off crash during startup threw away messages that would have gone through. What changes for the person.
Verified.
Not verified.
|
… in the Node runtime project
Main removed the desktop's saved outbox (#25959): the renderer's delivery notices now come from in-memory sends and the host's rows. PR A's start-row reading (a row keyed by its message speaks for it alone; otherwise any loaded row with the same failure, but a command's) is kept on the new shape, and the start-row test moves to the new API; its outbox-only and released-host-batch cases are gone. Main's new Claude sign-in prompt counts start-failure rows through their fact.
Contributor
Author
|
CI note at
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Size: net +415 production lines vs main (+756 added, −341 removed); tests +1.7k. This is the focused rebuild of #24340, which was far larger.
ELI5
When a chat's agent can't start, Orca used to throw away every message you had typed while it was starting, all with the same error. Now only the message that start was for fails. Each message after it gets its own try, so if the problem was a one-off crash, they go through.
What Changed
The problem. In native chat, if you send A, then B and C while the agent is still starting, and that start fails, main rejects A, B and C together with one " couldn't start" row. B and C never get a start of their own. When the failure was a one-off (the CLI crashed during startup, or Orca stopped a start that hung), you lose messages that would have gone through. Cause: the delivery loop's failure path recorded the start failure with
rejectsQueued, which rejects every queued message in one write. Its error handler did the same.What you see now.
/compactfails to startNo new strings or UI elements. Released clients, such as a v1.4.221 desktop paired to a newer remote host and the phone app, still get the durable error row, because the host keeps writing it.
The mechanism.
Delivery loop (
structured-agent-session-delivery-loop.ts):The three writers of a failed start each own one message state:
Each writes the rejection and the start's row together in one journal append (
rejectWithStartFailureRow). The row is keyed by the message it speaks for, so a client can tell which message a row is about.One row per run (
journal-start-failure-run.ts), decided on the journal's write lane. A message's start that fails with the same structured failure as the latest start-failure row writes no new row, provided no send was accepted since. "Same" is a field comparison of the failure fact (sameAgentSessionFailureFact, now shared by host and renderer). The CLI's own text counts only when it is meant for a person, not when it is log output. A/compact's row is always written and never starts a run.Renderer (
structured-agent-session-delivery-notices.ts): a rejected message under its own row, or under the run's row, reads "Your message was not sent." The row says why, once.Why
Differences from the common pattern:
Linked Issue
Supersedes #24340 (rebuilt to its core fix).
Visual Proof
Live run on a laptop rig: main
0acf039b5d5(before) vs this PR at0c1224262dc(after). Stand-in agents scripted the failures, and an external driver did the clicking.1. Claude crashes once while starting; A, then B and C sent during the start. Before: all three not sent. After: only A is not sent; B and C are answered by a new start.
2. Claude not signed in; three sends. Both builds: three "Your message was not sent." and one row. After: each message made its own start, and the row is not repeated.
3. Codex start crashes on each send, and its stderr carries a new timestamp each time. Before: three identical rows. After: one row.
Testing
Host tests:
structured-agent-session-start-failure-writer.test.ts(18 cases across every failure path: a refusal before spawn, an exit before and after handover, a dispatch fault while starting, a joined start, a cleanup failure, failures inside the error handler);start-failure-run-lane;command-start-failure-run;signed-out-run(scripted Claude rig, real not-signed-in path); exit settlement unit tests.Each new rule was checked by ablation: turning the rule off fails its test.
Typecheck (node, web, cli). The only errors are in files this PR does not touch, all present on main. Lint, changed-code quality, max-lines and reliability gates pass.
Live QA (laptop, main vs this PR): the three cases above, plus
/compactwhile signed out (its own row on both builds; the message after it gets its own row). Not covered: the phone (fix(native-chat): the phone keeps a message the host did not deliver, and a failed /compact is said once #24918), and Codex crashing while the chat first opens, which shows "Chat could not be started" with messages left "Sending…" on both builds. That is a separate path, unchanged here.I manually tested these changes locally (macOS laptop rig)
Automated tests added/updated
Review
Reviewed in six rounds by independent reviewers. The last round's findings are fixed or labelled above.
Agent skill upstream boundary
Notes
X: @BrennanKB5
Checklist