Skip to content

feat(native-chat): keep a message a Stop took back on screen, with one stop row after it - #25051

Merged
brennanb2025 merged 256 commits into
mainfrom
brennanb2025/stop-withdrawn-shown
Oct 6, 2026
Merged

brennanb2025 merged 256 commits into
mainfrom
brennanb2025/stop-withdrawn-shown

Conversation

@brennanb2025

@brennanb2025 brennanb2025 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor
Files Added Deleted Net
Test 17 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​2903 $\color{#cf222e}{\Huge{\mathbf{−}}}$​78 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​2825
Prod 29 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​966 $\color{#cf222e}{\Huge{\mathbf{−}}}$​95 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​871

Base main: #24864, #24369 ("Stopping…") and #25073 are all on main.

Builds on #25073 (the submittedSequence and answeredInTurn fields), which is on main as 9def4b9ba19.

ELI5

When you press Stop before the agent has started on your message, the message used to vanish from the chat and its text reappeared in your typing box. Other devices just lost it. Now the message stays in the chat, with one line saying it was stopped, and your typing box is left alone. A Stop never loses what you sent.

What Changed

The problem.

  • A Stop can take back a message before the agent starts it. Orca then removed the message from the transcript and pasted its text back into the composer of the window that sent it.
  • Nothing on screen said what had happened. On a phone, or on any desktop other than the sender's, the message simply disappeared.
  • With the message gone there was no run left for a stop row to describe. In live testing, "Cancellation requested." landed under the previous finished turn, so that earlier turn read as if it had been cancelled. fix(native-chat): a Stop binds only the turn it actually stopped #24864 stopped writing that row as a stopgap and named this follow-up.

Before / after:

Situation Before After
A Stop takes back a message before the agent starts it The message disappears; its text is pasted into the sender's composer The message stays where it was sent, followed by "Stopped before the agent started"; the composer is untouched
Several messages taken back by the same Stop All disappear; their text is pasted together All stay, with one "Stopped before the agent started" row after the last of them
The same chat on a phone or another desktop The message disappears Same as the sender: the message and the row
The agent had opened a turn for the message (echoed or not), and the Stop interrupted it The message disappears; the turn's "Interrupted after N" header sits on its own The message stays in that turn; its "Interrupted after N" header and the turn's own "Cancellation requested." are the stop, and no third row is added
A message sent into a turn the agent was already running (a steer) that the Stop took back The message disappears The message is drawn after that turn's "Cancellation requested.", followed by "Stopped before the agent started". The host now says which turn it joined, so it could be drawn inside that turn later; this PR does not do that
A message waiting for a running turn to finish (for example behind /compact), or sent just before the agent opened its turn, that the Stop took back The message disappears The message and its row are drawn after that turn's own rows, where the Stop took it back
A queued card's send taken back by a Stop The card returns, paused Unchanged
A message this window had not yet passed to the host when you pressed Stop Its text is pasted after whatever you typed Its text comes back into the composer if the composer is empty. Otherwise it stays below the conversation as "Message was not sent." with Retry, and the composer keeps what you typed
A message refused for any other reason "Message was not sent." with Retry Unchanged

To send a stopped message again, copy it (user messages already have a Copy button on desktop, and selectable text on the phone) and send it.

Mechanism.

  • Shared message projection (src/shared/structured-agent-session-message-projection.ts, used by desktop and phone):
    • A user message whose submission was rejected as withdrawn (the cancelled fact, or its older marker) is no longer hidden. A queued card's hand-off is still left out: the card holds its text.
    • It gets one client-drawn row after it (presentation: 'stopped-before-start', src/shared/native-chat-stopped-before-start.ts), shared by consecutive taken-back messages. The row is skipped when a turn opened for the message: that turn's interrupted end already says it was stopped.
    • A message the Stop took back is drawn at its own journal row. Since fix(native-chat): show a message Orca accepted and then failed to deliver as "Not sent" in the chat #24710 the host moves a rejected message's row to the row that rejected it, in no turn, so that is where the Stop took it back. It is drawn past the end of every turn whose opener the journal wrote before that row, so it never splits the exchange it waited on. That is one comparison in journal order, never a clock.
    • Several messages the Stop took back with their own stop rows keep the order they were sent in, by where each was sent (submittedSequence, feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073). The move erased that from the row itself.
    • A steer the Stop took back is no exception: the host takes a rejected message out of its turn, so it is drawn after that turn like any other. The host now names the turn it joined (answeredInTurn with via: 'steer', feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073), which would allow drawing it inside that turn later. This PR does not.
    • On a host that publishes neither field, the earlier placement rules apply instead (Temporary); see Mixed versions.
    • Other rejections stay hidden, as before.
  • Turn anchors (structuredAgentTurnAnchors in src/shared/native-chat-turn-membership.ts):
    • Codex reports a turn open before it echoes the send, so until the echo the turn record names the provider's key.
    • A send the Stop took back before that echo opens a turn when the host says it did: its submission names that turn and says Codex answered the send's start request with it (answeredInTurn, via: 'start', feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073). It is then drawn as that turn's opener, just before the record, with no row of its own. A send the host states was answered into no turn (null) opens none.
    • A send steered into the turn (via: 'steer') opens nothing, even when the turn's real first message is on an older page that is not loaded yet.
    • When the turn's record names a first message that is loaded, that message keeps the turn, whatever the host says about a stopped send.
    • A turn whose record names itself, one the provider resumed on its own, is opened by no send.
    • When the loaded history starts after that turn's record, and a row of that turn is loaded, the send is drawn just before that turn's first loaded row, still with no row of its own. If no row of that turn is loaded, it is drawn on its own with its stop row, until the older history loads.
    • A rejection with no answeredInTurn at all was written before the host was upgraded to name turns. It is matched to a turn by the rule used before:
  • Desktop draws the row through the same path as the Stop's own status rows ("Cancellation requested."), in its own words from all six catalogs. The row never folds behind a settled turn. The phone draws the row's English text, as it does every other status line.
  • Rail and outline:
    • The conversation outline the host serves for older history keeps its old rule: it leaves out sends a Stop took back.
    • The message rail gives them no tick either, so the rail and the outline agree.
  • Composer:
    • The host-withdrawal refill (useStructuredAgentSessionWithdrawnRestore().byHost) is removed.
    • The sender's own Stop step gives back entries the host never received, but only into an empty composer (no text and no images).
    • Entries the composer cannot take, including where no composer shows the chat, stay in the outbox as not sent, on their Retry. That outbox is their only copy.
    • The phone: a send whose answer says a Stop took it back is not handed back to the draft, which since fix(mobile-chat): a failed message's text is added to the draft, and an expired resend no longer blocks that text #25149 adds a failed message's text to it: the chat already draws it with its stop row. A first answer like that reads as sent, as on desktop. A resend the phone could not tell from a new send of the same text (its earlier answer was lost) is sent once more under a fresh id: the stopped one provably never ran. Every other refusal still goes back to the draft.
  • Host: unchanged apart from comments. fix(native-chat): a Stop binds only the turn it actually stopped #24864 already writes no "Cancellation requested." for a Stop whose only effect was taking the send back.

Why

The common pattern keeps a stopped message on screen with one stop row after it, and leaves the composer alone. There is no per-message Retry for a stopped message, and re-running it is the person's own action. This PR adopts that.

Alternatives considered:

  • Keep refilling the composer and only fix where "Cancellation requested." lands. The message would still vanish on every other device, and a row with no message above it still has nothing to describe.
  • Have the host write the stop row. See the first difference below.

Differences from the common pattern

Difference Label
The stop row is derived on each client from the stored rejection, not stored as its own host event. Intended. Orca records the withdrawal as the send's rejection, and every client reads it from that one record. A host-written row would reach clients that still hide the message, and live testing showed exactly that row landing under the previous finished turn, as if that turn had been cancelled. Deriving it keeps the row with the message it describes on every client that shows the message.
A message the Stop took back is drawn past the end of the exchange it waited on, not inside it at the row that took it back. The common pattern has the host assign the position and the client draw it there. Orca does the first part (the host moves the message's row to where it was taken back) but draws the message after that exchange rather than at the row. Intended. Drawn at that row, it sat between an earlier message's reply and that reply's own stop note, or above output that streamed after the Stop, which reads as the agent answering after the stop row. Live testing and review found exactly that, three times.
A message this window had not yet passed to the host comes back only into an empty composer; otherwise it is kept, across restarts, as "Message was not sent." with its own Retry. The common pattern puts it back into an empty draft and shows an error; nothing is stored. Intended. The common shape lost the text whenever the composer held something: an earlier version of this PR did exactly that, and its own test pinned the loss. A kept entry needs a way to be cleared, and the not-sent row's Retry is that way. Without it, nothing would ever remove the stored entry.
A message the Stop took back after its turn opened, but before the agent echoed it, is drawn as that turn's opener, just before its record, with no row of its own. The host moved its row past the turn's record, to where it was taken back. Intended. The turn's interrupted end is its stop. Drawn at the moved row, the turn's "Cancellation requested." sat above the message, and the message got a second stop row. As in the common pattern, the host says which run the message belongs to (answeredInTurn) and the client draws it there.
Plain sends the host accepted but had not yet passed to the agent are shown as taken back. The common pattern holds them in a paused queue. Temporary. Orca holds these as paused cards once the queued-messages capability rolls out (#21062). Until then, plain sends have no card form.

Follows the common pattern:

  • the stopped message stays in the chat: where it was taken back, in no turn (a steer included), or, if it was waiting on a turn, right after that turn;
  • one stop row follows it;
  • the composer is not refilled;
  • no Retry on a stopped message;
  • a queued card stays paused.

Mixed versions. The host publishes two more optional fields on each submission (#25073):

  • submittedSequence: where the journal wrote it, worked out fresh from the session's recorded history each time it is read and never stored.
  • answeredInTurn: on a rejected message, the turn Codex answered it into and how it joined it, or null when the host recorded no turn. It is stored on the rejection row, and absent from rejections written before it existed.

Rule 1 of the wire doc applies: an older client ignores them, and #25073's cross-version test shows the two latest released clients draw a page and an incremental update carrying them exactly as they draw the same without them. No RPC method or stream frame changes, and the host-served conversation outline keeps its old rule.

Known limit: a long chat loads its newest history first. If that loaded part starts after an earlier message whose turn is still on screen, and a later message the Stop took back follows, the stopped message and its row can show above that earlier turn's reply or stop note until you scroll up and the older history loads. On a host without #25073's fields, it can also show with no row of its own at the top of the earlier message's turn, as if it had started that turn. Review found these by loading every possible starting point of 47 end-to-end cases on the real host with the fake Codex, and re-ran that sweep on this change. For rejections a host with the fields wrote, only the first shape remains: the second, and a third (the turn's "Cancellation requested." above the message that started it), are gone. For rejections written before the host was upgraded, the second and third shapes remain, exactly as before this change.

Known limit (reachability not verified): a message Codex quietly folded into a turn already running, which Orca had not heard of yet, is recorded as having started that turn. If that turn has no user message of its own (for example, one Codex ran on its own), or its first message is on an older page that is not loaded yet, the stopped message is drawn as the head of that turn. With the first message loaded, the first message keeps the turn.

Known limit (not new): a Codex rewind rebuilds the transcript from the provider's own items plus Orca's turn rows. A message a Stop took back is neither, so it drops from the transcript before the rewind point, as every withdrawn message already did.

What this removes:

Linked Issue

Follow-up to #24864 and #24369. Closes STA-9337 (https://linear.app/stably/issue/STA-9337), with #25073: a stopped message is placed by journal order and the host's named turn instead of clocks wherever the host publishes the fields.

Visual Proof

Live desktop QA on a separate Mac, against a background build of this branch (9f3cd5e5), driving native Codex chats backed by a stand-in Codex. The stand-in can open a turn and never echo the message, and can hold a steer without echoing it. Desktop only; the phone draws from the same shared projection and is covered by tests.

A Stop after the turn opened but before Codex echoed the message. The message keeps its turn: "Interrupted after 11s" sits on it, with no "Stopped before the agent started" row. "Cancellation requested." is inside the turn, shown when it is expanded.

Collapsed Expanded
interrupted on the message, no stop row Cancellation requested. inside the turn

A message steered into a running turn, then taken back by the Stop. It is drawn after that turn's "Cancellation requested.", with its own row.

withdrawn steer after the turn, with its own stop row

A Stop before the agent started the message. The message stays where it was sent, with one "Stopped before the agent started" row, and the composer is not refilled. Before this PR the message disappeared and its text was pasted back into the composer.

never-started message with one stop row

An older chat, written before the host recorded these fields, opened on the new host. Its stopped message ("Phone1 m4 never opens") draws with one stop row, and no "Cancellation requested." sits above a message that owns its turn.

old chat: stopped message with one stop row

Also checked in the same run: two messages queued behind /compact and stopped before the echo (the first keeps its turn; the second, steered in, is drawn after it with its own row); every chat draws the same after a window reload; and the journal rows record answeredInTurn as start, steer or null as described above.

Testing

  • I manually tested these changes locally (live desktop QA on a separate Mac with a stand-in Codex, at 9f3cd5e5; see Visual Proof)
  • Automated tests added/updated, or explained why not below

New tests:

  • src/shared/structured-agent-session-message-projection-stopped-before-start.test.ts (25). Each case that depends on the host is built in the host's own shape: on a host that publishes submittedSequence, the taken-back message's row sits where it was taken back, in no turn.
    • the message stays with one row after it;
    • the typed fact and the older marker read the same;
    • consecutive messages share one row;
    • a message a turn opened for stays in that turn, with no row, before and after the provider's echo, including when its row moved past that turn's record to the take-back;
    • a turn does not claim it when the turn's record came before the message was sent, or after it was taken back;
    • a steer the Stop took back, named as steered into its turn, is drawn after that turn's "Cancellation requested.", with its own row, on both host versions;
    • a message never handed over is drawn after the turn it waited behind, and a follow-up sent before the agent opened its turn is drawn after that turn's rows;
    • two messages queued during /compact, the first passed to the agent and the second not, stay in send order after one Stop takes both back, and one a later Stop took stays below one an earlier Stop took;
    • a message the Stop took back stays below everything sent before it, including an earlier message's whole exchange, whether it had finished or was still streaming when the Stop came, on both host versions;
    • two messages keep the order they were sent in, whatever their clocks say;
    • where accept times tie and the client holds the stopped message first, its moved row still puts it below the exchange before the Stop;
    • a queued card's hand-off and other refusals stay hidden.
  • src/shared/native-chat-stopped-send-turn.test.ts (17), which turn a message taken back before its echo opens:
    • the message the host says started the turn opens it, not an earlier message taken back across it;
    • in history written before the host named turns, served by a host that publishes where each was sent: the message sent before the turn's record and taken back after it opens it, the first sent when two were, and a turn recorded before it was sent is not its;
    • a message the host states was answered into no turn opens none: a Claude rejection beside a turn Claude resumed on its own, and one before a later Codex record;
    • a message naming a way of joining this build does not know opens none;
    • an echoed first message the turn's record names keeps the turn, even from a stopped message the host says started it (a start Codex folded into a running turn);
    • a turn the provider resumed on its own is opened by no message;
    • a message the host states it named none for opens none, though it was sent before the record and taken back after;
    • a message taken back inside the turn another stopped message opened is drawn after that turn;
    • a steer: with an accepted first message, it is drawn after the turn with its own row; with the first message also taken back, the first message keeps the turn and the steer is drawn after it; with the first message on an older page, it still opens nothing; on a page that starts inside the turn, it keeps its own row;
    • on a page that starts after the turn's record, the message that started the turn is drawn before the turn's first loaded row with no row of its own, unless a loaded user message of that turn is there.
  • src/renderer/src/components/native-chat/NativeChatMessageList.stopped-before-start.test.tsx (7), rendered on desktop:
    • placement after the previous turn, with no header of its own;
    • no "Working for" clock survives;
    • one row for consecutive messages;
    • "Interrupted after N" and no row for a turn that opened, before and after the echo;
    • the row is styled as the Stop's own status rows;
    • the row stays visible under a settled turn on a host that states no turn scopes.
  • mobile/src/session/mobile-native-chat-stopped-before-start.test.tsx: the phone's projection, turn status and row text.
  • Two guard tests for the phone, mobile/src/session/mobile-native-chat-merged-snapshot-parity.test.ts and mobile-native-chat-merged-snapshot-parity-hooks.test.tsx. They use a made-up chat shaped like the one in QA: a long history, a message sent while a Stop was ending a turn, then a message the agent never started, taken back when the agent exited.
    • The phone's chat, built up from the host's live updates, shows exactly what a freshly opened phone shows, after every update and from every point the phone could have joined. That includes a chat long enough that the phone first loads only its latest part.
    • The "Stopped before the agent started" row is in the phone's drawn list at every update from the one that took the message back on.
    • The second test runs the phone's own chat screen code, for a phone that sent the message itself (with its own pending copy) and sees the host's "Stopping" status.
  • native-chat-rail-outline-parity.test.ts: the outline and the rail both leave out a send a Stop took back, while the transcript draws it.
  • End to end, src/main/runtime/structured-agent-session-codex-stopped-send-order.test.ts, with the real host and the fake Codex:
    • a second send made while Codex still holds the first send's answer, then a Stop: the first send's turn opens after the second was accepted and streams after the Stop, and the stopped second send and its row are drawn after that turn's output, never between its header and its output;
    • a send, a Stop while the turn has not opened, and a second send while Stopping: the stopped message and its row stay above the second exchange after a restart whose resume reports turn start times in whole seconds.
  • End to end, src/main/runtime/structured-agent-session-codex-turn-end-settlement.test.ts, with the real host and the fake Codex:

Updated tests:

  • structured-agent-session-withdrawn-message-restore.test.tsx:
    • a host withdrawal leaves the composer as it is;
    • the sender's own Stop gives back only into an empty composer;
    • when the composer holds text or an image, or no composer shows the chat, the message stays in the outbox as not sent, text and images.
  • use-structured-agent-session-outbox-withdrawal.test.tsx and use-structured-agent-session-outbox.draft-hand-off.test.tsx: updated to the same rules.
  • Three of fix(native-chat): show a message Orca accepted and then failed to deliver as "Not sent" in the chat #24710's tests in structured-agent-session-message-projection.rejected-in-place.test.ts pinned a withdrawn message as hidden, the rule this PR replaces:
    • "stays when the same body was only sent before it, or by a copy a Stop withdrew": the withdrawn copy is now drawn too, and it still does not hide the failed copy, which is drawn as not sent;
    • "keeps a message a Stop withdrew hidden" is now "draws a message a Stop withdrew, with its stop row, and no delivery notice";
    • the phone's "still hides every accepted-then-rejected message" is now "still hides a message the host failed to deliver, but draws one a Stop withdrew": the restart-rejected message stays hidden.
    • Each checks that the stop row comes right after the withdrawn message.

Each placement rule was checked by breaking it and watching its tests fail, at this head (6 files: the two placement test files, turn membership, desktop rendering, and the two end-to-end files):

  • the host's named turn ignored: 5, including the end-to-end test;
  • null read as "written before the field": 2;
  • a rejection without the field read as null: 5, in the two shared placement files and the desktop render test (the end-to-end tests stay green, since the real host always writes the field);
  • older history placed by times only, not journal order: 2;
  • any way of joining allowed to open a turn, not only start: 3;
  • a steered message allowed to open its turn: 2;
  • the host's named turn allowed to override a loaded first message: 1;
  • a turn that names itself claimable: 1;
  • the first message of a turn measured at its moved row, not where it is drawn: 1;
  • no placement on a page that starts inside the turn: 1;
  • the "past the end of every turn opened before it" rule off: 10;
  • older host, the floor from earlier sends off: 3;
  • the send-order rule off: 1.

The phone guard tests turn red when each layer is broken on purpose: the phone dropping part of an update (the taken-back message's new state, a removed row, a newer version of a message or status), the list hiding the stop row, and the chat screen's own list (rows held back until the next send, a list refilled in place, or two rows sharing a key).

Run locally:

  • At cc6d83826cf (fixes from a review of the main merge: the Spanish "Stopped before the agent started" line restored, a phone test that compared nothing now reads "Stopping…", and one stale comment): the changed phone test, the outbox test and 4 language-file tests pass; phone tsc and the phone tests-typecheck ratchet pass. The full 10 desktop + 7 phone file set was not re-run at this head. CI: 16 pass, 16 skipped, 0 fail.
  • At a200a857519 (base main, after feat(native-chat): the chat reads "Stopping…" from your Stop until the turn actually ends #24369 merged): CI 16 pass, 16 skipped, 0 fail, mergeable. This PR's own diff against main is unchanged line for line by the main merge (45 files). This head's tests: 10 desktop files (105 tests) and 7 phone files (67 tests) pass. tc:web, tc:node, phone tsc and the phone tests-typecheck ratchet pass. oxlint is clean on every changed file (desktop and phone configs). Merging main changed three tests mechanically for main's API: the dropped NO_CARDS argument, the three-argument projection call, and the new live-line mock.
  • At 38567b27758 (this head, after merging feat(native-chat): the chat reads "Stopping…" from your Stop until the turn actually ends #24369 at 970c6a10eb5, which carries main through Add a Chat settings page for structured native chat #25685): this PR's own root test files plus the rewind and Stop tests that share its code (15 files): 197 pass. tc:web passes from a fresh build cache. The phone's files were not re-run locally; CI runs them.
  • At ac6d23a3a03: this PR's own test files and their neighbours (22 files, including feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073's tests and the Codex echo, steer, options, adapter and late-settlement tests): 323 tests pass. The phone's 3 files: 13 pass.
  • Before that, after the send-order helper replaced the non-null assertions: the placement, turn membership, desktop render, rejected-in-place and both end-to-end files (7 files, 108 tests), and the phone's 3 files (13 tests), pass.
  • At ac6d23a3a03, the cross-version wire suite, run serially: 15 of 19 files pass, including feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073's downgrade test. The other 4 cannot load locally: this worktree's installed packages predate main's new stream-json version, and they fail on that import before running a test. CI runs them, and they pass there.
  • At ac6d23a3a03, tc:node and tc:web reported only that same missing stream-json module, in main's files. Mobile tsc and the mobile tests typecheck ratchet pass.
  • oxlint and the changed-code quality gate pass.

On a phone (Android emulator, with temporary diagnostic logging):

  • After a message the agent never started and a Stop from the desktop, the phone showed "Stopped before the agent started" under the message within about 6 seconds of the Stop. This held on a fresh chat and on a long chat with a message sent while stopping just before.
  • An earlier phone run once showed that row only after the next message was sent. The two runs with logging did not reproduce it, and its cause is not known.
  • The one live failure seen in these runs was the connection itself: under very heavy machine load, the desktop dropped the phone's connection and the phone did not reconnect. That earlier run's logs show no reconnect, so this does not explain it. It is tracked in STA-9496 (https://linear.app/stably/issue/STA-9496).

Projection cost. The message projection runs on every streaming update, on desktop and on the phone, so the Stop placement must cost nothing in a chat without a stopped message and grow only with what is loaded. A chat with no stopped message skips placement entirely, and the turn-claim index is built only when a taken-back message exists. Otherwise the turn anchors, the row lookup and (on an older host only) the earlier-send floor are built once per update, and each stopped message is placed with a binary search over the turns' furthest rows, not a rescan of every row. The earlier benchmark was run on the previous placement and has not been re-run on this one.

Review

Agent skill upstream boundary

  • Not applicable, or this change follows docs/reference/agent-skill-sharing-upstream-boundary.md and copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.

Notes

  • SSH/remote: the execution host computes the published position and records the turn from its own journal; clients only draw them. Nothing about where work runs changes.
  • Mobile: the phone gets the same projection and draws the row.
  • Cross-platform: no platform-dependent code.
  • Backwards compatibility: see Mixed versions above.

Checklist

  • This PR is small and focused
  • I explained what changed and why (ELI5, the user-facing before/after, the mechanism, and why over the alternatives)
  • Before/after screenshots or videos attached for UI changes, or N/A with reason
  • Self-reviewed for correctness, security, and performance
  • Cross-platform, SSH/remote, and path/shortcut impact considered (or N/A)
  • pnpm lint, pnpm typecheck, pnpm test, and pnpm build pass (or CI will cover; local preferred)

…journal rows

Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.

One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.

The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.

Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.
…ead of held ones

A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.
… in its transaction

The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.
A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.

Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
  again" at one instant, before a queue ignoring the pause re-sends. They
  now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
  whether or not a person's turn lifts it; it now reads the Stop's pause
  before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
  queued before a rewind.
…turn and caller

The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
… the at-start stop

Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.
…the Stop's timing

Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.
… Stop events

Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.
… a writable downgrade

The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.

Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.
A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.

It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.
A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.

An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.
… came before the turn showed

A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.

A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.

Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.
…ing a value no build writes

A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.

The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.
…ore the turn ends

Every stop that ends work now writes the Stop's event before it ends the child: a
person's close of the chat, an eviction (worktree teardown, orchestration stop, tab
cleanup) and the idle sweep's stop of a start that never landed. A stop that ends
nothing writes nothing, and quit writes none: its resume marker records why.

The turn-end write reads the latest Stop event where every turn row is built, so the
adapter's settle, the host's fallback and the relaunch's settle all agree: a turn a
person's Stop or close named, ending with no verdict of its own after that Stop, ends
as their cancellation. A relaunch's probe-bounded end is no earlier than a Stop that
found the turn running. When the provider refuses the interrupt and the turn runs on,
a refusal row answers the Stop, so a later crash still reads Failed; pressing Stop
again after a refusal is a new Stop.
…2025/stop-withdrawn-shown

# Conflicts:
#	mobile/src/session/mobile-structured-send-delivery.test.ts
#	mobile/src/session/mobile-structured-send-delivery.ts
#	src/renderer/src/components/native-chat/structured-agent-session-message-projection.rejected-in-place.test.ts
#	src/shared/structured-agent-session-message-projection.ts
…-state-on-b

# Conflicts:
#	mobile/src/session/MobileNativeChatView.tsx
#	mobile/src/session/use-mobile-structured-agent-session.ts
#	src/renderer/src/components/native-chat/NativeChatComposer.test.tsx
…e limits after the main merge

The Stop control now disables itself while Stopping, so the composer passes
the flag through and its test file stays as main has it. The phone chat
header's Stop moves to its own component.
…2025/stop-withdrawn-shown

# Conflicts:
#	src/shared/native-chat-turn-membership.ts
…n every snapshot

Every snapshot projects each item through the Stop-note read, and each new
snapshot rebuilds the index of Stop notes by turn. Both parsed every item's
key first; they now check the row's kind and turn scope (and, for the
projection, its unconfirmed-stop failure) before the parse. Every Stop note
is a status row, so what each finds is unchanged.
@brennanb2025
brennanb2025 changed the base branch from brennanb2025/stopping-state-on-b to main October 6, 2026 19:18
@brennanb2025

brennanb2025 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Review status: draft, base main, head cc6d838

The problem. If you pressed Stop before the agent started on your message, the message vanished from the chat and its text reappeared in your typing box. Other devices just lost it, and a "Cancellation requested." line could land under the previous, finished turn.

What changes for you.

  • A message a Stop takes back before the agent starts it stays in the chat where you sent it, with one line: "Stopped before the agent started". Your typing box is left alone, and every device shows the same thing.
  • Nothing you sent is lost. Text the host never received still shows as "Message was not sent." with Retry.
  • Several taken-back messages keep the order you sent them in.

The stack: #24864, #24369 ("Stopping…") and #25073 (the host publishes where each message was sent and answered) are on main. This PR is now based on main, and its diff is only its own change (45 files).

Verified.

  • CI at cc6d838: 16 pass, 16 skipped, 0 fail. One earlier run lost its package job to GitHub's setup-node action crashing before any code ran ("Unexpected end of JSON input"); the re-run passed.
  • Merging main, including feat(native-chat): the chat reads "Stopping…" from your Stop until the turn actually ends #24369 and fix(native-chat): keep a message accepted before a quit or crash as a held card #24660: fix(native-chat): keep a message accepted before a quit or crash as a held card #24660 (a message kept as a held card after a quit or crash) changed the same projection and phone-send code.
    • Its "kept as a card" rule runs first, so a card always owns its text. This PR's "taken back by a Stop" rule follows.
    • This PR's own lines are unchanged by the merge.
    • Three tests changed mechanically for main's API.
  • A review of the merge found one real bug, fixed at cc6d838. The merge left the Spanish translation file with two "notices" sections, and the later one won, so Spanish lost "Stopped before the agent started". The line now sits in the one section, as in English; no language file has a duplicate key, and all six carry the line.
    • Also fixed: a phone test meant to compare "Stopping…" between two ways of loading the chat read the wrong field, so it compared nothing. It now sees "Stopping…" during the Stop.
    • And one stale comment that said a taken-back message goes back to the typing box.
  • Tests: at cc6d838, the changed phone test, the outbox test and 4 language-file tests pass, and phone tsc and the phone tests-typecheck ratchet pass. The full set (10 desktop files, 7 phone files) was not re-run at this head; it last passed at a200a85, and the changes since are a language-file move, one comment and one test. tc:web, tc:node, phone tsc and the phone tests-typecheck ratchet pass, and oxlint is clean on every changed file.
  • Live desktop QA, on a separate Mac with a stand-in Codex, at 9f3cd5e (round 16, screenshots in the PR body). Cases 1 to 5 passed. In case 6, the first chat read empty for about 6 s after it was activated. That was filed separately as STA-9571 (low), and this PR's rows are drawn once the read fills. Since then only merges have touched this branch; this PR's own lines are unchanged.

Not verified.

  • Live QA on this head. The changes since round 16 are merges, covered by the tests above.
  • The phone, live. Earlier Android rounds are advisory until Brennan confirms them himself; one phone check (a steered message's stop row) was inconclusive. Phone behaviour is covered by unit tests.
  • A real Claude or Codex agent, and SSH hosts. Covered by tests, not run live.

@brennanb2025
brennanb2025 marked this pull request as ready for review October 6, 2026 20:35
@brennanb2025
brennanb2025 marked this pull request as draft October 6, 2026 20:36
@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: acea8f4d-0d34-4090-ad50-28968b548e6b
📥 Commits

Reviewing files that changed from the base of the PR and between a200a85 and cc6d838.

📒 Files selected for processing (3)
  • mobile/src/session/mobile-native-chat-merged-snapshot-parity-hooks.test.tsx
  • src/renderer/src/i18n/locales/es.json
  • src/shared/structured-agent-session-outbox-reconcile.ts
🚧 Files skipped from review as they are similar to previous changes (1)
  • src/renderer/src/i18n/locales/es.json

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 4 remain after this review.


📝 Walkthrough

Walkthrough

The change adds transcript projection and placement for sends withdrawn before the agent starts. It updates mobile send delivery and composer restoration behavior. Chat rendering displays a localized stopped-before-start notice and excludes those messages from rail ticks. New tests cover turn placement, outbox behavior, host journal frames, and parity between incremental mobile updates and fresh snapshots.

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to cc6d8

Withdrawn messages can remain visible with localized stop text. No actionable merge-blocking risk is evident; merge after normal checks.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 45.78% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 83 functions across 39 files. (1 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: Stop-withdrawn messages remain visible with one stop row after them.
Description check ✅ Passed The description covers the change, rationale, linked issue, visual proof, testing, compatibility, and notes. It reports validation gaps, leaves the Review section blank, and does not check the full li…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 45.78% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 83 functions across 39 files. (1 skipped: 1 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1


ℹ️ Review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 2868e28d-d9f9-4295-8a16-602d3585c45c
📥 Commits

Reviewing files that changed from the base of the PR and between 4c35516 and a200a85.

📒 Files selected for processing (45)
  • mobile/src/session/mobile-native-chat-host-journal-frames.test-fixture.ts
  • mobile/src/session/mobile-native-chat-merged-snapshot-parity-hooks.test.tsx
  • mobile/src/session/mobile-native-chat-merged-snapshot-parity.test.ts
  • mobile/src/session/mobile-native-chat-stop-journal.test-fixture.ts
  • mobile/src/session/mobile-native-chat-stopped-before-start.test.tsx
  • mobile/src/session/mobile-structured-agent-session-send-bypass.test.ts
  • mobile/src/session/mobile-structured-agent-session-send.ts
  • mobile/src/session/mobile-structured-send-delivery.test.ts
  • mobile/src/session/mobile-structured-send-delivery.ts
  • mobile/src/session/use-mobile-native-chat-drafts-glued-pending.test.ts
  • mobile/src/session/use-mobile-structured-agent-session-send.test.tsx
  • mobile/tsconfig.json
  • src/main/native-chat/agent-session-wire/structured-agent-session-stop-note-withdrawn-send.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-turns-cancel.ts
  • src/main/runtime/structured-agent-session-codex-stopped-send-order.test.ts
  • src/renderer/src/components/native-chat/NativeChatMessageList.stopped-before-start.test.tsx
  • src/renderer/src/components/native-chat/NativeChatMessageRow.tsx
  • src/renderer/src/components/native-chat/native-chat-message-rail-items.ts
  • src/renderer/src/components/native-chat/native-chat-rail-outline-parity.test.ts
  • src/renderer/src/components/native-chat/native-chat-stopped-before-start-row.ts
  • src/renderer/src/components/native-chat/native-chat-transcript-slots.ts
  • src/renderer/src/components/native-chat/structured-agent-session-message-projection.rejected-in-place.test.ts
  • src/renderer/src/components/native-chat/structured-agent-session-withdrawn-message-restore.test.tsx
  • src/renderer/src/components/native-chat/structured-agent-session-withdrawn-message-restore.ts
  • src/renderer/src/components/native-chat/use-structured-agent-session-outbox-ownership.ts
  • src/renderer/src/components/native-chat/use-structured-agent-session-outbox-withdrawal.test.tsx
  • src/renderer/src/components/native-chat/use-structured-agent-session-outbox.draft-hand-off.test.tsx
  • src/renderer/src/components/native-chat/use-structured-agent-session-outbox.ts
  • src/renderer/src/components/native-chat/use-structured-agent-session-stop.ts
  • src/renderer/src/i18n/locales/en.json
  • src/renderer/src/i18n/locales/es.json
  • src/renderer/src/i18n/locales/fr.json
  • src/renderer/src/i18n/locales/ja.json
  • src/renderer/src/i18n/locales/ko.json
  • src/renderer/src/i18n/locales/zh.json
  • src/shared/agent-session-conversation-outline.ts
  • src/shared/native-chat-send-order.ts
  • src/shared/native-chat-stopped-before-start.ts
  • src/shared/native-chat-stopped-send-turn.test.ts
  • src/shared/native-chat-turn-membership.ts
  • src/shared/native-chat-types.ts
  • src/shared/structured-agent-session-dispatch-rejection.ts
  • src/shared/structured-agent-session-message-projection-stopped-before-start.test.ts
  • src/shared/structured-agent-session-message-projection.ts
  • src/shared/structured-agent-session-outbox-stop-withdrawal.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 5 remain after this review.

Comment thread src/renderer/src/i18n/locales/es.json Outdated
… in the one notices block; the phone parity test reads Stopping from the live line
@brennanb2025
brennanb2025 marked this pull request as ready for review October 6, 2026 21:13
@brennanb2025
brennanb2025 merged commit e813fa5 into main Oct 6, 2026
56 checks passed
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
… typecheck passes

Main's typecheck fails here too: #25721 made resolveLaunchArgs required and
this test (#25051) predates it. Same line, same place as the open main fix
(#25977), so merging main after it lands is a no-op.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
…send-queue

Brings in #25051: a message a Stop took back stays in the chat with one
stop row after it, and the composer is left alone. The in-memory sender
now finishes a host withdrawal as recorded and hands nothing back, so a
withdrawn message, a note included, stays visible in the chat. The
saved-outbox files #25051 touched stay deleted; its shared projection
and phone side come in as they are. A send still in its pre-send checks
at the Stop never went out, so it still goes back to the draft.
brennanb2025 added a commit that referenced this pull request Oct 7, 2026
…otocol (#25225)

* Leave a stopped turn's running tools to the agent's own end

When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.

An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.

* Pin that a stopped turn's running tools hold budget until the agent's end

* Type the stopped turn's tool progress update as a tool body

* List every event the assembler hands to the decision step

The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.

* End a running call as its turn's journal row ends

A call still running when its turn ends takes the state of that turn's
row: a row another writer settled first (a person's Stop) stands, so its
calls read interrupted whatever the provider's later end reports. The
no-ending path that settled calls from the Stop row is gone, since a Stop
now leaves running calls to the provider. Adds the two Spanish strings.

* Say why a Grok turn failed, and keep task rows in Grok's own words

A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.

A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.

A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.

* Read a monitor from Grok's exact output prefix

* Word a failed Grok turn in Grok's own text, not as a refused message

A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.

* Settle a stopped turn's running call as its turn row ended after a restart too

The restart sweep ended every running call by the death evidence alone, so after
a person's Stop with no proof the child died the call read failed under a turn
that read interrupted. The sweep and the live dead-generation settlement now ask
the same rule the assembler does: a call in a turn already settled ends as that
row ended; only a turn still running leaves its calls to the evidence.

* Keep the dead-generation settlement under the line cap

* Register the ACP schema verify step in the PR preflight phase test

* refactor(agent-session): one required agent registry; declarations admit what they claim

A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.

/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.

Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).

* Read ACP permissions, session events and prompt errors through the protocol client's own types

The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.

* refactor(agent-session): the router applies the declared rewind itself

The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.

* test(agent-session): register the agents the merged-in tests now need

The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.

* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record

* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop

The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.

The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.

One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.

* fix(agent-session): a changed agent definition never hides that agent's chats

A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.

Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.

* refactor(agent-session): each agent's registration says where it runs and which account it pins

createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.

Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.

* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state

A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.

* fix(agent-session): a scoped dismiss-all persists no per-session fence

The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.

* fix(agent-session): refuse an attach whose agent is not the session's own

The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.

* fix(agent-session): offer to start a chat only when the start would accept it

The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.

* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it

A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.

* refactor(agent-session): the record store admits agent ids; comments say where transport is checked

The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.

* docs(agent-session): the record store admits the registered agents' ids

* refactor(native-chat): Grok's registration declares where it runs; ACP no longer borrows Codex's location rule

The rule a self-supervised agent child runs under (this machine, no WSL, Windows only with process
start-time proof) is its own module that Codex and the ACP adapter both use. Grok's registration
takes the full account-home resolver signature, and D3's tests build hosts with the agent registry.

* fix(native-chat): Grok follows the ACP runtime's request contract and the managed process's close

A request the agent or a Stop cancels is answered with the agent's own cancelled reply by the code that
owns it (the runtime no longer answers a silent handler), so a Stop needs no separate decline pass. A
permission answer still being saved when the agent stopped waiting is reported unconfirmed, since the
protocol already answered it cancelled. Cancelling the agent's own turn is the plain cancel. Request
rows are matched under their generation-scoped ids. A refusal's reason comes from the dialect's wording
path. The child drops its own stderr tail and close policy for the managed process's, and a close
whose process tree was not proven gone is reported as the adapter contract asks.

* fix(native-chat): a Grok chat Orca already holds resumes without writing what Grok replays

A chat with a saved Grok session reattaches with session/resume where the agent offers it, else
session/load. Either way the call runs inside the translator's load window, so what Grok sends while
it reattaches (its saved exchange, a task the dead process left running, ended by the restart) opens
no turn and writes no row; only context usage reads on. A reply an Orca or Grok crash cut short is no
longer completed from Grok's saved history: it reads like a Claude or Codex chat's, with the existing
notice. The attach window also closes after a failed attach, and a created session that session/resume
reports missing is replaced like one session/load reports missing.

The replay reconciliation is removed: the lane no longer reads the journal, and D3's replayed-input
grammar test and completed-turn check in the assembler go with it.

* refactor(native-chat): a failed Grok reattach needs no window close of its own; its lane is replaced

* test(native-chat): D3's merged tests use the shipped declarations and the launch options main requires

* fix(native-chat): typecheck fallout of the base merges; any agent's empty chat is reusable

Main's idle-empty-chat lookup and launch join now take any registered agent, as the rest of the
launch path does. The refusal check moved into the prompt turns and the prompt-block conversion beside
the turns that send it, keeping both files in their line limit.

* fix(native-chat): a Grok Stop ends the process once Grok settles its turn; the next send resumes

Grok's session/cancel ends only the running turn: work it already moved to the background keeps
running and can begin a turn of its own after the person pressed Stop. Stop is now a session
boundary, as it is for Claude: the cancel answers open requests and lets Grok end the turn, the host
waits a bounded grace for that, then ends the process; the next send relaunches and resumes.
The adapter's own bounded close of a turn Grok began is gone. Its named-turn check stays: the host
ends the session unless the provider declines a Stop naming a turn that has since ended.

* test(native-chat): a Grok Stop ends the process only after Grok answered the cancel

* fix(native-chat): Steer on a Grok card cancels the running prompt, then sends it

A send that reached Grok while a prompt ran was held in the adapter until that turn ended: Steer
on a queued card took the card out of the host's editable queue and meant 'send after this turn'.
It now cancels the running prompt (session/cancel; the session stays) and sends as the next prompt
once Grok answers the cancel, as the common pattern does; a steer behind another cancels it in
turn, so the last one runs. The adapter holds a send only while that cancel lands, so its general
held-send queue and its holdsDispatch report are gone (every send it holds has its turn open in
the journal). An older client's mid-turn send takes the same path. capabilities.steering is
unchanged and still unread.

* refactor(native-chat): a close or Stop cancels a start through the acquire's own abort signal

The host owns the acquire it runs, so it now owns its cancellation: each attach's acquire gets an
AbortSignal, aborted from outside the session's queue by a close and by a Stop admitted now (the
same admission rule as before). The optional abandonStart adapter hook, the router's fan-out to
every adapter and the ACP adapter's session-keyed start map are gone; the ACP adapter keeps an
unkeyed set of starts only so quit can prove their children gone, and keeps a failed start's
unproven child until its exit is proven.
The hook also let a later close ask that child again. The host now does that from state it holds:
a close of a chat with no live child whose record still names an owner process with no death
evidence asks the adapter to release it. The answer is not recorded as proof (the lease probe
does that), so an owner pid an earlier Orca left is never killed or marked gone. Claude and Codex
ignore the signal and hold no such child; their release is a no-op (tested).

* fix(native-chat): a Grok crash that closes stdout before its exit still ends with Grok's last words

On macOS and Linux the agent's stdout ends before its exit is observed, with or without the
supervisor's EOF forwarding, so the connection's loss closed the journal first and its error text
became the session's ended reason, dropping Grok's stderr. The reason is now read at the proven
exit: the agent's last words when it left any, else why the connection closed. The failure already
carried them. Comments that assumed the exit comes first, that early frames past the cap refuse the
start, and that dispatch re-checks image support are corrected.

* fix(native-chat): nothing Grok sends while a held chat reattaches is written, marked as replay or not

The reattach window relied on the dialect's replay verdict, and Grok's frames read as live unless
they carry isReplay, so an unmarked chat frame during session/resume opened a turn that never
ended. D3 now marks every frame inside the window as replay before the translator reads it, so the
translator keeps only context usage whatever the agent marked; options and commands are still
adopted. The translator's load semantics are unchanged.

* test(native-chat): a Stop after a resume finds no turn an unmarked old reply opened

* test(native-chat): a resumed Grok chat keeps its last context reading; the resume refreshes only the window

* test(native-chat): a Grok background task a Stop ended reads as stopped reporting

* refactor(native-chat): quit's stop of each start answers through one promise kind

* fix(native-chat): quit aborts every start the host has in flight before draining attaches

A Grok that never answered its handshake held quit until the start's own 60 s bound, past the
20 s quit deadline. The host's teardown now aborts each in-flight acquire (and any the drain
still begins), so the adapter's own quit controller and its map of starts are gone: a start
has one canceller, the host's signal.

* fix(native-chat): a Grok start's abort stops reaching its child once the start has returned

The listener stayed on the host's signal until the attach finished committing, so a Close in that
window killed the now-live child behind the host's back and it read as Grok crashing. The start
now detaches it when it ends; a later Close goes through the session's own stop.

* fix(native-chat): a close or Stop during any attach phase stops the start before it launches

The attach began its abort controller only after reconciling leases, resolving recovery and
probing the previous owner, so a close or admitted Stop in those phases reached nothing and Grok
launched anyway. The controller now begins first, and the acquisition checks it before asking the
adapter to start.

* test(native-chat): a close during the attach's owner probe asks no adapter to start

Also renames the close test after the hook it no longer exercises.

* test(native-chat): a close's re-ask closes a Claude or Codex child a failed cleanup left

The re-ask is not a no-op for them: when the adapter still holds the child its cleanup could not
prove gone, the close stops it again as a requested close, and Claude persists the handle of the
conversation it ran so the next send resumes it. Corrects the tests' and comment's wording; the
close awaits the re-ask, bounded by each adapter's kill ladder.

* fix(native-chat): Steer during a turn Grok began itself cancels it and sends once it ends

A send while Grok ran a turn of its own (a background task waking it) went straight to Grok, which
queued it behind that turn where Orca could no longer withdraw it, while Stop treated the same turn
as the running reply. The send now waits as a steer, the turn is cancelled once, and the message
goes when the turn ends; a Stop withdraws it and an exit rejects it as never sent.

* test(native-chat): a steer whose cancel Grok never answers ends Grok and is rejected as never sent

Pins the bounded steer cancel kept from the runtime: past the bound the connection closes, the
running reply reads unverifiable, Grok's end reads as its exit, and the waiting steer is rejected
as never sent.

* fix(native-chat): a Grok crash stays a crash when a stop lands before its exit is proven

After the connection broke and the close could not prove Grok's exit, any later stop Orca asked
for (the next start, a Stop, a Close) marked the child as closed by Orca, so the crash read as a
requested close and Grok's last words were dropped; a send meanwhile was recorded unconfirmed.
The connection loss now decides the cause, and a send on that session is rejected as never sent.

* test(native-chat): fixtures this PR's registered Grok and desktop capability made stale

CI's unit shards failed on tests outside the PR's own lists. Each encodes something this PR changes
on purpose: Grok is now a registered agent (the seam test's unregistered agent is now Cursor); the
desktop now advertises registered agents (the restart-offer tests' older client drops that
capability explicitly); the attach context carries the start's abort controllers (the forget-status
double gains them); and the ACP real-host test rig sends to the host directly (listed beside the
other real-host rig in the send ratchet).

* fix(native-chat): a start quit stops is not the queued message's start failure

With quit now aborting a start it would have waited for, the delivery step recorded the aborted
start as the message's failure ("couldn't restart"). After quit has stopped delivery, the step
leaves the message to quit, which settles it as a close does ("The chat closed before this message
was sent."). The test that pinned quit waiting for that start and stopping its child now pins that
nothing is launched behind quit.

* fix(native-chat): a message sent after a Stop or close aborted a start gets its own start

A start the host aborts (an admitted Stop, a close, or quit) returned its refusal to the delivery
loop, which then rejected whatever was queued at that moment with "couldn't restart", including a
message the user sent after the Stop. The attach now reports that the host aborted it, and the loop
re-derives from the journal instead: what the Stop or close withdrew is already settled, a message
accepted since gets a start of its own, and quit's next step stops the loop. This replaces the
quit-only carve-out with the same rule for every abort and every agent.

* test(native-chat): the message sent after an aborted start is answered, so no settlement outlives the test

* fix(native-chat): a Grok model pick Grok never answers no longer holds Stop or Close

The pick runs on the session's queue. It now registers in the host's out-of-queue
abort registry beside a start, so a close, an admitted Stop or quit abandons it, and
the ACP adapter bounds it at 30 s like Claude and Codex. A late answer is still adopted.

* fix(agent-launch): a phone's launch opens a terminal for an agent whose chat it cannot show

agent.launch now reads the caller's capabilities by the rule tabs and restart offers
use (clientRendersStructuredAgent). A phone without registered-agents.v1 gets Grok as
a terminal again, as on main; the host's own callers and desktop clients are unchanged.

* fix(acp): strip every agent hook variable from the ACP child, from the shared list

ACP_CHILD_ENV_TO_DELETE was a second copy of the hook runtime keys that missed
ORCA_AGENT_HOOK_TRANSPORT; it now spreads AGENT_HOOK_RUNTIME_ENV_KEYS beside the pane
identity keys.

* refactor(native-chat): the mutation context carries the provider-wait registry itself

Keeps the host file within its line limit; one field instead of two closures over it.

* fix(agent-launch): agent.launch.v2 still vouches for Claude and Codex chats

The caller rule from the previous commit also turned Claude and Codex into terminals
for a client advertising only agent.launch.v2, whose contract says it opens a chat
(mobile retry-authority tests). Only an agent beyond those two now needs the client to
read it (clientRendersStructuredAgent); the test fixtures go back to what they were.

* refactor(native-chat): drop saved-history adoption from the timeline assembler

The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.

* refactor(acp): drop session/load history adoption from the translator

The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.

* refactor(native-chat): a pending input is only Orca's send now

Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.

* test(acp): keep the task-result status table on live frames

Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.

* test(acp): a frame helper for a shell command Grok is running

* fix(acp): a Grok crash settles through the host's provider-exit batch, scoped to the turn it ended

A Grok crash ended the journal unverifiable before the adapter reported the exit, so the host's
provider-exit settlement found no running turn and wrote nothing: the adapter's failure (with
Grok's last words) never reached the journal, and a later stale-session pass wrote a bare,
thread-scoped cut-short row, so the partial reply was not folded as Claude's and Codex's are.

At a proven exit the ACP lane now ends its running turn interrupted at the exit instant, as the
host's exit contract expects of a child's own translator (Codex's does the same). When Grok's
stdout closed first (every POSIX crash), the turn is unverifiable only until the exit is proven:
the host's provider-exit settlement now takes the exit as proof naming the child's fence and
revises what that child left unverifiable in the same batch, with the turn-scoped row and the
adapter's failure. Claude and Codex write no unverifiable turn of a live child except a command
whose hand-off is in doubt; that turn is now revised at the exit instead of at the next open.

* test(acp): a crash seen first leaves the host no Grok turn to revise

* refactor(native-chat): what a gone generation left unfinished gets its own module

The settlement file passed 300 lines with the exit-proof revision. The unfinished-work reads
(capture, interrupted-by-the-exit, in-progress) are their own concept and move out unchanged,
apart from the exit proof they now take.

* refactor(native-chat): a watched exit revises what its child left unverifiable without reading Stop marks

An exit's own instant is the turn's end, so the revision needs only each row's fence: the
settlement's journal type gains itemFence alone, and the host test fakes say so.

* test(native-chat): drop the duplicate itemFence on the fake that already had one

* test(claude, codex): an exit whose stdout ended first still reports as it always did

The provider supervisor now ends Orca's stdout when the agent's ends, so on every crash EOF
arrives before the exit is seen. Claude's and Codex's connections report nothing at EOF and
report the exit, with its usual reason, once it is seen.

* fix(acp): reopen a chat with session/load, as the common pattern does

An agent that offers both now reloads its session instead of resuming it; the
reattach window still discards what it replays except context usage.

* fix(acp): drop the 60 s handshake bound; an abort fails the start's waits at once

Neither common design bounds an ACP handshake: Close, Stop and quit end a start
that never answers. The abort now also closes the connection, as a kill there
does, so the start settles even before the child's exit is proven. The
host-stopped start refusal only this bound produced goes with it; the idle
sweep keeps its words.

* fix(acp): a Stop naming an ended turn follows Claude's rule

It still stops nothing while another turn is live, but in the gap before a
follow-up's turn opens, which no client can name, it now stops what is in
flight and the session ends, as a Claude Stop does.

* fix(native-chat): a close no longer re-asks a failed start's unproven child

Neither common design retries that stop at Close, and Orca's Claude contract
re-asks only at the next start and at quit. The ACP adapter keeps the child
until its exit is proven and asks it again there, as Claude does.

* fix(acp): a message sent during a turn the agent began itself goes at once

Both common designs send it straight to the agent with no cancel; only Orca's
own running prompt is steered (cancelled, then re-prompted).

* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer

Uses an audience production sends (one that cannot show every agent), per review.

* fix(acp): launch Grok as `grok agent stdio`, without the update and leader flags

The common pattern passes neither --no-auto-update, --no-leader nor
GROK_DISABLE_AUTOUPDATER; full access still adds --always-approve.

* fix(acp): an agent that ends its stdout, or answers unreadably, is not a lost connection

As in the common pattern, only a broken stdin (or Orca's own close) ends the
agent; one that closed its output but can still be written to stays until a
Stop, a close or its exit. The provider supervisor goes back to its base
content, so Claude and Codex no longer get the forwarded stdout end either.

* fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts

On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.

Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).

* test(native-chat): Grok opens as a chat only behind the structured-chat setting

agent.launch and orchestration worker-start read the same setting as the
renderer route; pin both states for Grok on each. The setting's description no
longer names only Codex and Claude, in every catalog.

* docs(acp): generic ACP comments say what holds for every agent, not Grok

Stop ends the session for every ACP agent, as in the common pattern; the
adoption hook comment goes (adoption is not planned); a failed start's child is
retried at the next start or quit.

* test(claude, codex): type the EOF-before-exit test's streams; the supervisor no longer forwards EOF

The Claude test wrote to the child's stdout and stderr through their Readable
type, which the node typecheck rejects; it now holds its own PassThrough
streams. The comments no longer credit the reverted supervisor change.

* feat(acp): a steer's cancel asks once and never ends the agent

The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.

* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge

* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns

Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.

* fix(native-chat): drop the stopDelivery the A3 merge doubled

* fix(acp): a steer's cancel asks Grok once and never ends it

A steer now uses D1's notify-only cancel. Two messages sent during a reply Grok began itself
cut that reply, as the common pattern does, and then both run; before, the queued first
message could not answer the bounded cancel and Orca ended Grok although Grok answered.
A Stop keeps the bounded cancel and its 4 s grace.

* fix(acp): a permission Grok asks with no prompt of Orca's running is declined

During a turn Grok began itself nobody asked it to act, so the request is answered
cancelled at once instead of opening a card that waits, as the common pattern does.

* fix(acp): a Grok that dies while starting is reported with its own last words

A dying process's stdout ends before its exit is seen, so the start failed as a closed
connection and Grok's stderr was lost. A start whose connection closed now waits, bounded by
the Stop grace (or a Close/Stop), for the exit before it is told.

* test: a Stop after a steer sends its own cancel; drop the import the A3 merge doubled

* test(native-chat): main's Stop-note test builds its turn context with the agent registry

* test(claude): say why the close test's fake child cast is safe

* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.

* test(native-chat): build the Stop-opened-turn test's identity with the opaque handle

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 and this branch both added the import at different lines; the merge kept both.

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.

* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore

A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.

Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.

* test(native-chat): build this stack's journal identities with main's opaque provider handle

Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test
files from this stack still wrote the old shape. Same lines the downstream ACP branch uses.

* docs(acp): every reattach drops the agent's replay, not only for a chat the journal holds

* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row

Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.

* refactor(native-chat): read hosts' structured agents from the app-shell services

Main grew the startup hydration hook to its line limit; the host agents sync is an app-lifetime subscription like the structured session tabs sync beside it, so it moves there.

* Use current provider handles in transition tests

* Use current provider handles in timeline fixtures

* test(native-chat): prove replacement rows survive downgrade and re-upgrade

* Require the ACP directory in the runtime import check

* test(ratchet): require src/main/acp now that this PR lands it

* feat(acp): a saved session the agent cannot reopen continues in a new one, with one warning row

When session/load (or session/resume) of a saved ACP session fails, the chat starts a new session and records it as a creation that replaces the lost one (#25747's 'replaces' link), and writes one warning row that the agent no longer remembers the earlier messages. A created session the agent reports missing is still superseded silently; a signed-out agent or a start that is over (Close, Stop, a lost agent) still fails the start.

* chore(acp): rewrap the acquire header comment

* test(acp): a start closed while the agent reopens fails without opening or announcing a new session

* Let ACP connections own their supervised agent process

* Preserve ACP cleanup evidence and isolate exit observers

* Expose ACP cleanup observations and type the permission fixture

* refactor(native-chat): the registered-agents capability lives in its own module

Main's growth put protocol-version.ts one counted line over its 300-line limit once the capability
was added; like main's other per-feature capabilities, it now has its own module, and importers read
it from there.

* refactor(acp): one connection owns the Grok process and its protocol

D3 now opens each ACP agent through createAcpAgentConnection (ACP-ALIGN #25810): one object spawns the
process on the execution host, owns its stdio and protocol, and reports its proven exit. It is built and
tracked before the handshake, so a start's abort (Close, Stop, quit) still reaches it, and a failed start
keeps that same connection for the next close to retry rather than spawning another process.

Deleted: the spawnAcpStructuredChild wrapper and its test, the raw-stream runtime assembly, the caller's
exit -> runtime.close wiring, the stdout-EOF heuristic (the connection no longer treats stdout EOF as
exit), and the 10 s steer/Stop cancel bound with requestSteerCancel. Reader control maps to
pauseReading/resumeReading; a close is connection.close after the host's existing 4 s Stop grace.

The adapter owns what the protocol no longer does: one session/cancel per running prompt however many
steers arrive (cleared with that send's settlement, retried after a failed write), and a Stop or steer
answers every open agent request the person has not already answered with the agent's own cancelled
reply. An answer already being saved when the Stop lands is sent.

Tests: blocked cancel write never holds Stop's grace, two quick steers send one cancel, a failed cancel
write is retried, a real process exiting while a child holds its stdout ends the session, and the
existing start-abort, retention, crash, connection-loss and reload-failure suites on the new rig.

* fix(acp): Grok signs in on its own machine with its API key or cached sign-in

When Grok reports that it needs authentication, Orca now names a sign-in method on the machine Grok runs
on, read from the same environment Grok was launched with: xai.api_key when XAI_API_KEY is set there and
Grok offers that method, else cached_token when Grok offers it, else none and the chat keeps the existing
not-signed-in refusal. The rule lives in Grok's launch spec; the adapter applies any agent's rule for new
and reopened sessions through the protocol client's caller-named method (authenticate, then retry once).
No new sign-in UI; interactive methods are never chosen.

* fix(acp): the adapter decides which of Grok's requests reach the person

The turn owner now admits every agent request, permission or question, from its own turn state: a
request reaches the person only while Orca's prompt runs and no steer or Stop is cutting it short (a
question may also come from a turn Grok began itself, until a Stop). Anything else gets the agent's
own cancelled reply and opens no card, so a question arriving after Stop or during a steer never
appears. A steer, like a Stop, withdraws the requests already open; an answer already being saved is
still sent. The protocol client's abort-on-cancel path is no longer used: after the connection
change its request signal aborts only when the connection closes.

* fix(acp): a plan Grok proposes shows as a plan, with no approval card

When Grok leaves plan mode it asks the client to approve its plan (x.ai/exit_plan_mode). Orca showed a
blocking 'Approve plan / Request changes' card for it; the common pattern has no such gate. Now the
plan goes into the chat's existing Plan row (the plan-document status row Codex and ACP plan updates
already use) and the request is answered at once with 'abandoned' plus feedback telling Grok to stop and
wait for the person's feedback or a request to implement it in a later turn, so nothing is approved on
the person's behalf. Dialects gain settleRequest for requests answered without asking anyone.

* fix(orchestration): worker-start opens a Grok worker in a terminal, as before

With the structured chat setting on, worker-start decided 'structured' for Grok and then the structured
worker factory (Claude and Codex only) refused it, so the start failed; main opened a terminal Grok
worker. Worker-start now decides with no registered agents beyond Claude and Codex, so Grok gets a
terminal worker as before. agent.launch and the app's own launches still open Grok as a structured
chat. Temporary until structured workers take registered agents.

* fix(acp): a prompt answer Orca can't read ends the turn instead of hanging it

A session/prompt rejection that was not the agent's own error answer (an answer that fails Orca's
schema, or one too large to read) left the turn running: the next message became a steer with nothing
to cancel and was never sent or settled, and Stop waited its full grace. As in the common pattern, any
prompt failure now ends the turn as failed (a failed-turn row without words, since none are the
agent's) and settles the send, so the next message goes. Only a closed connection keeps the send
running, for the connection-loss path to settle.

* fix(acp): send Grok's prompt-identity extension only to agents that echo it

session/prompt carried _meta {promptId, requestId} for every ACP agent, though only Grok's dialect
echoes it (injectedPromptIdentity). Now only an agent whose dialect declares it gets the extension;
other ACP agents get a plain prompt.

* refactor(native-chat): the registered-agents capability lives in protocol-version again, as on main

This reverts 0ef6d21. That commit moved the capability to its own module only because main's
protocol-version.ts was then one counted line over its limit; main now defines it there itself within
the limit, and main's new restart test imports it from there. Main's test also reads the desktop's
capability list as an older client; on this branch the desktop advertises registered agents, so its
older client is that list without this one capability.

* test(acp): read the sign-in method with a schema, not a type assertion

* refactor(native-chat): composer transport and Stop control in their own modules

Main's rewind change (#19338) brought NativeChatStructuredSession.tsx and use-structured-agent-session.ts
to their line limits, leaving no room for this branch's image-acceptance and unpublished-Stop lines.
The composer's transport (sends, commands, options, image acceptance) moves to
use-native-chat-structured-composer-transport.ts, and whether Stop shows and what it does moves to
structured-agent-session-stop-control.ts. Behavior is unchanged; the runtime cast on the composer's
'local' | 'remote' is now a typed return.

* test(native-chat): read registered agents by id, as main's structuredAgentsReadBy now takes

Main's A3 squash changed structuredAgentsReadBy to take agent ids; this branch's test still passed
{ agent } objects (CI typecheck TS2322).

* fix(acp): a first reopen warns when Grok forgets a chat that exchanged turns

A Grok session the chat created was treated as one Grok never saved, so
when Grok reported it missing on the chat's first reopen, Orca swapped in
a fresh session silently, even after completed exchanges: the person saw
the old messages while Grok had forgotten them.

The launch now counts a created session as never saved only when the
chat's journal, read at the failed reopen, holds no turn of that Grok
session; anything else, an unreadable journal included, takes the normal
path: the fresh session is recorded as replacing the old one and the one
warning row is written.

* fix(acp): a start writes the warning row an earlier attach failure dropped

The row saying Grok forgot the chat was written only into the attach's
deferred sink, while the fresh session's link was saved earlier. An attach
failure, quit or crash in between dropped the row forever.

Every start now derives the owed rows: each conversation the chain says was
lost to a failed restore gets its row unless the chat's journal already holds
it. The row now names the lost conversation rather than the fresh session, so
a row an unused replacement wrote still counts after Grok supersedes it.

* fix(acp): a question during a turn the agent began itself gets its cancelled reply

A question or other card-opening request the agent sends while no prompt of
Orca's runs (a turn it began itself, as when a background task wakes it) now
gets the agent's own cancelled reply and opens no card, the same rule
permissions already follow there. Nobody is waiting on that turn. A plan the
agent shares in it is still shown.

* test(native-chat): import the unfinished-work capture from the module that owns it

Main's reasoning sweep test (#19221) imported it from the dead-generation settlement, which
this branch split it out of.

* refactor(native-chat): keep the session host within its line budget after main's Stop work

Main's #25949 left the host at exactly its 300-line budget, and this branch's acquire-abort wiring
adds one line. The reveal module now comes in as a namespace import, as the host already does for
its other helper modules, and the earlier reorder of two type imports is undone.

* test(codex): move the stdout-before-exit test into its own file

Main's connection test file is at its 800-line test budget, and this branch's exit-order test
pushed it over (CI lint, max-lines).

* fix(acp): cancel a running turn before a close, dispose or quit ends the agent

Closing a tab, disposing a session or quitting while an ACP agent's turn ran
killed the process without asking the agent to cancel first. A requested
close now does what a Stop does: withdraw the agent's open requests and held
steers, send session/cancel once, and wait for the turn to end, bounded by
the Stop's grace (4 s), before closing the process. A close that follows a
Stop sends no second cancel and waits only what is left of that Stop's grace.
An idle close, a lost connection and a sink-failure force close are
unchanged. The cancel-and-wait moves to acp-structured-stop.ts, shared by
Stop and close.

* test(codex): pass resolveLaunchArgs in the stopped-send-order test so typecheck passes

Main's typecheck fails here too: #25721 made resolveLaunchArgs required and
this test (#25051) predates it. Same line, same place as the open main fix
(#25977), so merging main after it lands is a no-op.

* test(native-chat): run the close-aborts-start test on the Node runtime, as its SQLite journal fixture requires

* refactor(native-chat): the Stop control reads the outbox itself, so the chat hook stays within its line limit

* Keep the session host under the line cap after main's two new delegates

Pass the lifetime's conversation opener to reveal directly; it is already passed unbound to the mutation context.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant