Repository navigation
feat(native-chat): keep a message a Stop took back on screen, with one stop row after it - #25051
Conversation
…op-keeps-steered-card
…journal rows Stop now appends one journal row where it takes effect, before the interrupt, whatever the queue holds; Resume appends its own. The pause is a pure function of the fold: the latest Stop with no later Resume and no later accepted turn a person asked for. A /clear's carried cards name their source, which is the replacement's 'cleared' pause. Host-origin turns never lift either. One predicate decides which cards a pause holds; by default every waiting card without a hold of its own, including one queued after the Stop. The drain's consume re-judges it inside its own transaction. The rows are tombstones of an id no item takes, carrying the mark: a released host reads an unknown row kind as corruption and truncates the journal there. Deletes the stored pause (recordPause, the retire hook on every appended row, the settle-before-record step, mayReturnToWaiting and its row overlay) and the tests that only proved it retires. The queued_message_pauses table stays in the schema, unread and unwritten, for downgrade safety.
…ead of held ones A Stop's pause now holds only the cards queued before its row, plus a steer it withdrew, which returns to its own place. Each card records the journal position it was queued at, and the one hold rule compares that with the Stop row. A card queued after the Stop is a new instruction: it sends as usual, but the drain still stops at the first held card, so it never overtakes them. /clear's pause holds the cards it carried. Holding every card again is a one-line switch in that rule.
… in its transaction The drain's pick and its consume now read one function, nextSendableQueuedCard, so a Stop row that lands between them holds a newer card behind an older held one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
…op-keeps-steered-card
The queue's pause is derived from journal rows, so nothing reads or writes queued_message_pauses. It was still created on every open "for downgrade safety", but an older build creates it itself when it opens the database, so the table only sat empty in every new database. The tests now pin that no pause table exists.
A Stop holds only the cards queued before it. The pause derivation still returned the Stop alone whenever it was in force, so the restart pause was never considered: a card queued after the Stop, written by a host process that has since exited, sent by itself after Orca restarted, with no pause header and no Resume. A /clear pause that held nothing could hide it the same way. Every pause in force is now derived. A card is held if any of them holds it, and it names the first that does. The drain's pick, the consume transaction's re-check and the published header all read that one rule; the header names the pause holding the first card Resume would send.
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent again" at one instant, before a queue ignoring the pause re-sends. They now wait for the stopped turn to end and re-check after a quiet window. - The deleted-card test read a card queued after the Stop, which sends whether or not a person's turn lifts it; it now reads the Stop's pause before and after that turn. - Unit cases pin that a Stop holds a card with no recorded position and one queued before a rewind.
…op-keeps-steered-card
…turn and caller
The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
…states the same Stop event
… the at-start stop Through the real host: the event names the turn and who asked and is in the journal when the interrupt reaches the agent; at an agent still starting it is there before the start is ended and holds a card queued before it; an idle Stop writes one only when it withdrew a send; and the queue's claim re-judges a pause that landed after its pick.
…the Stop's timing Also says precisely what the claim's in-transaction pause check defends against: the Stop and the drain share one serialized lane.
… Stop events Replays this build's rows from the released build's own journal database: every row is kept, the history after the Stop still folds, and an older client is sent only removed ids no item uses.
…op-keeps-steered-card
…op-keeps-steered-card
… a writable downgrade The Stop-event downgrade test ran in no CI lane: unit shards exclude the cross-version folder, and the cross-version lane runs a fixed file list that did not name it. It is now on that list. Its only case replayed the rows into a release's own fresh database, because that release cannot open the current host database. A second case opens the journal this build wrote with a main build that shares the database: it opens writable, keeps every row, appends, and this build then reopens it with the person's Stop still pausing the queue.
A Stop reaching a running agent wrote a Stop event on every press. Two presses before the first interrupt landed wrote two events, so a card queued between them counted as before the latest Stop and was held, though a card queued after a Stop should send normally. A Stop naming a turn that had already ended, as a phone sends late, also wrote an event for a turn it never stopped. It now writes one only when it withdrew a queued send, or stops something no event records yet: not a turn the journal no longer runs, and not the live turn a Stop still in force already names, unless a card was handed over into it since, which this Stop's interrupt sends back and must hold. The interrupt and the "already finished" note are unchanged. A Stop at a starting agent still always writes.
…r lifts a person's Stop
… say only user-stop is journaled
A person's Stop paused the queue until their next accepted turn or Resume, and a later Stop of another reason (the host stopping the agent, an eviction, a close) was ignored. Now the pause is the latest Stop event's: a later Stop of any reason ends a person's pause, and only a person's Stop pauses. The fold keeps the latest Stop event whatever its reason. An eviction of a resting chat writes no Stop event (a Stop that stops nothing writes nothing), so it cannot release held cards; a test pins that no event means no lift.
… came before the turn showed A Stop pressed before the agent's turn shows in the journal (before Claude's echo, or before Codex opens the turn) records no turn. A second press once the turn showed compared that missing turn with the live one, wrote a second Stop event, and held a card queued between the presses. A repeat is now judged by what was sent since the Stop in force: with nothing sent after it (a refused send aside), a Stop that named no turn, or named the live one, is repeated and writes nothing. Anything sent since and not refused, including a send whose fate is unknown, makes the new press write, since its interrupt may send that card back to waiting. Tests: the two-press case across the turn showing; a steer between the presses settled unknown; and a Stop naming a turn that ended while the next card is sent but shows no turn yet, which writes and holds that card. The fold test that claimed an eviction path is renamed.
…ing a value no build writes A Stop or Resume row's value is read from disk with no shape check, and the pause fold stored whatever it found. A stored `stopEvent: null` would then throw on every pause check for that chat: the queue's pick, its send, and every queue update to clients. No build writes such a row, so this is hardening. The fold now reads a Stop only when it is an object with a string reason and a finite time, and a Resume only when it is `true`. Anything else is ignored: it pauses nothing and ends nothing. The row is still not treated as malformed, which could cut the history short.
…ore the turn ends Every stop that ends work now writes the Stop's event before it ends the child: a person's close of the chat, an eviction (worktree teardown, orchestration stop, tab cleanup) and the idle sweep's stop of a start that never landed. A stop that ends nothing writes nothing, and quit writes none: its resume marker records why. The turn-end write reads the latest Stop event where every turn row is built, so the adapter's settle, the host's fallback and the relaunch's settle all agree: a turn a person's Stop or close named, ending with no verdict of its own after that Stop, ends as their cancellation. A relaunch's probe-bounded end is no earlier than a Stop that found the turn running. When the provider refuses the interrupt and the turn runs on, a refusal row answers the Stop, so a later crash still reads Failed; pressing Stop again after a refusal is a new Stop.
…op-keeps-steered-card
…2025/stop-withdrawn-shown # Conflicts: # mobile/src/session/mobile-structured-send-delivery.test.ts # mobile/src/session/mobile-structured-send-delivery.ts # src/renderer/src/components/native-chat/structured-agent-session-message-projection.rejected-in-place.test.ts # src/shared/structured-agent-session-message-projection.ts
…-state-on-b # Conflicts: # mobile/src/session/MobileNativeChatView.tsx # mobile/src/session/use-mobile-structured-agent-session.ts # src/renderer/src/components/native-chat/NativeChatComposer.test.tsx
…e limits after the main merge The Stop control now disables itself while Stopping, so the composer passes the flag through and its test file stays as main has it. The phone chat header's Stop moves to its own component.
…2025/stop-withdrawn-shown # Conflicts: # src/shared/native-chat-turn-membership.ts
…n every snapshot Every snapshot projects each item through the Stop-note read, and each new snapshot rebuilds the index of Stop notes by turn. Both parsed every item's key first; they now check the row's kind and turn scope (and, for the projection, its unconfirmed-stop failure) before the parse. Every Stop note is a status row, so what each finds is unchanged.
…2025/stop-withdrawn-shown
Review status: draft, base
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (3)
🚧 Files skipped from review as they are similar to previous changes (1)
Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 4 remain after this review. 📝 WalkthroughWalkthroughThe change adds transcript projection and placement for sends withdrawn before the agent starts. It updates mobile send delivery and composer restoration behavior. Chat rendering displays a localized stopped-before-start notice and excludes those messages from rail ticks. New tests cover turn placement, outbox behavior, host journal frames, and parity between incremental mobile updates and fresh snapshots. Priority: ➖ Normal Merge Risk: ⚪ Minimal · up to Withdrawn messages can remain visible with localized stop text. No actionable merge-blocking risk is evident; merge after normal checks. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 45.78% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 83 functions across 39 files. (1 skipped: 1 unsupported.)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
ℹ️ Review info
⚙️ Run configuration
- Configuration used: Repository UI
- Review profile: CHILL
- Plan: Advanced
- Run ID:
2868e28d-d9f9-4295-8a16-602d3585c45c
📒 Files selected for processing (45)
mobile/src/session/mobile-native-chat-host-journal-frames.test-fixture.tsmobile/src/session/mobile-native-chat-merged-snapshot-parity-hooks.test.tsxmobile/src/session/mobile-native-chat-merged-snapshot-parity.test.tsmobile/src/session/mobile-native-chat-stop-journal.test-fixture.tsmobile/src/session/mobile-native-chat-stopped-before-start.test.tsxmobile/src/session/mobile-structured-agent-session-send-bypass.test.tsmobile/src/session/mobile-structured-agent-session-send.tsmobile/src/session/mobile-structured-send-delivery.test.tsmobile/src/session/mobile-structured-send-delivery.tsmobile/src/session/use-mobile-native-chat-drafts-glued-pending.test.tsmobile/src/session/use-mobile-structured-agent-session-send.test.tsxmobile/tsconfig.jsonsrc/main/native-chat/agent-session-wire/structured-agent-session-stop-note-withdrawn-send.test.tssrc/main/native-chat/agent-session-wire/structured-agent-session-turns-cancel.tssrc/main/runtime/structured-agent-session-codex-stopped-send-order.test.tssrc/renderer/src/components/native-chat/NativeChatMessageList.stopped-before-start.test.tsxsrc/renderer/src/components/native-chat/NativeChatMessageRow.tsxsrc/renderer/src/components/native-chat/native-chat-message-rail-items.tssrc/renderer/src/components/native-chat/native-chat-rail-outline-parity.test.tssrc/renderer/src/components/native-chat/native-chat-stopped-before-start-row.tssrc/renderer/src/components/native-chat/native-chat-transcript-slots.tssrc/renderer/src/components/native-chat/structured-agent-session-message-projection.rejected-in-place.test.tssrc/renderer/src/components/native-chat/structured-agent-session-withdrawn-message-restore.test.tsxsrc/renderer/src/components/native-chat/structured-agent-session-withdrawn-message-restore.tssrc/renderer/src/components/native-chat/use-structured-agent-session-outbox-ownership.tssrc/renderer/src/components/native-chat/use-structured-agent-session-outbox-withdrawal.test.tsxsrc/renderer/src/components/native-chat/use-structured-agent-session-outbox.draft-hand-off.test.tsxsrc/renderer/src/components/native-chat/use-structured-agent-session-outbox.tssrc/renderer/src/components/native-chat/use-structured-agent-session-stop.tssrc/renderer/src/i18n/locales/en.jsonsrc/renderer/src/i18n/locales/es.jsonsrc/renderer/src/i18n/locales/fr.jsonsrc/renderer/src/i18n/locales/ja.jsonsrc/renderer/src/i18n/locales/ko.jsonsrc/renderer/src/i18n/locales/zh.jsonsrc/shared/agent-session-conversation-outline.tssrc/shared/native-chat-send-order.tssrc/shared/native-chat-stopped-before-start.tssrc/shared/native-chat-stopped-send-turn.test.tssrc/shared/native-chat-turn-membership.tssrc/shared/native-chat-types.tssrc/shared/structured-agent-session-dispatch-rejection.tssrc/shared/structured-agent-session-message-projection-stopped-before-start.test.tssrc/shared/structured-agent-session-message-projection.tssrc/shared/structured-agent-session-outbox-stop-withdrawal.ts
Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 5 remain after this review.
… in the one notices block; the phone parity test reads Stopping from the live line
…send-queue Brings in #25051: a message a Stop took back stays in the chat with one stop row after it, and the composer is left alone. The in-memory sender now finishes a host withdrawal as recorded and hands nothing back, so a withdrawn message, a note included, stays visible in the chat. The saved-outbox files #25051 touched stay deleted; its shared projection and phone side come in as they are. A send still in its pre-send checks at the Stop never went out, so it still goes back to the draft.
…otocol (#25225) * Leave a stopped turn's running tools to the agent's own end When another writer settles the open turn (a person's Stop), the assembler now only stops that turn's text and cancels its pending prompts. Running tool calls stay the agent's: a progress update or completion it reports after the Stop lands as reported, and whatever is still running settles at the agent's turn end for that turn, the next turn's open, or the session's end. An agent's end for an earlier turn while a newer one is open no longer clears the open turn's activity line or ends its anonymous reply. An unnamed end right after a Stop ends the stopped turn instead of being dropped. The test rig's restart no longer writes the dead assembler's window text, matching dispose. * Pin that a stopped turn's running tools hold budget until the agent's end * Type the stopped turn's tool progress update as a tool body * List every event the assembler hands to the decision step The type-aware lint requires an exhaustive switch with no default case. Also retitle a Stop test to say what it asserts. * End a running call as its turn's journal row ends A call still running when its turn ends takes the state of that turn's row: a row another writer settled first (a person's Stop) stands, so its calls read interrupted whatever the provider's later end reports. The no-ending path that settled calls from the Stop row is gone, since a Stop now leaves running calls to the provider. Adds the two Spanish strings. * Say why a Grok turn failed, and keep task rows in Grok's own words A failed Grok turn ended with no reason on screen: the translator dropped every copy of Grok's message. The failed turn now gets one status row in Orca's existing "provider did not accept this message" words with Grok's reason, read from whichever copy arrives first (the given-up retry, the turn's end, the prompt's completion notice, or the prompt's error answer); later copies only fill a reason the row still lacks. A running background command no longer reads "Background task <id> started": a task's summary is mapped only once it has settled. A monitor stays a monitor when the agent reads its output: a frame that names no kind keeps the known one, and a "[monitor" command is a monitor. A prompt's turn is marked started, so a late frame for an ended prompt neither reopens it nor becomes the active turn. A tool's turn is held in one place at a time. * Read a monitor from Grok's exact output prefix * Word a failed Grok turn in Grok's own text, not as a refused message A turn that started and then failed was told "The provider did not accept this message", Orca's sentence for a message refused before its turn. The row now reads as a Codex turn-ending error does: an error status row with the provider's own words. With no words, the dialect names the failure ("Grok ended this turn with an error." / "Grok usage limit reached."), else the agent's display name does. * Settle a stopped turn's running call as its turn row ended after a restart too The restart sweep ended every running call by the death evidence alone, so after a person's Stop with no proof the child died the call read failed under a turn that read interrupted. The sweep and the live dead-generation settlement now ask the same rule the assembler does: a call in a turn already settled ends as that row ended; only a turn still running leaves its calls to the evidence. * Keep the dead-generation settlement under the line cap * Register the ACP schema verify step in the PR preflight phase test * refactor(agent-session): one required agent registry; declarations admit what they claim A StructuredAgentRegistry, built once from the {definition, adapter} registrations, is now a required host dependency and the router routes with it. The adapter interface loses its optional router-only capabilities?() and definition?(); every reader (options at rest, thread goal, rewind, the /compact handover) asks the registry. A live session still narrows rewind through the adapter, and the host combines declared and narrowed in one helper. The registry refuses a registration that declares compact, a thread goal or rewind without the adapter method behind it. /compact is admitted by the declared capability, so an agent that declares compact:false gets the commandRefused fact instead of a thrown error. Also: the cut-turn notice names the agent from the catalog, create-support builds the account home through agentSessionAccountHome, and the turn status text older clients read names the session's agent instead of defaulting to Codex (Claude/Codex text unchanged). * Read ACP permissions, session events and prompt errors through the protocol client's own types The translator now reads a permission request with the client's lenient reader, a session update with its session-event reader, and takes only the agent's own error answer as a failed prompt's reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values newer than this build. * refactor(agent-session): the router applies the declared rewind itself The router holds the registry, so its rewindSupport answers the owner's declared rewind narrowed by the adapter, in one place no reader can bypass. The host-side combining helper is gone; readers ask the adapter they hold. * test(agent-session): register the agents the merged-in tests now need The adoption replay test builds its host over this build's registered agents instead of an adapter definition method, and the stored-form test passes the stored agents the record guard now requires. * chore(agent-session): the registry is the one lookup; drop the test-only empty capability record * fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop The desktop renderer calls the restart methods as a paired runtime client without the registered-agents capability, so it was handed a Claude/Codex audience and its "dismiss all" took the scoped path: no dismissal fence and no clearing of unwritten teardown witnesses, so a late teardown write could bring a dismissed offer back. The audience is now derived against this host's registry: a client that can show every agent the host registered gets none, and dismisses exactly as before (one fence for every offer). A client that truly cannot show some agent still gets a final dismissal for what it sees: each cleared session is fenced on its own (bounded, superseded by a dismiss-all fence) and only the witnesses of agents it sees are dropped. One predicate now answers which agents a client renders, for both tab projection and restart offers; the Claude capability rule is folded in. * fix(agent-session): a changed agent definition never hides that agent's chats A record was readable only if every handle used the transport its agent's definition declares today, and its account variable was declared by some registered agent. A later build that drives the same agent over another protocol, or under a renamed variable, would have set every existing chat of that agent aside: no tab, no history, no error. Readability now asks only that the record agree with itself: its agent is registered, every handle shares one namespace owned by that agent, and the account variable is a well-formed name. Whether this build can drive it (the chain's transport is the one the agent speaks, and the account variable is that agent's own, since it becomes the child's environment) is decided in performAttach, the one admission every start of every agent passes through, and refused there as hostUnsupported. A record pinning another agent's variable is refused the same way rather than passing because some agent declares it. * refactor(agent-session): each agent's registration says where it runs and which account it pins createSupport, create and the model catalog's account read still chose a location rule and an account resolver by name (Claude, Codex, else none), so a new agent would have needed a third branch beside the registration list that already decides storage, routing and publication. Each runtime registration now carries supportsLocation(location) and resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }), and Claude's and Codex's rules move into their entries unchanged: a read still syncs no home and starts no bridge, a Codex launch still trusts the workspace first, and WSL locations resolve as before. The list exists before the host is built, so none of these reads installs the host, and an agent the list does not hold is answered no without opening the journal. * test(agent-session): drop a duplicate registry key; keep fences out of the capsule state A spread input already carries the registry, so the explicit key was overwritten (TS2783). The session-fence helper now returns only the fence, not the capsule state it was given. * fix(agent-session): a scoped dismiss-all persists no per-session fence The per-session dismissal fence added a forever-persisted capsule field with an arbitrary cap that served no caller: only the local desktop calls the restart methods, and on a host whose agents it all shows it gets no audience and takes the unscoped, fenced dismissal. A scoped dismiss-all (a caller that cannot show some registered agent) keeps that agent's offers, clears only its audience's unwritten witnesses, and writes no fence; a late write from this process is serialized behind it. The field never left this unreleased branch. * fix(agent-session): refuse an attach whose agent is not the session's own The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal. * fix(agent-session): offer to start a chat only when the start would accept it The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired. * fix(agent-session): the model-catalog probe uses a record's account only if this build drives it A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record. * refactor(agent-session): the record store admits agent ids; comments say where transport is checked The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept. * docs(agent-session): the record store admits the registered agents' ids * refactor(native-chat): Grok's registration declares where it runs; ACP no longer borrows Codex's location rule The rule a self-supervised agent child runs under (this machine, no WSL, Windows only with process start-time proof) is its own module that Codex and the ACP adapter both use. Grok's registration takes the full account-home resolver signature, and D3's tests build hosts with the agent registry. * fix(native-chat): Grok follows the ACP runtime's request contract and the managed process's close A request the agent or a Stop cancels is answered with the agent's own cancelled reply by the code that owns it (the runtime no longer answers a silent handler), so a Stop needs no separate decline pass. A permission answer still being saved when the agent stopped waiting is reported unconfirmed, since the protocol already answered it cancelled. Cancelling the agent's own turn is the plain cancel. Request rows are matched under their generation-scoped ids. A refusal's reason comes from the dialect's wording path. The child drops its own stderr tail and close policy for the managed process's, and a close whose process tree was not proven gone is reported as the adapter contract asks. * fix(native-chat): a Grok chat Orca already holds resumes without writing what Grok replays A chat with a saved Grok session reattaches with session/resume where the agent offers it, else session/load. Either way the call runs inside the translator's load window, so what Grok sends while it reattaches (its saved exchange, a task the dead process left running, ended by the restart) opens no turn and writes no row; only context usage reads on. A reply an Orca or Grok crash cut short is no longer completed from Grok's saved history: it reads like a Claude or Codex chat's, with the existing notice. The attach window also closes after a failed attach, and a created session that session/resume reports missing is replaced like one session/load reports missing. The replay reconciliation is removed: the lane no longer reads the journal, and D3's replayed-input grammar test and completed-turn check in the assembler go with it. * refactor(native-chat): a failed Grok reattach needs no window close of its own; its lane is replaced * test(native-chat): D3's merged tests use the shipped declarations and the launch options main requires * fix(native-chat): typecheck fallout of the base merges; any agent's empty chat is reusable Main's idle-empty-chat lookup and launch join now take any registered agent, as the rest of the launch path does. The refusal check moved into the prompt turns and the prompt-block conversion beside the turns that send it, keeping both files in their line limit. * fix(native-chat): a Grok Stop ends the process once Grok settles its turn; the next send resumes Grok's session/cancel ends only the running turn: work it already moved to the background keeps running and can begin a turn of its own after the person pressed Stop. Stop is now a session boundary, as it is for Claude: the cancel answers open requests and lets Grok end the turn, the host waits a bounded grace for that, then ends the process; the next send relaunches and resumes. The adapter's own bounded close of a turn Grok began is gone. Its named-turn check stays: the host ends the session unless the provider declines a Stop naming a turn that has since ended. * test(native-chat): a Grok Stop ends the process only after Grok answered the cancel * fix(native-chat): Steer on a Grok card cancels the running prompt, then sends it A send that reached Grok while a prompt ran was held in the adapter until that turn ended: Steer on a queued card took the card out of the host's editable queue and meant 'send after this turn'. It now cancels the running prompt (session/cancel; the session stays) and sends as the next prompt once Grok answers the cancel, as the common pattern does; a steer behind another cancels it in turn, so the last one runs. The adapter holds a send only while that cancel lands, so its general held-send queue and its holdsDispatch report are gone (every send it holds has its turn open in the journal). An older client's mid-turn send takes the same path. capabilities.steering is unchanged and still unread. * refactor(native-chat): a close or Stop cancels a start through the acquire's own abort signal The host owns the acquire it runs, so it now owns its cancellation: each attach's acquire gets an AbortSignal, aborted from outside the session's queue by a close and by a Stop admitted now (the same admission rule as before). The optional abandonStart adapter hook, the router's fan-out to every adapter and the ACP adapter's session-keyed start map are gone; the ACP adapter keeps an unkeyed set of starts only so quit can prove their children gone, and keeps a failed start's unproven child until its exit is proven. The hook also let a later close ask that child again. The host now does that from state it holds: a close of a chat with no live child whose record still names an owner process with no death evidence asks the adapter to release it. The answer is not recorded as proof (the lease probe does that), so an owner pid an earlier Orca left is never killed or marked gone. Claude and Codex ignore the signal and hold no such child; their release is a no-op (tested). * fix(native-chat): a Grok crash that closes stdout before its exit still ends with Grok's last words On macOS and Linux the agent's stdout ends before its exit is observed, with or without the supervisor's EOF forwarding, so the connection's loss closed the journal first and its error text became the session's ended reason, dropping Grok's stderr. The reason is now read at the proven exit: the agent's last words when it left any, else why the connection closed. The failure already carried them. Comments that assumed the exit comes first, that early frames past the cap refuse the start, and that dispatch re-checks image support are corrected. * fix(native-chat): nothing Grok sends while a held chat reattaches is written, marked as replay or not The reattach window relied on the dialect's replay verdict, and Grok's frames read as live unless they carry isReplay, so an unmarked chat frame during session/resume opened a turn that never ended. D3 now marks every frame inside the window as replay before the translator reads it, so the translator keeps only context usage whatever the agent marked; options and commands are still adopted. The translator's load semantics are unchanged. * test(native-chat): a Stop after a resume finds no turn an unmarked old reply opened * test(native-chat): a resumed Grok chat keeps its last context reading; the resume refreshes only the window * test(native-chat): a Grok background task a Stop ended reads as stopped reporting * refactor(native-chat): quit's stop of each start answers through one promise kind * fix(native-chat): quit aborts every start the host has in flight before draining attaches A Grok that never answered its handshake held quit until the start's own 60 s bound, past the 20 s quit deadline. The host's teardown now aborts each in-flight acquire (and any the drain still begins), so the adapter's own quit controller and its map of starts are gone: a start has one canceller, the host's signal. * fix(native-chat): a Grok start's abort stops reaching its child once the start has returned The listener stayed on the host's signal until the attach finished committing, so a Close in that window killed the now-live child behind the host's back and it read as Grok crashing. The start now detaches it when it ends; a later Close goes through the session's own stop. * fix(native-chat): a close or Stop during any attach phase stops the start before it launches The attach began its abort controller only after reconciling leases, resolving recovery and probing the previous owner, so a close or admitted Stop in those phases reached nothing and Grok launched anyway. The controller now begins first, and the acquisition checks it before asking the adapter to start. * test(native-chat): a close during the attach's owner probe asks no adapter to start Also renames the close test after the hook it no longer exercises. * test(native-chat): a close's re-ask closes a Claude or Codex child a failed cleanup left The re-ask is not a no-op for them: when the adapter still holds the child its cleanup could not prove gone, the close stops it again as a requested close, and Claude persists the handle of the conversation it ran so the next send resumes it. Corrects the tests' and comment's wording; the close awaits the re-ask, bounded by each adapter's kill ladder. * fix(native-chat): Steer during a turn Grok began itself cancels it and sends once it ends A send while Grok ran a turn of its own (a background task waking it) went straight to Grok, which queued it behind that turn where Orca could no longer withdraw it, while Stop treated the same turn as the running reply. The send now waits as a steer, the turn is cancelled once, and the message goes when the turn ends; a Stop withdraws it and an exit rejects it as never sent. * test(native-chat): a steer whose cancel Grok never answers ends Grok and is rejected as never sent Pins the bounded steer cancel kept from the runtime: past the bound the connection closes, the running reply reads unverifiable, Grok's end reads as its exit, and the waiting steer is rejected as never sent. * fix(native-chat): a Grok crash stays a crash when a stop lands before its exit is proven After the connection broke and the close could not prove Grok's exit, any later stop Orca asked for (the next start, a Stop, a Close) marked the child as closed by Orca, so the crash read as a requested close and Grok's last words were dropped; a send meanwhile was recorded unconfirmed. The connection loss now decides the cause, and a send on that session is rejected as never sent. * test(native-chat): fixtures this PR's registered Grok and desktop capability made stale CI's unit shards failed on tests outside the PR's own lists. Each encodes something this PR changes on purpose: Grok is now a registered agent (the seam test's unregistered agent is now Cursor); the desktop now advertises registered agents (the restart-offer tests' older client drops that capability explicitly); the attach context carries the start's abort controllers (the forget-status double gains them); and the ACP real-host test rig sends to the host directly (listed beside the other real-host rig in the send ratchet). * fix(native-chat): a start quit stops is not the queued message's start failure With quit now aborting a start it would have waited for, the delivery step recorded the aborted start as the message's failure ("couldn't restart"). After quit has stopped delivery, the step leaves the message to quit, which settles it as a close does ("The chat closed before this message was sent."). The test that pinned quit waiting for that start and stopping its child now pins that nothing is launched behind quit. * fix(native-chat): a message sent after a Stop or close aborted a start gets its own start A start the host aborts (an admitted Stop, a close, or quit) returned its refusal to the delivery loop, which then rejected whatever was queued at that moment with "couldn't restart", including a message the user sent after the Stop. The attach now reports that the host aborted it, and the loop re-derives from the journal instead: what the Stop or close withdrew is already settled, a message accepted since gets a start of its own, and quit's next step stops the loop. This replaces the quit-only carve-out with the same rule for every abort and every agent. * test(native-chat): the message sent after an aborted start is answered, so no settlement outlives the test * fix(native-chat): a Grok model pick Grok never answers no longer holds Stop or Close The pick runs on the session's queue. It now registers in the host's out-of-queue abort registry beside a start, so a close, an admitted Stop or quit abandons it, and the ACP adapter bounds it at 30 s like Claude and Codex. A late answer is still adopted. * fix(agent-launch): a phone's launch opens a terminal for an agent whose chat it cannot show agent.launch now reads the caller's capabilities by the rule tabs and restart offers use (clientRendersStructuredAgent). A phone without registered-agents.v1 gets Grok as a terminal again, as on main; the host's own callers and desktop clients are unchanged. * fix(acp): strip every agent hook variable from the ACP child, from the shared list ACP_CHILD_ENV_TO_DELETE was a second copy of the hook runtime keys that missed ORCA_AGENT_HOOK_TRANSPORT; it now spreads AGENT_HOOK_RUNTIME_ENV_KEYS beside the pane identity keys. * refactor(native-chat): the mutation context carries the provider-wait registry itself Keeps the host file within its line limit; one field instead of two closures over it. * fix(agent-launch): agent.launch.v2 still vouches for Claude and Codex chats The caller rule from the previous commit also turned Claude and Codex into terminals for a client advertising only agent.launch.v2, whose contract says it opens a chat (mobile retry-authority tests). Only an agent beyond those two now needs the client to read it (clientRendersStructuredAgent); the test fixtures go back to what they were. * refactor(native-chat): drop saved-history adoption from the timeline assembler The common pattern discards the history a provider replays while loading a session, so the assembler has no use for an input.history event. * refactor(acp): drop session/load history adoption from the translator The common pattern discards the history an agent replays during session/load, keeping only what it says about the context window. Remove the adoption path (acp-history-adoption.ts, the adopt option, and the historical background-task liveness rewrite it fed) so load replay is always dropped except usage. * refactor(native-chat): a pending input is only Orca's send now Review follow-up to the adoption removal: drop the comment naming the provider's saved message, and make requestedAt required since every pending input comes from input.accepted. * test(acp): keep the task-result status table on live frames Review follow-up to the adoption removal: the result-status mapping was only tested through adopted history, so run the same table on live frames, and cover an unmarked task notice during a load being dropped. * test(acp): a frame helper for a shell command Grok is running * fix(acp): a Grok crash settles through the host's provider-exit batch, scoped to the turn it ended A Grok crash ended the journal unverifiable before the adapter reported the exit, so the host's provider-exit settlement found no running turn and wrote nothing: the adapter's failure (with Grok's last words) never reached the journal, and a later stale-session pass wrote a bare, thread-scoped cut-short row, so the partial reply was not folded as Claude's and Codex's are. At a proven exit the ACP lane now ends its running turn interrupted at the exit instant, as the host's exit contract expects of a child's own translator (Codex's does the same). When Grok's stdout closed first (every POSIX crash), the turn is unverifiable only until the exit is proven: the host's provider-exit settlement now takes the exit as proof naming the child's fence and revises what that child left unverifiable in the same batch, with the turn-scoped row and the adapter's failure. Claude and Codex write no unverifiable turn of a live child except a command whose hand-off is in doubt; that turn is now revised at the exit instead of at the next open. * test(acp): a crash seen first leaves the host no Grok turn to revise * refactor(native-chat): what a gone generation left unfinished gets its own module The settlement file passed 300 lines with the exit-proof revision. The unfinished-work reads (capture, interrupted-by-the-exit, in-progress) are their own concept and move out unchanged, apart from the exit proof they now take. * refactor(native-chat): a watched exit revises what its child left unverifiable without reading Stop marks An exit's own instant is the turn's end, so the revision needs only each row's fence: the settlement's journal type gains itemFence alone, and the host test fakes say so. * test(native-chat): drop the duplicate itemFence on the fake that already had one * test(claude, codex): an exit whose stdout ended first still reports as it always did The provider supervisor now ends Orca's stdout when the agent's ends, so on every crash EOF arrives before the exit is seen. Claude's and Codex's connections report nothing at EOF and report the exit, with its usual reason, once it is seen. * fix(acp): reopen a chat with session/load, as the common pattern does An agent that offers both now reloads its session instead of resuming it; the reattach window still discards what it replays except context usage. * fix(acp): drop the 60 s handshake bound; an abort fails the start's waits at once Neither common design bounds an ACP handshake: Close, Stop and quit end a start that never answers. The abort now also closes the connection, as a kill there does, so the start settles even before the child's exit is proven. The host-stopped start refusal only this bound produced goes with it; the idle sweep keeps its words. * fix(acp): a Stop naming an ended turn follows Claude's rule It still stops nothing while another turn is live, but in the gap before a follow-up's turn opens, which no client can name, it now stops what is in flight and the session ends, as a Claude Stop does. * fix(native-chat): a close no longer re-asks a failed start's unproven child Neither common design retries that stop at Close, and Orca's Claude contract re-asks only at the next start and at quit. The ACP adapter keeps the child until its exit is proven and asks it again there, as Claude does. * fix(acp): a message sent during a turn the agent began itself goes at once Both common designs send it straight to the agent with no cancel; only Orca's own running prompt is steered (cancelled, then re-prompted). * test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer Uses an audience production sends (one that cannot show every agent), per review. * fix(acp): launch Grok as `grok agent stdio`, without the update and leader flags The common pattern passes neither --no-auto-update, --no-leader nor GROK_DISABLE_AUTOUPDATER; full access still adds --always-approve. * fix(acp): an agent that ends its stdout, or answers unreadably, is not a lost connection As in the common pattern, only a broken stdin (or Orca's own close) ends the agent; one that closed its output but can still be written to stays until a Stop, a close or its exit. The provider supervisor goes back to its base content, so Claude and Codex no longer get the forwarded stdout end either. * fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts On main every close of the chat writes its own Stop and settle. Here a later close joins the first and writes no row, and the first's settle closed when its kill failed, so a turn that opened in between and was cut by the next close read as failed. A person's close joining a person's close whose Stop opened a settle now reopens that settle until its attempt is done. Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and the next reads as the person's cancellation (each fails without its half of the fix). * test(native-chat): Grok opens as a chat only behind the structured-chat setting agent.launch and orchestration worker-start read the same setting as the renderer route; pin both states for Grok on each. The setting's description no longer names only Codex and Claude, in every catalog. * docs(acp): generic ACP comments say what holds for every agent, not Grok Stop ends the session for every ACP agent, as in the common pattern; the adoption hook comment goes (adoption is not planned); a failed start's child is retried at the next start or quit. * test(claude, codex): type the EOF-before-exit test's streams; the supervisor no longer forwards EOF The Claude test wrote to the child's stdout and stderr through their Readable type, which the node typecheck rejects; it now holds its own PassThrough streams. The comments no longer credit the reverted supervisor change. * feat(acp): a steer's cancel asks once and never ends the agent The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle, then close the connection, which ends the agent. A steer used it too, so a slow agent lost its process just because the person added a message. requestSteerCancel() now sends session/cancel once per prompt, cancels the agent's open requests and answers later permissions cancelled, and never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel() stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel paths move into acp-prompt-cancel.ts over one cancel channel. * chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge * fix(acp): a repeated steer shares the cancel in flight; say what the caller owns Per review: a second steer before the first write lands returns that write instead of resolving early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that fails instead must not take the steer until the caller rebuilds the session; the Stop's says a prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives the runtime a handler that would allow: the open permission's signal aborts and the late one never reaches it. * fix(native-chat): drop the stopDelivery the A3 merge doubled * fix(acp): a steer's cancel asks Grok once and never ends it A steer now uses D1's notify-only cancel. Two messages sent during a reply Grok began itself cut that reply, as the common pattern does, and then both run; before, the queued first message could not answer the bounded cancel and Orca ended Grok although Grok answered. A Stop keeps the bounded cancel and its 4 s grace. * fix(acp): a permission Grok asks with no prompt of Orca's running is declined During a turn Grok began itself nobody asked it to act, so the request is answered cancelled at once instead of opening a card that waits, as the common pattern does. * fix(acp): a Grok that dies while starting is reported with its own last words A dying process's stdout ends before its exit is seen, so the start failed as a closed connection and Grok's stderr was lost. A start whose connection closed now waits, bounded by the Stop grace (or a Close/Stop), for the exit before it is told. * test: a Stop after a steer sends its own cancel; drop the import the A3 merge doubled * test(native-chat): main's Stop-note test builds its turn context with the agent registry * test(claude): say why the close test's fake child cast is safe * test(native-chat): build the Stop-opened-turn test's identity and turn context the current way The test (#25056) landed before the opaque provider handle (#24991), so main still built the old {kind, threadId} handle; the turn context also needs this branch's agent registry. * test(native-chat): build the Stop-opened-turn test's identity with the opaque handle The test (#25056) landed before the opaque provider handle (#24991), so main still built the old {kind, threadId} handle. * test(ratchet): require src/main/provider-process now that it has landed * test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test Main's #25706 and this branch both added the import at different lines; the merge kept both. * test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test Main's #25706 made the same fix as this branch at a different line; the merge kept both imports. * feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore A chat whose saved conversation the agent cannot reopen can now continue in a fresh one: the handle chain records the new conversation as a creation that replaces the lost one (which, why, and when), keeping every earlier link. Rows keep a shape older builds read: the stored chain starts at the latest replacement and carries the earlier links inside it. * test(native-chat): build this stack's journal identities with main's opaque provider handle Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test files from this stack still wrote the old shape. Same lines the downstream ACP branch uses. * docs(acp): every reattach drops the agent's replay, not only for a chat the journal holds * refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row Older builds only read Claude and Codex records, so the nested stored form protected rows no replacement can reach while adding a cap mismatch after a downgrade. Store the chain as held, refuse a replacement in a Claude or Codex chain until one has a stored shape older builds read, and refuse a supersession key on a replacement that names no creation in the chain. * refactor(native-chat): read hosts' structured agents from the app-shell services Main grew the startup hydration hook to its line limit; the host agents sync is an app-lifetime subscription like the structured session tabs sync beside it, so it moves there. * Use current provider handles in transition tests * Use current provider handles in timeline fixtures * test(native-chat): prove replacement rows survive downgrade and re-upgrade * Require the ACP directory in the runtime import check * test(ratchet): require src/main/acp now that this PR lands it * feat(acp): a saved session the agent cannot reopen continues in a new one, with one warning row When session/load (or session/resume) of a saved ACP session fails, the chat starts a new session and records it as a creation that replaces the lost one (#25747's 'replaces' link), and writes one warning row that the agent no longer remembers the earlier messages. A created session the agent reports missing is still superseded silently; a signed-out agent or a start that is over (Close, Stop, a lost agent) still fails the start. * chore(acp): rewrap the acquire header comment * test(acp): a start closed while the agent reopens fails without opening or announcing a new session * Let ACP connections own their supervised agent process * Preserve ACP cleanup evidence and isolate exit observers * Expose ACP cleanup observations and type the permission fixture * refactor(native-chat): the registered-agents capability lives in its own module Main's growth put protocol-version.ts one counted line over its 300-line limit once the capability was added; like main's other per-feature capabilities, it now has its own module, and importers read it from there. * refactor(acp): one connection owns the Grok process and its protocol D3 now opens each ACP agent through createAcpAgentConnection (ACP-ALIGN #25810): one object spawns the process on the execution host, owns its stdio and protocol, and reports its proven exit. It is built and tracked before the handshake, so a start's abort (Close, Stop, quit) still reaches it, and a failed start keeps that same connection for the next close to retry rather than spawning another process. Deleted: the spawnAcpStructuredChild wrapper and its test, the raw-stream runtime assembly, the caller's exit -> runtime.close wiring, the stdout-EOF heuristic (the connection no longer treats stdout EOF as exit), and the 10 s steer/Stop cancel bound with requestSteerCancel. Reader control maps to pauseReading/resumeReading; a close is connection.close after the host's existing 4 s Stop grace. The adapter owns what the protocol no longer does: one session/cancel per running prompt however many steers arrive (cleared with that send's settlement, retried after a failed write), and a Stop or steer answers every open agent request the person has not already answered with the agent's own cancelled reply. An answer already being saved when the Stop lands is sent. Tests: blocked cancel write never holds Stop's grace, two quick steers send one cancel, a failed cancel write is retried, a real process exiting while a child holds its stdout ends the session, and the existing start-abort, retention, crash, connection-loss and reload-failure suites on the new rig. * fix(acp): Grok signs in on its own machine with its API key or cached sign-in When Grok reports that it needs authentication, Orca now names a sign-in method on the machine Grok runs on, read from the same environment Grok was launched with: xai.api_key when XAI_API_KEY is set there and Grok offers that method, else cached_token when Grok offers it, else none and the chat keeps the existing not-signed-in refusal. The rule lives in Grok's launch spec; the adapter applies any agent's rule for new and reopened sessions through the protocol client's caller-named method (authenticate, then retry once). No new sign-in UI; interactive methods are never chosen. * fix(acp): the adapter decides which of Grok's requests reach the person The turn owner now admits every agent request, permission or question, from its own turn state: a request reaches the person only while Orca's prompt runs and no steer or Stop is cutting it short (a question may also come from a turn Grok began itself, until a Stop). Anything else gets the agent's own cancelled reply and opens no card, so a question arriving after Stop or during a steer never appears. A steer, like a Stop, withdraws the requests already open; an answer already being saved is still sent. The protocol client's abort-on-cancel path is no longer used: after the connection change its request signal aborts only when the connection closes. * fix(acp): a plan Grok proposes shows as a plan, with no approval card When Grok leaves plan mode it asks the client to approve its plan (x.ai/exit_plan_mode). Orca showed a blocking 'Approve plan / Request changes' card for it; the common pattern has no such gate. Now the plan goes into the chat's existing Plan row (the plan-document status row Codex and ACP plan updates already use) and the request is answered at once with 'abandoned' plus feedback telling Grok to stop and wait for the person's feedback or a request to implement it in a later turn, so nothing is approved on the person's behalf. Dialects gain settleRequest for requests answered without asking anyone. * fix(orchestration): worker-start opens a Grok worker in a terminal, as before With the structured chat setting on, worker-start decided 'structured' for Grok and then the structured worker factory (Claude and Codex only) refused it, so the start failed; main opened a terminal Grok worker. Worker-start now decides with no registered agents beyond Claude and Codex, so Grok gets a terminal worker as before. agent.launch and the app's own launches still open Grok as a structured chat. Temporary until structured workers take registered agents. * fix(acp): a prompt answer Orca can't read ends the turn instead of hanging it A session/prompt rejection that was not the agent's own error answer (an answer that fails Orca's schema, or one too large to read) left the turn running: the next message became a steer with nothing to cancel and was never sent or settled, and Stop waited its full grace. As in the common pattern, any prompt failure now ends the turn as failed (a failed-turn row without words, since none are the agent's) and settles the send, so the next message goes. Only a closed connection keeps the send running, for the connection-loss path to settle. * fix(acp): send Grok's prompt-identity extension only to agents that echo it session/prompt carried _meta {promptId, requestId} for every ACP agent, though only Grok's dialect echoes it (injectedPromptIdentity). Now only an agent whose dialect declares it gets the extension; other ACP agents get a plain prompt. * refactor(native-chat): the registered-agents capability lives in protocol-version again, as on main This reverts 0ef6d21. That commit moved the capability to its own module only because main's protocol-version.ts was then one counted line over its limit; main now defines it there itself within the limit, and main's new restart test imports it from there. Main's test also reads the desktop's capability list as an older client; on this branch the desktop advertises registered agents, so its older client is that list without this one capability. * test(acp): read the sign-in method with a schema, not a type assertion * refactor(native-chat): composer transport and Stop control in their own modules Main's rewind change (#19338) brought NativeChatStructuredSession.tsx and use-structured-agent-session.ts to their line limits, leaving no room for this branch's image-acceptance and unpublished-Stop lines. The composer's transport (sends, commands, options, image acceptance) moves to use-native-chat-structured-composer-transport.ts, and whether Stop shows and what it does moves to structured-agent-session-stop-control.ts. Behavior is unchanged; the runtime cast on the composer's 'local' | 'remote' is now a typed return. * test(native-chat): read registered agents by id, as main's structuredAgentsReadBy now takes Main's A3 squash changed structuredAgentsReadBy to take agent ids; this branch's test still passed { agent } objects (CI typecheck TS2322). * fix(acp): a first reopen warns when Grok forgets a chat that exchanged turns A Grok session the chat created was treated as one Grok never saved, so when Grok reported it missing on the chat's first reopen, Orca swapped in a fresh session silently, even after completed exchanges: the person saw the old messages while Grok had forgotten them. The launch now counts a created session as never saved only when the chat's journal, read at the failed reopen, holds no turn of that Grok session; anything else, an unreadable journal included, takes the normal path: the fresh session is recorded as replacing the old one and the one warning row is written. * fix(acp): a start writes the warning row an earlier attach failure dropped The row saying Grok forgot the chat was written only into the attach's deferred sink, while the fresh session's link was saved earlier. An attach failure, quit or crash in between dropped the row forever. Every start now derives the owed rows: each conversation the chain says was lost to a failed restore gets its row unless the chat's journal already holds it. The row now names the lost conversation rather than the fresh session, so a row an unused replacement wrote still counts after Grok supersedes it. * fix(acp): a question during a turn the agent began itself gets its cancelled reply A question or other card-opening request the agent sends while no prompt of Orca's runs (a turn it began itself, as when a background task wakes it) now gets the agent's own cancelled reply and opens no card, the same rule permissions already follow there. Nobody is waiting on that turn. A plan the agent shares in it is still shown. * test(native-chat): import the unfinished-work capture from the module that owns it Main's reasoning sweep test (#19221) imported it from the dead-generation settlement, which this branch split it out of. * refactor(native-chat): keep the session host within its line budget after main's Stop work Main's #25949 left the host at exactly its 300-line budget, and this branch's acquire-abort wiring adds one line. The reveal module now comes in as a namespace import, as the host already does for its other helper modules, and the earlier reorder of two type imports is undone. * test(codex): move the stdout-before-exit test into its own file Main's connection test file is at its 800-line test budget, and this branch's exit-order test pushed it over (CI lint, max-lines). * fix(acp): cancel a running turn before a close, dispose or quit ends the agent Closing a tab, disposing a session or quitting while an ACP agent's turn ran killed the process without asking the agent to cancel first. A requested close now does what a Stop does: withdraw the agent's open requests and held steers, send session/cancel once, and wait for the turn to end, bounded by the Stop's grace (4 s), before closing the process. A close that follows a Stop sends no second cancel and waits only what is left of that Stop's grace. An idle close, a lost connection and a sink-failure force close are unchanged. The cancel-and-wait moves to acp-structured-stop.ts, shared by Stop and close. * test(codex): pass resolveLaunchArgs in the stopped-send-order test so typecheck passes Main's typecheck fails here too: #25721 made resolveLaunchArgs required and this test (#25051) predates it. Same line, same place as the open main fix (#25977), so merging main after it lands is a no-op. * test(native-chat): run the close-aborts-start test on the Node runtime, as its SQLite journal fixture requires * refactor(native-chat): the Stop control reads the outbox itself, so the chat hook stays within its line limit * Keep the session host under the line cap after main's two new delegates Pass the lifetime's conversation opener to reveal directly; it is already passed unbound to the mutation context.
Base
main: #24864, #24369 ("Stopping…") and #25073 are all onmain.Builds on #25073 (the
submittedSequenceandansweredInTurnfields), which is onmainas9def4b9ba19.ELI5
When you press Stop before the agent has started on your message, the message used to vanish from the chat and its text reappeared in your typing box. Other devices just lost it. Now the message stays in the chat, with one line saying it was stopped, and your typing box is left alone. A Stop never loses what you sent.
What Changed
The problem.
Before / after:
/compact), or sent just before the agent opened its turn, that the Stop took backTo send a stopped message again, copy it (user messages already have a Copy button on desktop, and selectable text on the phone) and send it.
Mechanism.
src/shared/structured-agent-session-message-projection.ts, used by desktop and phone):cancelledfact, or its older marker) is no longer hidden. A queued card's hand-off is still left out: the card holds its text.presentation: 'stopped-before-start',src/shared/native-chat-stopped-before-start.ts), shared by consecutive taken-back messages. The row is skipped when a turn opened for the message: that turn's interrupted end already says it was stopped.submittedSequence, feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073). The move erased that from the row itself.answeredInTurnwithvia: 'steer', feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073), which would allow drawing it inside that turn later. This PR does not.structuredAgentTurnAnchorsinsrc/shared/native-chat-turn-membership.ts):answeredInTurn,via: 'start', feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073). It is then drawn as that turn's opener, just before the record, with no row of its own. A send the host states was answered into no turn (null) opens none.via: 'steer') opens nothing, even when the turn's real first message is on an older page that is not loaded yet.answeredInTurnat all was written before the host was upgraded to name turns. It is matched to a turn by the rule used before:submittedSequence), by journal order: sent before the turn's record and taken back after it. This is history an older host wrote, served by a host with feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073, which has moved its row to the take-back. It is a permanent compatibility branch, not a stopgap: old chats keep those rows, and no host version floor removes them. It is safe to keep because it draws those rows exactly as this client drew them before the host named turns; review confirmed that on 51 of 51 end-to-end cases.useStructuredAgentSessionWithdrawnRestore().byHost) is removed.Why
The common pattern keeps a stopped message on screen with one stop row after it, and leaves the composer alone. There is no per-message Retry for a stopped message, and re-running it is the person's own action. This PR adopts that.
Alternatives considered:
Differences from the common pattern
answeredInTurn) and the client draws it there.Follows the common pattern:
Mixed versions. The host publishes two more optional fields on each submission (#25073):
submittedSequence: where the journal wrote it, worked out fresh from the session's recorded history each time it is read and never stored.answeredInTurn: on a rejected message, the turn Codex answered it into and how it joined it, ornullwhen the host recorded no turn. It is stored on the rejection row, and absent from rejections written before it existed.Rule 1 of the wire doc applies: an older client ignores them, and #25073's cross-version test shows the two latest released clients draw a page and an incremental update carrying them exactly as they draw the same without them. No RPC method or stream frame changes, and the host-served conversation outline keeps its old rule.
mainand in no release tag. So the hosts a new client meets are:mainafter fix(native-chat): show a message Orca accepted and then failed to deliver as "Not sent" in the chat #24710 but before feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073: the row moves, but neither field is published. No released host is in this state.answeredInTurn, and are placed by journal order as described under Turn anchors (Intended, kept for good).mainin between, they are the best guess available.Known limit: a long chat loads its newest history first. If that loaded part starts after an earlier message whose turn is still on screen, and a later message the Stop took back follows, the stopped message and its row can show above that earlier turn's reply or stop note until you scroll up and the older history loads. On a host without #25073's fields, it can also show with no row of its own at the top of the earlier message's turn, as if it had started that turn. Review found these by loading every possible starting point of 47 end-to-end cases on the real host with the fake Codex, and re-ran that sweep on this change. For rejections a host with the fields wrote, only the first shape remains: the second, and a third (the turn's "Cancellation requested." above the message that started it), are gone. For rejections written before the host was upgraded, the second and third shapes remain, exactly as before this change.
Known limit (reachability not verified): a message Codex quietly folded into a turn already running, which Orca had not heard of yet, is recorded as having started that turn. If that turn has no user message of its own (for example, one Codex ran on its own), or its first message is on an older page that is not loaded yet, the stopped message is drawn as the head of that turn. With the first message loaded, the first message keeps the turn.
Known limit (not new): a Codex rewind rebuilds the transcript from the provider's own items plus Orca's turn rows. A message a Stop took back is neither, so it drops from the transcript before the rewind point, as every withdrawn message already did.
What this removes:
Linked Issue
Follow-up to #24864 and #24369. Closes STA-9337 (https://linear.app/stably/issue/STA-9337), with #25073: a stopped message is placed by journal order and the host's named turn instead of clocks wherever the host publishes the fields.
Visual Proof
Live desktop QA on a separate Mac, against a background build of this branch (
9f3cd5e5), driving native Codex chats backed by a stand-in Codex. The stand-in can open a turn and never echo the message, and can hold a steer without echoing it. Desktop only; the phone draws from the same shared projection and is covered by tests.A Stop after the turn opened but before Codex echoed the message. The message keeps its turn: "Interrupted after 11s" sits on it, with no "Stopped before the agent started" row. "Cancellation requested." is inside the turn, shown when it is expanded.
A message steered into a running turn, then taken back by the Stop. It is drawn after that turn's "Cancellation requested.", with its own row.
A Stop before the agent started the message. The message stays where it was sent, with one "Stopped before the agent started" row, and the composer is not refilled. Before this PR the message disappeared and its text was pasted back into the composer.
An older chat, written before the host recorded these fields, opened on the new host. Its stopped message ("Phone1 m4 never opens") draws with one stop row, and no "Cancellation requested." sits above a message that owns its turn.
Also checked in the same run: two messages queued behind
/compactand stopped before the echo (the first keeps its turn; the second, steered in, is drawn after it with its own row); every chat draws the same after a window reload; and the journal rows recordansweredInTurnasstart,steerornullas described above.Testing
9f3cd5e5; see Visual Proof)New tests:
src/shared/structured-agent-session-message-projection-stopped-before-start.test.ts(25). Each case that depends on the host is built in the host's own shape: on a host that publishessubmittedSequence, the taken-back message's row sits where it was taken back, in no turn./compact, the first passed to the agent and the second not, stay in send order after one Stop takes both back, and one a later Stop took stays below one an earlier Stop took;src/shared/native-chat-stopped-send-turn.test.ts(17), which turn a message taken back before its echo opens:src/renderer/src/components/native-chat/NativeChatMessageList.stopped-before-start.test.tsx(7), rendered on desktop:mobile/src/session/mobile-native-chat-stopped-before-start.test.tsx: the phone's projection, turn status and row text.mobile/src/session/mobile-native-chat-merged-snapshot-parity.test.tsandmobile-native-chat-merged-snapshot-parity-hooks.test.tsx. They use a made-up chat shaped like the one in QA: a long history, a message sent while a Stop was ending a turn, then a message the agent never started, taken back when the agent exited.native-chat-rail-outline-parity.test.ts: the outline and the rail both leave out a send a Stop took back, while the transcript draws it.src/main/runtime/structured-agent-session-codex-stopped-send-order.test.ts, with the real host and the fake Codex:src/main/runtime/structured-agent-session-codex-turn-end-settlement.test.ts, with the real host and the fake Codex:Updated tests:
structured-agent-session-withdrawn-message-restore.test.tsx:use-structured-agent-session-outbox-withdrawal.test.tsxanduse-structured-agent-session-outbox.draft-hand-off.test.tsx: updated to the same rules.structured-agent-session-message-projection.rejected-in-place.test.tspinned a withdrawn message as hidden, the rule this PR replaces:Each placement rule was checked by breaking it and watching its tests fail, at this head (6 files: the two placement test files, turn membership, desktop rendering, and the two end-to-end files):
nullread as "written before the field": 2;null: 5, in the two shared placement files and the desktop render test (the end-to-end tests stay green, since the real host always writes the field);start: 3;The phone guard tests turn red when each layer is broken on purpose: the phone dropping part of an update (the taken-back message's new state, a removed row, a newer version of a message or status), the list hiding the stop row, and the chat screen's own list (rows held back until the next send, a list refilled in place, or two rows sharing a key).
Run locally:
cc6d83826cf(fixes from a review of the main merge: the Spanish "Stopped before the agent started" line restored, a phone test that compared nothing now reads "Stopping…", and one stale comment): the changed phone test, the outbox test and 4 language-file tests pass; phonetscand the phone tests-typecheck ratchet pass. The full 10 desktop + 7 phone file set was not re-run at this head. CI: 16 pass, 16 skipped, 0 fail.a200a857519(basemain, after feat(native-chat): the chat reads "Stopping…" from your Stop until the turn actually ends #24369 merged): CI 16 pass, 16 skipped, 0 fail, mergeable. This PR's own diff againstmainis unchanged line for line by the main merge (45 files). This head's tests: 10 desktop files (105 tests) and 7 phone files (67 tests) pass.tc:web,tc:node, phonetscand the phone tests-typecheck ratchet pass. oxlint is clean on every changed file (desktop and phone configs). Merging main changed three tests mechanically formain's API: the droppedNO_CARDSargument, the three-argument projection call, and the new live-line mock.38567b27758(this head, after merging feat(native-chat): the chat reads "Stopping…" from your Stop until the turn actually ends #24369 at970c6a10eb5, which carriesmainthrough Add a Chat settings page for structured native chat #25685): this PR's own root test files plus the rewind and Stop tests that share its code (15 files): 197 pass.tc:webpasses from a fresh build cache. The phone's files were not re-run locally; CI runs them.ac6d23a3a03: this PR's own test files and their neighbours (22 files, including feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073's tests and the Codex echo, steer, options, adapter and late-settlement tests): 323 tests pass. The phone's 3 files: 13 pass.ac6d23a3a03, the cross-version wire suite, run serially: 15 of 19 files pass, including feat(native-chat): publish where each submission was sent and the turn a stopped one was answered into #25073's downgrade test. The other 4 cannot load locally: this worktree's installed packages predatemain's newstream-jsonversion, and they fail on that import before running a test. CI runs them, and they pass there.ac6d23a3a03,tc:nodeandtc:webreported only that same missingstream-jsonmodule, inmain's files. Mobiletscand the mobile tests typecheck ratchet pass.On a phone (Android emulator, with temporary diagnostic logging):
Projection cost. The message projection runs on every streaming update, on desktop and on the phone, so the Stop placement must cost nothing in a chat without a stopped message and grow only with what is loaded. A chat with no stopped message skips placement entirely, and the turn-claim index is built only when a taken-back message exists. Otherwise the turn anchors, the row lookup and (on an older host only) the earlier-send floor are built once per update, and each stopped message is placed with a binary search over the turns' furthest rows, not a rescan of every row. The earlier benchmark was run on the previous placement and has not been re-run on this one.
Review
Agent skill upstream boundary
docs/reference/agent-skill-sharing-upstream-boundary.mdand copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.Notes
Checklist
N/Awith reasonpnpm lint,pnpm typecheck,pnpm test, andpnpm buildpass (or CI will cover; local preferred)