Skip to content

feat(native-chat): Grok as a structured chat over the Agent Client Protocol - #25225

Merged
brennanb2025 merged 425 commits into
mainfrom
brennanb2025/acp-d3-grok
Oct 7, 2026
Merged

brennanb2025 merged 425 commits into
mainfrom
brennanb2025/acp-d3-grok

Conversation

@brennanb2025

@brennanb2025 brennanb2025 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor
Files Added Deleted Net
Test 57 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​5189 $\color{#cf222e}{\Huge{\mathbf{−}}}$​84 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​5105
Prod 111 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​4184 $\color{#cf222e}{\Huge{\mathbf{−}}}$​389 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​3795

Draft. Every base is on main now; this PR carries only its own changes. Merges only, no rebase, no force-push.

Stacked on

Nothing any more. Every PR this one was stacked on has been squash-merged into main: #24991, #25076, #25159 (registered-agents wire and storage), #24989, #24988, #25204 (shared provider process lifecycle), #25072, #25141, #25064 (shared timeline writer), #25181 (a stopped tool shows as interrupted), #24990 (ACP protocol client), #25090 (ACP → timeline translation and Grok's dialect), #25747 (record a fresh session after a failed reopen) and #25810 (the ACP connection owns the agent's process). This branch last merged main at 86d03ff9085. Where main's squash equals a branch head this branch had already merged (#25159, #25181, #25810, #25090), the merge was computed against main's tree with those heads as extra parents and committed as an ordinary merge of main, so every base file equals main's copy. What remains in the diff is this PR's own: Grok's registration, the ACP adapter, the client wiring, and the few shared changes listed below.

ELI5

Today Grok's native chat in Orca is a terminal running Grok, with Orca reading Grok's transcript file and typing into the terminal. Claude and Codex instead have a structured chat: Orca talks to the agent's own protocol, so replies stream, tool calls are real rows, approvals are real cards, and Stop actually stops. This PR gives Grok that structured chat. Orca now speaks the Agent Client Protocol (ACP, the open JSON-RPC protocol grok agent stdio implements), and Grok is the first, and only, agent on it. How Orca starts, stops, reopens and talks to Grok follows the common pattern for ACP agents; where Orca does something different, the list below says so and why.

What Changed

The problem. The PRs under this one built the parts: an ACP protocol client whose connection owns the agent's process (#24990, #25810), a translator from ACP traffic to Orca's chat timeline with Grok's protocol extensions (#25090), a shared timeline writer (#25064), a shared process lifecycle (#25204), a record that can hold a fresh session after a failed reopen (#25747), and a wire that can carry agents other than Claude and Codex (#25159). Nothing joined them. No agent was registered, the desktop never asked a host which agents it runs, and agent.launch, the account-home lookup and a dozen renderer gates still only knew Claude and Codex.

What you'll see. All of this is behind the existing setting Settings → Chat UI → Use updated structured native chat, the same switch Claude and Codex structured chats use, with the same fallbacks. With it off, Grok keeps the terminal-backed chat it has today, unchanged.

  • Opening a Grok chat. A Grok agent tab on this machine or on a paired Orca runtime opens as a structured chat. Grok's replies stream, its tool calls show as rows, and its permission requests show as approval cards with Grok's own choices ("Yes", "Yes, allow all edits during this session", "No, and tell Grok what to do differently"). The model and reasoning-effort pickers list what Grok reports, and Grok's / commands come from Grok. On SSH or WSL, Grok keeps the terminal chat.
  • Signing in. Grok signs in with what the machine running it already has: the XAI_API_KEY in Grok's own environment there, else the sign-in Grok cached on that machine. With neither, the chat shows the existing "Grok is not signed in… Sign in first." start failure. For a chat on a paired runtime, that runtime's environment and Grok sign-in are used, never this computer's. No new sign-in screen.
  • The setting's description used to read "Open new Codex and Claude agents as structured chats…". It now reads "Open new agents as structured chats where supported. Off opens them in the terminal-backed chat. Chats that already exist stay as they are." Every language has the new sentence.
  • Other ways to start Grok. agent.launch from a client that can show Grok chats (a desktop window on this version) starts a structured Grok chat only when that setting is on; with it off it starts a Grok terminal, as for Claude and Codex. Starting Grok from a phone still opens a Grok terminal, because phones cannot show Grok chats yet. A phone shows the existing "Update to view" row for a Grok chat started elsewhere. orca orchestration worker-start --agent grok opens a terminal Grok worker, as today, whatever the setting (structured workers only take Claude and Codex yet; see Differences, temporary).
  • Reopening a chat. Grok is asked to load its saved session (session/load). It replays the conversation, and Orca throws the replay away except the context-usage reading, because Orca's own history already holds the chat. The chat looks the same as before it closed. A long chat may take a little longer to reopen, because Grok replays it.
  • If Grok can't reopen its session, the chat continues in a fresh Grok session and shows one warning row in the chat: "Grok couldn't reopen its earlier session, so this chat continues in a new one. Grok doesn't remember the earlier messages." (in every language). Your earlier messages stay on screen; the next reopen loads the fresh session. A Grok that is signed out still fails the start with the existing sign-in failure, and a chat in which nothing was ever exchanged with Grok continues in its new session silently, because Grok had nothing to forget (see Differences). Before this round, a chat's first reopen after real exchanges could also take that silent path when Grok reported its session missing; it now shows the warning row. If Orca failed to open the chat, quit or crashed after the fresh session was saved but before the row was written, the row used to be lost for good; now the next reopen that works writes it, and never a second one for the same forgotten session.
  • Stop ends Grok. Stop asks Grok to cancel, waits up to 4 seconds for Grok to wind the turn down, then ends Grok's process. The next message starts Grok again and it reloads the same conversation, with no repeated history and no notice. Work Grok had moved to the background ends with it, so it cannot start a new reply after you pressed Stop; its row reads "No recent update", as a Claude background task's does after Stop.
  • Stop also withdraws what Grok asked. An open question or approval card closes and Grok hears it cancelled; a question Grok asks after you pressed Stop never shows. An answer you already gave before the Stop is still sent.
  • A message typed while Grok works steers, like the common pattern's default. Grok's current step is cancelled, the session stays, and your message runs next. Grok is asked to cancel once per running prompt: with two quick messages, the first is sent to Grok and cancelled at once, then the last one runs, as the common pattern does. Open questions from the step being cancelled close, and questions Grok asks while its step is being cancelled are declined; it can ask again once your message runs. (The host's queued-message cards are not turned on for any agent on main yet, so this is what a plain mid-turn message does.) However long Grok takes to answer that cancel, Orca waits; it never ends Grok for it. Stop still ends Grok after its 4 seconds.
  • A message sent while Grok is answering on its own (for example, when a background task it started finishes and wakes it) goes to Grok right away and does not cut that answer off; Grok runs it next. A second message sent before Grok gets to the first cuts Grok's own answer off, then both messages run, in the order you sent them, as the common pattern does. Grok is never ended for it.
  • A permission or question Grok asks while answering on its own is declined at once, with no card: no message of yours is running, so nobody is waiting to answer. Grok sees the request cancelled. Before this round a question there showed as a card. A plan Grok shares in such a reply still shows.
  • A plan Grok proposes shows as a plan. When Grok leaves plan mode, its plan appears in the chat's existing "Plan" card (the one Codex plans use), and Grok is told the plan was shown and to wait for your feedback or a request to implement it. There is no "Approve plan / Request changes" card any more, and nothing is approved for you; you reply in the chat.
  • A Stop pressed in a second window that still shows an older turn stops nothing while a newer turn runs, and it does stop the newer message in the moment before Grok has started answering it. This is Orca's rule for Claude.
  • No time limit on starting. A Grok that never answers its start stays "starting" until you end it. On a chat that already exists, Stop, closing the tab and quitting Orca each end it at once; on a brand-new chat, close the tab (next bullet). The "Grok never finished starting, so Orca stopped it." row after 60 seconds is gone. That sentence stays for the existing case where Orca's idle sweep stops a start that went quiet for 30 minutes.
  • Stop on a brand-new chat that is still starting takes the message back into the composer, but it does not stop the start yet; closing the tab does. So a brand-new chat whose Grok never answers its start needs Close, and a message sent meanwhile waits behind that start (see Differences, temporary).
  • Grok may update itself while Orca runs it, as it does when other tools run it. Orca passes no "don't auto-update" or "no leader" flags.
  • Grok's leader mode: if you turned on Grok's own leader mode ([cli] use_leader in Grok's config, off by default), Grok's chat process connects to Grok's shared leader process. Stop and Close then end Orca's connection to Grok, not that shared leader. With leader mode off (the default) nothing changes.
  • A Grok that closes its output but keeps running (seen only from a test wrapper) stays "Working" until you press Stop, which ends it. Orca stops Grok on its own only when Grok's input breaks or Grok exits. A Grok that exits while a process it started still holds its output open is seen as exited at once.
  • A reply Orca can't read (an answer to a message that doesn't match the protocol) ends that turn as failed, and your next message goes normally. Before, the turn stayed running and the next message never went.
  • If Grok crashes, the reply it was writing reads as cut short with the standard sentence ("Grok stopped while this response was in progress. You can continue in this conversation."), the same way Claude's and Codex's do. Grok's own last words (its error output) go into the chat's record and Orca's log, not onto the screen. A message sent before Orca has proven the crashed process gone reads "Grok stopped before this message was sent.".
  • If Grok dies while starting, the start failure is recorded as Grok stopping while it started, with Grok's last words in the record, instead of as a lost connection.
  • A Grok reply cut off by an Orca crash or quit shows the part Orca wrote and the existing "cut short" notice, like a Claude or Codex chat. Grok itself still remembers the whole exchange, so a follow-up question works.
  • Images are not offered for Grok chats (Grok reports it can't take them).
  • One new row, through the existing status-row surface: the "couldn't reopen its earlier session" warning above. No new notices, banners or toasts otherwise.

Mechanism.

  • Launch spec as data (src/main/acp/acp-launch-specs.ts), one row: grok agent [--always-approve] stdio, no extra environment. --always-approve is added only when the Agent Permissions setting gives Grok full access; in that mode Orca also answers any permission request with the agent's allow_once option. The account home is GROK_HOME (else ~/.grok), resolved on the runtime's own machine; the binary is searched on PATH, then in <GROK_HOME>/bin. The spec also carries Grok's sign-in rule: xai.api_key when XAI_API_KEY is set in the environment Grok is launched with and Grok offers that method, else cached_token when offered, else none.
  • One connection owns Grok's process and its protocol (Let ACP connections own their agent processes #25810's createAcpAgentConnection). The adapter builds it before the handshake and records it at once, so a Close, Stop or quit during a start reaches it. It spawns through the shared process lifecycle on the machine that runs the chat, with the chat's launch environment and working directory. Grok's stdout ending is not treated as exit; the connection reports Grok's exit only when the process is seen to exit, and reports a broken protocol (onClose) only while the process may still run. The adapter pauses and resumes reading through it when the journal is backed up, and Stop's end of the process is connection.close() after the host's 4-second grace. A start whose process Orca can't prove gone keeps that same connection, and the next start or quit closes it again; Orca never spawns a second Grok meanwhile.
  • AcpStructuredSessionAdapter (src/main/acp/acp-structured-*.ts) implements the same adapter contract Claude and Codex do:
    • Start (the adapter's acquire): initialize with the client file system and terminals off, then session/new for a new chat. A chat that already has a Grok session reopens it with session/load (an agent that can only resume gets session/resume). Inside that call Orca marks every frame Grok sends as replay before the translator reads it, so only context usage is kept and options and commands are still taken. If the reopen fails, the adapter starts session/new and records it as a fresh session that replaces the old one (feat(native-chat): record fresh sessions after failed restoration #25747's replaces link, reason restore-failed), then writes the one warning row. The row's id names the forgotten session. Every start that succeeds also writes the row for any forgotten session in the chat's session list whose row the chat's history lacks, so a row an attach failure, quit or crash dropped is written by the next start. Nothing is stored for it: the session list and the history already say whether the row is owed. A history that can't be read whole writes nothing. Exceptions: a session this chat created that Grok reports not found, on which the chat's history holds no turn, is replaced through the existing "replace an unsaved creation" link with no row (the history is read only at that failure, from the chat's own journal, with no stored flag; a turn on that session, or a history that can't be read whole, takes the normal path above); a sign-in failure, or a start already being closed, still fails the start. When Grok reports it needs to sign in, the adapter names the method the launch spec's rule picks and the protocol client signs in and retries once, for new and reopened sessions alike. The handshake has no time bound. A start that fails because the connection closed waits (up to the Stop's 4 seconds, or until a Close or Stop) for the process's exit before it is classified, so the failure carries Grok's own last words.
    • Close, Stop and quit during a start: the host owns the start it runs, so it owns its cancellation. Each attach creates an AbortController first and hands its signal to the acquire (StructuredAgentSessionAcquireInput.signal). A close, a Stop that names no turn and that admission would run now, and quit abort it from outside the chat's queue. The abort closes the connection: the process is stopped and every request Grok left unanswered fails at once, even before the exit is proven, as a kill does in the common pattern. An aborted start fails nothing: what the Stop or close withdrew is already settled, and a message accepted after it gets its own start.
    • A start that fails without proving its process gone: the adapter keeps that connection until its exit is proven. The next start asks it to close again and refuses while it can't prove it gone, and quit asks again. This is Claude's contract.
    • Sending: session/prompt. For an agent whose dialect echoes a prompt identity (Grok's does), Orca's prompt id goes in _meta.promptId, which the translator keys the turn on; other ACP agents get a plain prompt. Grok's first event for the turn, or its answer, settles the send as accepted. Grok's own error answer before the turn starts rejects the send with Grok's reason; after, it ends that turn as failed. Any other failure of the prompt (an answer Orca can't read, one too large) also ends the turn as failed and settles the send. Only a closed connection leaves the send running, for the connection-loss path below to settle.
    • Steer: a send while Orca's own prompt runs withdraws Grok's open requests, sends session/cancel once for that prompt, then sends the message as the next prompt once Grok answers. The adapter owns the "once": further steers before Grok answers send no second cancel, a cancel that could not be written lets the next steer ask again, and the mark clears when that prompt settles. The cancel is one notification with no time bound and never closes the connection. Steers wait in the adapter only for that answer; the last one runs, a Stop withdraws a waiting steer, and an exit rejects it as never sent. A send while Grok runs a turn it began itself goes as a prompt at once, with no cancel, and Grok queues it behind that turn; a second send then steers.
    • Stop: stopEndsSession is true for every ACP agent, as for Claude. Stop withdraws Grok's open requests, sends its own session/cancel (without waiting on that write) and asks Grok to end its turn; the host waits for that (4 s from the cancel), then closes the connection, which ends the process. A Stop naming a turn that has ended is declined only while another turn is live, which keeps the session. In the gap before a follow-up's turn opens it is taken, as Claude's adapter does.
    • Close, quit and dispose during a reply: as Stop does, Orca withdraws Grok's open requests, sends session/cancel once (none if a Stop already sent one), waits for the turn to end up to the same 4 s grace, then ends the process. An idle close ends it at once; a lost connection or a journal failure ends it without asking. The cancel-and-wait lives in acp-structured-stop.ts, shared by Stop and close.
    • Requests (permissions, questions): the adapter, which owns the turn, decides from its own turn state whether a request may reach the person. A request reaches the person only while Orca's prompt runs and no Stop or steer is cutting it short; a question or permission from a turn Grok began itself is declined like any other. Anything else gets the agent's own cancelled reply and opens no card. What a request answered at once carries (a plan) still shows in a turn Grok began itself. On Stop and on steer the adapter withdraws every request no answer has claimed; an answer claims the request, writes it to the journal only if the card is still unanswered, then replies, so a second window loses the write and Grok hears one answer, and an answer already being saved when a Stop lands is still sent. A permission request while no prompt of Orca's runs is answered cancelled before the full-access answer.
    • Plans: Grok's plan-mode exit (x.ai/exit_plan_mode) is answered at once, asking no one: the plan text goes to the chat's existing plan row (the plan-document status row Codex plans and ACP plan updates use), and Grok is answered abandoned with feedback to stop and wait for the person. Dialects gain settleRequest for requests like this one.
    • Options and commands: the model and effort lists come from Grok's configOptions, and a pick goes through session/set_config_option. A pick waits at most 30 seconds, as Claude's and Codex's do, and a close, Stop or quit ends the wait at once. Saved picks are re-applied at start. / commands come from available_commands_update, including ones Grok sends while it reloads.
    • Exit and connection loss: at the proven exit, open requests die, waiting steers are rejected, the running send's fate is unknown, the running turn ends interrupted, and the host's provider-exit settlement writes one turn-scoped "cut short" row carrying Grok's last words. When the protocol breaks while the process may still run (the connection's onClose: Grok's input broke, or a frame Orca can't accept), the journal closes, the running turn reads unconfirmed, the connection is closed, and the host's settlement revises that turn to interrupted when the exit is proven.
    • Where Grok runs: Grok's registration declares its location rule (this machine, no WSL, Windows only with process start-time proof).
  • Registration and capability record (structured-agent-runtime-registrations.ts): Grok is one more entry. It declares rewind, compact and goals off, context usage on, images off, steering: 'queue' (a steer is held behind a cancel, the common pattern's own meaning of that word), and approvalEnforcement: 'orca' (Orca answers what Grok asks; Grok's own settings decide when it asks). The registration resolves its own account home and opens its connection with createAcpAgentConnection.
  • One setting decides the surface everywhere: the renderer's route, agent.launch and orchestration worker-start all read experimentalStructuredNativeChat through prefersStructuredNativeChatByDefault / structuredAgentLaunchSupported (shared/structured-native-chat-launch-route.ts). agent.launch makes a chat for an agent other than Claude or Codex only when the calling client can show it (clientRendersStructuredAgent); the host's own callers (CLI, orchestration) carry no capability list. Worker-start decides with no registered agents beyond Claude and Codex, because the structured worker factory creates only those two.
  • Client: the desktop reads agentSession.agents per host (local and paired) into one cache stamped with each host's runtime id, only once structured chat is in use, and advertises agent-session.structured.registered-agents.v1 with that reader. The renderer's Claude/Codex gates on launch plan, provisional tab, persistence, tab bar, pane overlay, worktree creation and terminal seeding now trust the structured route.
  • Agent status, one producer: the ACP process is launched with every pane-identity and hook-endpoint variable removed, so Grok's own Orca status hooks stay silent and the structured session's feed is the only status for that chat.
  • Smaller fixes kept from live QA: the bypass-argument check matches a flag plus its value as a sequence (tuiAgentArgsBypassPermissions), which is what makes the full-access setting apply to Grok. Stop shows for a message waiting on a new chat that is not published yet, and takes it back.

Why

  • Follow the common pattern for ACP lifecycle. Each behavior was checked first-hand against the common pattern. Where it has a behavior, this PR uses the same mechanism; where it has none, the behavior is dropped (the 60 s start bound, the extra launch flags, reading closed output as a lost connection, re-asking at Close, a blocking plan-approval card). The exceptions are Orca contracts Claude and Codex already follow, listed below.
  • One owner for the process and its protocol. Before Let ACP connections own their agent processes #25810 the adapter assembled a process wrapper and a protocol runtime and wired one's exit into the other's close, with its own guess at what a closed stdout meant. One connection now owns both, so its exit and protocol failure are reported once, with one meaning.
  • The turn owner decides what reaches the person. The adapter already holds the open requests and knows whether a turn runs, is being cancelled, or has ended; the protocol client cannot. So withdrawal on Stop and steer, and refusing late requests, live there, with no second gate in the protocol layer.
  • Sign in on the machine that runs Grok, from Grok's own launch environment, so a remote chat never reads this computer's keys and no sign-in UI is needed for a key or a cached sign-in.
  • One adapter per protocol, rows of data per agent. A Grok-only adapter would repeat the protocol lifecycle for every later ACP agent. Grok's quirks stay in its dialect and launch spec.
  • On the shared timeline writer, not a third hand-written journal writer, so the turn and item rules that took several review rounds for Claude and Codex are not re-implemented.
  • Reload and discard instead of replaying into the journal. Orca's journal already holds every turn of a chat it created, so the replay adds nothing but risk of duplicates.
  • Stop ends the process, because a cancel ends only the running turn; work the agent moved to the background could start a turn after Stop.
  • Keep the unproven process itself, not a "cleanup owed" flag. The next start and quit re-ask it; its own exit or a proven close removes it.
  • Strip the pane environment instead of teaching the hook to skip structured sessions: the hook already refuses without a pane key.

Differences from the common pattern

  • intended: Orca signs Grok in only when Grok reports that it needs it, then retries once; it does not sign in right after Grok starts. Same methods, picked from the same environment; nothing is sent to a Grok that is already signed in.
  • intended: a Grok that is signed out still fails the start with the existing sign-in failure, instead of continuing in a fresh session. A fresh session would be signed out too, and the existing failure says what to do.
  • intended: a session this chat created that Grok reports not found, and on which the chat's history holds no turn, is replaced with no warning row. Nothing was exchanged on it, so Grok has nothing to forget; a warning there would describe a loss that didn't happen, and recording each such replacement would grow the chat's session list on every reopen of an unused chat. Any turn on that session takes the normal replacement with its warning.
  • intended: Orca sends Grok's prompt-identity extension (_meta.promptId and _meta.requestId) on session/prompt, only to agents whose dialect declares they echo it (Grok); the common pattern sends a plain prompt. Orca needs it to join each reply to the turn Orca opened for that send, and to keep Grok's own wake-up turns apart from the person's sends. Other ACP agents get a plain prompt.
  • intended: the plan row shows the plan Grok sends with its plan-mode exit, else a placeholder sentence; Orca does not look for a plan file among the turn's tool calls.
  • intended: the client file system is off (fs.readTextFile/writeTextFile: false). Common implementations differ here; off adds no new trust surface, and Grok edits files with its own tools either way.
  • intended: a failed start whose process Orca can't prove gone keeps its connection and re-asks it at the next start and at quit, instead of killing and forgetting it. Orca releases a process only on a proven exit; Claude follows the same contract.
  • intended: approvals are enforced by Orca (approvalEnforcement: 'orca'). Nothing in Orca reads that field yet: Orca has no per-chat approval policy for every agent, so full access is applied from the launch's own setting.
  • intended: a picked option waits at most 30 seconds, as Claude's and Codex's control requests do.
  • intended: Close, an admitted unnamed Stop and quit cancel a start through the acquire's own abort signal; the common pattern reaches a starting session through its registered session. Same effect: the process is stopped and the start's waits fail at once.
  • intended: with Grok's own leader mode on, Stop and Close end Orca's connection to Grok, not Grok's shared leader process. Orca passes no leader flag, as the common pattern passes none; the user's own Grok setting decides.
  • temporary: orca orchestration worker-start --agent grok opens a terminal Grok worker even with the setting on, because structured workers only take Claude and Codex. Follow-up: structured workers for registered agents.
  • temporary: Stop on a new chat that is still starting takes the message back but doesn't stop the start; the common pattern stops the start. The window could reach the start through agentSession.close, but the aborted create then comes back as a refusal and the window would show "Chat could not be started" with Retry, a failure notice for the person's own Stop. Follow-up: the host reports a start aborted by the person's own Stop as stopped, not failed, and the launch flow lets the next send start the chat again.
  • temporary: the sign-in command (grok login) is declared on the launch spec, but the existing "not signed in" words don't name it. Follow-up: carry the launch spec's sign-in command into those words.
  • temporary: images are off for Grok, as Grok reports. Follow-up: send ACP image blocks where an agent reports it can take them.
  • temporary: a Grok structured chat needs a readable process start time (on Windows, the creation-time addon); after fix(native-chat): start Windows chats without reading process creation times #25718 Claude and Codex no longer do, because main renews their held-process lease without it and the ACP adapter doesn't implement that renewal (holdsLiveProviderProcess) yet. Without a start time Grok falls back as before this PR's merge. Follow-up: give the ACP adapter holdsLiveProviderProcess, then drop ACP's start-time requirement.
  • temporary: no model list while a Grok chat is at rest (the picker fills once Grok runs). Follow-up: let ACP agents feed the host model catalog.
  • matches: Stop cancels, waits up to 4 s, then ends the process; the next send reloads the session.
  • matches: Close, quit and dispose during a reply cancel the turn and wait up to 4 s before ending Grok; with no reply running they end it at once.
  • matches: a reopen that fails continues in a fresh session with one warning row.
  • matches: a steer over Orca's own prompt is one cancel + re-prompt, the last of several runs, and its cancel never ends Grok; a message sent during a turn Grok began itself goes at once with no cancel, and a second one cuts that turn.
  • matches: Stop and steer withdraw the agent's open requests in the layer that owns the turn; the agent's own handlers send its cancelled replies.
  • matches: reopen with session/load, discarding the replay except context usage.
  • matches: no time bound on the start; only a broken input or the exit ends a Grok whose output closed.
  • matches: any prompt failure ends the turn; a permission request while no prompt of Orca's runs is answered cancelled, with no card. Questions follow the same rule (the common pattern has no Grok question support at all; its request gate is the permission one).
  • matches: a plan Grok proposes is shown as a plan and its approval request is answered at once, approving nothing.
  • matches Orca's rule for Claude: a Stop naming an ended turn is declined only while another turn is live.
  • Not built: adopting Grok's own saved history into a chat whose Orca history is empty. It is not planned in this PR.

Linked Issue

N/A (internal stack: native chat for more agents).

Visual Proof

Live QA round 6 with a real Grok (1.0.46, auto-update off) on a separate Mac, driven only through the hidden renderer, with the app's HOME, Codex and Claude folders isolated and the machine's real config files hashed after every step (nothing changed but Grok's own sign-in refresh). Every cell that ran passed. Cells ran on three heads as main moved: 940e05a5c42, 7a712ac2a4c and b081e2ddf49; later merges only brought in main and line-limit moves.

Cell What it shows Head
Setting off → terminal Grok; on → structured chat S 940e05a5c42
Signs in with Grok's cached sign-in, no prompt A 940e05a5c42
Reopen at idle: session/load, no duplicate history 7 940e05a5c42
Mid-turn message stops the step, runs next (one cancel) 5 940e05a5c42
Two quick messages: first cut off, last answered 5b 940e05a5c42
Plan shows as a plan, no "Approve plan" card P 940e05a5c42
Reopen fails → one warning row, new session, chat keeps working F 940e05a5c42
Never-used session missing on reopen → continues silently F3 7a712ac2a4c
Hung start: Close ends it (0.7 s), no 60 s limit 9 7a712ac2a4c
Hung restart: Stop ends it (1 s), message returned to the composer 9s 7a712ac2a4c
worker-start --agent grok opens a terminal Grok worker W 7a712ac2a4c
Stop during a reply: interrupted, Grok gone in 1.2 s, next prompt on a new Grok 3r b081e2ddf49
First reopen after one exchange, session missing → one warning row F2 b081e2ddf49
Orca killed mid-reply → one "Grok stopped…" notice, partial folded, follow-up remembers 8 b081e2ddf49
Grok's output closes while it runs → stays Working; Stop ends it 10 b081e2ddf49
Close the tab mid-reply → one session/cancel, Grok gone in 1.2 s T1 b081e2ddf49
Quit mid-reply → one cancel, Grok gone in 1.3 s, reopened chat shows it interrupted T2 b081e2ddf49
Close right after Stop → one cancel in total T3 b081e2ddf49

Testing

  • I manually tested these changes locally (live QA ran on a separate Mac; see Visual Proof)
  • Automated tests added/updated

Tests (explicit files, fake agents only). The adapter contract against a scripted ACP agent and against Grok's recorded sessions; the real host with a scripted Grok for Stop, steer, reopen, the reopen fallback, crash and start-abort paths; the launch spec, launch resolution and registration; the client's registered-agent route and its setting gate. The scripted agent now sits behind the same connection surface as #25810's (stdout ending is not exit; the exit closes the protocol with the agent's last words), and one test runs a real Node process through the real connection. Added or changed in this round, each removed in turn to see its tests fail:

  • Connection ownership: a Stop whose cancel write never completes still ends Grok at the grace (1 fails without it); two quick steers send one cancel (2 fail); a cancel that could not be written is asked again by the next steer (1 fails); a real process that exits while a child it started holds its output open ends the session (1 fails); Stop withdraws open requests (2 fail).
  • Sign-in: API key in Grok's launch environment, cached sign-in, key method not offered, neither offered, and on reopen (no sign-in: 3 fail; key ignored: 2; reopen without sign-in: 1).
  • Requests: a question after Stop or during a steer, and with no turn open, gets Grok's cancelled reply and no card; a steer withdraws an open question; a permission during a steer is declined (admission off: 3 fail; steer gate: 1; steer withdrawal: 1).
  • Plans: Grok's plan shows in the Plan row and its plan-mode exit is answered at once with no card (no settlement: 5 fail; no row: 1).
  • Worker-start: Grok gets a terminal worker with the setting on; other launches stay structured (1 fails).
  • Unreadable prompt answer: the turn ends as failed and the next message goes (1 fails).
  • Prompt identity: an agent whose dialect doesn't echo it gets no _meta (1 fails).
  • First reopen of a created session Grok reports missing, through the real host, journal and launch resolution: after a completed exchange it is replaced with one warning row, the earlier messages stay and the next reopen loads the fresh session; with nothing exchanged it is replaced silently; launch resolution reads the history only at that failure, and a turn of another Grok session, a damaged or newer history, or a failed read proves nothing (old rule "every created session": 2 fail; "never unsaved": 3 fail).
  • Warning row after a failed attach, through the real host, journal and launch resolution: the journal open fails after the fresh session is saved (no row), then the next reopen writes exactly one row and the one after adds none; the same when Grok then reports the fresh session missing and a newer one takes its place; and a row written once is not repeated when Grok forgets the unused fresh session (no derived row: 2 fail; row keyed by the fresh session instead of the forgotten one: 1 fails, two rows). Launch resolution: no lost session reads no history; a row written under any session counts; a damaged or unreadable history writes nothing.
  • A question during a turn Grok began itself gets Grok's cancelled reply and no card, and a plan there still shows (old admission rule: 1 fails; plan behind the narrowed rule: 1 fails).

Local run (this machine's shared install is older than the branch; I may not reinstall): 952 explicit test files on the repository configuration before main's newest test-runner change: 921 passed; the other 30 need stream-json, which that install lacks — 29 of them pass through a configuration that stubs it (the one real failure was main's new restart test, adapted here), and 2 whole-tree ratchet tests hit the same missing module. After main's Vitest-on-Bun change (#25840) the repository configuration needs a package this install lacks, so the last merges were tested with the previous configuration from a scratch folder: all 36 ACP test files (297 tests), the renderer native-chat files (61 files, 522), the settings, capability and merge-touched suites. This round: all 36 ACP test files plus timeline identity and the runtime registration and session-runtime suites (42 files, 338 tests), and after merging main 5cf3585b78d the same plus the held-send, accept-then-deliver, attach-retry and restart-ownership host suites (46 files, 415 tests), all passing. oxlint, oxfmt and the changed-code quality gate pass. Localization checks pass.

Typecheck: not run locally (this machine's install can't run it safely); CI's typecheck is the authority. Its first run on this round found 3 errors in one test (main's structuredAgentsReadBy now takes agent ids); fixed.

CI: At 27c63719b2a (after merging main 86d03ff9085): every check passed except unit shard 2 of 5, and the PR Checks verify job, which failed only because shard 2 did. Passed: static analysis and typecheck, unit shards 1, 3, 4 and 5, relay integration, cross-version wire compatibility, package (macOS and Windows), Mobile Checks, unit selection evidence and the test LoC check (PR Checks, Mobile Checks). Shard 2 is not a test failure: on three runs here (b081e2ddf49 twice and 27c63719b2a), GitHub stopped its runner partway through ("The runner has received a shutdown signal"), with no failing test in its log. Another open PR that has main's new unit-test scheduling (#25967) but none of this branch's code fails shard 2 the same way, so this comes from main. The 25 test files this branch adds or changes that no finished shard ran all pass locally (302 tests). Since the last green run (7a712ac2a4c), merging main changed the PR's own diff in four places. (1) A failed agent start carries both this branch's aborted flag and main's saved-Arguments problem (#25721). (2) The failure wording and its five translations keep main's retry-count strings (#25818) beside this branch's "couldn't reopen its earlier session" string. (3) The composer's send moved into this branch's composer-transport hook, so main's sending changes (#24514: a send that was queued reports it, and sending scrolls to the newest message) now live in that hook. (4) Line limits: the delivery loop names its refused-start type once, and Stop's control now reads the unsent messages itself, so the chat hook stays under its limit. This branch's test for a close during a start now runs on the Node runtime, as main's new check (#25967) requires of every test that opens the SQLite journal. A one-line test fix the branch had carried for main is gone: main's own copy (#25977) replaced it. Earlier merges (up to main 81a1968bc65) changed the PR's own diff in four places. (1) After main's Stop rework (#24369), this branch's Stop module adds only its own case: before the host publishes a new chat, Stop takes back a message that hasn't been sent yet. Every other Stop goes through main's hook, so the chat reads "Stopping…" as on main. (2) Codex's process-identity and location files are main's again: #25718 made Codex's start-time read best-effort, which relies on a lease-renewal path the ACP adapter doesn't implement yet. The shared helpers this branch added now serve only the ACP adapter, which keeps requiring a readable process start time, as before. (3) The session host and Codex's connection test stayed within their line limits: a namespace import in the host, and this branch's stdout-before-exit test moved into codex-app-server-connection-exit-order.test.ts. (4) Merge-only resolutions: the composer uses main's owner worktree, and main's floating-workspace comment and Settings copy are kept beside this branch's agent-neutral wording.

What I verified / didn't

  • Verified: each behavior above through tests with a scripted agent and the real host, and each new test fails without its change; one test runs a real (plain Node) process through the real connection that exits while a child holds its output open; that agent.launch, orchestration worker-start and the renderer route read the same setting; the common pattern's behavior for each item, read first-hand.
  • Not covered: a send Grok received but never answered before it died leaves no turn on that Grok session, so if Grok then reports that session missing, the reopen takes the silent path with no warning row (the person sees only their own unanswered message). Counting the send itself would need the send tied to a Grok session, which the history does not record.
  • Verified live (round 6, real Grok, see Visual Proof): the setting gate, cached sign-in, reopen, the reopen fallback and its warning row (first reopen after one exchange, and silent for a never-used session), Stop, steering with one and two messages, the plan card, an Orca crash mid-reply, Grok's output closing, hung starts, worker-start, and close, quit and close-after-Stop during a reply.
  • Not verified live: two windows on one chat, approval allow/deny cards, the model picker, a Grok refusal, background tasks, full access, a paired runtime, mixed versions, Windows, Linux and SSH; whether Grok's session start needs the sign-in step when a cached sign-in exists (Grok is only signed in when it asks); the mobile tests locally (CI runs them). The last merges of main (after b081e2ddf49) were not re-run live.

QA plan for live Grok (run on an allowed machine with an isolated HOME and GROK_HOME)

Set GROK_DISABLE_AUTOUPDATER=1 in the rig's own environment so QA never updates the machine's real Grok (Orca does not set it). Use a forcing hook for permission prompts.

  1. Prompt identity: "list the files in this folder": one turn row with the tool rows and the streamed reply.
  2. Approval allow/deny, and answering in one of two windows closes the card in the other.
  3. Stop: sleep 20; echo done > marker.txt, Stop after the shell starts: turn interrupted, no marker.txt, Grok's process gone within ~4 s, the next prompt reloads (session/load in the protocol log).
  4. Stale Stop: two windows; finish A, send B, Stop in the window still showing A: B keeps running.
  5. Mid-turn message: a second message while Grok works cancels the running step and runs next; two quick messages send two session/cancels (the first message is sent to Grok and cancelled at once), and the last one runs.
  6. Model/effort picker applies on the next turn and survives a restart.
  7. Reopen after restart: no duplicated history, the context meter keeps its last reading, session/load in the log.
  8. Crash mid-reply (Orca killed): the part Orca wrote, the existing "cut short" notice, no duplicate bubble.
  9. Start that never answers: a wrapper that never answers initialize stays starting past 60 s; Stop, Close and quit each end it at once; a send after a Stop during a restart is delivered with no start-failure row.
  10. Closed output: a wrapper closes Grok's stdout and keeps running: the turn stays Working; Stop ends Grok.
  11. Refusal before the turn starts: rejected with Grok's reason; the next send works.
  12. Background task, then Stop: Grok's process ends, no new turn appears later, the task row reads "No recent update".
  13. Background task wakes Grok, then a message: the message goes at once and Grok's own reply is not cut off. Then, during another such reply, send two messages: Grok's reply is cut off, both messages run, and Grok's process stays (same pid).
  14. Sign-in: with XAI_API_KEY in the rig's Grok environment, and separately with only a cached sign-in, the chat starts (protocol log: authenticate with xai.api_key / cached_token if Grok asks); with neither, the existing "not signed in" failure.
  15. Full access: no cards with it; no --always-approve without it (ps the argv).
  16. Status: one status for the Grok chat, none from Grok's own hooks.
  17. Setting off: Grok tabs and agent.launch open a Grok terminal.
  18. Paired runtime and folders; SSH/WSL keep the terminal chat.
  19. Mixed versions: an older desktop and a phone ("Update to view").
  20. Windows/Linux: Stop, close and quit leave no Grok process behind.
  21. Crash with last words: a wrapper prints to stderr and exits mid-reply: the row is the standard "Grok stopped while this response was in progress…" sentence; the words are in Orca's log, not on screen.
  22. Orchestration worker: worker-start --agent grok with the setting on opens a terminal Grok worker.
  23. Permission or question during Grok's own reply: with a forcing hook, a background task's completion wakes Grok and it asks permission, then (separately) a question: no card appears and Grok reports each request cancelled.
  24. Reopen fails: make the saved Grok session unloadable (remove it from the rig's GROK_HOME), reopen: the chat continues, one warning row "Grok couldn't reopen its earlier session…", earlier messages still on screen, the next reopen loads the fresh session with no second row. Run it once on a chat's first reopen after one completed exchange (that case used to be silent), and once on a chat closed before any message (no row).
  25. Plan mode: ask Grok to plan in plan mode: the plan shows in a "Plan" card, no "Approve plan" card, Grok stops and waits; your next message continues.
  26. Question then Stop: while Grok asks a question, press Stop: the card closes; Grok's process ends.

AI Disclosure

Review

Self-reviewed, then reviewed in five rounds and audited behavior by behavior against the common pattern, then two final reviews. The first (readiness checklist, the reopen fallback, the body) found the worker-start regression, the hung turn on an unreadable answer and the prompt-identity extension sent to every agent. A reference audit found a chat's first reopen after real exchanges could forget silently. The second final review found three wrong sentences in this description (two quick messages, questions during a steer, the last-merged main) and the warning row lost when an attach fails after the fresh session is saved; it also asked for a ruling on questions during a turn Grok began itself (now declined). All are fixed here. Live QA rounds 1–4 ran on a separate machine with a real Grok; round 4 passed every required cell at an earlier head. Live QA round 6 passed every cell it ran (see Visual Proof). A later review found that Close, quit and dispose during a reply ended Grok without cancelling its turn first; fixed (see What Changed).

Agent skill upstream boundary

  • Not applicable; no skill source is copied or translated.

Notes

Local runtime and paired runtimes only; SSH/WSL stay blocked as today. Cursor and user-defined ACP agents are not built. The terminal-backed Grok chat is unchanged. No new dependencies; no lockfile change. New persisted shape: Grok records use the neutral handle {transport: 'acp', agent: 'grok', nativeId} and GROK_HOME; after a failed reopen a Grok record's session chain also holds a created link with replaces: {key, reason: 'restore-failed', replacedAt} (#25747's shape, on main; Claude and Codex never write it), and the warning row is a journal status row whose id names the forgotten session. Claude/Codex records are byte-for-byte unchanged. Rollback: after a downgrade to a version without this PR, a saved Grok chat tab loses its agent (that version's tab schema does not know Grok); Claude and Codex tabs are unaffected.

Checklist

  • This PR is small and focused (it is the stack's first user-visible PR: adapter + registration + client wiring)
  • I explained what changed and why
  • Visual proof: live QA screenshots
  • Self-reviewed for correctness, security, and performance
  • Cross-platform, SSH/remote, and path impact considered
  • CI typecheck, lint and tests at this head (see Testing)

When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.

An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.
…e' into brennanb2025/acp-a2-agent-registry

# Conflicts:
#	src/renderer/src/lib/structured-agent-session-provisional-tab.ts
The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.
…embler' into brennanb2025/acp-c5-tool-interrupted
A call still running when its turn ends takes the state of that turn's
row: a row another writer settled first (a person's Stop) stands, so its
calls read interrupted whatever the provider's later end reports. The
no-ending path that settled calls from the Stop row is gone, since a Stop
now leaves running calls to the provider. Adds the two Spanish strings.
A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.

A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.

A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.
A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.
…start too

The restart sweep ended every running call by the death evidence alone, so after
a person's Stop with no proof the child died the call read failed under a turn
that read interrupted. The sweep and the live dead-generation settlement now ask
the same rule the assembler does: a call in a turn already settled ends as that
row ended; only a turn still running leaves its calls to the evidence.
…mit what they claim

A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.

/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.

Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).
…otocol client's own types

The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.
The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.
…ry' into brennanb2025/acp-a3-wire

# Conflicts:
#	src/main/native-chat/agent-session-wire/structured-agent-session-adapter-router.ts
#	src/main/native-chat/agent-session-wire/structured-agent-session-adapter.ts
#	src/main/native-chat/agent-session-wire/structured-agent-session-host-types.ts
#	src/main/runtime/orca-runtime-get-structured-agent-session-create-support.ts
#	src/main/runtime/structured-agent-session-runtime.ts
The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.
The PR was conflicting with main, so CI could not run.

- structured-agent-session-agent-start.ts: kept D3's `aborted` and main's
  `argumentProblem` (#25721) on the failed start outcome.
- en.json: kept D3's `sessionNotRestored` and main's saved-Arguments strings.
- structured-agent-session-delivery-loop.ts: main's growth plus D3's
  aborted-start guard went over the 300-line budget; named the refused-start
  type once so refusedStart's signature fits on fewer lines.
… typecheck passes

Main's typecheck fails here too: #25721 made resolveLaunchArgs required and
this test (#25051) predates it. Same line, same place as the open main fix
(#25977), so merging main after it lands is a no-op.
# Conflicts:
#	src/renderer/src/i18n/locales/es.json
#	src/renderer/src/i18n/locales/fr.json
#	src/renderer/src/i18n/locales/ja.json
#	src/renderer/src/i18n/locales/ko.json
#	src/renderer/src/i18n/locales/zh.json
#	src/shared/agent-session-failure-copy.ts
# Conflicts:
#	src/renderer/src/components/native-chat/NativeChatStructuredSession.tsx
@brennanb2025

brennanb2025 commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor Author

Review summary (review coordinator) — head 1d98732e682

The problem. Grok only had the terminal-backed chat. Orca typed into a terminal running Grok and read Grok's transcript file. Claude and Codex already have a structured chat, where Orca speaks the agent's own protocol: replies stream, tool calls are rows, approvals are cards, and Stop really stops. This PR gives Grok that structured chat over the Agent Client Protocol (ACP, the JSON-RPC protocol grok agent stdio speaks). Every PR it was stacked on is now on main.

What changes for you. All of this applies only with Settings → Chat UI → Use updated structured native chat on. With it off, Grok keeps today's terminal chat.

  • A new Grok tab opens as a structured chat, on this machine or a paired Orca runtime. You get streamed replies, tool rows, Grok's approval choices as cards, and Grok's model/effort pickers and / commands. SSH and WSL keep the terminal chat.
  • Grok signs in with what the machine running it already has: its XAI_API_KEY, else its cached sign-in. There is no new sign-in screen.
  • Stop cancels Grok's turn and ends Grok's process within about 4 seconds. The next message restarts Grok and reloads the same conversation.
  • Closing the tab, quitting Orca, or disposing the chat during a reply now does the same as Stop: it cancels Grok's turn, waits up to 4 s, then ends Grok. Before this round's fix, Grok was killed without being asked to cancel.
  • A message typed while Grok works stops Grok's current step and runs next. If you send two quick messages, the last one runs.
  • A plan Grok proposes shows in the existing plan card, with no "Approve plan" card. Nothing is approved for you.
  • Reopening a chat reloads Grok's session. If Grok can't reopen a session the chat actually used, the chat continues in a fresh session with one warning row. A chat that never exchanged anything continues silently.
  • A Grok crash reads like a Claude or Codex crash. A start Grok never answers stays "starting" until you close it.
  • orca orchestration worker-start --agent grok still opens a terminal Grok worker.

Fixed in this last round.

  • Teardown: close, quit and dispose now cancel a running turn first, the way bb's stopSession does, and a close right after Stop sends no second cancel. Eight tests cover it. Removing the fix fails 6 of them; removing the no-second-cancel guard fails 1.
  • Shard 2 test: a D3 test that opens the SQLite journal now runs on the Node runtime, which main's new check requires.

Deferred (stated in the PR body). Each is labeled temporary, with its follow-up:

  • Windows start time: a Grok structured chat still needs a readable process start time; on Windows that means the creation-time addon. After fix(native-chat): start Windows chats without reading process creation times #25718, Claude and Codex no longer do. The follow-up gives the ACP adapter holdsLiveProviderProcess (held-process lease renewal), then drops Grok's start-time requirement.
  • Stop on a still-starting new chat takes the message back but doesn't stop the start.
  • No model list at rest: a Grok chat shows no model list until Grok runs.
  • Images are off for Grok.

Verified.

  • Live QA round 6 with a real Grok on a separate Mac. The test app's own home folders were isolated, and the machine's real config files were hashed after every step. Every cell that ran passed, and screenshots for each are in the body's Visual Proof. The cells ran on three heads as main moved:
    • on 940e05a5c42: setting gate, cached sign-in, reopen, steering with one and two messages, plan card, reopen fallback with one warning row;
    • on 7a712ac2a4c: second reopen loads the new session, silent fallback for a never-used session, hung start (Close, Stop), worker-start;
    • on b081e2ddf49: Stop mid-reply, first reopen after one exchange, Orca crash mid-reply, Grok's output closing, and the three teardown cells. Close and quit each sent exactly one session/cancel, and Grok was gone in about 1.2–1.3 s.
  • CI at 1d98732e682: every check passed, typecheck and lint included, except unit-test shard 2 and the verify job that requires it.
    • Shard 2's runner received a shutdown signal about 5 minutes in, on the first run and on one re-run. No test had failed before the cut-off.
    • This happens across main right now, not just here. The same shutdown hit shard 2 on another open PR in the same hour. Open PR Fix unsafe test fixtures and the Bun version pin #26051 traces it to a runtime import check whose esbuild grows to about 12.75 GB on a 15.6 GB runner, and fixes it.
    • Locally, on the merge, these pass: the conflicted files' tests, the Claude and Codex child-exit tests, and D3's ACP stop and teardown tests (35 files, 331 tests). A wider set of main's new native-chat, Claude, Codex and runtime tests also passed: 91 of 99 files, 951 tests. The other 8 files could not load because this machine's installed packages are older than main's.

Not verified.

  • Merges since round 6d: the merges of main after b081e2ddf49 weren't re-run live. The latest (e1b7cf44081) brought in main's fix(claude): start a Claude chat with its saved options and send the first message at once #25152 (a Claude chat starts with its saved options). That needed a hand-merge in the shared child-exit and dead-generation settlement code, keeping both sides. A fresh review of the merge found no lost behaviour. 1d98732e682 only trims the session host back under the line limit after main's two new delegates. Per the fixer's diff, Stop changed only by a line-limit move, reopen picked up main's fix(native-chat): alert when a structured chat asks for approval or input mid-turn #25766, close got only a rename, and crash/exit and output-close are unchanged.
  • Never run live: two windows on one chat, approval allow/deny cards, the model picker, a Grok refusal, background tasks, full access, a paired runtime, mixed versions, Windows, Linux, SSH and mobile.
  • Typecheck: no local typecheck was run; CI's typecheck passed.

For the merger. At 1d98732e682 the PR is mergeable. If main moves again, merge current main right before merging.

Isolation note. During QA, one run's first launch rewrote the test Mac's real ~/.grok/hooks/orca-status.json. The hooks setting was turned off after launch instead of before. Brennan approved restoring it from its backup, and its hash was verified afterwards. The re-run seeded the setting first and changed nothing.

Conflicts: agent-launch.ts (main moved intent resolution to its own module; kept D3's
callerRendersLaunchedChat), its test (gemini stays a terminal: viewMode 'terminal'),
the capabilities test imports, and NativeChatStructuredSession.tsx (main's queue
hold/resume now flow through D3's composer-transport hook).
@brennanb2025
brennanb2025 marked this pull request as ready for review October 7, 2026 01:47
@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

📝 Walkthrough

Walkthrough

This change adds an ACP structured-session adapter and registers Grok as an ACP agent. It adds session acquisition, launch resolution, request and turn handling, options, restore behavior, and process lifecycle management. Host-registered agents now inform structured-chat launch routing and renderer capabilities. The host can abort pending acquisitions and revise settlement records when it observes a child exit. The renderer also uses agent capability data for image input, option catalogs, and Stop behavior. Launch argument matching now supports multi-token permission-bypass flags.

Priority: ➖ Normal

Merge Risk: 🔵 Low · up to 412d1

Closing or stopping Grok can be delayed while launch environment probing completes, and a rare process-exit timing can leave internal turn tracking unresolved. These are localized lifecycle risks, but should be considered before relying on prompt cleanup.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 94 functions across 65 files. (7 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: adding Grok as a structured chat over ACP.
Description check ✅ Passed The description is comprehensive and covers the user-facing changes, implementation, rationale, visual proof, testing, limitations, and review. The AI Disclosure is blank as the template permits for i…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 94 functions across 65 files. (7 skipped: 7 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
src/renderer/src/runtime/host-structured-agents.ts (1)

103-108: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Discard a host-agent reply when its host is no longer retained.

The .then handler writes entries.set(executionHostId, ...) without checking if the host is still paired. retainHostStructuredAgents can run while the read is in flight and drop the host. The late reply then adds the unpaired host back. For a paired host this does not cause a wrong route: readHostStructuredAgents compares runtime IDs, and the stale entry is pruned on the next status sync. The impact is a small, self-recovering cache leak. Change this only if the cache must stay exact.

Based on learnings: re-check the captured generation after an await before you write results.

Source: Learnings


ℹ️ Review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: d0834f72-f3f8-4224-9abd-bad3e59838e9
📥 Commits

Reviewing files that changed from the base of the PR and between 0c96550 and a74cd83.

📒 Files selected for processing (168)
  • config/scripts/vitest-sqlite-runtime-files.mjs
  • src/main/acp/acp-dialects/acp-dialect.ts
  • src/main/acp/acp-dialects/grok-dialect.ts
  • src/main/acp/acp-dialects/grok-requests.ts
  • src/main/acp/acp-launch-specs.ts
  • src/main/acp/acp-prompt-turns.ts
  • src/main/acp/acp-session-reopen-failure.ts
  • src/main/acp/acp-structured-acquire.ts
  • src/main/acp/acp-structured-adapter.test-support.ts
  • src/main/acp/acp-structured-agent-definitions.ts
  • src/main/acp/acp-structured-auth.test.ts
  • src/main/acp/acp-structured-broken-stream.test.ts
  • src/main/acp/acp-structured-connection.ts
  • src/main/acp/acp-structured-electron-free.test.ts
  • src/main/acp/acp-structured-fixture-replay.test-support.ts
  • src/main/acp/acp-structured-host-crash.test.ts
  • src/main/acp/acp-structured-host-lifecycle.test.ts
  • src/main/acp/acp-structured-host-option-hang.test.ts
  • src/main/acp/acp-structured-host-restore-failed.test.ts
  • src/main/acp/acp-structured-host-start-aborts.test.ts
  • src/main/acp/acp-structured-host-stop.test.ts
  • src/main/acp/acp-structured-host.test-support.ts
  • src/main/acp/acp-structured-lane.ts
  • src/main/acp/acp-structured-launch-resolution.test.ts
  • src/main/acp/acp-structured-launch-resolution.ts
  • src/main/acp/acp-structured-options.ts
  • src/main/acp/acp-structured-prompts.ts
  • src/main/acp/acp-structured-request-admission.test.ts
  • src/main/acp/acp-structured-session-adapter-deps.ts
  • src/main/acp/acp-structured-session-adapter-lifecycle.test.ts
  • src/main/acp/acp-structured-session-adapter-recordings.test.ts
  • src/main/acp/acp-structured-session-adapter-resume.test.ts
  • src/main/acp/acp-structured-session-adapter-teardown.test.ts
  • src/main/acp/acp-structured-session-adapter.test.ts
  • src/main/acp/acp-structured-session-adapter.ts
  • src/main/acp/acp-structured-session.ts
  • src/main/acp/acp-structured-starts.ts
  • src/main/acp/acp-structured-steer.test.ts
  • src/main/acp/acp-structured-stop.ts
  • src/main/acp/acp-structured-turns.ts
  • src/main/acp/acp-timeline-dialect.test.ts
  • src/main/acp/acp-timeline-recordings.test.ts
  • src/main/acp/acp-timeline-recovery.test.ts
  • src/main/acp/acp-timeline-requests.ts
  • src/main/acp/acp-timeline-translator.ts
  • src/main/agent-launch/agent-launch-executor.test.ts
  • src/main/agent-launch/agent-launch-executor.ts
  • src/main/agent-launch/agent-launch-mode.ts
  • src/main/agent-launch/agent-launch-surface-factories.ts
  • src/main/claude/claude-stream-json-connection-close.test.ts
  • src/main/codex/codex-app-server-connection-exit-order.test.ts
  • src/main/ipc/desktop-renderer-runtime-capabilities.test.ts
  • src/main/ipc/desktop-renderer-runtime-capabilities.ts
  • src/main/native-chat/agent-session-timeline/provider-timeline-identity.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-acquire-aborts.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-acquisition.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-adapter.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-agent-start.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-attach-flow.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-attach-orchestration.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-child-exit.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-close-aborts-start.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-codex-stop-row.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-conversation-lifetime.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-dead-generation-settlement.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-dead-generation-settlement.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-delivery-loop.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-forget-status.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-host-mutations.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-host-runtime-state.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-host-teardown.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-host-teardown.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-host.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-mutation-admits-now.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-mutation-context.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-provider-child-record.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-reasoning-sweep.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-stale-turn-verdict.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-startup-failure-exit.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-stop-aborts-start.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-turns-options.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-unexpected-exit.test.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-unfinished-work.ts
  • src/main/provider-process/provider-spawned-process-identity.ts
  • src/main/provider-process/supervised-provider-child-location.ts
  • src/main/runtime/orca-runtime-get-worktree-ps.ts
  • src/main/runtime/orca-runtime-structured-agent-registration-seam.test.ts
  • src/main/runtime/orchestration/send-agent-turn-boundary.test.ts
  • src/main/runtime/rpc/methods/agent-launch-caller-renders.test.ts
  • src/main/runtime/rpc/methods/agent-launch.test.ts
  • src/main/runtime/rpc/methods/agent-launch.ts
  • src/main/runtime/rpc/methods/orchestration-worker-start-mode.ts
  • src/main/runtime/rpc/methods/orchestration-worker-start-receipt-wording.test.ts
  • src/main/runtime/rpc/methods/orchestration-worker-start-registered-agent.test.ts
  • src/main/runtime/rpc/methods/structured-agent-session-restart-dismiss-fence.test.ts
  • src/main/runtime/rpc/methods/structured-agent-session-restart-resume.test.ts
  • src/main/runtime/rpc/methods/structured-agent-session-restart-unregistered-agent.test.ts
  • src/main/runtime/structured-agent-account-home.ts
  • src/main/runtime/structured-agent-runtime-registrations-acp.test.ts
  • src/main/runtime/structured-agent-runtime-registrations.ts
  • src/main/runtime/structured-agent-session-runtime.ts
  • src/main/runtime/structured-agent-shell-environment.ts
  • src/renderer/src/app-shell/use-app-shell-services.ts
  • src/renderer/src/components/native-chat-resume-on-restart-grouping.ts
  • src/renderer/src/components/native-chat/NativeChatComposer.tsx
  • src/renderer/src/components/native-chat/NativeChatStructuredSession.tsx
  • src/renderer/src/components/native-chat/StructuredAgentSessionPaneOverlayLayer.tsx
  • src/renderer/src/components/native-chat/StructuredAgentSessionStatusBridge.test.tsx
  • src/renderer/src/components/native-chat/agent-session-failure-words-text.ts
  • src/renderer/src/components/native-chat/native-chat-composer-types.ts
  • src/renderer/src/components/native-chat/structured-agent-registered-agent-surface.test.ts
  • src/renderer/src/components/native-chat/structured-agent-session-seed-catalog.ts
  • src/renderer/src/components/native-chat/structured-agent-session-stop-control.ts
  • src/renderer/src/components/native-chat/structured-agent-session-tabs.ts
  • src/renderer/src/components/native-chat/use-host-model-catalog-upgrade.ts
  • src/renderer/src/components/native-chat/use-native-chat-composer-attachments.ts
  • src/renderer/src/components/native-chat/use-native-chat-composer-image-input.test.tsx
  • src/renderer/src/components/native-chat/use-native-chat-resolved-path-attachments.ts
  • src/renderer/src/components/native-chat/use-native-chat-structured-composer-transport.ts
  • src/renderer/src/components/native-chat/use-structured-agent-session-option-state.ts
  • src/renderer/src/components/native-chat/use-structured-agent-session-options.ts
  • src/renderer/src/components/native-chat/use-structured-agent-session-provisional.test.tsx
  • src/renderer/src/components/native-chat/use-structured-agent-session.ts
  • src/renderer/src/components/settings/ExperimentalPane.test.tsx
  • src/renderer/src/components/settings/NativeChatExperimentalSetting.tsx
  • src/renderer/src/components/tab-bar/TabBarItemRow.tsx
  • src/renderer/src/components/terminal/initial-terminal-structured-launch.test.tsx
  • src/renderer/src/components/use-terminal-watcher-effects.ts
  • src/renderer/src/i18n/en-runtime-required.json
  • src/renderer/src/i18n/locales/en.json
  • src/renderer/src/i18n/locales/es.json
  • src/renderer/src/i18n/locales/fr.json
  • src/renderer/src/i18n/locales/ja.json
  • src/renderer/src/i18n/locales/ko.json
  • src/renderer/src/i18n/locales/zh.json
  • src/renderer/src/lib/agent-launch-route-input.ts
  • src/renderer/src/lib/agent-launch-route-registered-agents.test.ts
  • src/renderer/src/lib/agent-launch-routing.ts
  • src/renderer/src/lib/agent-session-launch-plan.test.ts
  • src/renderer/src/lib/agent-session-launch-plan.ts
  • src/renderer/src/lib/launch-structured-agent-session.ts
  • src/renderer/src/lib/structured-agent-launch-settlement.ts
  • src/renderer/src/lib/structured-agent-session-host-admission.ts
  • src/renderer/src/lib/structured-agent-session-idle-empty-chat.ts
  • src/renderer/src/lib/structured-agent-session-launch-draft.ts
  • src/renderer/src/lib/structured-agent-session-launch-persistence.test.ts
  • src/renderer/src/lib/structured-agent-session-launch-persistence.ts
  • src/renderer/src/lib/structured-agent-session-launch-registered-agent.test.ts
  • src/renderer/src/lib/structured-agent-session-launch-registry.ts
  • src/renderer/src/lib/structured-agent-session-launch-status.ts
  • src/renderer/src/lib/structured-agent-session-launch.ts
  • src/renderer/src/lib/structured-agent-session-paired-admission.ts
  • src/renderer/src/lib/structured-agent-session-provisional-tab.ts
  • src/renderer/src/lib/worktree-creation-flow-execute.ts
  • src/renderer/src/lib/worktree-creation-structured-session.ts
  • src/renderer/src/runtime/host-structured-agents-sync.ts
  • src/renderer/src/runtime/host-structured-agents.test.ts
  • src/renderer/src/runtime/host-structured-agents.ts
  • src/renderer/src/runtime/use-host-structured-agent.ts
  • src/shared/agent-session-failure-copy.ts
  • src/shared/agent-session-failure-words.ts
  • src/shared/agent-session-failure.ts
  • src/shared/agent-session-refusal-notice.test.ts
  • src/shared/electron-remote-runtime-client-capabilities.ts
  • src/shared/structured-agent-session-create.ts
  • src/shared/structured-agent-session-mutation.ts
  • src/shared/tui-agent-launch-defaults.test.ts
  • src/shared/tui-agent-launch-defaults.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 1 remain after this review.

Comment on lines +149 to +154
await input.beforeDispatch?.()
session.turns.dispatch({
clientMessageId: input.clientMessageId,
prompt,
requestedAt: input.requestedAt ?? this.now()
})

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win

Check the session again after beforeDispatch resolves.

dispatch reads the live session before it awaits input.beforeDispatch. The connection can break, or the child can exit, during that await. In that case, connectionLost or finish has already run closeAcpSessionJournal, which calls session.turns.end().

After that, session.turns.dispatch adds the send to unsettled and calls start:

  • lane.apply does nothing, because the lane is disposed.
  • connection.prompt rejects with AcpConnectionClosedError.
  • The rejection handler returns without settling, because it expects the session's end to settle the send. That end has already run.

Result: the adapter returns { state: 'admitted' } for a message that never reached Grok, and no later settlement arrives. The send stays unconfirmed in the journal.

To fix this, re-check that the session is still live after the await. If it is not, reject the send as never sent.

🐛 Proposed fix
     await input.beforeDispatch?.()
+    if (this.sessions.get(input.sessionId) !== session || session.journalClosed !== null) {
+      // The child ended while the send was prepared: the message never left Orca.
+      return this.rejected(session, 'providerExited')
+    }
     session.turns.dispatch({
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
await input.beforeDispatch?.()
session.turns.dispatch({
clientMessageId: input.clientMessageId,
prompt,
requestedAt: input.requestedAt ?? this.now()
})
await input.beforeDispatch?.()
if (this.sessions.get(input.sessionId) !== session || session.journalClosed !== null) {
// The child ended while the send was prepared: the message never left Orca.
return this.rejected(session, 'providerExited')
}
session.turns.dispatch({
clientMessageId: input.clientMessageId,
prompt,
requestedAt: input.requestedAt ?? this.now()
})

# Conflicts:
#	src/main/native-chat/agent-session-wire/structured-agent-session-child-exit.ts
#	src/main/native-chat/agent-session-wire/structured-agent-session-dead-generation-settlement.ts
Pass the lifetime's conversation opener to reveal directly; it is already passed unbound to the mutation context.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to GitHub limitations.

⚠️ Outside diff range comments (1)

🟡 Minor · Make launch resolution abortable. · acp-structured-acquire.ts:79-92

src/main/acp/acp-structured-acquire.ts:79-92
🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win

Make launch resolution abortable.

acquireAcpStructuredSession awaits deps.resolveLaunch before it tracks a connection. The production resolver awaits the login-shell environment, whose probe can take 5 seconds and can run a second 5-second fallback probe. Close, Stop, and quit abort only the tracked connection, so they cannot settle this wait. Host teardown can therefore wait several seconds for the attach to settle.

Thread attempt.signal into launch resolution and race the environment and workspace awaits against it. Return the existing pre-spawn cancellation error when the signal aborts.


ℹ️ Review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 702b71fe-318d-4627-887d-47fad10da60b
📥 Commits

Reviewing files that changed from the base of the PR and between 1d98732 and 412d159.

📒 Files selected for processing (13)
  • src/main/native-chat/agent-session-wire/structured-agent-session-adapter.ts
  • src/main/native-chat/agent-session-wire/structured-agent-session-host.ts
  • src/renderer/src/components/native-chat/agent-session-failure-words-text.ts
  • src/renderer/src/i18n/en-runtime-required.json
  • src/renderer/src/i18n/locales/en.json
  • src/renderer/src/i18n/locales/es.json
  • src/renderer/src/i18n/locales/fr.json
  • src/renderer/src/i18n/locales/ja.json
  • src/renderer/src/i18n/locales/ko.json
  • src/renderer/src/i18n/locales/zh.json
  • src/shared/agent-session-failure-copy.ts
  • src/shared/agent-session-failure-words.ts
  • src/shared/agent-session-failure.ts
🚧 Files skipped from review as they are similar to previous changes (8)
  • src/renderer/src/i18n/en-runtime-required.json
  • src/shared/agent-session-failure-copy.ts
  • src/renderer/src/i18n/locales/zh.json
  • src/renderer/src/i18n/locales/es.json
  • src/renderer/src/i18n/locales/ja.json
  • src/renderer/src/i18n/locales/fr.json
  • src/renderer/src/i18n/locales/ko.json
  • src/renderer/src/i18n/locales/en.json

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 6 remain after this review.

@brennanb2025
brennanb2025 merged commit 2f49377 into main Oct 7, 2026
53 of 55 checks passed
brennanb2025 added a commit that referenced this pull request Oct 7, 2026
Main's #25225 moved the composer transport into
use-native-chat-structured-composer-transport.ts and Stop into
structured-agent-session-stop-control.ts, both built on the saved outbox
this branch removes. Both now read the in-memory sender: the transport
sends through sendThroughLaunch (keeps image connectionIds, sendOut while
starting), and Stop before the chat is published gives the launch's
unsent text back to the composer (takeBackStructuredLaunchPrompts) and
asks nothing of the host, as main's rule does for its outbox.
Jinwoo-H added a commit that referenced this pull request Oct 7, 2026
…he restore-failed test (#26094)

#26038 removed JournalHostDatabase.legacyDirectoryFor while #25225's test still
spied on it, so typecheck fails on main. Use the replacement spy #26038 gave
its own tests.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant