Skip to content

feat(native-chat): open structured chat's wire and stored records to registered agents, behind a negotiated capability - #25159

Merged
brennanb2025 merged 79 commits into
mainfrom
brennanb2025/acp-a3-wire
Oct 6, 2026
Merged

brennanb2025 merged 79 commits into
mainfrom
brennanb2025/acp-a3-wire

Conversation

@brennanb2025

@brennanb2025 brennanb2025 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor
Files Added Deleted Net
Test 36 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​1981 $\color{#cf222e}{\Huge{\mathbf{−}}}$​77 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​1904
Prod 71 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​1010 $\color{#cf222e}{\Huge{\mathbf{−}}}$​447 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​563

Based on main after #25076 and #25204 merged. Review git diff origin/main...HEAD; it contains only A3's change, including saved-chat readability and restart-audience fixes.

ELI5

Orca's structured native chat (the chat view that drives an agent through its own protocol instead of reading a terminal) can only ever be Claude or Codex, because "Claude or Codex" is written into what the desktop and a paired server send each other and into what Orca saves to disk. The next agents (Grok first, over the Agent Client Protocol) can't be added until those places accept "any agent this server has", and until older Orca versions on either side are guaranteed not to break when they meet one. This PR opens those places, adds the negotiation that keeps old and new versions safe with each other, and lets a server tell clients which agents it runs and what each can do.

Saved chats stay readable when their agent is unavailable. A later build that removes an agent registration can still open its saved conversations and history; starting new work is refused until registration returns. No new agent or unavailable notice is added, and Claude/Codex stored bytes remain unchanged.

What Changed

The problem. After #25076 each agent describes itself in one definition, but the wire and storage still named the two agents directly:

  • The agentSession.create, createSupport and modelCatalog requests only accepted claude or codex.
  • A saved chat record was only readable if its agent was claude or codex, and a record's stored conversation pointer was checked against a two-way choice that treated every non-Claude agent as Codex.
  • Saved tab state only kept claude or codex as a chat tab's agent.
  • The account-home variable (the environment variable that points an agent at its config directory) came from a fixed two-entry table, and the create path picked it with agent === 'claude' ? … : ….
  • The model-catalog types, the published chat-tab type and several host-internal types were typed as the pair.
  • Nothing told a client which agents a server runs, and nothing stopped a server from sending an older client a chat tab it can't show. Today a server would send such a tab to every current desktop and phone (the per-agent gate treated any unknown agent like Claude, and every current client says it renders Claude), and both clients would list it with an empty pane.

The mechanism now.

  • Registered agents are the vocabulary. An agent id is a short slug (isStructuredAgentId, the same bound a stored conversation pointer's agent already had). The three requests above accept any agent id; the server refuses one it didn't register (createSupport answers supported: false and create refuses; modelCatalog cannot resolve an account for it and answers that it knows no catalog, origin: unknown, as it does for any account it can't resolve). Attaching by a client-supplied conversation pointer (agentSession.ensure and the attach form of create) stays Claude/Codex: only their pointers have a wire form, and every client creates other agents by intent, which the server resolves.
  • One list of registered agents (structured-agent-runtime-registrations.ts): each entry is an agent's definition, the factory that builds its adapter, where it can run (supportsLocation) and how this machine finds its account (resolveAccountHomePath, for both a read and a launch). The registry the router and host read, the published agent list, and the create path come from this one list, so an agent cannot be offered while its account cannot be found. Stored conversation admission is independent of this list.
  • Readable vs. startable. A saved chat record is readable when its bounded agent id, schema, account-home shape and conversation-pointer namespace agree with the record itself. Registration is not required to load, restore, reveal or read its journal. Only malformed or unsupported-schema rows are set aside, preserved byte for byte. Whether this build can start the chat (its pointers use the protocol this build drives the agent over, and its variable is that agent's own) is checked in one place (hostCanStartRecord), used by the start itself, the pre-send check and the restart offers. A chat this build can't start keeps its tab, history, close and delete, and a send gets the existing "not available on this host" refusal. A later build that changes how an agent connects or removes its registration therefore keeps the saved chat readable; execution becomes available again when a compatible registration returns.
  • Create path. The variable a new chat pins and where its account lives come from the agent's registration; the create path has no Claude/Codex branches. Claude and Codex resolve exactly as before (including WSL and Codex's side-effect-free read path). Adopting an existing conversation from history stays Claude/Codex, the only agents with transcript importers. An attach whose agent differs from its record's provider is refused, so the agent checked is the agent started.
  • Publication. New agentSession.agents returns every registered agent with the capability record its definition declares (rewind, compact, threadGoal, contextUsage, imagePrompts, steering, approvalEnforcement), so a client can show what an agent supports before any chat of it exists. A shared decoder reads the reply row by row: a row it can't read is dropped and the rest kept, an unknown value degrades to the one that claims least (steering: queue, approvalEnforcement: orca, a missing flag reads as false).
  • Negotiation: agent-session.structured.registered-agents.v1, used in both directions like agent-session.structured.v1:
    • A server advertising it accepts any agent it lists and serves agentSession.agents. Every server from this build advertises it; it runs only Claude and Codex until the ACP adapter registers more.
    • The client launch decision (resolveStructuredNativeChatSupport) offers an agent beyond Claude/Codex only when the server advertises the capability and listed that agent. No client passes a server's list yet, so nothing new is offered.
    • A client advertising it says it renders a chat tab of any agent its server lists. The server withholds every other agent's chat tabs from a paired client that doesn't (they are removed from the tab list and its groups, the same way Claude tabs are withheld from a client without the Claude capability), retitles them Update to view for a phone (the existing fallback row), refuses closing them from such a client, and keeps those agents' restart-resume offers (the "resume the chats that were working when Orca quit" prompt) away from it: it isn't listed them, its "resume all" and "dismiss all" reach only the offers it was shown, naming a hidden offer does nothing, and a resume's reply carries only offers and failures it can show. Whether a client can show an agent is answered by one rule, used for both tabs and restart offers, and against both current registrations and saved records: a client that can show that entire vocabulary takes the existing unscoped path and dismissal fence. A saved agent whose registration is gone still limits an older client's audience. For a client that can't, "dismiss all" forgets the offers it can show and keeps the others; it writes no dismissal fence (a filtered dismissal keeps hidden rows; a provider-scoped dismissal fence is deferred). No client advertises the capability yet.
  • Persisted tab state keeps any agent id; a value that isn't one degrades to absent (the tab survives and shows no chat), as before.
  • Model catalog types and its persisted rows take any agent id.

How an old client reads a row whose agent it doesn't know. It never receives one from a new server: the server withholds those tabs and offers from it. The one way an old build can still meet one is a downgrade on the same machine, reading what a newer build saved. Then its saved-tab parse degrades the agent to absent and keeps the tab and every other tab (it already used .catch(undefined)); its chat-pane check (isStructuredTab) is false for that tab, so it mounts nothing; and a release with the old record-admission rule refuses the record rather than mistaking it for Codex. A release that supports provider-independent storage can instead read that record. The cross-version test loads the release's own parsers, keeps Claude/Codex records readable, and verifies that this build reads the new agent's valid stored record.

Review fixes included here:

  • Removing an agent registration hid its saved chats (97020825d4e, c3b8f6f1387). Storage admission and read/reveal/restore used current registrations, so reopening a valid saved chat without its provider made the record disappear from the readable index. These paths now validate stored identity independently. Creation and acquisition still require a registered, compatible provider. The cross-version test no longer imports a registration fixture that later releases remove.
  • Dismissed restart prompts could come back (4f32dc9049e). The earlier restart-offer fix treated the local desktop as an older client (it doesn't advertise the new capability), so its "dismiss all" skipped the fence that stops a late teardown write from bringing a dismissed offer back. Today, with only Claude and Codex. The audience is derived from registered agents and valid saved records, so a desktop that reads all of them takes the original fenced path; one shared rule now answers "can this client show agent X" for tabs and restart offers.
  • A changed agent definition no longer hides its chats (212af519944, 7b52fb5c1c6). Readability used to require the agent's current protocol and variable, so a later build that changed either would silently set every existing chat of that agent aside. Readability now depends only on the record; startability is one check used by the start, the pre-send check and the restart offers, so an unstartable chat isn't offered "Resume" and its failure isn't marked retryable.
  • Account and location on each registration (e23946fe3a5). Claude's and Codex's account resolvers and location rules moved from named branches in the create path into their registration entries. createSupport no longer starts the host just to say no.
  • Smaller fixes. Attach refuses an agent that differs from the record's provider (f0e695f83eb); the model-catalog probe uses a chat's saved account only when this build can start that chat (38d510dbbc2); stale comments and an unused copy of the definition's fields are gone (1a569b82f77, 364f11fea9c); a short-lived per-chat dismissal fence that no client reached was removed before release (b557e5e4f18).

Why

A new agent should need only its adapter and its definition. That requires persisted storage to accept a well-formed agent id and creation to require an agent registered on the execution host, instead of both using the two current agents, and it requires old and new versions to keep working together, since users update the desktop, phone and paired servers independently. The common pattern does this with an open provider id on every row, a server-published list of providers with their capabilities, and a protocol version check.

Alternatives considered:

  • Use current registrations to decide which saved records can be read. Rejected: removing a registration must not hide an otherwise valid conversation. Stored identity is validated independently; the registration owns availability and launch configuration.
  • Keep checking each record against the agent's current protocol and variable when it loads. Rejected in review: a record lives forever but a definition can change, so a later build would silently hide every chat of that agent. Readability depends on the record; whether this build can start it is checked where it starts. Records of unregistered agents remain readable too. The variable, which becomes a child process's environment, is checked against the agent's own at the one place every start passes.
  • Let old clients receive the new tab and rely on their own fallback. Rejected: verified above that both current clients list it with an empty pane. Withholding it on the server is the existing pattern for Claude tabs.
  • Put the capability records on createSupport. Rejected: that answers one named agent for one workspace; a client needs to learn which agents exist.
  • Widen the attach-by-pointer request too. Not done: no client sends it, and a non-Claude/Codex pointer has no wire form (refactor(native-chat): keep the provider resume handle opaque to shared code #24991). Widening provider alone would admit a request that can't name its conversation.

Differences from the common pattern

  • Capabilities are negotiated per server instead of exact protocol versions being required (intended). The common pattern refuses a connection whose protocol version differs and updates the older side. Orca's paired clients and servers run mixed versions as the normal state (docs/reference/remote-wire-compatibility.md), so each new surface is advertised and checked per server.
  • The server withholds rows of agents a client can't render instead of every client looking agents up and degrading (intended). Shipped Orca clients can't render an unknown agent, so they must never receive one; this is the existing Claude-tab rule generalized.
  • The published list carries the agent id and its capability record, not display names or icons (intended). Both sides ship the same agent catalog for labels and icons, and a client offers only agents it knows. A field can be added later without negotiation.
  • No client reads the list or advertises the capability yet (temporary). The ACP adapter PR, which registers the first new agent, wires the desktop to read agentSession.agents and advertise the capability once it renders those chats.
  • Phones get the existing Update to view row for a new agent's tab (temporary), until a phone build renders those chats and advertises the capability.
  • Stored conversations are read independently of provider registration; registration gates creation and execution (matches the common pattern). Storage validates the record's own shape and namespace, while routing and launch configuration come from the runtime registry.
  • Account resolution and location support live on each agent's registration (matches the common pattern, where the registration carries how to launch the agent). The ACP adapter PR (feat(native-chat): Grok as a structured chat over the Agent Client Protocol #25225) added a narrower hook for this; it should adopt this one when it merges.
  • A dismiss-all from a client that can't show every agent writes no dismissal fence (temporary): a filtered dismissal preserves hidden rows without a global fence. A follow-up must add a provider-scoped dismissal fence before clients support filtered restart prompts.
  • The server's generic launch (agent.launch) still probes only Claude/Codex for structured support (temporary, agent-launch-mode.ts), and its preflight passes no registered-agent list. The ACP adapter PR (feat(native-chat): Grok as a structured chat over the Agent Client Protocol #25225) updates this.
  • Attach by a client-supplied pointer and adoption from history stay Claude/Codex (intended for now): no client needs either for another agent, and adopting needs a transcript importer.

Linked Issue

N/A (internal stack: native chat for more agents).

Visual Proof

N/A: no visual layout or UI element changed. Saved-chat readability is covered by fake-host save/reopen tests; no application was launched for this revision.

CI status

Merged main 4e64fa9940f5a4a9bfb7b021a3c405fd76643ca5 after the registry/lifecycle changes landed, resolving the later launch-support and model-catalog conflicts with main’s behavior preserved. Current head 40fa4aace610902c414042ec82c198313a59409f has fully terminal CI: typecheck passed, with only the confirmed main release-checkout and cache-scan assertions plus their final verifier red; details below. Earlier head f164943c5dd95a850aca0ceb8411ec927ab76033 passed static analysis/typecheck and 126 cross-version cases with only main #25731’s unchanged parser-dependency assertion failing. The second catalog merge exposed one new main Codex test fixture missing A3’s required startability callback; that fixture is corrected, all 15 dependency objects were scanned, and its 11 tests plus scoped lint pass. The catalog merge itself passed 54 targeted tests and scoped lint.

Testing

  • I manually tested these changes locally
  • Automated tests added/updated, or explained why not below

New tests:

  • tests/e2e/cross-version-wire/cross-version-registered-agents.unit.test.ts, against the selected release baseline, with the old side's capability lists read from its checkout minus the new capability:
    • old client, new host: a new agent's tab never reaches an old desktop, and the Claude and Codex tabs it gets equal what the release publishes to its own client; the release's saved-tab parse keeps every tab (a release that predates registered agents degrades only the new agent's id and mounts no chat pane for it; a later one keeps the id); the release classifies a new agent's record according to its own storage contract and still reads Claude and Codex records written by this build. This build explicitly accepts that valid new-agent record.
    • new client, old host: the launch decision does not offer the new agent when the capability is absent, while Claude/Codex are offered exactly when the old server serves structured chat; the release's server refuses a createSupport for the new agent (why the gate is needed).
    • Claude/Codex both ways: both builds accept createSupport for Claude and Codex; a new server sends today's desktop its Claude/Codex tabs unchanged.
  • agentSession.agents is in the existing cross-version manifest, so the existing suite asserts it reaches the host and answers on this build, and answers method_not_found on the release.
  • Unit tests: provider-independent records (read without a registration, reject malformed ids or inconsistent namespaces, readable but not startable after the agent's protocol or variable changes, Claude/Codex bytes unchanged), the account-home variable check, the widened request schemas (and that attach stays Claude/Codex), the reply decoder (salvage, degrade, dedupe), the launch decision, saved-tab parsing, tab withholding/retitling/close refusal per client, and a pin that the shipped Claude/Codex definitions declare exactly the storage older builds wrote.
  • Restart offers: structured-agent-session-restart-resume.test.ts (the review's reproductions through the real RPC handlers: an older client's list, dismiss-all, named dismissal of a hidden offer, continue-all and a named continuation's reply; a new client and the server's own process still act on every offer) and structured-agent-session-restart-audience.test.ts (the real host and recovery file: an offer and a recorded failure the caller can't show are neither listed, dismissed, reserved nor returned, named or not, and a caller that can show them still can). One capsule test covers forgetting every record except kept ones without a fence.
  • structured-agent-stored-definitions.test.ts pins the Claude/Codex protocol and account-home variables older builds wrote. A save/reopen regression restores and reveals a chat, reads its saved history with an empty registry, refuses acquisition without changing the fence, and resumes after registration returns.
  • Ablation: with the old tab-projection rule restored, the new projection and cross-version tests fail (4 tests), so they guard the old-client case.

What I verified in this revision:

  • CI on the second catalog merge found one new Codex fixture missing the required startability callback. The one-line correction supplies it; all 15 direct catalog dependency objects now supply that callback. The changed explicit test file passes 11 tests and scoped format/lint. No production behavior changed.
  • A second main conflict in the model catalog keeps main’s per-session listing behavior and extends A3’s general agent type through two new private methods. Every other A3 patch line is unchanged. Five explicit catalog/startability suites / 54 tests and scoped format/lint passed. New head 2ed6c25dee8facea6b4638c2c1c94730bf131809 has fresh CI pending; the preceding head passed typecheck and 126 cross-version cases, with only main test(cross-version): load releases that import a package main dropped #25731’s unchanged parser-dependency assertion failing.
  • The focused storage/host/reveal suites passed: 4 explicit files, 58 tests. The save/reopen regression checks history with no registration, restored visible state, reveal, refusal to acquire without changing the ownership fence, and successful resumption after registration returns. Fresh creation without registration is also refused before a record is written.
  • A broader run passed 199 tests; an obsolete ownership-status expectation was corrected and then passed in the focused run. Four suites could not import the stale shared stream-json dependency.
  • A temporary ablation restoring the reveal availability gate makes the absent-registration regression fail at reveal. The source was restored byte for byte afterward.
  • The fresh read-only review found one P2 in restart audience filtering after registration removal. It is corrected by deriving the audience from registered agents plus valid stored records, rather than registrations alone. New cases cover empty and Claude/Codex-only registries, and hidden failure/offer retention through list, named continuation, named dismissal and dismiss-all.
  • Follow-up local tests passed 15 host/policy cases. Scratch derivatives invoking the real schemas/handlers, host and disk-backed recovery capsule passed 11 cases, including the existing late-teardown dismissal fence; RPC transport was omitted because the aggregate dispatcher import hits the stale dependency. Restoring the old audience derivation in the scratch runtime fails both new cases by returning a hidden Grok failure.
  • After merging the landed registry/lifecycle changes, all 3511 A3-owned inserted/deleted patch lines match the prior A3 delta exactly. Scoped format/lint of 106 code files passed; 44 explicit targeted suites yielded 425 passing tests, 8 skipped assertions and 7 suites blocked on the stale shared dependency, with no assertion failure.
  • The merged head failed CI changed-code quality on four unsafe reveal-fixture casts. The follow-up narrows the reveal store dependency to its only used method and removes those casts; 2 explicit suites / 12 tests and the casting-rule lint passed. Runtime behavior is unchanged.
  • Cross-version host fixture now supplies the known-agent method used by restart handlers. Four isolated cases using the actual fixture, schemas and handlers passed; the full local dispatcher suites were blocked by the stale shared import (27 skipped). Scoped format/lint passed.
  • The later main merge preserves main's removal of terminal-command overrides from native-chat support. Only six now-unneeded adaptation lines leave the A3 patch; every other file's own patch lines match the pre-merge delta. 16 explicit suites yielded 181 passing tests, 27 skipped cases, five stale import failures and the same main-only parser-dependency assertion failure; 106 code files pass scoped format/lint. Main's failing test was preserved exactly and routed upstream.
  • Formatting, scoped lint and whitespace checks passed; no branch-owned lockfile change. The submission-position test retains exactly one used codexProviderHandle import and main's complete version.

What I did not verify locally:

  • The explicit cross-version registered-agents suite and three restart RPC suites could not start their tests because the symlinked install lacks stream-json/core/utils/flex-assembler.js. The cross-version suite's eight tests were skipped by that import failure; fresh-install CI is the authority for that suite.
  • No local typecheck, application, Electron build, or real agent CLI was run. The shared installation was not changed. No desktop, phone, SSH, Linux or Windows application reproduction is claimed.
  • Passing tests show what did not regress; the architecture argument is the separation of stored identity from execution availability described above.

Review

One fresh read-only review of the saved-chat fix found the restart-audience P2 described above and no P0/P1 findings. The follow-up correction was checked with the real handlers and disk-backed recovery store, including a regression ablation; the reviewer's original reviewed head preceded that correction.

Agent skill upstream boundary

  • Not applicable, or this change follows docs/reference/agent-skill-sharing-upstream-boundary.md and copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.

Notes

  • Security: the stored account-home variable reaches a child process only through the one start check, which requires the agent's own declared variable, and each adapter's launch still refuses any variable but its own. Request validation still bounds the agent id; the server refuses an unregistered one.
  • SSH / remote: agents, records and account homes resolve on the server that runs the agent; clients only learn the list. SSH hosts have no structured chat, unchanged.
  • Mobile: phones don't advertise the capability; a new agent's tab reaches them as the existing Update to view row. No new method is on the phone allowlist.
  • Backwards compatibility: covered by the cross-version tests above. Claude/Codex requests, replies, tabs and records are byte-identical.
  • Performance: bounded stored-shape validation on load/write; registration lookup at creation/start.

Checklist

  • This PR is small and focused
  • I explained what changed and why (ELI5, the user-facing before/after, the mechanism, and why over the alternatives)
  • Before/after screenshots or videos attached for UI changes, or N/A with reason
  • Self-reviewed for correctness, security, and performance
  • Cross-platform, SSH/remote, and path/shortcut impact considered (or N/A)
  • pnpm lint, pnpm typecheck, pnpm test, and pnpm build pass (or CI will cover; local preferred)

Current-head CI finished

At 40fa4aace610902c414042ec82c198313a59409f, all current-head runs and jobs are terminal: 12 checks passed, 17 were skipped and 3 failed. Static analysis/typecheck, mobile, both packaging jobs and four unit shards passed. A runner shut down during the remaining shard; one targeted retry completed with 21,333 tests passed, 177 skipped and only the previously confirmed main cache-scan assertion failing. Cross-version completed with 126 passing tests and only main’s release-checkout assertion (#25731) failing. Final verification follows those two results; no branch-owned test failure remains. The retry keeps its original merge source, so later main fixes to those assertions are not part of this run. No local typecheck was run.

…ed code

Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).

Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.

The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).

No user-visible change.
…l identity

The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.
A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.
…initions

The structured-chat shared layer named Claude and Codex in a closed router
type, in option reads, in frame classification, and in renderer label and
validation branches, so adding an agent meant editing shared types. Each
agent now has a definition (capabilities, resting option rules) owned by its
own module; the router is a registry of definition + adapter pairs built at
runtime setup, and shared code reads the definition instead of the name.

No behavior change for Claude or Codex; no wire or stored shape change.
The host types, conversation commands, provider-session ownership and the
runtime's create path spelled the structured agent list inline; they now use
the one shared type the record and wire already validate against.
…gent-registry

# Conflicts:
#	src/renderer/src/lib/structured-agent-session-provisional-tab.ts
… definition lookup

The at-rest option reads looked definitions up in a second, separately populated map, so a
registered definition could route and declare capabilities while the static map decided which
options a chat at rest accepted. The router now answers `definition(agent)` from its
registrations, the at-rest readers take it through the host's adapter, and the built-in
definitions are listed only where the runtime composes the router.
…o registered agents

A host's structured agents are the ones its runtime registered. Records, RPC
params, persisted tabs and the model catalog accept any registered agent instead
of naming Claude and Codex; each agent's definition declares the transport its
handles live in and the variable its account home pins. A new runtime
capability, agent-session.structured.registered-agents.v1, advertises that a
host accepts and lists its agents (agentSession.agents, with each agent's
capability record), and the host withholds any other agent's tabs and restart
offers from clients that do not advertise it.
…lient can show

A paired client too old to show an agent's chat was listed only the offers it could show, but
dismissing or continuing all reached every offer on the host, and a named continuation answered
with the host's whole remaining inventory. The client's audience now goes to the host with every
restart operation: only offers it sees are reserved, dismissed or returned. Without an audience
(this host's own process, or a client that shows every agent) nothing changes.
…s own terms

The registered-agents downgrade test assumed its baseline release predates registered agents: it
expected the saved-tab parser to erase an unknown agent and called the record reader without the
agents list. Once a release with this change becomes the baseline, both break. The expectations now
follow what the baseline host advertises, and an agent-registering baseline is handed its own
Claude and Codex storage.
…'s agent registrations

Which agents a stored record may name and which agents the router drives came from two lists in
the runtime, so a newly registered agent could be routed while its records were set aside. One
list of registrations now holds each agent's definition and the factory for its adapter: the
store's admitted agents are derived from it before the store opens, and the adapters are built
from it once it has.
…nership

Ownership rows name any registered agent since the stored records opened to them; the adoption
check compares by agent, so it takes the same open id. Only Claude and Codex still adopt.
…n suite

Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.
Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.

State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.
@brennanb2025

brennanb2025 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Earlier CI merges missed main's restored provider-handle import and the test-only parser dependency still needed by old release checkouts. Merged updated A2 db71a89d7438 (includes main 08a970a3f117); the branch's own changes are preserved, the submission-position test matches main with exactly one used import, and main's package/lock fix is inherited through the merge with no branch-authored dependency edits. The prior broad run passed 1,659 tests; fresh-main checks passed 118 tests, with remaining local failures limited to stale stream-json/stream-chain dependencies, and scoped oxlint passed; new-head CI is pending.

@brennanb2025

brennanb2025 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Saved valid chats remain readable when registration is unavailable; live adapter checks remain at creation/acquisition/start. Current head 40fa4aace610902c414042ec82c198313a59409f; scoped tests/lint pass and CI is fully terminal with passed typecheck. One runner retry completed 21,333 passing tests; only confirmed main release-checkout/cache-scan assertions and their final verifier are red (12 passed / 17 skipped / 3 failed), with no branch-owned failure.

brennanb2025 added a commit that referenced this pull request Oct 6, 2026
@brennanb2025

Copy link
Copy Markdown
Contributor Author

CI at 40fa4aa: every PR-owned check passes (static analysis + typecheck, mobile, packaging, cross-version registered-agent suite 126 passes). The red checks are main's own failures, identical on current main: (1) cross-version release-checkout.unit.test.ts expects @streamparser/json to be absent (#25731; main restored the package), (2) unit shard 2/5 file-explorer-watch-reconcile.test.ts bounded cache scan (2001 expected, 4002 seen; reproduced on main), (3) verify, which aggregates them. The A3-F1 readability fix and the restart-audience fix both passed a narrow reference re-audit.

@brennanb2025
brennanb2025 marked this pull request as ready for review October 6, 2026 07:35
@brennanb2025
brennanb2025 merged commit 9905765 into main Oct 6, 2026
60 of 66 checks passed
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
…-home

#25159 moved structured account-home resolution into per-agent registrations and put Codex's trust write inside resolveCodexAccountHomePath, before the home is picked or adopted. Registrations now get an optional afterAccountHomePinned hook, which the create resolver awaits once after the replay check with the final home (adopted or selected); a read never runs it. Codex supplies the trust write there; Claude supplies none.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
…at quit

Main's #25159 left the host at its line limit; quit's stop now disposes both in one
expression instead of a block.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
Main now carries A3 #25159 (squash 9905765) and C5 #25181 (squash c4ea14c); both squashes equal
their final branch heads, so the merge was computed against main's tree with those heads as extra parents
(scratch commit, never pushed) and committed as an ordinary two-parent merge of origin/main.

Resolutions: dead-generation settlement keeps D3's move of the unfinished-work helpers; provider support,
create support, the record's stored-agent check and its tests take main's A3 (records no longer admitted
per registration). D3 callers of the removed registry follow: agent-launch-mode reads the registered
agents from STRUCTURED_AGENT_RUNTIME_REGISTRATIONS; the Grok host rig opens the store without agents; the
"admits Grok's records" registration test is gone with the registry.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
…ach tests

Merged main (#25159) requires agents on attach and an agent definition for catalog access; the PR's new acquisition attach tests and Fast-mode tests now pass them.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
… held card (#24660)

* fix(native-chat): keep a message accepted before a quit or crash as a held card

A send the host accepted while the agent was still starting, and never handed
over, was rejected unseen at quit or at the next open after a crash. The next
open now keeps a person's message (typed, or a launch's first prompt) as a
waiting card at the head of the queue, held until Resume, Send now, Edit or
Delete; quit no longer rejects it. Each submission records its source so a
restart knows which leftovers to keep. A direct send's replay answers from its
own record, never the queued arm. Cards shown without the queue capability
hide Turn off queueing and the steer chord.

* fix(native-chat): keep an unsent message as a held card at every close, not only a restart

The host now keeps a person's message it accepted and never handed over with one
rule wherever it can no longer hand it over: a quit or crash (settled at the next
open) and a close of the chat (tab close, worktree teardown, orchestration stop).

- The hold is card state: a new per-card hold_reason 'kept', published as the
  existing pausedReason, instead of a fake host_instance value. host_instance
  means the owner again, and the pause clause, adoption filter and /clear carry
  special cases are gone. A kept card holds the cards behind it until the
  person sends, edits or deletes it; a send_failed card still does not.
- A card hand-off rejected by a restart or a close returns as a kept card
  (rejectedDraftSettlement), so a person's next message can never release it.
- One hold function, parameterized by cause (hostRestarted / chatClosed),
  replaces the close path's plain rejection.
- The phone shows published cards and per-card holds whatever the queue
  capability says; only queueing a new send stays gated.
- source gains 'dispatch' for the orchestration preamble (still rejected); an
  unknown source is kept as written and never takes the legacy rule.
- Kept cards from earlier settlements stay ahead of a batch's new ones.

* test(native-chat): the dispatch preamble records its source

* fix(native-chat): skip a kept card like a failed one, and make a quit leave the queue as a crash does

- A kept card is held on its own, as a send_failed one is: the queue sends the
  cards behind it, and the "a message ahead needs attention" caption no longer
  appears behind it (desktop and phone).
- Quit disposes the queue's drain together with delivery, so it mints no
  hand-off that only the next process could settle.
- Only a Send the person asked for (origin client) that a restart or close cut
  short returns kept; the queue's own hand-off returns where it stood, under the
  restart's pause, as on main.
- Comments that said only a capable host gets cards or the card actions now say
  the capability gates only queueing a new send.

* refactor(native-chat): one dispose gate in the queue drain's step

* test(native-chat): the downgrade test names the older build's /clear exception

* fix(native-chat): re-check the drain's quit gate right before it appends a hand-off

A drain step already past its first check when quit begins no longer makes a
hand-off. Adds regression tests for that, for Send now on an ordinary card cut
short by a quit, and for the open-time repair deriving kept from the hand-off's
origin; fixes the pause and settlement comments that called a kept card one the
queue never passes.

* fix(native-chat): a card held on its own starts no restart pause for the others

After a second restart a kept (or send_failed) card is another process's, but it
waits for its own Send, so it no longer pauses every other card under
"Queue paused because Orca restarted".

* fix(native-chat): record on a kept send which card holds it, so no trace survives an Edit or Delete

The rejection row of a send the host kept as a card now names that card
(`keptAsQueuedMessageId`), in the same transaction that writes the card, and the
fold publishes it on the submission. The shared projection draws such a send
only as its card: once the card is sent, edited or deleted, neither the send
nor the sending desktop's local copy of it shows.

* chore(native-chat): leave the unused submission schema as main has it

No client parses published submissions with it (they arrive as typed frames,
and the host's history pages carry the field, as the Edit/Delete test reads);
listing the field put the file over its line budget.

* fix(native-chat): retire a kept send's local copy instead of only hiding it

The outbox reconcile and the send disposition drop an entry whose submission
the host rejected as kept as a card, as they already do for a Stop's
withdrawal, so the copy never comes back as "Not sent / Retry" once the
submission falls out of the loaded page. The projection reads the reconciled
outbox, so its separate filter goes.

* test(native-chat): move the queued-message rig's scripted provider into its own fixture

The rig fixture grew past the 300-line limit once main's changes merged in.

* fix(native-chat): hide a kept send by its own record, not by its card still existing

The transcript hid a rejected send while a queued card held it under its id, so the card's Edit or
Delete brought back a "Not sent" row. It now reads the send's own keptAsQueuedMessageId, which the
host records with the rejection, and the live card list is no longer threaded to the transcript or
the delivery notices.

* refactor(native-chat): move a sent message's row writes into their own journal collaborator

The journal store went past its line limit once main's ledger receipt joined this branch's
transaction hook. The submission and dispatch-transition writes, and what commits in their
transaction, now live in JournalSubmissionWriter; the store's methods delegate to it unchanged.

* fix(native-chat): list a kept send's card id in the submission schema

The schema drops keys it does not list, so a reader that kept a parsed submission would lose
keptAsQueuedMessageId and source, both persisted with the row. The submission schema and the
failure fact it shares with item bodies move to their own modules, with room for both fields.

* test(native-chat): match the transcript and outbox hook signatures main and the swap changed

* test(native-chat): import the journal types once in the queue-delivery test

* fix(mobile): a resend the host kept as a card shows no error and returns no text

The host answers a resend of a message it kept as a card with that
message's rejected submission, marked keptAsQueuedMessageId. The phone read
it as any rejection: "Message not sent" and the text back in the composer,
while the card showed the same text. It now answers like a queued send, as
the desktop's send disposition does: the id is spent, no error, and the card
holds the text.

* test(native-chat): pin that quit's first step stops the queue's hand-off

Quit now stops delivery, the queue drain included, at its first step
(stopDelivery), before teardown drains recovery. The drain-step quit test
runs from that step as well as from the flush.

* test(native-chat): give cards their source and store unknown sources as another build would

#25078 made a card's source required, so the tests that insert a card pass the
person's. The hold's unknown-kind and unreadable-source cases now rewrite the
stored row the way a newer build would leave it, instead of casting a type.

* refactor(native-chat): stop delivery and the queue drain in one line at quit

Main's #25159 left the host at its line limit; quit's stop now disposes both in one
expression instead of a block.

* test(native-chat): hold the drain step without reading the call stack

Bun formats a method's stack frame without its class ("at step"), so the quit
test's caller check never matched, the step was never held, and both cases timed
out once CI ran Vitest on Bun (#25840). Only the drain step heals owed queue
bookkeeping, so the hold needs no caller check.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
…-history

- #24660: the row writer's planned write is public writeRows (the new step writer
  calls it); the transactional write that also stores the chat's status is
  commitRows, so every path still writes the status row once per transaction.
- #25181: the open settlement plan ends a running call as its turn's row ended
  (runningCallEnd / terminalAgentJournalBody) and revises calls an unproven settle
  closed once a proof names their owner; the status facts read
  requiresTerminalSettlement, the same facts as before. Liveness keeps main's
  lostLiveWorkJournalBody under the per-item roster revision.
- #24660: the open plan settles what a gone process left queued through
  holdUnsentSends (a person's message becomes a held card); its leftovers are
  exactly the queued sends written before the open.
- #25159: reading and settling a chat need no adapter. The persisted tab listing
  lists every tab with a record, startup settles any chat whose record exists,
  and adapterSupportsRecord / hostCanSettleRecord are gone, as on main.
- Line caps: the runtime's startup step moves into its own layer
  (orca-runtime-structured-agent-session-startup-step.ts), and the host exposes
  its restore as one `startup` surface instead of four pass-through methods.
- The golden digest is unchanged, so the status rules stay at 3.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
…agent-tree-job-object

Main (#25159) kept the Windows creation-time refusal in Codex launch
resolution and moved the pinned config variable onto the Codex agent
definition; the refusal stays removed and main's pin is taken.

Main's host constructor grew to its line limit, so the held-child read moves
into the runtime state: it now takes the host's conversations (required) and
builds the renewer's read from them and the adapter.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
* fix(native-chat): keep Codex chat actions independent of catalog listing

* fix(native-chat): accept a Codex option pick before the model list arrives

- A live model, effort or Fast pick with no stored model list is kept as the
  next turn's intent instead of refused; Codex judges the model on turn/start
  and Fast without a known tier still sends Standard.
- Restoring saved Fast checks the saved model the turn will send, not the
  resumed thread's old model; the thread's effort is only reused for its own
  model.
- turn/start reads the exact Fast tier from the account's stored catalog at
  send time, so a tier any writer stored applies; the per-session copy is gone.
- Revert the delivery-loop extraction; that file matches main again.

* test(native-chat): pass the agent registry to the new acquisition attach tests

Merged main (#25159) requires agents on attach and an agent definition for catalog access; the PR's new acquisition attach tests and Fast-mode tests now pass them.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
B0 switches automatic prompts to stacking (no step-aside) and brings main
forward to 4c35516, including the per-client restart-offer audience
(#25159). The paired listed dismiss now honours that audience too.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant