Skip to content

Translate ACP traffic into shared timeline events - #25090

Merged
brennanb2025 merged 97 commits into
mainfrom
brennanb2025/acp-d2-translation
Oct 6, 2026
Merged

brennanb2025 merged 97 commits into
mainfrom
brennanb2025/acp-d2-translation

Conversation

@brennanb2025

@brennanb2025 brennanb2025 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor
Files Added Deleted Net
Test 26 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​4469 0 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​4469
Prod 43 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​4657 $\color{#cf222e}{\Huge{\mathbf{−}}}$​33 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​4624

ELI5

Orca has a client for the Agent Client Protocol (the open protocol Grok and other agents speak), but nothing turns what an agent sends into native chat history yet. Recorded Grok traffic shows the translation needs explicit rules: Grok's internal notifications would flood the conversation, separate replies would merge, context usage would be overstated, and a background command would lose its status when the tool that launched it closes.

Reopening a saved chat is the other hard case. When an agent loads a saved session, it first replays that session's whole history. An earlier version of this branch tried to merge that replay into whatever Orca had already recorded, to finish a reply an Orca crash had cut off. That made the translator read the whole recorded chat on many frames, and it decided "is this chat new?" from a view that is always empty at that point in a real resume. This version drops the merge and, like the common pattern, never writes a replayed history at all: everything an agent replays while a saved session loads is dropped, except what it says about context usage.

No user-visible change. Nothing here is connected to an agent or to the chat UI yet; the Grok runtime adapter (#25225) does that. Once connected:

  • A reopened Grok chat shows Orca's own record of it (what Grok replays while loading adds nothing), and a cut-off reply reads the way Claude and Codex show one since fix(native-chat): show a reply cut off by an Orca crash or quit like a finished turn, with one explanation #25043.
  • When a Grok turn fails after it started (for example its API returns an error), the turn shows one error row with Grok's own words for why. If Grok sends no words, the row says "Grok ended this turn with an error.", or "Grok usage limit reached." for a rate limit. It never says the message was refused: Grok accepted it and started working. A message Grok refuses before its turn starts is the runtime adapter's (feat(native-chat): Grok as a structured chat over the Agent Client Protocol #25225) to report, on the message itself.
  • A running background command reads "Started background command """, not Grok's internal "Background task started", and a background monitor stays labelled a monitor after the agent reads its output.

Dependencies #25064, #25141 and #24990 have landed on main. This diff now contains only Agent Client Protocol event translation and its tests; the shared assembler, transitions and protocol client come from main. Its base is main.

A few terms used below: the journal is Orca's stored record of a chat; the assembler (#25064) is the shared code that turns an agent's events into journal rows; a turn is one request and everything the agent does for it.

What Changed

Translation (unchanged by this revision):

  • Replies, reasoning, partial tool results, file changes, plans and permission requests become shared timeline events. The shared assembler owns row identities, placement and settlement.
  • Grok's internal setup, hook and settings notifications are ignored. Unknown standard updates keep the bounded fallback row; reasoning signatures are never stored.
  • Background commands become one task row beside the tool that launched it, independent of the turn ending. Completion notices keep success, failure, signal and explicit stop outcomes; the session ending leaves an unfinished task "unverifiable".
  • A task's sentence uses the provider's description, falling back to its command, both while it runs and after. Grok's start summary ("Background task started") and the command's output never become the sentence: a summary is used only once the task has finished.
  • A task keeps the kind it was first given (command, monitor, …) unless a later notice names a kind itself; a task-output result whose command starts with [monitor] or [monitor: is a monitor, because that prefix is the only way those results name one.
  • A completed kill call ends the tasks it reports killed. When its result is not in that shape, a completed kill_command_or_subagent call still ends the tasks named in its input: an inference no captured kill confirms yet.
  • A prompt identity is returned before dispatch so the runtime can inject it. Context occupancy comes from the newest model call, not a sum over the turn (the nine-call recording reads 39,190 tokens, not 331,949).

Failed turns (this revision):

  • When a turn ends with Grok's error or rate_limit stop reason, the translator writes one status row inside that turn, before its end. It is a plain status row in the error tone whose text is Grok's own reason: the same shape the Codex lane writes for an error that ends a Codex turn already running. It carries none of Orca's "message was refused" wording, because that is for a message the provider never started.
  • If Grok sends no reason, the Grok dialect (the Grok-only part of the translator) supplies the sentence: "Grok ended this turn with an error.", or "Grok usage limit reached." for rate_limit. Any other agent gets " ended this turn with an error.", from a name the runtime adapter passes in ("The agent" until it does).
  • Grok sends that reason up to four times: in a given-up retry notice, in the turn's end notice, in a separate prompt-completion notice, and in the error answer to the prompt request. Whichever copy arrives first writes the row; a later copy only fills in a reason the row still lacks, so there is never a second row. The error answer is read only for a turn that already failed.
  • Grok-specific parsing (which notices carry the reason, and that the error answer keeps it under data.message) and the Grok-named sentences stay in the Grok dialect; the translator only knows "this turn failed, here is the provider's reason".

Saved history:

  • Everything the agent replays while the session loads is dropped, except what it says about context usage, which still updates the chat's latest turn. This is the common pattern.
  • Removed: writing the replay into an empty record ("adoption", which an earlier revision of this branch had and the common pattern never does), the partial-history merge, its per-turn checks of the recorded chat, and every place the translator searched the recorded chat (tool, background-task and request placement). The translator lives exactly as long as its agent process, so nothing has to be found again after a restart.
  • Bounds: the translator's tool snapshots now use the assembler's own limit on open work (128 entries, 1 MiB). The assembler counts each running tool at no fewer bytes, so a running tool's snapshot is only evicted after the assembler has already refused work past its limit, which ends the session. Finished tools are evicted first; an update for one the assembler drops as already settled. A tool's turn is held in one place at a time: on its snapshot while that lives, then in a small record (128 entries) that outlives the turn, so a late first notice of a background task still lands beside its launch tool.
  • A prompt's turn is remembered as started once it opens, so a late frame for a prompt that already ended neither reopens it nor becomes the turn later unlabelled frames join. Background-task snapshots keep their 128-entry bound: past it, a task untouched while 128 others updated loses its earlier label, parent tool and output file on its next update, but its row and placement stay.
  • Absorbed the assembler's latest revision (Add a shared timeline assembler for structured agent chats (not wired yet) #25064): after a person's Stop, only the stopped turn's text is dropped and its pending questions cancelled; its running tools now settle at the agent's own end of that turn. The translator needed no change for it.
  • Absorbed the protocol client's latest revision (Add standalone Agent Client Protocol client layer #24990), and dropped the translator's own copies of what the client now owns:
    • Permission requests: the client reads them leniently (it keeps what it can read and reports what it had to drop). The translator now reads them with that same reader instead of its own stricter parse, so a permission the client delivered always becomes a card rather than being cancelled for a field the card does not need.
    • A failed prompt's row comes only from the agent's own error answer (the client's agent-error type). The generic " ended this turn with an error." appears only when that answer carries no words. An error Orca raises itself (a timeout, an answer it could not read, a closed connection) never reaches the translator and writes no row here; showing it is the runtime adapter's (feat(native-chat): Grok as a structured chat over the Agent Client Protocol #25225) job.
    • Session updates are read with the client's own reader rather than a second parse.
    • The client's enum fields now accept values newer than this build. An unknown tool status keeps the tool's last known state until it settles, an unknown plan status reads as not done, and an unknown stop reason ends the turn with no success or failure verdict.
    • Extension notifications reach a new runtime option; delivering them to the translator is the runtime adapter's job (feat(native-chat): Grok as a structured chat over the Agent Client Protocol #25225).

Why

The agent parser should describe what it receives; the shared assembler enforces lifecycle rules. Dropping replayed history entirely, as the common pattern does, removes a whole class of merge bugs (overwriting a reply a person stopped, regressing a settled turn, duplicating text) instead of guarding each one. The alternative, finishing a crash-cut reply from the agent's saved history, would make Grok behave differently from Claude and Codex and needs the translator to read the whole record; the trade is that a Grok reply cut by an Orca crash now reads as cut off, like the other agents.

A failed turn's reason goes into a status row, which clients already render, rather than a new kind of notice, and it is worded the way the Codex lane words an error that ends a running turn: the provider's own text. Writing it from the translator, inside the failed turn, covers the real order Grok uses, where the turn has already started when it fails. A prompt refused before its turn starts is a different case: the runtime adapter reports it on the message, in Orca's refusal words.

Differences from the common pattern

  • Matches the common pattern — loading a saved session: what the agent replays never reaches the record; only context usage is kept.
  • Intended — unmarked frames during a load are translated as live: the common pattern drops every frame that arrives while a session loads. Grok marks each replayed frame, so a Grok frame without that mark is treated as live output. In the one recorded Grok load, loading a chat Orca had already recorded wrote nothing, so no replayed frame arrived unmarked. A generic agent that marks nothing keeps the common behaviour: everything during its load is history.
  • Partly matches the common pattern — task sentence: like the common pattern, a running task never shows Grok's start summary. Two intended differences: Orca names a task by Grok's description and falls back to its command, where the common pattern uses the command's first line (the description is what Grok wrote for a person); and Orca never turns the command's output into the finished task's sentence, where the common pattern shows its first non-empty line (output can be anything, including a 16 KB log; the output-file field keeps it reachable).
  • Matches the common pattern — monitors and known kinds: a task-output result whose command starts with [monitor] or [monitor: is a monitor, and a task keeps the kind it started with.
  • Partly matches the common pattern — kill evidence: the common pattern treats a completed kill call whose result reports the task killed as the task's end, and so does Orca. Intended difference: Orca also ends the tasks named in a completed kill call's input when its result has another shape, so a killed task still settles if Grok sends no completion notice for it (not captured either way yet); pending or failed kill calls prove nothing. A real kill capture should confirm this once the runtime adapter can record one.
  • Partly matches the common pattern — failed turns: the common pattern marks the turn failed and attaches an error message to it; Orca marks the turn failed too. Intended difference: the message is a status row inside the turn, holding the provider's own text, because that is the row Orca's chat already renders for an error that ends a running turn (the Codex lane's), so no new kind of notice is added.
  • Partly matches the common pattern — rate limits: with no text from Grok, the row says "Grok usage limit reached.", like the common pattern's "Usage limit reached" label. Intended difference: when Grok sends its own text, the row shows only that text, as for every failed turn, rather than the label with the text under it.
  • Intended — context occupancy: the last model call, not the turn's aggregate, which counts the same context repeatedly.
  • Intended — bounded snapshots: tool snapshots are bounded at the assembler's open-work limit and background-task snapshots at 128, as described above.
  • Temporary: a tool left running by a cancelled turn still shows as failed. Show a tool call you stopped as interrupted, not failed #25181 shows it as interrupted.
  • Temporary: nothing is connected: agent launch, prompt-identity injection, request delivery, extension notifications and per-host capability advertisement belong to the runtime adapter (feat(native-chat): Grok as a structured chat over the Agent Client Protocol #25225), which must land them before any client offers the feature.

Linked Issue

Part of the Agent Client Protocol integration following #25064, #25141 and #24990; no separate issue.

Visual Proof

N/A — nothing here reaches the renderer or a running agent.

Testing

  • Automated tests updated: 174 tests pass across the 19 test files in src/main/acp (translation, fixture privacy scan, protocol client), and 100 across the 13 assembler test files (72 files, 756 tests across every test file this PR's stack changes in src/main/acp and src/main/native-chat).
  • New with the protocol client update: a tool with a newer kind or status stays a tool row in its last known state and settles normally; a plan entry with a newer status reads as not done; a newer stop reason ends the turn without a verdict or an error row; a permission the client could only partly read still opens a card with what it could read, and one with no tool call id is refused.
  • New in this revision: a failed turn from Grok's real failure shape (scrubbed to a placeholder reason) gets exactly one row in the failed turn with the reason, and the next turn succeeds; each of the four copies of the reason is read when it is the only one; a later copy without a reason never erases it; the row is exactly Grok's text, with no refusal sentence; with no text it reads "Grok ended this turn with an error." or, for a rate limit, "Grok usage limit reached."; a generic agent's failed turn reads its own error text, or names the agent when it sent none; a turn ended on purpose (a token limit) gets none; a late frame for an ended prompt neither reopens it nor captures later frames; a running recorded background command's row reads its description; a monitor stays a monitor after its output is read, and a kind a notice names still wins.
  • New: history replayed while loading is dropped, keeping only context usage (generic agent and the recorded Grok resume); a replayed background-task start, or an unmarked task notice, during a load writes nothing, and a later live completion still lands; Grok's task-result statuses (running, completed, failed, stopped, unknown) are checked on live frames (that table used to run only through adopted history); a request id a new agent process reuses gets its own row.
  • Removed with adoption (acp-history-adoption.ts, the translator's adopt option, and the "replayed task start is unverifiable" rewrite only adoption reached): the adoption tests (generic, multi-turn user history, two on the recorded Grok resume, replayed task-output results, and the replayed-task outcome cases).
  • Removed (their scenarios no longer exist): adopting history into a partly recorded chat (two tests), re-running an adoption into a partly recorded chat, a test that relied on another agent's id scheme, recovering request rows and the open turn after a restart, task-metadata recovery from the record after eviction (replaced by a test that the task keeps its row and placement), and the restart variants of the background-task tests (now run within one agent process).
  • Formatting and lint pass on changed files. One gated node typecheck on this head (5c4a89f0613, which also merges Add a shared timeline assembler for structured agent chats (not wired yet) #25064's latest branch so it merges cleanly with current main): no errors in this branch's code; its only 5 errors are a dependency (stream-json) missing from the shared, stale local install.
  • No dependency or lockfile change.
  • Manual app testing: not applicable; nothing is wired.

What I verified / didn't

Verified by tests through the real assembler and serialized record, using the existing scrubbed Grok recordings plus synthetic traffic. Each fix in this revision was removed in turn and its test failed without it; the placement test fails when the record of a finished tool's turn is removed. I also replayed third-party live Grok recordings (Grok 1.0.41 and 1.0.44, 15 scenarios, outside this PR) through the translator: the failed-prompt recording now shows Grok's reason in the failed turn, and the monitor recording keeps its task a monitor. I read the assembler's open-work accounting to confirm it counts a running tool at no fewer bytes than the translator does.

Not verified:

  • That our Grok build (1.0.46) echoes a prompt id Orca injects. Every recording in this PR had its prompt ids rewritten by the recording harness; the only live evidence is the third-party recordings above.
  • Real Grok resume, kill and rate-limit traffic; only one shape of Grok's retry notice (a retry it gave up on) has been seen, so other retry notices are ignored.
  • Windows, Linux and SSH by hand. (CI on the previous head, c1f0c7a with main merged at d173514, was green: every check that ran passed, including typecheck, the five unit test shards, both packaging jobs and cross-version wire compatibility; the rest were skipped by path filters.)

No agent CLI, Electron app or headless runtime was run. Passing tests show nothing regressed; they don't prove the design is right.

Review

@BrennanKB5

Agent skill upstream boundary

  • Not applicable; no upstream skill source is copied.

Notes

This remains a draft. Connecting it is #25225; the cancelled-tool display is #25181.

Checklist

  • Focused on protocol translation and supporting tests
  • Explained the problem, user-facing change, mechanism and alternatives
  • N/A visual proof with reason
  • Self-reviewed correctness, privacy and memory bounds
  • Cross-platform and execution-host boundaries considered
  • CI on this head (5c4a89f0613): green, including "static analysis and typecheck", all five unit-test shards and cross-version wire compatibility (two jobs that never got a runner were re-run)

Earlier main refresh (67e12c3)

Merged main at 4e64fa9940f5a4a9bfb7b021a3c405fd76643ca5; head 67e12c3046b36da60175d59b5902b4d1d3a0c947. Main supplies the landed transition and protocol client code, including the current Stop and steer handling. All 69 remaining translation/assembler patch sections are unchanged. Four conflicts take main verbatim, including its required ACP boundary-check lane and seven-file fixture. Local targeted verification: 297 passing tests across 34 explicit files; two boundary-check cases cannot resolve stream-json/stream-chain from the stale symlinked node_modules. Scoped format/oxlint passed on 66 code files. The lockfile and submission-position test exactly match merged main with exactly one used provider-handle import. Fresh CI is pending; no local typecheck was run.

Refresh after the shared assembler landed

Merged current main at 1de3aa405f71ea693b7138c40dac62e76883a10f; current head 9bd27f9d16b524ffa2dd83ee3c60a9fc1b9c4aba. All 36 translation-only sections and 4116 added/deleted feature lines are preserved exactly, with no copies of landed dependency changes. All 14 explicit suites / 97 tests and scoped format/oxlint on 29 code files passed. The root lockfile and entire submission-position test match merged main, with exactly one used provider-handle import. Main’s corrections to the release-checkout and cache-scan tests were inherited through the merge. Current-head CI is fully terminal and green; no local typecheck was run.

Current-head CI finished

At 9bd27f9d16b524ffa2dd83ee3c60a9fc1b9c4aba, all four current-head runs are complete: 14 checks passed and 18 were skipped. Static analysis/typecheck, all five unit shards, both packaging jobs, mobile and relay integration passed; the compatibility check was skipped by its path filter. The unchanged main Windows host-job test missed its startup marker on the first run and failed cleanup; one targeted Windows retry passed the test, packaging and smoke checks without source changes. Final verification passed.

…al timeline folder

Pure moves so a shared timeline assembler can use them: Codex's message
ordinal counter becomes ProviderTurnMessageOrdinals and Claude's turn-row
revision becomes the provider-neutral agent-journal turn-row revision. Only
names and import paths change.
…found again after a restart

- A sink transition is admitted whole or not at all; its steps run back to
  back at their turn in the journal's write queue, and each resolver reads
  the fold with every earlier write landed. A resolver may also say where the
  row belongs (turn scope, provider reference), and the writer always hears
  how the transition landed. A resolved lifecycle batch chooses its
  settlement mutations from the fold at execution.
- New optional row field providerItemRef: the provider's own reference for
  the item a row is, written only where the row's identity cannot spell it
  (Codex keys messages by their place in the turn and renumbers its item ids
  on resume). Set by the creating write, kept by revisions, indexed by the
  journal fold, never read by clients. A downgrade test shows an older host
  and client render such rows unchanged.
- Provider timeline identity schemes (shared legacy arm, Codex) and the join
  index that resolves a provider item to its row from memory or the fold:
  ordinals and request incarnations are read back from the rows, so a
  restart or an evicted entry finds the original row instead of placing a
  new one.
@brennanb2025
brennanb2025 force-pushed the brennanb2025/acp-d2-translation branch 2 times, most recently from 09ac32c to bdbd547 Compare October 4, 2026 05:47
… joins read from a replaced epoch

A fresh join index continued a turn's messages at the first free place, so a journal holding only a
later ordinal (an imported or removed earlier row) had its sequence back-filled. The place is now
one past the highest ordinal any row or echoed send holds there, read through a pure scheme reader.
The join caches also drop what they read when the journal's epoch is replaced.
…t its turn in the journal

Adapters translate their provider's dialect into a small grammar (turns,
items, streamed text, requests, context facts, session end/reset); one shared
assembler turns it into the journal rows every structured lane writes.

Each event is planned as one sink transition. Which row a write lands on,
whether a replay writes anything, and every change to what the assembler
knows (its ledger) are decided by the transition's resolvers at the event's
turn in the journal's write queue, against the fold as it stands then. A
forecast (the ledger plus admitted events still queued) only answers apply()
at once. So a refused event allocates nothing, a write the journal rejects
leaves no trace in memory, and a restart or evicted cache finds the same rows
again. Text and full snapshots of one provider item share one row and one
lifecycle; reset always flushes text and settles the old session from the
journal; the open-work budget is derived from what is actually open.

Codex migration contracts compare against the existing Codex translator,
including a restart mid-stream and a repeat that outlives the join cache.
…nd background work

A third review found two blockers with the earlier rounds' cause, a remembered interpretation
trusted after the journal moved on:

- A reused request id was judged by its earlier prompt's settled turn before asking which turn the
  new one lands in, so a real approval in a later turn was dropped. The target turn now decides:
  the old turn again is a replay; a different live turn opens the next prompt beside it.
- A text stream checked its row's turn only on its first write, and a turn's end released streams
  by the turn planning expected. Every write now checks the row, a turn's end stops the streams
  whose rows are in it, and turn status reads the journal first, so another writer's Stop wins.

Also: a message boundary drawn by an event the journal held as a replay no longer splits an
anonymous message; a send naming a turn not yet open waits for that turn; the budget charges a
stream's thread and turn strings and the turn caches are byte-bounded; the open turn ends when the
journal shows it settled; a turn's opener is read from the journal's row.

Background work is now Orca's existing background-task row instead of a tool call flagged
`outlivesTurn` (a flag remembered only in memory, so a restart failed the task). A turn's end
never settles that row, so it survives restarts; session end leaves one in flight unverifiable.
Three tests that opened a background tool call with `outlivesTurn` now open a background-task row
and keep their original expectations about which turn the row stays in.
… streams

A row kept for the anonymous stream that may continue it, and the marker that a stopped stream's
queued writes write nothing, lived in the live-stream map and were never removed when no stream
followed. They now live in their own bounded maps, so the live map holds open streams only.
@brennanb2025
brennanb2025 force-pushed the brennanb2025/acp-d2-translation branch from c897071 to d52d38b Compare October 4, 2026 07:17
The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.
Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.
Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.
- journal-store.ts imports: main #24576 dropped journalStoreLoadedFields; this branch types
  appendResolvedItem off JournalItemAppender, so neither it nor JournalResolvedItem is imported.
- event-sink-queue: keeps this branch's journalItems; journalStopDecidesTurn takes main #24864's
  two arguments (openedBy plumbing removed on main).
…ansitions' into brennanb2025/acp-c1-timeline-assembler
brennanb2025 added a commit that referenced this pull request Oct 5, 2026
No textual conflicts. D2 deletes acp-history-adoption.ts and the translator's adopt path, and
finishLoad no longer returns events; the reattach start now only ends the load window.
@brennanb2025

Copy link
Copy Markdown
Contributor Author

Removed saved-history adoption (5c4a89f0613).

What was removed: acp-history-adoption.ts, the translator's adopt option and adoption branches, the "replayed background-task start becomes unverifiable" rewrite that only adoption reached, and the adoption tests. finishLoad() now only ends the load and returns nothing. Everything an agent replays during session/load is dropped, except what it says about context usage; live frames are translated exactly as before.

Why: the common pattern discards replayed history on load, so Orca does the same and has no adoption path.

Verified: 19 ACP test files (174 tests) and 13 assembler test files (100 tests) pass, plus the other changed tests (33 files, 288 tests in total); one gated node typecheck found only this machine's stale-install stream-json errors; CI on this head is green. A fresh reviewer checked the removal diff and found no behaviour change outside adoption and load replay dropped except usage on every path. Its coverage note is addressed: Grok's task-result status table now runs on live frames, and an unmarked task notice during a load is tested as dropped. This head also merges #25064's latest branch, so it merges cleanly with current main again.

For #25225: its attaching.apply(attaching.translator.finishLoad(now())) needs to become attaching.translator.finishLoad(), and its "Adoption would plug in here" comment goes.

@brennanb2025

brennanb2025 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Merged current main after #25064 landed; head 9bd27f9d16b524ffa2dd83ee3c60a9fc1b9c4aba, with only translation remaining in the diff. All 97 targeted tests and scoped format/oxlint passed; main lock/import remain exact. CI is fully terminal and green: 14 passed / 18 skipped, including typecheck, all five unit shards and both packages; one targeted retry cleared the unchanged Windows host-job test, with no source edit.

brennanb2025 added a commit that referenced this pull request Oct 6, 2026
@brennanb2025
brennanb2025 marked this pull request as ready for review October 6, 2026 07:53
@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

📝 Walkthrough

Walkthrough

Adds ACP-to-provider timeline translation for prompts, turns, session updates, tool calls, context usage, requests, and background tasks. Adds Grok-specific notification and request mapping. Adds bounded snapshots for tool and task updates, JSONL recordings, a fixture test rig, and tests for translation, replay, routing, and fixture privacy. Also adds a shared journal-body helper for background-task rows and uses it in the Claude background-task journal.

Priority: ⬇️ Low

Merge Risk: 🔵 Low · up to 9bd27

The ACP translation is not yet connected to any user-facing path. The remaining issue is narrow: when the active Grok model has no window metadata, context-usage rows may show another model's window. This can be handled as a follow-up and should not block the merge.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 15.56% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 45 functions across 29 files. (7 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the primary change: translating ACP traffic into shared timeline events.
Description check ✅ Passed The description is comprehensive and covers the change, rationale, testing, visual proof, review notes, and checklist. It explains that there is no separate linked issue and references related issues …
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 15.56% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 45 functions across 29 files. (7 skipped: 7 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 2
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
src/main/acp/acp-timeline-fixture.test-support.ts (1)

111-116: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Match the pending row by request identity, not by kind.

The rig resolves the first pending row whose kind matches the reply's kind. Suppose one fixture has two pending approvals at the same time. The rig can then resolve the wrong row, and the tests still pass. The current fixtures answer each request before the next one opens, so no test fails today. The request id is already stored in requests at Line 102. Match item.itemId against the row that this request opened, so the rig resolves the row the provider actually asked about.


ℹ️ Review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: b28b989a-f1d3-4cba-8d90-b6c724846a3d
📥 Commits

Reviewing files that changed from the base of the PR and between 1de3aa4 and 9bd27f9.

📒 Files selected for processing (36)
  • src/main/acp/acp-background-task-timeline.ts
  • src/main/acp/acp-context-usage.ts
  • src/main/acp/acp-dialects/acp-dialect.ts
  • src/main/acp/acp-dialects/grok-background-tasks.ts
  • src/main/acp/acp-dialects/grok-dialect.ts
  • src/main/acp/acp-dialects/grok-requests.ts
  • src/main/acp/acp-fixture-privacy.test.ts
  • src/main/acp/acp-prompt-turns.ts
  • src/main/acp/acp-session-update.ts
  • src/main/acp/acp-timeline-background-task-evidence.test.ts
  • src/main/acp/acp-timeline-background-task-results.test.ts
  • src/main/acp/acp-timeline-background-task-words.test.ts
  • src/main/acp/acp-timeline-background-tasks.test.ts
  • src/main/acp/acp-timeline-dialect.test.ts
  • src/main/acp/acp-timeline-fixture.test-support.ts
  • src/main/acp/acp-timeline-generic.test.ts
  • src/main/acp/acp-timeline-joins.test.ts
  • src/main/acp/acp-timeline-open-enums.test.ts
  • src/main/acp/acp-timeline-recordings.test.ts
  • src/main/acp/acp-timeline-recovery.test.ts
  • src/main/acp/acp-timeline-requests.ts
  • src/main/acp/acp-timeline-routing.test.ts
  • src/main/acp/acp-timeline-translator.ts
  • src/main/acp/acp-timeline-turn-failures.test.ts
  • src/main/acp/acp-tool-timeline.ts
  • src/main/acp/acp-turn-failures.ts
  • src/main/acp/acp-turn-messages.ts
  • src/main/acp/fixtures/s1-basic.jsonl
  • src/main/acp/fixtures/s1-full-notifications.jsonl
  • src/main/acp/fixtures/s2-permission.jsonl
  • src/main/acp/fixtures/s3-cancel.jsonl
  • src/main/acp/fixtures/s4-resume.jsonl
  • src/main/acp/fixtures/s5-plan-approved.jsonl
  • src/main/acp/fixtures/s6-background.jsonl
  • src/main/claude/claude-background-task-row-journal.ts
  • src/shared/native-chat-background-task-row.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 5 remain after this review.

Comment on lines +60 to +70
function contextWindow(models: unknown): number | undefined {
const parsed = modelsSchema.safeParse(models)
if (!parsed.success) {
return undefined
}
const { currentModelId, availableModels = [] } = parsed.data
return (
availableModels.find((model) => model.modelId === currentModelId)?._meta?.totalContextTokens ??
availableModels.find((model) => model._meta)?._meta?.totalContextTokens
)
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Use the fallback window only when Grok names no current model.

contextWindow falls back to the first model that has _meta.totalContextTokens, and it does this even when currentModelId is set. Suppose currentModelId names a model without _meta, or a model absent from availableModels. The function then reports the window of a different model. AcpContextTimeline.update stores that value. Every later context.usage event then carries a wrong window.tokens, so the occupancy ratio is wrong for the rest of the session. Grok's x.ai/models/update notifications can change models mid-session, so this state can happen.

When currentModelId is present, return undefined if that model has no window. An unknown window is safer than the window of another model. Keep the first-model fallback only for the case where Grok does not name a current model. If you keep the fallback on purpose, add a comment that explains why it is safe.

🐛 Proposed fix
   const { currentModelId, availableModels = [] } = parsed.data
-  return (
-    availableModels.find((model) => model.modelId === currentModelId)?._meta?.totalContextTokens ??
-    availableModels.find((model) => model._meta)?._meta?.totalContextTokens
-  )
+  if (currentModelId !== undefined) {
+    return availableModels.find((model) => model.modelId === currentModelId)?._meta
+      ?.totalContextTokens
+  }
+  return availableModels.find((model) => model._meta)?._meta?.totalContextTokens
 }

Based on learnings: "avoid silently swallowing errors, falling back to defaults ... If a fallback is truly intentional and safe, add a comment explaining why it's safe; otherwise flag the silent fallback as a bug."

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
function contextWindow(models: unknown): number | undefined {
const parsed = modelsSchema.safeParse(models)
if (!parsed.success) {
return undefined
}
const { currentModelId, availableModels = [] } = parsed.data
return (
availableModels.find((model) => model.modelId === currentModelId)?._meta?.totalContextTokens ??
availableModels.find((model) => model._meta)?._meta?.totalContextTokens
)
}
function contextWindow(models: unknown): number | undefined {
const parsed = modelsSchema.safeParse(models)
if (!parsed.success) {
return undefined
}
const { currentModelId, availableModels = [] } = parsed.data
if (currentModelId !== undefined) {
return availableModels.find((model) => model.modelId === currentModelId)?._meta
?.totalContextTokens
}
return availableModels.find((model) => model._meta)?._meta?.totalContextTokens
}

Source: Learnings

@brennanb2025
brennanb2025 merged commit 670c59d into main Oct 6, 2026
77 of 79 checks passed
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
D2's new commits are main merges and its ratchet change requiring src/main/acp, which D3 already has.
The D1 protocol files take D3's side (main's D1 plus ACP-ALIGN); D2's copies equal main's.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
Main now carries ACP-ALIGN #25810 (squash 0bcd49c) and D2 #25090 (squash 670c59d); both squashes
equal the branch heads D3 already merged (20f7f8a, 9bd27f9), so the merge was computed against
main's tree with those heads as extra parents (scratch commit, never pushed) and committed as an
ordinary two-parent merge of origin/main. No conflicts.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant