feat(agent-status): publish the main agent's own state beside the combined row state - #22452
Conversation
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — the complete diff (51 files, one commit) that adds the lead agent's own state beside the combined row state.
- A new
leadfact on every status row.AgentLeadStatus(state, optional provideroutcome,stateStartedAt) rides onAgentStatusPayload.lead, admitted bynormalizeAgentStatusPayloadon the relay wire, over IPC and from disk; a malformed value drops the field and keeps the row. - Child-only state is now derived, never stored. The persisted
claudeLeadBoundaryChildOnlyflag and structuredfromChildWorkare gone, replaced byisAgentStatusHeldOpenByChildWorkoverleadplus the row's child evidence; hydration maps a legacy flag onto an absentlead. - Every producer publishes it. Claude stores a verdict instead of an interrupt flag via
setClaudeLeadTurnState(now the sole writer), structured sessions addturnOutcometo the summary, Grok moves its stop-keeps-working rule onto the shared fold, and Codex publishes from its root record while deliberately keeping its own combine. - Parity and restart coverage. A six-story table drives all four lanes against the fold, alongside restart round-trip, OSC carry, inferred-interrupt, IPC/store and mobile-projection tests.
ℹ️ Persisting claudeRunningNonAgentTask is consistent with the design
The author flagged this as a persisted-shape change; I traced it and agree with the call. lead cannot express "a shell is running beside the child agents" (lead.state === 'done' holds either way), yet hydration needs that fact to decide whether to seed a settled lead, so recording the boolean is the only non-divergent option once the derived flag stops being written. The derived boundary treats a missing/undefined value as "no shell" (row.claudeRunningNonAgentTask !== true), which is exactly what the old persisted flag's absence encoded, so legacy rows hydrate to the same child-only decision. No change requested.
Technical details
# Persisted shell fact
## Affected sites
- `src/main/agent-hooks/server/server-types.ts` — `claudeRunningNonAgentTask` removed from the `Omit` list, so it now serializes.
- `src/main/agent-hooks/server/server-persistence-validation.ts` — reads the boolean back onto the hydrated entry and maps a legacy child-only flag onto an absent `lead`.
- `src/main/agent-hooks/server/server-hydration.ts` — calls `isClaudeLeadBoundaryHeldByChildrenOnly(entry)` to seed the lead record.
## Required outcome
- None. Recorded so the persisted-shape decision is on the record; the field is a boolean, is not a secret, and `lead`'s admission path already governs the payload it rides in.
## Note
- After restart the in-memory `state.claudeRunningNonAgentTaskPaneKeys` set is still not seeded from the persisted boolean, but every consumer either already consulted the row itself (`reapRestoredClaudeSubagentsWithoutLiveAgent`'s `payload.state !== 'done'`, `inferInterrupt`'s `restoredUnconfirmed` refusal) or reproduces the old absent-flag semantics, so the disagreement cannot leave a row stuck non-terminal.DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (4)
🚧 Files skipped from review as they are similar to previous changes (1)
Included review availability: Your plan provides up to 10 included reviews per hour; 6 remain after this review. 📝 WalkthroughWalkthroughAgent status rows now carry a Priority: ⬇️ Low Merge Risk: 🟡 Moderate · up to The new main-agent status can lose a Grok turn verdict or show a timestamp from an earlier Claude session. Correct these status facts before merging. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 2📝 Generate docstrings 💡
🛠️ Fix failing CI checks 💡
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 3
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Advanced
Run ID: d9a410b8-685f-4a1a-9b9b-058b50632bb7
📒 Files selected for processing (51)
docs/reference/agent-status-store.mdsrc/main/agent-hooks/server-ingest-structured-status.test.tssrc/main/agent-hooks/server-ingest-terminal-status.test.tssrc/main/agent-hooks/server-interrupt-inference-validation.test.tssrc/main/agent-hooks/server-last-status-lead-boundary.test.tssrc/main/agent-hooks/server-last-status-lead-fact.test.tssrc/main/agent-hooks/server-last-status-write.test.tssrc/main/agent-hooks/server/server-claude-status-rules.tssrc/main/agent-hooks/server/server-hydration.tssrc/main/agent-hooks/server/server-ingest-structured.tssrc/main/agent-hooks/server/server-ingest-terminal.tssrc/main/agent-hooks/server/server-persistence-validation.tssrc/main/agent-hooks/server/server-persistence.tssrc/main/agent-hooks/server/server-status-inference.tssrc/main/agent-hooks/server/server-status-update.tssrc/main/agent-hooks/server/server-types.tssrc/main/agent-hooks/terminal-handle-row-identity.test.tssrc/main/native-chat/agent-session-wire/structured-agent-session-status-feed.tssrc/main/runtime/runtime-agent-row-lead-projection.test.tssrc/renderer/src/components/native-chat/StructuredAgentSessionStatusBridge.test.tsxsrc/renderer/src/components/native-chat/StructuredAgentSessionStatusBridge.tsxsrc/renderer/src/hooks/ipc-events/normalize-agent-status-event.tssrc/renderer/src/store/slices/agent-status-lead-fact.test.tssrc/renderer/src/store/slices/agent-status-live-entry-builder.tssrc/renderer/src/store/slices/agent-status-live-entry-lead.tssrc/shared/agent-hook-listener-claude-subagents.test.tssrc/shared/agent-hook-listener-claude-turn-state.test.tssrc/shared/agent-hook-listener-codex-lead.test.tssrc/shared/agent-hook-listener-extraction-characterization.test.tssrc/shared/agent-hook-listener/lead-turn-state.tssrc/shared/agent-hook-listener/listener-state.tssrc/shared/agent-hook-listener/providers/claude-events.tssrc/shared/agent-hook-listener/providers/claude-lifecycle-events.tssrc/shared/agent-hook-listener/providers/claude-roster-state.tssrc/shared/agent-hook-listener/providers/claude-status-build.tssrc/shared/agent-hook-listener/providers/codex-events.tssrc/shared/agent-hook-listener/providers/codex-state.tssrc/shared/agent-hook-listener/providers/grok-events.tssrc/shared/agent-lead-status-fold.test.tssrc/shared/agent-lead-status-fold.tssrc/shared/agent-lead-status-parity.test.tssrc/shared/agent-lead-status.tssrc/shared/agent-session-journal-types.tssrc/shared/agent-session-wire.tssrc/shared/agent-status-types.test.tssrc/shared/agent-status-types.tssrc/shared/agent-turn-outcome.tssrc/shared/structured-agent-session-agent-status.test.tssrc/shared/structured-agent-session-agent-status.tssrc/shared/structured-agent-session-projection.test.tssrc/shared/structured-agent-session-projection.ts
Included review availability: Your plan provides up to 10 included reviews per hour; 1 remains after this review.
| state.claudeRunningNonAgentTaskPaneKeys.delete(paneKey) | ||
| state.claudeActiveSessionCronPaneKeys.delete(paneKey) | ||
| state.claudeLeadStateByPaneKey.set(paneKey, { state: 'done' }) | ||
| setClaudeLeadTurnState(state, paneKey, { state: 'done' }) |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Restart the lead clock at a Claude SessionStart boundary.
setClaudeLeadTurnState goes through continueAgentLeadStatus. That function keeps previous.stateStartedAt when the previous state is also done. Suppose the previous lead is already done, for example after /clear, a resume, or a relaunch on an idle pane. The new session then publishes lead.stateStartedAt from the old session's last turn. The Grok lane treats a new process as a new lead and drops its lead cache on session_start. The Claude lane should also restart the lead clock here.
Proposed fix
- setClaudeLeadTurnState(state, paneKey, { state: 'done' })
+ // A new process owns the pane: its lead clock starts now, not at the old session's last Stop.
+ setClaudeLeadTurnState(state, paneKey, { state: 'done', stateStartedAt: Date.now() })📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| setClaudeLeadTurnState(state, paneKey, { state: 'done' }) | |
| // A new process owns the pane: its lead clock starts now, not at the old session's last Stop. | |
| setClaudeLeadTurnState(state, paneKey, { state: 'done', stateStartedAt: Date.now() }) |
| // Why: the provider's verdict, never inferred — a plain `stop` stays absent. | ||
| const outcome = isGrokEvent(eventName, 'stop_cancelled') | ||
| ? ('cancellation' as const) | ||
| : isGrokEvent(eventName, 'stop_failure') | ||
| ? ('failure' as const) | ||
| : undefined | ||
| // Only a plain end-of-turn `stop` reports what it left running; a cancel, a failure and a | ||
| // session boundary settle the pane whatever the inventory says, as they always have. | ||
| const resolution = foldAgentLeadStatus({ | ||
| leadState, | ||
| interrupted: outcome === 'cancellation', | ||
| childWorkLiveness: | ||
| isGrokEvent(eventName, 'stop') && !sessionBoundary | ||
| ? grokChildWorkLivenessAfterStop(hookPayload) | ||
| : null | ||
| }) | ||
| const stateName = resolution.stateName | ||
| const lead = continueAgentLeadStatus( | ||
| state.grokLeadStatusByPaneKey.get(paneKey), | ||
| { state: leadState, outcome }, | ||
| Date.now() | ||
| ) | ||
| state.grokLeadStatusByPaneKey.set(paneKey, lead) |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Keep the Grok turn verdict when an idle_prompt notification restates done.
outcome comes only from stop_cancelled or stop_failure. A later notification/idle_prompt, or a repeated done from session_end, also maps to leadState = 'done'. That event passes outcome: undefined, so continueAgentLeadStatus removes the stored verdict. The published lead then loses cancellation or failure while the lead is still in the same finished turn. This breaks the AgentLeadStatus.outcome contract that "a new turn clears it". The Claude lane keeps the cached verdict on such refreshes. Carry the previous verdict forward when the lead stays done and the event is not a turn end.
Proposed fix
- const outcome = isGrokEvent(eventName, 'stop_cancelled')
- ? ('cancellation' as const)
- : isGrokEvent(eventName, 'stop_failure')
- ? ('failure' as const)
- : undefined
+ const previousLead = state.grokLeadStatusByPaneKey.get(paneKey)
+ const outcome = isGrokEvent(eventName, 'stop_cancelled')
+ ? ('cancellation' as const)
+ : isGrokEvent(eventName, 'stop_failure')
+ ? ('failure' as const)
+ : !isTurnEnd && leadState === 'done' && previousLead?.state === 'done'
+ ? previousLead.outcome
+ : undefined
...
- const lead = continueAgentLeadStatus(
- state.grokLeadStatusByPaneKey.get(paneKey),
- { state: leadState, outcome },
- Date.now()
- )
+ const lead = continueAgentLeadStatus(previousLead, { state: leadState, outcome }, Date.now())If you apply this change, check whether interrupted and the fold's interrupted input should also stay set on these refreshes. They currently read the same outcome value.
📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| // Why: the provider's verdict, never inferred — a plain `stop` stays absent. | |
| const outcome = isGrokEvent(eventName, 'stop_cancelled') | |
| ? ('cancellation' as const) | |
| : isGrokEvent(eventName, 'stop_failure') | |
| ? ('failure' as const) | |
| : undefined | |
| // Only a plain end-of-turn `stop` reports what it left running; a cancel, a failure and a | |
| // session boundary settle the pane whatever the inventory says, as they always have. | |
| const resolution = foldAgentLeadStatus({ | |
| leadState, | |
| interrupted: outcome === 'cancellation', | |
| childWorkLiveness: | |
| isGrokEvent(eventName, 'stop') && !sessionBoundary | |
| ? grokChildWorkLivenessAfterStop(hookPayload) | |
| : null | |
| }) | |
| const stateName = resolution.stateName | |
| const lead = continueAgentLeadStatus( | |
| state.grokLeadStatusByPaneKey.get(paneKey), | |
| { state: leadState, outcome }, | |
| Date.now() | |
| ) | |
| state.grokLeadStatusByPaneKey.set(paneKey, lead) | |
| // Why: the provider's verdict, never inferred — a plain `stop` stays absent. | |
| const previousLead = state.grokLeadStatusByPaneKey.get(paneKey) | |
| const outcome = isGrokEvent(eventName, 'stop_cancelled') | |
| ? ('cancellation' as const) | |
| : isGrokEvent(eventName, 'stop_failure') | |
| ? ('failure' as const) | |
| : !isTurnEnd && leadState === 'done' && previousLead?.state === 'done' | |
| ? previousLead.outcome | |
| : undefined | |
| // Only a plain end-of-turn `stop` reports what it left running; a cancel, a failure and a | |
| // session boundary settle the pane whatever the inventory says, as they always have. | |
| const resolution = foldAgentLeadStatus({ | |
| leadState, | |
| interrupted: outcome === 'cancellation', | |
| childWorkLiveness: | |
| isGrokEvent(eventName, 'stop') && !sessionBoundary | |
| ? grokChildWorkLivenessAfterStop(hookPayload) | |
| : null | |
| }) | |
| const stateName = resolution.stateName | |
| const lead = continueAgentLeadStatus(previousLead, { state: leadState, outcome }, Date.now()) | |
| state.grokLeadStatusByPaneKey.set(paneKey, lead) |
d78bbab to
3798c11
Compare
…bined row state
Every status producer folded the main agent's state together with live child
work into one `state`, so a lead that had finished while a subagent still ran
read `working` and its own state was lost. The row now also carries
`lead: { state, outcome?, stateStartedAt }`, admitted by the one payload
normalizer on the relay wire, IPC and disk, and published from the Claude hook
lane, the structured host ingest and renderer bridge, Grok (now on the shared
fold) and Codex (own combine kept). The persisted child-only boundary flag is
derived from `lead` plus child evidence and no longer written; old rows map
onto `lead` at hydrate. Combined `state` and `workingMode` are unchanged for
every reader; a cross-lane parity table pins that, with the cancelled-turn
watch-loop story recorded as a known divergence.
…of a Claude lead cancellation Current Claude Code sends no hook at all on a cancel and no is_interrupt on Stop, so the cancellation enters the lead record from the server's inferred interrupt and rides into the next real Stop; is_interrupt on a turn boundary stays as the secondary source for builds that send it. Comments, the store reference and the parity table say so; no suppression changes.
… persisted shell fact The old sentence said a hydrated row no longer carries the shell fact, which is the opposite of the mechanism: claudeRunningNonAgentTask is persisted precisely so hydration can read it, and only a pre-lead row lacks it — reading as shell-free, the same assertion its legacy flag made at write time.
3798c11 to
2b24918
Compare
…n agent, and the row verdict docs name its inferred source
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — the two commits pushed since the prior pullfrog review (0d9d5e0f), both small behavior fixes with supporting tests and doc/comment corrections.
- Kept an already settled main agent intact through an inferred interrupt.
server-status-inference.tsnow reusespayload.mainAgentwhen it is alreadydoneinstead of synthesizing a freshcancellationverdict, so a watch loop holding the row open cannot rewrite the settled main agent's clock or verdict. The branch is only reachable for Grok/custom agents: for Claude and Codex,state === 'working'alongsidemainAgent.state === 'done'implies child evidence, and the non-idle-subagent guard returns first. - A child-induced wait now publishes the main agent state it displaced.
claudeMainAgentStatusForPayloadselectsstateBeforeWaitwhen the record carries awaitingAgentId, so the row's combinedwaitingstate no longer masquerades as the main agent's own state;clearClaudeAnsweredQuestionWait'smainAgentbecomes optional to match. - Verdict provenance documented.
agent-turn-outcome.ts,main-agent-status.ts,main-agent-turn-state.ts,claude-events.ts,grok-events.ts,agent-session-journal-types.tsanddocs/reference/agent-status-store.mdnow name the inferred interrupt as a recordedoutcomesource. - Coverage updated. Exact assertions added for the settled-behind-watch-loop interrupt and for the child-vs-own wait distinction.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
…agent fact Restart seeds the Claude main agent from the row's saved mainAgent whenever it settled and no shell held the row, instead of re-deriving a child-only shape. OSC cannot settle or repaint a row child agents hold open, including a row waiting on a child's permission prompt. A sticky child permission prompt still records the main agent's own progress, and OSC repaints and inferred answers keep the shell fact beside the main agent they preserve.
…k with that turn A Claude SessionStart restarts the main agent's clock instead of inheriting the previous session's last Stop. A Grok idle prompt or session end, and a late Codex root Stop after an inferred cancel, restate the same finished turn, so they keep its recorded verdict; only a new turn clears it.
… the saved main agent fact
… so restart does not settle it
…ell fact and is pushed
There was a problem hiding this comment.
ℹ️ No critical issues — one actionable gap in the new Grok verdict fallback, inline.
Reviewed changes — the four commits pushed since the prior pullfrog review (8bbe6d4).
- Hydration now seeds the Claude main-agent record from the saved
mainAgentfact.seedClaudeLeadTurnFromPersistedStatusdropped itschildOnlyBoundaryoption and seeds any restored row whose persistedmainAgentisdonewith no shell holding it, so a restored row's children's drain can still settle it todone. The live-child, shell/cron and cancelled-verdict holds all still prevent a false settle. - A held child permission prompt now records the main agent's own progress.
withHeldChildWaitMainAgentrewrites the held row'smainAgentwhen the main agent moves behind a child's sticky prompt, so a staledonecan no longer let a later drain settle a working main agent. - OSC refuses any Claude row child agents hold open. The
server-ingest-terminalguard now reads the sharedisAgentStatusHeldOpenByChildWorkplus a working subagent — covering waiting child-prompt rows and dropping the old shell disqualifier — and carries the shell fact forward beside the preservedmainAgent. - Turn verdicts survive restatement. Codex keeps a recorded verdict across a late root Stop (
codexOutcomeRestatedByStop), Grok preserves it across idle/session-end restatements, and ClaudeSessionStartnow stamps its own main-agent clock. - Coverage. Two new server suites (child-held rows across restart, main-agent turn verdicts) and updated boundary/legacy-flag tests, with exact assertions.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
| // Why: an idle prompt or session end restates the same finished turn, so its verdict stands. | ||
| const mainAgentOutcome = | ||
| outcome ?? | ||
| (!isTurnEnd && leadState === 'done' && previousMainAgent?.state === 'done' |
There was a problem hiding this comment.
A trailing stop reason=shutdown (or channel_closed) arrives as a turn end, so !isTurnEnd skips the fallback and continueMainAgentStatus drops the verdict on the done→done restatement — even though isGrokSessionBoundary (line 125) classifies that same event as the session ending, not a new turn. The captured fixture grok-background-completion-hooks.jsonl:13-14 emits session_end and then a trailing stop reason=shutdown, so a cancelled turn's cancellation is preserved by the session_end and erased by the stop right after it. Including session-boundary turn ends keeps the recorded verdict.
Technical details
# Grok drops a recorded verdict on the trailing shutdown stop
## Affected sites
- `src/shared/agent-hook-listener/providers/grok-events.ts:224` — `!isTurnEnd` excludes `stop reason=shutdown`, which `isGrokSessionBoundary` (`:125`) already treats as the same finished session.
- `src/shared/agent-hook-listener/providers/grok-events.ts:222-226` — the fallback restores `previousMainAgent.outcome` only for non-turn-end events.
- `src/shared/agent-lead-status-fold.ts:63` — `continueMainAgentStatus` emits `outcome` only from `next.outcome`, so a done→done call with `outcome: undefined` clears it.
- `src/shared/__fixtures__/grok-background-completion-hooks.jsonl:13-14` — real `session_end` followed by a trailing `stop reason=shutdown` with no `promptId` (`grokTurnEndApplies` returns true).
## Required outcome
- A session-boundary turn end that restates a finished turn keeps that turn's recorded `outcome`, matching the comment's "an idle prompt or session end restates the same finished turn".
## Suggested approach (optional)
- Widen the guard to `(!isTurnEnd || sessionBoundary) && leadState === 'done' && previousMainAgent?.state === 'done'`, using the `sessionBoundary` value already computed at `:165`. This also covers a `stop_failure` followed by a shutdown stop.
## Open questions for the human (optional)
- Confirm a `stop reason=shutdown` is not intended to be a fresh-turn boundary that deliberately clears the prior verdict.There was a problem hiding this comment.
✅ No new issues in the latest commits.
Reviewed changes — the three commits pushed since the prior pullfrog review (fc0fb44).
- A held child permission row is now pushed when its
mainAgentchanges.server-status-update.tsemits the held row to enriched subscribers whenheld.payloaddiffers, so pushed readers see the main-agent update a snapshot reader already has; no push when only the unpublished shell fact changed. - A held child permission row now keeps the shell fact.
withHeldChildWaitMainAgentcarriesnext.claudeRunningNonAgentTask ?? previous.claudeRunningNonAgentTaskand returns early only when neither themainAgentnor the shell fact moved, so restart no longer seeds a shell-held held row. - Docs and tests. Both behaviors are pinned by exact new tests (one push on a main-agent change, none on an unchanged main agent; a shell-held held row stays unsettled across restart).
The inline concern from the prior review (Grok's trailing stop reason=shutdown erasing a restated verdict) is unchanged by these commits and remains open.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
…the one row-build point Every non-hook rewrite (terminal-title repaint, inferred answer, held child permission) had to re-carry the shell fact beside `mainAgent`, and each one that forgot let a restart settle a row while a shell still ran. The row builder now pairs the fact once: a listener event restates it, any other write keeps it only while `mainAgent` is unchanged. Restart seeds a settled main agent only when the row says no shell ran, and legacy child-only rows map to that explicitly. A held child permission now also accepts the main agent event's background evidence, as it already accepts its `mainAgent`, so the child's drain no longer settles a row a shell still holds. The renderer keeps a previous `mainAgent` only for writers that never carry one, so a hook row without it matches the host snapshot.
…row silent about its shell
…n agent, and restart seeds only on an explicit no-shell
There was a problem hiding this comment.
ℹ️ No new issues in the commits since the last review. This delta is a refactor that centralizes the Claude shell-fact pairing, tightens restart seeding to require an explicit no-shell, and narrows the renderer's
mainAgentreuse to blind writers — all documented and pinned by exact tests. This is not an approval because the earlier Grokgrok-events.tsnote is untouched by these commits and its thread remains open.
Reviewed changes — the three commits pushed since the prior pullfrog review (5ffe00f).
- Centralized the Claude shell-fact pairing.
attachClaudeChildOnlyBoundary/invalidateClaudeChildOnlyBoundaryare replaced bypairedClaudeNonAgentWork, which a listener event restates and any other write carries only whilemainAgentis unchanged. The per-write carries in the OSC ingest and inferred-answer paths are removed because the row builder now pairs once. - Tightened restart seeding.
seedClaudeLeadTurnFromPersistedStatusseeds a settled main agent only when the row statesclaudeRunningNonAgentTask === false; hydrate maps a legacy child-only row to that explicitly, while a row silent about the shell stays unseeded. - Held child permission takes the event's background evidence. The held branch now calls
onAcceptedfor a child-owned prompt, so a shell reported on the same event keeps holding the row open. - Renderer reuse narrowed to blind writers. Only
osc/title/launch/processorigins reuse a previousmainAgenton an unchanged state; a hook row without one drops it to match the host snapshot. - Coverage. Exact new tests for the shell-held held row, the silent-shell non-seed, the inferred-answer restart settle, and the renderer drop.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
Review status: ready for review at
|
|
One more independent review at |
Merging main brought the closure to 4220 against a pin of 4218. The two extra modules are `src/shared/agent-turn-outcome.ts` and `src/shared/main-agent-status.ts` from #22452, which the status projection this route already reaches import. That change was src/shared-only, so the mobile job never ran on it -- the same way the structured tool line slipped past, as the ledger above already records. Repinned here because this PR's file set is what next made the job run, not because this PR reaches either module. Verified: of the 28 source files this branch changes, none appear anywhere in the route's 4220-module closure.
…tatus modules #22452 changed only src/shared, so its CI never ran the page-closure suite; main now measures 4220 modules (1034 local) against a pin of 4218. This branch adds nothing to the closure: its own count matches main's.
… modules #22452 added src/shared/main-agent-status.ts and src/shared/agent-turn-outcome.ts, which agent-status-types.ts imports, so the session route's closure grew by two local modules (4218 -> 4220). That change was src/shared-only, so its own CI never ran this suite; main has been at 4220 since, and any PR that fires the mobile web app job fails on the stale pin. Measured on 4064653 and on origin/main 3ea15dd.
…reebuff agent) (#22601) This reverts commit 0677271. #18790 was merged as one squash commit that carried two unrelated changes: a process-incarnation fallback for reaping leaked orchestration worker terminals, and an unannounced "Freebuff" third-party agent (catalog entry, icon, locale strings, README rows). The Freebuff agent was never meant to ship, so the whole PR is reverted; the reap fix should be re-submitted on its own. Until that re-land, a worker whose durable terminal handle goes stale is again reported missing on release/stop instead of being re-found through its process incarnation, so its terminal can leak on Remote Server. The mobile session page closure pin moves 4218 -> 4219: the revert drops the freebuff icon #22119 pinned (-1), and #22452 had already added two src/shared modules without re-pinning (+2).
…and stop double taps (#22392) * fix(mobile): keep the streamed browser pane flipping on slow phones, and stop double taps Frame pacing. The pane decodes each frame on a hidden layer and flips to it on onLoad. While one frame decoded, every newer frame re-pointed that same hidden layer, which cancels the in-flight load. On a phone that decodes a frame slower than frames arrive (~10/s during page loads, menus, spinners), onLoad never fired for any of them and the pane sat on an old frame until the page went still. A decoding layer is now never re-pointed: only the newest frame is held, and it takes the layer once the decode settles. A 1.5s watchdog frees a layer whose decode never reports, and a frame the hidden layer already holds (a blinking caret alternating two frames) flips at once, since an unchanged source reloads nothing. Double taps. When browser.mouseClick failed, the pane replayed the tap as move/down/up. On a timeout the click is still queued on the host, so the replay landed a second tap on whatever the first one opened. The replay now runs only when the click definitely did not reach the host. * test(mobile): re-record the corpus without the timed-out tap replay The corpus certified the move/down/up replay after a transport-rejected browser.mouseClick, which the commit before removes. Scoped like #22179: baseline bumped by editing that one line, then --record. 788 files. Every changed line classified: - `baseline`: 787 files (786 goldens + pilot-scenarios.json), nothing else. - matrix-browser.pointer-click-browser.mouseclick-1.json: the transport-rejection and transport-rejection-no-message partitions of browser-pointer-click-fallback now send only browser.mouseClick#1. The refused partitions still replay, unchanged. * test(mobile): move the session closure pin past #22452's main-agent-status modules #22452 changed only src/shared, so its CI never ran the page-closure suite; main now measures 4220 modules (1034 local) against a pin of 4218. This branch adds nothing to the closure: its own count matches main's. * docs(mobile): say a re-pointed decoding layer loses its onLoad, as measured on Android * refactor(mobile): give the streamed browser pane's double buffer one owner The frame pacing, decode watchdog, layer flip and reset were spread over three hooks and a helper module, wired back through the stream hook and the pane. They now live in one plain pacer (browser-frame-pacer.ts) with one timer, and the pane binds each layer's View/Image straight to it. Behaviour fixed on the way, each with a failing test first: - A slow last frame with nothing newer queued was abandoned by the watchdog and never shown. The decode deadline now only applies when a newer frame waits. - Any pane re-render re-pointed both layers at the newest frame behind the pacer's back, so a blinking caret froze. The Image source prop is now only the mount-time frame; every later source write is the pacer's. - A frame that failed to decode left its layer marked as holding it, so an identical frame flipped to an undecoded layer. Giving up on a decode now clears the layer's source. - A native onLoad for a source the layer has since moved off could flip early. The flip now checks nativeEvent.source.uri, which Android and iOS Fabric both report as the raw source string; RN Web's own load event has none, so the web flips only through its decode probe. The session closure pin drops by the two modules this removes. * refactor(mobile): send the tap's mouseClick directly instead of through a flag The delivery-unknown check was a mutable flag set inside the request callback. The click now calls browser.mouseClick itself in a try/catch: a delivered click returns, a delivery-unknown failure returns without replaying, and a refusal or null result still replays as move/down/up. Same wire traffic; the corpus certifies it unchanged. * fix(mobile): never cut a streamed frame's decode short The 1.5 s decode watchdog abandoned a slow decode whenever a newer frame was queued and re-pointed its layer. On an Android emulator under load that is the original freeze again: noise frames decode in 2-10 s, every abandon starts a decode the next abandon cuts, Fresco reports the superseded loads (30 stale onLoads in one run) and the pane showed 11 of 41 applied frames. Without it, the same run flips every applied frame (16/16, no stale load), and a slow last frame is shown in every cycle. Nothing else needs it: with the layer never re-pointed mid-decode, native answers every load with onLoad or onError, and the web probe's decode() always settles. The pacer keeps one timer, for the interval. * fix(mobile): track what each frame layer's Image holds, so no write goes unanswered The pacer cleared a layer's source when it gave up on a decode (a reset mid-decode, or a failed decode) while the native Image still held it. The next identical frame was then written again, which is a native no-op on Android (ReactImageView.setSource returns on equal sources) and iOS (ImageShadowNode skips equal requests): no onLoad, no onError, and the pane stayed frozen until the stream restarted. Returning from the background to an unchanged page is enough to trigger it. Each layer now records the source its Image holds and whether that source has answered (loading, ready, failed). A layer is written only when it is not loading and only with a different source, so every write gets exactly one answer. A reset no longer abandons anything; the load under way still answers for its layer. A frame the hidden layer already holds flips at once if it decoded and is skipped if it failed. The pane test's native model now treats a same-source write as a no-op and answers each change once with the layer's current source; both new cases (reset mid-decode then the same frame, failed decode then the same frame) freeze on the previous head. Also, per review: the pacer no longer touches busy or metadata. The stream hook creates it and receives each frame as it goes on screen, so metadata (and with it touch mapping) now follows the visible frame rather than one still decoding. * test(mobile): hold the pane test's AppState listener without a type assertion * fix(mobile): never replay a tap the host answered A fulfilled browser.mouseClick ran on the host, but a null result still replayed it as move/down/up. The native bridge always answers { clicked }, while the external-Chromium provider returns agent-browser's `data` as is, which can be null, so a right-click there was a double tap. Only a refusal that is not delivery-unknown now replays. * test(mobile): re-record the corpus for the answered-tap rule and rename its checkpoint The commit before stops replaying a tap the host answered with a null result. The seed's checkpoint was named clicked-by-fallback, which several partitions no longer do, so it is renamed tap-settled in the same record. Baseline bumped to 2c2b84c by editing that line, then --record. 788 files, 841 changed lines each side. Every one classified: - `baseline`: 787 goldens + pilot-scenarios.json. - the checkpoint rename: pilot-scenarios.json (1), its id in browser-pointer-click-fallback.json (1) and the eleven partition ids in each of the four matrix-browser.pointer-click-*-1.json (44), plus `scenarioSha256` in those five goldens. - behaviour, one checkpoint: the result-null partition in matrix-browser.pointer-click-browser.mouseclick-1.json now sends only browser.mouseClick#1 (sender and payloads drop move/down/up). The refused, method-not-found and result-absent partitions still replay, unchanged. * fix(mobile): cover the frame layer remount, and trim the pacer's edges Nothing tested the remount path: a fresh Image loads the source it mounted with, so the pacer re-arms that load and puts back what the layer held. The pane test's native model now mounts each host fresh (an Image loads its mount source), and a new case remounts both Images mid-decode through a zero-size layout: the visible layer keeps its frame and the stream goes on. Deleting either the re-arm or the put-back fails it. Also per review: drop the dead mountedUri guard in drain, return early from attachImage when nothing is mounted, say that replace writes over a loading layer, unexport the unused pacer types, cover the modifier bail in the tap's comment, and move the frameUri state into the stream hook, which now returns the source the layers mount with.
* feat(search): bundle ripgrep for local, WSL, and SSH search Ship @vscode/ripgrep-universal's prebuilt rg for all six relay platforms in every desktop artifact. Local and WSL searches spawn the bundled binary and drop the git ls-files / git grep fallbacks; SSH deploys upload the remote's binary once per ripgrep version and the relay prefers it over PATH rg. * fix(search): address bundled ripgrep review findings - Key the SSH ripgrep cache on the binary's content hash; a package bump is the only update step - glibc verifier: read arch tokens below the slice root and accept static ELFs (arm64 release blocker) - Ship ripgrep/PCRE2/musl license notices; bundle rg with orcad - Packaged builds never spawn a bare rg; report fd pressure as transient - SSH: install rg before sweep/GC, size-validate installs, back off instead of disabling on launch failure - Scope Dependabot to @vscode/ripgrep-universal; revert unrelated lockfile churn * chore(search): drop bundled-ripgrep reference doc; assert full packaging layout parity * refactor(search): one entry point for spawning the bundled ripgrep Local Quick Open, Quick Open path search, the Explorer name filter, and runtime text search each repeated the same three steps: resolve the bundled command, spread in the WSL distro, spread in the WSL shell expression. Fold that into spawnBundledRipgrep so one place owns the rule that a bare 'rg' must never reach spawn, and simplify the resolver's command/packaged checks. Restore the AGENTS.md ripgrep rule dropped alongside its reference doc in 63f4dac, and note why the relay's availability probe may spawn a bare 'rg'. No behaviour change; verified by the existing suites plus a new test that pins the local, WSL-routed, and distro-routed-but-Windows-output cases. * refactor(search): drop the local install-ripgrep path; enforce the rg rule Bundling rg removed the local git/readdir fallback, so nothing can produce the "install ripgrep on the host running the Quick Open scan" guidance any more -- only a remote host an upload never reached still reaches the capped listing. Drop the host parameter, the renderer's local branch and its translation key, and the relay wrapper that existed only to pass 'remote'. Add a ratchet test for bare 'rg' spawns, since the AGENTS.md rule alone had nothing enforcing it. Its one allowlist entry is the relay's PATH probe, which asks about PATH by definition. Verified the guard catches a planted offender rather than passing vacuously. Also stop chaining the remote cleanup sweep behind the ripgrep upload: on a cold host that is a multi-MB transfer, and stale upload stages and superseded version dirs were left on the remote for its whole duration. The two touch different trees, so they now run concurrently. * test(ssh): pin that the cleanup sweep does not wait on the ripgrep upload * fix(search): derive rg spawn types instead of importing node:child_process A type-only import still counts against the child_process ratchet, whose pin and allowlist only ever shrink. Derive both types from wslAwareSpawn instead. * fix(search): surface an unreachable WSL workspace instead of an empty result Inside `bash -c`, a failed `cd` exits 1 -- the same code ripgrep uses for "no matches" -- so a WSL workspace whose directory had gone away reported an empty listing as a successful scan. main did not have this hole: checkRgAvailable ran the same `cd` wrapper first and settled on `code === 0`, diverting to the git fallback that this PR deletes. The WSL wrapper now takes an optional cwdFailureExitCode; rg passes 97, and all four close handlers reject with a clear error before the unavailable check can blame the install. Also from review: - Bound the fire-and-forget ripgrep upload with deploySignal. The controller aborts only on the deploy timeout, never on success, so this cancels a still-running upload when the deploy gives up. - Run the stale-stage sweep before the installed check rather than inside its else branch. Once rg was installed every later deploy took the PRESENT path, so a stage orphaned by a dropped connection was never collected again. - Note in orcad-remote-deploy.ts why wiring it up needs ripgrep work first: build-orcad.mjs copies only the build host's rg, and orcad reports isPackaged() === true, so a remote of another platform would find nothing. ssh-relay-deploy.test.ts sat at the max-lines cap, so any edit to it failed the gate. Split the four Windows named-pipe deploys into their own file (926 -> 737 + 333); both are now well clear of it. * fix(search): name the unreachable root in every handler, not three of four Round-two review caught that the missing-cwd branch in scanRipgrepPaths sat AFTER isRipgrepUnavailableExit, which classifies any code above 2 as a broken install -- so for exit 97 it was dead code and Quick Open still told the user to reinstall Orca. Reordered; all four handlers now check it first. Also from review: - A vanished workspace makes spawn fail with ENOENT, which read as a damaged install on every local path. Confirm the cwd with isRipgrepSpawnCwdUsable -- the guard the relay already applies -- before blaming the binary. The async continuation re-checks `resolved`, because finish() drops its argument once settled and the rejected promise would otherwise go unhandled. - bundledRipgrepCommand returned a bare 'rg' for an arch outside the bundled set, bypassing the guard that exists so Windows cannot resolve a bare name against the repo cwd. A packaged app now always names an absolute path. Drop ci-shards/unit-assignment.json, a 9,425-line CI artifact swept in from reproducing a shard locally, and gitignore the directory that produced it. The "rg genuinely cannot start" test pointed at a synthetic /repo, which the new guard correctly reports as unreachable; it now resolves to a real root so it still tests what its name says. * fix(search): let the error handler own the spawn-failure verdict A failed spawn emits 'error' and THEN 'close' with a negative code. The cwd check added in the error handler did not settle, so the close handler settled first -- synchronously, with the reinstall message -- and won the race every time. The branch was not merely flaky, it was unreachable in all four handlers: it is guarded by pid === undefined, which is exactly the case that always produces a following close(code < 0). Verified against a real spawn: 3/3 runs give error(ENOENT) -> close(-2). The error handler now detaches 'close' before the probe, so it owns the outcome. The probe also had no rejection handler, so a probe that rejected left the search unsettled forever -- a hang, not just a wrong message. It now falls back to the prior verdict rather than inventing one. Tests: filesystem-search-rg-timeout and orca-runtime-files-search already cover error-first and close-first, but against synthetic roots that the new guard correctly calls unreachable; they now resolve to a real root, keeping each test's stated intent. Added a Quick Open case for the vanished-workspace path and confirmed it fails with the old ordering. * test(search): cover exit code 97 in all four ripgrep close handlers Round-four review found the missing-cwd branch had zero handler coverage: no test anywhere emitted close(97), only -2/0/1/2/127. Ordering was correct, but guarded by source-line order alone -- and that exact ordering was wrong in three of four handlers two commits ago. Each suite now drives close(97) through its real handler and expects the unreachable-root message. Verified the tests earn their place: neutering the missing-cwd check fails exactly four tests, one per handler. Also drop a Reflect.get the anti-slop gate rejects, in favour of `in` narrowing. * docs(search): stop claiming the close handler always wins the race The previous commit asserted close "would beat this threadpool round-trip every time", from an n=3 sample that measured event ordering -- which was never in dispute -- rather than probe-vs-close. Two later measurements disagree with each other: 50/50 close-first here, 30/50 probe-first in review. Either way it is a race on a sub-millisecond margin, and the detach is what makes the verdict deterministic. Why this wording matters: "close wins every time" is an argument for deleting the detach as a guard against an impossible race. No test would catch that -- the suites emit error and close in the same synchronous tick. * chore(search): ship the jemalloc and libunwind notices the Linux rg needs The statically linked Linux builds carry jemalloc (BSD-2-Clause) and LLVM libunwind (Apache-2.0 WITH LLVM-exception) in addition to PCRE2 and musl, and both require their notice on binary redistribution. Confirmed with `strings`: their symbols are present in linux-x64 and linux-arm64 and absent from the darwin and win32 builds. Texts taken from the upstream canonical sources. extraResources already copies the whole licenses directory, so these ship without a packaging change. * fix(relay): stop spawning a bare rg, name unreachable roots, collect old builds Three gaps the reviews surfaced on the remote side, all pre-existing on main. Bare `rg` on Windows remotes. Both relay spawn sites pass the user's repo as cwd, and CreateProcessW searches the cwd before PATH -- the same hijack the desktop side already fixes. The relay now walks PATH itself and spawns an absolute rg.exe, skipping relative PATH entries because those resolve against the cwd. No rg on PATH yields null, which callers treat as "ripgrep unavailable" rather than handing spawn a bare name. POSIX keeps the bare name: execvp never consults the cwd, so there is nothing to resolve and nothing to gain. With the last probe converted, the bare-spawn ratchet allowlist is empty. Empty results for an unreachable root. settleLaunchFailure resolved an empty, successful-looking scan when the root was gone but PATH rg existed, and the git/readdir chain never engaged because it only triggers on RipgrepUnavailableError. Both relay paths now reject naming the root, matching local workspaces. Missing-rg keeps precedence over a missing root, because only that verdict engages the fallback chain -- two tests pinned that deliberately and it would have been wrong to flip it. Unbounded ~/.orca-remote/ripgrep/. Nothing collected this tree; the relay's version GC only matches `relay-*`, so every rg bump left another ~5 MB per host forever. The probe command now also drops sibling builds older than two weeks, sparing the current one and live upload stages, on POSIX and PowerShell alike. Two weeks because a client pinned to an older build may still be using it; the cost of collecting one early is that client re-uploading once. * fix(relay): probe the rg that failed, and close the drive-relative PATH hole Five review findings against the previous commit, all reproduced first. The launch-failure classifier probed PATH rg, but the spawn that failed was the bundled binary. On the normal remote setup -- no rg on PATH, which is why Orca uploads one -- the probe failed and a moved workspace was reported as a missing ripgrep, telling the user to install what Orca already ships. So the fix was inert on exactly the hosts the uploader exists for. It now takes a candidate list and asks the binary that actually failed first, then PATH. path.win32.isAbsolute accepts `\tools` and `/tools`: rooted, but carrying no drive, so they resolve against whatever drive the process is on. The probe would have validated one against the relay's drive while the spawn, running with the user's repo as cwd, resolved it against the repo's -- the same cwd-dependence this lookup removes, narrowed from directory to drive. A real drive letter or UNC root is now required. probeRipgrepVersion had lost the timeout's kill in the rewrite, leaking a live process and a ref'd handle per launch failure -- for a hang, which is the very case the bundled-rg back-off exists for. It also spawned without windowsHide, which would flash a console; fixing that made an allowlist entry stale, so the entry is gone and the pin ratchets down 63 -> 62. `windowsPathRipgrep ??= …` never memoised a miss, because null is nullish. The caching was inverted against cost: a hit stops at the first directory, a miss stats every one, and only the miss was repeated -- per spawn. The bare-spawn ratchet claimed "nothing in production spawns a bare rg", which is false on POSIX. It now also matches PATH_RIPGREP_COMMAND at a spawn site, and the comment states plainly what a textual guard cannot see: the POSIX bare name reaches spawn as a parameter, and is safe because execvp ignores the cwd. The drive-rooted predicate is tested directly rather than through the filesystem -- a temp dir on a POSIX CI host has no drive letter to exercise win32 semantics with, so the filesystem test could never have caught this. * test(mobile): repin the session closure past #22452's two shared modules Merging main brought the closure to 4220 against a pin of 4218. The two extra modules are `src/shared/agent-turn-outcome.ts` and `src/shared/main-agent-status.ts` from #22452, which the status projection this route already reaches import. That change was src/shared-only, so the mobile job never ran on it -- the same way the structured tool line slipped past, as the ledger above already records. Repinned here because this PR's file set is what next made the job run, not because this PR reaches either module. Verified: of the 28 source files this branch changes, none appear anywhere in the route's 4220-module closure. * fix(search): preserve remote binaries and complete runtime packaging * test(relay): pin the probe's env now that it inherits the relay's PATH 8d6759a threaded the relay env into probeRipgrepVersion -- correctly, since the probe decides whether a launch failure was the binary or the root and so has to resolve the same rg the failed spawn would have. It left the assertion that pins the probe's spawn arguments behind, which is what CI caught. Asserting buildRelayCommandEnv() rather than loosening the match to any object: under process.env the probe could resolve a different rg, or none, which is the regression the change exists to prevent. * feat(ssh): collect remote ripgrep builds by reference, not by age Nothing collected `~/.orca-remote/ripgrep/`: the version GC matches only `relay-*`, so every change to the shipped bytes left another ~5 MB on every SSH host, permanently. The age window this replaces was the wrong instrument -- a directory's mtime is when it was written, not when it was last used, so it cannot tell a superseded build from the one a live relay was launched against. Deleting the latter is not graceful degradation: without a PATH ripgrep remote text search rejects outright, and listing drops to the capped walk this PR exists to remove. So the question is reference. Each relay directory now records the build it runs against in `.ripgrep-ref`, written only once that binary is confirmed present, and the GC collects a build only when no installation names it. The discipline is ssh-relay-native-deps-cache-gc.ts': anything the pass cannot account for blocks the whole pass. A relay directory with no readable marker is an older Orca's, possibly running right now against a binary it never recorded, so the pass declines rather than guessing. Those directories are removed by the version GC in time, which is what makes their builds collectable -- hence running after it, not beside it. Deletion is the same tombstone, recheck under the rename, then remove, so a deploy that takes a reference mid-pass gets its tree restored. Windows has no pass yet, matching the native-deps cache's gate. One test note: the first version of the "unaccountable blocks the pass" test passed against a deliberately broken guard, because the tombstone recheck masked its absence. The test now puts a readable recheck behind an unreadable first scan, which is the only shape that fails when that guard is removed. Recording the reference lives inside ensureRemoteBundledRipgrep rather than at the call site: it is the same concern, and it keeps the deploy's ripgrep surface to one call for the tests that mock it to protect their exec queues. * feat(ssh): collect Windows remotes too, and ship the Rust crate notices Three items previously left documented-but-open. Windows remote accumulation. The cache GC was POSIX-gated, so the leak did not go away -- it moved to the platform with the larger binary (rg.exe is 5.43 MB on win32-x64, against 4.77 MB for linux-arm64). The PowerShell dialect now does the same reference scan: entries and references carry token prefixes, because PowerShell writes every uncaptured value to stdout and an untokenised listing would feed Remove-Item whatever a cmdlet happened to emit. Verified on a real Windows host rather than a mock: the listing emits its ENTRY/LIST_OK tokens, a relay directory carrying a marker yields REF <entry>, and a relay directory without one yields REFS_ERR -- the safety path, on the real interpreter. Rust crate notices. The crate set was read out of the shipped binary's symbols and the licence identifiers taken from crates.io rather than assumed. Where a crate offers the Unlicense, Orca elects it: a public-domain dedication carries no notice obligation, and that covers eight of them. The four that do not offer it get their MIT text reproduced. encoding_rs carries a BSD-3-Clause notice for its WHATWG-derived encoding data that is joined by AND, not OR, so electing MIT does not discharge it. Release-only validation, corrected rather than repeated. Linux AppImage/deb/rpm already runs in CI's package job on every PR, and Windows signing was already rehearsed on this branch. macOS notarization is the only item a release must still exercise, and the exposure is narrow: notarization requires signatures on Mach-O binaries, and of the six bundled builds only the two darwin ones are Mach-O -- `file` reports ELF for linux and PE32+ for win32 -- so signIgnore excludes only files the notary never asks about. orcad-artifacts.test.ts caught the new notice file missing from the standalone runtime's shipped list, which is exactly the gap that test exists to catch: a notice committed to the repo but never actually shipped. * fix(search): protect relay cache references and handle failed spawns * fix(ripgrep): close review gaps and repair deployment fixtures * test(mobile): refresh merged session module census * fix(ssh): preserve ripgrep caches with empty legacy references * test(mobile): assert bundle boundaries instead of global module count
…#22586) * fix(terminal): serialize only the visible width after a column shrink xterm does not reflow the alternate buffer (or a normal buffer under pre-21376 ConPTY), so after a shrink each line keeps its old length. SerializeAddon walked every non-final row to line.length, so any snapshot taken after a shrink carried stale right-hand cells that wrapped into extra rows on replay; restores repainted that garbage and a differential TUI such as OpenCode never cleared it. Clamp the row walk and the wrap-boundary lookups to the terminal's columns in Orca's addon-serialize source patch, and regenerate the bundles, maps and lockfile hash per docs/reference/xterm-patch-regeneration.md. * test(terminal): read shrink-snapshot fixtures through public APIs Drops the private-terminal casts the casting gate flags; the normal-buffer case now drives a plain pre-21376 ConPTY terminal and its SerializeAddon directly. * fix(terminal): blank a wide glyph clipped by a column shrink when serializing After a non-reflowing shrink a width-2 glyph can have its lead half in the last column and its trailing half past the grid. Serializing the lead half makes the replay wrap it to the next row and shift every row below, so serialize that cell as a blank and keep the row exactly the grid's width. A glyph ending exactly at the edge is unchanged. * test(mobile): move the session closure pin past the main agent status modules #22452 added src/shared/main-agent-status.ts and src/shared/agent-turn-outcome.ts, which agent-status-types.ts imports, so the session route's closure grew by two local modules (4218 -> 4220). That change was src/shared-only, so its own CI never ran this suite; main has been at 4220 since, and any PR that fires the mobile web app job fails on the stale pin. Measured on 4064653 and on origin/main 3ea15dd. * test(terminal): differential serialize round-trip fuzz and transcript replay Seeded VT streams (text, CJK/emoji/combining, SGR, cursor/edit ops, scroll regions, DECAWM/IRM, alt-screen variants, DECSC, and shrink-heavy resizes) drive a source terminal in three modes: reflowing normal buffer, alternate buffer, and a non-reflowing pre-21376 ConPTY normal buffer. At each checkpoint every SerializeAddon build under test serializes it, and each output is replayed into a fresh terminal of the same size and compared cell by cell, plus cursor, active buffer and modes. CI runs 25 seeds per mode against a pinned list of pre-existing divergences, and replays the committed PTY transcripts (the existing agent fixtures plus new vim, less, pico, and OpenCode captures) under four resize schedules. Point ORCA_OLD_SERIALIZE_ADDON at a previous patched build to also check byte identity when no line is wider than the grid, and that no checkpoint regresses. * test(terminal): build the differential serialize baseline from any git ref config/scripts/build-serialize-addon-at-ref.mjs reverse-applies the patch that produced the installed @xterm/addon-serialize dist, applies the ref's patch, and verifies each step against the patches' blob hashes, so the fuzz can use origin/main (or any fix commit) as its baseline without a second install. ORCA_NEW_SERIALIZE_ADDON swaps in a built dist for the build under test, and a seed-pinned test replays the nine I3 regressions found against origin/main. * fix(terminal): serialize a clipped wide glyph as a width-1 blank The stand-in for a wide glyph clipped by a column shrink came from getNullCell(), whose width is 0. _nextCell skipped it as a wide trailer, and the row-end wrap check counted the width-0 _backgroundCell as content, so a soft wrap after the clipped column was taken as natural and replayed one column early (conpty seed 1149: `abcdefghi中WRAPPED` at 12 -> 10 cols replayed as `abcdefghiR`/`APPED`). Blank the cell in place instead: width 1, no codepoint, its own attributes, so it counts as one empty column and forces the wrap. Differential sweep vs origin/main, 7000 cases per mode: I1 0 byte diffs, seed 1149 fixed; the remaining I3 regressions are the trailing background-row seeds. * test(terminal): neutral paths in the serialize fixtures and usage comment The OpenCode transcript carried this machine's lane paths in its footer; replace them with same-length neutral paths so the recorded cursor layout is unchanged. * fix(terminal): keep trailing background-only rows when serializing Without scrollback, _serializeString trims rows after the last content cursor. A row made only of background-colored blanks emits its erase in _rowEnd but never moved that cursor, so two or more such rows at the bottom were dropped (4x3 `r1\r\n\e[48;5;157m\e[J\e[0m\e[3;1H` replayed with the last row blank). Track the erase separately and extend the kept rows to it, except when the cursor is wrap-pending: relative moves back from those rows cannot re-create that state, and doing so regressed normal 508, alt 1425/6647, conpty 4699. Harness: I1 now exempts checkpoints with a background row after the last text row, the one place this fix changes bytes on purpose (scope helpers move to serialize-grid-variant-scope.ts); conpty seed 5 leaves the pinned pre-existing list. Sweep vs origin/main, 7000 cases per mode: I1 0, I3 regressions 0; fixed/both-fail normal 967/3914, alt 2628/8316, conpty 6657/2937 (was 323/4558, 2296/8648, 5466/4128). * test(terminal): narrow the I1 carve-out to where the serializer keeps background rows Trailing background-only rows change bytes only when the serialized range has no scrollback (the trimming path) and the cursor is not wrap-pending; checkpoints with scrollback or a wrap-pending cursor are held to byte identity again. The 7000-per-mode sweep against origin/main stays at I1 0 and I3 0. * test(native-chat): pin that screen scrapers ignore kept background rows The serializer now keeps trailing background-only rows, so a painted TUI's screen read as text ends in \r\n\x1b[NX rows. Serialize the same frame with and without them and check the Claude option scrape, the empty-prompt check and the fork transcript read the same thing. * test(terminal): read xterm core internals through a checked parser, not Reflect.get The fuzz oracle reached xterm's private _core with Reflect.get, which the low-evidence gate rejects. Narrow _core, writeSync and the DECSTBM bounds with in/typeof checks into one named XtermCoreInternals shape instead. * test(terminal): keep captured serialize transcripts byte-exact on Windows checkouts
…undled Freebuff agent) (stablyai#22601) This reverts commit 0677271. a process-incarnation fallback for reaping leaked orchestration worker terminals, and an unannounced "Freebuff" third-party agent (catalog entry, icon, locale strings, README rows). The Freebuff agent was never meant to ship, so the whole PR is reverted; the reap fix should be re-submitted on its own. Until that re-land, a worker whose durable terminal handle goes stale is again reported missing on release/stop instead of being re-found through its process incarnation, so its terminal can leak on Remote Server. The mobile session page closure pin moves 4218 -> 4219: the revert drops the freebuff icon stablyai#22119 pinned (-1), and stablyai#22452 had already added two src/shared modules without re-pinning (+2). (cherry picked from commit 7a4f080)

ELI5
The sidebar,
worktree ps, mobile and the dashboard each show one status per agent: working, waiting or done. When Claude's main agent has finished but a subagent it started is still running, that status says "working", which is right for the user, but it throws away the fact that the main agent itself is done. This change records that fact next to the status. Along the way it replaces an old saved marker that Orca used to remember "only subagents are keeping this row busy", which fixes a few cases where a row got stuck on "working" after a restart.What Changed
Before and after, as the user experiences it. Normal flows look the same: every row still shows the same combined status, the same "monitoring" label for a settled agent with a background watch loop, and the same timers. These edge cases change, all around Claude or Codex background subagents:
The mechanism. Claude, Codex and Grok hook rows and structured-session rows now also carry the main agent's own state (other agents' rows and terminal-title-only rows carry none, and readers fall back to
state):stateis the main agent's own state before child work is folded in. While a wait belongs to a subagent (a subagent's permission prompt), it is the main agent's state that the wait displaced, usuallydone, notwaiting.outcomeis the recorded verdict on the main agent's most recent finished turn, present only whilemainAgent.stateisdone; a new turn clears it and restating the same finished turn keeps it. Absent means unknown, never success. It is reported by the agent (ClaudeStopFailure→failure,is_interrupton builds that send it, Grokstop_cancelled/stop_failure, the structured journal's own verdict) or, for a cancel, inferred from the user's Ctrl+C: current Claude Code sends no hook on a cancel, so Orca's inferred interrupt is the main source in the Claude hook lane. A bare Esc is never read as a Claude cancel.stateStartedAtis the main agent's own clock, separate from the row's clock. A new Claude session restarts it.Those producers publish it: the Claude hook lane from the main-agent turn record it already kept; structured sessions (host ingest and the renderer bridge), with the session summary gaining an optional
turnOutcome; Grok, whose inline "keep working after stop" logic moves onto the shared fold with identical results; and Codex from its root record (its own combining rule stays, because a waiting child wins there and the shared fold has no waiting-child input yet).mainAgentlives inside the normalized payload, so one function admits it on the relay wire, over IPC and from disk; a malformed one drops the field and keeps the row.What replaced the old saved marker. Orca used to persist
claudeLeadBoundaryChildOnly, a flag recording "the main agent's last stop happened while only subagents were running, and nothing but subagent events arrived since". It is no longer written:mainAgentthat isdone, but only when the row also says no shell was running (claudeRunningNonAgentTask: false, now persisted). A row that does not say stays unseeded, because a shell's liveness is not restored after a restart. Codex seeds its root from a savedmainAgenttoo. Older rows that still carry the flag load asmainAgent: donewith no shell.mainAgentwhere every row is built: a Claude hook event restates the shell fact, and any other rewrite (a terminal-title repaint, an inferred answer) keeps it only whilemainAgentis unchanged. This replaces three places that each had to remember to copy it.mainAgentchanges.mainAgentonly for writers that never carry one (terminal-title, launch and process rows), so a hook row without it matches what the host serves mobile and the CLI.A parity table drives all four lanes through the same six stories and asserts the published
{ state, workingMode, mainAgent }. One is pinned as a known divergence: after a cancelled turn with a shell still running, the hook lane readsdoneand the structured lane readsmonitoring; the cancel policy change removes it.Why
The combined status was answering two questions at once: "is the main agent working?" and "should the user see working?". Answering only the second destroyed the first at publish time, and every consumer that needed it grew a guard that rebuilt a fragment: the saved child-only flag, the structured
fromChildWorkclock switch, keep-awake special cases. Storing the fact once and deriving from it removes that class of guard.Alternatives considered:
fromChildWorkon the row. Derivable frommainAgentbut not the other way round, it cannot say "the main agent is blocked on a question while a subagent runs", and once persisted it is permanent.state, so the host keeps publishing the combined value andmainAgentis a new optional field.mainAgent. It reproduces the old behavior exactly, including a terminal title hiding a subagent that started after a normal stop, and it is a second persisted copy of "the main agent is done" that can disagree withmainAgent.mainAgentis accurate during subagent waits, the marker is derivable, and a permanent field that duplicates a derivable fact is the thing this change removes.Linked Issue
Follow-up to #22295, which added the fold this change makes explicit. No separate issue.
Visual Proof
Isolated dev instance, real Claude Code 2.1.280, sidebar plus the row read from the renderer store. Captured at
5ffe00ffa0; the commits after it change how the shell fact is kept and do not change these three flows.Main agent done, background subagent still working: sidebar working,
mainAgent.state: done.A background subagent's permission prompt after the main agent finished: sidebar waiting,
mainAgent.state: done. Claude's terminal title was its idle✳form for the whole wait; the row still stayed waiting (below: 34 seconds later), then settled to done after approval.After quitting and relaunching Orca while the subagent ran, the same session came back working with
mainAgent.state: done.That subagent then died on an API usage-limit error, and the row stayed working with the dead subagent listed. That happens on
maintoo, with or without a restart (Claude sends no subagent-stop event for it), so it is left for a separate change.Testing
pnpm tcexit 0;pnpm exec oxlinton every changed file clean;pnpm run check:code-quality:changedpassed.pnpm test(withORCA_STRUCTURED_SESSIONunset) oversrc/main/agent-hooks,src/shared,src/relay,src/renderer/src/store/slices, the structured status bridge, the runtime mobile projection andsrc/main/native-chat/agent-session-wire: 1511 files, 15753 tests passed.Server-level regression tests drive the real hook server with the events Claude, Codex and Grok send, including restarts through a persisted profile:
server-last-status-child-held-main-agent.test.ts,server-main-agent-turn-verdicts.test.ts,server-last-status-main-agent-fact.test.ts. Each fix was proven by deleting it and watching its test go red for the pinned reason: restart seeding, the terminal-title refusal (both narrowed and removed), the shell-fact pairing rule, fail-closed seeding, the legacy-row mapping, the held-prompt background evidence, the held-row push, the renderer carry, the Claude session clock, and the Grok and Codex verdict carry (Codex both local and relayed).The restart and permission-prompt stories were also run against
main: the casesmainalready handled pass there, and the "subagent started after a normal stop" restart case fails onmain(stuck working), as expected.Mobile checks: no file under
mobile/changes; the runtime agent-row projection carriesmainAgentas an optional field old apps ignore.Platforms: macOS. No platform branches; relay and SSH paths go through the same normalizer the tests drive via
ingestRemote.I manually tested these changes locally (Electron QA above)
Automated tests added/updated, or explained why not below
AI Disclosure
Author: @BrennanKB5
Review
Persisted shape. Two fields reach
last-status.json:payload.mainAgentand the top-levelclaudeRunningNonAgentTask. The shell fact is the one piece of child workmainAgentcannot express, and restart needs it to avoid settling a row a shell still holds. It is Claude-only and stays outside the status payload readers receive.Two combining rules deliberately remain outside the shared fold and are named in
docs/reference/agent-status-store.md: Codex's own combine, and the cancelled-turn watch-loop divergence between the hook and structured lanes.Known limits, unchanged from
main: a background subagent that dies without a stop event stays listed as working; an inferred Claude cancel is not restored as settled after a restart; an old SSH/WSL relay that predatesmainAgentgets neither the restart seeding nor the terminal-title refusal until it updates.Agent skill upstream boundary
docs/reference/agent-skill-sharing-upstream-boundary.mdand copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.Notes
mainAgenton the status payload andturnOutcomeon the session summary are new optional fields. Old clients ignore them; new clients fall back tostatewhen an old host sends none. Theworktree psrow keeps its shape.mainAgentand sends the shell fact with each event; main re-runs the canonical normalizer on receipt. Codex rows re-publishmainAgentfrom main's longer-lived cache after a relay restart.Checklist
N/Awith reasonpnpm lint,pnpm typecheck,pnpm test, andpnpm buildpass (or CI will cover; local preferred)