feat(native-chat): a Claude subagent waiting on a permission prompt reads as waiting - #22634
Conversation
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — the producer lane that makes a Claude chat subagent blocked on a permission request read waiting until it is answered (5 commits, 17 files; stacked above #22614).
- Request records the asking agent.
ClaudePendingPrompt.agentIdcarriesoptions.agentIDfrom the SDKcanUseTooloptions; set inbuildClaudePermissionCallbacks, in-memory only, not persisted or put on the wire. - Waiting is derived per drain.
drainClaudeChildWorkrecomputes the blocked set fromsession.prompts.pending()(claudeWaitingChildIds,agentIdfirst with achildToolOwnerfallback), emits start/stop flip edges in the newClaudeChildWorkQueue, and marks every live/inventory edge viawithClaudeChildWorkWaiting. - New publishes on the answer path.
answerPromptis now async and publishes child work in afinally; cancel and close already published. - Live error text. A live
task_updatedcarrying onlypatch.errornow reaches the record aslastMessage, with no ending and no new state. - Refactors and docs.
makeRoomForClaudeTaskextracted with its own test;producer()harness moved toclaude-structured-child-work-test-support.ts;agent-status-store.mdand the liveness comment updated for the child-waits/parent-blocked split. - Tests. A scrubbed real-CLI capture fixture replayed through the real adapter (9 tests), plus unit tests for waiting-id derivation, flip edges, inventory marking, live error text, and the capacity rule.
I read the full production diff plus the reconciliation, prompt registry/ownership/control actions, adapter, tracker, and session-close paths, and dispatched two targeted falsification passes (stuck-waiting, and an omitting roster settling a blocked background child); both came back unreachable. I also ran the five changed test files against the checkout — 37 tests pass. The main-allow case does assert no child record is ever created, so it is not theatre.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — incremental pass on the one commit pushed since the prior review (317a3471d8).
- An interrupted subagent now settles
cancelled, notfailed.claudeSpawnResultssettled a failed spawn-calltool_resultas a definitefailed, but a capture shows an interrupt writes that same error result before the child's owntask_updated {status: killed}. It now settlesunknown, and the child's own terminal status refines it through the existing "unknowncan be refined, a definite outcome latches" rule inapplyEnded(agent-status-child-work-reconciliation.ts). A genuine failure (captured with a non-existent subagent model) still endsfailed. - Capture and tests updated. The fixture adds a scrubbed
fg-failscenario (real 404 viaCLAUDE_CODE_SUBAGENT_MODEL) and scrubs the remaining API request ids; the interrupt test now asserts thecancelledoutcome and exact post-cancel timeline, and a new test asserts the genuine failure settlesfailedwith the API-error message.
I re-ran the outcome-affected suites against the new head: 25 tests across the three changed test files, plus 89 across the tracker, reconciliation, admission, view and waiting suites — all pass. The interrupt and failure tests assert exact outcome values, so they fail on revert. The residual (a spawn error with no terminal status following stays unknown, which maps to display idle) is deliberate and already called out in the PR's Review section, so I have nothing to add.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — the two commits pushed since the prior pullfrog review (852adb1b0f), on a PR that makes a Claude chat subagent blocked on a permission request read waiting. The delta is the parent-row behavior change plus its supporting refactor, and a one-line fixture note.
- A subagent's prompt makes the parent row wait, not block.
projectStructuredAgentSessionStatusgained anasking: 'anyone' | 'main-agent'parameter; the sharedprojectStructuredAgentSessionStatusSummarynow projects the session's own status from root prompts only, so a subagent's pending prompt no longer makes the sessionattention. The waiting child record still folds in, so the parent row readswaiting. Every other reader (teardown, restart candidates, pointerawaitingHuman) keeps the'anyone'default and still treats a child request as a human gate. - Prompt rows name the asking agent. Approval/question rows now carry producer linkage via
producerOf→claudeChildToolQueries.promptProducer, using the same join (agentIdfirst, else the gated tool call's owner) that decides which child readswaiting. - An answered prompt keeps its linkage. The answer path revises the row in place with
agentJournalLinkageFields(item), so the subagent's row does not re-root. - Refactor for file size. Child-tool queries moved to
claude-child-tool-queries.tsand theClaudeJournalTranslatorcontract toclaude-journal-translator-contract.ts(still re-exported from the translation module, so existing importers resolve); the tracked-task capacity rule moved tomakeRoomForClaudeTaskwith its own test. - Fixture note. The permission-capture
orderstring now states which ids are scrubbed and which are the capture's own.
I read the incremental range-diff in full plus the authoritative full-diff regions for every changed production file, and ran two read-only falsification passes. Both came back refuted: (1) the 'main-agent' summary change regresses no named consumer — worktree ps/dashboard bucket blocked|waiting both map to attention, mobile reads the already-folded row, restart/teardown/pointer keep 'anyone', and Codex prompt rows carry no agentId so Codex is unchanged (the updated doc claim is accurate); (2) the prompt-row linkage mis-attributes no isRootAgentJournalItem consumer in either direction. I also ran the affected suites against the checkout: 87 tests across the waiting/projection/host files and 41 across parity and child-work evidence, all passing. The src/main/runtime/orchestration failures in this sandbox are environmental (AppEnvironment not initialized, timeouts) and untouched by this PR.
ℹ️ Nitpicks
- The PR description's Visual Proof and Review sections are now stale: they still say the parent row reads
blockedwith the red attention dot and cite the removed testfeeds the parent row's fold a waiting child…. The new behavior (and new test) is the opposite, so worth updating the description before merge. - Two comments outside this diff are now inaccurate: the
KNOWN DIVERGENCEnote insrc/shared/main-agent-status-parity.test.ts:238-244(a structured producer can now report a waiting child) and the "delivery gate and the visible 'working' state can never disagree" comment insrc/main/runtime/orchestration/structured-session-pointer-delivery.ts:62-70(the chat view reads'main-agent'while the gate reads'anyone'). Neither is user-visible.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
619639b to
8dac1b1
Compare
b8d1a9e to
c89c4dd
Compare
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — incremental pass on the one commit pushed since the prior pullfrog review (b8d1a9ed14), on a PR that makes a Claude chat subagent blocked on a permission request read waiting. No PR-authored content changed: the head commit is a rebase of the prior reviewed head, and the only PR file the rebase moved is base-branch churn.
- The head commit is a rebase, not a change.
c89c4dd2ad("say which ids the permission capture scrubs and which are its own") has the same subject, the same+1 / -1stats and a byte-identical patch as the previously reviewedb8d1a9ed14. All branch SHAs moved because the stack base (brennanb2025/c5-read-child-records) was rebased forward. - The only PR file the rebase touched is base churn. Diffing the two heads over this PR's paths leaves just
claude-structured-session-adapter.tsandclaude-structured-session-exit-lifecycle.ts; the adapter delta is the base branch's extraction ofsettleUnexpectedExitintosettleClaudeUnexpectedExitin the latter (not in this PR's diff), disjoint from this PR'sanswerPrompthunk atclaude-structured-session-adapter.ts:239-246. - Prior description nitpick is moot. The PR's Visual Proof / Review sections now describe the parent row reading
waitingand no longer cite the removed test, so the earlier stale-description note is resolved.
I re-read the full diff end-to-end (26 files) in addition to the range-diff, and checked the prior thread state: all three prior pullfrog reviews were body-only (threadCount: 0), so nothing needed retiring and the previous review stands.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
254db2d to
20bff9f
Compare
c89c4dd to
136031a
Compare
There was a problem hiding this comment.
ℹ️ No critical issues — one documentation-drift note inline.
Reviewed changes — incremental pass on the 3 commits pushed since the prior pullfrog review (c89c4dd2ad), on a PR that makes a Claude chat subagent blocked on a permission request read waiting.
- The status clock now reads only the session's own asks.
oldestPendingPromptAt(structured-agent-session-status-started-at.ts:64) datesattentionfrom root pending prompts only, matching the'main-agent'-scoped status the summary already used; the old subagent fallback was unreachable. Test expectations updated: a subagent's ask leaves a running sessionworkingon its turn's clock, and a Codex child's ask leaves the settled parentidle/done. - Latest-message readers extracted.
latestStructuredAgentSessionPrompt/...UserItem/...AssistantMessagemoved intostructured-agent-session-latest-messages.tsand re-exported from the projection; the logic is byte-identical. invokeCanUseToolextracted. Moved verbatim intoclaude-can-use-tool-test-support.tsand re-exported from the session test support. The new module importsFakeConnectionas a type only, so there is no runtime import cycle.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
20bff9f to
02631b8
Compare
136031a to
47341b9
Compare
02631b8 to
04870e9
Compare
47341b9 to
5d960cf
Compare
04870e9 to
dbf9315
Compare
5d960cf to
2d7eeed
Compare
There was a problem hiding this comment.
ℹ️ No new issues in the code — one note on the PR description below.
Reviewed changes — incremental pass on the delta since the prior pullfrog review (136031ae4f): a restack that folds the child-work waiting producer into the task decoder and drops the now-redundant supporting refactors.
- Waiting now lives on the decoder.
ClaudeChildWorkDecoder.observeWaiting(waiting)holds the set of children a pending request blocks, emits one live edge when a child starts or stops waiting, and stampsstate: 'waiting'insideobservationso every live edge a drain carries agrees. The standaloneClaudeChildWorkQueueandwithClaudeChildWorkWaitingare gone;drainClaudeChildWorkre-derives the set fromsession.prompts.pending()before every drain. - A live
task_updatederror reaches the record aslastMessage, with no ending and no state (the decode path now reports an existing live task when the patch carries onlyerror). - The status clock and summary read the session's own asks only.
oldestPendingPromptAttakes root prompts only and the summary projects withasking: 'main-agent'; the Codex e2e assertion now expects the settled parent to stayidle/donewith noblockedevent. - Readers extracted for size:
structured-agent-session-latest-messages.ts(byte-identical) andclaude-can-use-tool-test-support.ts. - Dropped by the restack:
claude-child-work-queue.ts,makeRoomForClaudeTaskand its capacity test. The advanced base (dbf9315d33) owns that structure now.
I read the full 24-file diff plus the prior-reviewed implementation at 5d960cfca0 and confirmed the restructure is behavior-preserving: every live edge still comes from the decoder, the flip edge never carried facts in either shape, and this lane has no inventory-edge producer. I ran the six changed test files (84 tests) plus the tracker/decoder suites (43) against the checkout — all pass. Four failures in claude-agent-sdk-exit-proof/claude-stream-json-connection are process-exit-monitor flakes in this sandbox, in files this PR does not touch.
ℹ️ Nitpicks
- The PR description's mechanism and testing sections still describe the pre-restack implementation:
ClaudeChildWorkQueue,withClaudeChildWorkWaitingandmakeRoomForClaudeTaskno longer exist, and the capacity test it lists was dropped. Worth refreshing so a reviewer is not sent after removed symbols.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
dbf9315 to
7d02ed4
Compare
2d7eeed to
79e525e
Compare
7d02ed4 to
a4b5007
Compare
79e525e to
864aed4
Compare
|
There was a problem hiding this comment.
ℹ️ No new issues in the new commit — the delta is a small, correctly-scoped fix.
Reviewed changes — incremental pass on the one commit pushed since the prior pullfrog review (2d7eeed1ea), which restores the turn-completion feed's "a human is being asked" signal after the session status was scoped to the main agent.
- The completion feed reads
awaitsUser, not the main-agent status.StructuredAgentSessionTurnCompletionFeed.observestampedawaitingUserfromstate.summary.status === 'attention'; once that status became the session's own agent's, a subagent's pending approval stopped reaching the completion. It now reads the newstate.awaitsUser. awaitsUseris the "anyone" derivation.projectStructuredAgentSessionStatusStateaddsawaitsUser('main-agent'or'anyone'attention), withfalseon the no-request early return. It stays internal to the host state — only the completion feed reads it, so the wire is unchanged.- Test answers through the current adapter shape. The waiting-subagent replay now uses
response: { kind: 'option', optionId }.
This is an exact restoration of the pre-scoping behavior ('anyone' attention = any pending prompt). I confirmed it is load-bearing — reverting the feed line makes the two pre-existing subagent-prompt assertions in structured-agent-session-turn-completion-feed.test.ts fail ([false] vs [true]) — and ran the four affected suites at the head: 88 tests pass.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
864aed4 to
a61fede
Compare
1aebdca to
c01de97
Compare
a61fede to
304f26f
Compare
| if (row.providerId && agentChildRunStateFor(row.displayState) === 'waiting') { | ||
| into.add(row.providerId) |
There was a problem hiding this comment.
Nested wait marks parent waiting When a nested subagent asks for permission, its still-working parent can inherit a waiting display state because it owns that child. This loop adds both IDs to the waiting set, so the transcript makes the parent appear to be asking too. That makes it harder to tell which subagent needs approval.
Knowledge Base Used: Agent workspaces and sessions
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — incremental pass on the one commit pushed since the prior pullfrog review (6fb3da30f9): the transcript's subagent block now reads the same waiting state the composer strip shows.
- The transcript is handed the strip's waiting set.
NativeChatStructuredSessionwrapsNativeChatMessageListinNativeChatWaitingSubagentsProvider; the newnative-chat-waiting-subagents.tsderives the waiting provider ids frombackgroundTasks.childrenthrough the strip's ownbuildBackgroundTaskGroupsFromViews(children, useStructuredSessionChildRowContext(paneKey)), recursingrow.ownedand keying onrow.providerId. - Journal entries read waiting only while running.
nativeChatSubagentDisplayStatepromotes aworkingjournal entry whoseidis in the set towaiting;NativeChatSubagentRun,NativeChatSubagentEntriesandNativeChatSubagentSectionHeadconsume it. The group verdict countsrunning = working - waitingCount, surfaces the wait as the alert only beside a still-working sibling, draws the strip'sbg-agent-questiondot, and pulses only while something still runs. - Tests.
NativeChatSubagentRun.waiting.test.tsx(5 tests) covers the single waiting child, waiting beside working, returning to the journal state once the host stops reporting the wait, a settled child never reading waiting, and the section head's dot.
I verified the waiting set is built identically to the strip's (same backgroundTasks.children source and same child-row context, so a row reads as waiting on both surfaces), and ran the new file (5 pass) plus the two pre-existing suites it touches — NativeChatSubagentRun.test.tsx and NativeChatBackgroundTasksStatus.test.tsx (33 pass). The new assertions are discriminating: they require the label to read waiting and explicitly not working, so they fail if the display map is dropped.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
…iled sibling A failed sibling took the group row's one alert slot, so a group with a waiting, a working and a failed child read "1 working +1 failed" and hid the wait; it now reads "1 working +1 waiting +1 failed". The waiting set keeps its identity while a child frame changes no wait, so the transcript's subagent rows do not re-render on every frame, and the test of a wait ending now updates one mounted row instead of remounting it.
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — incremental pass on the one commit pushed since the prior pullfrog review (57e9328aa3): a collapsed subagent group now shows a pending wait alongside a failed sibling, plus a keyed memo for the waiting set.
- The collapsed group surfaces both the wait and the failure.
NativeChatSubagentRunreplaced the single adverse-only alert slot with analertsarray that renders the waiting label (running > 0 && waitingCount > 0) and the adverse label independently, so a mixed group reads e.g.1 working +1 waiting +1 failedinstead of letting the failure hide the wait. The dot followsadverseState ?? (waitingAlert ? 'waiting' : verdictState). - The waiting set is keyed by its members.
useNativeChatWaitingSubagentsbuilds a sort-joined id key and returns theSetthroughuseMemoon that key, so a child frame that changes no wait keeps the same reference and leaves the transcript rows memoized. - Tests.
NativeChatSubagentRun.waiting.test.tsxgained the mixed wait+failed case and a hook test pinning the set's reference stability across a no-op frame.
I read the incremental range-diff and the full-diff regions for the three changed files, then ran the new file (7 tests pass) plus NativeChatSubagentRun.test.tsx and NativeChatBackgroundTasksStatus.test.tsx (33 pass); oxlint is clean on all three. The mixed-group assertion is discriminating — the previous single-alert shape would render 2 working +1 failed and fail it. The remaining open greptile P2 (4152371333, nested parent marks both rows) is untouched by this commit and was already surfaced on an earlier head, so not re-raised.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
…e-child-waiting Committed with --no-verify: pre-commit max-lines fails on structured-agent-session-adapter.ts, byte-identical to origin/main 4bd55c0, which #24235 put at 301/300; not touched by this PR. Full oxlint reports that file as its only error. Conflicts kept main's shape with this PR's additions on top: the Claude child decoder keeps main's stopLive beside observeWaiting, and the prompt registry and permission callback take main's removal of turnId while keeping the asking agent. Main's new dismissPrompt path now republishes child work as answerPrompt does, so a declined subagent request stops reading waiting at once.
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — incremental pass on the merge commit a6b58d6b76 (origin/main up to 4bd55c07ea merged into the branch). The only PR-authored content the merge touches is how a settled prompt republishes child work, plus its test.
- A dismissed or declined card now frees the child it blocked. The PR's
answerPrompttry/finallychild-work publish was extracted intoprivate freeingAsker(sessionId, settled) => settled.finally(() => this.publishChildWork(...)), and the newly mergeddismissPrompt(from #24235) now routes through it too, so a card dismissed without an explicit answer still recomputes the blocked set instead of leaving a subagent stuck onwaiting.answerPromptkeeps its pre-mergefinallysemantics. - Decoder merge adaptation only.
claude-child-work-decoder.ts's header comment adopted main'sstopLivewording; the PR'sobserveWaiting/report/waitingfield are unchanged. - Test.
claude-structured-child-work-waiting.test.tsgained the dismiss-and-decline case. It is discriminating: I ablated thedismissPromptwrap and the assertion fails (waitingvsworking); restoring it passes 15/15.
DeepSeek Flash (free via Pullfrog for OSS) | 𝕏
Review status: tests rewritten and trimmed to the common pattern, ready for a merge decision (head
|
…e-child-waiting # Conflicts: # src/main/ssh/orcad-remote-node-runtime.ts
…agent alone reads waiting Drop the split of the main agent's own status from a session-wide "someone must answer" fact: awaitsUserSince, the agent-session.status-awaits-user.v1 capability and its old-client downgrade, and every reader change that only consumed them (fold, equality, ingest, delivery gate, turn-completion feed, quit snapshot, resume headline, status clock, status bridge, attention dispatch) go back to main. The parent row again reads one needs-input state for a pending request whoever asked, dated as before. Kept: a request's owner recorded once on its prompt row with full producer linkage; the asking subagent's own record reads waiting, re-derived on every update; the chat history's subagent block reads that same state; an answered subagent request stays in its subagent's group. A subagent now waits only on a request the user can still answer (its card open, no answer underway), and the adapter frees it before the host records an answer or dismissal. So a waiting child record always sits beside the pending card, and main's fold never reads the parent as waiting on it: no window after an answer, and no ~3 s wait after a card dismissed by Stop.
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — the incremental delta is a merge of main plus b7d03a2a35, which reverts the session-wide "someone must answer" split and returns the parent row to a single needs-input state.
- The
awaitsUserSincesplit is gone. The summary field, theagent-session.status-awaits-user.v1capability and its pre-split downgrade, and every reader that existed only to consume them go back tomain's shape:projectStructuredAgentSessionStatusreportsattentionfor any pending request,foldAgentLeadStatusdrops itsawaitsUserinput, and the pointer gate, completion feed, quit snapshot, status clock, host ingest, renderer bridge and resume headline all read the main-agent status again. The parent row therefore readsblockedwhile a subagent asks, exactly as before the split. - The asking subagent still reads
waiting, now gated on an answerable card.claudeWaitingChildIdsbecame module-private and derives the set fromprompts.unclaimed()(no answer claim held) intersected withjournalPrompts.holdsOpen(promptKey)(the card is still in the journal). The decoder'sobserveWaitingre-derives this every drain, so a child that starts or stops waiting emits its own live edge. - The child is freed before its card closes.
freeingAskernow wrapscommitsofree()republishes child work before the host records the answer/dismissal, and again infinally. Answer takes the claim beforecommit; dismiss callshandOverbeforecommit; a failed host write restores the card so the child waits again. No waiting child can outlive its card. - Docs, parity and tests reverted/retargeted.
agent-status-store.mdnames the structured child-prompt case a known divergence;main-agent-status-parity.test.tsexpects the structured child ask asblocked; the Codex Activity e2e asserts['done', 'blocked', 'done']; the new two-host replay inclaude-structured-child-work-waiting.test.tsasserts the parent row is identical whether or not a child waits.
I read the incremental range-diff and the full commit diff for b7d03a2a35, traced the production answer/dismiss/provider-cancel/teardown lifecycles against the new gate, and dispatched one reviewfrog pass on it: the unclaimed() ∩ holdsOpen combination is a strict subset of the old pending() set, every card-close transition mutates one of the two inputs and immediately re-drains (or clears the decoder), and during the one async gap the prompt is claimed, so no live wait is dropped and no stale wait survives. The parity refold helper and all callers of structuredAgentSessionAgentStatus were updated; a repo-wide grep finds no dangling reference to the removed field, capability, module or pending().
ℹ️ Nitpicks
- The PR description now describes the opposite of this head: its "What this changes for you" table, the
awaitsUserSince/capability mechanism, the before/after screenshots and the "Older clients" sections all document the split thatb7d03a2a35removes, and still claim the parent row readswaiting. Worth rewriting before merge so a reviewer is not sent after removed symbols and reversed behavior.
deepseek-v4.1-flash (free via Pullfrog for OSS) | 𝕏
A subagent's wait was pushed to the host as soon as its request arrived, while the request's card row reached the journal at least a microtask later. So every subagent request published the parent row as waiting before blocked (the main agent's own fold reads a waiting child that way), and Activity got an extra unread "waiting" event that main never shows. The card is now the one record of an open request. The translator records the asker on the card once (its row's linkage) and counts the card open only after the sink confirms its rows landed, then publishes the wait; anything that closes the card (an answer underway, a dismissal handed to the host, Claude's own withdrawal, the session's end) frees the subagent first. So every publish that shows a subagent waiting also shows its pending card, and the parent reads one needs-input state, exactly as on main. This retires the registry's view of pending requests (unclaimed(), the asking-child join) and the translator's holdsOpen. The prompt row's linkage takes one rule: the agent the provider names, else the gated call's owner. The parent-row proof now runs through the real deferred sink, durable journal and status feed, publishing as production does, and checks at every publish that waiting subagents have pending cards and that the parent row matches a host fed no waits.
| if (this.deps.sink.tryAppendItem) { | ||
| admitted &&= this.deps.sink.tryAppendItem(identity, body, options).accepted |
There was a problem hiding this comment.
Rejected prompts hide waiting When the journal queue rejects a permission prompt under backpressure, this code still registers the prompt but never marks its card as written. The subagent then stops showing waiting even though Claude still needs an answer. Draining the queue does not retry the rejected write or restore the waiting state.
Knowledge Base Used: Agent workspaces and sessions
The sink's written() resolved ok when the sink was closed with writes still queued, so a subagent's prompt card could count as open with no row in the journal. written() now reports a close that dropped writes admitted so far as not landed; drained() and lifecycleBarrier() keep reading a closed sink as settled. Tests: a card never opens when its sink closes first; a card Claude withdraws while the sink holds the cancelled row back closes at once; two subagents asking at once, and the main agent asking beside a subagent, keep the parent row as before with each waiting subagent beside its own card; a process that dies mid-request leaves no subagent waiting.
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — incremental pass on the one commit pushed since the prior pullfrog review (b7d03a2a35): a Claude subagent now reads waiting only beside a card the journal has actually written.
- Waiting is gated on the committed card.
claudeWaitingChildIds(claude-child-work-evidence.ts) derives the set fromjournalPrompts.openCards()— entries whose rows landed, carry anasker, and are not mid-cancellation — intersected with the prompt registry'sansweringclaim. This replaces the previousprompts.unclaimed() ∩ holdsOpengate. - The journal reports when a card's rows land.
ClaudeJournalPromptsentries now trackwritten/landed/asker;markWrittenwaits on a new optional sinkwritten()barrier,whenWrittenexposes that settlement, andopenCardsyields only written, open subagent cards.handleappends through the admission-returningtryAppendItem, andhandOverrestores the same entry object so a still-landing write marks the card it hands back. - The adapter publishes a wait only after the card lands.
emitschedulesclaudePromptCardWritten(...).then(() => drainChildWork(...)), so the child-work publish that first shows a wait follows the same barrier the card's row landed on. - Attribution no longer depends on the gated tool call.
promptProduceruseslinkageForAgent(prompt.agentId)whenever the provider names the asker, falling back to the gated call's owner only when it does not — so a request arriving before its tool call is read still files under the right agent. - Docs, parity and tests.
agent-status-store.mdand the parity comment describe the card-gated wait;claude-structured-journal-prompt-cards.test.tspins the open/close/hand-back lifecycle, andclaude-subagent-request-parent-row.test.tsreplays the captures through a real deferred sink and durable journal, asserting at every publish that each waiting child has a pending card and that the parent row is unchanged.
I read the incremental range-diff and the full-diff regions for every changed file, traced the answer/dismiss/provider-cancel/teardown lifecycles against the new gate, and dispatched two falsification passes: one on the written-card barrier (no path publishes a wait without its card, and no wait survives the request closing) and one on the linkageForAgent attribution change (the agent id is invariant old-to-new whenever the provider names one, so the waiting child still matches its record). Both came back refuted. The awaitsUserSince/capability PR-description staleness was already flagged on the prior head and is unchanged, so it is not re-raised.
deepseek-v4.1-flash (free via Pullfrog for OSS) | 𝕏
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — incremental pass on the 2 commits pushed since the prior pullfrog review (c5651f3e71): a sink-closed race in the card-write barrier, plus a merge of origin/main.
- A closed sink's dropped writes no longer read as landed.
StructuredAgentSessionSinkQueuetracksranSequence(the newest operation that actually ran) and addswritten(), which awaits the existingbarrier()and then reports not-landed when the barrier settled without those writes running — the case a close produces when it clears the queue. The deferred sink'swrittenis nowqueue.written(wasqueue.barrier), soClaudeJournalPrompts.markWrittenkeeps a card closed when its rows were dropped.drained()keeps its prior "a close still settles" semantics. - Tests.
structured-agent-session-event-sink.test.tspinswritten()vsdrained()across a close;claude-structured-journal-prompt-cards.test.tsasserts a card never opens when its deferred sink closes before the rows land and that a backpressured cancellation still closes it;claude-subagent-request-parent-row.test.tsadds multiple-asker-then-session-own-request, and mid-request process-death settle cases. - Merge of
origin/main. Carries no PR-authored content.
This closes the exact window the prior run's barrier pass flagged (a close settling waiters ok:true while clearing queued writes). I verified written() captures through before awaiting the barrier, that ranSequence advances only when an operation's run completes, and that the new multi-asker and process-death assertions are discriminating. Nothing to anchor.
deepseek-v4.1-flash (free via Pullfrog for OSS) | 𝕏
…e-child-waiting # Conflicts: # src/main/native-chat/agent-session-wire/structured-agent-session-event-sink-queue.ts # src/main/native-chat/agent-session-wire/structured-agent-session-event-sink.test.ts # src/main/native-chat/agent-session-wire/structured-agent-session-event-sink.ts
…card closes Main's sink now hands each write to the journal as it is submitted, and an idle journal commits it and runs the publication at once. Claude's own withdrawal of a subagent's request therefore closed the card and published the parent row before the child's wait was freed, so one publish showed the subagent waiting beside no pending card (fg-interrupt replay). The child's wait now also requires the request to still be open in the registry, and a withdrawal republishes child work before the journal takes the close.
… and its essential tests The chat history no longer marks a subagent block as waiting: the approval card itself carries the request, and the asking subagent's row in the sidebar and composer strip reads waiting, as before. NativeChatWaitingSubagentsProvider, native-chat-waiting-subagents.ts and their renderer changes go. A subagent waits while its request is still open and unanswered in the prompt registry and its card has landed in the journal. The registry check also covers a withdrawal under backpressure, so the card list no longer filters pending cancellations itself. Tests: one integration file replays the captured CLI frames through the real adapter, sink, journal and status feed (renamed claude-subagent-permission-request.test.ts), with the asking subagent's state timeline, attribution, nested linkage and a card write that waits for the journal. The producer-harness waiting test, the redundant prompt-card cases, the harness reducer swap and three unused captures (deny, interrupt, main agent, failed subagent) are removed.
…, and narrow the oracle's claim
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — incremental pass on the 5 commits pushed since the prior pullfrog review (a6d2db9c61): two origin/main merges plus a withdrawal race-fix, a refactor that trims the feature to the sidebar/strip, and a test pin.
- Fixed a withdrawn-request race. Main's rewritten sink now commits each write as it is submitted, so Claude's own withdrawal closed the card and published the parent row before the child's wait was freed — one publish showed a waiting child beside no pending card (the
fg-interruptreplay).claudeWaitingChildIdsnow also requires the request to still be open in the prompt registry (awaitsAnswer), and the adapter republishes child work on aprompt-cancelledevent before the journal takes the close. answeringbecameawaitsAnswer. The old!answeringread true when the prompt was absent from the registry (a withdrawn/forgotten request);awaitsAnswerreads false there, which is what frees that child.openCards()dropped its owncancellationPendingfilter, since the registry check already covers a withdrawal under backpressure.- Trimmed to the sidebar and composer strip. The chat-history subagent block no longer reads waiting:
NativeChatWaitingSubagentsProvider,native-chat-waiting-subagents.tsand their renderer changes are gone. - Trimmed tests. The integration replay was renamed to
claude-subagent-permission-request.test.tsand keeps thefg-allow/bg-allowcaptures; redundant prompt-card cases and the deny/interrupt/main/failed captures were removed (the decoder suite still covers the outcome vocabulary and the closed-sinkwritten()case). - Pinned the parent row's dating. The renamed integration test now asserts the parent row stays
blocked, dated by the session's own ask, even when a subagent asked first — main's rule.
I read the incremental range-diff, the new commits' own diffs, and the authoritative full-diff regions for every changed production file. I traced the withdrawal/answer/dismiss lifetimes against the new registry gate: each path that closes a card claims or forgets the registry prompt before the card closes, so no wait outlives it, and the written() gate still excludes a card whose row a close dropped. The one open greptile note (a backpressured initial prompt write is never retried, so its child never reads waiting) is a pre-existing drop this lane already intends.
deepseek-v4.1-flash (free via Pullfrog for OSS) | 𝕏
* fix: preserve registered account credentials after enrollment errors * refactor(orchestration): one function for six agent-to-agent sends (#24901) * refactor(orchestration): send every agent-to-agent message through one sendAgentTurn The structured mail-pointer lane and the structured worker preamble each carried a copy of "send, wait out a pending start, read the verdict", and four dispatch-preamble sites typed into a terminal directly. They now share sendAgentTurn, built from host.send, the host's settlement waiter and sendTerminalAgentPrompt. Every caller passes delivery 'now', so nothing sends differently; a source-scan ratchet keeps new direct sends out. * test(orchestration): the refusal fixture uses a real wire refusal code * fix(orchestration): derive the send fingerprint inside sendAgentTurn and pair target with turn A `queue` send was refused by a real host: callers supplied a fingerprint over the body alone while the host digests body and delivery. sendAgentTurn now builds the envelope with the composer's own builder (extracted to structured-agent-session-send-mutation), so the fingerprint is always over exactly the fields sent; `now` sends keep the identical digest. The pointer lane no longer carries a fingerprint it cannot get right. sendAgentTurn takes one argument, a union on kind that carries its own target and turn, resolved by an exhaustive switch; the terminal turn names its purpose (only the dispatch preamble today) instead of every terminal send inheriting the task lead line. A queued outcome keeps the host's draft receipt, and both structured callers read outcomes exhaustively with unchanged behaviour. The boundary test now also fences direct structured host sends, catches optional, bracket and bound member uses, has a planted-offender self-test, and pins the unmoved terminal mail pointer and agent-teams tmux senders. * fix(orchestration): keep federation.ts under max-lines and satisfy prefer-template in the send ratchet * fix(orchestration): wait on the answered submission's id when a replayed queue turn was already handed off * test(orchestration): pin that a replayed queue turn waits for its hand-off to settle * test(orchestration): fence call-result and cast host sends, raw terminal writes and the launch-prompt helpers * feat(native-chat): a Claude subagent waiting on a permission prompt reads as waiting (#22634) * feat(native-chat): a Claude subagent waiting on a permission prompt reads as waiting A subagent's permission request reaches the parent session's callback naming the subagent that asked (agent_id) and the tool call it gates (tool_use_id). The pending request is recorded with the asking agent. On every drain the child-work producer re-derives which children a pending request blocks and hands that set to the Claude child decoder, the one owner of each child's live edges: a blocked child reads waiting on every live edge it reports, and a child that starts or stops waiting is a live edge of its own. Answering, denying or cancelling the request returns the child to its prior live state; nothing is stored beyond the pending requests. A live task_updated carrying an error now reaches the record as the child's last message, without an ending or a new state. The replay test drives a scrubbed capture of the real CLI (foreground allow, deny, interrupt, background allow, and the main agent's own request) through the real adapter into the host's child records. * test(native-chat): a subagent's request names it before its tool call is read * docs(agent-status): a subagent asking for approval waits in every lane; the parent row keeps the session's own attention * test(native-chat): hand canUseTool the asking agent without widening the helper's cast * test(native-chat): an interrupted Claude subagent settles cancelled, not failed A captured interrupt shows the spawn call's error result ("The user doesn't want to proceed…") arriving before the subagent's own `task_updated {status: killed}`. The spawn result ends nothing (the child ends only on its own terminal frame), so the child stays live until its `killed` status settles it cancelled. A genuine failure, captured with the subagent on a model that does not exist, sends its `failed` status before the error result and still ends failed. Both captures now replay through the real adapter into the host's records. * fix(native-chat): a Claude subagent's prompt makes the parent row wait, not block A subagent's pending prompt made the whole session `attention`, which reads as the main agent's own `blocked` and outranks the fold's waiting arm, so the parent row read blocked where a CLI Claude parent reads waiting. The main agent's state now reads only its own pending prompts. - A Claude prompt row carries the linkage of the agent that raised it: the one the permission request names, or the owner of the tool call it gates. The same join decides which child reads waiting, so the two cannot disagree. - The status summary projects the session's own status from root prompts only; every other reader (delivery gates, teardown, restart) still asks whether anyone is waiting on a human. - An answer keeps the prompt row's linkage by the journal's own rule: a revision that names no producer keeps the row's existing one. - The child-tool queries gain the prompt's producer, so a prompt row and a child record answer "which agent" from the same join. * test(native-chat): say which ids the permission capture scrubs and which are its own * fix(native-chat): the status clock dates attention by the session's own asks only The session's status is now `attention` only for its own pending prompt, so the clock's fallback to a subagent's ask could no longer be reached, and it read the journal by a different rule than the status it dates. Both now read root prompts. The journal also stamps a Codex subagent's prompt with its thread (#22532), so a Codex child's approval is that child's wait in the Codex lane too. Two tests written for the earlier rule are updated: a subagent's ask leaves a running session `working` on its turn's clock, and a Codex child's answered approval leaves the settled parent's Activity row done with nothing unread. * fix(native-chat): a completion still says the user is asked when a subagent asks The turn-completion feed marked a completion `awaitingUser` from the status summary's `attention`. The status now means the session's own agent is waiting, so a subagent's pending approval stopped reaching the completion. The projection now also says whether anyone is waiting on the user, as the delivery gates, teardown and restart ask it, and the completion reads that. The waiting-subagent replay answers its prompt with the adapter's current response shape. * revert(native-chat): a live Claude task's error stays out of the child's last message No capture shows a live task_updated carrying an error, and it is unrelated to a subagent waiting on a permission request; it leaves this PR. * test(native-chat): settle the Claude session's startup before replaying a subagent's request A startup frame drained child work during the first await, so answering a request freed the child even with the answer's own republish removed. * fix(native-chat): a Claude subagent's prompt row names it as its other rows do The prompt row stamped only the asking agent's id, so a nested subagent's request lost the agent that spawned it, its spawn call and its run. It now takes the linkage the asker's own rows take: the gated tool call's, when that names the same agent, else the one resolved through the agent's spawn call. The provider's agent id stays the asker's id. * fix(native-chat): a subagent's request makes the parent row wait without a child record The parent row learned that a subagent needed the user only from that subagent's child record, so a request no record carried (a Codex child the host never registered, a Claude task past the live cap) left the row working or done while the approval card sat in the chat. "Someone in this session must answer" is now one derived session fact. The projection names two facts instead of a mode flag: the main agent's own status (attention only for its own request) and structuredAgentSessionAwaitsUser (any pending prompt). The status summary publishes the second as an optional awaitsUser, and the shared fold reads it: the main agent's own ask is blocked, otherwise awaitsUser or a waiting child record makes the row wait. Every caller picks the fact it means: the completion edge's awaitingUser and the delivery gate read awaitsUser; the quit snapshot folds the same two inputs the sidebar does. * fix(native-chat): a client that predates awaitsUser still reads a subagent's request as attention A status summary's status is now the main agent's own, so a client built before the split would read a subagent's request as working (or idle) and fold it with code that has no awaitsUser input. Clients advertise agent-session.status-awaits-user.v1; at agentSession.subscribeStatus the host sends any client that does not the pre-split summary: attention whenever awaitsUser is set, without the main agent's own tool line, verdict and clock. The feed and every in-process reader keep the canonical summary. Transitional, like the turn-item downgrade. * test(native-chat): a Codex subagent's approval makes its settled parent's Activity row wait The test pinned the parent row done while a Codex child asked, through a harness that fed no child records, so it proved nothing about the ask. It now drives the ask twice through the real host status store: with no child record (the session's awaitsUser alone) and with the child's own record waiting from thread/status/changed. Both read waiting with needsAttention while the ask is open, then done with nothing unread. * docs(agent-status): a subagent's request reaches the parent row through awaitsUser in every structured lane The store reference said a Codex child's request still read as the main agent's blocked and that only the Codex hook lane fed a waiting child. Both structured lanes stamp the asking child and feed child records, and awaitsUser carries the request when no record does. The liveness comment goes back to main's: a child's blocked is a failed task on an older host's legacy rows. * fix(native-chat): the restart dialog still headlines a subagent's pending approval The quit snapshot now records the main agent's own state, so a subagent asking while the main agent worked recorded `working` and the dialog said "Was mid-reply" where it used to say "Waiting for your approval". The headline now comes from the snapshot's pending prompt, whoever raised it, with the existing copy; `state` stays the main agent's own. * test(orchestration): a subagent's pending approval holds structured mail delivery Scoping the delivery gate to the main agent's own request left every gate test green; a subagent's request now has its own case. * fix(native-chat): a subagent's request is dated by when it was raised, on every client Since the summary's clock became the main agent's own, nothing dated a wait that only a subagent's request held: a pre-split client was sent attention with no clock, where the old host dated it by the subagent's prompt, and a new client's waiting row fell back to the time it first saw it, so after a reload a request the user had already read could read unread again. The session fact is now when someone started being asked: awaitsUserSince, the oldest pending prompt whoever raised it, and its presence is what awaitsUser meant. A row waiting on someone else's request takes that as its clock; the downgrade for a client without the capability dates its attention by it, which is what the old host published. A cross-version test pinned to the last pre-split release runs the same journals through that release's projection and through this one plus the downgrade, and compares the whole summary. The Codex end-to-end test also reads the host's own status row, and keeps a read ask read through a later row and a reload. * test(native-chat): the pre-split parity check compares only the fields the split owns An additive summary field is safe for old clients, so comparing whole summaries against the pinned release would redden on one. The wire comment now says how the downgrade dates attention: the main agent's own oldest ask, else awaitsUserSince. * test(runtime): an aged host-held working summary states that nobody is asked The test built its working summary by overriding the status of a published approval summary, which still carried awaitsUserSince, so the row correctly read waiting. It now drops the request as its scenario says. * fix(native-chat): the chat's subagent block says waiting when the strip does While a Claude subagent's request was open, the sidebar and the composer strip read waiting but the subagent block in the chat history a few pixels above still read "Kicked off 1 subagent working": it shows the journal's roster state, and the journal records no wait. The structured chat now hands its transcript the subagents the strip shows waiting, read from the host's child records through the strip's own row model and matched by the provider id the roster names each one by. A running entry the host says is waiting reads waiting in the group row, its entry and its section head, with the strip's word and the question colour; it reads the journal's state again as soon as the host stops reporting the wait. * fix(native-chat): a collapsed subagent group shows a wait beside a failed sibling A failed sibling took the group row's one alert slot, so a group with a waiting, a working and a failed child read "1 working +1 failed" and hid the wait; it now reads "1 working +1 waiting +1 failed". The waiting set keeps its identity while a child frame changes no wait, so the transcript's subagent rows do not re-render on every frame, and the test of a wait ending now updates one mounted row instead of remounting it. * refactor(claude): one needs-input state on the parent; the asking subagent alone reads waiting Drop the split of the main agent's own status from a session-wide "someone must answer" fact: awaitsUserSince, the agent-session.status-awaits-user.v1 capability and its old-client downgrade, and every reader change that only consumed them (fold, equality, ingest, delivery gate, turn-completion feed, quit snapshot, resume headline, status clock, status bridge, attention dispatch) go back to main. The parent row again reads one needs-input state for a pending request whoever asked, dated as before. Kept: a request's owner recorded once on its prompt row with full producer linkage; the asking subagent's own record reads waiting, re-derived on every update; the chat history's subagent block reads that same state; an answered subagent request stays in its subagent's group. A subagent now waits only on a request the user can still answer (its card open, no answer underway), and the adapter frees it before the host records an answer or dismissal. So a waiting child record always sits beside the pending card, and main's fold never reads the parent as waiting on it: no window after an answer, and no ~3 s wait after a card dismissed by Stop. * fix(claude): a subagent waits only beside its committed card A subagent's wait was pushed to the host as soon as its request arrived, while the request's card row reached the journal at least a microtask later. So every subagent request published the parent row as waiting before blocked (the main agent's own fold reads a waiting child that way), and Activity got an extra unread "waiting" event that main never shows. The card is now the one record of an open request. The translator records the asker on the card once (its row's linkage) and counts the card open only after the sink confirms its rows landed, then publishes the wait; anything that closes the card (an answer underway, a dismissal handed to the host, Claude's own withdrawal, the session's end) frees the subagent first. So every publish that shows a subagent waiting also shows its pending card, and the parent reads one needs-input state, exactly as on main. This retires the registry's view of pending requests (unclaimed(), the asking-child join) and the translator's holdsOpen. The prompt row's linkage takes one rule: the agent the provider names, else the gated call's owner. The parent-row proof now runs through the real deferred sink, durable journal and status feed, publishing as production does, and checks at every publish that waiting subagents have pending cards and that the parent row matches a host fed no waits. * fix(native-chat): a closed sink's dropped writes never read as landed The sink's written() resolved ok when the sink was closed with writes still queued, so a subagent's prompt card could count as open with no row in the journal. written() now reports a close that dropped writes admitted so far as not landed; drained() and lifecycleBarrier() keep reading a closed sink as settled. Tests: a card never opens when its sink closes first; a card Claude withdraws while the sink holds the cancelled row back closes at once; two subagents asking at once, and the main agent asking beside a subagent, keep the parent row as before with each waiting subagent beside its own card; a process that dies mid-request leaves no subagent waiting. * fix(claude): a withdrawn subagent request frees its child before its card closes Main's sink now hands each write to the journal as it is submitted, and an idle journal commits it and runs the publication at once. Claude's own withdrawal of a subagent's request therefore closed the card and published the parent row before the child's wait was freed, so one publish showed the subagent waiting beside no pending card (fg-interrupt replay). The child's wait now also requires the request to still be open in the registry, and a withdrawal republishes child work before the journal takes the close. * refactor(claude): trim subagent request waiting to the common pattern and its essential tests The chat history no longer marks a subagent block as waiting: the approval card itself carries the request, and the asking subagent's row in the sidebar and composer strip reads waiting, as before. NativeChatWaitingSubagentsProvider, native-chat-waiting-subagents.ts and their renderer changes go. A subagent waits while its request is still open and unanswered in the prompt registry and its card has landed in the journal. The registry check also covers a withdrawal under backpressure, so the card list no longer filters pending cancellations itself. Tests: one integration file replays the captured CLI frames through the real adapter, sink, journal and status feed (renamed claude-subagent-permission-request.test.ts), with the asking subagent's state timeline, attribution, nested linkage and a card write that waits for the journal. The producer-harness waiting test, the redundant prompt-card cases, the harness reducer swap and three unused captures (deny, interrupt, main agent, failed subagent) are removed. * test(claude): pin the parent row's dating when a subagent asked first, and narrow the oracle's claim * Turn desktop notifications on or off per machine (#24518) * feat(notifications): turn desktop notifications on or off per machine Settings > Notifications lists each machine (this computer, SSH targets, paired Orca servers) with a switch. Muted machines are stored as an opt-out list so a newly added machine still notifies. Both notification senders now name the machine a workspace runs on, and main skips the desktop banner for a muted one; phone push is unchanged. * fix(notifications): preserve phone alerts and honor machine mutes * fix(notifications): bind machine mutes to configured sources * fix(notifications): preserve chat completion subscriptions on owner collisions * fix(notifications): reduce machine settings clutter with a collapsed section * fix(notifications): align machine disclosure with settings rows * fix(notifications): clarify machine switches affect only this computer * fix: distinguish unavailable account metadata from absence * fix(accounts): resolve registered profile UUIDs consistently * Replace agentArgsOverride with unified removeAgentArgs approach (#25091) Consolidate override detection and removal into removeAgentArgs. Previously catalogs used agentArgsOverride to detect conflicts and removeAgentArgs to strip them; now removeAgentArgs handles both by returning stripped tokens. This eliminates redundancy and makes the intent clearer. Enhance removeAgentArgOption with optional value filtering to support selective removal for complex cases like Codex config overrides. * Reuse the ancestor path while building mobile agent rows (#24539) Preserve traversal order and cycle guards using one call-local path Set rather than a copy at every depth. * Release waiting terminal output when a mobile subscription fails to start (#24547) Dispose the failed subscription record’s existing terminal backlog before rethrowing its original start error. * Avoid cloning terminal agent owners twice in relay listings (#24713) Capture the existing host-age expression at its original time, then use one call-local fresh owner clone for the length check and unchanged published owner field. * Reuse prepared statements for internal worker checks (#24573) * Reuse prepared statements across remaining Dispatch reads Use the existing schema-pinned complete Dispatch column list in five remaining read modules, enabling the existing bounded statement cache without caching results. * Reuse prepared statements for internal worker checks Use the existing complete Dispatch projection only for named-field internal consumers; preserve original wildcard reads for every returned failure snapshot. * Prepare the Activity search query once per filtered list (#24701) Lazily reuse a call-local search predicate inside the existing filter, using unchanged byte gates and thread text cache. * Avoid rescanning shared chunks in the plain Node build guard (#24717) Reuse a successful unchanged chunk-code verdict within one synchronous writeBundle invocation, while preserving traversal and current-code reads. * Select the latest stable release tag in one pass (#24786) Replace filter/sort/last selection with one traversal using the unchanged stable-tag regex and numeric comparator; update on equality to preserve the original last spelling. Both actual callers own private plain arrays from Git stdout. * Avoid Resource Manager renders when terminal removal leaves its inventory unchanged (#24806) Reuse the private React state object only when the existing single/bulk removal helper returns its identical inventory; keep every lifecycle revision, tombstone, known-ID write and helper invocation in its original order. * Cancel pending cursor updates when a terminal closes (#24566) Reuse the existing pane frame tracker to cancel owned deferred focus-class updates during existing cleanup, with disposed guards against late or reentrant delivery. * Cancel terminal preview fit frames when the preview closes (#24712) Reuse the existing frame tracker inside the box-fit owner, suppress scheduling after disposal and cancel pending frames at the start of the existing effect cleanup. * Skip new mobile toast work after feedback owner cleanup (#24759) Reuse the existing mobile mounted-ref lifecycle pattern at toast presentation entry, preserving admitted clipboard outcomes and every live animation/sequence/timer operation. * Release terminal side effects after they have been delivered (#24548) Clear consumed queue slots after successful apply or overflow carry; preserve backing identity across reentrant callbacks and report actual retained slots. * Release relay handshake timers when a host connection leaves (#24554) Reuse the existing first-frame finish pattern to remove stage-owned timer/message/close callbacks on receipt, timeout or close; check OPEN after successful async assignment verification. * Avoid restarting error timers on closed mobile relay links (#24567) Check the existing irreversible closed flag before scheduling the existing missing-close fallback timer. * Avoid restarting error timers after mobile relay pairing closes (#24568) Check the existing pairing owner closed flag before allocating its missing-close fallback alarm. * Delete unreturned clipboard cache files after a failed write (#24599) Reuse existing provider-copy best-effort deletion for a newly created clipboard cache file whose write fails before its URI reaches the caller. * Skip new legacy file inventories after mobile search cleanup (#24792) Reuse the existing mobile mounted-ref pattern only at legacy fallback entry after completed passive cleanup; preserve admitted work and all live search/authority/cache paths. Correct only the strict fully-unmounted inventory scenario and its sole golden. * Discard completed AI Vault cancellation IDs in the relay worker (#24820) Use the relay child's existing pending Set to ignore cancellation IDs after completion, matching its desktop sibling; preserve admitted active/queued cancellation and the bundled private sender's complete replies, errors and ownership. * Release completed updater setup timers and callbacks (#24920) Cancel completed deferred updater fallbacks in finally, clear only their exact pending callback, and drop the captured timer before the original fallback guard runs; preserve destroyed admission, successor ownership and updater/quit callbacks. * Check daemon idle state without building discarded terminal inventories (#24749) * Check daemon idle state without building discarded terminal inventories Replace the private idle predicate full public inventory with a host-local full pass that still reads every Session liveness getter in original order. * Remove overwritten daemon test mock assignment * Preserve the registered account ID when selecting UUID variants * fix(updater): keep macOS Orca open when background instances block updates (#24952) * fix(updater): guard macOS installs against running app instances * fix(updater): match native app blockers and preserve quit lifecycle * fix(updater): keep ordinary macOS quit on Squirrel's install-on-exit path Converting every quit with a staged update into quitAndInstall made Cmd+Q relaunch Orca, refused the quit when background instances existed, and hijacked app.relaunch()+app.quit() restart flows (profile switch, admin restart) into an update install racing the relaunched old app. Only Update & Restart runs the running-instance preflight now; the quit-without-install allowance is no longer reachable and is removed. * fix(updater): preserve quit intent through macOS staging * test(native-chat): explicitly model legacy published tab ownership --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> Co-authored-by: m4air <m4air@m4airs-Air.localdomain> --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> Co-authored-by: OrcaWin <alpha-eng@stably.ai> Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

The problem
In a native (structured) Claude chat, when a subagent stopped to ask for permission ("may I run this command?"), the subagent's own row kept saying working in the sidebar and in the composer's subagent strip. The session's row turned blocked as it should, but nothing showed which agent was asking, so with several subagents running you could not tell which one was stuck.
Why it happened: Orca never recorded which agent raised a Claude permission request. The Claude CLI names the asking subagent when it asks, but Orca ignored that, filed every request as the main agent's, and the Claude subagent decoder had no way to report a child as waiting.
Scope: native (structured) Claude chats. Agents running in a terminal already showed a subagent's request as waiting, and Codex subagents already read waiting in the sidebar and strip.
What this changes for you
worktree ps, mobile, the dashboard, notifications, Activityterminal read/worker-read)What Changed
agentID, the same id as that subagent's task); a request from the main agent carries none. If no agent is named, the owner of the tool call the request gates is used. The request's journal row is stamped with that agent's "produced by" fields, the same ones its other rows carry, which is why an answered request sits in the subagent's group.written()on the session's event sink, built on the existing write barrier). This keeps the subagent from reading waiting a moment before the session reads blocked, which happens when the card's write is still queued.Why this approach
How this compares with other agent apps
It matches the common pattern: the conversation shows one needs-input state whoever asked, the request is shown on the approval itself, and the subagent's block in the chat history is not marked waiting. Differences:
worktree psall read one child-record store, so the wait has to be written there; ordering it after the card keeps both consistent.Linked Issue
None. Part of the structured chat status work; builds on #22614 (merged). Based on main.
Screenshots
Live, on a remote Mac: a native Claude chat in the default permission mode (it asks), real Claude Code 2.1.288. Before = main
84a4e8a9b0a, after = this PR at42ae9b22d6d. The "Add a setup script" popup in every frame comes from the fresh test profile.A subagent asks for permission. Before: the session row reads Blocked, and the asking subagent reads working everywhere (sidebar child, strip
1 agent — working).After: the session row still reads Blocked, as on main. The asking subagent reads waiting in its sidebar row (
Waiting for input · Bash: touch qa-sub-marker.txt) and in the strip (1 agent waiting — needs approval). The chat-history subagent block readsworking, the same as main.After Allow, subagent group collapsed. Before: the answered
Bash … Allow · Resolvedrecord stays at the top level.After: it belongs to the subagent's group, so collapsing
Ran 1 subagent completedhides it.Same, expanded. Before: the record sits below the subagent entry, outside it. After: it is nested under the subagent's own
touch qa-sub-marker.txt ✓row.Cancelling a subagent's pending request with the card's X. While a card is up it takes the composer's place, so the card's X is the stop control (same on main). On the PR the subagent left
Waiting for input128 ms after the click and never came back. The session row went Blocked → Working → Done, the same labels in the same order as main. The cancelled record sits in the subagent's group like an answered one.The main agent asks for itself (unchanged): Blocked, no subagent row, no strip. Before / after:
Also checked live, on both builds: the session row's label sequence, sampled every 250 ms, was
Done → Working → Blocked → Working → Doneand never read waiting; the chat-history subagent block never read waiting; Activity recorded the same five states with no extra event.Testing
Automated, at the right level for each behaviour:
claude-subagent-permission-request.test.tsis the one integration test. It replays captured real-CLI frames (a foreground and a background subagent asking) through the real Claude adapter, event sink, on-disk journal and status feed. At every status update it checks that each waiting subagent has a pending card in the same journal, and that the session row equals a second host fed the same evidence with no subagent ever waiting: that is the one parity check that the session row is unchanged. Its 16 cases pin: the asking subagent's state from the request until it is answered, including a background subagent that keeps waiting after the main agent's turn ends; blocked from the moment the request arrives, with the tool it waits on; waiting only once the card is written, even while the write waits for the journal; freed when the request is allowed, denied, dismissed or withdrawn by Claude; waiting again if Orca fails to save the answer; the answered request in the subagent's group; attribution through the gated tool call when the CLI names no agent; a nested subagent's request carrying the same linkage as its own rows; the agent the CLI names winning over the tool call's owner; two subagents at once; the main agent and a subagent at once (including its dating); and the Claude process dying while a subagent waits.claude-structured-journal-prompt-cards.test.ts: a card never opens when its write is refused or fails, and a card handed back before its write lands opens once it lands.written()reports writes a close dropped as not landed (and writes already handed to the journal as landed).tc:nodeandtc:webpass; oxlint, the changed-lines gate and React Doctor are clean.Not verified: mobile, SSH and WSL live; an older paired client live; the real-Claude-CLI test suites (they run only with
ORCA_REAL_CLAUDE_CLI_TEST=1, which CI doesn't set; the captured-frame replay covers this path).AI Disclosure
Written with Claude (Opus).
Agent skill upstream boundary
docs/reference/agent-skill-sharing-upstream-boundary.mdand copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.Notes
Checklist
N/Awith reasonpnpm lint,pnpm typecheck,pnpm test, andpnpm buildpass (or CI will cover; local preferred)