Skip to content

Share the managed process lifecycle for structured providers - #25204

Merged
brennanb2025 merged 54 commits into
mainfrom
brennanb2025/acp-a6-provider-lifecycle
Oct 6, 2026
Merged

brennanb2025 merged 54 commits into
mainfrom
brennanb2025/acp-a6-provider-lifecycle

Conversation

@brennanb2025

@brennanb2025 brennanb2025 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor
Files Added Deleted Net
Test 16 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​822 $\color{#cf222e}{\Huge{\mathbf{−}}}$​140 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​682
Prod 14 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​497 $\color{#cf222e}{\Huge{\mathbf{−}}}$​253 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​244

ELI5

The problem. A structured chat runs the agent (Claude, Codex, and soon Grok) as a process Orca starts and later stops. Stopping one safely takes several steps: close its input, wait, force it if it doesn't exit, check whether any processes it started are still running, and retry if Orca can't confirm the stop. Claude and Codex each had their own copy of these steps. The Grok chat over the Agent Client Protocol (ACP, #25225) would have added a third, and every copy is a place to get a step wrong. In the third copy, for example, the ACP adapter never read the field that says "the agent's leftover processes could not be confirmed gone", so that warning was lost.

What changes for you. Nothing you can see. Claude and Codex chats start, stop and recover exactly as before, with the same timings and messages. This is groundwork, so the Grok chat (and any later agent) starts from one stop, rather than a third copy of the steps. The Grok branch still has to read the descendant answer described below (#25225).

How. One shared "managed provider process" now starts the agent, sees it exit, stops it, and retries a stop it couldn't confirm. Every stop reports two separate answers about what Orca actually observed, using Orca's usual three words (live, unverifiable, exited), or "no observation" when the stop looked at nothing.

  • The agent's own process ("root").
  • The processes it started ("descendants").

What Changed

  • One stop result for every agent. A stop returns { root, tree }.
    • tree is null when that stop made no check of the descendants.
    • When Orca had to force-kill an agent that has no checker of its own, tree is what the force-kill observed:
      • exited only when the processes it captured were checked and found gone;
      • live / unverifiable as that check found them;
      • null when it signalled without observing anything. That covers the Windows tree kill (whose result can't be read), a macOS/Linux process table that couldn't be read, and a process-group kill.
      • A delivered kill is never reported as an exit.
    • The old optional side field (teardownAccepted) is gone. The answer is a required part of the result, though a caller still has to read it.
    • The three-word type is declared once, reusing the existing DescendantTreeVerdict instead of a copy.
  • Codex's "leftover processes not proven gone" report keeps its meaning. It is now read as "root exited and tree unverifiable or live", which fires in exactly the cases the old flag did. A new Codex test pins each case.
  • The shared process owns stderr. It drains the agent's error output, so a full pipe can never block the agent, and keeps the same last 8 KiB for error messages. Claude and Codex dropped their own copies.
  • Root-only agents get defaults. An agent with no descendant checker of its own (Codex, and Grok next) needs to pass nothing for the stop policy (1.5 s wait on Windows; the 5.5 s supervisor wait on macOS and Linux; then 1 s after a force-kill) or for the rule "done once the root is gone". Only Claude overrides them.
    • There is one "already exited" check, inside the shared stop. Its answer is recorded when no earlier stop ran. An earlier stop's findings are never overwritten by a later no-op stop.
  • A spawn that failed (for example, a missing binary) is never treated as "the agent's process was seen to exit".
    • The shared process has rootExitObserved, which is never true for a failed spawn.
    • Claude's claudeRootExitObserved agrees, whatever order its callers check things in.
  • Unchanged from the earlier rounds of this PR:
    • Root and descendant answers stay separate.
    • A running agent reports live.
    • An unconfirmed stop retries on the same process.
    • A launch under the Unix supervisor (the helper that stops the agent if Orca dies) is rejected before spawning if its wait is shorter than the supervisor's own 5.5 s.

For the ACP branch (#25225)

These are in-memory API changes only. No stored data or messages between devices change.

  • Read close().tree instead of teardownAccepted. Log it when it is unverifiable or live, as Codex does. null means the stop observed nothing about the descendants, not that they are gone.
  • Drop the local stop policy, completion rule and "already exited" pre-check. The defaults cover them.
  • Use managed.stderrTail() instead of its own stderr buffer.
  • Use rootExitObserved where "the agent's process exited" is meant, so a failed spawn doesn't read as an exit.

Why

Making the descendant answer a required part of the one result, rather than an optional field beside it, makes it harder for a new agent to forget, and the answer only claims what was observed. The common pattern also returns a descendant answer from every stop. Moving stderr and the default policy into the shared process removes three copies of each, and gives each agent fewer things to get wrong. Modelling a failed spawn separately keeps callers correct by construction rather than by the order of their checks.

Relationship to other PRs

Differences from the common pattern

  • Intended: a stop that can't be confirmed can be retried on the same process, rather than keeping its first answer forever. fix(native-chat): the agent's exit ends its record, and an unconfirmed stop is joined instead of held #24862 needs a later request to be able to retry the held process without starting a second one for the same chat.
  • Intended: watching for the agent's exit belongs to the shared managed process, not to each agent connection, so every connection shares one observer and one stop.
  • Intended: Codex starts a stop by closing input; Claude also signals the Unix supervisor straight away. The supervisor then sends the later termination signals itself, and the existing waits keep it alive long enough to finish.
  • Intended: descendant checks keep the existing process-identity checks and snapshots, which protect unrelated processes that reuse a process number.
  • Temporary: only Claude has a real descendant checker (snapshot plus identity, in claude/). Codex and Grok get the simpler force-kill fallback. It reports what it observed. On Windows, and when the macOS/Linux process table can't be read, it reports tree: null ("no observation"). For an unreadable table, Claude's checker reports unverifiable instead.
    • Why it's kept for now: Codex's "leftover processes not proven gone" check treated an unreadable table as fine. Reporting unverifiable would newly turn those Codex stops into errors, which is a behavior change for a lane that has already shipped.
    • What it costs: a reader can't tell this null from a graceful stop where nobody needed to look.
    • Follow-up:
      • move that checker into the shared provider-process code;
      • make it the only teardown;
      • report an unreadable table as unverifiable, taking Codex's behavior change deliberately.
    • This is not in this PR because it changes when Codex reads the process table during a stop, which is a behavior change for a lane that has already shipped.
  • Temporary, inherited: on Windows, the agent's processes are still not cleaned up automatically if Orca's runtime dies. The planned Windows process-containment follow-up removes this; this PR keeps today's Windows tree cleanup.

Linked Issue

Foundation follow-up from the #24989 review, stacked on #24862 and #24989. Internal work; no new issue opened.

Visual Proof

N/A: no UI or interaction change.

Testing

  • I manually tested these changes locally
  • Automated tests added/updated, or explained why not below

What I verified / didn't

Verified on macOS at 52becf72cb2:

  • 62 explicit test files: 617 passed, 1 skipped (Windows-only). This includes the real-process Codex teardown integration test. Fakes and plain Node child processes only; no real agent CLI and no app launched. The files are:
    • the lifecycle, connection and supervisor suites;
    • every test file that imports a changed module;
    • the native-chat stop and exit suites.
  • New tests:
    • managed-provider-process-root-only.test.ts, which covers:
      • a graceful stop reports tree: null;
      • a forced stop writes whatever the fallback teardown observed (exited, live, unverifiable or null) into tree, and a repeat stop doesn't re-run the teardown;
      • the default 1.5 s wait;
      • the recorded "already exited" answer;
      • the bounded stderr tail;
      • a failed spawn is never an observed exit.
    • A Claude test that a failed spawn is never "root exit observed".
    • managed-provider-process-fallback-tree.test.ts, which runs the real fallback teardown with only operating-system calls faked. It covers:
      • a Windows forced stop and an unreadable macOS/Linux process table both report tree: null;
      • exited appears only after captured descendants were verified gone.
    • Teardown unit cases for live, unverifiable, an empty process group (no observation) and a child that never spawned.
    • codex-app-server-connection-tree-unproven.test.ts: Codex's diagnostic for each teardown answer.
  • Scoped typechecks (tsconfig.node, tsconfig.tc.cli): the only errors are the known missing-module errors for stream-json / stream-chain, in files this PR doesn't touch, caused by a stale local install.
  • Lint and format pass on the changed files. The repository's changed-code quality gate passes locally. The reliability-gate manifest check passes.

Not verified:

  • Real agent CLIs, Electron, native Windows or Linux runs, and SSH.
  • CI on the previous head b629fcf3a07 failed one check: the changed-code quality gate flagged a test cast without a SAFETY: rationale. That is fixed in eb94d9b1615, where static analysis and typecheck passed. CI on 52becf72cb2 was still running when this was written, and CI is the authority for the full typecheck and test shards.
  • Passing tests show nothing regressed; they don't prove the design is right.

Review

Each item from the previous review round is either fixed above or labelled Temporary in the differences list. The shared provider-process code has no Electron dependency.

Agent skill upstream boundary

  • Not applicable; no skill source is copied or translated.

Notes

No new UI, stored data, runtime capability or remote message. Folder workspaces and SSH are unaffected: process work stays on the machine that runs the agent.

Checklist

  • This PR is small and focused
  • I explained what changed and why (ELI5, the user-facing before/after, the mechanism, and why over the alternatives)
  • Before/after screenshots or videos attached for UI changes, or N/A with reason
  • Self-reviewed for correctness, security, and performance
  • Cross-platform, SSH/remote, and path/shortcut impact considered (or N/A)
  • pnpm lint, pnpm typecheck, pnpm test, and pnpm build pass (CI will cover; changed-file checks, scoped typechecks and explicit tests passed locally)

…se, and bookkeeping after it never reads as unproven

- Both connections report the root process's exit once, with `expected` set when a close had
  begun. A close that came back unproven and whose root exits later is finished by the adapter,
  and its end reaches the host like any other.
- A Claude close whose resume-point write fails after the exit was proven, and a Codex close whose
  terminal row is refused, now end the session and report the failure, instead of keeping a dead
  child indexed as if its exit were unproven.
- A Codex close whose forced tree kill can't prove the descendants gone but saw the root exit
  reports the descendants and counts the root exit.
- Every child exit with an identity, expected or not, is forwarded to the host.
…p is the child's own close, which everyone joins

- The host keeps no stored "stop still owed" record any more. A stop begins the child's close
  (`child.close`), which lives on the child and ends with it. A second Stop, the idle reaper,
  quit, a send and an option/answer/goal/rewind all join that close instead of retrying a
  separate obligation.
- A caller waits on the close only as long as the step deadline; the close itself is never
  abandoned. A proof that lands after every caller stopped waiting reaches the host as the
  adapter's report of that exit, which ends the record through the same handler.
- Once the exit is proven, draining, settling, the lease release and the adapter's
  acknowledgement are each attempted and reported on failure; none keeps the child on record.
  A start, and the handle's close, write a release that failed from this host's proof of that
  exit, so a failed write never refuses a send.
- A start that meets a close still unverifiable is refused with `previousExitUnverifiable`, so the
  queued message is rejected with a send-again reason; nothing is held and nothing starts beside
  the old process.
- The idle sweep goes back to idle reaping only.
- Removes #24333's retry entry points, the wait row and its hold rule, the ask/failure cursors on
  the stored record, and the stop's own wake.

Tests replace the #24333 unproven-stop test: a send joining an unproven close and an in-flight
one, a late proof past the caller's bound, a root exiting after its close gave up, a proven exit
whose resume-point write and lease release both failed, an unverifiable close rejecting the send
and refusing an option change, a surviving descendant, quit and the idle reaper; and Codex's
unverifiable, late-exit and joined-close cases.
… unverifiable says so, and to send again

The start failure for a refusal with reason `previousExitUnverifiable` reads "Orca couldn't
confirm Claude's previous process ended. Send your message to try again." instead of "Claude
couldn't restart." The status-row kind and the refusal reason stay in the shared lists for rows
and hosts that still carry them; the catalogs keep one sentence for both.
…t follows it is logged

- A Claude close resolves as soon as the root's exit is proven: the session ends and its `ended`
  report goes out then. Saving the resume point runs afterwards and a failure is logged, so a slow
  or hung write never reads as an unproven exit or keeps a dead child on record.
- A root that exits after its close came back unproven finishes that close through the same path
  as any close, so the session's child work is published as ended (background tasks and subagents
  no longer stay shown running for a dead agent), and a failure there is logged.
- Codex logs a refused final row, and reports a root exit whose forced tree kill could not prove
  the rest of the tree gone the way Claude does, so the host logs it and blocks nothing.
- Both adapters take the host's logger for this bookkeeping.
…ts on the adapter's own close

- One exit handler (`structured-agent-session-child-exit`) ends a child's record for an exit
  expected or not. `expected` only changes what the chat is told: the stop's cause, its end at the
  stop's ask, the settlement id, and no crash outcome row. The lease release keeps the exit's
  evidence; the handoff guard, lifecycle barrier, sink release and adapter acknowledgement apply to
  both. A Claude journal-sink failure ends in the same step as its stop, as Orca's own fault.
- Joining a close is asking the adapter, whose close is memoized while it runs and bounded by its
  own kill escalation; the host keeps no attempt of its own and no 10 s caller bound. An ask after
  a close came back unproven runs the stop again.
- A close's end is stamped where its stop was asked for (a repeated ask moves it), so the closed
  chat and failed start checks order a message accepted meanwhile after it.
- A start refused because the old exit is unverifiable rejects what was queued in the same step.
- The end of a close the host asked for no longer waits on the cross-session recovery chain.
- The kill no longer waits for the stop event's write; the journal writes rows in order.
…f its reason

A crash's reason can carry kilobytes of the provider's stderr, and a lease whose death detail is
over 512 characters fails the store's own check. The exit handler cut it, but the release a start
or the chat handle's close re-derives did not, so after a crash whose own release failed every
message was refused as not resumable until restart. The record's builder now cuts the detail to
the record's bound, so no writer can hand it one too long.
…dy wrote

Once the root's exit is proven, the close still waited for the SDK's output reader to end. Something
outside the process tree that holds the output open would keep that close, and every send, Stop
and quit joining it, waiting with no bound. The wait is now bounded; past it the close resolves as
proven and the open output is logged.
…a closed it for

The connection reports the app-server's exit inside the close that ends it, so that report ended
every Codex close and replaced the close's own reason (for example, a provider frame that could not
be recorded) with the connection's stderr text in the ended record and the lease's exit evidence.
The session now records Orca's close with its reason, and the exit it ends keeps that reason. The
test connection reports its exit inside close the way the real one does.
Every exit now wakes delivery, and teardown drained exit recovery before it stopped delivery, so an
exit settled in that window could start a fresh agent that teardown then killed. Teardown stops
delivery first; queued messages wait for the next launch.
…ller's wait

The caller's bounded wait was removed; the comment describes the close as it is now.
…ext ask kills again

When a close's kill leaves the agent's root running, the host now logs it. Tests pin what a later
ask does: each connection runs its whole stop again (Codex sends SIGKILL a second time), refuses
input meanwhile, and proves the exit once the kill takes.
…dn't stop it

The host reaches an unverifiable verdict only after its own kill left the agent's root running, on
the machine that runs the agent, so the sentence now says that: "Orca couldn't stop {agent}'s
previous process." The refusal reason, failure kind and wire shapes are unchanged. The host test
also checks the failed kill is logged.
…ved the kill

The host's close runs where the agent runs, so lost contact never yields this verdict; the comment no longer says it does.
…er met it

The log added at the close fired beside a Stop's own failure report for the same event. A stop
still reports it through its failure; a send or option change refused over it now logs it at the
refusal, the only place it is otherwise invisible.
…p-exit-ends-record

# Conflicts:
#	src/main/codex/codex-app-server-connection.test.ts
…p-exit-ends-record

# Conflicts:
#	src/main/native-chat/agent-session-wire/structured-agent-session-claude-unproven-stop-send.test.ts
#	src/main/native-chat/agent-session-wire/structured-agent-session-eviction.test.ts
#	src/main/native-chat/agent-session-wire/structured-agent-session-unexpected-exit.ts
…p-exit-ends-record

# Conflicts:
#	src/main/native-chat/agent-session-wire/structured-agent-session-claude-unproven-stop-send.test.ts
#	src/renderer/src/i18n/en-runtime-required.json
#	src/renderer/src/i18n/locales/en.json
#	src/renderer/src/i18n/locales/es.json
#	src/renderer/src/i18n/locales/fr.json
#	src/renderer/src/i18n/locales/ja.json
#	src/renderer/src/i18n/locales/ko.json
#	src/renderer/src/i18n/locales/zh.json
#	src/shared/agent-session-failure-copy.ts
…inly

The rejection now reads "Couldn't stop {{agent}} from before. Send your message again to try once more."
This kind has its own send-again step; every other failure keeps "Send your message to try again."
…cess' into brennanb2025/acp-a6-provider-lifecycle
@brennanb2025

Copy link
Copy Markdown
Contributor Author

Review summary (head 52becf72cb2)

The problem. Claude and Codex each had their own code for starting the agent process, watching for it to exit, and stopping it only once that exit was proven. The two copies had already drifted (different waits, different signals). The Agent Client Protocol lane (Grok, #25225) would have added a third copy. That stop path is where Orca's subtle process bugs have come from: a stop reported as finished while the process still ran, or a supervisor killed too early so its children were orphaned.

What changes for users. Nothing visible. Claude and Codex chats start and stop with the same timings and order as before. This gives every provider one shared way to start and stop its process.

How it fits the stack. It is stacked on #24862 (the stop and exit handling it builds on) and merges in #24989 (the shared provider-process folder). Only this PR's own commits were reviewed. #24862's own changes were not touched.

What the review rounds fixed

  • A stop now reports the main process and its child processes as separate results. The main process is running, exited, or never started. The children are running, exited, unverifiable, or "nothing was checked". Before, one combined result said "unverifiable" for a healthy running process. For Codex and the coming Agent Client Protocol lane, the child result was a side flag that was easy to miss, and the Agent Client Protocol branch already ignored it.
  • The forced cleanup reports children as exited only when it actually saw them gone. A Windows kill command whose result isn't read, an unreadable process table, or an empty process group now report "nothing was checked" instead of claiming they exited.
  • A process that never started no longer counts as "exit observed".
  • The shared managed process now drains the agent's error output and keeps the last 8 KiB, so a new provider can't forget to drain it (an undrained pipe blocks the agent). It also supplies the default stop policy, so only Claude overrides it.
  • The minimum wait for a supervised process is enforced in code, not just stated in a comment. Options that existed only to keep old tests unchanged were removed, so the real-process Claude stop test runs the production path again.

Temporary differences, labelled in the description: the child-process cleanup code still lives in the Claude folder, because moving it touches a file #24862 changes, and making it the only cleanup would change Codex's stop behavior. A follow-up moves it into the shared folder and adopts it everywhere. Until then, an unreadable process table reads "nothing was checked" for Codex and Agent Client Protocol, where Claude's checker says "unverifiable".

For #25225 (Grok): it compiles unchanged, but should drop its own copies of the stop policy and error-output buffer and read the child result from close(). The exact call sites are in the review notes.

What was verified

  • Three review rounds plus a readiness checklist. A side-by-side test of the previous and new cleanup code over 16 cases shows Codex's "child processes not proven gone" result, the signals sent and the kill log unchanged.
  • 62 explicit test files (617 tests) on the final head.
  • CI on 52becf72cb2: typecheck, static analysis, packaging (Linux and Windows), cross-version wire, the real Codex contract job and 3 of 5 unit shards pass. Shards 1 and 4 fail only in renderer translation tests (source-control-discard-localization.test.ts, NativeChatSupportedAgents.test.tsx). Those fail the same way on unrelated branches and come from current main; main has since fixed them.

Not verified: no live app or real agent; Windows and Linux behavior covered only by platform-parameterized tests, not real hosts.

main #24864 (a Stop binds only the turn it stopped): recordStopEvent now resolves to the settle a
person's close opens, and the stop closes it once the child's end is done. Re-applied on this
branch's child close: the close's `recorded` carries that settle (null when no row is written),
and the stop closes it in a finally around the restart snapshot and the close join (which awaits
the exit settlement), proven or not. A second stop joining the same close closes it again, which
is a no-op.
…he turn its child end cuts

On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.

Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).
This branch carried an earlier copy of #24862's commits; main now has its final, squashed form.
Main's #24862 wins for #24862's code; this branch's own changes (shared managed provider process,
close result, stderr tail) are re-applied on top:
- host-lifetime.ts, host-types.ts, stop-settle-binding.test.ts: main's version. Final #24862
  already holds this branch's earlier merge resolution (the close's recorded settle) and its later
  fix (a repeated close reopens the settle), with its own test for it.
- claude-stream-json-connection.ts, codex-app-server-connection.ts: this branch's managed-process
  version; main's side of these files equals the earlier #24862 copy this branch replaced (checked
  with a three-way merge over old-#24862-on-main as base). Drops a duplicate processTreeUnproven
  getter the plain merge left outside the conflict.
@brennanb2025

Copy link
Copy Markdown
Contributor Author

Merged current main into this branch twice (now 636f6f1bdc4).

brennanb2025 added a commit that referenced this pull request Oct 5, 2026
No conflicts: A6's new head adds main's newer commits on top of what D3 already carried.
… and #24991)

The six conflicts were this branch's copy of the provider-process extraction (#24989)
against main's squash of it, which is byte-identical. Resolved by re-merging with
old-main plus that extraction as the base, so the result is main plus exactly this
branch's own change.
…e opaque handle

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.
…d-turn test

Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.
…p-a6-provider-lifecycle

Main's Electron-import ratchet still allowed src/main/provider-process to be absent after #24989
created it, which fails its own retirement test on every branch; #25710 makes the lane required.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
Conflicts were only in files another base owns: A6's Claude/Codex/provider-process files take A6 #25204's head, the journal schemas take C5 #25181's head, and main's Stop-note test takes A3's (A3 now fixes it for its own API, so D3's copy of that fix is dropped).
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
@brennanb2025

brennanb2025 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Earlier CI merges missed main's restored provider-handle import and the test-only parser dependency still needed by old release checkouts. Merged main 08a970a3f117; the branch's own changes are preserved, the submission-position test matches main with exactly one used import, and main's package/lock fix is inherited through the merge with no branch-authored dependency edits. The prior broad run passed 993 tests; fresh-main checks passed 118 tests, with remaining local failures limited to stale stream-json/stream-chain dependencies, and scoped oxlint passed; new-head CI is pending.

@brennanb2025
brennanb2025 marked this pull request as ready for review October 6, 2026 03:30
@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 2cefad0c-930c-4241-a22b-becbe848a1b9
📥 Commits

Reviewing files that changed from the base of the PR and between 08a970a and 5df4db2.

📒 Files selected for processing (30)
  • src/main/claude/claude-agent-sdk-exit-proof.test.ts
  • src/main/claude/claude-agent-sdk-exit-proof.ts
  • src/main/claude/claude-agent-sdk-process-spawn.test.ts
  • src/main/claude/claude-agent-sdk-process-spawn.ts
  • src/main/claude/claude-child-exit-proof-fixture.ts
  • src/main/claude/claude-child-exit-proof-ladder.test.ts
  • src/main/claude/claude-child-exit-proof-ladder.ts
  • src/main/claude/claude-stream-json-connection.test.ts
  • src/main/claude/claude-stream-json-connection.ts
  • src/main/claude/claude-structured-event-delivery.ts
  • src/main/claude/claude-structured-session-acquisition-processless.test.ts
  • src/main/claude/claude-structured-session-adapter.ts
  • src/main/claude/claude-structured-session-close.test.ts
  • src/main/claude/claude-structured-session-close.ts
  • src/main/claude/claude-structured-session-startup-state.ts
  • src/main/claude/claude-supervised-stop.integration.test.ts
  • src/main/codex/codex-app-server-connection-tree-unproven.test.ts
  • src/main/codex/codex-app-server-connection.test.ts
  • src/main/codex/codex-app-server-connection.ts
  • src/main/codex/codex-structured-session-cancel.test.ts
  • src/main/provider-process/managed-provider-process-fallback-tree.test.ts
  • src/main/provider-process/managed-provider-process-grace.test.ts
  • src/main/provider-process/managed-provider-process-root-only.test.ts
  • src/main/provider-process/managed-provider-process.test.ts
  • src/main/provider-process/managed-provider-process.ts
  • src/main/provider-process/provider-process-close.ts
  • src/main/provider-process/provider-process-teardown.test.ts
  • src/main/provider-process/provider-process-teardown.ts
  • src/shared/child-process/retryable-process-exit-proof.test.ts
  • src/shared/child-process/retryable-process-exit-proof.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 4 remain after this review.


📝 Walkthrough

Walkthrough

The changes add a managed provider-process lifecycle that tracks root exits, processless spawns, stderr, and descendant-tree verdicts. Claude and Codex connections now use that lifecycle for spawning, exit handling, and close operations. Teardown reports whether the process tree was verified or whether its status remains unknown. Claude structured-session startup waiting and event delivery now use dedicated helpers.

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to 5df4d

No merge-blocking behavior change was established. The PR is mergeable after normal checks.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 5df4d

The reviewed integrations preserve their existing shutdown guarantees and launch permissions. The main risk is applying shared lifecycle rules consistently across integrations and platforms; coverage of those wider behaviors remains incomplete.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — A defect in the shared lifecycle could affect both integrations and processes launched beneath them at their existing execution privileges. The reviewed construction sites retain the same launch inputs and child ownership; they do not establish a new remotely callable entrypoint or broader execution authority.

Trust Boundaries and Controls

  • observed — Claude still supplies an explicitly constructed child environment that excludes inherited account configuration and applies its authentication filtering. The managed wrapper receives that environment explicitly and routes launch through the existing inherit, overlay, then strip rule. SDK abort signaling remains outside the retained shutdown owner.

Resilience and Maintainability Implications

  • observed — Fallback teardown now exposes uncertainty instead of equating a delivered signal with observed descendant exit. Windows taskkill and a missing POSIX snapshot return no descendant observation; other teardown failures return unverifiable. Codex preserves its prior root-only completion and diagnostic treatment of these cases, so these observation limitations are not established as a new loss of containment.

Hardening Proposals

  • proposed — For future integrations, consider binding descendant proof to the managed child at construction or carrying an explicit owner identity. The current tree interface permits arbitrary pairing, but the inspected Claude caller constructs its verifier from the matching child; no mismatched production call was found.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 45 functions across 30 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes the main change: sharing the managed process lifecycle across structured providers.
Description check ✅ Passed The description covers the required sections, explains the change and its rationale, describes testing and limitations, and marks visual proof as not applicable with a reason. It also explains why no …
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@brennanb2025
brennanb2025 merged commit 3285f82 into main Oct 6, 2026
63 checks passed
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
Main squash-merged A2 #25076 (db71a89) and A6 #25204 (5df4db2), both already in D3; resolved
against those heads, so only zh.json conflicted: structuredCopy keeps D3's agent-neutral meaning with
main's new 智能体 term.
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
…agent-tree-job-object

Main moved the provider close ladder into the shared managed-process close
(#25204) and kept the Windows creation-time refusal in Codex launch resolution.
Resolved by keeping that refusal removed, taking main's pinned config variable,
and re-expressing this branch's Windows close rules on the shared close: the
Windows teardown reports taskkill's verdict ('exited' on exit 0, else
'unverifiable'), and a Claude root that leaves on its own after stdin end with
no forced reap on its tree is a proven close (Claude's close policy on win32).
brennanb2025 added a commit that referenced this pull request Oct 6, 2026
Main's #25204 moved every structured provider's close onto one shared close
(closeProviderProcess, driven by a close policy) and spawn (spawnManagedProviderProcess).
That is the same concept as PR1's stopSupervisedProvider and requestProviderClose, so PR1's
copies are removed and its callers move onto main's:
- The Codex connection and the Claude spawn and exit proof take main's versions.
- The close request a gone owner gets is derived from the provider's close policy inside
  spawnManagedProviderProcess (signalSupervisorOnClose -> 'stdin-end-and-sigterm', else
  'stdin-end'), so one source drives both the owner's close and the supervisor's.
- stopSupervisedChildProcess (agent one-shots) runs main's closeProviderProcess with the
  root-only policy and the caller's close request.
- The Codex close-request constant and its CLI build entries go: nothing uses them now.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant