Conversation
* test(e2e): await renderer recovery after worker exit * test(e2e): publish worker recovery through authenticated hooks * test: keep retired background worker dormant before activation
…yai#23678) * test(claude): expect typed cancellation in queued-send settlements * fix(ci): respect disabled terminal links and await browser recovery * test(e2e): give legacy close client a profile authority * test(e2e): account for frame pacing in pointer latency budgets * test: compare pointer timing on isolated visible display
…d shared hooks dir (stablyai#23500) * fix(opencode): stop OpenCode 2 loading a stale plugin from the retired shared hooks dir Before 1.4.209 Orca pointed OPENCODE_CONFIG_DIR at <userData>/opencode-hooks/shared and wrote a server()-only status plugin there. 1.4.209 moved the plugin to OpenCode's global config dir and 1.4.210 added the v2 setup() export, but nothing rewrote the old file. Shells, daemon-persisted panes and OpenCode 2 background services that still carry that OPENCODE_CONFIG_DIR load only that dir under OpenCode 2 (it replaces the global dir), so the v2 loader rejects the stale plugin with "Plugin must export a default definition with an id and an effect or setup function" and pane status dies. - Refresh the plugin in the retired shared dir (only when it already exists and its content differs) so OpenCode processes started later from old shells load the dual v1/v2 export. Runs on OpenCode pane spawns and on any spawn that inherits the retired dir, even with agent status hooks off. - Drop an inherited OPENCODE_CONFIG_DIR / ORCA_OPENCODE_* marker that points at the retired dir when building a new pane env, so new panes use global discovery. Limitation: an OpenCode 2 background service already running from an old pane keeps its cached copy of the stale module even after the file is rewritten (verified with opencode2 v2.0.18). It must be restarted (`opencode service restart`); a restart from a new Orca pane then picks up the global config because the env is stripped. * fix(opencode): harden legacy plugin repair and inherited config cleanup * fix(opencode): preserve daemon-owned user config during legacy cleanup * fix(opencode): sanitize inherited sources and repair unseen legacy copies * test(opencode): update shared PTY mocks for legacy repair * test(opencode): annotate shared repair mock signature --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
… status store (stablyai#22553) * refactor(native-chat): the Codex acquire names its turn-boundary methods as a set Behavior-neutral: the same two methods stamp receipt time. Keeps the file under the size limit once the child-work sink lands. * feat(native-chat): Codex sessions write their subagents into the host status store A Codex child thread and each persistent command become host child records, fed through the same delivery, ingest and reducer the Claude lane uses. The child's own turn decides it: turn start is live, turn completion settles it with the outcome Codex reports, and a follow-up turn reopens the same record as a new run. Its open tool call, last message, usage and waiting-on-user flag come from its own thread's frames. A parent turn ending settles nothing. * fix(native-chat): close a Codex child's tool call by its item id alone A completion frame need not restate the tool it ran, so reading the tool name before closing left the call open and the record naming a finished tool. * test(native-chat): pin the Codex child-work evidence and every hop to the host's records Child turn start/end/follow-up, open tool call, last message, usage, waiting, the persistent command a child owns and its monitoring display, a primary turn end settling nothing, and session end. End to end through the real adapter: evidence after the journal and the legacy republish, and the parent state the records imply equals today's at every frame of a scripted session. Through the production runtime: a Codex session's child work reaches the status sink under its own address, and a provider exit ends it there. * test(native-chat): a Codex child's new run never inherits the last run's open call * test(native-chat): a Codex session with no child-work sink holds no evidence * test(native-chat): deliver a Codex child's announcement twice, as Codex does, before counting edges * refactor(native-chat): hand the Codex producer's pending edge over directly * fix(native-chat): name every Codex turn state in the outcome map; type the runtime test's fake opener * fix(native-chat): a Codex child's turn ends on the error that ends it, or on its thread closing Codex can end a child's turn with no turn/completed: an error it will not retry is that turn's own end (the verdict the transcript already settles the same turn on), and a closed thread ran its last turn. The executions, the one owner of child turn state, now end the turn on both, so the strip drops the child and its record settles (failed, or unknown for a close) together, instead of reading working for the life of the session. A systemError status is not an ending: Codex raises it for errors that leave the turn running. A child fact whose frame names no turn now belongs to the turn the child is running, instead of counting for every run. * test(native-chat): a Codex child's turn ending by fatal error or thread close settles strip and record together * test(native-chat): the Codex parity script reads a waiting child through the shared fold's waiting arm * test(native-chat): a Codex child row's journal attempt is its record's generation The journal numbers a Codex child's runs by the turns it observed on the child's thread; the host record numbers them by the runs its evidence opened. Both are keyed by the child's own turn id, so they must agree run for run, including when Codex reports the child's first turn before the spawn that announces it. * test(native-chat): a Codex session's end settles its live children and keeps the ended ones The host no longer erases a session's children when its provider goes away: a child still running settles with an outcome nobody reported, and a child that had already ended keeps what it said. The producer tests now expect exactly that, from the close path and from an unexpected exit. * fix(native-chat): a Codex subagent's shell is its open tool until the process exits Codex runs every agent shell through unified exec, so every subagent shell arrives with the source the persistent-command tracker keys on. The producer skipped those items, so a working subagent never named its shell, and an approved command (started on the approval path, completed from unified exec) stayed its open tool until the turn ended. The tracker still records the process separately, so a command that outlives the turn reads as monitoring. * fix(native-chat): a Codex shell becomes a subagent's own work only once it outlives its turn Codex runs every agent shell through unified exec and never says when one is left running, so the producer turned every shell, even a millisecond `rg`, into a command record the moment it started. Each settled into the session's pool of 32 settled records, so a busy turn evicted a finished subagent's record (its outcome row would vanish) and listed dozens of finished shells beside it. A command now becomes a record at the first turn boundary of the thread that launched it while its process still runs: until then it is the agent's open call. A shell that exits within its turn never becomes a record. * refactor(native-chat): child records keep every settled child and can be removed outright Settled child records now stay until the host drops the session's row; the 32-record trim is gone. A producer can say work stopped with nothing to report, and its record (and the handles it answered to) goes instead of settling. Evidence stays host-internal: the producer and the store share one process. * fix(native-chat): a Codex command is live work from its start until its process stops The command tracker is now the one owner of a Codex command's lifetime. It admits every command whatever `source` Codex tags it with (the approval path starts one as `agent`), and ends it when its process exits, when its thread closes (Codex stops the processes first, so no exit ever arrives), or when the session ends. The producer mirrors that one-to-one: a live record from the start, removed when the command stops, never settled. This removes the turn-boundary rule: a command that was only recorded at its turn's end left the parent reading done for one publish when the main agent's turn ended with a shell still running. The parity script now checks the parent at every journal write, not only at frame end. * fix(native-chat): a Codex command whose approval its turn abandoned never ran Codex starts an approval's command item before it asks, and when the turn ends with the question unanswered (the user stops at the approval), it drops the question and never completes the item. The command tracker admitted that start as a running process, so the strip kept a phantom command row and the session row read working until the session ended. The prompt registry, which owns which approvals are still unanswered, reports the command approvals a turn ended without; the tracker ends those commands with the frame that ended the turn. An answered approval keeps its command. * test(native-chat): start the Codex child-work runtime test without the removed hold Main no longer has host.hold: creating the session starts its child, and nothing a viewer does keeps it running. The test attaches and asserts the one child that attach started, then drives it as before.
stablyai#22568) * feat(orchestration): inject the Orca session id into structured children and let the CLI act as it Every structured session's child (native Claude, native Codex, and the terminal view) carries ORCA_AGENT_SESSION_ID and reaches the Orca CLI. The CLI sends the id in the orchestration envelope; when present it is the caller, and a caller flag naming anyone else is refused before any request. The id is stripped from inherited PTY env and from the SSH host-CLI passthrough, and crosses into WSL so the host can refuse the cross-host claim. * test(orchestration): pin session id injection for native Claude, native Codex, the terminal view, WSL, PTY inheritance and SSH * test(orchestration): pin one caller precedence rule across every CLI verb that names its caller Adds the per-verb table (flagless acts as the session; a conflicting --from or --terminal is refused before any request; the session's own spellings are accepted), the enumerated guess population with its positive control, the structured worker's own handle, the identity-less refusal for an older child, the unchanged terminal agent, and the envelope. dispatch-show's --from only fills preview text, so it passes through unfenced and a session's flagless preview names the address the real dispatch writes. * refactor(orchestration): keep the identity-less marker reader to the marker; the id is checked first * test(orchestration): pin that a host refusal of the session surfaces verbatim from the CLI * fix(orchestration): keep the identity-less marker beside the id for CLIs that predate it A CLI older than the id, reached through a global install when a shell rc resets PATH, would otherwise guess a sibling's terminal in a chat that no longer carries the marker. It refuses on the marker instead; a current CLI checks the id first, so the marker never makes a session with an id identity-less. * fix(orchestration): refuse a conflicting --from on gate-list and task-list scoped by --run A --run listing needs no caller, so both handlers skipped the resolver and a --from naming another actor was dropped silently under a session. The conflict check now runs on that branch too; terminal callers are unchanged. * fix(orchestration): name this app's CLI by absolute path for a structured session's login shells A provider can run each command in a login shell: Codex runs zsh -lc, and the profile rebuilds PATH, putting a global install (possibly an older Orca) ahead of the directory Orca prepended. ORCA_CLI_COMMAND, which an agent resolves the CLI from first, is now the absolute launcher in that directory (the native launcher on Windows), so no shell's startup files can swap it. The PATH prepend stays for shells that read no profile. Found by the live coordinator run of the next PR. * test(orchestration): pin a structured worker's CLI command as this app's absolute launcher * test(orchestration): run the zsh login-shell arm in the real-shell lane that installs zsh The ordinary Linux unit lane has no /bin/zsh, so the zsh arm failed there with ENOENT. It moves to a live-shell file registered in the shell-contracts lane; the bash arm keeps running in every lane. The lane guard's detector now also sees a zsh spawned through the ProcessSpec program field, which is how this test escaped it. * fix(orchestration): omit a structured child's CLI command when no launcher resolves, and pin its instance A bare `orca` fallback named GNOME's screen reader on packaged Linux, and an inherited value named another app's CLI. The builder now deletes any inherited value, sets the absolute launcher only when one resolved, and pins ORCA_USER_DATA_PATH so a current CLI dials the instance that minted the id. Renames the marker reader to hasStructuredSessionMarker and records why the terminal view carries the id without the marker. * fix(terminal): name this app's CLI launcher by absolute path in every local terminal ORCA_CLI_COMMAND meant three things by lane: an absolute launcher for a structured session, a bare name for WSL, and nothing for any other terminal, so a structured session's terminal view lost it. Local terminals now get the same absolute launcher the structured lane gets; WSL keeps its guest command name, and a terminal whose launcher does not resolve still gets none. * feat(cli): hand a command to the session's own CLI when another Orca CLI was invoked A login shell can reorder PATH behind a global install, and an agent or its helper script can run bare `orca`, so the binary that answered depended on the agent following instructions. Orca's packaged launchers and bare-orca shims now export ORCA_CLI_SELF (outermost wins). At the CLI entry, when it names a different launcher than ORCA_CLI_COMMAND, the command re-runs once through the named launcher with ORCA_CLI_REEXEC=1 and exits with its status; both variables are consumed so no child inherits them. Dev launchers export no self on purpose, WSL and SSH names never qualify, and a launcher that cannot start leaves the command to run here. The Windows launcher no longer rewrites ORCA_CLI_COMMAND; the legacy ask protocol normalizes its resume command itself. * refactor(orchestration): declare which flag names the caller on each spec and refuse at the CLI entry Each handler hand-classified its --from/--terminal as the caller or a target, and the refusal of a conflicting caller flag ran inside the caller resolver plus two standalone calls for --run listings, so a new verb that read its flag raw would pass a sibling's handle to a pre-session host. Specs now declare identityFlagRoles, the CLI entry refuses a conflicting caller flag once from the spec, the resolver only applies the id-wins rule, and a test fails any orchestration verb that accepts --from or --terminal without classifying it. * perf(cli): keep the session caller check off the actor codec's module graph The check runs at the CLI entry for every command, and the actor codec pulls zod through the session record. Compare the session's own spellings as plain strings instead. * refactor(cli): spell a session's address from the one prefix constant, off the codec's module graph The Orca session address prefix moves to a leaf module with no imports, re-exported by the address codec, so the CLI entry check derives `session:<id>` from that constant instead of re-typing it and still stays off the codec's zod graph. Prose and test names say caller or Orca session id, not actor. * refactor(orchestration): drop the session id's terminal-view spawn now that the handoff is gone The terminal handoff was removed, so no terminal is ever a structured session: - delete the terminal-view identity env and its WSL passthrough, and their tests; - strip the session caller keys from every terminal's env unconditionally; - the CLI's own-address spelling moves beside the injected id in src/shared, with a test pinning it to the address the host's party resolver gives that session. * fix(terminal): run the Codex launch preflight through the CLI the terminal names Packaged Linux names the userData shim in ORCA_CLI_COMMAND, while the preflight ran the bundled launcher behind it. The CLI saw a different launcher and handed the preflight off to the shim, booting Electron twice before every codex launch. * revert(terminal): keep terminals on main's ORCA_CLI_COMMAND and Codex preflight Only a structured session needs an absolute ORCA_CLI_COMMAND; local terminals go back to naming none (WSL keeps its guest command), and the Codex launch preflight goes back to the bundled launcher. The CLI handoff is scoped to sessions, so a terminal's preflight can no longer be handed off and start Electron twice. This reverts commit d2cefb6 and commit dd2853a. * fix(cli): hand off to the session's CLI only inside a structured session The handoff ran whenever an Orca launcher's ORCA_CLI_SELF differed from an absolute ORCA_CLI_COMMAND, so any process with both - a terminal, a script - ran another install's CLI instead of the one invoked: a beta's --version lied, and an AppImage command from a terminal that outlived its Orca failed. It now requires the injected session id, the identity it exists to deliver. The launcher variables are still consumed in every process. * fix(cli): name the packaged Windows command after the handoff decision The launcher stopped writing orca/orca-ide over ORCA_CLI_COMMAND so the handoff could see a session's absolute launcher, which also changed what every Windows terminal's CLI read. The CLI entry now applies the launcher's rule itself once the handoff is decided, so terminals and the legacy ask resume command see exactly what they saw before, and the resume-command reader goes back to its original form. * refactor(cli): decide the session handoff from the CLI's own entry, not a launcher export Every packaged launcher, shim and dispatcher exported ORCA_CLI_SELF so the CLI could tell which launcher ran it, and compared that with the session's ORCA_CLI_COMMAND. Two launchers of the same app are different files, so a session that reached its own app through a global orca-ide on Linux still handed off and started Electron twice, and the export rode artifacts every terminal uses. A structured session now also names the JS entry its launcher runs (ORCA_SESSION_CLI_ENTRY), and the CLI compares its own argv entry with it: any launcher of the same app stays, another install hands off. The launcher scripts, Linux shim and dispatcher go back to main; the Windows launcher keeps only leaving ORCA_CLI_COMMAND for the CLI to name after the handoff decision. * refactor(cli): drop the session CLI handoff; the pinned instance and injected id already bind any current CLI Every current Orca CLI dials the instance ORCA_USER_DATA_PATH names and sends the injected session id in the orchestration envelope, so a bare `orca` that reaches another install's current CLI already acts as the session. An older CLI has no handoff code and refuses on the marker. The handoff only lined up versions between two current CLIs, and comparing two separately derived paths kept misfiring (an AppImage's mount against its registered extraction started the CLI twice on every call). Removes the re-exec, ORCA_SESSION_CLI_ENTRY and ORCA_CLI_REEXEC, and the CLI-side Windows command naming; the packaged Windows launcher rewrites ORCA_CLI_COMMAND again, as on main, inside its own process only. resolveHostCliEntryPath goes back to the SSH passthrough. * test(orchestration): say why the registered worker case pins the handle, now that every session's env is populated
… spawn tag (stablyai#23460) * fix(native-chat): stop signalling processes that only inherited a spawn token A spawn token is an environment variable, so every descendant of a provider child carries it. The Linux-only startup scan treated any carrier no lease claimed as a lost provider child and sent it SIGTERM, which also hit editors, tmux servers and nested Orca processes the agent had started. Remove that scan's killing consumer; the token scan stays for the reservation probe, and recorded owners are still stopped by identity during recovery. * fix(codex): remove the token-scan kill path from app-server teardown Every descendant inherits the spawn token, so killing each pid that carries it can reach processes the agent started that are not the provider. Production never injected this path; teardown always uses the process-group and descendant-snapshot proof. Drop it, its deps, and the now-unused spawn-token argument.
…, never the parent's (stablyai#23605) * fix(native-chat): a subagent's words are presented as that subagent's, never the parent's The journal already names the agent that produced every row, but the transcript projection dropped it, so a subagent's prose rendered as the parent's reply, its tool calls folded into the parent's runs, and a settled turn could fold down to a subagent's words as its only visible answer. The transcript message now keeps the row's producer. The fold keeps each agent's calls in that agent's own run, a turn's answer is the session's own agent's last prose, and a subagent's row names the subagent on desktop, mobile and a worker's transcript text. * test(native-chat): give the window fixture's slot the attribution field it now carries * fix(mobile): read the subagent label the row is given, and pin the caption * fix(native-chat): keep interleaved agents in order and each agent's own run live Review follow-ups: - the fold is main's adjacency fold plus one condition: a row never folds into another agent's run, so an agent's later call stays below its subagent's work instead of jumping back into its earlier row - each agent has its own live frontier, so a parent still inside its spawn call reads as running while its subagent works below it - mobile names no one on a row whose only content is hidden behind its settled turn - a pending question from a subagent keeps its producer - worker reads serve only the producing agent's id, bounded like the roster key that names it, and drop the provenance fields - the single-message worker formatter is private, so no caller can drop names
…ly while visible (stablyai#23592) * fix(tab-group): measure fallback pane geometry once per tab group, only while visible * refactor(tab-group): derive the shared resize listener's lifetime from the source map The map already drops empty groups, so a separate counter was a second copy that could disagree.
…zing both (stablyai#23585) Status-row change detection stringified two full IPC payloads on every status write, including an up-to-8 KB lastAssistantMessage re-posted on every OpenCode streamed part. Compare the same published field set with the existing structural-equality helper, with a same-reference fast path. Linear: STA-7432
…ablyai#22644) * fix(ssh): don't overwrite remote agent config after a failed read A flaky read was treated as an empty file, wiping the user's config. Fixes stablyai#22638 * test(runtime): model remote missing-config reads as relay ENOENT errors The runtime harness stubbed isENOENT as code-only, and the remote Codex startup specs rejected with a generic error that only passed while any read failure seeded an empty config. Use the real isENOENT and the message-only shape the relay actually delivers. --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
…stopped (stablyai#23466) * fix(codex): the provider supervisor outlives its provider group when stopped A signalled supervisor forwards the signal to the provider group, escalates to SIGKILL after the grace, and exits only once the group is gone, so recovery's proof that the recorded pid is dead also proves the provider is. It refuses to spawn when its parent is already not the owner named in its spec, and watches that owner rather than whichever parent it first saw. The grace is a spec field. Recovery's SIGTERM stage now outlasts the supervisor's own stop, since a SIGKILL that lands first cannot be handled and leaves the group running. * fix(codex): a closed owner pipe no longer ends the supervisor before its provider group When Orca dies, the supervisor's stdout pipe has no reader. Provider output in the window before the parent-death watch fired raised an unhandled EPIPE that exited the supervisor with the provider group still running. * fix(codex): bound the supervisor grace so recovery's SIGTERM stage always covers it Recovery sized its SIGTERM stage from the default grace, so a launch with a longer grace would be SIGKILLed mid-stop and orphan its group with no test noticing. The spec now refuses any grace above one exported maximum, and recovery derives its SIGTERM stage from that maximum. * fix(codex): every supervisor stop asks the provider with SIGTERM first Owner death, stdin end after the grace, and a signal to the supervisor now all take one path: SIGTERM the provider group, SIGKILL it after the grace, and exit only once it is gone. The signal handlers are registered before the provider is spawned, so a stop that lands in the spawn window still reaps it. The longest stop grows to two graces plus the reap wait, and both recovery's SIGTERM stage and the connection's graceful close now wait that long before forcing, since forcing the supervisor sooner can orphan its group. * fix(codex): give the provider 1 s after stdin end and 3 s after SIGTERM to flush before SIGKILL The supervisor's stop was stdin end, 1.25 s, SIGTERM, 1.25 s, SIGKILL. Codex now gets 3 s after SIGTERM to flush its state. The two graces are separate constants, the longest stop they derive becomes 5.5 s, and a test keeps it inside quit's 8 s child-eviction bound. * test(codex): count eviction's pre-stop drain in the quit budget test Eviction drains the sink for up to 1 s before it stops the child, inside the same 8 s bound.
…ai#23492) * fix(runtime): retire an exited terminal before its stream end An exit's durable retirement became asynchronous, so onPtyExit released the terminal stream before the retirement landed. A paired client answers a stream end by re-activating its pane; that activation still found the exited leaf, materialized it under the same session id, and registerPty dropped the pending retirement. The exited split pane came back as a fresh shell. The exit now stages the retirement into the in-memory session and publishes it synchronously, then notifies exit listeners, and only then makes it durable. A failed durable write is logged and left in memory for the next profile write instead of being rolled back, since the process is gone either way. This removes the pending-retirement latch and its post-await incarnation fence: there is no longer a window for them to guard. * test(runtime): a failed exit retirement still reaches disk Pins the no-rollback contract through a real Store and SQLite authority: when the retirement's own durable write fails, the in-memory retirement is carried by the next unrelated profile write, and by the app-quit flush when no other write happens. The delayed authority fixture can now fail its next write, and the acknowledged-retirement fixture reads the database a relaunch would load and models the quit flush. * test(runtime): a stream end observes the exit retirement already published The re-activation check alone passes with the listener ordering reverted, because activation awaits before its lookup. Record the session binding and publication count at the moment the exit listener fires so the ordering itself is pinned. * fix(runtime): an exit cleanup fault still ends the terminal stream * perf(runtime): exits retired together share one durable write * test(runtime): a refused staging write still retires the pane and ends the stream * refactor(runtime): describe exit retirement as staged, not durably accepted The retirement result is staged in memory before any write, and the removable-surface comment and the replacement-admission test name still described the old publish-after-durable rule.
…ai#23701) * fix(terminal): retain typing while a parked remote pane reattaches * test: persist restored remote terminal screenshots * fix(remote): buffer recovery reconnect input * fix(remote): retain input across restored pane attach * fix(remote): flush attach input after subscription * fix(remote): flush reattach input after attach readiness * fix(remote): stop buffering after reattach readiness * test(remote): trace parked reattach input lifecycle * test(remote): forward paired client lifecycle diagnostics * fix(remote): preserve restored typing before connect starts * chore(i18n): refresh runtime required catalog * fix(i18n): ship compact agent runtime label * fix(i18n): merge required label into existing sidebar catalog
…th no other rest signal (stablyai#23598) * fix(runtime): reopen the quiet-foreground tui-idle lane for agents with no other rest signal A tui-idle wait could never settle on a pane running amp, goose, crush, kimi, qwen-code, rovo, auggie and other agents whose titles Orca cannot classify: the quiet-foreground lane was closed for every launched agent, and it was the only lane those agents could reach, so worker start failed at agent_readiness after 60s. Model each agent's rest signal, derived from the tables that already encode it (synthetic ready titles, the title classifier, the DSH hook and Muse ready screen lanes). The lane stays closed where a stronger signal will arrive and reopens for agents with none. On a reopened lane, silence counts only after the TUI has painted: an agent that has painted nothing is still booting. Linear: STA-7440 * fix(runtime): count only the command's own output as an agent's paint on the tui-idle foreground lane The after-paint lane accepted any output, and the shell's prompt and echoed launch command always land before the agent starts, so a silently booting agent could still settle and lose its first prompt. The runtime now reads the shell integration's command-start marker and requires visible output after it; panes whose shell emits no marker keep the any-output rule. Also skip the foreground-process read while the pane cannot settle, and register the new title-classifier call site in the pane agent identity inventory. * fix(runtime): classify Freebuff's rest signal and skip the backward marker scan on chunks without one Main added the Freebuff agent after this branch point, so the full per-agent rest-signal table no longer matched on the merge ref. Freebuff derives `none`: its screen reports a first-party `done`, which tui-idle trusts only for DSH, so the quiet-foreground lane is its only one. The command-start scan ran a backward search over every PTY chunk; a forward check first cuts that to the cost of a plain substring test on chunks with no marker. * fix(runtime): classify Qoder's rest signal after merging main Main added Qoder with its own readiness branch returning a boolean quiet-foreground flag; map it to the lane type and classify Qoder by its ready screen so the full rest-signal table and lane-agreement check stay exhaustive. Say what `none` actually means: no stronger lane tui-idle trusts, not no hooks at all. * refactor(runtime): track command paint with the shared OSC 133 scanner The command-paint tracker had its own split-unsafe 133;C parser. Reuse the chunk-boundary-safe scanner, which now reports where in the chunk the marker ended, so a marker split across reads is still found. Correct the unmarked-launch list: bash and zsh mark typed launches after the echo. * fix(runtime): drop command-paint state on an output gap or a new process A dropped chunk can cut a command-start marker in half, and the scanner's carry then completes it on unrelated output after the gap, leaving the pane waiting for a paint that already happened. Reset it with the other cross-chunk carries. * fix(terminal): keep the command-start offset out of renderer lifecycle callbacks
…i#23738) * test: wait for revealed remote PTY grid convergence * test: retain reveal diagnostics on geometry failure
…tablyai#23725) * fix(editor): bound offscreen combined diff rendering by height * perf: scope native chat relational styles to their ancestor
* feat(agents): integrate CodeBuddy launch, status and session history * docs: record CodeBuddy lifecycle verification * fix(codebuddy): backfill scoped history and negotiate remote resume * test(cli): include CodeBuddy in known search agents
…2762 (stablyai#23757) stablyai#22762 squash-merged as 29c7d5d with the corpus pinned to its branch commit 03995ae, which the squash left unreachable from main. Repin baseline to main's tip and re-record the full corpus in place; only the baseline header moves on every golden. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
…d of behind it (stablyai#23743) A caller's `needs` gate the whole called workflow, so while the plan job lived in unit-tests.yml it could not start until static analysis and typecheck had both finished and passed -- and the shard matrix then waited on it. The two hops were serial when they did not need to be: planning reads the checkout, a git diff against HEAD^1, the import graph and the checked-in timing baseline in config/scripts/ci-shard-timings.json, and consumes nothing that static analysis, typecheck or the native-cache primer produce. Planning moves to its own reusable workflow so pr.yml can run it against code_paths alone, overlapping it with the gate. Measured across 99 runs, the shard matrix is created a median 93s earlier (p25 47s, p90 241s, never later). Planning stays a required predecessor of the shards, so an empty assignment cannot expand the matrix. The gate itself is deliberately left in place. It fires on 22% of runs, and the shard queue wait knees hard above ~9 concurrent ARM jobs -- 4s median below that against 218s at 15-19 -- so admitting 8 doomed shards per failed run would cost more in queue pressure than it returns in latency. Cost is one 37s ubuntu-latest job, which does not touch the ARM pool the shards contend for. A planning failure still fails the PR: the shards are skipped, and verify's check_job requires success whenever the classifier says tests should run, so it reports `test: expected success, got skipped`.
…lyai#23755) * feat(mobile): tell users when a newer app binary is installable With OTA page updates, store releases get rare and users stop looking. The shell now asks the channel that installed it. Android sideload reads GitHub's mobile-android-v* tag refs and proves the release has an APK. iOS reads the App Store lookup. A home card above Desktops, dismissible per version, and Settings rows surface the result. The check runs on the desktop updater's cadence: cold start, foreground once 24 h have passed, and a 1 h retry after a failure. The releases atom feed was not used because it lists only the 10 newest releases, which are all desktop builds, so it never carries a mobile tag. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): parse update replies with zod schemas The anti-slop gate refuses Reflect.get on dynamic input. The GitHub refs, the release, the App Store lookup and the stored update record are now parsed into named schemas before they are read. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * docs(mobile): say why the Android update source reads tag refs Record why the Android source reads tag refs. The releases atom feed and /releases?per_page=100 are both newest-first windows that desktop releases fill. Either would silently report "current" once a run of desktop builds pushes the newest mobile release out. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): load update state once and apply review rulings Every check and dismissal now awaits one shared store load. This replaces the merge that guessed whether a check had landed during the load. A manual check that fails while the store loads therefore keeps its 1 h retry instead of re-checking at once. - checking is derived from the in-flight check. - start() uses a per-start flag, so a StrictMode double start applies one load. - A check that finishes after stop() writes nothing. - A corrupt stored update record loses only itself. - Tag refs are parsed with a single schema. - The runtime wiring is folded into one file, and the card moves to home/. - The recorded App Store fixture is oxfmt-formatted, with the same parsed value. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): keep the update timer armed across a stop and restart A check that stayed in flight across stop and restart returned 'failed' without rescheduling. The restart skipped arming because a check was in flight, which left a live checker with no timer until the next foreground. The stop counter is removed. schedule() already arms nothing while no start is active, and saving a real result after a stop is harmless. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * chore: retrigger CI after the RPC recording repin (stablyai#23757) landed on main Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): trim the update checker and Settings rows - The load sets prefs and the due time only. start() re-arms the schedule after it. - The Settings result hides through one effect keyed on the result. - onUpdate receives the URL. - The version row is bound once. - The retry and timeout constants are no longer exported. - The unused AppUpdateChecker type is deleted. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): run the update check when its timer fires The armed timer is the due time. Re-checking the wall clock when the timer fired meant a clock stepped back skipped the check and re-armed nothing. The due-time guard now applies only on foreground. The binary version still comes from expoConfig.version. SDK 55 removed Constants.nativeAppVersion, so the no-expo-updates invariant is now named in the comment. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): recover from a future check time and use Apple's page URL If the device clock was ahead when a check ran and was corrected later, the stored check time is in the future. Cold starts then armed a timer for the whole skew, and foreground never came due. The stored state now reads as never checked in that case. The iOS link is the lookup's trackViewUrl instead of a URL built from trackId. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * style(mobile): fit the future-check-time comment in the print width Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* feat(agent-launch): every launch carries the surface that started it The host now attributes every agent it builds to the surface that asked for it, resolving a missing or unrecognized surface to 'unknown' in one place instead of silently skipping it. The CLI names itself on worktree.create and orchestration workers name themselves host-side. * fix(agent-launch): attribute the agent a startup-draft create launches The host builds a third kind of agent launch: a worktree.create with a startupDraft and no startupAgent, where the host picks the agent itself. It carried no launch record at all and ignored the caller's launchSource. Route it through the same resolver as the other two builders, and derive the startupAgent terminal record only from the resolver so no prebuilt record can stand in for it. * fix(agent-launch): attribute the agent a host-built agent session launches terminal.createAgentSession builds a fresh agent's launch on the host, like the other startup builders, but spawned it with no launch record, so those launches were never counted. Record them through the same resolver; the request names no surface, so they count as unknown. * test(agent-launch): require an attribution decision for every host-built agent startup
…blyai#23682) * fix(native-chat): the sink queue keeps a settlement's first batch, as the journal does The journal applies a lifecycle batch's settlement id once and skips any later batch with the same id. The deferred sink queue coalesced the same key the other way: a second batch replaced the first while it was still queued. So which record survived depended on whether the first had drained yet. A lifecycle batch now keeps the queued operation with its key, and a later one is accepted and dropped, which is what the journal does once the first is written. * fix(native-chat): only turn/completed ends a Codex turn Codex follows every turn-ending `error` (willRetry=false) with a failed `turn/completed` for the same turn, 0-32 ms later. That was captured from the real app-server on 0.141.0 and 0.158.0 across eight failure scenarios, and it is how Codex builds a failed turn: it records the error as the turn's last error, records any pending input, and then derives `failed` from that error when it completes the turn. The translator ended the turn twice: once on the error, and again on the completion, with a guard to make the first end final. Ending on the error threw away what only the completion carries: Codex's duration, and the completion's receipt time. It also forgot the turn before Codex recorded the turn's pending input. Now the error is only the row the user reads, inside the still-open turn, and `turn/completed` is the turn's only live end. A process exit between the two is the existing exit sweep's observed end, recorded as interrupted. A failed completion is stored as completed with outcome failure, live and on restore alike. Only `interrupted` maps to the interrupted state. The first-end-final guard is gone. Codex sends one completion per turn, the only redelivery Orca has is the retry of a refused frame (which changes nothing), and the settlement id already keeps the first record in the queue and the journal. * refactor(codex): delete the unreachable oversized-notification settlement The translator settled a streamed item when the transport rejected its notification as oversized. Nothing can produce that frame. The Codex stdio reader frames with `maxLineBytes: Number.POSITIVE_INFINITY` (codex-app-server-record-reader.ts), which it has done since the app-server records were uncapped. With an infinite limit the framer never reports `line-too-long`: no line, pending suffix or paused queue can exceed it. So the dispatcher never emits `frame:oversized-notification`, and the arm that settles it never runs. The arm, its helper module and its test go. In place of the test, the connection test now proves the reason: a notification past the old 16 MiB wire limit arrives whole, and no oversized frame is reported. * test(codex): replace the captured ids in the turn-endings fixture with synthetic ones The replay reads ids only to group frames, so the real thread, turn and response ids from the capture account carry nothing the test needs. The fixture moves beside the Codex tests that read it. * test(codex): use a neutral made-up status as the unknown-status example 'cancelled' read as a stop being recorded as a completion. * test(codex): a restored turn with a status Orca cannot place ends with no verdict Codex's history carries the same status field as the live completion, so the restore path is pinned to the same mapping: completed, and no outcome.
…stablyai#23666) * fix(native-chat): stop flashing "still starting" on every chat launch Every structured chat passes through a short startup phase, and the pane showed "<agent> is still starting…" for all of it, so a normal launch flashed the notice for a fraction of a second. The notice now goes through a keyed delayed status: it appears only once startup outlasts a grace period, stays up for a minimum time once shown, and resets per session. * fix(native-chat): reset startup notice for each provider child
…ly (stablyai#24320) * fix(relay): gate the Asia canary on region fallbacks against a pre-canary baseline Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): gate the Asia canary on region fallbacks against a pre-canary baseline Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
…promotion (stablyai#24318) Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
…move Bun (stablyai#24128) * feat(ai-vault): read OpenCode history with the pinned Node instead of Bun SSH hosts whose Node lacks node:sqlite (or its backup(), which 22.13-22.15 omit) now get the pinned Node in the shared ~/.orca-remote/runtimes/node-<sha> store orcad uses: POSIX hosts receive the official archive and extract and hash-verify it on the host; Windows hosts receive the verified node.exe the client extracted, promoted by host Node with the same hash check. WSL distros use the same layout and checks under ~/.cache/orca/runtimes/. The Bun release pin table and its materializer are deleted. Old relays keep reading their vault-sqlite/<sha>/bun references; nothing deletes those files. An unconfirmed runtime upload now keeps its stage instead of removing it. * refactor(sqlite): drop the Bun SQLite adapter; node:sqlite is the only backend Nothing outside Electron runs on Bun any more (design D4), so SyncDatabase loses its Bun branch, and bun-sqlite-database, bun-sqlite-statement and bun-readonly-wal go, with the relay's bun:sqlite external. The profile-state backup worker admits Electron or an entry that exists, and startup errors name the pinned Node. The D7 cross-runtime gate still runs Bun 1.4.2, now reaching Bun's SQLite through its node:sqlite. * test(native-chat): drop the Bun SQLite driver case now that node:sqlite is the only backend --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
… store GC (stablyai#24130) * fix(ssh): relay version GC deletes only on an exited verdict and keeps the previous build The relay records .relay-pid in its version dir once it owns its socket. GC calls a relay version dir exited only when that PID is provably dead and every relay-*.sock refuses a connection; a dir without a PID file keeps the test -S rule. The most recently completed other relay build is pinned like orcad's rollback target. Design D5 GC liveness. * feat(ssh): collect the shared runtimes/ Node store and give it its own owner runtimes/ gets its own owner in the install model, so no version-dir GC (new or old clients, whose listings are prefix-scoped) can list or delete it. A store pass removes node-<sha> only when no retained dir references it, it is neither a current pin nor the newest other verified runtime, and a ps or /proc check ran and found no process using it. Legacy relay-*/orcad-* dirs are read for references and reported as diagnostics only (design D10 two-step). Wired behind orcad GC's nodeRuntimePins. * fix(ssh): runtime store process check holds runtimes reached through a symlinked home /proc exe resolves symlinks and argv keeps whatever spelling launched the runtime, so filtering on the exact $root path missed in-use runtimes on hosts like /home -> /var/home. Filter on the store segment instead; the parser already attributes holds root-agnostically. * test(ssh): wait for the holder process to spawn instead of a fixed delay --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…compat slot (stablyai#24134) * build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25 through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps its Ubuntu 20.04 (2.31) default. Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014 with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in DT_NEEDED, and smokes it under the glibc-217 Node. * refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…tablyai#24129) * feat(relay): runtime self-test flag and informational runtime on handshake-ok relay.js --orca-runtime-selftest <nonce> dlopens pty.node, opens and closes a PTY, and prints one JSON line (nonce, node, napi, glibcVersionRuntime) for the client to classify before it launches a daemon on a runtime (design D5). handshake-ok gains an optional runtime {kind, version}; bridge and daemon already match exactly on version, so it is informational only (D8.1). * feat(ssh): opt-in pinned-Node relay with prebuilt addons (D5, D6 rung A, D8.1) SshTarget.remoteRuntime (legacy | pinned-node, default legacy; env ORCA_SSH_REMOTE_RUNTIME for development) selects the runtime. On POSIX hosts the pinned path resolves the target with its glibc major.minor, ensures ~/.orca-remote/runtimes/node-<sha>/bin/node, uploads the relay bundle plus the target's node-pty slot and @parcel/watcher from the orcad artifact (no npm or node-gyp on the host), writes .runtime-ref-node-<sha>, and folds the runtime and addon digests into the relay version so pinned and host-Node builds never share a dir or socket. A 30 s self-test (node --version, then the relay self-test) gates .install-complete. Timeouts and lost channels are unverifiable and never step down; noexec, missing_lib, libc_floor, illegal_instruction and wrong_libc refusals fall back to the untouched host-Node path with a logged reason, remembered for the session. * fix(ssh): only an answered libc probe steps the pinned relay down A lost channel during target detection says nothing about the host; descending would launch a host-Node daemon beside a running pinned one and strand its sessions. * test(ssh): mark the mocked SSH connection casts in the pinned relay tests * test(ssh): resolve the pinned runtime mock to an executable path --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…-agent-session-adapter.ts so main's lint passes (stablyai#24323) structured-agent-session-adapter.ts reached 301 lines once two changes on main met, one over the 300-line limit, so oxlint fails on main. The file is the adapter contract: its types, its interface and the typed failure verdicts. The only logic in it read what a cleanup or a stop proved about the provider child's exit: rethrowing a failed acquisition with the verdict its cleanup reached, and reading whether a stop left the provider root gone. That logic moves unchanged to structured-agent-session-provider-exit-proof.ts, which imports the verdicts from the adapter; nothing imports back. Its six importers now import from the new module. No behavior change.
…setting (stablyai#24133) * feat(ssh): complete the relay runtime fallback ladder (D6 rungs B slot, C, D) Rung C runs the relay on the host's Node >= 18 with Orca's prebuilt N-API addons and no npm (addon-only probe mode). Rung B is a data-driven slot chosen only when a compat runtime is listed. Rung D fails the connect with a classified reason carried as a TerminalUnavailableCause. The ladder steps down only on classified refusals; unanswered probes throw. The rung decision is persisted per host keyed by (glibc, runtime hash, Orca major), and ssh_remote_runtime_resolved reports it once per host per session. * feat(settings): SSH host runtime choice (Auto | Orca-managed Node | Host Node) * docs(telemetry): describe ssh_remote_runtime_resolved * fix(ssh): let a passing rung C disprove a remembered noexec; allow glibc-less compat runtimes A remembered rung A noexec was re-persisted even after rung C self-tested addons from the same ~/.orca-remote tree, so rung A stayed skipped until the key changed. Rung B's evaluator also could never match a musl compat runtime. * test(ssh): import node:fs once in the host-node addon test --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…lly receives (stablyai#24343) The parent workflow collapses canary-apply and batch-apply into the job mode apply, so the headroom step's canary-apply/batch-apply condition never held and the gate was skipped on every real roll. Run it wherever the drain runs (apply, rollback before its restart) and in read-only verify; skip only a resumed rollback, which drains nothing. A new workflow-shape test fails on any job step comparing against a mode the parent cannot pass, and on a drain that can run without the headroom check. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* feat(ssh): pinned-Node relay on Windows SSH hosts (D5 Windows, D2) Windows hosts opted into remoteRuntime 'pinned-node' now get the same rung A relay POSIX hosts do, instead of an early host-Node fallback. - Runtime store: the official node-v24.21.0-win-<arch>.zip is uploaded to a stage under %USERPROFILE%\.orca-remote\runtimes, verified against the pinned archive hash, node.exe extracted with System32 tar.exe (Expand-Archive fallback), hashed with Get-FileHash, run once, and published with node.exe + .verified by one Directory.Move. One powershell.exe per phase via the existing powerShellCommand helper; the probe also creates the stage. No new -EncodedCommand site, no -ExecutionPolicy, no Add-Type. node.exe keeps its real name at runtimes\node-<sha>\node.exe. - Bytes that change or vanish after Orca wrote and verified them are reported as ORCA_NODE_RUNTIME_SECURITY_MODIFIED and become a remembered 'security_software' refusal (fallback to the host-Node relay); application control blocks classify as 'noexec'. - Addons: the win32 slot's conpty.node, conpty_console_list.node, conpty\conpty.dll + OpenConsole.exe, watcher and windows-process-tree.node ride with the relay; the orcad template now carries the win32 targets. - Self-test on Windows is one powershell.exe running relay.js on node.exe; the report must name the pinned Node. The relay self-test loads conpty.node and opens a PTY with useConptyDll, and reports a missing bundled ConPTY file as a load failure. A pinned relay's terminals use the bundled ConPTY too; host-Node relays are unchanged. - describeRelayRuntime recognizes the Windows store layout. * fix(ssh): skip the redundant stage-cleanup powershell.exe after a Windows runtime promote The promote script already removes its stage on every path, so the client-side cleanup only runs when promote never returned (upload failure, abort, timeout). * test(ssh): expect the ladder's remembered flag and pin check on Windows pinned relays --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…ack (stablyai#24136) * feat(ssh): run runtime-store GC after a pinned relay launch, under a store lock (D5) The pinned-Node relay deploy now collects runtimes/ after a successful launch, keeping the pin it runs. Promotion in ensureRemoteOrcadNodeRuntime and GC deletion both hold runtimes/.store-lock (install-lock primitives, 20-minute stale rule); GC only tries the lock and skips when busy. A cold pinned install re-checks its runtime under the lock once the relay ref is visible, closing the ensure-then-ref window. GC also sweeps runtimes/.stage-* dirs nothing has written to within the stale rule. * feat(ssh): stream relay and runtime uploads over exec stdin when SFTP is refused (D5) On the bundled ssh2 transport to a POSIX host, a definite SFTP refusal (subsystem refused, sftp-server exited during the handshake, or a chrooted view answering NO_SUCH_FILE for a shell-created path) now falls back to writing through an exec channel's stdin, reusing makePosixWriteFileCommand with a byte-count check and atomic rename. Transport loss, timeouts and aborts never select the fallback. execCommand gains a stdin option. * test(ssh): answer the runtime store lock in the OpenCode runtime setup test Promotion now runs under runtimes/.store-lock, so the mocked host must grant the lock and the stage-exhaustion case makes four more round trips. * test(ssh): rung C relays never take the runtime store lock --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…ials (stablyai#24347) * fix(relay): let a draining cell pass restart-safe through refused redials A draining cell refuses every control and host proof, so once no session, splice, or queued byte remains, an in-flight or reserved connection unit can only belong to a dial the cell is about to refuse. The restart-safe wait no longer resets its pace-window streak on those units, and its progress line now prints every counter the gate reads. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * test(relay): pin fail-closed parsing of handshake counters on a draining cell Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
…4146) * fix(ssh): classify a musl host missing libstdc++ as missing_lib, not wrong_libc musl's loader follows each missing-library line with one 'Error relocating ... symbol not found' per unresolved symbol, and the relocation pattern was checked first. Check missing libraries before relocation errors; the ld-linux/ld-musl interpreter case stays wrong_libc. * build(orcad): allow a partial deployment template for CI build-orcad-template --targets a,b builds and verifies only the named slots, so a CI job that can fill just the x64 Linux prebuild slots can still materialize rung A/C addons. Without the flag every target is still built and verified. * ci(ssh): hostile-host matrix for the relay runtime ladder Drives the real client-side relay deploy against Docker sshd targets and asserts the design D6 rung each lands on: Debian 10 and AlmaLinux 8 (glibc 2.28) and Alpine (musl) on rung A; Alpine without libstdc++ refused missing_lib down to D; Ubuntu 22.04 with a host Node 20 and a noexec home straight to D (home_noexec); CentOS 7 (glibc 2.17) refused libc_floor at A and C, falling to a host-npm path with no Node; and a no-egress Debian 10 still on rung A. Launched cells also prove the terminal echoes, no npm or compiler ran, a second connect reuses the uploaded runtime, and runtime GC keeps the in-use runtime while collecting an idle one. New workflow ssh-hostile-hosts.yml runs on dispatch and on path-filtered PRs. * test(ci): pin the hostile-host workflow to the headless-server builder images The matrix builds its runtime slots in copies of the node-server lanes' Alpine and manylinux images; this contract fails when NODE_RUNTIME_PIN or either builder digest moves in one workflow and not the other. * test(ssh): reconnect as the same client and retry a grace-held PTY owner in the hostile-host matrix --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
… can run (stablyai#24147) * feat(ssh): connect in plain SSH mode when no Orca runtime can run on the host Runtime ladder rung D (design D6): instead of failing the connect, register an ssh2 shell-channel PTY provider and an SFTP-only filesystem provider and publish the classified reason on the SSH connection state. * fix(ssh): harden plain SSH mode against stale reconnects, host sleep and tilde cwd - Only the current connect or reconnect attempt may enter plain SSH mode; a superseded reconnect whose ladder ends at rung D no longer registers a second provider set. - Host-sleep resume probes a plain session over SFTP instead of always reconnecting, which ended every open plain shell. - A home-relative cwd keeps its tilde outside the quotes so the shell expands it. - The SFTP provider implements folder download, which the connect state advertises. * docs(ssh): rung D now means plain SSH mode, not a failed connect --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain>
… Windows (stablyai#24149) * fix(ssh): collect the pinned-Node runtime store on Windows hosts Windows SSH hosts now run runtime-store GC instead of skipping it: one PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from every version dir, and one Get-CimInstance Win32_Process query filtered on an image path under runtimes\ adds process holds (never by image name; a failed query keeps everything). Stale upload stages are swept with the same rule as POSIX. Promotion and the post-upload hold check now take the store lock on Windows too, and the lock's own commands run unwrapped there. Windows relay version-dir liveness now honours .relay-pid (design D5): a live PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing pipes is exited, anything else is unverifiable. The runtime probe adopts a pinned node.exe an earlier vault reader left without a .verified marker after running it. * fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke helper when the relay runs on Orca's verified pinned node.exe: the stage commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It prints the legacy helper's vol:high:low lowercase hex, and identity files are compared after normalising hex spelling, so old and new clients recover each other's stages. Host-Node relays keep the legacy helper; the choice is documented in windows-edr-posture.md. The Windows OpenCode vault reader now installs the pinned runtime through ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified, store lock) instead of uploading a client-extracted node.exe, and the relay dir gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses. * test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane The legacy/node.exe identity compatibility test and the Win32_Process hold path were gated to win32 but no CI lane ran them. Add both files to the Windows package lane and a real running-node.exe hold test. * test(ssh): tear down Windows-lane temp trees through removeTreeSync * test(ssh): grant the store lock to the Windows OpenCode runtime setup test The Windows promote now runs under runtimes/.store-lock, so the mocked host must answer the lock's CreateNew step. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain>
…uthorisation, not job startup (stablyai#24349) * fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup The same-cap gate now verifies the dry-run on its own clock and records the authorisation instant in the single-use consumed marker; each cell job checks the evidence age at that instant and bounds its own start after it. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): refuse a re-run same-cap gate before it consumes evidence; tighten order tests Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
…ablyai#24155) * build(orcad): merge per-runner prebuild slot trees into one matrix Each node-server lane builds only its own node-pty slot. Release CI needs their union before `build:orcad-prebuilds --require-slots` and the template build can run; merge-orcad-prebuilds.mjs verifies every lane's files against its own manifest, refuses duplicate slots and mismatched node-pty/N-API/Node-header builds, then writes one merged manifest. * build(orcad): keep agent-browser out of the desktop deployment template The template rides inside every desktop build (design D2). Seven ~10 MB agent-browser binaries would be ~76 MB, more than the rest of the template; design D2's package contents never listed it, and a slot without one already reports no headless browser. ORCAD_OMIT_AGENT_BROWSER=1 skips the copy; standalone build:orcad still includes it. * feat(packaging): ship the orcad deployment template in desktop builds Design D2: the server JS and every target's addons ship inside the app, as out/relay does; the ~120 MB Node runtimes stay excluded and are downloaded on demand. electron-builder copies out/orcad-template to Resources/orcad-template on every desktop OS, which is the first path materializeOrcadArtifact tries (process.resourcesPath). Platform signing rewrites native bytes the template manifest hashes: - macOS: the tree is signIgnored (codesign rejects its ELF/PE payloads); afterPack signs the darwin targets' Mach-O files with the app identity, as notarization requires, then reseals only those manifest entries. - Windows: SignPath signs after packaging, so release CI reseals from the inner-signing list (packaged-orcad-template.cjs --reseal-signed). Every other file must still match the build's hashes; afterPack verifies. ORCA_REQUIRE_ORCAD_TEMPLATE=1 makes a missing template fail beforePack and afterPack; without it a build ships none and SSH relays keep the legacy path. verify-packaged-orcad-template.test.mjs's "unused, excluded" contract is reversed on purpose. * ci(release): build the orcad template from qualified lanes and package it node-server-tests.yml becomes callable with a ref and build_template. With build_template, each lane that owns a release slot (macOS, Windows, the glibc 2.28 and Alpine lanes, and the glibc 2.17 compat lane) uploads its qualified out/orcad-prebuilds, the Windows lane also uploads both process-table addons, and desktop_template merges them, gates the full matrix plus the compat slot with --require-slots, runs build:orcad-template and uploads the orcad-template artifact. release-cut calls it at the release tag beside the other gates. The build and build-mac jobs wait for it, download it into out/orcad-template (the mac workflow from the parent run), and require it via ORCA_REQUIRE_ORCAD_TEMPLATE. The Windows signing staging skips the template's Linux/macOS payloads, and a reseal step records SignPath's bytes before the installer rebuild. A template-scoped concurrency group keeps a release call and main's push runs from cancelling each other. * test(orcad): keep the packaged-lookup imports clear of the compat-slot import edits * ci(orcad): let a rerun lane replace its template artifacts upload-artifact v4 refuses a second upload under an existing name in the same run, so rerunning a flaky node-server lane during a release would fail at the upload instead of re-qualifying the slot. * ci(node-server): build the template's Windows addons before the lane switches to Node 18 The addon build script imports TypeScript, which Node 18 cannot load, so every build_template run (release-cut included) failed on windows-2022. * fix(build): ship the orcad template's shared node_modules electron-builder's extraResources filter always drops the root node_modules of a source directory, so packaged apps lost orcad-template/node_modules and the afterPack verify failed. Copy it through its own resource entry. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…ned relay (stablyai#24180) * ci(ssh): import the private Windows OpenSSH provisioning harness Copied unchanged from origin/OrcaWin/np-windows-ssh-provider-diagnostic (config/ci/windows-ssh-provider/preview-ssh/ at 1242f3c, commits 78b3857, c24adcc, 069aa7b): a private LocalSystem sshd service on 127.0.0.1 for a dedicated standard user, either the Microsoft-signed Win32-OpenSSH 10.0.0.0p2-Preview ZIP (archive and every binary pinned by sha256) or the inbox OpenSSH.Server capability binaries. The following commits extend it for the pinned-Node relay host lanes. * test(ssh): run hostile-host cells through a host-agnostic driver The Docker matrix drove the relay deploy and inspected the container with inline docker exec calls, so no other host could reuse it. Split it into: - ssh-hostile-host-test-harness.ts: the deploy, ladder observation, terminal echo (per-shell probe), runtime reuse and GC-keeps-in-use assertions, now also capturing every command the deploy sent the host. - ssh-hostile-host-observer.ts: how a driver inspects the host outside SSH; docker exec for containers, the local filesystem for a loopback host. - a legacy_opt_out outcome: the ladder never runs and nothing enters the pinned store, whatever the host-Node path does. Launched cells now also check the runtime's sha256 on the host and that every slot file (the Windows bundled ConPTY pair included) landed in the relay dir. * ci(ssh): Windows SSH-host lanes for the pinned-Node relay Phase 2 exit gate, Windows half: the real deployAndLaunchRelay through a real SshConnection against Win32-OpenSSH on 127.0.0.1, on windows-2022 (x64) and windows-11-arm (arm64), for both the inbox OpenSSH.Server capability and the Microsoft-signed 10.0.0.0p2-Preview release (ZIP; archive and each binary pinned by sha256 and Authenticode, as in the imported harness). Builds on the provisioning harness from origin/OrcaWin/np-windows-ssh-provider-diagnostic (previous import commit): - one private standard account per cell, so every cell starts from an empty runtime store; - -HiddenTools: the private sshd service's own Environment carries a PATH without any machine PATH entry holding node/npm/compilers, led by logging .cmd shims; a session probe fails the job if node.exe still resolves; - DefaultShell set per cell by invoke-pinned-relay-cells.ps1 and restored at cleanup (dispatch proven per cell via %COMSPEC%). Cells (src/main/ssh/ssh-windows-host-cells.ts): pinned-cmd (stock sshd), pinned-powershell (DefaultShell = Windows PowerShell) expect rung A on the pinned node.exe with the relay self-test passing, terminal echo, runtime reuse, GC keeping the in-use runtime, stage identity through node.exe and no Add-Type in any decoded session command; legacy-opt-out expects the ladder never to run and an untouched pinned store. * fix(ssh-ci): tolerate absent-drive PATH entries and retry Windows userData teardown Join-Path throws on a machine PATH entry naming a drive the runner lacks, which would abort provisioning before any cell ran; the toolchain split now probes with [IO.File]::Exists over [IO.Path]::Combine, and the self-test covers an absent drive. The hostile-host harness removes its throwaway userData with removeTreeSync so a transient Windows lock cannot fail the lane's afterAll. * fix(ssh-ci): stop the account list rebinding the typed -Accounts param PowerShell variable names are case-insensitive, so $accounts=[List[hashtable]] assigned into the [int]$Accounts parameter and every Windows host job died before provisioning. Rename the list and make the provisioning self-test reject script-scope assignments that shadow a param. * fix(ssh-ci): hide the host toolchain by ACL, since sessions ignore the service PATH Win32-OpenSSH builds a session's PATH from the machine and user registry values, so the private service's Environment never reached SSH sessions and host node.exe stayed visible. Deny the private accounts the toolchain PATH directories, put the logging shims on each account's own PATH, and record failing sshd and client log lines so a refused login is diagnosable from the receipt. * fix(ssh): resolve the upload root before checking entries stay inside it uploadDirectory compared each entry's realpath against the root as given, so a root reached through a symlink, junction or Windows 8.3 short name (C:\Users\RUNNER~1 in TEMP) rejected every entry as escaped and the pinned runtime upload never started. * fix(ssh-ci): fail cells on a vitest failure and give each account its own keys file The cells script read $LASTEXITCODE under the workflow's GetNewClosure callback, which sees a stale captured copy, so failed cells reported exit 0 and the job passed. Read the global value. Inbox sshd 8.1 checks authorized_keys with read_ok=0, refusing a file other accounts can read; use one keys file per account via %u. * fix(ssh-ci): keep the account name in inbox mode and surface the WMI launch gap The inbox binary-verification loop reused $name, so later SSH and SFTP probes logged in as 'sftp-server.exe'. Before the cells run, probe whether a standard SSH user can call WMI Win32_Process.Create (the Windows relay launch path); when refused, warn and grant the cell accounts Remote Enable on root\cimv2 for the run so the remaining assertions execute. * test(ssh): keep the first terminal session answering keepalives through GC The hostile-host driver disposed the first session's multiplexer before the GC and reconnect steps, so a slow Windows GC let the relay reap the silent owner as 'local' and the reconnect then waited out the full owner grace. Keep the session live until the connection closes, as the app does, and resend the terminal probe until the shell evaluates it: ConPTY PowerShell can drop typeahead sent before its first prompt. * ci(ssh): keep each cell's relay logs in the receipts * fix(relay): detach an ended socket client as peer-closed before destroying it The listener destroyed a socket on 'end' but detached its client only on 'close'. A relay write in that window failed with 'Relay socket is closed', and the dispatcher closed the client as 'local', so its PTY owner kept the full 30s grace instead of the peer-closed floor and a quick reconnect was refused. The Windows host lanes logged this race on the named-pipe endpoint. * test(relay): drive the peer-end listener test with a real dispatcher instead of a cast stub The stub was an unchecked 'as unknown as RelayDispatcher' that failed the changed-code casting gate. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…with EPIPE or ECONNRESET (stablyai#24209) On the Windows arm64 lane a relay write to a named-pipe client could fail with EPIPE or ECONNRESET before EOF arrived. The dispatcher's writer then closed the client as 'local', so its PTY owner kept the full 30s grace instead of the 250ms peer-closed floor and a quick reconnect was refused. The reconnect listener now detaches the client as peer-closed before settling a write that failed with EPIPE/ECONNRESET, or a write attempted after a peer reset destroyed the socket (so ERR_STREAM_DESTROYED counts only via socket.errored). ECANCELED, ECONNABORTED and a plain destroy() stay 'local', since our own teardown can raise them. Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…a symlinked root (stablyai#24179) * test(ssh): upload a root reached through a symlinked parent The upload-root realpath fix landed with stablyai#24180; this keeps macoshost's case where the root is passed explicitly beneath a symlinked parent. * ci(ssh): macOS hostile-host lane on a loopback user-level sshd Adds local-sshd cells for darwin-arm64 (macos-14) and darwin-x64 (macos-15-intel): a non-root sshd on 127.0.0.1 logs in as the runner user with SetEnv PATH=<shims>:/usr/bin:/bin:/usr/sbin:/sbin and an empty HOME, so no rc file restores Homebrew. The driver asserts rung A, terminal echo, cached runtime reuse, GC keeping the in-use runtime, no toolchain or xattr calls, and that the SFTP-uploaded Node carries no quarantine and runs as uploaded. Docker cells are unchanged; each machine runs only cells it can host. * test(ssh): fail a hostile-host run that would skip every named or hostable cell A cell named for the wrong OS or arch was silently skipped, so a macOS job on a mismatched runner went green having deployed nothing. Named cells must now be hostable here, and a gated run must select at least one cell. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…ded cells without rewriting the MIG (stablyai#24373) * fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG The stranded-rollback recovery ran a gcloud rolling action, which renames the MIG version outside Terraform. Every later plan for that cell then reverted the label, and the capacity-plan validator refused the revert as an unreviewed MIG change, so the cell could be neither rolled nor rolled back. The validator now accepts a MIG field moving back to what relay-gce-cells.tf declares (version name and update policy), in every mode, and a test pins those values to the Terraform file. The stranded branch recreates the cell's single instance with recreate-instances, which leaves the MIG untouched. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): let a label-only MIG plan through and recreate on it in a stranded rollback A stranded rollback whose template is already in place plans only the version name revert. The validator still required the MIG template to move, so that plan was refused, and the recreate gate (changes == 0) would have skipped a plan of one change and left the drain flag set. Require the template move only when no declared field reconciles, and recreate whenever the template was not replaced (changes < 2). Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
…t node:sqlite (stablyai#24148) * feat(ssh): wire rung B to the glibc 2.17 compat runtime; gate rung C vault on full node:sqlite - COMPAT_RELAY_RUNTIMES lists linux-x64-glibc217; rung B plans the compat slot and compat pinned Node when glibc is below 2.28 or rung A refused with libc_floor/missing_lib. - The relay version folds the compat runtime's executable hash; refusals are cached per runtime. - The orcad template stages an optional linux-x64-glibc217 target (base package + compat node-pty slot + compat runtime marker); the verifier and materializer accept it. - node-pty slot loader falls back to the compat slot when the default slot is missing or needs a newer glibc. - Runtime store GC keeps the compat pin beside the default one on every relay connect. - hasNodeSqliteReaderApi (DatabaseSync + backup) gates relay session search and the relay OpenCode reader, which now names the host Node version in its unavailable reason; the SSH vault reader installs the compat Node on old-glibc hosts and uploads nothing when no pinned Node can run. - Rung D: a remembered noexec reports home_noexec and never advises installing Node. * fix(ssh): re-prove a replayed noexec after rung D so allowing exec recovers the host * fix(ssh): keep the rung B compat runtime pinned in the relay-connect store GC * test(ssh): mock deployment-target facts in the Windows OpenCode runtime tests * ci(ssh): build the glibc 2.17 compat slot for the hostile-host matrix; CentOS 7 lands on rung B --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain>
…rs work (stablyai#24224) * fix(ssh): launch the Windows relay outside sshd's job without WMI Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js gains a one-shot launcher mode that starts the detached relay with CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a relay without the addon, and a refusal there is named. The Windows SSH-host lanes drop their WMI grant and assert the breakaway route and adoption. * fix(ssh): find runtime holds without WMI on a standard-user Windows host The store GC read held runtimes through Get-CimInstance Win32_Process, which WMI refuses to a standard user's SSH logon, so the pass kept every runtime. On a refusal it now reads this account's own process image paths through Get-Process. * build(relay): ship the Windows relay launcher addon in every desktop package macOS and Linux packages carried Windows relays without windows-process-tree.node, so a legacy-runtime relay they uploaded to a Windows SSH host could not launch outside sshd's job and fell back to WMI, which a standard user is refused. A reusable Windows job now compiles the x64 and arm64 addons once and uploads them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds download them before build:release and require both arches. Staging now rejects a binary with the wrong PE machine, the ReadProcessMemory import, or no spawnOutsideJob export, so a stale pre-launcher build cannot ship. * ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change The staging and gyp-rebuild scripts decide which windows-process-tree addon the relay ships, so a change to either must re-prove the Windows host cells. * test(ci): find the mac orcad-template download by artifact name The release mac job now also downloads the relay Windows process-tree addons, so the first download-artifact step is no longer the template's. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
A connectionless executionHostId is the group's host. Placement still fell back to the focused server, so the group moved when focus changed. Visibility already honored the stamp. Fixes stablyai#13944
The host-section cases skipped the project-grouping flags the sidebar passes, so they could pass without the filtered view placing the group under its owner. Co-authored-by: Cursor <cursoragent@cursor.com>
innocarpe
force-pushed
the
fix-13944-project-group-host
branch
from
October 1, 2026 13:08
82741b4 to
af9ecb1
Compare
Owner
Author
Sync update (
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upstream
Summary
Description A project group that belongs to one Orca server was rendered under whichever server was focused. Runtime-owned groups usually carry
executionHostIdand no SSHconnectionId. The sidebar filter already treated that stamp as the group's host, but placement did noNote
innocarpe/orcamainuntil the upstream PR is merged.