Conversation
…o status works over WSL (stablyai#10328) * fix(agent-hooks): install OpenCode status plugin into the WSL guest so status works over WSL OpenCode reports agent status via a JS plugin dropped into OPENCODE_CONFIG_DIR (unlike Claude/Codex, which use managed hooks.json scripts). Over the WSL runtime that plugin was never materialized inside the guest and OPENCODE_CONFIG_DIR never crossed into the guest, so OpenCode status never reached Orca's sidebar (the workspace stayed green). Codex already worked; this was OpenCode-specific. Mirror the SSH plugin-overlay path for WSL: - Guest relay registers AGENT_HOOK_INSTALL_PLUGINS_METHOD, byte-caps the source, and materializes an OpenCode config overlay via PluginOverlayManager (the same electron-free path the SSH relay uses), returning overlayDirs.opencode. - Host manager ships the plugin source over the existing stdio channel after installers run (and on mid-session reinstall) and records the guest overlay dir; -32601 / CONNECTION_LOST / DISPOSED are swallowed like ssh-relay-session. - PTY env points OPENCODE_CONFIG_DIR/ORCA_OPENCODE_CONFIG_DIR at the guest overlay; until the relay reports it (first spawn / older guest bundle) it drops those vars rather than crossing the Windows overlay path into WSL — so in-guest OpenCode falls back to its own config (pre-fix behavior, no regression). - WSLENV passes OPENCODE_CONFIG_DIR/ORCA_OPENCODE_CONFIG_DIR through (/u for guest paths). SSH is untouched: the same JSON-RPC constant is reused and the guest response merely gains an optional overlayDirs field the SSH host ignores. Runtime repro: native opencode launched in a WSL Orca terminal had ORCA_AGENT_HOOK_PORT/ORCA_PANE_KEY but no OPENCODE_CONFIG_DIR and no Orca plugin in ~/.config/opencode, so agentStatusByPaneKey stayed empty. * fix(agent-hooks): stop the WSL OpenCode overlay leaking Windows paths and churning under running agents Review fixes on top of the WSL OpenCode plugin install: - Never cross a Windows OPENCODE_CONFIG_DIR into the guest. The /p flag was not a defensive default but WSLENV's translate-and-deliver flag, and buildWslRelaySpawnEnv spreads process.env while the daemon merge resurrects keys buildPtyHostEnv only deleted -- so a Windows value reached the guest as /mnt/c/... and was adopted as its OpenCode config root. Register the two vars only when the value is already a guest POSIX path. - Make guest materialization idempotent. The overlay id is instance-scoped, and the host re-ships on every reinstall (60s after connect, and on later pane spawns), so materializeOpenCode's remove-and-rebuild wiped the config root under running agents and raced panes spawning against the path just handed to them. Rebuild only when the shipped source changed or the overlay went missing. - Mirror the guest's default ~/.config/opencode (honouring XDG_CONFIG_HOME) when no explicit dir is discoverable, so pointing OPENCODE_CONFIG_DIR at the overlay no longer silently drops the user's models/agents/skills/mcp. - Carry opencodeOverlayDir across relay relaunch; it is instance-keyed and on the distro's persistent filesystem, so dropping it only blanked status on panes spawned mid-relaunch. * fix(agent-hooks): stop advertising a WSL OpenCode overlay the guest failed to rebuild Round-2 review fixes: - materializeOpenCode wipes before rebuilding, and every failure path after the wipe returns null leaving the dir present but plugin-less. The host treated that null the same as "no handler / teardown" and silently kept the previous value, so a pane could be pointed at an empty config root -- worse than the documented fallback of dropping the var. requestGuestOpenCodeOverlayDir now distinguishes 'none' (guest answered, no dir) from 'unavailable', and the manager clears the recorded dir on 'none'. - The handler cache keyed only on plugin source, so a ~/.config/opencode created after the relay connected was never mirrored for the relay's lifetime. Key on the resolved source dir too, and validate the cache by the plugin file rather than the directory -- the directory is exactly what a failed rebuild leaves behind, so checking it alone made the bad state stick. * fix(agent-hooks): don't mirror the XDG default OpenCode config into the WSL overlay OPENCODE_CONFIG_DIR is APPENDED to OpenCode's config-dir list, not a replacement for it. Verified against the shipped binary: the list is built as [Path.config, ...project .opencode dirs, ...OPENCODE_CONFIG_DIR ? [it] : []], and Path.config is derived independently from XDG_CONFIG_HOME/$HOME/.config. So ~/.config/opencode is read whether or not Orca overrides the var, and the earlier fallback that mirrored it into the overlay made OpenCode load the user's config -- and their plugins -- twice. Resolve only an explicitly-set dir, which is the one case that genuinely leaves the list when Orca overwrites the variable. This also restores parity with the SSH and local paths. * docs(agent-hooks): correct the WSL install-plugins cache comment and test framing The per-call source-dir re-resolution comment still described the XDG default branch that 6745eba removed. The relay's env is fixed for its lifetime and the rc scan behind it is memoized, so the sourceDir cache key is defensive rather than live -- say so, and retitle the test that simulates it by mutating the captured env, so neither reads as coverage of a production scenario.
…hboard (stablyai#10531) The companion-board mutual-exclusion Effect keyed on `workspaceBoardRenderedOpen`, which is `workspaceBoardOpen || workspaceBoardDragPreviewOpen`. WorktreeList sets the drag preview at the start of every card drag — including a pure reorder within a group that never opens the board — so dragging any card closed the Agent Dashboard drawer, and cancelling the drag never restored it. Key the Effect on `workspaceBoardOpen` so only the user actually opening the board evicts the dashboard. The reciprocal Effect is unchanged.
… as freezes (stablyai#10530) * feat(diagnostics): record OS suspend/resume so sleep gaps aren't read as freezes Renderer timers stop across OS sleep, so a multi-hour `renderer_memory` heartbeat gap looks identical to a wedged renderer. That ambiguity sent the uber-crash investigation down a deadlock path that the telemetry later disproved -- and a healthy machine's own trace shows a 427-minute mid-session gap with reason=interval, so the gap alone proves nothing either way. powerMonitor 'resume' was already wired for renderer wake recovery but left no breadcrumb. Stamp suspend and report the measured span on resume so the next freeze report can be told apart from a laptop lid. * fix(diagnostics): only record sleeps long enough to hide a heartbeat Adversarial review flagged that the first cut would flood the 30-entry breadcrumb ring and evict the crash evidence it exists to explain. Measured on 7 days of pmset history: 70 user-visible sleep cycles, worst 60-min burst of 7. Median span is 2 SECONDS -- only ~24% run past 60s, so most of that traffic could never explain a gap anyway. Cycles are counted Sleep -> next FULL Wake, since powerMonitor's resume maps to NSWorkspaceDidWake, which does not fire for dark wake. Drop the suspend breadcrumb (suspend now only stamps a timestamp) and emit a single `system_slept` on resume, gated at 60s. That cuts 70 cycles to 17 over the same week (worst burst 3) while still catching every sleep long enough to swallow a 60s renderer heartbeat. * test(diagnostics): assert resume listeners detach by identity The off mock deleted by event name alone, so teardown detaching a different closure than the one registered still passed -- a leak of the real powerMonitor listener would have gone unnoticed. For 'suspend' that leak has no other observable effect through the public API. Co-authored-by: Orca <help@stably.ai> * fix(diagnostics): span from the first suspend across dark wake powerMonitor 'resume' maps to NSWorkspaceDidWake, which does not fire for dark wake, so macOS can deliver suspend -> suspend -> resume. Overwriting the stamp reported only the trailing segment, and when that segment fell under the 60s gate a 90-minute sleep recorded nothing at all -- leaving the gap looking like the unexplained freeze this is meant to rule out. Co-authored-by: Orca <help@stably.ai> * docs(diagnostics): correct the threshold rationale to match measurement The comment claimed maintenance sleeps would flood the ring. Re-measuring pmset over 6 days (78 sleeps, 29 full wakes, 51 dark wakes) shows they resolve as DarkWake, which never fires 'resume', so they never recorded a breadcrumb at all. Real rate is 29 breadcrumbs / 6 days, worst 60-minute burst 4 against a 30-entry ring. The gate's actual job is narrower: skip sleeps shorter than the 60s heartbeat, which cannot open a gap to explain. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…s hidden (stablyai#10528) startGitCommonNarrowWatch was the only watch entry point that never received WorktreePollerWindowVisibility, so its `worktrees/` existence poll kept stat'ing every repo without a linked worktree forever in the background (0.5 stat/sec/repo at the 2s default). Its siblings — the primary-metadata snapshot poller and the non-darwin git-common polling — already park. Threads visibility through and matches the snapshot poller's park/re-arm pattern: the poll stops on the first hidden tick, and re-checks immediately on onWindowBecameVisible so a worktrees dir created while hidden still upgrades to the native stream and emits its create event. The visibility listener is dropped in the dispose path. darwin-only: the narrow watch is the `platform === 'darwin'` branch.
…tablyai#10526) * fix(terminal): make primary-selection paste suppression single-shot Middle-clicking in the terminal armed a 750ms window that swallowed every native paste event, not just Chromium's one follow-up — so a real Ctrl+V inside that window was silently dropped on Linux. Consume the deadline on first use (`shouldSuppress*` -> `consume*`) so the arm owes exactly one event; 750ms stays as that event's expiry bound. * test(terminal): cover the paste event xterm actually forwards to the PTY Review follow-ups on the single-shot suppression change; no source change. The new real-module file only synthesized `beforeinput`. An Electron probe (Chromium 150) showed `paste` fires first, is cancelable, and cancelling it at document capture suppresses `beforeinput` entirely — and xterm registers `handlePasteEvent` for `paste` only, never `beforeinput`. So `paste` is the sole event that can double-write the PTY, and it was the one event the end-to-end file did not exercise. Add it; it fails with the fix reverted. Consuming mutates, so `isTerminalNativePasteTarget(...) && consume()` is now load-bearing: swapped operands would burn the arm on an unrelated paste and let the real follow-up double-paste. The guarding test asserted only `defaultPrevented`, which survives the swap — assert `consume` is never reached instead. Rename the mocked-file case that claimed to prove single-shot. With the module mocked its sequence is dictated by the mock; it pins that the hook re-asks per event rather than caching, so name it that. * test(terminal): name the beforeinput dispatcher for the event it dispatches Round-2 review nit on my own round-1 change: once `dispatchClipboardPaste` existed alongside it, a helper named `dispatchPaste` that dispatches `beforeinput` inverted the reader's expectation. Match the sibling file's `dispatchNativePasteBeforeInput` convention.
… locale (stablyai#10536) Settings search indexed the macOS privacy-toggle name as English-only aliases on the LAN keyword, so a Chinese user typing the term macOS System Settings actually shows them (本地网络) got no hit — zh's catalog value had been changed to 局域网 (LAN). Split the two wordings onto their own catalog keys so each locale carries both: 87620e6416 = LAN, fa3239cd42 = Local Network (its true content hash). Localized values come from the repo's own LAN title translations and from macOS 26's SecurityPrivacyExtension Localizable.loctable (LOCAL_NETWORK), so nothing is invented. ja/ko/es already matched macOS and are unchanged apart from gaining the LAN key.
…ai#10543) * feat(diagnostics): record OS suspend/resume so sleep gaps aren't read as freezes Renderer timers stop across OS sleep, so a multi-hour `renderer_memory` heartbeat gap looks identical to a wedged renderer. That ambiguity sent the uber-crash investigation down a deadlock path that the telemetry later disproved -- and a healthy machine's own trace shows a 427-minute mid-session gap with reason=interval, so the gap alone proves nothing either way. powerMonitor 'resume' was already wired for renderer wake recovery but left no breadcrumb. Stamp suspend and report the measured span on resume so the next freeze report can be told apart from a laptop lid. * fix(diagnostics): only record sleeps long enough to hide a heartbeat Adversarial review flagged that the first cut would flood the 30-entry breadcrumb ring and evict the crash evidence it exists to explain. Measured on 7 days of pmset history: 70 user-visible sleep cycles, worst 60-min burst of 7. Median span is 2 SECONDS -- only ~24% run past 60s, so most of that traffic could never explain a gap anyway. Cycles are counted Sleep -> next FULL Wake, since powerMonitor's resume maps to NSWorkspaceDidWake, which does not fire for dark wake. Drop the suspend breadcrumb (suspend now only stamps a timestamp) and emit a single `system_slept` on resume, gated at 60s. That cuts 70 cycles to 17 over the same week (worst burst 3) while still catching every sleep long enough to swallow a 60s renderer heartbeat. * test(diagnostics): assert resume listeners detach by identity The off mock deleted by event name alone, so teardown detaching a different closure than the one registered still passed -- a leak of the real powerMonitor listener would have gone unnoticed. For 'suspend' that leak has no other observable effect through the public API. Co-authored-by: Orca <help@stably.ai> * fix(diagnostics): span from the first suspend across dark wake powerMonitor 'resume' maps to NSWorkspaceDidWake, which does not fire for dark wake, so macOS can deliver suspend -> suspend -> resume. Overwriting the stamp reported only the trailing segment, and when that segment fell under the 60s gate a 90-minute sleep recorded nothing at all -- leaving the gap looking like the unexplained freeze this is meant to rule out. Co-authored-by: Orca <help@stably.ai> * docs(diagnostics): correct the threshold rationale to match measurement The comment claimed maintenance sleeps would flood the ring. Re-measuring pmset over 6 days (78 sleeps, 29 full wakes, 51 dark wakes) shows they resolve as DarkWake, which never fires 'resume', so they never recorded a breadcrumb at all. Real rate is 29 breadcrumbs / 6 days, worst 60-minute burst 4 against a 30-entry ring. The gate's actual job is narrower: skip sleeps shorter than the 60s heartbeat, which cannot open a gap to explain. Co-authored-by: Orca <help@stably.ai> * fix(window): stop viewport reflow from moving screen geometry Blink's ScreenMetricsEmulator::Apply checks screen_size and view_position before the desktop/mobile branch, so screenPosition:'desktop' does not make them inert -- despite Electron documenting both as mobile-only. Passing the content size and 0,0 overrode screen.width/availWidth and the window origin for the whole 32ms hold, so a browser context menu opened mid-reflow would translate against 0,0 and land in the wrong place (BrowserPane.tsx:3167). Empty screenSize means 'no override', and an omitted viewPosition stays nullopt in Electron's converter, so the real position survives. Only the scale factor moves now, which is what the reflow actually needs. Also record a breadcrumb when the restore exhausts its attempt budget: the renderer is left at the wrong scale factor until some later reveal fixes it, and that was previously silent. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…i#10525) * fix(release-cut): gate an explicit RC against its own series semver_gt compares through strip_pre(), so the explicit-version override only ever checked the stable line: 1.4.156-rc.0 read as 1.4.156, cleared a 1.4.155 stable, and republished an RC below what clients already run. Anchor a prerelease request on highest_rc_for_base -- the same rc history the kind path uses -- so the override can only advance the series. Two sibling gaps in the same block: - version_suffix was silently dropped when version was set, because the append lives in the kind branch the override skips. - the shape regex rejected X.Y.Z-rc.N.suffix, so a suffixed RC the rc path can produce could never be re-cut explicitly. * fix(release-cut): close both ends of the rc-number range the gate compares The new explicit-rc gate compares with `[[ -le ]]`, i.e. bash machine-width integers, and the author closed only the low end. Past INTMAX bash saturates, so `version=1.4.156-rc.99999999999999999999` reads as "above the published rc.3" and the gate falls open — then the tag it cuts pins highest_rc_for_base at 1e20 for that base forever, and every later cut wraps to a lower rc the fleet never updates to. Bound the rc number to nine digits. Also reject leading zeros on an all-digit prerelease identifier. `npm version` renormalizes rc.4.01 to rc.4.1 while the tag step keeps the literal input, so the shipped package.json version and its own release tag name different releases. The explicit path's embedded identifier now goes through the same validator the kind path uses instead of only the shape regex. * fix(release-cut): stop the refusal pointing minor/major RCs at the wrong series kind=rc derives its base from bump(latest_stable, patch), so the remedy the refusal suggested only works when the requested base *is* that next patch. A 1.5.0-rc.N series exists only because this override created it, so an operator resuming a stuck 1.5.0-rc.2 was told to dispatch kind=rc, which would have cut an unrelated 1.4.156-rc.4. Spell the condition out and give the fallback that does work for a non-patch base. Also correct the mechanism in the comment I added in 698c5be: bash wraps two's-complement, it does not saturate, which is why the hole is value-dependent (rc.10000000000000000000 wraps negative and failed closed, rc.99999999999999999999 wraps to 7766279631452241919 and sailed through). And name both inputs in the suffix error, which now serves version_suffix and the trailing identifier in version. * fix(release-cut): count a suffixed RC from its commit subject, not just its tag The new explicit-version gate only fails closed on a deleted tag because highest_rc_for_base also reads `release: v<base>-rc.N` subjects. That fallback did not parse the suffixed form: rcNumberFromTag accepts an optional .identifier, rcNumberFromReleaseSubject did not, so `4.perf` failed its `(\d+)(\s|$)` anchor and returned null. So deleting a v1.4.156-rc.4.perf tag dropped the series back to rc.3, and an explicit 1.4.156-rc.4 was waved through — below the rc.4.perf build perf-channel clients already run. Same under-count already made kind=rc recompute rc.4 over a deleted suffixed tag. Mirror the tag form's optional identifier. Covered by a unit assertion and a git-fixture test that both fail with this reverted. * docs(release-cut): correct four operator-facing claims in the explicit path All four are wording or consistency, no behavior change (harness: 26/26 before and after, on bash 3.2 and bash 5.2). - The trailing-identifier comment justified itself as preserving a shape that "can never be re-cut through the override", but re-cutting a suffixed rc at or below the series head is exactly what the new gate refuses. State what it actually admits: a second spelling of version=X.Y.Z-rc.N + version_suffix. - version_suffix's input description still said "rc kind only" after this PR made it apply to an explicit bare X.Y.Z-rc.N. - The suffix guard's own rc pattern was unbounded while the shape check twelve lines up is bounded to nine digits; reuse the bounded one so a later edit to either cannot silently drift. - "which recovers the existing tag" was unconditional, but kind=rc recovery is also gated on tag_matches_current_ref, so a tag cut from a ref main has moved past advances to rc.N+1 instead.
* fix(git): read core.sparseCheckout the way git does Sparse-checkout detection parsed git config line-by-line and only accepted a section header alone on its line, so git's legal same-line form `[core] sparseCheckout = true` matched neither branch and was silently skipped: a genuinely sparse worktree lost its badge and partial-checkout warning. It also read `config.worktree` unconditionally, although git honors that file only while extensions.worktreeConfig is on, so a stale worktree config could override the repo's real setting. Headers are now consumed left-to-right off each line (further headers and one assignment may follow), and config.worktree is read only behind the extension gate. Every new expectation was confirmed against real `git config --get`. * test(git): correct what git actually does with a trailing-junk config value Git does not reject `[core] sparseCheckout = true bogus = false` outright: it parses the line and takes the whole tail as one value (`git config --list` reports `core.sparsecheckout=true bogus = false`), then fails only the boolean coercion. The expectation is unchanged; the comment now matches the binary.
…tarve terminal.wait (stablyai#10529) * fix(runtime): reserve long-poll headroom so orchestration.ask can't starve waits orchestration.ask joined the long-poll set, which also opted it into the single server-wide activeLongPolls counter. Because ask blocks on a reply for its full timeout (600 s default, previously unbounded via a caller timeoutMs), 16 asking workers could hold every slot and shed terminal.wait and check --wait with runtime_busy for every other client — mobile, web, CLI, SSH and relay all share this runtime. Meter ask as its own long-poll class with a sub-cap of half the budget, and clamp the caller-supplied timeoutMs at 30 min. The keepalive and abort-signal wiring that motivated the original change is unchanged. * test(runtime): cover the ask sub-cap and counter release on the WebSocket path The admission fence is shared by both transports but only the Unix-socket path was exercised, so a WS-only regression in admitLongPoll/releaseLongPoll would have shipped silently. Drives handleWebSocketMessage with a 'runtime' scoped device (orchestration.ask is absent from the mobile allowlist) and asserts the overflow ask is shed without burning a reserved slot, that check --wait still gets the other half, and that both counters return to zero when the socket closes.
…cker (stablyai#10527) * fix(tasks): keep repos with a pending remote-identity probe in the picker Task-repo eligibility filtered on `hasProjectRemoteIdentity`, which is populated by a background `git remote -v` probe. When the probe could not reach git — an SSH-hosted repo whose connection is not up yet, a cold launch — the repo silently vanished from the Tasks picker and stayed hidden for the full 5-minute negative-cache TTL, even after the host came back. GitHub repos were largely shielded because a persisted `upstream` satisfies the identity projection through a different route; GitLab and other providers depend on the probe. Distinguish unknown from settled instead of hiding both: - `probeGitRemoteIdentity` reports `resolved` / `no-remote` (git answered, no usable remote) / `unavailable` (never reached git). - Enrichment persists `gitRemoteIdentity: null` only on `no-remote`, mirroring the existing `upstream: null` "not a fork" marker. An unreachable host leaves the identity undefined. - Persistence keeps the explicit `null` instead of dropping it. - `getTaskEligibleRepos` keeps a repo whose identity is still pending; folders and settled remote-less repos stay filtered out. * test(tasks): cover the SSH probe exec paths for remote-identity status Addresses CodeRabbit review: the unavailable-on-error case only exercised the local git runner. Adds a connected-provider whose exec rejects, and an SSH repo git answered for with no remotes. * test(tasks): pin that a settled no-remote repo still resolves once it gains a remote Three independent reviewers flagged that the candidate filter's `!repo.gitRemoteIdentity` looks like an oversight next to the new null marker. Tightening it to `=== undefined` would silently stop detecting a remote added after the marker landed. Document that the re-probe is deliberate and pin the behavior with a test.
…stablyai#10547) The old anchored regex matched neither branch on a `[section "sub"]key = value` line, so the parser never left `[core]` and credited the next indented line to it — reporting sparse for a worktree git says is not. Fails on the pre-fix parser (returns true where git reports unset).
…t freeze workspace creation (stablyai#10540) * fix(worktree): bound the .worktreeinclude copy so a huge include can't freeze creation `.worktreeinclude` copying was bounded in entry count (1000) but unbounded in bytes and files, and awaited inline during worktree creation. A repo listing `node_modules` froze creation for minutes behind the create dialog on Linux and Windows, where the fallback is a full `fs.cp` (macOS gets a cheap APFS clone). Measure each copy-mode source against a cumulative budget (2 GB / 50k files) before the first byte is written, and refuse the entries that bust it. Refused entries ride the existing `CreateWorktreeResult.warning` channel so a workspace never silently comes up missing its included files. Pre-measurement rather than mid-copy abort: `fs.cp` ignores its `signal` option, so a started copy cannot be cancelled and would strand a partial tree. Refusing up front means there is no partial state to clean up. * fix(worktree): don't charge bytes for copy-on-write clones, and bound the sizing walk Two defects in the copy budget, both found by review: - The byte limit was applied on macOS, where the copy is an APFS clonefile. Measured: a 2.7 GB tree clones in 22 ms and consumes no disk. Refusing it on a 2 GB byte ceiling denied work that was already free — a regression on the one platform this bound was never meant to touch. Bytes are now charged only when a byte-for-byte copy will actually run; the volume probe that decides this is the same cached df+diskutil pair the clone runs, and writes nothing, so the "refuse before the first byte" invariant holds. The entry limit still applies everywhere: inodes are real work even on the clone path. - A refused entry consumed no budget, so a `.worktreeinclude` listing many over-budget directories paid a fresh full-limit walk for each one — up to 1000 x 50,000 lstat calls, re-creating the stall this bounds. The walk is now charged against its own ceiling whatever the verdict. Also documents that `admit()` must be awaited sequentially (CodeRabbit). * fix(worktree): give the sizing walk headroom so one huge entry can't starve the rest The walk ceiling added in the previous commit was seeded with maxEntries, the same number the entry limit uses. Sizing an entry that busts the file-count limit walks maxEntries + 1, driving the ceiling negative, so every later `.worktreeinclude` entry was refused without being measured at all. That regressed the common case: a repo listing `node_modules` plus `.env` used to get `.env`; it silently got nothing. Reproduced, and now covered by a test that fails when the headroom is removed. The walk now gets 5x the entry budget, so total sizing work stays bounded (<=250k lstat per materialization, vs the 1000 x 50k this ceiling exists to prevent) while ordinary lists never reach it. Entries refused because earlier ones exhausted the walk report a distinct 'sizing' reason, so the warning stops quoting size limits at a 4-byte file that was never measured. * fix(worktree): bill a failed clone's bytes, and blame the right ceiling Two follow-on defects from the copy-on-write fix: - A predicted APFS clone that then failed mid-copy (EPERM, ENOSPC) fell through to a real `fs.cp` whose bytes were never charged, because the entry had been admitted on the premise that cloning is free. That reopened the unbounded copy on macOS. The measured size is already known, so the fallback now bills it and refuses if it no longer fits, reporting the entry as skipped instead of silently copying gigabytes. A clone that was never viable (ApfsCloneUnavailableError) was already charged as a real copy, so that path keeps falling back as before. - The walk ceiling is also applied inside the measurement via min(remainingEntries, remainingWalk), and when the walk term bound, the refusal was still reported as 'entries' — telling the user a 3-file directory busted a 4-file limit. It now attributes to whichever ceiling actually bound. Also fixes the singular warning text, which said "entry X was not copied ... copying them would exceed ... Copy them in manually". * fix(worktree): flag a partial clone leftover, cap the warning, cover two branches - A clone that fails partway only removes an *empty* reservation, so leftovers can survive at the target. Reporting that entry as simply "not copied" sent the user to copy it in manually, straight into a half-populated directory. Those skips now carry mayBePartial and the warning says to check the path first. Cleaning up the leftovers stays the deferred follow-up it already was. - The warning enumerated every skipped path. `.worktreeinclude` allows 1000 entries and all of them can be skipped, so it now names five and counts the rest — an unbounded string is a poor look in a PR about bounds. - Two load-bearing branches had no test, both proven by surviving mutants: the `bytesAreCopied` short-circuit (reachable when a wedged df/diskutil makes the volume probe answer "no clone", so bytes are charged up front and must not be billed twice), and chargeBytes actually consuming budget for later entries. * fix(worktree): only flag directory clones as partial, and cap that list too - mayBePartial was set for every refused clone fallback, but only a *directory* clone can leave anything behind: the file path clones into a temp name and publishes with link(2), so a failure leaves nothing at the target. Sending the user to inspect a path that does not exist is its own small lie. - The partial-copy sentence sliced to five names without the "and N more" that the other sentence appends, so entries past the fifth were surfaced nowhere. Both sentences now share one nameList helper.
…t misreported (stablyai#10550) The 30-min clamp was silent: the ask result carried no timeout figure, so the CLI printed the value the caller *sent*. A worker passing --timeout-ms 3600000 was told "ask timeout after 3600000ms" after only 30 min of real waiting — off by 2x, and accurate before this PR added the clamp. Echo the effective budget on every ask return and print that. Additive optional field; older clients fall back to the requested value.
- Add E2E_FORCE_DAEMON_HEALTH_UNREACHABLE env to simulate failed health checks - Log when replacing a failed daemon, but stay silent on cold starts - Simplify daemon-slow-health-check-preservation: use forced-unreachable health instead of SIGSTOP/SIGCONT - Add --no-sandbox flag to electron launch args for Ubuntu CI - Support extraEnv option in restart session launches
…10615) * fix(diff): keep scroll restore armed through layout shifts * test: add scroll-restore convergence and user-scroll disarm cases Verify that a converging restore withstands layout shifts and continues retrying, while unmarked user scroll disarms the restore attempt. Stabilize marks and offset objects across renders to preserve the bookkeeping state that guards restoration retries.
* fix(editor): preserve Markdown focus handoffs * fix(editor): scope focus requests to panes via viewStateId When opening a file to focus it, tag the pending request with the pane's viewStateId. This prevents split siblings from claiming each other's requests and stops later remounts from stealing focus. Both Monaco and rich-markdown editors now retire requests on mount.
… picker (stablyai#10472) * fix(worktree): collapse duplicate "Local Mac" run targets in the host picker A linked worktree added as its own project projects a second ready host setup on the same project+host, so the run-target picker rendered N identical "Local Mac" rows differing only by path. Only the first was reachable — resolveWorkspaceCreationTarget takes the first project+host match — so the extras pointed at paths that may no longer exist. - Dedupe ready setup options by host in the picker (display fix for profiles that already hold duplicates). - Canonicalize a stale draft's setup id to the setup the picker shows, so the displayed path is the path the workspace is created in. - Reject a linked worktree at repos:add when its main checkout is already tracked, preventing new duplicates. * fix(worktree): only dedupe a linked worktree against a git main checkout Review follow-up: the repos:add guard matched any tracked repo on the main checkout path, including a folder-kind record. A folder repo does not project onto the same project as the git worktree, so matching it would suppress a legitimate add without deduping anything.
…ing, alt-screen snapshot) (stablyai#10614) * test(e2e): register a real runtime host and publish the alt-screen frame as its snapshot Two long-running scheduled-E2E failures on main were stale test setup, not product defects. `onboarding.spec.ts:420` seeded the Active Server by faking a runtime environment in the renderer store and writing `activeRuntimeEnvironmentId` through the generic `settings:set` IPC. Since stablyai#10011 that setter strips the key, and the dedicated `settings:set-active-runtime-environment-preference` handler resolves the id against the main-process environment store — CI logged `RuntimeEnvironmentStoreError: Unknown environment: env-e2e` from `runtimeEnvironments:subscribe`/`:call` alongside the assertion failure. Register the host for real via `runtimeEnvironments:addFromPairingCode` (offline; no live server) and write the preference through its own channel. `terminal-tab-switch-visual-restore.spec.ts:604` wrote alt-screen frames straight into the renderer's xterm, so those bytes never transited the PTY and main's model could not contain them. On cycle 0 the freshly spawned shell still has queued startup output, so hiding the pane makes main's hidden-delivery gate drop bytes and latch a reveal restore, which repaints main's snapshot over the fabricated frame; later cycles run against an idle shell and survive. Arm the existing `setHiddenSnapshotOverride` seam (already used by sibling tests in this file) with the same frame so the live-write and restore paths render identically, and keep the `markerPresent` assertion. Co-authored-by: Orca <help@stably.ai> * test(e2e): keep the alt-screen restore path observable Numbering the snapshot frame one higher than the live-written frame keeps the marker assertion path-agnostic while leaving the frame number on screen as the signal for which path painted. An unrecognised frame now fails, and the per-cycle path is recorded rather than asserted because which cycles latch a restore is load-dependent. Frame authoring and readback move to a helper module; the additions crossed the spec's max-lines cap. Co-authored-by: Orca <help@stably.ai> * Escape regex metacharacters in alt-screen marker pattern Marker is treated as a literal string, so escape regex metacharacters to prevent them from being interpreted as regex syntax. --------- Co-authored-by: Orca <help@stably.ai>
…lyai#10628) Terminal and Window & Sidebar sections now expand alongside Interface by default so users don't miss advanced settings. Sections remain independently collapsible and each can be force-open on deep-link navigation without collapsing siblings. Search disables toggles to prevent unexpected collapse when query clears. Remove unused "ghostty" translation key; product name stays untranslated for search consistency.
…ty (stablyai#10632) Preserve object identity in pane-title overlay rect state to fix infinite re-renders. A fresh {} literal creates a new reference on every call, causing React to treat the state as changed even when logically equivalent, triggering the layout effect to continuously re-measure and re-set state (React #185). Add cold-park verdict flip telemetry to diagnose crash cluster C5: records whether parking state churns in the field for next crash bundle analysis. Document that AddressPicker crash (cluster C6) originates in radix-ui's SelectItem unmount cleanup, not in our component code.
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
* Move language setting to primary interface section Language is a first-class presentation preference alongside Theme and Zoom, so it shouldn't be buried in the Advanced disclosure. Advanced now contains only platform-chrome controls. * Show selected language in Interface section summary Extracts the interface summary logic into a dedicated module that includes theme, language, and font selections. This makes the chosen language visible in the collapsed Interface section summary, addressing the discoverability issue where language was hidden until Advanced was expanded.
* Fix ssh-relay install on hosts with split shell/SFTP namespaces On Synology DSM and similar hosts, the SSH shell and SFTP subsystem expose different absolute paths for the same directory (e.g., /var/services/homes/alice vs /homes/alice). The relay installer silently picked the wrong path and failed discovery. This fix implements SFTP namespace detection: each install creates an unguessable ownership marker and probes both namespaces to detect divergence. When paths differ, SFTP writes redirect to the candidate namespace while shell commands keep the canonical path. Markers are random tokens redacted from logs. * Fix ssh-relay install on hosts with split shell/SFTP namespaces Strengthen path validation to catch traversal and empty segments in absolute POSIX paths, preventing security issues. Improve split-namespace handling with comprehensive wire tests for uploads and file writes. Ensure system SSH connections bypass namespace mapping entirely rather than attempting incorrect retargeting.
…blyai#10640) * feat(source-control-ai): add {linkedIssue} recipe variable for commit and PR prompts Custom commit-message and pull-request recipes can now reference the GitHub issue linked to the workspace, so a template like "Fixes #{linkedIssue}" lands the closing trailer without the user retyping the number. - register `linkedIssue` on the commitMessage and pullRequest actions only, with the VARIABLE_INFO entry the chip hover card requires - substitute unconditionally via `formatLinkedIssueTemplateValue` (empty string when nothing resolves) so the token never survives into a prompt; enrich the draft context conditionally via `withLinkedIssueDraftContext` so unlinked workspaces keep their existing context shape - attach at the 7 call boundaries (runtime commit x2, runtime PR shared, IPC commit x2, IPC PR x2); the pure git gather stays pure - validate the renderer-supplied worktreeId against the request path and repoId before any meta read, comparing SSH paths as raw strings so a Windows host cannot rewrite a remote POSIX path - built-in prompts are unchanged; no GitLab dual-read and no default trailer * fix(source-control-ai): resolve {linkedIssue} adversarial review findings Addresses 13 of the 14 findings from the {linkedIssue} code review (6 minor, 8 nit, 0 critical, 0 major); Issue 5 (GitLab provider naming) is deferred to design Open Question 3 as product expansion. Behavior: - Dialog previews the workspace's real linked issue instead of the synthetic 123, in both the chip hover card and the plan preview, so an unlinked workspace previews the `Fixes #` it will actually generate. Settings dry-runs stay fully synthetic. - Reject non-positive, fractional and unsafe-integer issue numbers at the IPC resolver via a shared isLinkedIssueNumber predicate, so corrupt meta never reaches a draft context (previously -7 rendered `Fixes #-7` and 1e21 rendered `Fixes #1e+21`). - Fail closed on an empty-string repoId instead of skipping the cross-check. Structure: - Split the variable registry into source-control-ai-action-variables.ts and re-export it, restoring max-lines headroom with no consumer churn and no lint disable. - Constrain withLinkedIssueDraftContext to contexts declaring linkedIssue. - Move the misplaced shared imports into their import group. Docs and tests: - Document that the IPC id/path validator guards relay/CLI/future callers, not the renderer (whose path is id-derived), and rename the three tests that read as proof of a protection that cannot fire. - Add PR-side coverage that was missing: three git:generatePullRequestFields handler tests, a built-in PR prompt no-leak guard, and the runtime PR unlinked case. - Replace the coincidental '42' assertion with a fixture-unique sentinel. - Type the runtime worktree fixture with satisfies, which surfaced and fixed pre-existing drift in its git sub-object. - Add an e2e case covering the preload -> main -> meta -> template chain. Co-authored-by: Orca <help@stably.ai> * fix(source-control-ai): resolve {linkedIssue} adversarial re-review findings Addresses all 8 findings from the {linkedIssue} code re-review (2 minor, 6 nit, 0 critical, 0 major); none deferred. Behavior: - Revert the variableOverrides parameter on planSourceControlTextGeneration. Its result is a Save/Generate gate, not a preview, and the recipe it validates is saved repo- or globally scoped -- so rendering it against the active workspace disabled both buttons with "Command input is empty." for a {linkedIssue}-only template on any unlinked workspace, blocking a global settings write. Validation is synthetic again; chip previews are unchanged. - Make the chip hover card additive instead of either/or. A supplied preview now appends a "This workspace" sample below the description and Example rather than replacing them, so the GitLab-empty and dangling `Fixes #` warning survives on the two dialogs where recipes are actually authored. basePrompt keeps its preview-only shape, where the preview is the content. Structure: - Drop the registry re-export from source-control-ai-actions.ts and move the last two consumers onto source-control-ai-action-variables, so one import path per symbol keeps a grep of the registry's consumers complete. - Split the registry/helper suites into source-control-ai-action-variables.test.ts so each test file mirrors its module. Tests: - Cover the Save/Generate gate at the canRunGeneration level for a bare {linkedIssue} recipe on linked and unlinked workspaces, with a negative control proving the buttons can still be disabled. - Cover the chip hover card directly; the dialog tests mock it away. - Guard the PR mismatched-id test with toHaveLength(1) so it cannot pass vacuously on an unrelated early return. - Add an unlinked-workspace e2e case (saw-issue:empty), which is what distinguishes a real resolver from one that always returns a number. Spec now runs green: 3 passed. - Rename the dialog test that claimed a synthetic-fallback assertion it did not make, and route its renders through one shared helper. Docs are worktree-local (.gitignore:84 ignores docs/**): the design doc's plan-preview and chip-surface claims, the manual QA rows, and both reviews' statements about pre-existing PR-handler tests are corrected there. * fix(source-control-ai): make the {linkedIssue} e2e guard and dialog test falsifiable The e2e unlinked case extracted the echoed issue with `ORCA_E2E_ISSUE=(\d*)`, which matches zero digits in front of an unexpanded `{linkedIssue}` and reported it as `empty` — so the case that exists to catch a literal token surviving into a prompt passed on exactly that regression. Capture the whole line instead: a literal now arrives as `saw-issue:{linkedIssue}` and fails, verified by dropping the substitution key for unlinked contexts and watching the case go red. Also drop the inert `not.toContain('Command input is empty.')` assertion — that copy is click-driven `generationError` state and this suite renders statically, so it could never fail; the claim it reached for is carried by the plan test. Rename two plan tests off the "plan preview" framing the design now rejects. Local review artifacts (design doc, implementation notes, final review) were swept to match the tree in the same pass; they are gitignored here. * Resolve {linkedIssue} from live metadata, not cache Resolved worktrees are cached for a second, causing commit and PR generation to use stale linked-issue state. Hosts now implement getWorktreeLinkedIssue to provide fresh issue metadata by worktree id, with proper fallback for unlinked workspaces. Updates both commit message and PR field generation paths; includes integration and e2e coverage. * Keep cached linkedIssue when metadata is unavailable Return undefined from getWorktreeLinkedIssue when live metadata cannot be read (store not ready), distinguishing it from null (unlinked). The caller now falls back to the cached worktree value instead of treating unavailable as unlinked. Also extract the linked-issue echo generator as a shared e2e test helper. --------- Co-authored-by: Orca <help@stably.ai>
* fix(terminal): recover wedged panes after input * refactor(terminal): make wedge probes explicit * fix(terminal): require quiet parser before recovery
…tablyai#10588) * fix(terminal): make Zellij/TUI OSC 52 clipboard copy work by default Zellij and other multiplexers copy via OSC 52. Empty Pc is a valid XTerm default for clipboard, but we rejected it, and the feature defaulted off so copy silently failed inside Zellij. Accept empty Pc as clipboard, default the setting on (query still blocked; size capped), and surface Zellij in settings. Closes stablyai#10567 * fix(review): make the OSC 52 default actually reach existing installs Review fixes for stablyai#10588: - Persistence: profiles saved under the old off default persisted `false`, which is indistinguishable from a real opt-out, so the default flip never reached stablyai#10567's reporter. Added the repo's one-shot stamp (terminalAllowOsc52ClipboardDefaultedOnForAllUsers) so unmigrated profiles flip once and a later opt-out sticks. - Replay: reattach/cold-restore re-writes recorded PTY bytes through the same parser, so a stale `\e]52;c;...` silently clobbered the clipboard on every restart. Gated behind isPaneReplaying via a new resolveOsc52ClipboardGate. - Blocked toast latches once per renderer session and could be burned by a pre-hydration read; it now fires only for a real opt-out. - An empty Pd decoded to '' and, with the gate default-on, silently blanked the clipboard. Now rejected as invalid. - Localization: en.json is bundled and the catalog beats the code fallback, so all three copy changes were inert. Resynced across five locales. - Corrected the empty-Pc rationale: tmux (not Zellij) emits `\e]52;;<b64>`. * test(terminal): cover the OSC 52 gate wiring and settings copy Extracts createOsc52OscHandler so the replay/hydration gate wiring is covered, not just the pure gate — dropping the isReplaying getter now fails a test instead of passing silently. Adds catalog assertions for the two OSC 52 settings strings. Only the toast key was pinned, so the same inert-copy regression (code fallback edited, bundled en.json not) could still ship for the settings pane. * docs(settings): note that the OSC 52 default only covers new profiles Co-authored-by: Orca <help@stably.ai> * fix(terminal): migrate the web settings store to the OSC 52 default-on flip The default-on flip only reached the Electron store. The web/remote client keeps its own settings in localStorage, so a profile that persisted the old `false` there stayed opted out — the same bug the Electron migration fixed, in the second store. Extract the migration into shared/osc52-clipboard-settings.ts and call it from both stores. Also coalesce OSC 52 writes onto a microtask so a hostile chunk of ~15-byte sequences cannot fan out into a million clipboard writes, and latch the blocked-write toast after it renders rather than before. * feat(terminal): tell users when the OSC 52 flip overrides their opt-out The default-on migration cannot distinguish a deliberate opt-out from a profile that simply never touched the setting — both persisted `false` under the old default. Flipping everyone is the only way to fix stablyai#10567 for existing installs, but doing it silently reverses a security choice the user made. Arm a one-shot notice at load when the migration overrides a persisted `false`, on both settings stores, and show it once the renderer hydrates. Profiles that never opted out are never notified. * fix(terminal): clear the OSC 52 notice after it renders, not before Co-authored-by: Orca <help@stably.ai> * fix(terminal): keep the web OSC 52 notice armed against an unmigrated host The host store always projects osc52ClipboardDefaultOnNoticePending, so the plain spread in the web client's runtime UI merge overwrote an arm raised by its own localStorage settings migration — flipping the opt-out in silence. Co-authored-by: Orca <help@stably.ai> * fix(terminal): stop the OSC 52 notice overclaiming, and cover it Round-3 review fixes: - Rename the arming predicate to osc52ClipboardDefaultOnOverridesPersistedOff. Both stores rewrite the whole settings object on every save, so every profile saved under the old off default holds `false` — the deliberate-opt-out cohort is not distinguishable on disk. Name, docs and test names now say so. - Read settings before the UI snapshot in readLocalWebUIState: getStoredSettings() arms the notice, so reading first snapshotted a pre-arm state that callers wrote back, erasing an arm the stamp can never raise again. - Give the notice toast a stable id; StrictMode re-runs the effect against the same closure, so the early return cannot catch the second pass. - Restore guardParserHandler parity in the coalescer microtask. - Drop the unverified Zellij claim justifying all-selections routing; that routing predates this branch and PRIMARY routing stays an open question. - Cover the notice hook (order, single-fire, deep-link), the armed flag reaching disk and surviving a clear, and pin the notice catalog to its code fallbacks. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin OSC 52 setting discovery by product name The migration notice says to turn it off in Terminal settings, so searching Zellij/Grok/tmux has to find it. Also note why the OSC 52 write-back clauses stay despite an unrelated always-true clause in the same condition. Co-authored-by: Orca <help@stably.ai> * fix(terminal): consume the OSC 52 notice on close, and cover the guards it relies on The notice was cleared the moment the toast was enqueued, so a quit inside its 15s window spent the profile's only warning on a launch where nothing was ever seen — and the settings stamp means it can never re-arm. Clear on onAutoClose/onDismiss instead, plus explicitly in the action handler, because sonner's action path deletes the toast without firing onDismiss. Also closes three coverage gaps a review found: - ui.set must accept osc52ClipboardDefaultOnNoticePending. The update schema is strict, so dropping the key rejects the whole call rather than stripping it, and the renderer only logs that failure — every paired client would re-toast forever with nothing red. - the coalescer's try/catch and .catch had no test; the rejection case needs a plain function because vi.fn tracks settled results and hides the leak. - pin that every selection kind (including bare `p`) lands in the system clipboard, so routing PRIMARY separately later is a deliberate break. Co-authored-by: Orca <help@stably.ai> * test(web): pin that ui.get arms the OSC 52 notice when it runs the migration readLocalWebUIState reads settings before the UI blob so the migration's arm is in place before the snapshot every caller writes back. Seeding localStorage after install is what makes ui.get the first settings read, and therefore what makes swapping those two lines fail. Co-authored-by: Orca <help@stably.ai> * test(store): cover the OSC 52 notice clear and its hydration The clear sets local state before persisting so a rejected ui.set cannot leave the toast re-firing for the rest of the session; losing the persist only re-arms the notice next launch. Co-authored-by: Orca <help@stably.ai> * docs(terminal): state the real residual risk of default-on OSC 52 Three comment corrections from review: - the safety note claimed exfil was the risk; queries are blocked, so it isn't. The actual accepted risk is execute-on-paste: decoded text goes to the clipboard verbatim, newlines included. Filtering here would break multi-line TUI copies, which is the feature; bracketed paste is where that is handled, and kitty/Ghostty take the same posture. - the coalescer bounds a flood per parse yield, not overall. - the replay gate reads at parse time while queued live bytes are drained before the guard engages, so a copy racing a reattach is dropped silently. Co-authored-by: Orca <help@stably.ai> * test(terminal): close the four OSC 52 gaps a full revert walked through Mutation testing found four assertions that stayed green against the very change they were written to pin. The notice suite passed 8/9 against a complete revert to clear-at-enqueue: `calls[0][1][callback]?.()` is a silent no-op when the option is absent, and the call count was already satisfied by the enqueue-clear, so nothing separated "cleared by this callback" from "cleared earlier". Assert the option exists and the notice is unspent before invoking it. The stable toast id was deletable with all 9 green despite the adjacent comment calling it load-bearing for StrictMode. Pin it. The blocked toast's latch-after-throw fix was unproven: both orderings pass when `toast.info` succeeds. Only a throwing first call tells them apart. Deleting the hook call in App.tsx silenced the desktop notice with every suite green. Pin it alongside the static Toaster import, since sonner drops a toast enqueued before any Toaster subscribes and never replays it. Also retone the coalescer-latch comment, which claimed the reset ordering was load-bearing on its own; the try/catch reaches the same end, so the test binds the pair. All four verified green->red by mutation, then restored. * test(terminal): cover the OSC 52 notice and its guards Add tests pinning the static Toaster mount required to prevent notice dropout (stablyai#10567), the stable toast ID deduping StrictMode double-invokes, that the notice stays unspent on toast throws, and that flush-latch guards prevent silent consumption across error boundaries. --------- Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> Co-authored-by: Orca <help@stably.ai>
…lyai#11221) The per-scan issue budget kept an issue only when it explained a candidate or truncated the walk. Neither set intersects the attention set, so 'io-error' — the sole reason a plugin-cache scan can raise "Needs attention" — was droppable. Once 16 ordinary issues filled the budget (16 'outside-root' vendor symlinks is an install shape the scan itself documents as normal), a later read failure was evicted for a generic 'issue-limit' row that raises neither attention nor truncation, and the dialog headline read "All installed Orca skills are up to date" over a path that could be hiding a stale copy. Attention issues now outrank the budget, capped at a small reserve so an adversarial tree of unreadable folders cannot pin one issue per folder.
…ai#10647) * fix(activation): stop relaunching the creation-time agent on workspace activation Activating a workspace with zero renderable tabs launched the agent it was created with, unprompted and in approval-bypass mode. Navigation is not consent to start a process: the same fallback fired from post-delete focus handoff, the jump palette, keyboard cycling, CLI/relay activation, and notification clicks. The mechanism was superseded. stablyai#1814 added it when relaunching the created agent *was* the resume feature; stablyai#4706 later added real provider-session resume six lines above and left the fallback in place. What remained fired whenever a workspace had no renderable tabs -- including when nothing had ever slept -- and reported itself as `request_kind: 'resume'` while resuming nothing, discarding any resumable session a plain tab close had already purged. No caller depends on it. All seven intent-carrying callers pass an explicit `startup` on the branch where they intend a launch, and every no-startup branch either declined an agent, already has one running (host `didSpawnStartup`), or is this same defect arriving over IPC. Drops the now-orphaned imports, retargets the stale comment in launch-work-item-direct that cited reopen-relaunch as the reason to persist `createdWithAgent`, and moves the WSL default-args quoting assertion to launch-agent-in-new-tab, whose launch path still resolves those args. Regression tests are revert-sensitive -- all four fail if the fallback returns. * test(activation): name the relaunch regression tests after what they reach Three tests were named after scenarios they never invoked, which is the failure mode that lets a coverage gap read as closed. - The "host-originated" test's `notifyHostRuntime: false` is inert here: both gates resolve through `isWebRuntimeSessionActive`, false with no runtime environment seeded, so it was byte-identical to the plain reopen test. It no longer claims to cover the host `didSpawnStartup` leg, which lives in main and is unreachable from this layer. - The "post-delete focus handoff" test never deleted anything and never touched `prepareActiveWorktreeFocusAfterDelete`. That caller is asserted directly in active-worktree-focus-after-delete.test.ts, which locks out any opts. - The activate/close loop resets state instead of calling `closeTab`, so it does not exercise the sleeping-record purge its comment claimed. Also folds the primary reopen test onto `seedEmptyActivatableWorktree` — the fixture extracted for exactly that state, which its inline copy had drifted from by hardcoding a POSIX repo path. `preflight` is dropped from the launch-work-item-direct comment: the trust preflight reads the create-time argument (worktree-remote.ts), not the persisted meta. Removal safety and ownership do read the field and remain accurate. Renames the ported quoting test to what it pins. Under vitest's node environment `navigator.userAgent` carries no "Windows", so platform resolution bails before the WSL branch and the WSL preference is inert — the real coverage is single-quote escaping of user-configured agentDefaultArgs. * transfer large terminal history seeds across bounded protocol messages - Oversized cold-restore snapshots (>1MB) now upload via chunked startHistorySeedTransfer/appendHistorySeedTransfer protocol instead of inline, avoiding NDJSON line-size violations - Checkpoints automatically trim oldest rows to fit within configured byte limit (200MB) before commit - Protocol v30 required for chunked transfers; v29 daemons gracefully fall back to renderer-only recovery - NDJSON encodeNdjson() validates line size and rejects oversized payloads; notifications silently swallow encoding errors * fix(daemon): drop held output when teardown checkpoint fails to serializ When a final snapshot checkpoint fails to serialize (returns retryable), the pending output records must not be appended later—doing so would splice them over the seq gap left by the failed snapshot, defeating gap detection. Drop the records and retry the checkpoint instead. * Bump daemon protocol version to 30 * Bump daemon protocol version to 30
…unt's home (stablyai#11224) * fix(ai-vault): resume a bridged Codex session under the selected account's home The account session bridge hardlinks every rollout into each per-account CODEX_HOME, and vault dedup keeps the lexicographically-smallest alias, so Resume could pin an inline CODEX_HOME naming a peer account — running the session under that account's auth.json and quota. At resume time the owning host now substitutes the selected account's home when it holds the same rollout at the same sessions-relative path, declining on any uncertainty so resume degrades to today's behavior instead of failing. * fix(ai-vault): repin dropped sessions without a cwd instead of resuming under the wrong account The drag payload only carried sessionCwd when session.cwd was truthy, so a null-cwd codex session dropped onto a pane silently fell back to the prebuilt command - which pins the wrong account's CODEX_HOME, the exact defect this PR eliminates on the other resume surfaces. - Serializer always sends sessionCwd (null when the session has no cwd), so absence now only means an older-serializer payload. - The repin rebuild accepts a null cwd (the builders already omit the cd prefix), matching the sidebar Resume/Copy paths which repin regardless of cwd. - An unrepinnable payload (absent sessionCwd) now fails loudly with guidance instead of silently resuming under the wrong account's home.
stablyai#11228) The startup sweep asks main which panes are stale, and main answers by account id. The renderer then threw that away: it resolved both ids to labels and let the store's A -> B -> A collapse compare the strings. Two accounts can share a label — doAddAccount has no duplicate-email check, so one OpenAI login used in two ChatGPT workspaces gives both the same email, and a failed roster read collapses every account to 'Codex account'. Either way the notice was deleted for a pane that really is running under the account the user switched away from. The sweep then made it permanent: it marked every stale pane notified, including the ones whose notice had just been dropped, and a notified pane is suppressed for the rest of the session. Relaunching cleared the set but the deletion recurred, so the prompt never came back and the pane kept running on the other account's auth.json and quota, silently. Carry the account ids into the notice and decide on them, falling back to labels only for callers that have none; report which panes were left holding a notice so a dropped one cannot claim suppression. The prompt also names the ChatGPT workspace when that is what tells two same-email accounts apart, which is what the two duplicated getCodexAccountLabel copies now share.
* fix(ssh): resync after watcher terminal retry * fix(ssh): resync after watcher terminal retry - Coalesce repeated recovery resyncs within 5s to reduce SSH refreshes during link flaps - Abort in-flight watcher installs when a replacement provider registers, preventing duplicate watchers from old and new transports - Clear resync state when removing watcher snapshots or on provider change to prevent stale retry timers * fix(ssh): resync after watcher terminal retry Avoid logging spurious warnings when a remote watcher is already closed or suspended. Move the console.warn call in handleRemoteWatcherTerminalError() to after the early-return checks. Refactor createSender() in tests to properly simulate the destroyed event for better coverage of retry-cancellation behavior.
…y to users macOS is prompting (stablyai#9756) (stablyai#9910) * fix(macos): add a Full Disk Access nudge to reduce recurring TCC prompts (stablyai#9756) macOS shows the "Orca wants to access other apps' data" (kTCCServiceSystemPolicyAppData) prompt and it can keep reappearing. The reappearing loop is not a fixable app bug: it is TCC identity churn — an unsigned local rebuild mints a new code identity each build, so macOS treats each as a new app — and Orca's other-app reads are already gated behind opt-in settings or explicit user actions. The durable remedy for the population we can help (release users) is Full Disk Access, a superset macOS grant that stops these prompts for a stable identity. Surface it with an ambient, dismissable sidebar card that reuses the existing developer-permissions IPC. macOS-only; probes FDA status at most once per renderer session (the probe itself reads protected data, so it must not repeat on focus/remount); "Open System Settings" opens the Full Disk Access pane; permanent localStorage dismissal. * fix(macos): stop the FDA nudge promising macOS will stop asking The card said Full Disk Access makes "macOS stop asking", but the grant covers this app while terminals are spawned by the detached PTY daemon (daemon-init.ts forks execPath with ELECTRON_RUN_AS_NODE + detached:true, reparented to launchd), which macOS treats as its own TCC identity. A user who followed the card would grant FDA and still be prompted from terminals. Scope the claim to reducing prompts and name the terminal caveat. * fix(macos): drop stale focus refreshes in the FDA nudge refreshFullDiskAccessStatus() applied whichever getStatus() round-trip resolved last. Rapid blur/focus puts several in flight, so an earlier pre-grant 'unknown' landing after a newer 'granted' un-hid the card and also wrote 'unknown' into the module-level session cache, re-nagging a user who already has Full Disk Access for the rest of the session. The adjacent FullDiskAccessSetupPrompt already guards this with a refresh sequence; mirror it here. Also unmount React roots in afterEach: clearing document.body left them mounted, leaking each test's window focus listener into later tests. * test(macos): unmount the StrictMode FDA nudge root between tests The afterEach unmount added in 5a0f717 only covers roots created through renderNudge(). The StrictMode probe test builds its own root, so it was never unmounted and its component stayed live for the rest of the file. Today that component has no window focus listener, so nothing breaks; add a CTA click to it and the same contamination 5a0f717 fixed comes back — the two tests after it see extra getStatus() calls and fail. Register the root so the fix covers every mount site. * fix(macos): attribute the FDA prompts to agent activity, not Orca's own reads The card said the prompts happen "when this copy of Orca reads protected app data", but Orca's own reads are small and gated; stablyai#9756's trigger is agent find/grep sweeps into ~/Library/Containers, which macOS bills to Orca because Orca is the responsible process for every terminal child. Blaming Orca reads as an accusation and hid why FDA works at all — the grant attaches to Orca rather than to each churning child binary. Name agents as the trigger, keep the "reduce" hedge and the terminal caveat, and drop the "this copy of Orca" dev-build hedge that cost a clause. Assert the causation wording so it can't silently regress. * fix(macos): explain the TCC prompts on the settings row, drop the sidebar card The sidebar nudge added in 344d466 was premised on FDA being reachable "only inside onboarding". It isn't: Settings > macOS Permissions has had a full-disk-access row all along (searchable), the Setup Guide hosts the same prompt from both a settings pane and a re-openable modal, and the sidebar already links to that modal via the "Onboarding checklist" entry. The card added a fifth affordance to the same sidebar that already had the fourth, so it bought prominence rather than access - shown to every macOS user without FDA, most of whom never hit stablyai#9756. Keep the part that was actually new. The settings row still described the prompts as something projects and worktrees trigger, which is the same misattribution the card carried: the reads come from the agents Orca runs, and macOS names Orca only because it is the responsible process for every terminal child. It also never mentioned that the grant has to cover Orca Helper, or that the preserved daemon keeps stale TCC state until restart. Non-English catalogs get the English string as a placeholder; the bootstrap translators key their cache on the English value, so a changed string is re-translated on the next run. * feat(macos): nudge Full Disk Access only after macOS repeatedly prompts The FDA hint is only worth showing to users macOS is actually prompting. tccd emits one AUTHREQ_PROMPTING line per consent dialog it displays, carrying the service and both identities, so a narrow log-stream predicate detects the real thing without correlating across lines or guessing whether a dialog appeared. Verified against a captured dialog: the predicate matched 1 line out of 1436 TCC lines in ~28s, because routine preflight checks - the overwhelming majority of TCC traffic - do not emit it. Count dialogs where Orca is the responsible process, persist across launches, and tell the renderer on the third one. The event separates the accessing binary from the responsible app, which is the crux of stablyai#9756, so the toast can name the tool that triggered it rather than blaming Orca generically. One toast per user, with a permanent opt-out; it deep-links to the FDA row in Settings > macOS Permissions rather than restating the guidance. macOS-only: the watcher no-ops elsewhere, the web client stubs the API, and the child is killed on before-quit since log stream ignores a closed stdout. * test(macos): pin the platform so the TCC watcher tests exercise the darwin path start() is darwin-gated, so on Linux CI it no-opped and the stream/kill assertions passed vacuously against a watcher that never spawned. Pin process.platform per the existing convention (shared/secure-file.test.ts), and cover the gate itself with an explicit non-darwin case. * fix(macos): start the TCC watcher from app bootstrap, not the window wiring attachMainWindowServices is called directly by its own unit test, so wiring initTccPromptNotice there made `vitest src/main/window/` spawn real `log stream` children that outlived the run - two orphaned watchers were left behind by a single test session. Only the IPC handler registration stays there; the spawn moves to the real app bootstrap in index.ts, which tests never execute. Verified: running the suite that leaked now leaves the watcher count unchanged. * fix(macos): clarify repeated permission notice * fix(macos): keep TCC notice lifecycle safe * fix(macos): retain pending TCC notice delivery * fix(macos): acknowledge TCC notice delivery * fix(macos): release failed TCC notice claims * fix(macos): retry transient TCC notice display * fix(macos): contain TCC notice IPC failures * fix(macos): harden TCC notice renderer lifecycle * fix(macos): contain TCC notice dismissal failures * test(macos): satisfy promise executor lint * fix(macos): detect helper-attributed TCC prompts * fix(macos): align TCC watcher lifecycle and helper identity * perf(macos): defer TCC log reader until first paint * fix(macos): recover deferred TCC watcher startup * fix(macos): recover TCC watcher from deferred quit * fix(macos): localize recurring file access notice * fix(macos): preserve TCC watcher and localized guidance * fix(macos): avoid duplicate TCC watcher recovery * fix(macos): wait for locale before TCC notice * perf(macos): isolate TCC notice subscriptions --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
…tablyai#11087) * fix(release): stop packaging plugin authoring examples into app.asar electron-builder's `files` is an all-negation list, so its default `**/*` packs anything without an explicit `!` entry. examples/ arrived with the plugin system in stablyai#8549 and never got one, so 1.4.160-rc.3 shipped examples/plugins/hostile-panel/panel.html — the adversarial fixture the panel containment tests point at, complete with its fetch-exfiltration probe — plus hello-orca, inside every user's app.asar. Verified against the installed 1.4.160-rc.3 artifact, not just the config. The two orchestration design docs landed at the repo root in the same span and shipped the same way; fold them into the existing root-doc negation. Neither has a runtime consumer: bundled plugins ship via extraResources from resources/plugins/launch/, which is already excluded from the asar for exactly this reason. * test(release): assert the examples exclusion through the real file matcher The added case mapped each negation to a bare top-level token, so it passed under '!examples/README.md' — a pattern that still ships the whole tree. Drive app-builder-lib's FileMatcher instead so the assertion matches the test name, and pin the root anchoring so the negation cannot grow into '!**/examples'.
…ds (stablyai#11223) A local-build check (Option+click "Check for Updates" on macOS) pins activeUpdateSource to 'local' for the rest of the process. The 'update-available' success path never restores it, and runBackgroundUpdateCheck early-returns on it, so every wake-from-sleep check, window-focus daily check and nudge poll became a no-op once a local build reached 'available'. The one-shot automatic timer fired into that early return and nothing re-armed it, so the scheduling chain died too and lastUpdateCheckAt froze. Restoring the source when 'update-available' fires would break the flow the user just started — the pending download still needs the local feed and allowDowngrade. Instead the release source is restored when the user closes the offered card, which main previously never learned about, and only while status is exactly 'available': downloadUpdate() flips status to 'downloading' synchronously before it calls into electron-updater, so this cannot fire once a download is under way. The automatic timer now re-arms when a check is deferred rather than launched, so a deferral can no longer end automatic checks for the process lifetime.
* feat(dashboard): add agent status search board * fix(dashboard): keep idle controls reachable * chore: drop merge-only formatting drift * fix(dashboard): compare sparse subagent snapshots safely * fix(dashboard): satisfy settings handler lint * fix(dashboard): address review feedback * fix(dashboard): complete search and localized status copy * fix(dashboard): pad active filter row * fix(dashboard): keep idle control in board settings * fix(dashboard): source filters from workspace state * fix(dashboard): clarify PR and MR status filter * fix(dashboard): preserve review and board parity
…sh (stablyai#11089) * perf(dashboard): stop re-sending repo icon data URLs on every republish stablyai#11012 put repo icons on the dashboard snapshot keyed by repoId. Image icons are data URLs capped at MAX_REPO_ICON_DATA_URL_LENGTH (400KB) and every repo contributing a card ships one, while the snapshot republishes up to 4x/sec (PUBLISH_THROTTLE_MS = 250) for as long as the pop-out is open. Icons change about never, so that structured-clones megabytes per second across the window boundary for bytes the pop-out already has. Publish the map only when it actually changed, comparing by reference since icons come off immutable store repo records. The two paths where the pop-out could be starting from nothing — it opened, or it mounted and asked — still force a full send, so the retained copy can never be the only one. The pop-out keeps the last map it was given when a republish omits the field. An explicitly empty map still clears, so removing an icon works. repoIconsByRepoId was already optional on DashboardSnapshot and isDashboardRepoIcons already returns true for undefined, so the main-process validator needed no change. * fix(dashboard): keep repo icons in the main-process snapshot cache The bridge now omits an unchanged repoIconsByRepoId from republishes, so the cached snapshot main replays to a mounting pop-out could be icon-less, blanking the board's repo glyphs until the forced publish landed. Carry the last map into the cache; the forwarded payload is unchanged. Also covers the forced full sends (open, reopen, snapshot request) that no test exercised. * test(dashboard): pin the icon omit on the throttled trailing republish * fix(dashboard): keep the popout bridge effect off the react-doctor gate The changed-code quality gate reports react-doctor findings that overlap added lines, and effect-needs-cleanup spans the whole publish effect — so this PR's edits inside it turned a pre-existing false positive into a red static-analysis check. Hoisting the store subscriber leaves the effect owning one disposable; behaviour is unchanged. * docs(dashboard): correct why watchSnapshotInputs sits outside the effect The effect owns four disposables (offOpenChanged, offRequested, the store unsubscribe, and the trailing timer), not one. State the real reason the subscribe is hoisted so nobody inlines it back and re-reds the gate. * test(dashboard): pin that the bridge subscribes only while the pop-out is open The lazy wiring exists so an enabled-but-closed pop-out costs nothing — a live subscriber would rebuild a cross-worktree snapshot on unrelated store writes. Nothing pinned the unsubscribe on close.
* fix(runtime): surface desktop RPC startup failures * fix(runtime): isolate RPC failure telemetry * fix(runtime): satisfy the changed-code quality gate and kill vacuous dialog tests The `no-floating-promises` label span covers the whole `app.whenReady().then()` callback, so adding lines inside it made a long-standing finding overlap changed code. `void` is the linter's own suppression; no `.catch()` on purpose. The startup-failure tests were vacuous: mutation runs showed the wait-for-show deferral, the destroyed-window guard, the `closed` companion event, listener cleanup, the cause walk, the cycle guard, and the truncation bound could all be deleted with every test still green. The "not called yet" assertion ran before any microtask, so it passed either way. * test(runtime): de-brittle the desktop RPC-failure source assertions Anchoring the slice on the full destructure and matching the whole dialog call expression made an innocuous rename break the test with a cryptic 'expected -1'. Match the shape that is actually the contract instead. * test(runtime): repair the silently-unbounded desktop startup slice The desktopEnd anchor comment lost a word in 98b00d3, so indexOf returned -1 and slice(start, -1) covered index.ts to EOF. Moving the dialog call to a path that never runs at startup still passed. Anchor on code instead, and assert both bounds so a future reword fails loudly. * test(runtime): bound the attach anchors in the startup ordering slice Round 3 bounded the desktop pair but left attachStart/attachEnd unguarded in the same test: deleting the PTY startup barrier from attachMainWindowServices() and breaking the rateLimits.attach(window) end anchor still left the case green. * test(startup): bound the last two unguarded slice anchors in this file Rounds 3 and 4 fixed the desktop and attach pairs; two instances of the same class survived in the same file, both proven vacuous by mutation: - it #3 never bounded readyEnd. Renaming the `pairing:` payload key makes it -1, widening readyPayload from 372B to ~52KB. Moving the reconciliation status out of the serve-ready payload (its whole point) but leaving it later in index.ts then kept all 6 cases green. - it #2 bounded desktopWindowStart against reconciliationStart rather than serveEnd. An earlier `Promise.resolve(openMainWindow())` steals the anchor, collapsing desktopStartup to '' while every existing guard still passes, so its only assertion — a negative — succeeds against an empty string. Both mutants now fail. `src/main/ipc/pty-startup-barrier-ordering.test.ts:11` has the same latent shape; left alone as out of scope for this PR. * fix(runtime): keep walking the cause chain past an unmapped code getErrorCode returned the first code it found, so an outer wrapper carrying an unrecognised code masked a nested EACCES/ENOSPC and classified it unknown. Only a mapped code ends the walk now; every other input classifies as before. Unreachable today (writeSecureFile rethrows raw fs errors with .code intact), but the classifier's job is surviving whatever error shape reaches it. * fix(runtime): tell the user what to fix, not just to restart The dialog's only advice was "Restart Orca to try again", which is true for address_in_use and wrong for the rest: permissions, a full or read-only disk, and a missing data folder all survive a relaunch, so the user restarted, hit the same failure and had no next step. Route the error class we already compute into the copy so each cause names the thing the user has to change. Guidance and telemetry now derive from the same classifier, so they cannot drift apart. * fix(runtime): guide users through long RPC paths * fix(runtime): avoid false window listener warning * fix(runtime): guard destroyed window before web contents
…detach (stablyai#10698) * fix(agent-status): prevent ghost sidebar row on completed split-pane detach Detaching a done-state split pane into its own tab migrated the agent paneKey from oldTab:leaf to newTab:leaf. useRetainedAgentsSync only saw the old key vanish and, finding no suppressor, resurrected it as an unclickable duplicate sidebar row (and inflated the worktree count). Plant a one-shot retention suppressor on the source key during transferAgentPaneAuthority, but only when the source actually held a live agent, so a suppressor is never leaked for a pane that had none. Fixes stablyai#10675 Co-Authored-By: Claude <noreply@anthropic.com> * fix(agent-status): annotate suppressor record type and condense retention comments Type the migrated retentionSuppressedPaneKeys as Record<string, true> so a computed-key `true` isn't widened to boolean, which broke the web typecheck. Also condense the retention rationale comments per review. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
…#10256) A deadlocked main thread never crashes, so it leaves no crash report and no artifact — incidence has been unmeasurable (n=1 confirmed, macOS 26.5.1, FB24004458 / electron#52437). This forks a plain-Node watchdog sibling under ELECTRON_RUN_AS_NODE that survives the deadlock, listens for a 2s heartbeat, and after 45s of silence writes a marker to userData. The next launch consumes it, records a durable crash breadcrumb, and emits a main_thread_hang_detected telemetry event carrying unresponsive_ms and self_recovered. Observes only — it never kills or relaunches the parent. A true positive recovers nothing force-quitting wouldn't, while a false positive would SIGKILL a live main thread mid-write. self_recovered counts exactly the stalls such a killer would have gotten wrong, so recovery can be built on evidence if the field numbers justify it. macOS-only, packaged-only (ORCA_HANG_WATCHDOG_FORCE=1 to test), with sleep-gap suppression and idempotent shutdown on will-quit.
…t opens it (stablyai#11128) * fix(skills): stop the skill review dialog contradicting the badge that opens it A skill whose only fault was an edited copy or one Orca could not read turned the setup-rail badge amber and offered Details — and Details opened a dialog headlined "All installed Orca skills are up to date." over an empty list. The badge says something is wrong, the dialog it points at says nothing is. The grouping only returned skills with an out-of-date copy, so those two states produced no row and the summary fell through to the all-clear headline. Include a skill when a copy needs attention as well, using one shared predicate so the badge and the dialog cannot disagree again. A plugin's own copy of a same-named skill stays out: that is the vendor's, not the user's drift. * test(skills): pin that a routine outdated copy raises no attention marker
…erge (stablyai#11248) Re-lands stablyai#11129, which was merged into stablyai#11128's branch rather than main and so never reached main. Content is identical to the reviewed and live-QA'd head ac5ec5b (1775d83 + ac5ec5b, minus the intermediate merge).
* fix(perf): correct three 07-27 perf regressions Traversal capacity cap no longer scales with worker concurrency (stablyai#11026). retainWorkspaceSpaceScanEntry charged a traversal-wide entry counter, so N workers each holding a listing multiplied the live charge. At concurrency 48 a 48x2,100 tree (100,848 entries) hit the 100,000 cap while 100x1,500 (150,100 entries, 50% more) passed, and scanLocalWorktree treats the capacity error as terminal, reporting an intact worktree as "Unavailable" with sizeBytes 0. The cap is now per directory listing -- the only quantity fixed by directory shape -- restoring the invariant docs/workspace-space-scan-resource-bounds.md already states. Aggregate live retention stays bounded by the unchanged 64 MiB byte cap. Note: releasing each entry's charge at dispatch (the originally suggested fix) was measured and does not help; the peak is set at admission, before any entry is dispatched. Repo image icons are no longer fully base64-decoded on every snapshot publish (stablyai#11012). sanitizeRepoIcon reached decodeBase64Prefix, which sized its buffer to the whole payload to read a 24-byte header, running synchronously inside ipcMain.handle at a 250 ms throttle. Validation is now memoized on source+src in a BoundedMap. Measured for 10 icons x 256 KB: 37.34 ms -> 0.67 ms per publish. One over-long card label no longer discards the entire snapshot (stablyai#11012). isDashboardSnapshot was all-or-nothing and dashboard-popout returned early with no log while replaying lastSnapshot, so `orca terminal rename --title "<1025+ chars>"` froze the pop-out board on its last good paint with nothing surfaced. Labels are truncated at the producer, the validator drops only the offending card, and both the rejection and the drop are logged. The bound now lives in the shared snapshot contract so producer and validator cannot drift. Co-authored-by: Orca <help@stably.ai> * fix(perf): charge a scan listing's parent path once, not per entry The 4.1 fix made the entry cap per-listing but left the 64 MiB byte cap charging parentPath.length for every entry in a listing. Because a listing's entries all share one parent-path string, that multiplied the path by the directory's width, so the byte cap measured checkout depth rather than live heap. The reported symptom therefore still reproduced at the production default limits: 48 x 2,100 @ concurrency 48 raised a capacity error once the worktree path passed ~58 characters, while the same layout at concurrency 1 succeeded. The shipped regression test could not see this because it passes maxRetainedBytes: Number.MAX_SAFE_INTEGER, disabling the only cap still in play. Measured at a real 65-char worktree root, 3 of the report's 4 documented layouts still failed. The parent path is now charged once per listing, with its first entry, so an empty listing strands no charge. Per-entry overhead is unchanged at 512 B + name, which still dominates the estimate, so the OOM protection the original PR added is preserved. Adds a production-default-limits case covering the report's layouts under a deep root, plus an assertion that a short and a deep root reach the same verdict -- the path independence docs/workspace-space-scan-resource-bounds.md requires and which no existing test enforced. Co-authored-by: Orca <help@stably.ai> * fix(perf): prove the icon cache by decode count, not wall clock The caching test asserted a per-publish millisecond budget, which failed on CI at 5.64 ms against a 5 ms ceiling. Any threshold flakes on a loaded box, so count real sanitizeRepoIcon entries instead: 10 repos x 20 publishes is 200 icon checks against exactly 1 decode. Added cases pin the cache key (payload and source both re-decode; a cached image verdict never answers for an emoji) and that a rejection is cached too. Also drops budget.entries, which the per-listing cap left as a traversal-wide counter no check reads -- exactly the shape a future guard could reintroduce the concurrency bug from. Co-authored-by: Orca <help@stably.ai> * fix(dashboard): bound the project filter label the whole board rides on stablyai#11042 added snapshot-level filterOptions whose project labels are repo.displayName -- the same unbounded source this PR already bounds for card.repoName, but one level up where dropping a card cannot recover it. An over-long project name would fail isDashboardFilterOptions and take the entire snapshot with it, which is the exact frozen-board failure the per-card drop was added to end. Workspace-status labels are already capped at 32 by workspace-statuses.ts, so only projects needed this. Co-authored-by: Orca <help@stably.ai> * fix(dashboard): disambiguate the repo icon cache key The memoization key joined `source` and `src` with a space, but the sanitizer's base64 pattern admits whitespace inside a valid `src`. A rejected icon can therefore split the same concatenation differently and inherit an accepted icon's cached verdict, reaching the pop-out's `<img src>` without ever being sanitized. Length-prefix the source. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…ted locales (stablyai#11241) * fix(ui): preference sync, picker arming, zoom, chat status, and reverted locales 7.1 ui.set rejected whole preference payloads on enum drift. The new AssertNoMissingKeys guard is key-only, so it could not see that LegacyWorktreeCardProperty omitted 'cli' (in DEFAULT_WORKTREE_CARD_PROPERTIES) or that rightSidebarTab omitted 'workspaces'/'pr-checks' and every plugin tab. UiUpdate is .strict(), so one bad value failed the entire batch and silently dropped sidebarWidth/groupBy/sortBy/filterRepoIds riding the same debounced write. Both enums now derive from the shared unions, AssertNoMissingValues catches value drift by name, and UiUpdate drops an unknown value instead of rejecting the batch around it. Unknown KEYS still reject. 7.2 The SSH shell-ready fallback moved from first-output to spawn, so a remote shell needing >1.5s to prompt got the bracketed-paste startup command before readline armed it, with no recovery afterward. The short deadline now applies only once output proves the shell is talking; a silent-since-spawn shell gets a longer budget and still delivers eventually. 7.3 The project picker armed in rank order but rendered in section order, so with a folder group present the BOTTOM row was armed on open and Enter created the workspace in the wrong place. Row keys now derive from the same sections that render. The folders bucket also gains the recent-exclusion guard the projects bucket has; that duplicate was unreachable, so this is symmetry, not a live bug fix. 7.4 setBrowserPageZoomLevel now compares before writing, so a pane reasserting a level the host already holds no longer emits a redundant host-wide HostZoomMap write. The user-applied level also moved to a module-level map keyed by page id: the guest webview outlives its React pane, so the pane-local ref re-seeded from the shared Settings default on every remount and let a later default retroactively hijack an already-zoomed tab. See PR notes on the part of this finding that could not be fixed as prescribed. 7.5 A non-null sessionId short-circuited the live-work escape hatch, forcing 'loading' over hook 'working' and rendering an idle pane mid-turn: Send instead of Stop, no typing indicator, no streaming preview. Status stays 'working'; the empty-transcript loading SURFACE moves to selectNativeChatViewState, which keeps 7.6 stablyai#10770 merged from a base predating stablyai#8549, reverting 182-187 translated strings per locale to English (es 182, ja/ko/zh 187) plus en.json's recipesHelp. Restored by script, only where the English source is unchanged between the two shas, so later legitimate edits are preserved: 0 keys added or removed, every value sourced from 97e4776, and the four other English-source changes since 7.7 Match highlighting indexed by UTF-16 code unit but rendered by code point, so an emoji-named folder showed marks one glyph late. Cosmetic. Co-authored-by: Orca <help@stably.ai> * fix(ssh): keep fast startup delivery on the short fallback deadline The 15s no-output budget added for the shell-ready fallback was applied to every SSH launch, including 'fast' delivery. Fast delivery waits for no marker and pastes nothing prompt-sensitive, so it gained a 10x startup delay for nothing. Co-authored-by: Orca <help@stably.ai> * fix(rpc): generalize the ui.set value-parity guard to every shared key Naming worktreeCardProperties and rightSidebarTab left the next field to drift exactly as unguarded: dropping 'pr-status' from groupBy typechecked clean. Check the value domain of every shared key instead, against z.input (what a client may send) rather than z.infer (post-transform). Also move the pure mergeNativeChatLiveSession suite beside the module it covers; the hook's test file owns an IO harness and was at the max-lines cap. Co-authored-by: Orca <help@stably.ai> * fix(i18n): re-apply only the locale strings still reverted at HEAD stablyai#10770 merged from a base predating stablyai#8549, so its stale locale copies overwrote ~185 already-translated strings per locale back to English. A present catalog value always beats the English translate() fallback, so those strings render English with nothing to signal the loss. Since that finding was written, stablyai#11205 and other upstream translation passes independently re-covered most of ja/ko/zh. Replaying stablyai#8549's catalogs wholesale would now overwrite that newer work, so this re-applies a key ONLY where all of the following hold at origin/main: it was translated at 97e4776, stablyai#10770 reverted it, its English source is unchanged since, and no upstream commit has touched it since the revert. es 182 ja 32 ko 17 zh 55 Everything else is left to upstream. Verified zero upstream translations reverted: every changed key still matches its stablyai#10770 value at main. Keys upstream deleted are not resurrected, and keys whose English source was edited since are skipped as legitimate source changes rather than reverts (this is what keeps zh CPU on stablyai#11205's deliberate "CPU" over stablyai#8549's "中央处理器"). Key count and order are unchanged in all five catalogs. en.json's own recipesHelp was reverted by the same stale base and no upstream commit has touched it since, so it is restored to match the live source at EphemeralVmsPane.tsx:252. * fix(rpc): keep null in the ui.set value-parity guard NonNullable stripped null as well as undefined, so dropping .nullable() from a `| null` field passed the guard while still rejecting the batch at runtime -- the exact drift class the guard exists to catch. Proven: making visibleWorkspaceHostIds non-nullable typechecked clean before, now errors by name. Also pins the 15s silent-shell budget so it cannot silently shrink back toward the short deadline. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…tablyai#11249) After a successful headless update the CLI installs source-repo HEAD, which legitimately runs ahead of any shipped bundle. The scan classified those bytes 'unrecognized', so the row went amber ('may be modified… remove it') seconds after our own Update button ran, and the advice looped: remove + reinstall lands the same newer content. Scan half: a canonical/alias placement whose observed git tree sha equals the updater lock's skillFolderHash is the CLI's own install, not a user edit — reclassify it 'newer-known'. Display half: 'newer-known' is recognized official content ahead of this build with nothing to fix, so it no longer marks the copy blocked. Eligibility is deliberately unchanged: ahead of the bundle means there is nothing this build can update to, and offering one risks the provably-unperformable update (stablyai#11110) when source HEAD still equals the lock. Copies whose sha does not match the lock, copies with no lock entry, same-name copies outside the placements the command writes, and plugin-cache behavior all stay flagged exactly as before.
…clobber (stablyai#11230) * fix: stop notification loss, credentialed cache reuse, and clipboard clobber Mobile catch-up (stablyai#8591): fetchMissed swallowed the RPC failure while deliverLive kept advancing and persisting lastDeliveredSeq, so the next successful catch-up asked from above the abandoned range and the desktop cut it. Sessions are module-scope, so an unchanged epoch never resets it. Quarantine the watermark at the last contiguously-delivered seq and hold it there until some later catch-up actually drains — not just one retry. A batch cut short by a teardown quarantines at the last event it settled. Jira attachment cache: currentEpoch summed two independent counters, so a site at siteEpoch 1 read the same value before and after a global clear. The mid-flight guard passed and re-inserted credentialed image bytes that "disconnect all" had just purged — resident for the process lifetime since pruneExpired has no timer. One monotonic ticker, compared by max. Web copy fallback: the handler registered in the capture phase, so xterm's bubble-phase listener overwrote text/plain with the terminal selection afterwards; served was already true, so the copy reported success. Every Orca copy affordance over plain HTTP (Copy Pane ID, Copy Path, commit SHA, PR URL) pasted the terminal selection. Bubble phase with stopImmediatePropagation. Covers the secure-context retry branch too, which shares the same helper. * fix: roll back the persisted watermark on catch-up failure; cover stopImmediatePropagation Adversarial review of a98d7f4d5d found two gaps. 1. The quarantine clamped only writes made AFTER the failure. getMissedSince waits up to 30s, so a live event routinely persists a higher seq while the request is still outstanding; that value stayed on disk, and the next launch read it back and resumed past the abandoned range -- the original bug, reached through a restart. quarantineCatchUpWatermark now re-persists the clamped seq, so the stored value never outlives the gap it guards. 2. web-clipboard-copy-terminal-selection's second test registered its "late" document handler BEFORE the fallback's, so it lost on registration order alone and stopImmediatePropagation was never exercised -- the test passed with that line deleted. Bubbling reaches the document before the window, so a window-level listener is what actually requires it. * fix(mobile): mark a notification seen only once its show lands A pre-marked seen key made a rejected show unrecoverable: the next catch-up re-fetched the seq and the dedup guard dropped it, and the first later event to drain the batch lifted the quarantine past it. Also contains the rejection so it does not escape the un-awaited 'ready'/live handlers as an unhandled rejection. Co-authored-by: Orca <help@stably.ai> * test(web-clipboard): pin stopImmediatePropagation with a same-target handler Both existing cases passed with plain stopPropagation, and with the listener back in the capture phase — neither half of the fix was actually pinned. The window-level clobber is on a different target, so stopPropagation suppresses it too. Registering the clobber on the document, ordered after the fallback's own listener, is the only shape stopPropagation cannot cover. Addresses the review comment posted after the last commit. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…tablyai#11232) * fix(plugins): close trust-boundary holes in the plugin system Move five security decisions to their chokepoints rather than leaving them enumerated at individual call sites. - Kill-list revocation reaches content packs: PluginContentPackRegistry now takes an isKilled predicate and intersects it with any caller-supplied approval, so a killed plugin's VM recipes can no longer reach spawn(..., { shell: true }) through either reconcile() call site. - Bound kill-list generatedAt to a 24h future skew at the parse chokepoint. A far-future timestamp previously made every genuine later list look "older" and disabled revocation permanently, persisted across restarts. - Protect the whole auto.components.settings.Plugin* translation subtree instead of an enumerated prefix list, so language packs cannot forge the consent provenance badge or rewrite install-error security copy. - Resolve manifest panel icons by own-key only; "constructor"/"__proto__" previously yielded non-component prototype members that crashed the right sidebar to its error boundary. - Give panel liveness frames a reserved control budget so a panel that saturates its action budget can still answer the watchdog. Co-authored-by: Orca <help@stably.ai> * fix(plugins): keep the kill-list future bound off the cache read path The schema-level generatedAt bound re-judged the on-disk cache against the device clock at every launch, so a client whose clock ran behind the last genuine publication discarded its whole cached kill list and started with zero revocations. Move the bound to the two fetch chokepoints instead. Co-authored-by: Orca <help@stably.ai> * fix(plugins): remove the reserved-lane starvation window and the revocation TOCTOU Review follow-ups on the trust-boundary fixes: - The reserved liveness lane had a per-window count equal to the ping interval, so a panel's own pong-shaped traffic could spend it and drop the next genuine reply — reintroducing the starvation the lane exists to prevent. The lane is now size-bounded only; rate stays bounded because every pong is also charged to the data budget. - Only schema-valid pongs take the lane now, so near-miss pong-shaped junk cannot drain it. readPanelPongId replaces the zod parse on this guest-controlled path (a rejected safeParse allocates an issue list, ~90x the accepted-path cost) and is pinned to the schema by a parity test. - Re-read the kill list inside approveAtomically: approvedKeys is snapshotted before an awaited verification phase, so a plugin killed during that wait could still publish VM recipes and language packs. - Assert the curated icon resolves to FileText; the old equality also passed when both sides fell back to Plug. Co-authored-by: Orca <help@stably.ai> * fix(plugins): match zod's safe-integer bound in the pong reader readPanelPongId used Number.isInteger, but zod's .int() rejects anything above 2**53-1, so pingIds like 1e100 took the reserved lane the schema would have refused. The parity test never probed that boundary. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
Register Z.ai ZCode (GLM) as a launchable TuiAgent with PATH detection (`zcode`), agent catalog / labels, telemetry agent_kind, and managed status hooks that POST to /hook/zcode. Hook installer writes zcode-hook.sh under ~/.orca/agent-hooks and registers Claude-compatible lifecycle events in ~/.zcode/cli/config.json (hooks.enabled + hooks.events), bridging ZCode's local subprocess hook protocol to Orca's HTTP endpoint. Residual blockers (documented for follow-up; not fixed here): - ZCode hooks are process/command only (no native HTTP/webhook). The managed script is the bridge; status only works when ZCode actually fires configured hooks. - zai-org/feedback#32: native/built-in ZCode agent may not fire hooks.events; only external CLI sub-agents currently trigger them. - Desktop installs often lack a standalone `zcode` binary on PATH (app-only); detection then leaves the agent available but not auto-picked until the CLI is present or a cmd override is set. Closes stablyai#10564
CodeRabbit: SessionStart is a documented ZCode lifecycle event, and remove() left hooks.enabled forced true after managed install. Stash the pre-install value and restore it when managed hooks are removed.
innocarpe
force-pushed
the
fix/zcode-first-class-agent
branch
from
July 29, 2026 01:03
83fa941 to
4b8066a
Compare
Owner
Author
Sync update (
|
CodeRabbit/Greptile: unterminated string on openClaudeHookService import broke the whole test file; indent falls-through comment for kimi→zcode. Co-Authored-By: Grok Companion <noreply@x.ai>
Owner
Author
Sync update (
|
Owner
Author
|
Closing this portfolio mirror because upstream PR stablyai#10654 is superseded by current-main refresh PR stablyai#13965. The active ZCode implementation is now reviewed in that newer upstream lane, so this mirror no longer needs to remain open. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upstream
Summary
Summary - Register ZCode as a
TuiAgent(detect/launchzcode, catalog, labels, telemetry). - Managed hook bridge: installzcode-hook.shand wire ZCode local process hooks to Orca's HTTP listener (POST /hook/zcode). - Claude-compatible lifecycle events normalize to `agenNote
innocarpe/orcamainuntil the upstream PR is merged.