Conversation
* perf(mobile): gate host polling on foreground The mobile host screen ran two 3s polls (routed + embedded), each firing worktree.ps AND repo.list, with no foreground/background gate — so a connected phone kept pinging every 3s (worktree.ps is a full multi-repo process scan) plus a radio wakeup, including brief background windows while the socket stays parked. Consolidate both into one startHostWorktreeRefresh lifecycle and AppState-gate the interval so BOTH polls stop while backgrounded and refresh immediately on foreground return. worktree.ps keeps its 3s cadence while foregrounded (it carries live agent status/preview/unread that no push event replaces). repo.list stays on the interval as an AppState-gated, self-throttling (REPO_METADATA_REFRESH_MS=60s) convergence safety-net — desktop Settings repo edits notify only the renderer, not the runtime clientEvents stream, so it can't be made purely event-driven without going stale — and additionally gets a reposChanged/worktreesChanged fast-path and reconnect-replay refetch. Verified in a deps-installed mobile checkout: full mobile suite 2232 pass, typecheck, oxlint (within the frozen max-lines budget), and oxfmt --check all clean. Co-authored-by: Orca <help@stably.ai> * chore(mobile): drop stale fetchRepoMetadata dep from the reconnect effect Address CodeRabbit nitpick: the reconnect effect no longer calls fetchRepoMetadata (that refetch moved into startHostWorktreeRefresh), so it shouldn't remain in the effect's dependency array. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Orca <help@stably.ai>
stablyai#9979) * fix(mobile): keep quick-commands button steady while capabilities load The tab-row quick-commands button only rendered once the capability probe resolved true, so it popped in after the row was already visible (and vanished during reconnect re-probes). Render it whenever support is not confirmed absent and disable it until the probe settles — pre-quick-commands hosts strip agentPrompt, so the action (not the button) must wait for confirmation. Confirmed-unsupported hosts still hide it entirely. * fix(mobile): explain unsupported quick commands on tap instead of hiding Per feedback on the disabled/hidden states: the button now always renders and stays tappable. Tapping against a desktop that confirmed no support shows "Desktop update required for quick commands" (mirroring the browser streaming copy); tapping while the capability probe is still resolving says to try again in a moment. The sheet still opens only once support is confirmed, since pre-quick-commands hosts strip agentPrompt. * docs(pr): add QA screenshots for quick-commands button states * test(mobile): lock quick-commands button stability Add a focused source-contract test for the always-mounted tab action and confirmed-support sheet gate. Keep the non-obvious safety comment concise, and remove PR screenshots now hosted as GitHub user attachments. * test(mobile): structurally guard quick-command action mount
…9843) react-markdown's <Markdown> has no internal memoization: it rebuilds the whole unified remark->rehype->highlight->katex processor and re-parses the document on every render. MarkdownPreview re-renders on internal state that does not affect the rendered output — most visibly, every keystroke in Find (query/match-index state) — so a large doc re-ran the full parse + syntax-highlight + KaTeX pass per keypress, making Find laggy. Hoist the two fully-static plugin arrays to module scope (a fresh array identity per render would defeat the memo) and render the body through a React.memo'd MarkdownBody keyed on content + components. The pipeline now re-runs only when the rendered content or the components map actually changes; Find/review-pulse/copied- note re-renders skip it. The components map was already memoized, so its identity is stable across those re-renders. Behavior unchanged: 106 existing MarkdownPreview tests pass. Co-authored-by: Orca <help@stably.ai>
…i#9842) The iOS MJPEG and Android scrcpy device streams are gated only on the pane being the active tab (isActive, PR stablyai#7382). When the emulator tab is frontmost but the whole Orca window is hidden/minimized/occluded/display-asleep, the full-fps pipeline keeps running: main-process socket read + JPEG/H.264 decode + IPC + renderer decode. Renderer background-throttling (stablyai#9395) cannot stop it because the pipeline is IPC-push driven from main. Gate showStream additionally on window visibility via a new occlusion-safe hook that honors the terminal stale-visibility latch (so a display-sleep occlusion wedge can't freeze the emulator on a black frame) and delays the visible->hidden park by 500ms so a quick Cmd+Tab round-trip doesn't renegotiate the device stream. Co-authored-by: Orca <help@stably.ai>
…#9995) Zone.js patches the global Promise with a non-native thenable. When a bare `new Promise(...)` crosses the Electron executeJavaScript boundary, it's serialized as-is, losing { page, target } and exposing __zone_symbol__* fields instead. Wrap in an async IIFE to return a native promise that Electron always unwraps correctly.
…blyai#9985) Production crash diagnostics measured ~128 `git worktree list` execs/min (9,400 in one 80-minute session, ~16% of wall-clock in git subprocesses): the resolved-worktree scan fans out over every registered repo on a 30s cache TTL, and most registered repos on the affected installs were agent-CLI scratch repos (~/.codex-tmp capsules, vendor imports, skill checkouts) that need no freshness. Classify agent-scratch repo roots with a curated shared matcher and stamp their scan-cache entries with a 5-minute TTL instead of 30s. Orca-driven mutations still bypass the TTL via the per-repo generation bump, so only passive pickup of external changes slows for scratch repos. Expected steady-state reduction on the measured install: ~82% fewer git spawns.
…7936) (stablyai#9826) * fix(daemon): retire macOS daemons whose login session died (stablyai#7936) A daemon that survives a full macOS logout is unsalvageable: its PAM context can no longer host login(1) spawns (every new PTY becomes a 'Login incorrect' prompt zombie) and its Mach bootstrap namespace has lost the system DNS resolver, so terminals it hosts have no egress. Today it also keeps the stablyai#9301 preflight's cached 'accepted' verdict, so it keeps wrapping spawns in login(1) forever; only a manual daemon restart recovers. GUI-spawned daemons now watch for login-session death from the inside: a fresh cache-bypassing PAM probe (triggered by PTY-exit bursts, fresh client hellos, and a slow periodic timer) must conclusively reject three consecutive times AND the in-process system resolver must be degraded; then the daemon exits crash-style so session meta stays unclean and the replacement daemon cold-restores scrollback. A conclusive rejection also flips the spawn-wrapper cache off immediately. Headless serve/SSH daemons never get the watch (they must survive their spawning session ending), and a session that never conclusively accepted login(1) never arms it — a PAM anomaly alone can't kill a healthy daemon (fast user switching keeps accepting, so switched-away sessions are preserved). * test(daemon): e2e seam to drive login-session death oracles from a verdict file A dead macOS login session cannot be fabricated without root (PAM owns audit-session teardown), so live lifecycle QA drives the death watch's probe and resolver oracles from ORCA_E2E_LOGIN_SESSION_PROBE_FILE: 'alive' → accepted/healthy, 'dead' → rejected/unhealthy, anything else inconclusive — with compressed watch timing. Mirrors the existing ORCA_E2E_DAEMON_INIT_DELAY_MS seam; inert unless the env var is set. * fix(daemon): close the hang-shaped gap in login-session death detection The conclusive-PAM-verdict trigger had one blind failure shape: login(1) hanging at the prompt past the probe bound (killed → inconclusive forever → the watch never fires). Three changes close it: - The death-watch probe gets its own 4s bound (the 500ms preflight bound exists for spawn-path latency, which doesn't apply off-path), so a slow-but-answering PAM stack isn't misread as a hang. - An inconclusive pipe probe escalates to a PTY-hosted probe via script(1) — a dead session's PAM stack may only misbehave under a real tty (the pipe-vs-PTY fidelity limit the preflight documents). - A streak of timeout-killed probes (which a live session never produces) is a second retirement trigger, at a higher threshold (5) and still gated on the degraded resolver, logged with a distinct cause so field logs discriminate the two paths. Every dead-session behavior — fast reject, prompt-then-EOF, or hang — now fires retirement; all inconclusive states still fail toward preserving the daemon. * fix(daemon): keep login-session retirement conclusive * fix(daemon): stop login watch before clean shutdown * fix(daemon): make login-session PTY probe reliable * fix(daemon): close login-session watch races * fix(daemon): ignore health probes for login watch activity
…ai#9980) * test(e2e): verify Claude is prefilled with issue URL on start Regression test for stablyai#6613: when starting a workspace from a newly created GitHub issue, ensure the issue URL is passed to Claude via `--prefill` and `--dangerously-skip-permissions` flags. This prevents context loss after issue creation. * test(e2e): fix GitHub-issue-start prefill test flakiness - Reorder mock API handlers to ensure `/labels` and `/assignees` paths match before the specific issue endpoint - Replace regex heading matcher with exact string for more reliable assertions - Refactor terminal content polling to capture text once and reuse in subsequent assertions
* fix(terminal): defer remote output ACKs until parse * test(terminal): document synchronous credit claims --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
…tablyai#10000) The workspace-options filter row used the Workflow icon while the Automations nav item and page use CalendarClock. Match them so the filter clearly maps to automation-created workspaces.
… a remote runtime is active (stablyai#6628) * fix(worktrees): refresh local worktrees in the sidebar while a remote runtime is active When a remote runtime is active, a local `worktrees:changed` event for an unbound repo was dropped by the renderer guard in useIpcEvents. Worktrees created outside Orca for that repo (e.g. `orca worktree create` from a CLI or automation flow) therefore stayed invisible in the sidebar until an app restart, even though their sessions were already running. The guard existed because an unbound repo's list fetch routes to the active runtime (settingsForKnownRepoOwner's unbound fall-through), so refreshing with local worktree ids could query — and purge against — the remote host. Instead of dropping the event, pin the refresh to the local host (forceLocalOwner): fetch the worktree list against the local owner and merge additively. The merge is host-scoped and the deletion-purge is skipped on this path, so it only ever adds local-host worktrees and never overwrites the active runtime's worktree state. A genuinely-removed local worktree is reclaimed by the next unguarded full refresh. * test(e2e): regression — CLI-created worktree visible while a remote runtime is active Drives the real `orca worktree create` path: the CLI RuntimeClient calls `worktree.create` over the app's socket, registering a managed worktree and firing the `worktrees:changed` IPC the renderer listens for. Stages a remote runtime as active by injecting `activeRuntimeEnvironmentId` into the renderer store, so no real remote host is needed. Fails on the prior behavior (the worktree never appears while a runtime is active) and passes with this fix. * fix(worktrees): pin local lineage refresh during runtime activity Co-authored-by: Orca <help@stably.ai> * review: trim comments to house style, normalize queue coalescing to booleans * review: sweep rename-grace expiry before early returns in worktrees:changed handler * review: document accepted workspace-space gap, drop imprecise 'additive' wording * fix(worktrees): route duplicate local repo events locally * fix(worktrees): tag local worktree events at origin, gate purge skip on runtime overlap * test: pin origin-based forceLocalOwner with a no-runtime local event assertion --------- Co-authored-by: brennanb2025 <brennankbenson@gmail.com> Co-authored-by: Orca <help@stably.ai> Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
…tablyai#9996) * feat(agent-status): show question glyph for needs-you state everywhere Replace the amber attention dot with the dashboard's MessageCircleQuestion chat glyph for the waiting/permission "needs you" state across all surfaces: the agent dashboard cards, in-app dashboard rows, the sidebar's Agent Dashboard quick-indicator counts, the worktree-level status dot, and the shared agent-row/terminal-tab indicators. On the dashboard card the header glyph is suppressed when a question summary pill is present so "needs you" reads once, not twice. * test(agent-status): assert amber question glyph, not amber dot The needs-you unification replaced the amber dot with the amber MessageCircleQuestion glyph, so update the remaining state assertions in DashboardAgentRow, WorktreeCardStatusSlot, and TerminalTabLeadingIcon to match (lucide-message-circle-question + text-amber-500). * test(agent-status): cover needs-you glyph surfaces
…ablyai#9883) * perf(web): pause runtime heartbeat while hidden Co-authored-by: Orca <help@stably.ai> * fix(web): preserve inbound-liveness baseline across a hidden heartbeat re-arm The visible re-arm rebaselined lastInboundFrameAt=now, so a socket that went silent while the window was hidden looked freshly-heard-from and its death was masked for another full idle window (~25s). Move the fresh-connect baseline into startHeartbeat (the real 'we just connected' moment) and have the visible re-arm only reset the tick clock + clear an in-flight probe, preserving lastInboundFrameAt so the next visible tick probes a stale connection promptly and closes if unanswered. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…nected (stablyai#9885) * perf(runtime): idle websocket heartbeat without clients Co-authored-by: Orca <help@stably.ai> * fix(runtime): probe immediately when the WS heartbeat arms Arming the heartbeat on the first accepted connection started a fresh interval, so the first liveness ping was a full interval (~15s) out — a socket that died right after connecting went unprobed for that window. Run one sweep synchronously in start() so the first ping goes out at arm time; the seeded socket is pinged (never reaped on the arm sweep) and reaped on the next tick only if it never pongs. Tests updated for the earlier first probe. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…blyai#9888) * perf(mobile): coalesce overlapping home requests Co-authored-by: Orca <help@stably.ai> * fix(mobile): queue a trailing follow-up for triggers during an in-flight read Single-flight returned the in-flight promise to any trigger that arrived mid-read, so a distinct refresh requested while a slow read was on the wire was silently answered by the older response and never re-read the latest state (UI could stay one refresh cycle stale). Coalesce mid-flight triggers into exactly one trailing follow-up (latest params win) whose fresh result is delivered to those callers. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…ingle-flight (stablyai#9892) * fix(mobile): gate dictation setup polling Co-authored-by: Orca <help@stably.ai> * fix(mobile): fence a stale dictation refresh against a newer setPolling intent An in-flight setup read resolving 'keep polling' after an explicit setPolling(false) wrote polling=true and rescheduled, resurrecting a poll the caller had just stopped. Snapshot a pollingRevision when each read starts and only apply its result if no explicit setPolling superseded it mid-flight — so a late true can't restart a stopped poll (nor a late false cancel a restart). Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…ablyai#10007) * Enable accessibility tree (`ax`) command on iOS emulator sessions Fetch the accessibility tree from serve-sim's /ax endpoint, which requires an active session but provides the same UI snapshot capability as Android's uiautomator output. Derive the endpoint from the stream URL when not explicitly provided by the helper, and route through the bridge to pass session context to the backend. * Add ax command routing and backend integration tests Tests verify accessibility tree routes through EmulatorBridge, Android backend ignores iOS-specific ax URLs, and ax endpoints are derived from serve-sim stream URLs.
…rant lane (stablyai#10001) * feat(telemetry): classify codex trust-grant fallbacks and attribute grant lane * fix(telemetry): tighten codex trust-grant classification
Allow the existing "Open in" entries to launch a configured VS Code
launcher against an SSH-backed worktree via Remote-SSH:
code --remote ssh-remote+<authority> <remote-path>
- Split the blanket SSH/runtime block into a capability model: file
managers and non-VS Code launchers stay local-only (disabled with
"Local only" metadata); a recognized VS Code command is enabled and
forwarded with connectionId over a typed object IPC.
- Main process stays authoritative: rejects active/owned runtimes,
resolves the SshTarget from the persisted Store, derives the authority
(config alias, or username@host on port 22, or ssh-alias-required on a
non-default port), validates POSIX/Windows absolute remote paths without
local stat/normalize, and rejects non-VS Code and compound commands
before spawn.
- Authority and remote path are passed as separate argv; getSpawnArgsForWindows
remains the cmd/bat shim boundary and fails closed on metacharacters.
- Same capability rules across the worktree menu, Explorer overflow, and
the source-control entry context menu.
Refs STA-2386
Closes stablyai#9999
…ar (stablyai#8047) * fix(rate-limits): use session cookies for OpenCode Go * fix(rate-limits): harden OpenCode session setup --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> Co-authored-by: OrcaWin <alpha-eng@stably.ai>
…ablyai#9475) * fix(codex): promote [tui] settings so they survive the managed-home remirror Codex TUI preferences (/statusline, theme, terminal title) are written into the [tui] table of the managed runtime config.toml, but the write-back promotion allowlist only covered four top-level scalars — so the next mirror pass rewrote the runtime config from ~/.codex and silently discarded them. Extend promotion to the [tui] keys the Codex TUI persists (status_line, status_line_use_colors, terminal_title, theme), keyed as structured tui.* paths so the same three-way merge (runtime vs baseline vs ~/.codex) applies: in-Codex changes promote into ~/.codex before the mirror, and outside edits to ~/.codex still win over stale runtime values. The byte-preserving upsert moves to codex-config-settings-upsert.ts (max-lines) and learns [tui] placement: replace an existing bare or dotted key in place, insert into the first [tui] body, insert dotted beside existing dotted tui.* keys, or create one [tui] table at EOF — never defining tui twice, including when the system config holds an inline tui = {...} table. * Add codex-config-settings-upsert to the CLI tsconfig file list * fix(codex): keep tui upserts out of array tables * fix(codex): handle quoted tui config paths during promotion * fix(codex): harden tui promotion writes
… c5d87873) (stablyai#10024) Co-authored-by: Orca <help@stably.ai>
…stablyai#10484) * fix(terminal): verify Windows PTY root identity before taskkill /T /F killWithDescendantSweep guarded its Windows tree kill with ownsRoot() alone, which is JS state only. node-pty's ConPTY exit watcher closes the last shell handle before it queues the JS exit callback, so Windows can recycle the PID while the session map still looks live — force-killing an unrelated process and its whole descendant tree. Walk the recycled PID's ancestry back to this process before taskkill: skip the sweep when the root is gone or resolves to a stranger, and keep the sweep when identity is unknown so stablyai#10004 orphan cleanup still runs. Also gate the local provider's ownsRoot on observed physical exit. * fix(terminal): dedupe the Windows root-identity scan, drop dead exit gate Review fixes on the PID-identity guard. The probe read the process table through a new uncached export, bypassing the reader that worktree teardown depends on: worktree-teardown.ts fans out 32-wide inside a 10s deadline, so a delete forked 32 powershell cold-starts (the churn windows-foreground-process-rows.ts:25-32 warns about, stablyai#6288/stablyai#6667). getFreshSnapshot() already guarantees a scan that starts after the request -- the exact property the bypass existed for -- and coalesces concurrent callers, so use it. Measured on the new test: 32 scans -> 1. The PhysicalExitTracker.hasExited gate could never fire. markExited() is only reached at local-pty-provider.ts:985/:1431, and both are followed synchronously by clearPtyState(), which deletes the ptyProcesses entry -- so ownsRoot's map check is already false whenever hasExited is true. Reverting it broke no test. Drop it and the shared getter it added; the identity probe already covers every ownsRoot caller from inside killWithDescendantSweep. Also point the Windows terminal-restart E2E job at the files that own this behavior, so a change to the new Windows-only module runs the one job that executes on a real Windows host. * docs(terminal): state what the Windows root probe actually proves The probe checks subtree membership, not root identity: a recycle that lands on another Orca descendant (another pane's shell, an agent CLI, a git.exe we spawned) still reads `own`, and that is not remote during teardown when Orca is itself allocating pids. It bounds the blast radius rather than closing the class. Say so at the type and at classifyWindowsTreeKillTarget, and name what a real close would need (a CreationDate baseline -- the analogue of the POSIX lstart check already used here -- or an inherited handle / Job Object). Also note why our own pid must classify `foreign`. * ci(windows): trigger the terminal-restart E2E on the shared snapshot reader The Windows root-identity probe now reads through getFreshSnapshot, so an edit to that module changes Windows teardown behavior without touching any path the job already watches. * test(terminal): guard the teardown probe against a reintroduced scan bypass The existing volume guard covers queryWindowsProcessRowsFresh directly, but the identity-probe cases all inject readRows, so nothing exercised the DEFAULT reader wiring -- a bypass reintroduced inside windows-pty-root-identity would have gone unnoticed. Drive verifyWindowsTreeKillTarget 32-wide through the real reader and assert one scan. Verified it fails at 32 when the bypass is put back.
…stablyai#10478) * fix(daemon): keep agent-completion detection alive on pre-v27 daemons DaemonPtyAdapter.inspectProcess() threw terminal_liveness_unavailable when the connected daemon predated protocol v27. The intended provider-level fallback only fires when a provider lacks inspectProcess, so for the daemon adapter the throw propagated: agent-completion-coordinator swallowed it into consecutiveInspectionErrors and retried forever, killing process-exit completions and pending-title validation. Daemons intentionally survive app updates, so updating in place with agent terminals open routes those PTYs to a legacy adapter and permanently disables agent-finished notifications until the terminal is recreated. Compose the inspection client-side from getForegroundProcess, which v26 fully supports. No new wire traffic and no new daemon capability. * test(daemon): pin null-foreground semantics on the pre-v27 inspect fallback The legacy composition had no coverage for a null foreground, which is the one daemon response shape whose semantics diverge from v27: there inspectProcess goes through getAliveSession() and throws for a vanished session, while getForegroundProcess is deliberately null-not-throw. It is also the only shape that reaches a user-visible completion, so reading it as idle is a deliberate choice that should not change silently.
…ckpoints (stablyai#10479) * fix(terminal): stop cold restore dropping all scrollback on large checkpoints The checkpoint read cap shipped without its write bound. history-reader.ts reads checkpoint.json through a 16MiB cap that throws past the limit, and the catch swallows it to checkpoint=null; history-manager.ts still writes the checkpoint with an unbounded JSON.stringify. Every fallback then collapses (stale-generation log, unlinked legacy scrollback.bin), so the terminal reopens empty with nothing surfaced. Raise the read cap to cover the largest checkpoint the writer can legitimately emit, derived from the scrollback policy's 50k-row max preset so it cannot drift back under the writer. A bound is kept so a corrupt file still cannot OOM the main process. Bounding the writer instead would not recover the scrollback: the stringify throw lands in handleWriteError, which adds the session to disabledSessions and permanently stops history recording for it. * fix(terminal): anchor checkpoint read cap to its own reasoning The cap was derived as 2 * LEGACY_TERMINAL_SCROLLBACK_BYTES_100_MB, but that constant is a legacy byte-preset setting value with no other consumer, and the 50k-row bucket it was attributed to has no upper byte bound. Same value, stated without the false policy linkage. * fix(terminal): assert the checkpoint byte cap, correct its rationale The retained oversized-checkpoint test passed with the byte guard removed entirely — 200MB of NUL fails JSON.parse, so detectColdRestore returned null either way. Assert the bounded reader directly, as the amplification test does. Also: a 50k-row max preset of ordinary text measures ~14MB serialized, not 'far below' by an unbounded margin — per-cell-colored output can still exceed the cap, which is what a writer-side snapshot trim has to fix.
…nmounts (stablyai#10480) * fix(mobile): heal an orphaned native-chat image paste across screen unmounts The stale-input marker lived in a per-screen `useRef`, but the condition it tracks — a bracketed image paste sitting unsubmitted on the agent's composer line — lives on the host and outlives the screen. Backing out of a session and returning remounted the hook with an empty Set, so the next message submitted on top of the orphaned paste and the agent received `<image path><text>`. Move the marker to a module-level store keyed by terminal handle, and consult and consume it from every write path that can submit the composer: the image hook's text-only send, the controller send (which the chat overlay's question card reaches directly, bypassing the image hook), and the ask-answer send. Permission choices and the Escape cancel deliberately do NOT heal: they are `enter: false` keys for an active overlay that swallows the clear, so healing there would consume the marker without clearing the line and leave the next real message corrupted. Desktop scopes its Ctrl+U the same way. * fix(mobile): stop the ask heal from burning the marker on selector answers The heal ran on every ask answer, but Claude's and Codex's selector shapes cannot submit the composer: a single-select answer is a bare option digit and every stepping group is written `enter: false` (the host coerces it), so the clear is swallowed by the live overlay while the host still acks the write. That consumed the one-shot marker and left the orphaned paste to corrupt the next real message — the same failure this PR exists to fix, through a new door that main did not have. Scope the heal to the pasted-label shape, which does commit the composer. Desktop splits it the same way: use-native-chat-interactive-send.ts routes only the non-stepping answer through the clearing sender and never pre-clears sendNativeChatAskAnswer. Also pin the three deliberate skips (selector answer, permission choice, Escape cancel) with tests, so the PR's central design argument is an invariant rather than a comment, and guard the failed-heal toast with the generation check every other error surface in answerAsk already uses.
…tablyai#10464) * fix(cli): explain SIGABRT serve exits instead of naming the signal (stablyai#10461) `orca serve` reported only "Orca serve exited via SIGABRT", which sent a P0 investigation down a code-signature path while a diagnostic crash report sat unread on disk. On darwin + SIGABRT the signal-exit path now names the macOS application-startup abort, its usual sandbox/SSH/CI causes, and points at ~/Library/Logs/DiagnosticReports/Orca-*.ips via the existing nextSteps channel. Other platforms and signals get a clear message with no invented cause. * fix(cli): stop asserting the SIGABRT exit happened at startup * fix(cli): stop steering macOS SIGABRT users away from SSH serve
…ns freeze (stablyai#10483) * fix(skills): advance the release ledger at the cut so shipped revisions freeze stablyai#10340 made the released-skill registry a function of the committed ledger instead of a git tag walk, and stablyai#10460 reverted the cut step that advances that ledger because it violated the stablyai#9119 contract (a version-only cut must not regenerate or stage the content-addressed skill artifacts). Both were right; the result is a ledger that never advances. generate-skill-bundle-manifest.mjs:390 derives releasedCount solely from release-mapping.json and :461 assigns a changed skill releaseRevision = releasedCount + 1, while :518 protects only committedReleasedCounts[name] — so index releasedCount is unprotected. A tag ships that tail revision, nothing records it, and the next skill change rebuilds the same revision number over different bytes. Installs carrying the shipped digest then match no snapshot and degrade to unrecognized, which cannot be updated. Restore the advance in a form the stablyai#9119 contract can keep enforcing: --release now verifies that current-manifest.json and snapshot-registry.json already match the ref being tagged, appends the mapping row, and writes only release-mapping.json. The cut stages just that file, so it still cannot move a content-addressed artifact — the failure stablyai#9119 guarded against — and now fails loudly instead of recording a revision the tag does not ship. The contract test is narrowed to match: it asserts the cut runs --release (never --write) and stages exactly package.json and release-mapping.json. * test(release-cut): close the staging bypasses the narrowed gate left open The narrowed contract test anchored its `git add` scan to line start and only inspected staged paths, so three ways to reintroduce stablyai#9119 stayed green: a `git add` chained after `&&`, a write that never calls `git add` at all, and `pnpm run generate:skill-bundle-manifest` — the package.json alias for `--write`, which the hyphenated ban never matched. That last one also passed the pre-stablyai#10460 assertions, so it was never covered. Drop the anchor, require every `resources/skills` mention in the step to be exactly what is staged, and ban the alias and `commit -a`. Comments are stripped first so prose cannot trip a ban. Verified each bypass fails and the real workflow passes. * fix(release-cut): make the new provenance failure actionable to an operator Verifying the content-addressed artifacts is the only new way the cut can block, and it fails inside a step named "Bump package.json and tag" with a lint-shaped message. That names the files and the command but not the two things the operator needs: the regeneration has to land on main, and the cut is safe to re-run afterwards. Say so. Also pin down why assertReleasedHistoryPreserved takes the pre-append mapping. It pairs with artifacts.releasedSnapshotCounts, which seeding fixed before the row existed; handing it the post-append mapping makes every cut throw "Released snapshot history is incomplete", which points at tag fetching rather than the real cause. Nothing enforces the pairing. * test(release-cut): gate the whole cut job, not just the bump step Round-2 review defeated the previous gate twice, both proved by running the full contract file green with stablyai#9119 reintroduced. Every step in the cut job shares one workspace and one index, but the contract test only inspected `Bump package.json and tag`. A step inserted earlier could run --write and `git add resources/skills`, and the bump step's own commit swept it into the version commit and the tag. Assert job-wide instead: only the bump step may name the directory, and no step may regenerate under either the flag or its package.json alias. That lives in the generator suite because the contract file is at its max-lines cap. Two regexes were also evadable. The mention scan required a trailing slash, so a path held in a variable was invisible; it now matches the directory itself. The `commit -a` ban matched nothing at all — `commit\s` ate the only separator, so `-a`, `-am`, and `--all` all survived while only a trailing `-a` was caught. `--allow-empty` stays allowed. * fix(release-cut): assert the index, not the workflow text, before committing Round-3 review defeated the job-wide grep three ways, each proved by running both test files green with stablyai#9119 reintroduced into the tagged commit: an `env:` block holding `--write` and `resources/skills`, a composite action whose steps the workflow never spells out, and plain shell concatenation (`root=resources; leaf=skills`). Grepping shell source for path literals is inherently evadable, and the previous fix only relocated round-2's variable-indirection hole one step over. Move the invariant to where it cannot be dodged: immediately before committing, the cut diffs its own index and refuses anything that is not package.json or the release-mapping row. That does not care which step staged what, or how the path was spelled. The workflow grep stays as a cheap tripwire for literal spellings, now paired with a positive assertion that the index guard exists and precedes the commit — indirection cannot hide a missing guard. Mention matching dedupes and trims quotes, since the guard names the row a second time. * fix(release-cut): match the staged-path allowlist literally `grep -vx` treats its patterns as regexes, so the `.` in `package.json` matched any character: a staged `packageXjson` or a `resources/skills/release-mappingXjson` was silently accepted by the index guard. Verified both slip through `-vx` and are caught by `-vxF`. Exercised the guard against a legitimate cut, an empty index, a staged content-addressed artifact, paths containing a space and a non-ASCII character (git quotes the latter, so it fails closed), and a staged deletion. Only the two allowed paths pass. * test(release-cut): assert the index guard aborts, not just that it exists The positive assertion pinned the guard's shape and its position before the commit, but not its effect: replacing `exit 1` with `:` left both test files green while the cut logged the error and shipped the artifact anyway. That is the same failure this whole gate keeps having — asserting the shape of a defense rather than what it does. Pin the abort too. Verified the neutered guard now fails the suite. * test(release-cut): scope the abort check and catch clustered commit flags Two holes in the guards this PR added, both in the same shape-not-effect class the previous commit was meant to close. The abort assertion's lazy match was not scoped to the guard's own block, so it could borrow an `exit 1` from any later `if ... fi` in the step. Degrading the guard to a warning while adding a plausible HEAD precondition left every test green. Stop the match at the guard's `fi`. The `commit -a` ban only matched when `a` led the flag cluster, so `-vam`, `-va`, `-qam` and `-sam` all survived. That matters more than it looks: `commit -a` stages at commit time, after the index guard has already inspected a clean index, so it is the one way to defeat that guard. Match `a` anywhere in a short-flag cluster; `--allow-empty` and `--amend` stay allowed. Verified both mutants now fail. * fix(release-cut): validate the commit, not the index, before tagging The index guard asserted the wrong thing. `git commit` has a family of forms that commit the working tree rather than the index — `-a`, `-i`, `--only`, and a bare pathspec — so a rogue earlier step could leave regenerated artifacts unstaged and any of those forms would carry them into the tagged commit while the guard saw a clean index and passed. Reproduced end to end: `git commit -i resources` put current-manifest.json and snapshot-registry.json in the tag with all gates green, and `--only resources` additionally dropped package.json from the tag. Banning those flags one by one is the same enumeration game the earlier rounds kept losing. Assert the outcome instead: after committing and before tagging, diff-tree HEAD and refuse anything that is not package.json or the release-mapping row. That is indifferent to which step staged what and to how the commit was spelled. Verified the whole family is now blocked (-i, --only, -a, -am, -vam, pathspec, and an alias expanding to `commit -i`), that a stock commit and an --allow-empty re-cut still pass, and that deleting, neutering, un-anchoring, or relocating the guard each fails the suite. * fix(release-cut): make the commit guard fail closed on a merge commit Plain `git diff-tree` prints nothing for a merge commit, so the guard would have passed silently instead of failing closed — the one direction that matters on a release path. `-m --first-parent` reports the diff against the first parent; verified byte-identical output for an ordinary commit and still empty for the `--allow-empty` re-cut, so nothing else changes. Not reachable today (nothing in the cut job creates a merge, and npm version has no lifecycle hooks defined), but the failure mode is a guard that looks like it ran. Pin the flags in the assertion too, so neither dropping -m nor slipping in a `--diff-filter` can weaken it without failing the suite.
stablyai#10500) * test(terminal): pin checkpoint-only cold restore of a large checkpoint Triaging a report that checkpoint-only cold restore renders blank at every checkpoint size found no v1.4.156 regression: an A/B of v1.4.155 against origin/main returned byte-identical ColdRestoreInfo at 1.01, 5.56, 11.30 and 20.61 MiB, all 300 marker lines intact on both refs. Blankness tracks the meta.endedAt eligibility gate, not size — a cleanly ended session refuses to cold-restore at any size, which is by design and unchanged between the refs. What the triage did surface is a coverage gap. stablyai#10179's 16MiB checkpoint read cap sat under what the unbounded writer emits, so a large checkpoint threw, was swallowed to checkpoint=null, and the pane reopened empty; stablyai#10479 raised the cap but nothing pinned the round trip it had broken. history-reader-memory covers the bounded reader at its limit, not writer→reader recovery. Adds that round trip through the real HistoryManager writer and HistoryReader over a header-only log, and pins the endedAt gate that has now been mistaken for a size regression twice. Fails at the pre-stablyai#10479 cap with the exact production symptom (detectColdRestore returns null). * test(terminal): fail loudly when writeSync is unavailable writeLargeScrollback ignored writeSync's boolean, which is false when xterm's private _core.writeSync goes away. The size assertions catch that in the large-checkpoint test (416 bytes vs 16MiB), but the endedAt gate is size-independent, so that test silently passed on an empty snapshot — pinning the gate over no scrollback at all. * test(terminal): cut large-checkpoint fixture peak RSS from 1.2GiB to 800MiB The filler colored per line, not per cell as its comment claimed, so the serialized seed only tracked plain-text size and needed 26k buffer rows to clear 16MiB. The xterm buffer costs rows x cols, which made this single file raise the whole src/main/daemon/ suite's peak RSS 5.2x (241MiB -> 1244MiB) — a real OOM risk on CI right after stablyai#10299 bounded readers for that reason. Carrying the bytes in SGR runs instead of rows reaches a larger seed from 5.5k rows: suite peak RSS 1244MiB -> 793MiB, and the margins improve too (seed 25.16MiB = 1.57x the threshold vs 1.19x before, checkpoint.json 46.1MiB = 2.88x vs 1.20x). Also makes the fixture match its own comment and be more representative of real colored agent output. Re-verified all three mutations still fail: cap at 16MiB -> "expected null not to be null"; endedAt gate deleted -> test 2 fails; writeSync unavailable -> both fail. * test(terminal): tighten timeout, reuse prod dir-name helper, dispose first Review follow-ups, all test-only: - Import getHistorySessionDirName instead of hand-rolling encodeURIComponent. Equivalent today, but that helper exists to absorb encoding changes, so the hand-rolled copy would silently diverge from the writer it is checking. - Index FILLER_ROWS by its own length so growing the array cannot leave rows unused. - 300_000ms -> 60_000ms. The file runs in ~2.8s and vitest's own default is 30s; a 5-minute ceiling turns a hung regression into a stalled job rather than a failure. 60s keeps Windows headroom. - Take the snapshot, dispose, then checkpoint, so a throwing checkpoint() cannot leak the buffer. Note this is memory-neutral, not a saving: peak RSS is set at getSnapshot(), where the buffer and the serialized strings coexist, and measured ~800MiB either way. Mutations re-verified: cap at 16MiB -> "expected null not to be null"; endedAt gate deleted -> test 2 fails; writeSync unavailable -> both fail.
…ad path (stablyai#10499) * perf(terminal): drop the JSON structural pre-scan from the history read path readTerminalHistoryJson/Async walked every character of checkpoint.json in interpreted JS before handing the same string to native JSON.parse. On a 26.8MB checkpoint that scan measured 63-330ms — 2.7-12x the JSON.parse it guards, and 57% of the whole read path — all of it synchronous main-thread work. readTerminalHistoryJsonAsync exists so cold-restore reads do not block the main thread; running the scan inline right after the async read defeated its own stated purpose. The scan also had a correctness cost: TERMINAL_HISTORY_JSON_MAX_STRUCTURAL_TOKENS was 1_000_000, and oscLinks is unbounded at ~10 structural tokens per link, so roughly 100k OSC-8 hyperlinks tripped the assert, history-reader swallowed the throw to a null checkpoint, and the terminal restored blank — the same user-visible loss stablyai#10479 just fixed for the byte cap. checkpoint.json and meta.json are our own SerializeAddon output, not untrusted input: the byte cap still bounds the read, and a corrupt file fails JSON.parse into the same catch. The shared helper stays for the call sites that do handle untrusted JSON. * test(terminal): pin iterative JSON.parse for the dropped nesting-depth cap The retired pre-scan enforced two limits; the new tests only covered the structural-token half. Dropping the 128-level nesting cap is safe solely because V8 parses JSON iteratively — depth costs heap, not stack — so a deeply nested checkpoint parses instead of overflowing. Verified: 10M-deep arrays and objects parse without throwing on V8 14.6. That property is an engine guarantee this code now silently depends on, and a recursive parser would abort the daemon outright rather than throw into the callers' catch. The test fails on main with "JSON nesting exceeds 128 levels" and costs 4ms. * docs(terminal): drop the rot-prone benchmark figure from the reader comment Keeps the durable rationale for skipping the pre-scan (self-authored input, byte cap still bounds the read, corrupt files land in the callers' existing catch) and moves the "~57% of the read path" measurement to the PR body, where it cannot go stale against a later change to this path. Addresses the CodeRabbit comment-length nitpick against the repo's one-line-if-possible comment guideline.
…o status works over WSL (stablyai#10328) * fix(agent-hooks): install OpenCode status plugin into the WSL guest so status works over WSL OpenCode reports agent status via a JS plugin dropped into OPENCODE_CONFIG_DIR (unlike Claude/Codex, which use managed hooks.json scripts). Over the WSL runtime that plugin was never materialized inside the guest and OPENCODE_CONFIG_DIR never crossed into the guest, so OpenCode status never reached Orca's sidebar (the workspace stayed green). Codex already worked; this was OpenCode-specific. Mirror the SSH plugin-overlay path for WSL: - Guest relay registers AGENT_HOOK_INSTALL_PLUGINS_METHOD, byte-caps the source, and materializes an OpenCode config overlay via PluginOverlayManager (the same electron-free path the SSH relay uses), returning overlayDirs.opencode. - Host manager ships the plugin source over the existing stdio channel after installers run (and on mid-session reinstall) and records the guest overlay dir; -32601 / CONNECTION_LOST / DISPOSED are swallowed like ssh-relay-session. - PTY env points OPENCODE_CONFIG_DIR/ORCA_OPENCODE_CONFIG_DIR at the guest overlay; until the relay reports it (first spawn / older guest bundle) it drops those vars rather than crossing the Windows overlay path into WSL — so in-guest OpenCode falls back to its own config (pre-fix behavior, no regression). - WSLENV passes OPENCODE_CONFIG_DIR/ORCA_OPENCODE_CONFIG_DIR through (/u for guest paths). SSH is untouched: the same JSON-RPC constant is reused and the guest response merely gains an optional overlayDirs field the SSH host ignores. Runtime repro: native opencode launched in a WSL Orca terminal had ORCA_AGENT_HOOK_PORT/ORCA_PANE_KEY but no OPENCODE_CONFIG_DIR and no Orca plugin in ~/.config/opencode, so agentStatusByPaneKey stayed empty. * fix(agent-hooks): stop the WSL OpenCode overlay leaking Windows paths and churning under running agents Review fixes on top of the WSL OpenCode plugin install: - Never cross a Windows OPENCODE_CONFIG_DIR into the guest. The /p flag was not a defensive default but WSLENV's translate-and-deliver flag, and buildWslRelaySpawnEnv spreads process.env while the daemon merge resurrects keys buildPtyHostEnv only deleted -- so a Windows value reached the guest as /mnt/c/... and was adopted as its OpenCode config root. Register the two vars only when the value is already a guest POSIX path. - Make guest materialization idempotent. The overlay id is instance-scoped, and the host re-ships on every reinstall (60s after connect, and on later pane spawns), so materializeOpenCode's remove-and-rebuild wiped the config root under running agents and raced panes spawning against the path just handed to them. Rebuild only when the shipped source changed or the overlay went missing. - Mirror the guest's default ~/.config/opencode (honouring XDG_CONFIG_HOME) when no explicit dir is discoverable, so pointing OPENCODE_CONFIG_DIR at the overlay no longer silently drops the user's models/agents/skills/mcp. - Carry opencodeOverlayDir across relay relaunch; it is instance-keyed and on the distro's persistent filesystem, so dropping it only blanked status on panes spawned mid-relaunch. * fix(agent-hooks): stop advertising a WSL OpenCode overlay the guest failed to rebuild Round-2 review fixes: - materializeOpenCode wipes before rebuilding, and every failure path after the wipe returns null leaving the dir present but plugin-less. The host treated that null the same as "no handler / teardown" and silently kept the previous value, so a pane could be pointed at an empty config root -- worse than the documented fallback of dropping the var. requestGuestOpenCodeOverlayDir now distinguishes 'none' (guest answered, no dir) from 'unavailable', and the manager clears the recorded dir on 'none'. - The handler cache keyed only on plugin source, so a ~/.config/opencode created after the relay connected was never mirrored for the relay's lifetime. Key on the resolved source dir too, and validate the cache by the plugin file rather than the directory -- the directory is exactly what a failed rebuild leaves behind, so checking it alone made the bad state stick. * fix(agent-hooks): don't mirror the XDG default OpenCode config into the WSL overlay OPENCODE_CONFIG_DIR is APPENDED to OpenCode's config-dir list, not a replacement for it. Verified against the shipped binary: the list is built as [Path.config, ...project .opencode dirs, ...OPENCODE_CONFIG_DIR ? [it] : []], and Path.config is derived independently from XDG_CONFIG_HOME/$HOME/.config. So ~/.config/opencode is read whether or not Orca overrides the var, and the earlier fallback that mirrored it into the overlay made OpenCode load the user's config -- and their plugins -- twice. Resolve only an explicitly-set dir, which is the one case that genuinely leaves the list when Orca overwrites the variable. This also restores parity with the SSH and local paths. * docs(agent-hooks): correct the WSL install-plugins cache comment and test framing The per-call source-dir re-resolution comment still described the XDG default branch that 6745eba removed. The relay's env is fixed for its lifetime and the rc scan behind it is memoized, so the sourceDir cache key is defensive rather than live -- say so, and retitle the test that simulates it by mutating the captured env, so neither reads as coverage of a production scenario.
…hboard (stablyai#10531) The companion-board mutual-exclusion Effect keyed on `workspaceBoardRenderedOpen`, which is `workspaceBoardOpen || workspaceBoardDragPreviewOpen`. WorktreeList sets the drag preview at the start of every card drag — including a pure reorder within a group that never opens the board — so dragging any card closed the Agent Dashboard drawer, and cancelling the drag never restored it. Key the Effect on `workspaceBoardOpen` so only the user actually opening the board evicts the dashboard. The reciprocal Effect is unchanged.
… as freezes (stablyai#10530) * feat(diagnostics): record OS suspend/resume so sleep gaps aren't read as freezes Renderer timers stop across OS sleep, so a multi-hour `renderer_memory` heartbeat gap looks identical to a wedged renderer. That ambiguity sent the uber-crash investigation down a deadlock path that the telemetry later disproved -- and a healthy machine's own trace shows a 427-minute mid-session gap with reason=interval, so the gap alone proves nothing either way. powerMonitor 'resume' was already wired for renderer wake recovery but left no breadcrumb. Stamp suspend and report the measured span on resume so the next freeze report can be told apart from a laptop lid. * fix(diagnostics): only record sleeps long enough to hide a heartbeat Adversarial review flagged that the first cut would flood the 30-entry breadcrumb ring and evict the crash evidence it exists to explain. Measured on 7 days of pmset history: 70 user-visible sleep cycles, worst 60-min burst of 7. Median span is 2 SECONDS -- only ~24% run past 60s, so most of that traffic could never explain a gap anyway. Cycles are counted Sleep -> next FULL Wake, since powerMonitor's resume maps to NSWorkspaceDidWake, which does not fire for dark wake. Drop the suspend breadcrumb (suspend now only stamps a timestamp) and emit a single `system_slept` on resume, gated at 60s. That cuts 70 cycles to 17 over the same week (worst burst 3) while still catching every sleep long enough to swallow a 60s renderer heartbeat. * test(diagnostics): assert resume listeners detach by identity The off mock deleted by event name alone, so teardown detaching a different closure than the one registered still passed -- a leak of the real powerMonitor listener would have gone unnoticed. For 'suspend' that leak has no other observable effect through the public API. Co-authored-by: Orca <help@stably.ai> * fix(diagnostics): span from the first suspend across dark wake powerMonitor 'resume' maps to NSWorkspaceDidWake, which does not fire for dark wake, so macOS can deliver suspend -> suspend -> resume. Overwriting the stamp reported only the trailing segment, and when that segment fell under the 60s gate a 90-minute sleep recorded nothing at all -- leaving the gap looking like the unexplained freeze this is meant to rule out. Co-authored-by: Orca <help@stably.ai> * docs(diagnostics): correct the threshold rationale to match measurement The comment claimed maintenance sleeps would flood the ring. Re-measuring pmset over 6 days (78 sleeps, 29 full wakes, 51 dark wakes) shows they resolve as DarkWake, which never fires 'resume', so they never recorded a breadcrumb at all. Real rate is 29 breadcrumbs / 6 days, worst 60-minute burst 4 against a 30-entry ring. The gate's actual job is narrower: skip sleeps shorter than the 60s heartbeat, which cannot open a gap to explain. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…s hidden (stablyai#10528) startGitCommonNarrowWatch was the only watch entry point that never received WorktreePollerWindowVisibility, so its `worktrees/` existence poll kept stat'ing every repo without a linked worktree forever in the background (0.5 stat/sec/repo at the 2s default). Its siblings — the primary-metadata snapshot poller and the non-darwin git-common polling — already park. Threads visibility through and matches the snapshot poller's park/re-arm pattern: the poll stops on the first hidden tick, and re-checks immediately on onWindowBecameVisible so a worktrees dir created while hidden still upgrades to the native stream and emits its create event. The visibility listener is dropped in the dispose path. darwin-only: the narrow watch is the `platform === 'darwin'` branch.
…tablyai#10526) * fix(terminal): make primary-selection paste suppression single-shot Middle-clicking in the terminal armed a 750ms window that swallowed every native paste event, not just Chromium's one follow-up — so a real Ctrl+V inside that window was silently dropped on Linux. Consume the deadline on first use (`shouldSuppress*` -> `consume*`) so the arm owes exactly one event; 750ms stays as that event's expiry bound. * test(terminal): cover the paste event xterm actually forwards to the PTY Review follow-ups on the single-shot suppression change; no source change. The new real-module file only synthesized `beforeinput`. An Electron probe (Chromium 150) showed `paste` fires first, is cancelable, and cancelling it at document capture suppresses `beforeinput` entirely — and xterm registers `handlePasteEvent` for `paste` only, never `beforeinput`. So `paste` is the sole event that can double-write the PTY, and it was the one event the end-to-end file did not exercise. Add it; it fails with the fix reverted. Consuming mutates, so `isTerminalNativePasteTarget(...) && consume()` is now load-bearing: swapped operands would burn the arm on an unrelated paste and let the real follow-up double-paste. The guarding test asserted only `defaultPrevented`, which survives the swap — assert `consume` is never reached instead. Rename the mocked-file case that claimed to prove single-shot. With the module mocked its sequence is dictated by the mock; it pins that the hook re-asks per event rather than caching, so name it that. * test(terminal): name the beforeinput dispatcher for the event it dispatches Round-2 review nit on my own round-1 change: once `dispatchClipboardPaste` existed alongside it, a helper named `dispatchPaste` that dispatches `beforeinput` inverted the reader's expectation. Match the sibling file's `dispatchNativePasteBeforeInput` convention.
… locale (stablyai#10536) Settings search indexed the macOS privacy-toggle name as English-only aliases on the LAN keyword, so a Chinese user typing the term macOS System Settings actually shows them (本地网络) got no hit — zh's catalog value had been changed to 局域网 (LAN). Split the two wordings onto their own catalog keys so each locale carries both: 87620e6416 = LAN, fa3239cd42 = Local Network (its true content hash). Localized values come from the repo's own LAN title translations and from macOS 26's SecurityPrivacyExtension Localizable.loctable (LOCAL_NETWORK), so nothing is invented. ja/ko/es already matched macOS and are unchanged apart from gaining the LAN key.
…ai#10543) * feat(diagnostics): record OS suspend/resume so sleep gaps aren't read as freezes Renderer timers stop across OS sleep, so a multi-hour `renderer_memory` heartbeat gap looks identical to a wedged renderer. That ambiguity sent the uber-crash investigation down a deadlock path that the telemetry later disproved -- and a healthy machine's own trace shows a 427-minute mid-session gap with reason=interval, so the gap alone proves nothing either way. powerMonitor 'resume' was already wired for renderer wake recovery but left no breadcrumb. Stamp suspend and report the measured span on resume so the next freeze report can be told apart from a laptop lid. * fix(diagnostics): only record sleeps long enough to hide a heartbeat Adversarial review flagged that the first cut would flood the 30-entry breadcrumb ring and evict the crash evidence it exists to explain. Measured on 7 days of pmset history: 70 user-visible sleep cycles, worst 60-min burst of 7. Median span is 2 SECONDS -- only ~24% run past 60s, so most of that traffic could never explain a gap anyway. Cycles are counted Sleep -> next FULL Wake, since powerMonitor's resume maps to NSWorkspaceDidWake, which does not fire for dark wake. Drop the suspend breadcrumb (suspend now only stamps a timestamp) and emit a single `system_slept` on resume, gated at 60s. That cuts 70 cycles to 17 over the same week (worst burst 3) while still catching every sleep long enough to swallow a 60s renderer heartbeat. * test(diagnostics): assert resume listeners detach by identity The off mock deleted by event name alone, so teardown detaching a different closure than the one registered still passed -- a leak of the real powerMonitor listener would have gone unnoticed. For 'suspend' that leak has no other observable effect through the public API. Co-authored-by: Orca <help@stably.ai> * fix(diagnostics): span from the first suspend across dark wake powerMonitor 'resume' maps to NSWorkspaceDidWake, which does not fire for dark wake, so macOS can deliver suspend -> suspend -> resume. Overwriting the stamp reported only the trailing segment, and when that segment fell under the 60s gate a 90-minute sleep recorded nothing at all -- leaving the gap looking like the unexplained freeze this is meant to rule out. Co-authored-by: Orca <help@stably.ai> * docs(diagnostics): correct the threshold rationale to match measurement The comment claimed maintenance sleeps would flood the ring. Re-measuring pmset over 6 days (78 sleeps, 29 full wakes, 51 dark wakes) shows they resolve as DarkWake, which never fires 'resume', so they never recorded a breadcrumb at all. Real rate is 29 breadcrumbs / 6 days, worst 60-minute burst 4 against a 30-entry ring. The gate's actual job is narrower: skip sleeps shorter than the 60s heartbeat, which cannot open a gap to explain. Co-authored-by: Orca <help@stably.ai> * fix(window): stop viewport reflow from moving screen geometry Blink's ScreenMetricsEmulator::Apply checks screen_size and view_position before the desktop/mobile branch, so screenPosition:'desktop' does not make them inert -- despite Electron documenting both as mobile-only. Passing the content size and 0,0 overrode screen.width/availWidth and the window origin for the whole 32ms hold, so a browser context menu opened mid-reflow would translate against 0,0 and land in the wrong place (BrowserPane.tsx:3167). Empty screenSize means 'no override', and an omitted viewPosition stays nullopt in Electron's converter, so the real position survives. Only the scale factor moves now, which is what the reflow actually needs. Also record a breadcrumb when the restore exhausts its attempt budget: the renderer is left at the wrong scale factor until some later reveal fixes it, and that was previously silent. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>
…i#10525) * fix(release-cut): gate an explicit RC against its own series semver_gt compares through strip_pre(), so the explicit-version override only ever checked the stable line: 1.4.156-rc.0 read as 1.4.156, cleared a 1.4.155 stable, and republished an RC below what clients already run. Anchor a prerelease request on highest_rc_for_base -- the same rc history the kind path uses -- so the override can only advance the series. Two sibling gaps in the same block: - version_suffix was silently dropped when version was set, because the append lives in the kind branch the override skips. - the shape regex rejected X.Y.Z-rc.N.suffix, so a suffixed RC the rc path can produce could never be re-cut explicitly. * fix(release-cut): close both ends of the rc-number range the gate compares The new explicit-rc gate compares with `[[ -le ]]`, i.e. bash machine-width integers, and the author closed only the low end. Past INTMAX bash saturates, so `version=1.4.156-rc.99999999999999999999` reads as "above the published rc.3" and the gate falls open — then the tag it cuts pins highest_rc_for_base at 1e20 for that base forever, and every later cut wraps to a lower rc the fleet never updates to. Bound the rc number to nine digits. Also reject leading zeros on an all-digit prerelease identifier. `npm version` renormalizes rc.4.01 to rc.4.1 while the tag step keeps the literal input, so the shipped package.json version and its own release tag name different releases. The explicit path's embedded identifier now goes through the same validator the kind path uses instead of only the shape regex. * fix(release-cut): stop the refusal pointing minor/major RCs at the wrong series kind=rc derives its base from bump(latest_stable, patch), so the remedy the refusal suggested only works when the requested base *is* that next patch. A 1.5.0-rc.N series exists only because this override created it, so an operator resuming a stuck 1.5.0-rc.2 was told to dispatch kind=rc, which would have cut an unrelated 1.4.156-rc.4. Spell the condition out and give the fallback that does work for a non-patch base. Also correct the mechanism in the comment I added in 698c5be: bash wraps two's-complement, it does not saturate, which is why the hole is value-dependent (rc.10000000000000000000 wraps negative and failed closed, rc.99999999999999999999 wraps to 7766279631452241919 and sailed through). And name both inputs in the suffix error, which now serves version_suffix and the trailing identifier in version. * fix(release-cut): count a suffixed RC from its commit subject, not just its tag The new explicit-version gate only fails closed on a deleted tag because highest_rc_for_base also reads `release: v<base>-rc.N` subjects. That fallback did not parse the suffixed form: rcNumberFromTag accepts an optional .identifier, rcNumberFromReleaseSubject did not, so `4.perf` failed its `(\d+)(\s|$)` anchor and returned null. So deleting a v1.4.156-rc.4.perf tag dropped the series back to rc.3, and an explicit 1.4.156-rc.4 was waved through — below the rc.4.perf build perf-channel clients already run. Same under-count already made kind=rc recompute rc.4 over a deleted suffixed tag. Mirror the tag form's optional identifier. Covered by a unit assertion and a git-fixture test that both fail with this reverted. * docs(release-cut): correct four operator-facing claims in the explicit path All four are wording or consistency, no behavior change (harness: 26/26 before and after, on bash 3.2 and bash 5.2). - The trailing-identifier comment justified itself as preserving a shape that "can never be re-cut through the override", but re-cutting a suffixed rc at or below the series head is exactly what the new gate refuses. State what it actually admits: a second spelling of version=X.Y.Z-rc.N + version_suffix. - version_suffix's input description still said "rc kind only" after this PR made it apply to an explicit bare X.Y.Z-rc.N. - The suffix guard's own rc pattern was unbounded while the shape check twelve lines up is bounded to nine digits; reuse the bounded one so a later edit to either cannot silently drift. - "which recovers the existing tag" was unconditional, but kind=rc recovery is also gated on tag_matches_current_ref, so a tag cut from a ref main has moved past advances to rc.N+1 instead.
* fix(git): read core.sparseCheckout the way git does Sparse-checkout detection parsed git config line-by-line and only accepted a section header alone on its line, so git's legal same-line form `[core] sparseCheckout = true` matched neither branch and was silently skipped: a genuinely sparse worktree lost its badge and partial-checkout warning. It also read `config.worktree` unconditionally, although git honors that file only while extensions.worktreeConfig is on, so a stale worktree config could override the repo's real setting. Headers are now consumed left-to-right off each line (further headers and one assignment may follow), and config.worktree is read only behind the extension gate. Every new expectation was confirmed against real `git config --get`. * test(git): correct what git actually does with a trailing-junk config value Git does not reject `[core] sparseCheckout = true bogus = false` outright: it parses the line and takes the whole tail as one value (`git config --list` reports `core.sparsecheckout=true bogus = false`), then fails only the boolean coercion. The expectation is unchanged; the comment now matches the binary.
…tarve terminal.wait (stablyai#10529) * fix(runtime): reserve long-poll headroom so orchestration.ask can't starve waits orchestration.ask joined the long-poll set, which also opted it into the single server-wide activeLongPolls counter. Because ask blocks on a reply for its full timeout (600 s default, previously unbounded via a caller timeoutMs), 16 asking workers could hold every slot and shed terminal.wait and check --wait with runtime_busy for every other client — mobile, web, CLI, SSH and relay all share this runtime. Meter ask as its own long-poll class with a sub-cap of half the budget, and clamp the caller-supplied timeoutMs at 30 min. The keepalive and abort-signal wiring that motivated the original change is unchanged. * test(runtime): cover the ask sub-cap and counter release on the WebSocket path The admission fence is shared by both transports but only the Unix-socket path was exercised, so a WS-only regression in admitLongPoll/releaseLongPoll would have shipped silently. Drives handleWebSocketMessage with a 'runtime' scoped device (orchestration.ask is absent from the mobile allowlist) and asserts the overflow ask is shed without burning a reserved slot, that check --wait still gets the other half, and that both counters return to zero when the socket closes.
…cker (stablyai#10527) * fix(tasks): keep repos with a pending remote-identity probe in the picker Task-repo eligibility filtered on `hasProjectRemoteIdentity`, which is populated by a background `git remote -v` probe. When the probe could not reach git — an SSH-hosted repo whose connection is not up yet, a cold launch — the repo silently vanished from the Tasks picker and stayed hidden for the full 5-minute negative-cache TTL, even after the host came back. GitHub repos were largely shielded because a persisted `upstream` satisfies the identity projection through a different route; GitLab and other providers depend on the probe. Distinguish unknown from settled instead of hiding both: - `probeGitRemoteIdentity` reports `resolved` / `no-remote` (git answered, no usable remote) / `unavailable` (never reached git). - Enrichment persists `gitRemoteIdentity: null` only on `no-remote`, mirroring the existing `upstream: null` "not a fork" marker. An unreachable host leaves the identity undefined. - Persistence keeps the explicit `null` instead of dropping it. - `getTaskEligibleRepos` keeps a repo whose identity is still pending; folders and settled remote-less repos stay filtered out. * test(tasks): cover the SSH probe exec paths for remote-identity status Addresses CodeRabbit review: the unavailable-on-error case only exercised the local git runner. Adds a connected-provider whose exec rejects, and an SSH repo git answered for with no remotes. * test(tasks): pin that a settled no-remote repo still resolves once it gains a remote Three independent reviewers flagged that the candidate filter's `!repo.gitRemoteIdentity` looks like an oversight next to the new null marker. Tightening it to `=== undefined` would silently stop detecting a remote added after the marker landed. Document that the re-probe is deliberate and pin the behavior with a test.
…stablyai#10547) The old anchored regex matched neither branch on a `[section "sub"]key = value` line, so the parser never left `[core]` and credited the next indented line to it — reporting sparse for a worktree git says is not. Fails on the pre-fix parser (returns true where git reports unset).
…t freeze workspace creation (stablyai#10540) * fix(worktree): bound the .worktreeinclude copy so a huge include can't freeze creation `.worktreeinclude` copying was bounded in entry count (1000) but unbounded in bytes and files, and awaited inline during worktree creation. A repo listing `node_modules` froze creation for minutes behind the create dialog on Linux and Windows, where the fallback is a full `fs.cp` (macOS gets a cheap APFS clone). Measure each copy-mode source against a cumulative budget (2 GB / 50k files) before the first byte is written, and refuse the entries that bust it. Refused entries ride the existing `CreateWorktreeResult.warning` channel so a workspace never silently comes up missing its included files. Pre-measurement rather than mid-copy abort: `fs.cp` ignores its `signal` option, so a started copy cannot be cancelled and would strand a partial tree. Refusing up front means there is no partial state to clean up. * fix(worktree): don't charge bytes for copy-on-write clones, and bound the sizing walk Two defects in the copy budget, both found by review: - The byte limit was applied on macOS, where the copy is an APFS clonefile. Measured: a 2.7 GB tree clones in 22 ms and consumes no disk. Refusing it on a 2 GB byte ceiling denied work that was already free — a regression on the one platform this bound was never meant to touch. Bytes are now charged only when a byte-for-byte copy will actually run; the volume probe that decides this is the same cached df+diskutil pair the clone runs, and writes nothing, so the "refuse before the first byte" invariant holds. The entry limit still applies everywhere: inodes are real work even on the clone path. - A refused entry consumed no budget, so a `.worktreeinclude` listing many over-budget directories paid a fresh full-limit walk for each one — up to 1000 x 50,000 lstat calls, re-creating the stall this bounds. The walk is now charged against its own ceiling whatever the verdict. Also documents that `admit()` must be awaited sequentially (CodeRabbit). * fix(worktree): give the sizing walk headroom so one huge entry can't starve the rest The walk ceiling added in the previous commit was seeded with maxEntries, the same number the entry limit uses. Sizing an entry that busts the file-count limit walks maxEntries + 1, driving the ceiling negative, so every later `.worktreeinclude` entry was refused without being measured at all. That regressed the common case: a repo listing `node_modules` plus `.env` used to get `.env`; it silently got nothing. Reproduced, and now covered by a test that fails when the headroom is removed. The walk now gets 5x the entry budget, so total sizing work stays bounded (<=250k lstat per materialization, vs the 1000 x 50k this ceiling exists to prevent) while ordinary lists never reach it. Entries refused because earlier ones exhausted the walk report a distinct 'sizing' reason, so the warning stops quoting size limits at a 4-byte file that was never measured. * fix(worktree): bill a failed clone's bytes, and blame the right ceiling Two follow-on defects from the copy-on-write fix: - A predicted APFS clone that then failed mid-copy (EPERM, ENOSPC) fell through to a real `fs.cp` whose bytes were never charged, because the entry had been admitted on the premise that cloning is free. That reopened the unbounded copy on macOS. The measured size is already known, so the fallback now bills it and refuses if it no longer fits, reporting the entry as skipped instead of silently copying gigabytes. A clone that was never viable (ApfsCloneUnavailableError) was already charged as a real copy, so that path keeps falling back as before. - The walk ceiling is also applied inside the measurement via min(remainingEntries, remainingWalk), and when the walk term bound, the refusal was still reported as 'entries' — telling the user a 3-file directory busted a 4-file limit. It now attributes to whichever ceiling actually bound. Also fixes the singular warning text, which said "entry X was not copied ... copying them would exceed ... Copy them in manually". * fix(worktree): flag a partial clone leftover, cap the warning, cover two branches - A clone that fails partway only removes an *empty* reservation, so leftovers can survive at the target. Reporting that entry as simply "not copied" sent the user to copy it in manually, straight into a half-populated directory. Those skips now carry mayBePartial and the warning says to check the path first. Cleaning up the leftovers stays the deferred follow-up it already was. - The warning enumerated every skipped path. `.worktreeinclude` allows 1000 entries and all of them can be skipped, so it now names five and counts the rest — an unbounded string is a poor look in a PR about bounds. - Two load-bearing branches had no test, both proven by surviving mutants: the `bytesAreCopied` short-circuit (reachable when a wedged df/diskutil makes the volume probe answer "no clone", so bytes are charged up front and must not be billed twice), and chargeBytes actually consuming budget for later entries. * fix(worktree): only flag directory clones as partial, and cap that list too - mayBePartial was set for every refused clone fallback, but only a *directory* clone can leave anything behind: the file path clones into a temp name and publishes with link(2), so a failure leaves nothing at the target. Sending the user to inspect a path that does not exist is its own small lie. - The partial-copy sentence sliced to five names without the "and N more" that the other sentence appends, so entries past the fifth were surfaced nowhere. Both sentences now share one nameList helper.
…t misreported (stablyai#10550) The 30-min clamp was silent: the ask result carried no timeout figure, so the CLI printed the value the caller *sent*. A worker passing --timeout-ms 3600000 was told "ask timeout after 3600000ms" after only 30 min of real waiting — off by 2x, and accurate before this PR added the clamp. Echo the effective budget on every ask return and print that. Additive optional field; older clients fall back to the requested value.
The Note textarea only grew on onInput, so programmatic PR title prefills left the box at one row with overflow hidden. Resize whenever the note value changes and allow scroll under max-height (fixes stablyai#10575).
… pass The prefill bug is the failure mode of imperative sizing: the height is only recomputed at the events someone remembered to hook, so a programmatic setNote (and a pane resize, and a font reflow) leaves it stale. Let the layout engine own the height, matching NativeChatComposerField and LinearIssueTextEditor. Drops the extracted helper and its unit test — that test asserted the two assignments it wrote and stayed green with the bug present. The composer test now pins the class contract and goes red on the pre-fix markup.
Owner
Author
|
Upstream PR stablyai#10580 has been merged, so this portfolio mirror is no longer needed. Closing the stale mirror; the merged upstream PR remains the source of truth. |
Owner
Author
|
Upstream stablyai#10580 is merged; keeping this portfolio mirror open would duplicate the completed contribution. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upstream
Summary
Summary Fixes stablyai#10575 — Create Worktree Note no longer clips long PR titles after Smart-tab prefill. Prefill uses
setNote(...)(noonInput), so the old height logic never ran and the field stayed atrows={1}withoverflow-hidden. ### Changes - Autosize the Note tNote
innocarpe/orcamainuntil the upstream PR is merged.