Skip to content

[upstream #10588] fix(terminal): make Zellij/TUI OSC 52 clipboard copy work by default - #77

Closed
innocarpe wants to merge 280 commits into
mainfrom
fix/zellij-osc52-clipboard
Closed

innocarpe wants to merge 280 commits into
mainfrom
fix/zellij-osc52-clipboard

Conversation

@innocarpe

Copy link
Copy Markdown
Owner

Portfolio mirror of my contribution to upstream stablyai/orca.
Exhibition only — the real review/merge target is upstream.

Upstream

Summary

Summary - Inside Zellij (and similar multiplexers), copy goes through OSC 52 to the host terminal. Orca had OSC 52 default off, so Zellij copy failed while bare-shell copy still worked (stablyai#10567). - Empty Pc (\e]52;;base64) is a valid XTerm default meaning clipboard

Note

  • Do not merge this into innocarpe/orca main until the upstream PR is merged.
  • After upstream merges: sync fork from upstream, then close this mirror PR.
  • This open PR exists so visitors see in-flight work on this fork's Pull requests tab.

nwparker and others added 30 commits July 22, 2026 16:07
…tablyai#10000)

The workspace-options filter row used the Workflow icon while the
Automations nav item and page use CalendarClock. Match them so the
filter clearly maps to automation-created workspaces.
… a remote runtime is active (stablyai#6628)

* fix(worktrees): refresh local worktrees in the sidebar while a remote runtime is active

When a remote runtime is active, a local `worktrees:changed` event for an
unbound repo was dropped by the renderer guard in useIpcEvents. Worktrees
created outside Orca for that repo (e.g. `orca worktree create` from a CLI or
automation flow) therefore stayed invisible in the sidebar until an app
restart, even though their sessions were already running.

The guard existed because an unbound repo's list fetch routes to the active
runtime (settingsForKnownRepoOwner's unbound fall-through), so refreshing with
local worktree ids could query — and purge against — the remote host.

Instead of dropping the event, pin the refresh to the local host
(forceLocalOwner): fetch the worktree list against the local owner and merge
additively. The merge is host-scoped and the deletion-purge is skipped on this
path, so it only ever adds local-host worktrees and never overwrites the active
runtime's worktree state. A genuinely-removed local worktree is reclaimed by
the next unguarded full refresh.

* test(e2e): regression — CLI-created worktree visible while a remote runtime is active

Drives the real `orca worktree create` path: the CLI RuntimeClient calls
`worktree.create` over the app's socket, registering a managed worktree and
firing the `worktrees:changed` IPC the renderer listens for. Stages a remote
runtime as active by injecting `activeRuntimeEnvironmentId` into the renderer
store, so no real remote host is needed. Fails on the prior behavior (the
worktree never appears while a runtime is active) and passes with this fix.

* fix(worktrees): pin local lineage refresh during runtime activity

Co-authored-by: Orca <help@stably.ai>

* review: trim comments to house style, normalize queue coalescing to booleans

* review: sweep rename-grace expiry before early returns in worktrees:changed handler

* review: document accepted workspace-space gap, drop imprecise 'additive' wording

* fix(worktrees): route duplicate local repo events locally

* fix(worktrees): tag local worktree events at origin, gate purge skip on runtime overlap

* test: pin origin-based forceLocalOwner with a no-runtime local event assertion

---------

Co-authored-by: brennanb2025 <brennankbenson@gmail.com>
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
…tablyai#9996)

* feat(agent-status): show question glyph for needs-you state everywhere

Replace the amber attention dot with the dashboard's MessageCircleQuestion
chat glyph for the waiting/permission "needs you" state across all
surfaces: the agent dashboard cards, in-app dashboard rows, the sidebar's
Agent Dashboard quick-indicator counts, the worktree-level status dot, and
the shared agent-row/terminal-tab indicators.

On the dashboard card the header glyph is suppressed when a question
summary pill is present so "needs you" reads once, not twice.

* test(agent-status): assert amber question glyph, not amber dot

The needs-you unification replaced the amber dot with the amber
MessageCircleQuestion glyph, so update the remaining state assertions in
DashboardAgentRow, WorktreeCardStatusSlot, and TerminalTabLeadingIcon to
match (lucide-message-circle-question + text-amber-500).

* test(agent-status): cover needs-you glyph surfaces
…ablyai#9883)

* perf(web): pause runtime heartbeat while hidden

Co-authored-by: Orca <help@stably.ai>

* fix(web): preserve inbound-liveness baseline across a hidden heartbeat re-arm

The visible re-arm rebaselined lastInboundFrameAt=now, so a socket that went
silent while the window was hidden looked freshly-heard-from and its death was
masked for another full idle window (~25s). Move the fresh-connect baseline into
startHeartbeat (the real 'we just connected' moment) and have the visible re-arm
only reset the tick clock + clear an in-flight probe, preserving lastInboundFrameAt
so the next visible tick probes a stale connection promptly and closes if unanswered.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
…nected (stablyai#9885)

* perf(runtime): idle websocket heartbeat without clients

Co-authored-by: Orca <help@stably.ai>

* fix(runtime): probe immediately when the WS heartbeat arms

Arming the heartbeat on the first accepted connection started a fresh interval,
so the first liveness ping was a full interval (~15s) out — a socket that died
right after connecting went unprobed for that window. Run one sweep synchronously
in start() so the first ping goes out at arm time; the seeded socket is pinged
(never reaped on the arm sweep) and reaped on the next tick only if it never pongs.
Tests updated for the earlier first probe.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
…blyai#9888)

* perf(mobile): coalesce overlapping home requests

Co-authored-by: Orca <help@stably.ai>

* fix(mobile): queue a trailing follow-up for triggers during an in-flight read

Single-flight returned the in-flight promise to any trigger that arrived mid-read,
so a distinct refresh requested while a slow read was on the wire was silently
answered by the older response and never re-read the latest state (UI could stay
one refresh cycle stale). Coalesce mid-flight triggers into exactly one trailing
follow-up (latest params win) whose fresh result is delivered to those callers.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
…ingle-flight (stablyai#9892)

* fix(mobile): gate dictation setup polling

Co-authored-by: Orca <help@stably.ai>

* fix(mobile): fence a stale dictation refresh against a newer setPolling intent

An in-flight setup read resolving 'keep polling' after an explicit setPolling(false)
wrote polling=true and rescheduled, resurrecting a poll the caller had just stopped.
Snapshot a pollingRevision when each read starts and only apply its result if no
explicit setPolling superseded it mid-flight — so a late true can't restart a stopped
poll (nor a late false cancel a restart).

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
…ablyai#10007)

* Enable accessibility tree (`ax`) command on iOS emulator sessions

Fetch the accessibility tree from serve-sim's /ax endpoint, which requires an
active session but provides the same UI snapshot capability as Android's
uiautomator output. Derive the endpoint from the stream URL when not explicitly
provided by the helper, and route through the bridge to pass session context to
the backend.

* Add ax command routing and backend integration tests

Tests verify accessibility tree routes through EmulatorBridge,
Android backend ignores iOS-specific ax URLs, and ax endpoints
are derived from serve-sim stream URLs.
…rant lane (stablyai#10001)

* feat(telemetry): classify codex trust-grant fallbacks and attribute grant lane

* fix(telemetry): tighten codex trust-grant classification
Allow the existing "Open in" entries to launch a configured VS Code
launcher against an SSH-backed worktree via Remote-SSH:

    code --remote ssh-remote+<authority> <remote-path>

- Split the blanket SSH/runtime block into a capability model: file
  managers and non-VS Code launchers stay local-only (disabled with
  "Local only" metadata); a recognized VS Code command is enabled and
  forwarded with connectionId over a typed object IPC.
- Main process stays authoritative: rejects active/owned runtimes,
  resolves the SshTarget from the persisted Store, derives the authority
  (config alias, or username@host on port 22, or ssh-alias-required on a
  non-default port), validates POSIX/Windows absolute remote paths without
  local stat/normalize, and rejects non-VS Code and compound commands
  before spawn.
- Authority and remote path are passed as separate argv; getSpawnArgsForWindows
  remains the cmd/bat shim boundary and fails closed on metacharacters.
- Same capability rules across the worktree menu, Explorer overflow, and
  the source-control entry context menu.

Refs STA-2386
Closes stablyai#9999
…ar (stablyai#8047)

* fix(rate-limits): use session cookies for OpenCode Go

* fix(rate-limits): harden OpenCode session setup

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
Co-authored-by: OrcaWin <alpha-eng@stably.ai>
…ablyai#9475)

* fix(codex): promote [tui] settings so they survive the managed-home remirror

Codex TUI preferences (/statusline, theme, terminal title) are written into
the [tui] table of the managed runtime config.toml, but the write-back
promotion allowlist only covered four top-level scalars — so the next mirror
pass rewrote the runtime config from ~/.codex and silently discarded them.

Extend promotion to the [tui] keys the Codex TUI persists (status_line,
status_line_use_colors, terminal_title, theme), keyed as structured tui.*
paths so the same three-way merge (runtime vs baseline vs ~/.codex) applies:
in-Codex changes promote into ~/.codex before the mirror, and outside edits
to ~/.codex still win over stale runtime values.

The byte-preserving upsert moves to codex-config-settings-upsert.ts (max-lines)
and learns [tui] placement: replace an existing bare or dotted key in place,
insert into the first [tui] body, insert dotted beside existing dotted tui.*
keys, or create one [tui] table at EOF — never defining tui twice, including
when the system config holds an inline tui = {...} table.

* Add codex-config-settings-upsert to the CLI tsconfig file list

* fix(codex): keep tui upserts out of array tables

* fix(codex): handle quoted tui config paths during promotion

* fix(codex): harden tui promotion writes
Coalesce identical and sub-pixel fallback overlay measurements so ResizeObserver and xterm fit cannot sustain a render feedback loop, while preserving precise committed geometry.

Adds regression coverage for stable measurements, sub-pixel jitter across integer boundaries, and genuine resizes.
…+ add glibc/libstdc++ packaging gate (stablyai#9902) (stablyai#10019)

* fix(linux): restore Ubuntu 20.04 launch by pinning node-pty glibc symbols (stablyai#9902)

The bundled node-pty pty.node is compiled from source in release CI on
ubuntu-latest (glibc 2.39). glibc's 2.32-2.34 libpthread/libutil merge
relocated openpty/forkpty (GLIBC_2.34) and pthread_sigmask (GLIBC_2.32)
into libc under new symbol versions, so the from-source build bound to
versions absent on Ubuntu 20.04 (glibc 2.31). The main process imports
node-pty at startup, so the app crashed on launch. pty.node is the sole
blocker (Electron needs GLIBC_2.25; other native modules <= 2.17).

- Patch node-pty: a .symver shim pins the 3 symbols to their pre-merge
  version (GLIBC_2.2.5 x64 / GLIBC_2.17 arm64), and Linux-only ldflags
  force libutil.so.1/libpthread.so.0 back into DT_NEEDED. Guarded to
  Linux; macOS/Windows untouched.
- Add a packaging gate (verify-linux-glibc-floor.cjs, afterPack): reads
  each bundled native binary's objdump -p version needs and fails the
  Linux build if any strong GLIBC_/GLIBCXX_/CXXABI_ node exceeds stock
  Ubuntu 20.04 (glibc 2.31 / GLIBCXX_3.4.28 / CXXABI_1.3.12). Catches
  GLIBC_ABI_DT_RELR, rejects GLIBC_PRIVATE, skips weak needs, fail-closed.
- Docs + tests; the lazy sherpa-onnx speech prebuilt (GLIBCXX_3.4.29,
  never loaded at launch) is a documented libstdc++-floor exemption.

* fix(linux): assert DT_NEEDED provider deps in the glibc-floor gate

Harden the packaging gate (flagged in adversarial re-eval): the version-floor
check alone can false-pass if the patch's forced `-l:libutil.so.1` ever silently
drops — the pinned openpty@GLIBC_2.2.5 still resolves from libc's compat alias at
build time, but fails to load on Ubuntu 20.04 where openpty/forkpty live only in
libutil. The gate now also asserts that any binary importing openpty/forkpty
keeps libutil.so.1 in DT_NEEDED. Validated on a real symver-pinned .so with
libutil dropped (now fails) vs. present (passes). Documents the recommended
real-host smoke-test follow-up.
Update desktop experimental settings, mobile settings/onboarding, i18n
(en/zh/ja/ko/es), and user-visible error strings. Keep internal APIs and
identifiers as nativeChat.
…, not failed (stablyai#10021)

* fix(mobile): report interrupted native chat sends as delivery-unknown, not failed

A terminal.send interrupted mid-flight showed a definite "Message not sent"
even when the desktop may have already delivered the text. Three paths were
misclassified as definite failures:

- Logical relay/direct cutover: migrateTo rejects in-flight requests with
  LogicalClientCutoverError, which mapped to 'rejected'. Now maps to 'unknown'
  (held unconfirmed + transcript-echo verification; never retried since
  terminal.send is non-idempotent).
- Suspend/close of a half-open session: the stable logical client blanket-
  rejected in-flight pendings with plain 'Client suspended'/'Client closed',
  preempting the physical layer's delivery-unknown marking. It now lets the
  physical close settle them, so post-write failures stay marked and pre-write
  failures stay definite.
- Relay path: mobile-relay-rpc-session never marked delivery ambiguity at all
  (timeout, close, link failure). Post-write rejections are now marked;
  pending entries only exist after the frame reached the authenticated link.

Permission, ask-answer, and cancel-Escape surfaces now show "unconfirmed —
check chat before retrying" instead of a definite "not sent" on ambiguous
outcomes (still not-accepted, never retried). Also consolidates a private
copy of isLogicalClientCutoverError in worktree-create-retry.

Co-authored-by: Orca <help@stably.ai>

* chore(skills): regenerate skill-bundle manifest artifacts

---------

Co-authored-by: Orca <help@stably.ai>
…#10040)

The daemon health-check guard logs during main-process startup, which can
complete before the renderer window resolves. Moved stderr listening to the
launch options so early logs aren't missed. Also made the assertion regex
pattern-based instead of exact-string matching to tolerate benign log
rewording, and added a check that the replace path stayed off.
…tablyai#9988)

* fix(mobile): keep native chat from resizing the covered terminal PTY

Native chat reads the agent transcript stream and never renders the
terminal grid, but two paths still pushed phone dimensions into the
covered PTY, reflowing the desktop terminal for no benefit:

- The covered lease-only subscribe carried the cached viewport, and
  handleMobileSubscribe phone-fits the PTY whenever a viewport is
  present. The lease now omits the viewport so the host keeps the
  desktop baseline and late-binds on return to the terminal tab.
- useTerminalViewportRefit measured the still-mounted WebView under
  the chat overlay and sent terminal.updateViewport on rotation,
  keyboard, text-scale, reconnect, and iOS-resume triggers. Refits
  are now suppressed while native chat covers the active terminal;
  the triggers already mark the viewport stale, and the
  return-to-terminal resubscribe re-measures.

* fix(mobile): harden native-chat resize suppression
…tablyai#10050)

Grok transcripts carry no timestamps. Previous logic excluded these rows
from matching, leaving pending sends and launch prompts unmatched and
causing the seeded bubble to appear rank-pinned at the list tail — which
reads as conversation reordering.

Now pending sends, launch prompts, and their pruning rules treat null
timestamps as matching-eligible, allowing echo suppression and proper
cleanup of delivered messages.
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
The type-aware switch-exhaustiveness lint rule requires explicit cases
for 'current', 'duplicate', and undefined; they fall through to the
existing generic skipped message, so behavior is unchanged.

Co-authored-by: Yuris Auzins <zuz666@users.noreply.github.com>
nwparker and others added 28 commits July 25, 2026 03:11
… as freezes (stablyai#10530)

* feat(diagnostics): record OS suspend/resume so sleep gaps aren't read as freezes

Renderer timers stop across OS sleep, so a multi-hour `renderer_memory`
heartbeat gap looks identical to a wedged renderer. That ambiguity sent the
uber-crash investigation down a deadlock path that the telemetry later
disproved -- and a healthy machine's own trace shows a 427-minute mid-session
gap with reason=interval, so the gap alone proves nothing either way.

powerMonitor 'resume' was already wired for renderer wake recovery but left no
breadcrumb. Stamp suspend and report the measured span on resume so the next
freeze report can be told apart from a laptop lid.

* fix(diagnostics): only record sleeps long enough to hide a heartbeat

Adversarial review flagged that the first cut would flood the 30-entry
breadcrumb ring and evict the crash evidence it exists to explain. Measured on
7 days of pmset history: 70 user-visible sleep cycles, worst 60-min burst of 7.
Median span is 2 SECONDS -- only ~24% run past 60s, so most of that traffic
could never explain a gap anyway.

Cycles are counted Sleep -> next FULL Wake, since powerMonitor's resume maps to
NSWorkspaceDidWake, which does not fire for dark wake.

Drop the suspend breadcrumb (suspend now only stamps a timestamp) and emit a
single `system_slept` on resume, gated at 60s. That cuts 70 cycles to 17 over
the same week (worst burst 3) while still catching every sleep long enough to
swallow a 60s renderer heartbeat.

* test(diagnostics): assert resume listeners detach by identity

The off mock deleted by event name alone, so teardown detaching a
different closure than the one registered still passed -- a leak of the
real powerMonitor listener would have gone unnoticed. For 'suspend' that
leak has no other observable effect through the public API.

Co-authored-by: Orca <help@stably.ai>

* fix(diagnostics): span from the first suspend across dark wake

powerMonitor 'resume' maps to NSWorkspaceDidWake, which does not fire for
dark wake, so macOS can deliver suspend -> suspend -> resume. Overwriting
the stamp reported only the trailing segment, and when that segment fell
under the 60s gate a 90-minute sleep recorded nothing at all -- leaving
the gap looking like the unexplained freeze this is meant to rule out.

Co-authored-by: Orca <help@stably.ai>

* docs(diagnostics): correct the threshold rationale to match measurement

The comment claimed maintenance sleeps would flood the ring. Re-measuring
pmset over 6 days (78 sleeps, 29 full wakes, 51 dark wakes) shows they
resolve as DarkWake, which never fires 'resume', so they never recorded a
breadcrumb at all. Real rate is 29 breadcrumbs / 6 days, worst 60-minute
burst 4 against a 30-entry ring. The gate's actual job is narrower: skip
sleeps shorter than the 60s heartbeat, which cannot open a gap to explain.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
…s hidden (stablyai#10528)

startGitCommonNarrowWatch was the only watch entry point that never received
WorktreePollerWindowVisibility, so its `worktrees/` existence poll kept
stat'ing every repo without a linked worktree forever in the background
(0.5 stat/sec/repo at the 2s default). Its siblings — the primary-metadata
snapshot poller and the non-darwin git-common polling — already park.

Threads visibility through and matches the snapshot poller's park/re-arm
pattern: the poll stops on the first hidden tick, and re-checks immediately on
onWindowBecameVisible so a worktrees dir created while hidden still upgrades to
the native stream and emits its create event. The visibility listener is
dropped in the dispose path.

darwin-only: the narrow watch is the `platform === 'darwin'` branch.
…tablyai#10526)

* fix(terminal): make primary-selection paste suppression single-shot

Middle-clicking in the terminal armed a 750ms window that swallowed every
native paste event, not just Chromium's one follow-up — so a real Ctrl+V
inside that window was silently dropped on Linux.

Consume the deadline on first use (`shouldSuppress*` -> `consume*`) so the
arm owes exactly one event; 750ms stays as that event's expiry bound.

* test(terminal): cover the paste event xterm actually forwards to the PTY

Review follow-ups on the single-shot suppression change; no source change.

The new real-module file only synthesized `beforeinput`. An Electron
probe (Chromium 150) showed `paste` fires first, is cancelable, and
cancelling it at document capture suppresses `beforeinput` entirely — and
xterm registers `handlePasteEvent` for `paste` only, never `beforeinput`.
So `paste` is the sole event that can double-write the PTY, and it was
the one event the end-to-end file did not exercise. Add it; it fails with
the fix reverted.

Consuming mutates, so `isTerminalNativePasteTarget(...) && consume()` is
now load-bearing: swapped operands would burn the arm on an unrelated
paste and let the real follow-up double-paste. The guarding test asserted
only `defaultPrevented`, which survives the swap — assert `consume` is
never reached instead.

Rename the mocked-file case that claimed to prove single-shot. With the
module mocked its sequence is dictated by the mock; it pins that the hook
re-asks per event rather than caching, so name it that.

* test(terminal): name the beforeinput dispatcher for the event it dispatches

Round-2 review nit on my own round-1 change: once `dispatchClipboardPaste`
existed alongside it, a helper named `dispatchPaste` that dispatches
`beforeinput` inverted the reader's expectation. Match the sibling file's
`dispatchNativePasteBeforeInput` convention.
… locale (stablyai#10536)

Settings search indexed the macOS privacy-toggle name as English-only
aliases on the LAN keyword, so a Chinese user typing the term macOS
System Settings actually shows them (本地网络) got no hit — zh's catalog
value had been changed to 局域网 (LAN).

Split the two wordings onto their own catalog keys so each locale carries
both: 87620e6416 = LAN, fa3239cd42 = Local Network (its true content
hash). Localized values come from the repo's own LAN title translations
and from macOS 26's SecurityPrivacyExtension Localizable.loctable
(LOCAL_NETWORK), so nothing is invented. ja/ko/es already matched macOS
and are unchanged apart from gaining the LAN key.
…ai#10543)

* feat(diagnostics): record OS suspend/resume so sleep gaps aren't read as freezes

Renderer timers stop across OS sleep, so a multi-hour `renderer_memory`
heartbeat gap looks identical to a wedged renderer. That ambiguity sent the
uber-crash investigation down a deadlock path that the telemetry later
disproved -- and a healthy machine's own trace shows a 427-minute mid-session
gap with reason=interval, so the gap alone proves nothing either way.

powerMonitor 'resume' was already wired for renderer wake recovery but left no
breadcrumb. Stamp suspend and report the measured span on resume so the next
freeze report can be told apart from a laptop lid.

* fix(diagnostics): only record sleeps long enough to hide a heartbeat

Adversarial review flagged that the first cut would flood the 30-entry
breadcrumb ring and evict the crash evidence it exists to explain. Measured on
7 days of pmset history: 70 user-visible sleep cycles, worst 60-min burst of 7.
Median span is 2 SECONDS -- only ~24% run past 60s, so most of that traffic
could never explain a gap anyway.

Cycles are counted Sleep -> next FULL Wake, since powerMonitor's resume maps to
NSWorkspaceDidWake, which does not fire for dark wake.

Drop the suspend breadcrumb (suspend now only stamps a timestamp) and emit a
single `system_slept` on resume, gated at 60s. That cuts 70 cycles to 17 over
the same week (worst burst 3) while still catching every sleep long enough to
swallow a 60s renderer heartbeat.

* test(diagnostics): assert resume listeners detach by identity

The off mock deleted by event name alone, so teardown detaching a
different closure than the one registered still passed -- a leak of the
real powerMonitor listener would have gone unnoticed. For 'suspend' that
leak has no other observable effect through the public API.

Co-authored-by: Orca <help@stably.ai>

* fix(diagnostics): span from the first suspend across dark wake

powerMonitor 'resume' maps to NSWorkspaceDidWake, which does not fire for
dark wake, so macOS can deliver suspend -> suspend -> resume. Overwriting
the stamp reported only the trailing segment, and when that segment fell
under the 60s gate a 90-minute sleep recorded nothing at all -- leaving
the gap looking like the unexplained freeze this is meant to rule out.

Co-authored-by: Orca <help@stably.ai>

* docs(diagnostics): correct the threshold rationale to match measurement

The comment claimed maintenance sleeps would flood the ring. Re-measuring
pmset over 6 days (78 sleeps, 29 full wakes, 51 dark wakes) shows they
resolve as DarkWake, which never fires 'resume', so they never recorded a
breadcrumb at all. Real rate is 29 breadcrumbs / 6 days, worst 60-minute
burst 4 against a 30-entry ring. The gate's actual job is narrower: skip
sleeps shorter than the 60s heartbeat, which cannot open a gap to explain.

Co-authored-by: Orca <help@stably.ai>

* fix(window): stop viewport reflow from moving screen geometry

Blink's ScreenMetricsEmulator::Apply checks screen_size and view_position
before the desktop/mobile branch, so screenPosition:'desktop' does not make
them inert -- despite Electron documenting both as mobile-only. Passing the
content size and 0,0 overrode screen.width/availWidth and the window origin
for the whole 32ms hold, so a browser context menu opened mid-reflow would
translate against 0,0 and land in the wrong place (BrowserPane.tsx:3167).

Empty screenSize means 'no override', and an omitted viewPosition stays
nullopt in Electron's converter, so the real position survives. Only the
scale factor moves now, which is what the reflow actually needs.

Also record a breadcrumb when the restore exhausts its attempt budget: the
renderer is left at the wrong scale factor until some later reveal fixes it,
and that was previously silent.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
…i#10525)

* fix(release-cut): gate an explicit RC against its own series

semver_gt compares through strip_pre(), so the explicit-version override
only ever checked the stable line: 1.4.156-rc.0 read as 1.4.156, cleared
a 1.4.155 stable, and republished an RC below what clients already run.
Anchor a prerelease request on highest_rc_for_base -- the same rc history
the kind path uses -- so the override can only advance the series.

Two sibling gaps in the same block:
- version_suffix was silently dropped when version was set, because the
  append lives in the kind branch the override skips.
- the shape regex rejected X.Y.Z-rc.N.suffix, so a suffixed RC the rc
  path can produce could never be re-cut explicitly.

* fix(release-cut): close both ends of the rc-number range the gate compares

The new explicit-rc gate compares with `[[ -le ]]`, i.e. bash machine-width
integers, and the author closed only the low end. Past INTMAX bash saturates,
so `version=1.4.156-rc.99999999999999999999` reads as "above the published
rc.3" and the gate falls open — then the tag it cuts pins
highest_rc_for_base at 1e20 for that base forever, and every later cut wraps
to a lower rc the fleet never updates to. Bound the rc number to nine digits.

Also reject leading zeros on an all-digit prerelease identifier. `npm version`
renormalizes rc.4.01 to rc.4.1 while the tag step keeps the literal input, so
the shipped package.json version and its own release tag name different
releases. The explicit path's embedded identifier now goes through the same
validator the kind path uses instead of only the shape regex.

* fix(release-cut): stop the refusal pointing minor/major RCs at the wrong series

kind=rc derives its base from bump(latest_stable, patch), so the remedy the
refusal suggested only works when the requested base *is* that next patch. A
1.5.0-rc.N series exists only because this override created it, so an operator
resuming a stuck 1.5.0-rc.2 was told to dispatch kind=rc, which would have cut
an unrelated 1.4.156-rc.4. Spell the condition out and give the fallback that
does work for a non-patch base.

Also correct the mechanism in the comment I added in 698c5be: bash wraps
two's-complement, it does not saturate, which is why the hole is
value-dependent (rc.10000000000000000000 wraps negative and failed closed,
rc.99999999999999999999 wraps to 7766279631452241919 and sailed through).
And name both inputs in the suffix error, which now serves version_suffix and
the trailing identifier in version.

* fix(release-cut): count a suffixed RC from its commit subject, not just its tag

The new explicit-version gate only fails closed on a deleted tag because
highest_rc_for_base also reads `release: v<base>-rc.N` subjects. That fallback
did not parse the suffixed form: rcNumberFromTag accepts an optional
.identifier, rcNumberFromReleaseSubject did not, so `4.perf` failed its
`(\d+)(\s|$)` anchor and returned null.

So deleting a v1.4.156-rc.4.perf tag dropped the series back to rc.3, and an
explicit 1.4.156-rc.4 was waved through — below the rc.4.perf build
perf-channel clients already run. Same under-count already made kind=rc
recompute rc.4 over a deleted suffixed tag.

Mirror the tag form's optional identifier. Covered by a unit assertion and a
git-fixture test that both fail with this reverted.

* docs(release-cut): correct four operator-facing claims in the explicit path

All four are wording or consistency, no behavior change (harness: 26/26 before
and after, on bash 3.2 and bash 5.2).

- The trailing-identifier comment justified itself as preserving a shape that
  "can never be re-cut through the override", but re-cutting a suffixed rc at
  or below the series head is exactly what the new gate refuses. State what it
  actually admits: a second spelling of version=X.Y.Z-rc.N + version_suffix.
- version_suffix's input description still said "rc kind only" after this PR
  made it apply to an explicit bare X.Y.Z-rc.N.
- The suffix guard's own rc pattern was unbounded while the shape check twelve
  lines up is bounded to nine digits; reuse the bounded one so a later edit to
  either cannot silently drift.
- "which recovers the existing tag" was unconditional, but kind=rc recovery is
  also gated on tag_matches_current_ref, so a tag cut from a ref main has moved
  past advances to rc.N+1 instead.
* fix(git): read core.sparseCheckout the way git does

Sparse-checkout detection parsed git config line-by-line and only accepted a
section header alone on its line, so git's legal same-line form
`[core] sparseCheckout = true` matched neither branch and was silently skipped:
a genuinely sparse worktree lost its badge and partial-checkout warning. It also
read `config.worktree` unconditionally, although git honors that file only while
extensions.worktreeConfig is on, so a stale worktree config could override the
repo's real setting.

Headers are now consumed left-to-right off each line (further headers and one
assignment may follow), and config.worktree is read only behind the extension
gate. Every new expectation was confirmed against real `git config --get`.

* test(git): correct what git actually does with a trailing-junk config value

Git does not reject `[core] sparseCheckout = true bogus = false` outright: it
parses the line and takes the whole tail as one value (`git config --list`
reports `core.sparsecheckout=true bogus = false`), then fails only the boolean
coercion. The expectation is unchanged; the comment now matches the binary.
…tarve terminal.wait (stablyai#10529)

* fix(runtime): reserve long-poll headroom so orchestration.ask can't starve waits

orchestration.ask joined the long-poll set, which also opted it into the
single server-wide activeLongPolls counter. Because ask blocks on a reply
for its full timeout (600 s default, previously unbounded via a caller
timeoutMs), 16 asking workers could hold every slot and shed terminal.wait
and check --wait with runtime_busy for every other client — mobile, web,
CLI, SSH and relay all share this runtime.

Meter ask as its own long-poll class with a sub-cap of half the budget, and
clamp the caller-supplied timeoutMs at 30 min. The keepalive and abort-signal
wiring that motivated the original change is unchanged.

* test(runtime): cover the ask sub-cap and counter release on the WebSocket path

The admission fence is shared by both transports but only the Unix-socket
path was exercised, so a WS-only regression in admitLongPoll/releaseLongPoll
would have shipped silently. Drives handleWebSocketMessage with a 'runtime'
scoped device (orchestration.ask is absent from the mobile allowlist) and
asserts the overflow ask is shed without burning a reserved slot, that
check --wait still gets the other half, and that both counters return to
zero when the socket closes.
…cker (stablyai#10527)

* fix(tasks): keep repos with a pending remote-identity probe in the picker

Task-repo eligibility filtered on `hasProjectRemoteIdentity`, which is
populated by a background `git remote -v` probe. When the probe could not
reach git — an SSH-hosted repo whose connection is not up yet, a cold
launch — the repo silently vanished from the Tasks picker and stayed
hidden for the full 5-minute negative-cache TTL, even after the host came
back. GitHub repos were largely shielded because a persisted `upstream`
satisfies the identity projection through a different route; GitLab and
other providers depend on the probe.

Distinguish unknown from settled instead of hiding both:

- `probeGitRemoteIdentity` reports `resolved` / `no-remote` (git answered,
  no usable remote) / `unavailable` (never reached git).
- Enrichment persists `gitRemoteIdentity: null` only on `no-remote`,
  mirroring the existing `upstream: null` "not a fork" marker. An
  unreachable host leaves the identity undefined.
- Persistence keeps the explicit `null` instead of dropping it.
- `getTaskEligibleRepos` keeps a repo whose identity is still pending;
  folders and settled remote-less repos stay filtered out.

* test(tasks): cover the SSH probe exec paths for remote-identity status

Addresses CodeRabbit review: the unavailable-on-error case only exercised
the local git runner. Adds a connected-provider whose exec rejects, and an
SSH repo git answered for with no remotes.

* test(tasks): pin that a settled no-remote repo still resolves once it gains a remote

Three independent reviewers flagged that the candidate filter's `!repo.gitRemoteIdentity`
looks like an oversight next to the new null marker. Tightening it to `=== undefined`
would silently stop detecting a remote added after the marker landed. Document that the
re-probe is deliberate and pin the behavior with a test.
…stablyai#10547)

The old anchored regex matched neither branch on a `[section "sub"]key = value`
line, so the parser never left `[core]` and credited the next indented line to
it — reporting sparse for a worktree git says is not. Fails on the pre-fix
parser (returns true where git reports unset).
…t freeze workspace creation (stablyai#10540)

* fix(worktree): bound the .worktreeinclude copy so a huge include can't freeze creation

`.worktreeinclude` copying was bounded in entry count (1000) but unbounded in
bytes and files, and awaited inline during worktree creation. A repo listing
`node_modules` froze creation for minutes behind the create dialog on Linux and
Windows, where the fallback is a full `fs.cp` (macOS gets a cheap APFS clone).

Measure each copy-mode source against a cumulative budget (2 GB / 50k files)
before the first byte is written, and refuse the entries that bust it. Refused
entries ride the existing `CreateWorktreeResult.warning` channel so a workspace
never silently comes up missing its included files.

Pre-measurement rather than mid-copy abort: `fs.cp` ignores its `signal`
option, so a started copy cannot be cancelled and would strand a partial tree.
Refusing up front means there is no partial state to clean up.

* fix(worktree): don't charge bytes for copy-on-write clones, and bound the sizing walk

Two defects in the copy budget, both found by review:

- The byte limit was applied on macOS, where the copy is an APFS clonefile.
  Measured: a 2.7 GB tree clones in 22 ms and consumes no disk. Refusing it on
  a 2 GB byte ceiling denied work that was already free — a regression on the
  one platform this bound was never meant to touch. Bytes are now charged only
  when a byte-for-byte copy will actually run; the volume probe that decides
  this is the same cached df+diskutil pair the clone runs, and writes nothing,
  so the "refuse before the first byte" invariant holds. The entry limit still
  applies everywhere: inodes are real work even on the clone path.

- A refused entry consumed no budget, so a `.worktreeinclude` listing many
  over-budget directories paid a fresh full-limit walk for each one — up to
  1000 x 50,000 lstat calls, re-creating the stall this bounds. The walk is now
  charged against its own ceiling whatever the verdict.

Also documents that `admit()` must be awaited sequentially (CodeRabbit).

* fix(worktree): give the sizing walk headroom so one huge entry can't starve the rest

The walk ceiling added in the previous commit was seeded with maxEntries, the
same number the entry limit uses. Sizing an entry that busts the file-count
limit walks maxEntries + 1, driving the ceiling negative, so every later
`.worktreeinclude` entry was refused without being measured at all.

That regressed the common case: a repo listing `node_modules` plus `.env` used
to get `.env`; it silently got nothing. Reproduced, and now covered by a test
that fails when the headroom is removed.

The walk now gets 5x the entry budget, so total sizing work stays bounded
(<=250k lstat per materialization, vs the 1000 x 50k this ceiling exists to
prevent) while ordinary lists never reach it. Entries refused because earlier
ones exhausted the walk report a distinct 'sizing' reason, so the warning stops
quoting size limits at a 4-byte file that was never measured.

* fix(worktree): bill a failed clone's bytes, and blame the right ceiling

Two follow-on defects from the copy-on-write fix:

- A predicted APFS clone that then failed mid-copy (EPERM, ENOSPC) fell through
  to a real `fs.cp` whose bytes were never charged, because the entry had been
  admitted on the premise that cloning is free. That reopened the unbounded
  copy on macOS. The measured size is already known, so the fallback now bills
  it and refuses if it no longer fits, reporting the entry as skipped instead
  of silently copying gigabytes. A clone that was never viable
  (ApfsCloneUnavailableError) was already charged as a real copy, so that path
  keeps falling back as before.

- The walk ceiling is also applied inside the measurement via
  min(remainingEntries, remainingWalk), and when the walk term bound, the
  refusal was still reported as 'entries' — telling the user a 3-file directory
  busted a 4-file limit. It now attributes to whichever ceiling actually bound.

Also fixes the singular warning text, which said "entry X was not copied ...
copying them would exceed ... Copy them in manually".

* fix(worktree): flag a partial clone leftover, cap the warning, cover two branches

- A clone that fails partway only removes an *empty* reservation, so leftovers
  can survive at the target. Reporting that entry as simply "not copied" sent
  the user to copy it in manually, straight into a half-populated directory.
  Those skips now carry mayBePartial and the warning says to check the path
  first. Cleaning up the leftovers stays the deferred follow-up it already was.

- The warning enumerated every skipped path. `.worktreeinclude` allows 1000
  entries and all of them can be skipped, so it now names five and counts the
  rest — an unbounded string is a poor look in a PR about bounds.

- Two load-bearing branches had no test, both proven by surviving mutants: the
  `bytesAreCopied` short-circuit (reachable when a wedged df/diskutil makes the
  volume probe answer "no clone", so bytes are charged up front and must not be
  billed twice), and chargeBytes actually consuming budget for later entries.

* fix(worktree): only flag directory clones as partial, and cap that list too

- mayBePartial was set for every refused clone fallback, but only a *directory*
  clone can leave anything behind: the file path clones into a temp name and
  publishes with link(2), so a failure leaves nothing at the target. Sending
  the user to inspect a path that does not exist is its own small lie.

- The partial-copy sentence sliced to five names without the "and N more" that
  the other sentence appends, so entries past the fifth were surfaced nowhere.
  Both sentences now share one nameList helper.
…t misreported (stablyai#10550)

The 30-min clamp was silent: the ask result carried no timeout figure, so
the CLI printed the value the caller *sent*. A worker passing
--timeout-ms 3600000 was told "ask timeout after 3600000ms" after only
30 min of real waiting — off by 2x, and accurate before this PR added the
clamp. Echo the effective budget on every ask return and print that.

Additive optional field; older clients fall back to the requested value.
Zellij and other multiplexers copy via OSC 52. Empty Pc is a valid XTerm
default for clipboard, but we rejected it, and the feature defaulted off so
copy silently failed inside Zellij. Accept empty Pc as clipboard, default the
setting on (query still blocked; size capped), and surface Zellij in settings.

Closes stablyai#10567
Review fixes for stablyai#10588:

- Persistence: profiles saved under the old off default persisted `false`,
  which is indistinguishable from a real opt-out, so the default flip never
  reached stablyai#10567's reporter. Added the repo's one-shot stamp
  (terminalAllowOsc52ClipboardDefaultedOnForAllUsers) so unmigrated profiles
  flip once and a later opt-out sticks.
- Replay: reattach/cold-restore re-writes recorded PTY bytes through the same
  parser, so a stale `\e]52;c;...` silently clobbered the clipboard on every
  restart. Gated behind isPaneReplaying via a new resolveOsc52ClipboardGate.
- Blocked toast latches once per renderer session and could be burned by a
  pre-hydration read; it now fires only for a real opt-out.
- An empty Pd decoded to '' and, with the gate default-on, silently blanked
  the clipboard. Now rejected as invalid.
- Localization: en.json is bundled and the catalog beats the code fallback,
  so all three copy changes were inert. Resynced across five locales.
- Corrected the empty-Pc rationale: tmux (not Zellij) emits `\e]52;;<b64>`.
Extracts createOsc52OscHandler so the replay/hydration gate wiring is
covered, not just the pure gate — dropping the isReplaying getter now
fails a test instead of passing silently.

Adds catalog assertions for the two OSC 52 settings strings. Only the
toast key was pinned, so the same inert-copy regression (code fallback
edited, bundled en.json not) could still ship for the settings pane.
…n flip

The default-on flip only reached the Electron store. The web/remote client
keeps its own settings in localStorage, so a profile that persisted the old
`false` there stayed opted out — the same bug the Electron migration fixed,
in the second store.

Extract the migration into shared/osc52-clipboard-settings.ts and call it
from both stores. Also coalesce OSC 52 writes onto a microtask so a hostile
chunk of ~15-byte sequences cannot fan out into a million clipboard writes,
and latch the blocked-write toast after it renders rather than before.
The default-on migration cannot distinguish a deliberate opt-out from a
profile that simply never touched the setting — both persisted `false` under
the old default. Flipping everyone is the only way to fix stablyai#10567 for existing
installs, but doing it silently reverses a security choice the user made.

Arm a one-shot notice at load when the migration overrides a persisted
`false`, on both settings stores, and show it once the renderer hydrates.
Profiles that never opted out are never notified.
… host

The host store always projects osc52ClipboardDefaultOnNoticePending, so the
plain spread in the web client's runtime UI merge overwrote an arm raised by
its own localStorage settings migration — flipping the opt-out in silence.

Co-authored-by: Orca <help@stably.ai>
Round-3 review fixes:
- Rename the arming predicate to osc52ClipboardDefaultOnOverridesPersistedOff.
  Both stores rewrite the whole settings object on every save, so every profile
  saved under the old off default holds `false` — the deliberate-opt-out cohort
  is not distinguishable on disk. Name, docs and test names now say so.
- Read settings before the UI snapshot in readLocalWebUIState: getStoredSettings()
  arms the notice, so reading first snapshotted a pre-arm state that callers wrote
  back, erasing an arm the stamp can never raise again.
- Give the notice toast a stable id; StrictMode re-runs the effect against the
  same closure, so the early return cannot catch the second pass.
- Restore guardParserHandler parity in the coalescer microtask.
- Drop the unverified Zellij claim justifying all-selections routing; that routing
  predates this branch and PRIMARY routing stays an open question.
- Cover the notice hook (order, single-fire, deep-link), the armed flag reaching
  disk and surviving a clear, and pin the notice catalog to its code fallbacks.

Co-authored-by: Orca <help@stably.ai>
The migration notice says to turn it off in Terminal settings, so searching
Zellij/Grok/tmux has to find it. Also note why the OSC 52 write-back clauses
stay despite an unrelated always-true clause in the same condition.

Co-authored-by: Orca <help@stably.ai>
…ds it relies on

The notice was cleared the moment the toast was enqueued, so a quit inside its
15s window spent the profile's only warning on a launch where nothing was ever
seen — and the settings stamp means it can never re-arm. Clear on
onAutoClose/onDismiss instead, plus explicitly in the action handler, because
sonner's action path deletes the toast without firing onDismiss.

Also closes three coverage gaps a review found:
- ui.set must accept osc52ClipboardDefaultOnNoticePending. The update schema is
  strict, so dropping the key rejects the whole call rather than stripping it,
  and the renderer only logs that failure — every paired client would re-toast
  forever with nothing red.
- the coalescer's try/catch and .catch had no test; the rejection case needs a
  plain function because vi.fn tracks settled results and hides the leak.
- pin that every selection kind (including bare `p`) lands in the system
  clipboard, so routing PRIMARY separately later is a deliberate break.

Co-authored-by: Orca <help@stably.ai>
…gration

readLocalWebUIState reads settings before the UI blob so the migration's arm is
in place before the snapshot every caller writes back. Seeding localStorage
after install is what makes ui.get the first settings read, and therefore what
makes swapping those two lines fail.

Co-authored-by: Orca <help@stably.ai>
The clear sets local state before persisting so a rejected ui.set cannot leave
the toast re-firing for the rest of the session; losing the persist only re-arms
the notice next launch.

Co-authored-by: Orca <help@stably.ai>
Three comment corrections from review:
- the safety note claimed exfil was the risk; queries are blocked, so it isn't.
  The actual accepted risk is execute-on-paste: decoded text goes to the
  clipboard verbatim, newlines included. Filtering here would break multi-line
  TUI copies, which is the feature; bracketed paste is where that is handled,
  and kitty/Ghostty take the same posture.
- the coalescer bounds a flood per parse yield, not overall.
- the replay gate reads at parse time while queued live bytes are drained
  before the guard engages, so a copy racing a reattach is dropped silently.

Co-authored-by: Orca <help@stably.ai>
Mutation testing found four assertions that stayed green against the very
change they were written to pin.

The notice suite passed 8/9 against a complete revert to clear-at-enqueue:
`calls[0][1][callback]?.()` is a silent no-op when the option is absent, and
the call count was already satisfied by the enqueue-clear, so nothing
separated "cleared by this callback" from "cleared earlier". Assert the
option exists and the notice is unspent before invoking it.

The stable toast id was deletable with all 9 green despite the adjacent
comment calling it load-bearing for StrictMode. Pin it.

The blocked toast's latch-after-throw fix was unproven: both orderings pass
when `toast.info` succeeds. Only a throwing first call tells them apart.

Deleting the hook call in App.tsx silenced the desktop notice with every
suite green. Pin it alongside the static Toaster import, since sonner drops
a toast enqueued before any Toaster subscribes and never replays it.

Also retone the coalescer-latch comment, which claimed the reset ordering
was load-bearing on its own; the try/catch reaches the same end, so the
test binds the pair.

All four verified green->red by mutation, then restored.
Add tests pinning the static Toaster mount required to prevent notice dropout (stablyai#10567), the stable toast ID deduping StrictMode double-invokes, that the notice stays unspent on toast throws, and that flush-latch guards prevent silent consumption across error boundaries.
@innocarpe

Copy link
Copy Markdown
Owner Author

Upstream stablyai#10588 merged on stablyai/orca — closing portfolio mirror.

@innocarpe innocarpe closed this Jul 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.