Skip to content

[upstream #10079] feat(workspace): Upstream/Origin issue source switch on branch create - #19

Open
innocarpe wants to merge 463 commits into
mainfrom
fix/composer-issue-source-switch
Open

innocarpe wants to merge 463 commits into
mainfrom
fix/composer-issue-source-switch

Conversation

@innocarpe

Copy link
Copy Markdown
Owner

Portfolio mirror of my contribution to upstream stablyai/orca.
Exhibition only — the real review/merge target is upstream.

Upstream

Summary

Summary Fork-based contributors often need to pick issues from upstream when creating a branch, not only from their fork. Tasks and Create Issue already expose IssueSourceSelector; the New Workspace smart name field (branch-from-issue create) did not. This PR: - Shows

Note

  • Do not merge this into innocarpe/orca main until the upstream PR is merged.
  • After upstream merges: sync fork from upstream, then close this mirror PR.
  • This open PR exists so visitors see in-flight work on this fork's Pull requests tab.

brennanb2025 and others added 30 commits July 24, 2026 13:25
…ry tab (stablyai#10437)

* fix(runtime): stop the headless hydration repo gate from dropping every tab

stablyai#9343 gated headless mobile-session hydration on the live repo list, but read
it as `this.store?.getRepos?.() ?? []`. A store that cannot report repos then
yields an empty set, which reads as "every repo is gone" and skips every
parseable session key — so no tabs hydrate at all. It also called getRepos on
every hydrate, including the hot floating-tab poll path that is contractually
free of repo/provider inventory work.

Resolve the inventory lazily and only for keys that parse to a repoId, and keep
`null` (unavailable) distinct from an empty list (all repos really gone). An
absent list now fails open; a known list still prunes as stablyai#9343 intended.

Fixes 3 tests that have been failing on main since stablyai#9343 landed:
orca-runtime-terminal-retirement (2) and orca-runtime (1).

* refactor(editor): extract the pending-focus effect to clear the max-lines cap

stablyai#8083 pushed RichMarkdownEditor.tsx to 404 counted lines against the 400-line
.tsx cap, failing lint for every PR that merges current main. The Explorer
find-focus request is self-contained, so it moves to its own hook.
…ablyai#9127)

* fix(codex): preserve runtime config without system source

* fix(codex): retain baseline when mirror is skipped

* refactor(codex): extract deprecated hook-flag normalization

Why: codex-config-mirror.ts sat at the 300-line cap, so the missing-source
guard could not land without a max-lines disable.

* fix(codex): bootstrap a baseline when the mirror is skipped

Why: a runtime home seeded outside the mirror (WSL, per-account) never got a
baseline while the source was missing, so promotion stayed inert and silently
reverted the in-Codex change once the source returned.

* fix(codex): stop a synthesized source config from wiping runtime settings

Two routes still reached the stablyai#9073 data loss after the missing-source guard:

- Promotion runs before the guard and, with no ~/.codex/config.toml, created
  one holding only the promoted keys. The next mirror treated that skeleton as
  authoritative and deleted every other runtime setting. It needs no missing
  file: `codex mcp add` inside an Orca-launched Codex plus /model was enough to
  drop the MCP server for good. Promotion now seeds a brand-new system config
  from the runtime's ordinary settings, so the mirror round-trips them.
- A 0-byte source (half-written, or an unhydrated cloud-synced home) still read
  as an authoritative empty config and advanced the baseline, making the loss
  unrecoverable. A blank source is now treated like a missing one.

Moves the TOML section model out of codex-config-mirror.ts so promotion can
share it without a cycle.

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
…tablyai#10443)

stablyai#9343 broke two contracts at once and stablyai#10437 fixed both, but neither is
asserted: stablyai#10429 repaired the failing tests by deleting the poll assertion and
by handing the retirement store a live repo, so a revert would land silently.

Restore `expect(getRepos).not.toHaveBeenCalled()` on the floating-tab poll, and
add a hydrate case for a store that cannot report repos. Verified both fail
against the pre-stablyai#10437 code: the poll assertion reports getRepos "called 2
times", and the hydrate case returns [] instead of the persisted tab.
…lyai#10340)

* fix(skills): source released history from the committed ledger, not a tag walk

verify:skill-bundle-manifest rebuilt the entire released-skill history by
walking every local refs/tags/v* on each run and demanded byte-equality with
the committed artifacts. Output was therefore a function of (skill bytes x
local tag set x release timing), so any clone holding stray, deleted, or fork
tags the committed artifacts predate rebuilt a divergent registry and failed
lint. This was the 4th instance of one failure class (stablyai#8637 -> stablyai#9119 version
bumps -> stablyai#9778 new tags -> local tag drift), each patched with a new tolerance
rather than removing the tag coupling.

Fix: the committed snapshot-registry + release-mapping ARE the released history;
trust them instead of re-deriving from tags.

- releasedHistoryFromCommitted() seeds generation from the committed ledger,
  dropping the floating unreleased tail (entries beyond what the mapping names).
  verify and --write are now pure functions of working-tree bytes with zero tag
  access. The tag walk survives only behind --rebuild-from-tags (disaster
  recovery), off the everyday path.
- --release <version> + appendReleaseRow() perform the O(1) append of one
  mapping row at release cut (dedupes vs the last row, strips the v-prefix) --
  the single authoritative point where working-tree bytes become an immutable
  released revision.
- release-cut.yml runs generate --release "$VERSION" before the release commit
  (Node built-ins only, no install needed); pr.yml drops fetch-depth: 0 from the
  lint job since verify no longer needs tag history.

Recognition is unaffected: the runtime uses knownSnapshots = registry.skills
(all entries, incl. the tail committed at PR-merge time), so a missing mapping
row only loses a version label, never recognition or the update nudge.

Trade-off: lint no longer cross-checks committed historical snapshots against
tags. A hand-edit to an old released entry is still caught by the runtime
manifest<->registry consistency check when the current manifest points at it,
and can be audited anytime with --rebuild-from-tags.

Verified: verify passes committed-sourced; --write is zero-diff (byte parity);
a planted stray v-tag no longer changes output; edit-stub -> --write -> --release
appends the correct single row; double --release is idempotent;
--rebuild-from-tags reproduces the committed artifacts. Generator tests 14 pass/
1 skip; runtime skill-bundle-artifacts + freshness-inventory 14 pass; bundled
skill guides verify passes.

* fix(skills): keep one release-mapping row per version on a re-cut

A cut that pushed the version bump to main but died before pushing the
tag is re-cut at the same version. If skills changed in between, the
second --release appended a duplicate row, and the stale one named
revisions that tag never ships — which verify-skill-update-roundtrip
then pairs with the tag's real bytes.

Overwrite the trailing row instead (the tag is absent, so that version
was never published). Refuse only when an earlier row claims the
version, which the cut workflow already rejects upstream, so this
cannot wedge a recovering cut.
…yai#10339)

* fix(win): taskkill plain-shell PTY trees on immediate teardown

Deleting a Windows worktree that still has a live process in a terminal
tab (a `pnpm i`, or any command that spawns a child tree) failed with
"Failed to physically stop every PTY for worktree" after the full 10s
teardown deadline.

Root cause: on Windows, closing a plain shell's ConPTY does not reap its
orphaned children — node-pty's `useConptyDll` skips the console-process
reap. A live `pnpm i`/`node` child survives the shell's exit, keeps the
ConPTY console non-empty (so the daemon still reports the session alive),
and holds the worktree cwd handle. The destructive-removal physical-stop
check then fails closed. stablyai#10100 fixed this for agent sessions
(killWithDescendantSweep -> taskkill /T /F) but plain shells were never
swept.

Fix: extend the Windows taskkill /T /F descendant tree-kill to non-agent
shells on the immediate/destructive teardown path, in both the daemon
(TerminalSessionTeardown) and local provider. Gated to win32 + immediate
with the same ownsRoot guard so a naturally-exited/recycled PID is never
signalled; POSIX shells keep reaching their child pgroup via forceKill
and are unchanged.

Reproduced and verified end to end via the Electron dev build: a worktree
with a live node child previously failed to delete after 10s; with the
fix the child tree is taskkilled, the directory is removed, and the
delete succeeds in ~1.8s.

* fix(win): claim plain-shell termination before the taskkill sweep

forceKillAndWaitForExit sets _isTerminating in its synchronous prologue.
Awaiting the Windows descendant sweep ahead of it left createOrAttach's
doomed-session guard open for the taskkill's duration, so a concurrent
attach could bind a pane to a session about to be tree-killed.
* feat(mobile): add safe Codex rate-limit resets

* fix(mobile): address reset credit review feedback

* review: purge removed-account reset attempts, shared capability constant, rebase test mocks

* review: preserve host compatibility and reset durability

* fix(mobile): recover reset capability after cutover

* fix(mobile): validate runtime capability payloads

* fix(mobile): enforce capability payload contract

* fix(mobile): route mock terminals to selected worktree

* test(mobile): pin malformed probe retry behavior

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
…s) for copying gitignored files into worktrees (stablyai#9791)

* feat(worktrees): copy project-level .worktreeinclude paths into new worktrees

Read .worktreeinclude at the repo root (gitignore syntax) and copy matching
gitignored paths from the primary checkout into each newly created local
worktree, so .env and other local config carry over with zero per-user setup.

- Literal patterns resolve by direct stat; globs match against
  ls-files --others --ignored --exclude-standard --directory (collapsed
  dirs keep huge repos fast); every candidate is re-verified with
  check-ignore so tracked or unignored files are never copied.
- Copy semantics, never symlink: APFS clone-copy on macOS, real copy
  elsewhere, so each worktree owns its files (unlike repo.symlinkPaths,
  which it merges with rather than replaces).
- Failures never block worktree creation.
- Remote (SSH) creation skips it, same as symlinkPaths.
- Split APFS clone helpers into worktree-apfs-clone.ts (max-lines).

Closes stablyai#7549

* fix(worktrees): harden worktree include copying

* fix(worktrees): support nested includes on Git 2.25

* fix(worktrees): bound include copy costs

* fix(worktrees): close include correctness and perf gaps

* fix(worktrees): preserve included copy semantics

* fix(worktrees): harden include resolution

* fix(worktrees): preserve bounded include resolution

* fix(worktrees): bound include filesystem resolution

* fix(worktrees): harden include matching

* fix(worktrees): tighten include matching and scan bounds

* fix(types): use concrete filesystem stat types

* fix(worktrees): harden included path materialization

* perf(worktrees): stop include parsing at resolver budgets

* chore(skills): refresh bundled skill manifests

* refactor(worktrees): reduce .worktreeinclude to focused literal-only scope

The reviewed implementation grew well past the ticket (stablyai#7549), which asks for a
size-M feature that reuses existing worktree machinery. Trim back to the minimal
change that solves the reported problem safely:

- Resolver now supports literal files and directories only. Glob/negation lines
  are skipped with a warning (documented follow-up), which removes the entire
  user-controlled-regex ReDoS surface, the CPU/byte budgets, the git enumeration
  scan, and the case-sensitivity engine. The filesystem + git check-ignore
  handle existence and case for free.
- Copy layer folded back into worktree-symlinks.ts (link/copy modes share one
  loop); dropped worktree-path-copy.ts, worktree-target-safety.ts, the
  descendant-dedup/realpath/target-parent machinery, and the per-materialization
  APFS filesystem cache. Kept the df/diskutil probe timeout.
- Reverted unrelated changes: check-ignored-paths timeout param and the
  git-binary-compatibility enumeration tests.

Net: -1903/+172 across the include+copy code. Behavior for the ticket's cases
(.env, .env.local, .vscode/, node_modules, config/secrets.json) is unchanged;
gitignored-only + copy-not-symlink semantics preserved.

Closes stablyai#7549

* fix(worktrees): dereference symlinked .worktreeinclude entries + cache APFS volume probe

Two issues found by review + perf audit of the copy path:

- Correctness (HIGH): a listed entry that is itself a gitignored symlink was
  copied AS a symlink (fs.cp dereference:false), and the darwin APFS branch was
  skipped for all symlink sources. Editing the worktree's copy then wrote through
  to the shared/primary target — inverting copy-mode's 'each worktree owns its
  files' guarantee, and escaping the worktree entirely if the link pointed
  outside it. Now resolve realpath for a top-level symlink in copy mode so we
  copy content; nested symlinks inside a copied dir stay as-is (cp -R semantics).

- Perf: assertSameApfsVolume ran df+diskutil per copied path (4 subprocesses
  each), so an N-entry include spawned ~4N short-lived processes on the macOS
  create hot path, all re-probing one volume. Add a per-materialization
  device-keyed cache: one probe per distinct volume (4N -> ~4).

Tests: symlinked-file and symlinked-dir dereference regressions (no leak to
primary); APFS volume probed once regardless of copied-path count.
…stablyai#10436)

* fix(skills): stop promising a skill update the command cannot deliver

An "Update available" badge could never clear: pressing Update ran
`npx skills update <name> --global`, which reported "All global skills are
up to date" and wrote nothing, while the badge stayed on.

Freshness marked a name updatable whenever ANY placement was outdated,
including a standalone duplicate in an agent home. The global command only
converges the canonical copy and its symlink aliases, so a stale duplicate
kept the badge lit with no command that could clear it.

- Count only reliably-convergent placements toward the update promise, so a
  stale duplicate no longer advertises an update that cannot land.
- Give the duplicate chip a real skipped-reason instead of falling through to
  the generic sentence, naming the copy and how to resolve it.
- Surface the existing freshness review dialog from the setup rails via a
  Details link, so a blocked or duplicate copy is explainable where the user
  actually sees the badge. Covers orchestration, Computer Use, Ephemeral VMs,
  Linear, the CLI section, the floating orchestration modal, the Browser Use
  card, and the Mobile Emulator row.

* fix(skills): say when a skill copy needs attention instead of reading as all-clear

A skill with an out-of-date copy the update command cannot reach rendered as a
green "Installed" pill. That is honest about the main copy but reads as
all-clear, so real drift in an agent home stayed invisible — the user had no
reason to suspect there was anything to click.

- Add a needs-attention display status (amber) for a placement that is not
  current but has no eligible update: stale duplicates, edited copies,
  read-only, inaccessible, and broken or external links. Presence-only stays
  green, and an unscanned inventory stays quiet so nothing flashes amber on
  launch.
- State the reason inline on the setup rails, so the cause is readable without
  opening the review dialog; Details still opens the full per-location list.
- Extract the skipped-reason sentence into its own module so the rails and the
  dialog share one source and can never drift apart.
- Carry the state through the settings sidebar badge so the nav and the card
  cannot disagree.

* fix(skills): give the inline skill warning something to point at

The shared reason sentences are deictic — "this copy", "the copy here" — because
they were written for the review dialog, where the location rows they describe sit
directly beneath them. On the setup rails there are no rows, so "this" referred to
nothing and the sentence read as if it were about the skill itself.

Name the offending paths above the sentence on the rails, so the referent is
present before the wording that depends on it. Every copy sharing the blocking
reason is listed, not just the first, so resolving one does not leave the badge
unexplained. The dialog keeps its existing wording and rows unchanged.

* fix(skills): stop withholding the update over copies it cannot reach

Eligibility is now decided purely over the placements the global command actually
converges — the canonical copy and its symlink aliases.

A project skill, plugin cache, standalone duplicate, or unreadable copy in another
agent's home used to withhold the update from the whole name. `skills update
--global` provably never writes any of them, so that refused work the command could
have done over a copy that was never at stake. A blocked *convergent* copy still
withholds it: that is the placement the command writes to, and overwriting it is the
real data-loss case.

Also drop the inline reason from the setup rails. The reasons are written to sit
beside the location rows they describe, so on a card they had nothing to point at and
could only ever name one cause. The rails now mark Details with a warning icon when a
copy needs the user's own hands, and the dialog does the explaining with every
location and cause it knows about.

* fix(skills): make the skill Details affordance read as a control

The review link rendered as bare text, so nothing but hover said it could be
clicked — the badge told the user something was wrong and then gave them no visible
way in.

Use the ghost variant so it carries a hover/focus background and a real hit target
instead of a zero-padding text run, and add a chevron so it reads as a control at
rest. The chevron points right, not down: this opens the review dialog, while a down
chevron already means the in-place expander inside that dialog.
…ablyai#8643)

* fix(rate-limits): keep Codex PTY reset text for weekly-only plans

The PTY /status fallback parses '5h limit' and 'Weekly limit' lines by
label, but the extracted reset text was only ever attached to the
session window. Codex plans without a 5h session bucket (e.g. current
Pro) produce a weekly-only parse, so the reset time the CLI printed was
silently dropped. Fall back to the weekly window when no session window
exists.

* review: parse Codex PTY reset text per window into resetsAt

* review: make Codex PTY status fallback work on codex >=0.145

* review: harden PTY status parse against model-scoped rows and styled output

* fix(rate-limits): strip private PTY control sequences

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
…ai#10460)

stablyai#10340 added a ledger-advance step to the release cut, which violates the
contract test stablyai#9119 added: the cut must not run generate-skill-bundle-manifest
or stage resources/skills. Both PRs were green on their own branches and only
conflicted once merged, so nothing failed until main had both — main and every
open PR have been red since.

Revert the two workflow lines. The script's --release implementation stays: it
is correct and harmless when unused, and the root fix in stablyai#10340 — verify no
longer walking git tags — does not depend on the cut step.

This leaves stablyai#10340 semantically incomplete and that must not be dropped. With
the registry seeded from the committed ledger instead of a tag walk, nothing
advances the ledger at cut, so each new skill change re-uses the same unreleased
tail revision for different bytes; older installs then match no known snapshot
and degrade to unrecognized, which reads in the UI as a skill that needs
attention and cannot be updated. Follow-up is to reintroduce the advance
narrowly — stage only resources/skills/release-mapping.json and narrow the
assertion to forbid mutating the content-addressed artifacts while permitting
the provenance row.
…i#10432)

Add closeBeforePress flag to Rename, Browser, and Refresh actions to
defer modal opening until the action sheet closes. Eliminates the race
condition that caused the mobile app to freeze when opening these modals.
…ablyai#10445)

* fix(ssh): re-arm remote file watches when the provider reconnects

Remote file changes stopped being detected over SSH until the file was
reopened. The watch pipeline itself was fine — nothing ever re-established
the subscription after the transport it was made on went away.

Two paths left an editor tab permanently stale:

- A reconnect kills the relay's watch registrations, and the previous
  provider's unwatch handle belongs to the dead transport.
- A connect slower than the 60s retry window made installRemoteWatcher
  give up for good; first deploy to a new host far exceeds that.

Neither recovered, because installRemoteWatcher is only reachable from the
fs:watchWorktree handler, the retry timer, and the removal-restore path,
and the renderer only issues a watch for newly added targets. Reopening
the file just re-read it — the watcher stayed dead.

Give the layer that owns the transport the job of re-arming:
registerSshFilesystemProvider now notifies subscribers, which covers both
establish and reconnect since registerProviders runs on both. The watcher
keeps the intent to watch in a registry that outlives any single
connection, and on registration drops the stale entry (installRemoteWatcher
treats an existing entry as installed and would otherwise hand back a
watcher that can never fire), reinstalls, and emits overflow so consumers
resync the gap. Intent is dropped on unwatch and sender destroy so a
closed tab is never resurrected.

Verified over SSH to a Rocky Linux 10 host: after the reconnect that
previously killed it, a remote append lands in the editor in ~2s with a
live remote watcher process, and a second edit in ~1.5s.

* fix(ssh): drop watch intent when a remote worktree is removed

A removed worktree kept its entry in the intent registry, so a reconnect
landing before the renderer's unwatch would re-watch a deleted path —
60s of retries against the host and then a bogus overflow.

Also covers two reinstall cases: several senders on one connection must
collapse onto a single relay watch (and all of them resync), and a
destroyed renderer must not be reinstalled.

* fix(ssh): resync when a reconnect's watch only lands on a retry

The reinstall emitted the overflow only for listeners whose first install
returned 'installed'. If that attempt failed — relay-watcher.js still
spawning on the fresh transport, a transient fs.watch rejection — the 1s
retry restored the watch but never signalled the gap, so everything that
changed while the transport was down stayed invisible: the STA-2525
symptom reappearing inside the fix for it.

Thread the resync intent through the retry record so the overflow lands
when the retry does. All pre-existing callers default to false, so the
watch/terminal-error/restore paths are unchanged.

* test(ssh): cover the resync merge when a fresh watch claims the retry slot

The `resyncOnInstall ||=` merge was uncovered: removing it left every
existing test green. It is load-bearing — if a second renderer joins the
reinstall's failing install and reaches the retry slot first with
resync=false, the whole chain stays false and the retry restores the
watch without ever signalling the gap.

New sibling file rather than an append: filesystem-watcher.test.ts is
~10 counted lines from the 800-line lint cap, and AGENTS.md forbids a
max-lines disable.
…id (stablyai#9743) (stablyai#10455)

* fix(editor): read out-of-worktree SSH paths without a stamped target id (stablyai#9743)

Route external absolute-path tabs by their resolved SSH connection instead of
the externalSshTargetId stamp, which only the terminal-link open path sets. The
stamp keeps its fail-closed role when present.

* fix(editor): keep client-local live-tail log tabs off the worktree SSH host

AI Vault "View Log" tabs are opened client-local by construction (readOnly +
liveTail, no runtime env), so inferring the external SSH owner from the
worktree connection made them read the remote host instead of the granted
client path. Restrict that inference to non-live-tail tabs; an explicit
externalSshTargetId stamp still routes remotely.
…i#10470)

* fix(ssh): re-arm remote file watches when the provider reconnects

Remote file changes stopped being detected over SSH until the file was
reopened. The watch pipeline itself was fine — nothing ever re-established
the subscription after the transport it was made on went away.

Two paths left an editor tab permanently stale:

- A reconnect kills the relay's watch registrations, and the previous
  provider's unwatch handle belongs to the dead transport.
- A connect slower than the 60s retry window made installRemoteWatcher
  give up for good; first deploy to a new host far exceeds that.

Neither recovered, because installRemoteWatcher is only reachable from the
fs:watchWorktree handler, the retry timer, and the removal-restore path,
and the renderer only issues a watch for newly added targets. Reopening
the file just re-read it — the watcher stayed dead.

Give the layer that owns the transport the job of re-arming:
registerSshFilesystemProvider now notifies subscribers, which covers both
establish and reconnect since registerProviders runs on both. The watcher
keeps the intent to watch in a registry that outlives any single
connection, and on registration drops the stale entry (installRemoteWatcher
treats an existing entry as installed and would otherwise hand back a
watcher that can never fire), reinstalls, and emits overflow so consumers
resync the gap. Intent is dropped on unwatch and sender destroy so a
closed tab is never resurrected.

Verified over SSH to a Rocky Linux 10 host: after the reconnect that
previously killed it, a remote append lands in the editor in ~2s with a
live remote watcher process, and a second edit in ~1.5s.

* fix(ssh): drop watch intent when a remote worktree is removed

A removed worktree kept its entry in the intent registry, so a reconnect
landing before the renderer's unwatch would re-watch a deleted path —
60s of retries against the host and then a bogus overflow.

Also covers two reinstall cases: several senders on one connection must
collapse onto a single relay watch (and all of them resync), and a
destroyed renderer must not be reinstalled.

* fix(ssh): resync when a reconnect's watch only lands on a retry

The reinstall emitted the overflow only for listeners whose first install
returned 'installed'. If that attempt failed — relay-watcher.js still
spawning on the fresh transport, a transient fs.watch rejection — the 1s
retry restored the watch but never signalled the gap, so everything that
changed while the transport was down stayed invisible: the STA-2525
symptom reappearing inside the fix for it.

Thread the resync intent through the retry record so the overflow lands
when the retry does. All pre-existing callers default to false, so the
watch/terminal-error/restore paths are unchanged.

* test(ssh): cover the resync merge when a fresh watch claims the retry slot

The `resyncOnInstall ||=` merge was uncovered: removing it left every
existing test green. It is load-bearing — if a second renderer joins the
reinstall's failing install and reaches the retry slot first with
resync=false, the whole chain stays false and the retry restores the
watch without ever signalling the gap.

New sibling file rather than an append: filesystem-watcher.test.ts is
~10 counted lines from the 800-line lint cap, and AGENTS.md forbids a
max-lines disable.

* fix(ssh): re-arm a remote watch that died without a reconnect

The 60s fast-retry window gave up with a single overflow and no further
trigger — provider registration is the only re-arm, and a watch killed by
remote OOM/inotify exhaustion leaves the SSH link perfectly healthy. Back
off from 1min to a 30min ceiling instead, and stand down entirely when the
provider is gone so registration owns that case.

* fix(ssh): re-arm the remote file-explorer watch on reconnect

SshFilesystemProvider.dispose() stops each watch registration without
invoking its terminal callbacks, so a dropped transport left the runtime
file-explorer watch silently dead: the lease's restart only runs from the
failed-removal path, and the renderer's runtime subscription stays open
because it rides a different link. Reinstall on the connection's next
provider registration and emit overflow so clients resync.
stablyai#10449)

* feat(codex): surface a stalled config sync instead of failing silently

Why: the mirror keeps serving the last synced settings when ~/.codex/config.toml
is missing, blank, or unreadable. That is the right call for data safety, but it
is invisible — a downed WSL distro or an unhydrated cloud-synced home leaves
"Orca ignores my config edits" with no log line and no UI to diagnose.

Status is derived on demand from the same predicates the mirror uses, so the two
cannot disagree. The stall is logged once per episode rather than on every launch
and quota poll, and the Codex account section names the file and what to do.

* fix(codex): latch an unreadable source and stop over-claiming recovery

An unreadable source throws out of the mirror, so reporting only on the success
path left that stall latch-less: it logged the raw failure on every launch and
quota poll while its reason never reached the surfaced status. Report from the
catch path too.

The clear message also claimed the source was "readable again", which is false
when the stall ended because the runtime config was removed rather than because
the source came back.

Restoring console.warn now happens in afterEach — an inline mockRestore is
skipped by a failing assertion, and the leaked spy made every later case in the
block fail spuriously.

* fix(codex): latch the stall promotion hits first, and scope it to the host

Review round 1 findings:

- The unreadable-source latch still never fired in the steady state. Once a
  baseline exists, promotion reads the source before the mirror does, so it
  throws first and `!promotionPlan` returned before any reporting — logging a
  reasonless failure every launch and quota poll, which is exactly what the
  previous commit claimed to fix. Report from that branch too. The test only
  passed because its fixture had no baseline; it now seeds one first and fails
  without the fix.
- The banner named the host's ~/.codex while a WSL or per-account runtime was
  selected, whose real source is a different file entirely. Gate it to the host
  scope, matching how the sign-in warning is already gated.
- Three new translate keys were missing from the locale catalogs, failing the
  localization gate in `pnpm lint`.
- The registrar mock was never asserted, so deleting the registration left the
  suite green.
- `codexConfigSyncStatus` hung off the `agentHooks` namespace despite having
  nothing to do with agent hooks; moved to its own `codexConfigSync.status`
  while it is still a four-file change.

* fix(codex): report sync health for the home the selection actually mirrors

Review round 2:

- The status resolved the shared runtime home, but the system default now runs
  Codex directly against ~/.codex and managed accounts get their own home. So a
  stalled per-account mirror showed no banner at all, while a stale shared home
  could warn about a config the active lane never reads. Resolve the mirrored
  home from the current selection, and report synced when the lane has no mirror
  to fall behind.
- The round-1 report on the promotion failure path could clear the latch on a
  pass where no mirror ran, claiming a recovery that never happened and
  silencing every later pass. Only ever latch a stall there; leave clearing to
  the path that actually mirrored.

* fix(codex): refetch sync status when the active Codex account changes

Review round 3:

- Resolving the status per selection made the fetch account-dependent, but the
  effect was not keyed on the active account. Switching accounts left the banner
  describing the previous one — and switching INTO a stalled account showed
  nothing at all, which is the silence this change exists to remove.
- Pin the home resolution itself: it had no direct test, and its shared-home
  path was a hand-copied literal that could drift from the real helper and
  silence the banner with every other test still green.
- Narrow the handler's dependency to the one method it calls, which also drops
  an `as unknown as` cast from its test.
- Skip the chmod-based test on Windows, where a read-only directory does not
  block writes so the scenario cannot be constructed; matches the convention
  already used in config-settings-promotion.test.ts.

* chore(codex): restore the handler docstring and isolate the resolver suite

Round 4 returned clean; these are its two non-blocking nits.

Narrowing the handler param left its JSDoc stranded above the new type, so the
function had no hover doc. The resolver suite also read the developer's real
CODEX_HOME and shell rc, so anyone exporting one would see it fail locally.
…stablyai#9648)

* fix(terminal): resume hibernated agents that reattach with no payload

When agent hibernation is on, a stopped (done) agent's PTY is killed and a
passive sleeping record is kept; returning to the worktree relies on the pane
reattaching on remount. On the daemon path (Windows/local worktrees) the daemon
can reattach the hibernation-killed session as already-live (isReattach,
isNew:false) and return no snapshot/replay/coldRestore — and, being a reattach
rather than a fresh spawn, it silently drops the --resume command passed on
connect. The renderer adopted that empty session, leaving a blank terminal with
nothing running and the sidebar history still pointing at the dead tab (the
sleeping record never cleared).

A reopened pane that owns a resumable slept session must always be re-driven
with its resume command, never left as a bare empty attach. handleReattachResult
now discards a contentless isReattach for a pane with a hibernation record and
re-drives the prepared resume. This excludes the healthy cases: a fresh session
the daemon created (isReattach falsy — it already ran the command) and a live
reattach (carries a snapshot/replay). Forward the isReattach signal the
transport was dropping so the two cases are distinguishable.

Adds a deterministic regression test that fails without the guard and passes
with it.

* fix(terminal): preserve provider ownership on resume

* chore(skills): refresh generated skill bundle manifests

Regenerate the skill bundle artifacts against the full release-tag set so
the freshness verify check passes. Append-only additions for the newer
release tags; no released snapshot history is rewritten.
…e the reveal reflow (stablyai#10473)

* fix(window): restore the macOS 26 reflow without touching the native frame

stablyai#10253 stopped the main-thread deadlock by skipping the repaint size nudge on
macOS 26, but invalidate() repaints without reflowing, so the h-dvh root kept a
stale viewport height and the status bar stayed clipped off-screen (STA-2383) on
every Tahoe reveal, restore and wake.

Drive the reflow through device emulation instead: a +1px emulated viewport,
reverted a frame later, makes the renderer recompute layout without any NSWindow
mutation, so the FrontBoardServices re-entrancy that wedged the main thread for
109 minutes is still never triggered.

Verified against real Electron 43.1.0 on macOS 26.3.1 (Darwin 25.3.0): the
renderer sees the resize and relayouts, the native frame is untouched, the
viewport and devicePixelRatio restore exactly across zoom levels, overlapping
calls collapse to one cycle, and a 1px delta never crosses a terminal cell
boundary so no pane reports new geometry (no SIGWINCH to running shells).

Also close two gaps in the surrounding code:
- the pre-Tahoe size jiggle now clears its WeakSet latch in a finally block, so a
  throwing setSize can no longer suppress every later repaint for that window
- cover powerMonitor 'resume' under the Tahoe guard, the other AppKit dispatch
  context implicated in the freeze

* fix(tray): keep NSStatusItem scene updates off the AppKit callout stack

The main-window repaint was only one of the two doors into the macOS 26
FrontBoardServices deadlock. Showing or restoring the window calls
setTrayAttention(false) straight from the window event handler, and
tray.setImage/setToolTip drive an NSStatusItem scene update — the same
re-entrant scene mutation from inside AppKit's own dispatch, matching the
stackshot in openai/codex#23695.

Defer the native mutation to a fresh event-loop turn so the callout frame is
vacated first. The attention flag itself still flips synchronously: rapid
show/hide would otherwise mis-dedupe against a value that had not landed yet.

Bursts collapse to a single repaint, and because the deferred pass reads current
module state rather than a captured value, a coalesced schedule can never apply
a stale icon. applyTrayImage already no-ops on a destroyed tray, so a repaint
still queued when the tray goes away is harmless.

* fix(window): reflow maximized and fullscreen windows on macOS 26 too

The maximized/fullscreen bail-out predates the Tahoe path and exists only to
keep the size nudge from resizing a window out of those states. Emulation never
touches the frame, so that guard was suppressing the reflow for no reason — and
a maximized window strands its dvh layout exactly like a normal one.

Run the Tahoe branch before the guard. Verified on macOS 26.3.1 that the
emulated viewport reflows a maximized and a fullscreen window while leaving both
states intact.

* fix(window): retry the viewport restore instead of stranding the renderer

If disableDeviceEmulation threw while the webContents was still alive, the
previous code swallowed the error and cleared the latch anyway, leaving the
renderer pinned at the emulated 1px-taller viewport for the rest of the window's
life — and letting the next reveal stack a fresh cycle on top of it.

Retry the restore on a bounded schedule and hold the latch while a retry is
pending. A destroyed webContents still short-circuits, since the emulated
viewport dies with it, and the attempt budget keeps a permanently failing
restore from pinning the latch forever.
…yai#10294)

Files: 18 applied, 0 deleted (from 6eb70d8)

Co-authored-by: Orca <help@stably.ai>
* oom(01): A1-shared-readers — reintroduce stablyai#10179 subset

Files: 18 applied, 0 deleted (from 6eb70d8)

Co-authored-by: Orca <help@stably.ai>

* oom(02): A2-shared-image-media — reintroduce stablyai#10179 subset

Files: 7 applied, 0 deleted (from 6eb70d8)

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
…tablyai#10485)

The +1px emulated viewport changed the CSS box, so a terminal sitting one
pixel under an xterm row boundary gained a row. The pane fit observer's
two-frame stability check reads that transient grid as stable well inside
the 32ms hold, forwards a real PTY resize, then reverses it on restore —
two SIGWINCHes per reveal, measured on 3/54 window heights (~1/cellHeight).

Nudging the device scale factor instead re-runs layout with byte-identical
CSS geometry. Same reflow, 0/54 SIGWINCH, no WebGL atlas rebuild or context
loss, and webview guests stop seeing spurious native resizes too.

Co-authored-by: Orca <help@stably.ai>
)

* fix(ui): stop the reveal reflow from dismissing popovers

The main process reflows the renderer on every reveal, resume, and restore.
On macOS 26 that nudges the emulated device scale factor, and Chromium fires
a real window resize for it even though innerWidth/innerHeight are identical.
Anything bound directly to resize treated that as a user resize.

The selection copy menu and the markdown link bubble both dismiss on resize
unconditionally, so they closed on their own whenever the window was revealed
or restored. Gate both on an actual dimension change.

Only fixes the macOS 26 path: the pre-Tahoe repaint jiggles the native frame
by a real pixel, so the renderer sees a genuine dimension change and no
change-detection guard can — or should — filter it.

Co-authored-by: Orca <help@stably.ai>

* test(ui): cover the reveal-reflow guard at the component level

Pins both halves of the wiring: a same-size resize must not dismiss, a real
one still must. Mutation-checked — reverting to a bare resize listener, or
dropping the listener entirely, each fails the suite.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
…ve (stablyai#10299)

* oom(01): A1-shared-readers — reintroduce stablyai#10179 subset

Files: 18 applied, 0 deleted (from 6eb70d8)

Co-authored-by: Orca <help@stably.ai>

* oom(02): A2-shared-image-media — reintroduce stablyai#10179 subset

Files: 7 applied, 0 deleted (from 6eb70d8)

Co-authored-by: Orca <help@stably.ai>

* oom(03): A3-shared-fs-listing — reintroduce stablyai#10179 subset

Files: 21 applied, 0 deleted (from 6eb70d8)

Co-authored-by: Orca <help@stably.ai>

* oom(04): A4-shared-remote-relay — reintroduce stablyai#10179 subset

Files: 8 applied, 0 deleted (from 6eb70d8)

Co-authored-by: Orca <help@stably.ai>

* oom(05): A5-shared-misc — reintroduce stablyai#10179 subset

Files: 28 applied, 0 deleted (from 6eb70d8)

Co-authored-by: Orca <help@stably.ai>

* oom(06): B-shared-wiring — reintroduce stablyai#10179 subset

Files: 81 applied, 0 deleted (from 6eb70d8)

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
…stablyai#10484)

* fix(terminal): verify Windows PTY root identity before taskkill /T /F

killWithDescendantSweep guarded its Windows tree kill with ownsRoot()
alone, which is JS state only. node-pty's ConPTY exit watcher closes the
last shell handle before it queues the JS exit callback, so Windows can
recycle the PID while the session map still looks live — force-killing an
unrelated process and its whole descendant tree.

Walk the recycled PID's ancestry back to this process before taskkill:
skip the sweep when the root is gone or resolves to a stranger, and keep
the sweep when identity is unknown so stablyai#10004 orphan cleanup still runs.
Also gate the local provider's ownsRoot on observed physical exit.

* fix(terminal): dedupe the Windows root-identity scan, drop dead exit gate

Review fixes on the PID-identity guard.

The probe read the process table through a new uncached export, bypassing
the reader that worktree teardown depends on: worktree-teardown.ts fans out
32-wide inside a 10s deadline, so a delete forked 32 powershell cold-starts
(the churn windows-foreground-process-rows.ts:25-32 warns about, stablyai#6288/stablyai#6667).
getFreshSnapshot() already guarantees a scan that starts after the request --
the exact property the bypass existed for -- and coalesces concurrent callers,
so use it. Measured on the new test: 32 scans -> 1.

The PhysicalExitTracker.hasExited gate could never fire. markExited() is only
reached at local-pty-provider.ts:985/:1431, and both are followed synchronously
by clearPtyState(), which deletes the ptyProcesses entry -- so ownsRoot's map
check is already false whenever hasExited is true. Reverting it broke no test.
Drop it and the shared getter it added; the identity probe already covers every
ownsRoot caller from inside killWithDescendantSweep.

Also point the Windows terminal-restart E2E job at the files that own this
behavior, so a change to the new Windows-only module runs the one job that
executes on a real Windows host.

* docs(terminal): state what the Windows root probe actually proves

The probe checks subtree membership, not root identity: a recycle that lands
on another Orca descendant (another pane's shell, an agent CLI, a git.exe we
spawned) still reads `own`, and that is not remote during teardown when Orca
is itself allocating pids. It bounds the blast radius rather than closing the
class. Say so at the type and at classifyWindowsTreeKillTarget, and name what
a real close would need (a CreationDate baseline -- the analogue of the POSIX
lstart check already used here -- or an inherited handle / Job Object).

Also note why our own pid must classify `foreign`.

* ci(windows): trigger the terminal-restart E2E on the shared snapshot reader

The Windows root-identity probe now reads through getFreshSnapshot, so an edit
to that module changes Windows teardown behavior without touching any path the
job already watches.

* test(terminal): guard the teardown probe against a reintroduced scan bypass

The existing volume guard covers queryWindowsProcessRowsFresh directly, but the
identity-probe cases all inject readRows, so nothing exercised the DEFAULT
reader wiring -- a bypass reintroduced inside windows-pty-root-identity would
have gone unnoticed. Drive verifyWindowsTreeKillTarget 32-wide through the real
reader and assert one scan. Verified it fails at 32 when the bypass is put back.
…stablyai#10478)

* fix(daemon): keep agent-completion detection alive on pre-v27 daemons

DaemonPtyAdapter.inspectProcess() threw terminal_liveness_unavailable when
the connected daemon predated protocol v27. The intended provider-level
fallback only fires when a provider lacks inspectProcess, so for the daemon
adapter the throw propagated: agent-completion-coordinator swallowed it into
consecutiveInspectionErrors and retried forever, killing process-exit
completions and pending-title validation.

Daemons intentionally survive app updates, so updating in place with agent
terminals open routes those PTYs to a legacy adapter and permanently
disables agent-finished notifications until the terminal is recreated.

Compose the inspection client-side from getForegroundProcess, which v26
fully supports. No new wire traffic and no new daemon capability.

* test(daemon): pin null-foreground semantics on the pre-v27 inspect fallback

The legacy composition had no coverage for a null foreground, which is the
one daemon response shape whose semantics diverge from v27: there
inspectProcess goes through getAliveSession() and throws for a vanished
session, while getForegroundProcess is deliberately null-not-throw. It is
also the only shape that reaches a user-visible completion, so reading it
as idle is a deliberate choice that should not change silently.
…ckpoints (stablyai#10479)

* fix(terminal): stop cold restore dropping all scrollback on large checkpoints

The checkpoint read cap shipped without its write bound. history-reader.ts
reads checkpoint.json through a 16MiB cap that throws past the limit, and the
catch swallows it to checkpoint=null; history-manager.ts still writes the
checkpoint with an unbounded JSON.stringify. Every fallback then collapses
(stale-generation log, unlinked legacy scrollback.bin), so the terminal
reopens empty with nothing surfaced.

Raise the read cap to cover the largest checkpoint the writer can legitimately
emit, derived from the scrollback policy's 50k-row max preset so it cannot
drift back under the writer. A bound is kept so a corrupt file still cannot
OOM the main process.

Bounding the writer instead would not recover the scrollback: the stringify
throw lands in handleWriteError, which adds the session to disabledSessions
and permanently stops history recording for it.

* fix(terminal): anchor checkpoint read cap to its own reasoning

The cap was derived as 2 * LEGACY_TERMINAL_SCROLLBACK_BYTES_100_MB, but that
constant is a legacy byte-preset setting value with no other consumer, and the
50k-row bucket it was attributed to has no upper byte bound. Same value, stated
without the false policy linkage.

* fix(terminal): assert the checkpoint byte cap, correct its rationale

The retained oversized-checkpoint test passed with the byte guard removed
entirely — 200MB of NUL fails JSON.parse, so detectColdRestore returned null
either way. Assert the bounded reader directly, as the amplification test does.

Also: a 50k-row max preset of ordinary text measures ~14MB serialized, not
'far below' by an unbounded margin — per-cell-colored output can still exceed
the cap, which is what a writer-side snapshot trim has to fix.
…nmounts (stablyai#10480)

* fix(mobile): heal an orphaned native-chat image paste across screen unmounts

The stale-input marker lived in a per-screen `useRef`, but the condition it
tracks — a bracketed image paste sitting unsubmitted on the agent's composer
line — lives on the host and outlives the screen. Backing out of a session and
returning remounted the hook with an empty Set, so the next message submitted
on top of the orphaned paste and the agent received `<image path><text>`.

Move the marker to a module-level store keyed by terminal handle, and consult
and consume it from every write path that can submit the composer: the image
hook's text-only send, the controller send (which the chat overlay's question
card reaches directly, bypassing the image hook), and the ask-answer send.

Permission choices and the Escape cancel deliberately do NOT heal: they are
`enter: false` keys for an active overlay that swallows the clear, so healing
there would consume the marker without clearing the line and leave the next
real message corrupted. Desktop scopes its Ctrl+U the same way.

* fix(mobile): stop the ask heal from burning the marker on selector answers

The heal ran on every ask answer, but Claude's and Codex's selector shapes
cannot submit the composer: a single-select answer is a bare option digit and
every stepping group is written `enter: false` (the host coerces it), so the
clear is swallowed by the live overlay while the host still acks the write.
That consumed the one-shot marker and left the orphaned paste to corrupt the
next real message — the same failure this PR exists to fix, through a new door
that main did not have.

Scope the heal to the pasted-label shape, which does commit the composer.
Desktop splits it the same way: use-native-chat-interactive-send.ts routes only
the non-stepping answer through the clearing sender and never pre-clears
sendNativeChatAskAnswer.

Also pin the three deliberate skips (selector answer, permission choice, Escape
cancel) with tests, so the PR's central design argument is an invariant rather
than a comment, and guard the failed-heal toast with the generation check every
other error surface in answerAsk already uses.
xianjianlf2 and others added 29 commits July 27, 2026 21:43
…0091)

Refresh the persisted Windows PATH during preflight without blocking Electron's main thread. Bound and deduplicate registry reads, preserve the last good cache on failure, skip host refresh for WSL, and add Windows regression coverage.
…11062)

* feat(new-workspace): make the project picker a type-ahead field

The Create-worktree Project slot read as bulky and unpolished: a label row,
an add-project icon, a 36px outline trigger and a chevron, all spent before
choosing anything — then a popover carrying its own *second* search box,
two-line rows, and a footer that scrolled out of reach.

The field is now the search. Typing filters in place, so the nested search
box is gone. Exactly one row is armed at any time and Enter takes it;
hovering arms, so pointer and keyboard drive one cursor rather than two
competing highlights. Armed is tracked by row key, not index, so a list
arriving late over SSH cannot slide a different project under a keypress
the user already aimed.

Rows are single-line at 28px with an on-row Enter cap that takes space only
while armed, and "Add a new project" is pinned to the popover edge so it
survives every state — scrolled, filtered to nothing, or no projects at all.

Long names and deep paths degrade deliberately: the name keeps up to half
the row and the path elides from its middle, so two monorepo siblings stay
distinguishable as …/services/checkout-api vs -web where a flat truncate
rendered both identically.

Recency is derived from when each project last had a workspace created,
which is the action this picker is about to repeat — no new store field.

The shell keeps data-project-combobox-root + role=combobox and stays
focusable, so the composer's initial-focus and project-required handlers
still land on it.

* fix(new-workspace): align, scroll and loosen the project picker

Five fixes to the type-ahead picker, three reported and two found while
checking for related breakage.

Alignment: the name and its smaller detail line were centred as boxes, so
the 12px path sat visibly high against the 14px name. Both now share a
baseline, in the committed field and in every row. The dot mark and the
Enter cap are chips rather than text, so they stay centred on the row.

Scrolling: the mouse wheel did nothing over the list. The composer is a
Radix Dialog, and react-remove-scroll cancels wheel events for portaled
content outside the dialog's DOM tree — the scrollbar dragged fine but the
wheel was dead. The old cmdk list carried a shim for exactly this; the
plain scroll pane that replaced it did not, so it has its own now.

Density: rows go 28px -> 32px, row text 13px -> 14px and detail 11px ->
12px, with a taller Add row and more air above section headings.

Escape stranded a query: with the list closed but text still typed, Escape
was ignored (it was gated on the list being open), leaving the field
showing text that matched nothing and hid the committed project. Escape now
always restores the committed display, and only bubbles when there is
nothing to undo.

Listbox ownership: options sat inside unroled section and scroll wrappers,
which breaks the listbox -> option relationship assistive tech relies on.
Sections are groups carrying the heading as their label, and the scroll
pane is presentational.

Both new behaviours are covered by tests verified to fail without the fix.

* chore(tools): keep the project-picker design lab

The exploration harness behind the picker rewrite: 16 interactive design
variants rendered against the app's real tokens and shadcn primitives, so a
prototype is a drop-in ProjectCombobox rather than a mockup.

Worth keeping because the frames encode bugs that only reproduce in
context. DialogFrame renders the picker inside a real Radix Dialog, which
is the only way the react-remove-scroll wheel bug shows up; the fixtures
carry duplicate display names and deep sibling paths that a naive truncate
renders identically.

Run with: npx vite --config tools/wt-picker-lab/vite.config.ts

* fix(new-workspace): stop the project list flashing open, shrink its empty state

Opening the picker read as a double flash. The shared popover surface is
translucent and fades 0 -> 1, which is right over the app canvas but wrong
here: this popover lands directly on the composer dialog, so for the length
of the fade the Name field underneath showed straight through the list and
you saw two layers at once. The list now uses an opaque surface and zooms
without fading, so it is solid from the first frame. Every other popover
keeps the blur and fade.

The "No projects match your search." state was a 60px centred block sitting
next to 32px rows, which read as a different kind of surface and made an
empty result feel like an error. It is now sized and aligned like a row.

The lab's dialog frame focused whatever Radix picked first, which popped the
Add-project tooltip on open and masked the real problem; it now focuses the
name field the way the real composer does.

* fix(new-workspace): square mark, centred empty state, and keep the list open on tab-focus

Four fixes, three reported and one found while sweeping for others.

Square mark: the option dot had a `rounded-full` override, so a project read
as a circle here and a square everywhere else (jump palette, sidebar). Drop
the override and use RepoBadgeMark's own shape.

Centred empty state: "No projects match your search." was left-aligned after
being shrunk to row height; centre it.

Tab-focus blinked the list shut: the field lives in the popover's anchor, not
inside its content, so Radix's dismissable layer saw focus land "outside" and
closed the list the instant you tabbed in. Focus and pointer events within
this control no longer dismiss it; genuine outside events still do.

Junk text could strand the field: typing a query that matched nothing and
then clicking away left the text sitting there with the list closed, showing
no project and no error. A query only means something while the list is open,
so closing without committing now clears it.

On pressing Create with no project: no change needed. The create gate has not
depended on project selection since stablyai#4991, and both submit paths already call
showProjectRequiredError(), which sets the inline message and turns the field
red via aria-invalid. Verified end to end: the button is pressable, the press
paints the field destructive, and the message appears beneath it.

* feat(new-workspace): rebuild the Run-on picker to match the project picker

"Run on" was the last composer field still built the old way: an outline
trigger wrapping a cmdk list, two-line rows, and no way to search. It now
matches the project picker, so the two fields in the same form read as one
control.

The field is the search — type to filter hosts, paths and recipes with no
nested search box. Exactly one row is armed at a time and Enter takes it;
hovering arms, so pointer and keyboard drive one cursor. Rows are 32px with
the label and its path on a shared baseline, the path eliding from its
middle so two deep sibling paths stay distinguishable. The popover surface
is opaque and unfaded because it lands on the composer dialog, where a
translucent fade shows the form underneath.

Two behaviours the project picker doesn't have are preserved. Disconnected
hosts keep their inline Connect action, tracked per host so one stalled
connect never blocks the others, and the list stays open so the connecting
state is visible. Two rows open nested lists rather than committing: VM
recipes, and "Add host" pinned to the popover edge so it survives every
state — scrolled, filtered to nothing, or with no hosts at all. Enter and
ArrowRight open a submenu; Escape backs out one layer at a time.

Extracted from NewWorkspaceComposerCard (-563 lines) into files that each
stay under the line limit without a suppression.

Tests: the run-target cases asserted cmdk internals (`[cmdk-item]`,
aria-disabled, cmdk-separator) that no longer exist. Rewritten against
behaviour and the listbox roles instead. All 22 composer tests pass,
plus a live sweep of 11 interactions in a real dialog.

* fix(new-workspace): drop the Enter cap, fix submenu hover, match the Add rows

Three follow-ups on the two composer pickers.

The ↵ cap on the hovered row is gone from both. On a run-target row it sat
next to the Connect action and read as a second, competing affordance; the
highlight already says what Enter will take.

Submenu rows never highlighted under the pointer. They passed a hardcoded
`armed={false}`, so the recipe list and the Add-host choices were the only
rows in either picker with no hover state. They now track their own hover.

"Add a new project" used a chunky FolderPlus where "Add host" uses a plain
Plus. Both rows were already the same height and type, so matching the glyph
is the whole difference.

* fix(new-workspace): restore the folder glyph, two-line Add-host cards, quiet Connect rows

Three follow-ups.

"Add a new project" goes back to FolderPlus — matching "Add host"'s plain
Plus made the two consistent but lost the glyph that says which kind of
thing is being added.

A disconnected host row no longer repeats its status. The Connect button
already says the host isn't connected, so "Connect this host to set up
projects" beside it was saying it twice. Rows without a Connect action keep
their detail, since there it explains why the host can't run.

The Add-host choices go back to two-line cards. Their descriptions explain
what you're picking ("Use an existing machine over SSH" vs "Pair another
Orca runtime"), unlike a host row's detail, which just labels a host you
already recognise. RunTargetRow grows a `stacked` variant for that rather
than making the single-line row do both jobs.

* fix(new-workspace): give Run on the same vertical rhythm as the other fields

Run on is nested inside the Project block so the two share its error and
empty states, which also put it on that block's 4px internal spacing. It
reads as its own field, so it sat noticeably tighter than the 16px gap
every other field in the composer gets. Pad it to match.

* refactor(new-workspace): share the type-ahead machinery between both pickers

Project and Run on were built one after the other, so each grew its own copy
of the same mechanics: query and open state, arming by row key, the
arrow-key walk, scroll-the-armed-row-into-view, the react-remove-scroll
wheel shim, and the closes-drops-the-query rule. Two copies of subtle
behaviour is two places for it to drift.

useTypeAheadCombobox now owns all of it. Callers pass a function that turns
a query into row keys and get back the query, the armed key, and the
movement helpers. Run on layers its submenu state on top by wrapping
`close`, which is the only part that isn't shared.

The two long class strings both files repeated verbatim — the field shell
and the opaque unfaded popover surface — are named constants now, so the
reason they differ from the stock popover recipe is written down once
instead of implied by a duplicated literal.

No behaviour change: 16,456 renderer tests pass, plus the 22-check live
interaction sweep across both pickers in a real dialog.

* fix(new-workspace): drop aria-expanded from option rows, remove the design lab

`aria-expanded` isn't a supported prop on `role="option"`, so the submenu
rows were claiming a state screen readers can't interpret there.
`aria-haspopup` alone already says the row opens a menu.

Removes tools/wt-picker-lab. It was the harness for exploring this redesign
— 12 interactive variants — and it did its job, but the 11 that lost are
dead code, and its prototypes were the only thing failing the react-doctor
gate (5 errors, all in throwaway variants; the shipped pickers had none).
…al parking (stablyai#11091)

The Developer submenu shipped visible on every worktree right-click, and its
Park terminal action always refused with "These terminals cannot be parked
safely."

- reveal the Developer submenu only when Option/Alt is held at right-click,
  captured at open time so it can't shift rows mid-menu
- stop a settled pendingActivationSpawn tag from refusing a manual park: first
  activation stamps it on every tab and only a fresh updateTabPtyId consumes it,
  so a reattached tab kept it forever
- resolve the single leaf of a rootless layout for parked watcher coverage, so a
  workspace whose panes never mounted is no longer permanently uncoverable
- restore the !isVisible park guard dropped in stablyai#11016, which let the workspace
  being viewed unmount its own terminals
- split manual-park eligibility out of the automatic cold-park policy module
)

Add translations for OrchestrationPage coordinator and child PR names, plus agent workflow status messages (initial states and progress beats) across all supported languages (English, Spanish, Japanese, Korean, Simplified Chinese).
…tablyai#11105)

The Update skills dialog showed "The update didn't finish" / "Some skills
could not be updated" with an armed Retry, directly above the runner's own
log line saying "All global skills are up to date" — on a clean exit 0.

`skills update` compares its lock's recorded hash against the source and
never reads disk (dist/cli.mjs: `latestHash !== entry.skillFolderHash`).
Once the lock has advanced past the installed bytes it prints up-to-date,
exits 0 and writes nothing. The copy left behind is a recognised older
revision, so the post-run re-scan sees `outdated`, skillUpdateFailedNames
counted that as a failed run, and skill-update-run settled to state 'error'.
Retry re-ran the same command, which no-op'd again — a closed loop.

Reclassify `outdated` as "the command did not converge this", not "the run
failed". The freshness badge still marks the copy not-current, so nothing is
hidden; the run just stops being blamed for it.

A botched write is still caught: a half-written bundle hashes to
`unrecognized`, a wholly-degraded or removed copy leaves no convergent
placement, and process-level failure still surfaces via the spawn error.

Reachable by anyone who updated during the stub conversion window — every
bundled skill has a stub -> full -> stub oscillation in its last three
registry revisions.
…yai#11045)

* perf(terminal): eliminate dense control and frame gate regressions

* test(terminal): keep gate labels in valid expect shape

* test(terminal): expose the surviving sub-threshold control-density case

The only adverse strip fixture sat at 50% control density, which is exactly
where the fallback fires and wins. A shape at 31 controls per 64-unit block
evades the trigger and still loses to the per-character legacy (0.67x), so the
benchmark structurally could not show it.

Add that fixture, pin both density literals in the staleness guard so a retune
fails loudly instead of silently measuring a boundary that moved, and export
the probe constant the equivalence test was hardcoding.
* fix(agent-status): track Codex rollout subagents

* fix(agent-status): resolve cross-day Codex child rollouts and unblock CI gate

Codex files each rollout under its own local start date, so a session that
runs past midnight spawns children into a sibling day directory. Scanning
only the parent's directory left 13% of real subagent spawns (48/371 across
local rollouts) permanently unresolved, which pinned a phantom "working" row
and re-ran readdirSync every poll tick forever. Resolve the child's own day
directory from occurred_at_ms, and time-box a child whose rollout stays
unreadable so a deleted or never-written file can't leak a working row.

Also make the hook HTTP handler return void: the changed-code quality gate
keys findings by span overlap, so this PR's added line inside the pre-existing
async createServer callback resurfaced no-misused-promises as a new finding.

Tests cover cross-day resolution, grace-period retirement, and that the poll
re-arms across successive roster changes (the prior tests passed even when
the poll died after its first change).

* fix(agent-status): keep the Codex subagent poll alive across nested hooks

A nested non-codex CLI inherits its parent's ORCA_PANE_KEY, so its hook
POST reached scheduleCodexSubagentPoll and tore the timer down before the
source guard, silently ending polling while a rollout child was still live.
…tablyai#11110)

* Revert "fix(skills): stop reporting a failed update when the CLI succeeded (stablyai#11105)"

This reverts commit a866083.

* fix(skills): stop offering an update the CLI provably cannot perform

The Update skills dialog reported "The update didn't finish" / "Some skills
could not be updated" with an armed Retry, directly above the runner's own
"All global skills are up to date" log line, on a clean exit 0.

`skills update` decides what to do by comparing its lock's `skillFolderHash`
against the source tree and never reads disk (published CLI, dist/cli.mjs:
`latestHash !== entry.skillFolderHash`). So once the lock records a revision
the filesystem does not actually have, it reports up-to-date, exits 0 and
writes nothing. No retry converges it.

Gate eligibility on that instead of blaming the run afterwards: a name stays
updatable only while the lock's hash and the revision the DISK hashes to
agree. Those that disagree fall through to `needs-attention`, which
skill-freshness-display-status.ts already documents as the state for a copy
"out of date somewhere the update command cannot reach". The dialog no longer
offers them, so it can neither claim failure nor claim success.

Deliberately compared against disk, not the bundled manifest: the updater
pulls from the source repo, which legitimately runs ahead of what a build
ships, and gating on the bundle would withhold real updates. Both sides must
also be positively identified — an unplaceable lock hash is not evidence.

Also reverts a866083, which forgave `outdated` post-run. That turned the
false failure into a false success ("Updated 2 skills" over copies nothing
wrote) and let an empty verdict swallow spawnError, so an offline run
published green.

Reachable by anyone who updated during the stub-conversion window — every
bundled skill has a stub -> full -> stub oscillation in its last three
registry revisions.

* fix(skills): do not gate a skill when a placement is unidentifiable

`diskTreeShas` drops digests matching no known revision, so `every` ran over
only the resolved half — one stale copy beside one unidentifiable copy gated
the name, contradicting the unknown-stays-eligible rule the function documents.
Require every observed digest to resolve before treating all placements as
mismatched.

Caught by CodeRabbit on stablyai#11110.

* fix(skills): judge convergence only over placements the update command writes

A same-name plugin-cache or repo copy could defeat the gate two ways: an
unidentifiable repack read as an unresolved placement, and a cache copy
parked at the lock's own revision read as an anchor — both re-arming the
unwinnable update on a drifted canonical. Filter to
SUPPORTED_GLOBAL_SKILL_TOPOLOGIES, matching eligibility and outcome.
Consolidate standalone code-quality scanners into Oxlint, preserve focused native/type-aware enforcement, add custom plugin coverage, and harden deferred PTY test cleanup.
…lyai#11098)

* fix(browser): keep the address bar editable in a narrow toolbar

Every other browser toolbar control is shrink-0, so the address bar was
the only flexible item and absorbed the entire squeeze: below roughly
420px of pane width it collapsed to the leading globe icon with a
zero-width input. Clicking it only opened the suggestion dropdown, which
inherits `--radix-popover-trigger-width` and so rendered at icon width —
there was no way to type or edit a URL in that tab.

Focusing a squeezed bar now lifts the form out of the toolbar flow and
overlays the row edge to edge, giving a full-width editable field that
navigates on Enter (and a full-width suggestion list for free). A
measured slot stays in flow so the overlay cannot feed back into its own
width, and the slot keeps a min width so the globe remains a real hit
target instead of being overlapped by neighbouring buttons.

Fixes stablyai#11090

Claude-Session: https://claude.ai/code/session_01Mx53f7erbtw5NraS8HdXKE

* fix(browser): use the documented floating shadow for the expanded bar

STYLEGUIDE.md defines exactly three elevation levels and forbids a
fourth; shadow-md was not one of them. The overlaid address bar is a
floating surface, so it takes the documented floating shadow already
used by the other floating surfaces in this pane.

Claude-Session: https://claude.ai/code/session_01Mx53f7erbtw5NraS8HdXKE

* test(browser): make the narrow-toolbar regression deterministic

The spec passed only from a clean profile. Two preconditions it set once are
actively undone by the app:

- BrowserPane re-focuses a blank tab's address bar across several animation
  frames plus the blank-url did-finish-load handler, so a single blur() was
  reverted and the bar never reached its squeezed resting state.
- Startup paths re-open the right sidebar. At a fixed 700px window that leaves
  the pane ~70px, so the overlay had nowhere to go and the field measured 0px.

Settling these separately let whichever settled first drift back while the next
one ran. Re-assert them in one loop until they hold simultaneously, and size the
window from the chrome actually measured instead of assuming a fixed 700px.

Verified 8/8 green, and still fails at the overlay assertion when the fix is
disabled, so the regression coverage stays real.

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
…yai#11107)

* fix(orchestration): clarify legacy migration safety

* fix(cli): sanitize legacy formatted messages

* test(runtime): allow near-cap fuzz under shard load

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Show group 5 and group 6 QR codes side by side so people can join group 6 if group 5 is full.
…tablyai#11139)

stablyai#8321 mixed the selected card's wash into the opaque --worktree-sidebar
surface, lifting dark mode to 16% (#4b4b4b). Card text lost too much
contrast against it.

Return both modes to a translucent wash (light 8%, dark 10%) so the
brighter selection border added by stablyai#8321 carries the selected state
instead of the fill.

Co-authored-by: Orca <help@stably.ai>
* fix(macos): avoid redundant focus on app activation

* test(macos): cover passive app activation

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
…i#11131)

* ci(pr): run E2E when a PR touches tests/e2e paths

Regression specs under tests/e2e never ran on PR CI — only schedule and
release called e2e.yml — so a red regression test could merge green.
Path-filter and workflow_call the E2E suite when E2E-relevant files change.

Use merge-base diffs so base-branch drift does not false-trigger E2E, fail
the detector when git diff cannot compute the PR range, and pin
least-privilege contents:read on both the detector and reusable E2E workflow.

Closes stablyai#10518

Co-authored-by: Wooseong Kim <innocarpe@gmail.com>

Co-authored-by: Orca <help@stably.ai>

* ci(pr): make the E2E path gate actually block, and match the real config path

Two fixes to the new path-filtered E2E job.

The gate did not gate. pr.yml's `verify` job is the required check, and it
enumerates its dependencies explicitly — `e2e` was in neither `needs` nor the
result list, so a failing shard left `verify` green. That reproduces the exact
hole this job exists to close: a red spec merges green, just with a red box
further down the page. Add `e2e` to both.

Because the job is path-filtered, `skipped` is the normal result on a PR that
touches no E2E files and has to keep passing. That allowance is checked after
the strict loop rather than inside it, so it can never leak to the six jobs
that are always required.

The `playwright.` pattern matched nothing. The config is
tests/playwright.config.ts — beside tests/e2e/, not inside it — so no tracked
file starts with `playwright.` and editing the runner config would silently
skip E2E. Anchor it at `tests/playwright.`.

Adds a contract test alongside the existing release-e2e one. Verified it fails
when either fix is reverted, and simulated the gate across
success/skipped/failure/cancelled plus the skip-must-not-mask-a-real-failure
case.

* test(ci): close two gaps in the E2E gate contract

CodeRabbit was right on both counts — verified by reverting each and watching
the contract stay green.

The path filter was unasserted, so `e2e` could lose its `if:` and run on every
PR — the cost the filter exists to avoid — without failing anything.

The strict-loop check hardcoded four of the six required jobs, so dropping
GIT_COMPATIBILITY or SHELL_CONTRACTS left them unenforced while the contract
passed. Derive the list from verify.needs instead, so a newly added required
job that misses the loop fails here rather than silently going unchecked.

* ci(pr): land the E2E path gate advisory instead of blocking

The E2E suite is currently failing every scheduled run on main — 22 of the last
22 — so making verify depend on it would block any PR touching tests/e2e/**,
including the PRs that fix the suite. This PR's own run reproduced that: 3 of 12
shards failed on specs unrelated to it (agent-session resume, Jira linking,
plugin containment, terminal artifacts).

So the job runs and reports on E2E-path PRs but is left out of verify.needs for
now. The detector, the tests/playwright. path fix, and the contract tests are
unaffected — those stand on their own and were the substance of the review.

Flipping to blocking is a three-line change once the suite is green; the exact
wiring, including why the skipped allowance must sit outside the strict loop, is
recorded on verify's Require-successful-checks step. The contract test pins the
advisory choice so it reads as deliberate rather than as the unwired-gate bug it
originally caught, and still fails if the path filter, the strict-loop coverage,
or the config path regress.

---------

Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
The cadence picker now resolves Hourly, Daily, Weekdays, Weekly, and Custom cron through the renderer i18n catalog, with corrected zh/ko/es translations.

Supersedes stablyai#10044 (thanks @innocarpe) and stablyai#10045 (thanks @fsdwen).

Fixes stablyai#10043
Removes the lowercase hourly/weekly/custom entries from the AutomationSchedulePicker catalog node in all five locales. They were superseded by keys derived from the rendered labels in stablyai#11068, and the extractor only adds keys, so they stayed behind unreferenced.

Follow-up to stablyai#11068.
…eate

Expose the existing IssueSourceSelector on the New Workspace smart name
field when a fork has a divergent upstream remote, and re-fetch smart
GitHub results when the preference flips. Tasks already had this control;
branch-from-issue create did not (stablyai#9281).

Closes stablyai#9281
… disabled

CodeRabbit: match only host/repo segments; hide selector when repo sources off.
Compact U/O only changes issue routing; show "Showing issues from …"
so the control is not mistaken for PR routing.
@innocarpe
innocarpe force-pushed the fix/composer-issue-source-switch branch from 2a1bd48 to 2a88ff4 Compare July 28, 2026 12:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.