Skip to content

[upstream #24452] fix(monaco): highlight Svelte block closers inside markup - #300

Closed
innocarpe wants to merge 5402 commits into
mainfrom
fix-19918-svelte-block-close
Closed

innocarpe wants to merge 5402 commits into
mainfrom
fix-19918-svelte-block-close

Conversation

@innocarpe

Copy link
Copy Markdown
Owner

Portfolio mirror of my contribution to upstream stablyai/orca.
Exhibition only — the real review/merge target is upstream.

Upstream

Summary

Description Svelte block closers such as {/if} and {/each} were tokenized as HTML once markup had entered the html embed. Monaco only consults parent rules that pop the active embed, and the closer was a plain string token, so the html tokenizer swallowed it. ## Focused

Note

  • Do not merge this into innocarpe/orca main until the upstream PR is merged.
  • After upstream merges: sync fork from upstream, then close this mirror PR.
  • This open PR exists so visitors see in-flight work on this fork's Pull requests tab.

Jinwoo-H and others added 30 commits September 30, 2026 02:16
stablyai#24026)

Delivery collapsed on 2026-09-29 once send volume doubled: the worker, the
retention pruner and the request path share a two-connection pool, and the
database transaction rate pinned at ~140/s regardless of how many notifications
were delivered. Raise the pool to six so worker and pruner stop serialising on
one connection. The budget precondition stays satisfied (2 x 6 x 3 = 36 <= 64).

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
…ecorded requests against the desktop's params rules (stablyai#23732)

* test(mobile): add rpc:diff to decode what a recording change moved

The RPC recording goldens are content-addressed JSON, so their raw git diff is
pool hashes. `pnpm --dir mobile rpc:diff [<base>]` decodes both sides and prints,
per golden, the checkpoint, field and JSON path that moved with both values,
grouped across checkpoints, plus added and removed goldens. `--summary <file>`
appends a Markdown report capped for GitHub's step-summary limit.

It reads any pooled format, so it can prove the next commit's format change
moves no recorded value. Checkpoints are matched by occurrence because an id can
repeat within one golden.

This commit adds files under the recorder directory, which moves the header
digest every golden pins; the next commit removes that header.

* test(mobile): record RPC goldens without a pinned commit or input digests

Every golden carried a pinned `baseline` commit plus digests of the recorder,
its mount adapter and its scenario, and the record script refused to run unless
the product tree matched the pin. So every behaviour change repinned to its own
branch commit and rewrote all ~790 files, the squash made that commit
unreachable, and main's pin job stayed red until a hand-made repin pull request
landed (22 of them in 12 days). The digests could only fail when an input moved
and the recording did not, which is exactly the change that carries no
information; every run already re-derives each golden from the current tree and
compares it.

Format 6 keeps the format version, operation, family, named deltas, the value
pool and the recording. Removed: the pin and fence, the three digest modules and
their test, the pin guard and its CI job, and the dead scenario `version` field
(the manifest reader now refuses `baseline` and `version` with a message).

- `pnpm --dir mobile rpc:record [<golden-id>...] [--prune]` records all or some
  goldens; orphans are listed, and deleted only with `--prune`. Every derived
  test title now starts with its golden id so an id selects it.
- `compareGolden` reports every difference in one failure (identity fields by
  name, the checkpoint list, each checkpoint/field/path grouped), keeps the
  final byte compare, and ends with the command to re-record that golden.
- `unhandled-recording.test.ts` now drives a detached rejection through
  `runRecording` into a checkpoint and the cleanup checkpoint; no golden carries
  one, and disconnecting the capture passed every suite before.
- Seam rules that existed only to keep a digest honest are gone; the
  mutant-reachability, register-completeness and one-exposure rules stay.
- CI: `Mobile tests on main` runs the whole mobile suite on every merge that
  touches mobile/, src/shared/, the root lockfile or the host RPC paths, since
  `verify` never runs on main. A new `Mobile RPC Recording Replay` workflow
  replays the recordings on pull requests that touch src/shared/ or the root
  lockfile without touching mobile/. `verify` writes the `rpc:diff` report to
  the job summary.

Proof: `rpc:diff` against the parent reports no recorded behaviour moved; each
golden only loses its ten header lines.

* test(mobile): check every recorded request against the host's params contract

The goldens script the host's replies, so a scenario could record a success
for a request the real host would refuse, and a desktop change that tightens a
params schema moved no golden at all.

`recorded-request-params.test.ts` parses every distinct request the corpus puts
on the wire with the host dispatcher's own `parseRpcRequestParams` and the
schema `rpc-params-catalog.generated.ts` binds to that method. It fails on a
method the host lacks, params it refuses, params sent to a method that takes
none (the dispatcher never reads them), and keys the schema silently strips
unless an inventory entry gives the reason; a stale entry fails too. Each rule
is also shown firing on a made-up request, since the corpus has no instance of
three of them. It imports the desktop dispatcher, so it sits beside the other
Node-side tests outside the RN test program, and the params-contract boundary
now exempts test files, which are never bundled.

It found twelve requests the host would refuse, all from invented fixture
values, not product code, fixed at their source:
- git.branchDiff sent `base-oid`/`head-oid`/`merge-base` where the host needs
  full object ids (diff-review and source-control adapters, and the branch
  compare replies in the manifest that feed them);
- an iOS push registration without `apnsEnvironment`, which a real iOS token
  always carries (`push-token.ts`); the adapter now defaults to `production`;
- `settings.update` given Linear's `assigned` filter as a GitHub preset, which
  the product type forbids; the scenario now picks `my-issues`;
- GitLab `projectRef` as a string where the host and the product type take
  `{ host, path }` (7 methods, 5 adapters and the manifest).

46 goldens move, and a decoded comparison of every one of them shows no change
other than those substitutions; `rpc:diff` lists them.

* ci(mobile): detect a mobile change without a SIGPIPE-prone grep pipe

Under the runner's pipefail, grep -q exiting on its first match SIGPIPEs git
diff on a long file list, so a large pull request touching mobile/ read as
uncovered and replayed the recordings a second time.

* test(mobile): drop comments that still describe the golden header and digests

Eleven adapters justified an import rule by the header a golden no longer
carries, and that rule's test is gone. The census failure now names the
rpc:record and --prune commands.

* test(mobile): refuse a golden that keeps a key no recording writes

Decoding dropped unknown top-level keys, so an old header left behind by a
hand-resolved merge conflict passed every compare unseen.

* ci(mobile): summarize RPC recording changes after a failed test step too

* test(mobile): stream rpc:record output instead of capturing it

A captured run stayed silent for its whole duration and clipped its tail,
where the failure summary sits, past 8 MB.

* test(ci): let the Ruby-gate contract skip the always-run RPC summary step

fef088d gave the summary step an `if: ${{ !cancelled() }}`, and this test lists every gated step in `verify` and expects each to be gated on the Ruby scope.

* test(mobile): replay only a golden file that is exactly what rpc:record writes

Replay compared two re-encodings of decoded values, so anything decoding drops (a leftover header
key, a hand edit) sat in the committed file uncompared; a key allow-list covered one case of that.
Replay now passes only if the file text equals the formatted golden for the run, sharing one
formatter with writeGolden, and keeps the field-level report as the failure message. The allow-list
goes; the value-based compareGolden stays for the bridged run, which has no file.

* test(mobile): end a corrupt or hand-edited golden's failure with the re-record command

A hand edit to a pooled value failed in decode with only "Golden value <hash> does not hash to its
pool key": no golden id and no command to fix it. readGolden now prefixes parse and decode failures
with the golden id and ends them with the rpc:record command. The rpc:diff header also said it
always exits 0; it exits non-zero when git or a golden cannot be read, and now says so.

* ci(mobile): run Mobile Checks on every src/shared and root lockfile change

Replaces the replay-only workflow: mobile imports hundreds of shared modules, so a shared edit can
move a golden or break mobile's typecheck, and the full job catches both before merge. A root
lockfile-only change skips the Ruby release checks, which read no root Node dependency.

* test(mobile): list or prune orphaned goldens even when the recording run fails

Orphans come from the manifest, not the run, so a failed or timed-out rpc:record still reports
them; the exit code stays non-zero. README: say what a failed replay reports (first differing
path per field, capped groups) and what rpc:diff compares with and without a base.

* test(mobile): pin the RPC recording goldens to LF so a CRLF checkout still replays

Replay now requires the committed golden text to equal exactly what rpc:record writes, which is
LF. A Windows checkout with core.autocrlf=true converted every golden to CRLF and failed all 790
with "holds the same recording but is not the file rpc:record writes for it".
…lyai#24027)

resources/relay is the only relay copy a packaged build resolves, but out/relay
was also packed into app.asar — 14.2MB of unreachable duplicate. Kaspersky
flagged app.asar as a compound object precisely because relay.js was inside it,
so one script-heuristic verdict on relay.js gutted the whole install. Excluding
it decouples app.asar from that verdict and drops the duplicate bytes.
stablyai#24033)

stablyai#23900 probes codex --help before each launch; the fake counted it as a worker spawn, breaking two specs.
…test harness (stablyai#23986)

* test: open, seed and read the agent-session record store through one harness

Tests that open the durable agent-session record store, seed it, or read
back what it persisted now go through agent-session-record-store-test-harness.ts
instead of calling AgentSessionRecordStore.open or touching agent-sessions.json
themselves. A later change that moves the store into the chat database then
changes the harness instead of every test. No production code changes.

Tests whose subject is the JSON file itself (its .bak recovery, salvage,
schema versions, permissions, and what older builds read back) keep reading
and writing the file directly; the storage move rewrites or deletes them.

* test: address the record-store harness by the host's state directory

The harness took the store's own folder, so each caller picked one
(join(root, 'store'), or 'agent-sessions' where a test read the store the
runtime owns). A later change that moves the store into the state
directory's journal database could not tell those apart, and would have
had to edit every caller again.

Every harness function now takes the state directory, the one the test's
journal database and recovery capsule already live in, and keeps the
store in the same subfolder the runtime uses. Callers pass that directory;
store-only tests pass their temp directory unchanged. Format tests that
share a directory with harness calls take the file path from
testAgentSessionStoreFilePath.

The folder name moves from a private constant in the runtime to
AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness
shares it without importing the runtime. Its value and every path built
from it are unchanged.
Four drains, each holding one provider round trip of ~100 ms plus its
database statements, capped delivery near 30/s. Production inflow reached
35/s on 2026-09-30, so the backlog aged past the five-minute TTL and
notifications expired. Twelve drains lift the ceiling to roughly 90/s; the
pool is now six per instance, so the extra drains queue on connections
instead of starving the request path.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
Correct typo in PR template regarding issue linking for outside contributors.
… in the parent's conversation (stablyai#23752)

* refactor(native-chat): a subagent's rows live with that subagent, not in the conversation

A subagent's rows were drawn in its parent's conversation, each captioned with
the subagent's name. They now belong to the subagent: the transcript projection
keeps the session's own rows as the conversation and each subagent's rows apart,
keyed by the agent id its roster entry already carries, folded on their own.

Desktop: a subagent's rows open in a section under the roster row that names it,
from that agent's roster entry, and are windowed like any other rows. A subagent
no loaded roster names opens where its first row happened, inside the section of
the agent that spawned it or in the conversation. Its edits still count in the
turn they were made, and revealing one opens the sections around it.

Mobile shows the conversation, with each spawn's roster line. Worker reads and
structured terminal reads serve the worker's own rows.

Removes what the move makes redundant: the per-row caption and its copy, the
producer check in the tool fold and the turn answer, the per-agent frontier
interleaved in the conversation, worker-text subagent tags, and the agent id on
worker-read messages.

* refactor(native-chat): a diff target names the sections its row sits in

Revealing a subagent's edit opens the sections around it from the target the
rollup already holds, instead of looking the row up at click time. The section
head keeps to the agent's name and dot; its state in words stays on the roster
entry. The worker page test stubs the host through its module rather than a cast.

* fix(native-chat): a working subagent's section is open; a worker page windows its own rows

A subagent's section is open while its agent works and closes once it settles,
the way the turn's own live run does; a section the reader opened or closed by
hand keeps that choice. A subagent another subagent spawned opens inside that
one's section, so a working grandchild shows inside its working parent. Openness
is derived from the roster's state and the reader's choices; nothing stores an
automatic open.

A worker page is now the newest page of the worker's own rows. The host windows
the read over them before the limit, so a subagent's burst can no longer crowd
the worker's rows off the page, and "older" still means older worker rows. The
scope is an in-process argument of the host's history read; no wire request
carries it.

* fix(native-chat): a subagent section head names the turn it sits in, for the outline rail

* fix(native-chat): a subagent section's rows sit in the turn the section is shown in, for the outline rail

A background subagent's rows written during a later turn carried that later
turn onto their slots, so scrolling through its section lit the later turn's
rail tick and then snapped back. The rollup still counts each edit in the turn
it was made; only the slot, which the rail reads, takes the shown turn.

* fix(mobile): Load earlier reads past pages that hold only a subagent's rows

Mobile draws only the session's own rows, so an older page made entirely of a
subagent's rows landed as nothing: the reader tapped Load earlier, saw the
spinner, and got the same transcript back. One load now reads on (up to 8 pages)
until a page holds a row of the session's own, then applies the pages in order.

* test(mobile): stub the RPC client the way the other structured-session hook tests do

* perf(native-chat): order subagent rows for the changed-files rollup once per change to them

The rollup flattened and re-sorted every subagent row on each update, including
every token the parent streamed. The ordering now keys on the projection's
subagent rows, which keep their identity while only the conversation changes.

* refactor(native-chat): order subagent rows in the sections hook, keeping the list under its line limit

* fix(mobile): a transcript whose newest page is only a subagent's rows reads back on its own

Opened while a subagent is busy, the newest page can hold nothing but that
subagent's rows. Mobile draws none of them, so the reader saw an empty chat with
a Load earlier button, and an empty list cannot be scrolled to page. The hook now
reads back once from each such head, and the read runs on to the session's own rows.

* fix(native-chat): count the live window in the session's own rows, so a subagent's burst keeps its roster

The live window kept the newest 1,024 rows of every agent. A subagent writing
more than that trimmed its own spawn's roster row and the prompt, and its
section fell back to a closed, unnamed header. The window now keeps the newest
1,024 of the session's own rows and everything after, with an 8,192-row cap on
every agent's rows as the memory backstop. A transcript with no subagent rows
trims exactly as before.

* fix(agent-session): window history pages by the session's own rows, with a subagent's rows riding along

A history page held the newest 200 rows of every agent, so a subagent's burst
could fill a page on its own: the phone opened on an empty chat and "Load
earlier" landed nothing. A page now starts at the oldest of the newest `limit`
rows of the session's own and serves every row from there, so the subagent's
rows come with the conversation they happened in. The page stays contiguous,
the cursor still names its first row, and the byte bound still applies. A
transcript with no subagent rows gets the same pages as before.

Clients already take a page larger than its limit: both reducers raise their
retained window to the page's size. The mobile read-on and read-back stay for
older hosts.

* test(agent-session): a page reaches back to the start rather than leaving a subagent-only page

* fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it

The live window trimmed to just after the own row it dropped, so a subagent
whose roster row went kept its rows at the top as an unnamed section until
the parent wrote again. Trim to the oldest own row kept instead; it still
fires only once an own row passes the limit, so a paged-in run of subagent
rows at the head stays until then. With no subagent rows nothing changes.

* perf(native-chat): cap the live window at 4,096 rows, bounding each delta's re-derivation

Every live batch re-derives the transcript over every retained row. On the
largest real window (7,374 rows) that cost 7-8 ms a delta on desktop against
0.6 ms at the old 1,024-row window, and held about 26 MB of row content.
4,096 halves both. The most rows any local journal puts between a roster and
its subagent's last row, with the parent inside its own-row limit, is 3,005,
so no observed subagent loses its roster to the lower cap.

* fix(native-chat): a subagent section opens only while its roster is the running scope's live frontier

A section used to open whenever its roster said the subagent was working, anywhere
in the transcript and whether or not the session was running, so a background
subagent's section stayed open and grew mid-transcript while the parent moved on.

It now opens by default only while the session runs and the roster row naming the
subagent is the newest thing the parent produced, user rows aside. Newer parent
output closes it even while the subagent still works; the roster row keeps
showing that live state. A subagent still working is a running scope of its own
for the sections it spawned; a settled one closes its scope. Derived every
render, no latch; the reader's own open or close still wins.

* fix(native-chat): name a subagent's section from a client roster the window never trims

A section took its name and state from a roster row in the loaded window. Once a
burst trimmed that row, or the row sat on an older page, the section fell back to
an unnamed, closed "Subagent" header.

The shared reducer now keeps a roster keyed by agent id, folded from every roster
row and revision the client receives: pages, older pages and live batches,
including revisions of roster rows outside the window, which live batches already
carry. The first roster naming an agent wins and its revisions update it; a
removed roster row drops its entries; it is rebuilt on every page that replaces
the window and bounded to 512 agents. Sections take their name, state and
live-frontier place from it; placement stays under the loaded roster row, else
at the section's first loaded row. Only a subagent no roster ever named stays
unnamed.

* feat(agent-session): a history page names the subagents whose roster row is older than it

A page is a contiguous run of the journal whose older-page cursor is its first
item, so it cannot pull an older roster row in without skipping the rows between.
When a page held a subagent's rows but not the roster row naming it (about 11% of
the moments a reader could open a session on local journals), that subagent drew
as an unnamed "Subagent" header.

History and hydration pages now carry an optional `subagentRoster`: the first
roster entry naming each subagent whose rows are on the page and whose roster row
is not, with the row's id, sequence and revision; bounded to 64 entries and
16 KB. Items and cursor are unchanged. The client seeds its roster from it.

Rule 1 in docs/reference/remote-wire-compatibility.md: an optional field on an
existing frame, no capability gate. An older client ignores it (the released
reducer reads a page with it exactly as one without); against an older host the
field is absent and the section falls back to an unnamed header.

* Revert "fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it"

This reverts commit 2207865.

Its only purpose was to stop a subagent whose roster row an own-row trim had
dropped from showing at the head of the window as an unnamed section. The client
roster now names that section whatever the window holds, so the cut is back at
just after the own row the limit passes. The retention test that pinned the
unnamed-section case now asserts the section at the head keeps its name.

* chore(native-chat): state the retention limits' own reasons, now that no name depends on the window

Own-row retention keeps the conversation a reader sees from being crowded out by
rows drawn as a one-row section on desktop and not at all on mobile; the
every-agent cap bounds memory and each live delta's re-derivation. Neither is
about keeping a roster row loaded any more.

* fix(native-chat): hold the roster fold's draft map where type narrowing can see closure writes

* fix(native-chat): a roster row's newer revision replaces it in the client roster too

A revision that stops naming an agent (the host drops an entry it learns is not a
subagent, or re-keys a provisional one) left the client roster holding the old
entry, often still "working", with nothing to re-derive it. The section then read
as working forever and could auto-open, while a fresh read of the same journal
left it unnamed. The fold now drops an entry when a newer revision of the row that
named it no longer does, before any roster takes it over.

* fix(native-chat): a parent's spawn and wait calls keep the subagent they name open

A subagent section auto-opened only while its roster row was the running session's
newest row, so any later row closed it: a Codex wait on the agent, or the parent's
text before its next spawn call. Now a row that is part of delegating to a subagent
keeps that subagent open:

- a Codex collab call (spawn, wait, resume, message, close) opens each agent its
  receiver thread ids name; one naming none is ordinary output;
- a Claude spawn call names no agent, so it counts toward the roster announcing it;
- a roster row at the frontier opens its most recently added agent, not all of them.

A roster or call naming only agents one subagent spawned is that subagent's output,
so a grandchild's roster, which the host journals as the session's row, no longer
closes the spawner's section.

* fix(native-chat): a parent's call right after the roster closes its subagent's section

A parent's tool calls after a roster row fold into the tool run drawn above
the roster, so the roster stayed the newest drawn row and its section stayed
open while the parent was already reading or running commands. The fold now
records the newest journal position among the rows it merged, and the live
frontier orders rows by that newest part. The layout is unchanged. A spawn
call folded there still counts as part of the roster announcing it.

* fix(native-chat): a Codex call naming several subagents delegates to the first

A Codex collab call that names several agents opened every one of their
sections. It now counts as delegating to the first agent it names, so one
section opens, the same as a call naming one agent.

* fix(native-chat): closing a roster's list closes the sections under it

Collapsing a roster row's list of subagents left their open sections drawn,
so the section's own head became the only way to close them. And the list's
open state lived in the row, so a row the window unmounted came back
collapsed.

The transcript now holds each roster list's open state beside the section
choices. A closed list hides every section it anchors; each section keeps
its own open or closed choice for when the list reopens. With no choice from
the reader, a list is open while a section under it is open. Closing a
section from its entry keeps the list open, and revealing a subagent's edit
opens the list it sits under.

* perf(native-chat): a reveal finds the roster lists it opens with one set lookup per entry

* fix(native-chat): a subagent's roster entry heads its own rows

An open section drew the agent's name twice: its entry in the roster's list,
then a separate section head above its rows. The entry is now the head. The
roster row draws its entries through the first open one, that agent's rows
follow, then the entries after it, each run in its own windowed slot. A
section no loaded roster row holds (an older page, a grandchild, an unnamed
agent) keeps its own head.

A roster list is open while the live frontier or a reader's choice is on
one of its agents, unless the reader closed the list, so closing an agent
from its entry no longer needs to pin the list open.

The section emitter moves to its own module, and the trailing-run
predicates it shares with the slot builder to theirs, to keep the slot
builder under its line limit.

* fix(native-chat): the entries after an open subagent's rows set in its roster's type

The roster row's list inherits the system row's small muted type; the entries that
follow an open section sit outside that row, so they now carry the same type.

* docs(native-chat): a current host can also serve a page of only a subagent's rows

A page is bounded by bytes after it is windowed by the session's own rows, so a
burst that fills the bound yields a page, or an opening page, with none of the
session's own rows. Mobile's read-on and read-back therefore serve current hosts
too, not only older ones; the comments said otherwise. The retention comment
still described a closed section as a row of its own; it now sits behind its
roster entry.

* fix(native-chat): a section's prose keeps its copy/timestamp controls inside the section

An assistant row's hover controls (copy, scroll-to-top, timestamp) hang 20px
below the row into the gap before the next one (`-mb-5`). Inside a subagent's
section that put them below the section's left border, and on the section's
last row they touched the parent's next row with no gap.

Inside a section the controls now stay in flow, so the border covers them and
the next row sits the normal gap below. The row-height estimate reserves the
same 20px for a section's prose so windowing does not jump on measure.

* test(agent-session): state each appended row's turn scope, as the journal now requires

* refactor(native-chat): the client's journal retention policy lives in its own module
…er runs journaled (stablyai#23758)

* fix(claude): a subagent resumed after a restart keeps its canonical id

A backgrounded Claude subagent resumed by a message after its session's provider
restarted parents its frames to its ORIGINAL spawn call, while the announcement the
new provider sees names only the resuming call. The spawn call's alias lived only
in the old provider's memory, so every row the resumed child wrote fell through to
its raw call id: one child shown as two, a named roster entry with no rows and an
unnamed section holding them.

The alias table now also recalls what an earlier run of the session resolved,
re-derived from the agent rows it journaled (canonical id beside the call its frames
arrived under), read once per bound journal epoch. Nothing new is persisted.

* fix(claude): a restarted provider continues the subagent roster earlier runs journaled

The roster's state lived in one provider process while everything it writes is
the session's. After a restart a resumed child was re-rostered in a second group
row with its attempts restarted, its resumed frames lost the alias they still
carry, and the one group row no turn owns was rewritten from empty, erasing the
children an earlier run had listed there.

A run now reads what earlier runs journaled, once per bound journal epoch: group
rows give each child's entry and group, agent rows give the spawn-call aliases
and the latest attempt. A group is inherited only when this run's own events
reach it, and inheriting writes nothing. An inherited child is reopened by an
announcement exactly as an in-process resume reopens it, and by nothing else.
The alias recall is one facet of that read. Nothing new is persisted.

* fix(claude): an earlier run's subagent takes Claude's restart verdict in one row

At restart Claude reports each agent the previous session left running as
stopped ("didn't finish before the previous session ended"). The roster ignored
it, leaving the entry unverifiable, while the background-task lane, which never
saw that agent announced, opened a second, unnamed row for the same agent.

The roster now records the verdict on the inherited entry, keeping the time the
earlier run lost contact, and still reopens the entry on an announcement. The
background-task lane leaves any agent an earlier run rostered to the roster,
from the same journal-derived reading.

* fix(claude): a restart verdict on a child whose host died invents no stop time

A journal reopened after its host died leaves a child unverifiable with no stop
time. Stamping Claude's later restart verdict with the current time would show
the whole outage as how long the child ran, so the earlier run's stamp, or its
absence, is kept. The journal-liveness comment no longer describes the roster as
unable to continue from the journal.

* fix(claude): a child two journaled rows list stays live in one of them

An older build could re-roster a resumed child in a later turn's row, so two
group rows list it (67 children in 39 local sessions). Inheriting the second
row re-pointed the child to that row's copy, so the copy the resume had
reopened was left at working and swept to unverifiable, while the outcome
landed on the stale copy. The first row reached keeps the child.

* fix(native-chat): no run length for a settled child with no stop time

A subagent whose host died without sweeping it has no stop time. After a
restart, Claude's own verdict on it ("stopped") is now recorded, but the rule
that hides the group's duration only covered `unverifiable`, so a mixed group
showed the sibling's duration as the whole group's. The rule now keys on the
missing stop time, whatever the settled state.

* fix(claude): a twice-listed child resumes in the row the journal reading chose

A child an older build listed in two rows was placed in whichever row this
run's frames reached first, so a sibling's frame reaching the older row made
the resumed child reopen there. Placement now follows the reading's one
tie-break, and evicting a row drops only placements that point at it.

* fix(native-chat): a resumed subagent's clock never counts the idle gap

A reopened Claude child now starts a new run: its startedAt is reset to the
reopening frame's time on every reopen path (resume after a restart and a
same-process reactivation). The group clock reads the union of the children's
latest runs, so an idle gap between runs is never counted; an ordinary
overlapping fan-out reads the same as before.
…i#23768)

* fix(editor): keep saved untitled notes when their tab closes

Saving cleared the draft and dirty flag, so close-time cleanup took a saved
note for an untouched placeholder and deleted it.

Fixes stablyai#23688

* fix(editor): decide untitled-note cleanup from disk state, not a per-save flag

Drop setUntitledFileHasSavedContent and its save-queue hook so
deleteUntouchedOnClose keeps its creation-time template meaning. A clean
untitled tab now counts as an empty placeholder only while its last-known
disk content is empty or not yet loaded; the on-disk size check still gates
every actual delete. Notes an agent filled while their tab was showing now
stay reopenable with Cmd+Shift+T.

Also test both Don't Save dialogs end to end, mirror the shared discard in
the floating panel test mock, and remove markFileDirty selectors left
without consumers.

* test(editor): wait for untitled cleanup before asserting disk state

Drain the stat-then-delete chain before the untitled-note tests assert the
file survived, so the in-flight Don't Save case fails if the size check is
removed. Type the main-window close dialog hook by the controller fields it
reads, which lets its test drop a type assertion.

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
…tablyai#24010)

* fix: keep active tab visible by docking to viewport edges

Makes the current tab easier to locate in many-tab scenarios. The active tab
now sticks to a viewport edge via sticky positioning when it would scroll out
of view, with a full-foreground indicator bar for better visibility and arrow
animation when a background tab opens off-screen.

* fix(tab-bar): reveal offscreen tabs instead of nudge animation

When a background tab opens beyond the visible area, automatically scroll
to reveal it (unless hovering the tab strip). This replaces the previous
arrow-nudge animation with direct visibility. revealTabStripElement now
handles keeping the active tab visible alongside the revealed tab when
both fit, or docks the active tab when needed.

* fix(tab-bar): reveal tabs by identity, not count increase alone

Detect opened tabs by comparing tab identities independently of count
changes. Newly opened tabs are now revealed even when the total tab count
stays the same—e.g., when a tab closes as another opens.

* fix(tab-bar): track tabs by identity for reliable reveal on open/close

Replace count-based tab detection with identity tracking so the strip
correctly reveals tabs when they're added, replaced, or when the active
tab closes and switches to a far-back history tab. Removes the tabCount
parameter and simplifies overflow navigation by using identity sets.

* fix(tab-bar): defer revealing tabs until pointer leaves

When a background tab opens while the pointer hovers the tab strip, defer
its reveal until the pointer leaves. This prevents the active tab from
sliding away mid-interaction. Also support client-hosted rows taking
active state while maintaining tab dock positioning.
…ai#24040)

* fix(push): size the claim-attempt budget from the drain count

With twelve drains, up to eleven peers can hold device heads, so a
four-attempt claim budget can run out while claimable rows remain and the
drain exits idle for a tick. Move the drain count into one module and derive
the attempt budget from it.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* style(push): keep the worker and store in repo formatting

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* style(push): drop the stray semicolon in the concurrency constant

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
… settings (stablyai#23979)

* fix(agent-trust): write per-user trust under the home the launched agent reads

The Cursor, Copilot, Qoder and Antigravity writers and the local Codex config list
resolved ~ with os.homedir() at write time, so any test that reached them wrote into
the developer's real ~/.codex, ~/.cursor, ~/.copilot or ~/.gemini. Each writer now
takes the home, derived once from the launch env (HOME, or USERPROFILE on Windows,
else this host's home) by launchedAgentHome, which the SSH relay already used.

* test: give tests that wrote the real agent or Orca home a temp one

The structured Codex adoption replay pre-trusted /repos/workspace-1 in the real
~/.codex/config.toml; it now runs with a temp HOME and userData. The Codex
session-resume and WSL hook tests created Orca's managed Codex home under the live
userData, and the Claude Agent Teams tests wrote their tmux shim into ~/.orca; each
now runs against a temp userData or HOME.

* test: fail any unit test that writes the real agent or Orca home

A vitest setup file wraps the node:fs mutating calls and refuses a target under the
account's real ~/.codex, ~/.claude(.json), ~/.orca, ~/.cursor, ~/.copilot, ~/.gemini,
~/.qoder or Orca userData, found through os.userInfo() so a test that swaps HOME
cannot hide it. The refusal is recorded and rethrown after the test, since trust
writers swallow errors. Reads are untouched. It stands down only while an opted-in
real-agent suite's own switch is set. It also unsets what an Orca terminal exports
toward the live app (userData, Codex and Claude homes, and the Codex launch preflight
CLI, which a shell test would otherwise run), so a local run matches CI.

* test: type the guarded fs call from its narrowed original
…tablyai#24043)

Completes the `src/main` audit. Sweep over `src/main/runtime` (flat, rpc,
orchestration, relay, push) and flat `src/main`. 42 case declarations removed
across 24 files, 783 lines gone. No file deleted whole, no production code touched.

Highest-yield area by far was `runtime/orchestration` (21 cases from 151 files);
the rest of runtime measured under 1%.

What went, by pattern:

- Tests whose subject is the test itself: a source-scanning boundary test with no
  production import at all, three of whose cases checked its own regex against
  strings it declares; and a benchmark whose own simulation contains the
  short-circuit it asserts, with `expect(wouldHaveBeen).toBe(60000)` comparing the
  test's own arithmetic.
- Identity copiers: seven db cases where every asserted value is the literal input
  (`type: 'question'` in, `type === 'question'` out).
- Permanently skipped tests for behavior that does not exist — two cases carrying
  `// TODO: inline restore on re-subscribe not yet implemented`. A skipped test for
  an unimplemented feature can never fail; it is a note in test syntax.
- Duplicate invocations of a contract owned at a stronger boundary, including three
  reset scopes and two dependency-promotion cases owned by dedicated suites.
- Registration manifests whose every entry is referenced by other tests.
- Table rows and cases varying a field production never reads: a mobile tab-restore
  case varying `clientCapabilities`, which the mobile branch does not consult, and
  a Windows worker case in a module with no platform input at all.
- Names promising more than the input exercises: a case titled for forged AppImage
  variables whose body only removes `AppRun`, byte-identical to a row of the
  `it.each` table twenty lines above.
- Negative controls passing for an unrelated reason: a foreign-pane rejection whose
  fixture also differs in sender handle, so the handle guard can reject it.
- Self-comparisons, including `format(m, { authority: 'current' }) === format(m)`.

Two deletions were justified by the wrong argument and kept only after checking a
better one. "A generated-catalog check gates registration" is false: that script
prevents the catalog and dispatcher from drifting apart, so a developer who removes
a method regenerates the catalog and the check passes. Like a type annotation over
an interface and its implementing class, it verifies internal consistency and
cannot see a declaration and its use removed together. Both inventories go on
per-entry evidence instead — every name is referenced by other tests.

The same rule kept a 37-entry terminal-method inventory in the same wave, because
some of its entries are pinned solely by it. An inventory is a ratchet if and only
if at least one entry is pinned solely by it; that is evidence per entry, not a
verdict by shape.

Kept deliberately: everything a gate cites, checked by case title and not only by
file path; source-scanning ratchets that pair their negative assertion with a
positive one against real source (the surviving boundary case asserts the pattern
still matches the writer module, so a silently-broken regex goes red); a
destructive-delete PTY waiver guard; `it.fails` markers, which go red if the bug is
fixed; and `it.skipIf(platform)` cases, which do run on other hosts.

Coverage is partial and stated as such: 1,195 files in scope, roughly 660 read
case-by-case, the remainder title-scanned and mechanically triaged. Unread paths are
recorded for a later sweep rather than assumed clean.

Verified: `check-reliability-gates.mjs` (140 gates), `check:code-quality:changed`
(0 new findings). No `pnpm tc` needed since no production file changed. A local run
of `src/main` surfaced 42 failures in three files this change does not touch
(`browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers`); all 42 reproduce on a pristine `origin/main`
worktree, so they predate this wave and are environment-dependent locally — CI was
green on main at `2d85fdc753e2`.
…pair Orca's duplicates (stablyai#22592) (stablyai#23958)

* fix(codex): recognise Codex's quoted project-trust spellings in config.toml (stablyai#22592)

Codex's settings screen writes project trust as ["projects"."/p"] and
"trust_level" = "trusted". Orca's matchers only knew the bare spelling, so a
trust write appended a second [projects."/p"] table (or a second trust_level
line) and every codex command then failed with "duplicate key". The config
mirror kept both spellings in Orca-managed homes for the same reason.

- Project table headers are now read through the existing TOML key-path
  parser, so bare, quoted, literal-quoted, mixed and spaced spellings are the
  same table for trust writes and the managed-home mirror/dedupe.
- trust_level is found by decoded key, in both the trust writer and the
  mirror's trust reader, and an existing key is rewritten, never duplicated.
- On the next trust write, a table older Orca appended (exactly
  [projects."<p>"] holding only trust_level = "trusted") that duplicates the
  user's table, or the bare line it inserted under a quoted "trust_level", is
  removed; the user's table wins and the atomic writer keeps config.toml.bak.
  Any other duplicate, or a repair that would still leave one, leaves the file
  untouched and logs once.

* build(cli): list the new Codex trust modules in the CLI project

* fix(codex): recognise Codex's quoted hooks.state spellings and repair Orca's copies (stablyai#22592)

Codex writes hook trust as ["hooks"."state"."<key>"] (and the parent as
["hooks"."state"]). Orca's hook-trust writer, parent-table check and mirror
only knew the bare spelling, so a hook-trust write appended a bare copy and
the file failed to parse with "Cannot declare ... twice".

- The hooks.state header, parent-table and mirror checks now use the TOML
  key-path parser, like project tables.
- The duplicate repair now also removes Orca's own hooks.state tables (an
  exact [hooks.state."<k>"] with only enabled + trusted_hash, or an empty
  [hooks.state]) that repeat a table in another spelling, and runs on hook
  trust writes too, so a file with both project and hook duplicates is fully
  repaired. The Orca-shaped copy is removed whichever order the two tables are
  in, only when exactly one other table (the user's) remains; anything else is
  left untouched and logged once.

* fix(codex): carry plain-Codex plugin and project hook trust into Orca's Codex homes (stablyai#22592)

Codex keeps hook trust in $CODEX_HOME/config.toml under hooks.state, keyed
by the hook's source. Plugin keys (`id@mkt:path`) and project keys
(`<repo>/.codex/...`) are the same in every home, but the mirror dropped
every hooks.state table from ~/.codex, so Codex inside Orca asked users to
re-trust plugin and project hooks they had already trusted in plain Codex.

- classifyHookTrustKey splits keys into home-scoped (the home's own
  hooks.json/config.toml, re-keyed by install as before) and shared.
- The mirror now carries shared hook trust from ~/.codex in every spelling.
  A key the managed home already holds keeps the managed copy, a key
  repeated in ~/.codex is carried once, and the parent [hooks.state] table
  is never copied, so the result never declares a table twice.
- mergeSystemCodexConfigIntoRuntime moves to codex-config-mirror-merge.ts
  to keep codex-config-mirror.ts under the line limit.
- Tests cover plugin/project carry in each spelling, user-hook keys staying
  out, repeated launches, managed-copy precedence, Windows key spellings,
  parent tables, and user-hook trust re-keying (trusted_hash and enabled)
  from every ~/.codex spelling.

* fix(codex): carry session_end and interrupt hook trust into Orca's Codex homes (stablyai#22592)

The shared-trust classifier parsed hook keys with Orca's own trust-key
parser, which only knows the ten events Orca installs hooks for. Keys for
Codex's session_end and interrupt events did not parse, so their plugin and
project trust was treated as home-scoped and left out of the managed home.

The classifier now reads the source path from Codex's key shape
`{source}:{event}:{group}:{handler}` for any event label. A key without
that shape is still never carried. Tests cover both events for plugin and
project keys in both spellings, user-layer keys for both events, and five
unattributable key shapes.
Group 10 is full; swap the README QR code and copy (all locales) to the new group 11 invite.
…tle (stablyai#23948)

* fix(sidebar): keep hook-less agent rows while the agent runs, whatever its title

Codex retitles its pane to the project name, so the sidebar's title-derived
row (which required the title to name an agent) vanished while Codex kept
running (stablyai#23767). Rows now take identity from the canonical pane resolver
over the pane's foreground-process read and launch record, then the title;
the title only decides idle/working/needs-input. The row still goes away when
the PTY exits, the process tracker proves the shell is back, or the title is a
shell or default title.

* test(dashboard): justify the partial store fixture's type assertion

* fix(sidebar): only a live process read keeps a plain-title agent row

Review of the previous commit found ghost rows: the tab launch record is a
latch nothing clears on WSL, after an SSH exit, or for a launch that never
started, and a parked pane's process read went stale because only the
mounted tracker re-derives it.

- The launch record returns to main's role: a fallback only for titles that
  show activity, ranked below a title naming another agent (pane reuse),
  matching the tab icon's order.
- A parked pane's command boundary retires its unconfirmable process read,
  like the mounted ladder's unavailable path; reveal re-reads it.

* fix(sidebar): confirm before a parked marker retires an agent; read Git Bash prompt titles as the shell

- A parked pane's end-of-command marker can be a nested shell's leak under a
  still-running full-screen agent, so confirm the foreground first (as the
  mounted ladder does) and retire the process read only on a shell or no
  answer. SSH/remote parked panes hold no incarnation to fence a host read
  with, so they still retire.
- Git Bash emits no command marks; its `$MSYSTEM:$PWD` prompt title
  (MINGW64:/c/repo) is now shell evidence, so a stale Codex read there no
  longer keeps a ghost row after Codex exits.

* fix(sidebar): trust only process-read agents for plain-title rows; per-worktree foreground selector groups by tab

A daemon reattach seeds the pane's foreground entry with its launch agent, which
can outlive the process while Orca is closed. The entry now records where its
agent came from (agentEvidence), and the sidebar/dashboard title-derived rows
only keep a plain-title row on an actual process read. Routing and the tab icon
are unchanged.

selectPaneForegroundAgentsForWorktree grouped every pane key per worktree; it now
groups by tab once per map identity and skips worktrees with no tabs.

* fix(sidebar): a parked pane's reattach keeps its own process read of the same agent

The reattach seed marked a returning parked Codex pane as launch-record
evidence, over the process read this session already took, so its row
blinked out on reveal and stayed hidden if the user left the tab before
the visible read landed. Keep the read when it names the same agent; the
seed still drops byte-routing trust.

* test(terminal): foreground confirmation publishes process-read evidence

* fix(sidebar): a cleared pane title retires the agent's process read

Codex clears its title when it exits, and the tab then shows its default
title. A pane without shell command marks never re-reads its foreground
process, so the retained read kept a "Codex · Idle" row after /quit
(permanently for a hand-typed Codex; about 15 s while the marked-pane
confirm ladder ran). Treat a blank title like the default title it shows.

* fix(sidebar): the pane's process monitor retires an exited agent's process read

A hook-less pane keeps its sidebar row from the tracker's foreground-process
read, but nothing re-derived that read in a pane without OSC 133 command
marks. After Codex exited there, a "Codex · Idle" row stayed: permanently
when the shell titles its prompt, or when a killed Codex leaves its last
title.

The pane's agent-completion process monitor already confirms an agent's
exit (no agent and no child processes, held past its settle window). It now
reports that exit to the tracker, which retires its own process read and
runs the confirmed-shell path the visible-pty read uses. A tracker read that
names an agent seeds the monitor, so hidden panes and panes the monitor had
not polled yet are watched too. A command read in flight still decides the
pane, and launch records or other agents' reads are left alone.

* fix(sidebar): a monitor-confirmed exit leaves the next agent in an unmarked pane identifiable

The process-exit retire published shellForeground:true and left the one-shot
visible sample settled; a pane without command marks has no command start to
lift either, so a Codex typed again after quitting was never read and lost its
row on retitle. Publish shellForeground:false and reopen the sample.

* test(terminal): justify the pane binding cast in the process-exit relaunch test
… agent.launch (stablyai#22954)

* fix(mobile): start + menu, quick command and diff-note agents through agent.launch

The session screen's + menu, agent quick commands and diff notes' New agent
session now ask the host to start the agent with agent.launchReplay, so the
host picks chat or terminal from the desktop's default and delivers any
prompt. Hosts without the launch capabilities keep today's paths.

The phone's pending tab choice is one value (a tab, a terminal by handle, or a
launched surface) instead of two refs, and a launched chat is found by its
session id in the next snapshot rather than a predicted tab id. A launched
surface waits a bounded number of snapshots for its tab.

* test(mobile): add the + menu and diff-note launch scenarios to the recording corpus

* test(mobile): repin bridged-parity tallies for the four launch goldens; drop test casts

The corpus grows from 790 to 794 goldens; all four new ones replay identically.

* fix(mobile): show a refused agent launch as a toast beside open tabs

The inline create error renders only in an empty session, so a host refusal
(for example a disabled agent) from the + menu in a session with tabs showed
nothing. Always toast the failure: the caller's own copy when it gave one,
otherwise the host's reason.

* test(mobile): type the launch reply helper with the shared launch outcome types

* fix(mobile): record a launched agent's tab as this device's pick on the host

A launch carries no navigation, so the phone selected the new tab only
locally while the host kept this device on the tab it had before. Leaving
the session and coming back, or a reconnect that reset the screen, reopened
that old tab. The "+" terminal path this replaced asked the host to select
the tab for the caller.

When a launched surface's tab lands in a snapshot, activate it for the
caller exactly as a tap does. The resolver now names the landed tab in
place of the unused `missed` flag. Route parity re-pinned for the new
activation body, identity payload and strings.

* fix(mobile): land on a launched agent's tab without a 500 ms wait or a blank pane

The host publishes a launched tab before it replies, so the tab list the
phone already holds usually has it by the time the reply arrives. The
launch paths still waited for a refetch 500 ms later, leaving the phone on
the old tab for that long after every launch. Read the tab list at once.

On hosts without agent.launch, the chat path also unsubscribed the open
terminal and cleared its handle before the chat's tab landed, while the old
terminal tab stayed selected: a blank pane until the next tab list. Leave the
open tab live until the chat lands, as the launch path does; applying that
tab list tears the old terminal down.

Route parity re-pinned for the two bodies.

* fix(mobile): keep a tab the user picked while a prompted launch was still replying

A quick command or review-notes launch now waits for the host to deliver the prompt, which can take up to a minute. The launched tab shows up in the tab row well before that, so a user who tapped another tab meanwhile was pulled back onto the launched one when the reply arrived, and that pick was recorded on the host.

The launch now remembers which tab the phone was on when it started and only takes focus if the phone is still there when the reply lands. Any move made in between, by a tap or by the computer navigating this phone, wins. Session route parity re-pinned for the handleCreateTerminal body only.

* fix(mobile): name a launched agent's tab before asking, and land on it when it is listed

A "+" menu, quick-command or review-notes launch now reserves its tab before it asks the host: a fresh pane key (tab and leaf UUIDs) and, for an agent the host may start as a chat, a session id. Both are minted once per launch and sent unchanged on every replay, since the host's replay fingerprint covers them.

The phone arms its pending selection with that reservation before sending, so it lands on the terminal (matched by pane halves) or chat (matched by session id) as soon as the tab is listed. For an agent whose prompt is pasted after start, that is long before the reply, which waits for delivery. Landing also frees the "+" lock; the lock holds the create's id, so an older launch's reply cannot free a newer one's. The reply now only adds its own handle or session id (an older host ignores the reservation), starts the fallback countdown, and reports prompt delivery.

A tab the user picks mid-launch replaces the pending selection, so the launch-start tab check is gone. A reservation the host refuses as already taken reads "Couldn't start the agent. Try again." on the first send, and as unconfirmed after a replay. The mobile UUID fallback now yields a v4 UUID, because a pane key's leaf must be one. The host launch path moved to new-tab-agent-host-launch.ts; session route parity re-pinned for that move and the landing's lock release.

* test(mobile): expect the launch reservation in the four launch scenarios

The four launch scenarios now expect the pane key and session id the phone sends (the scripted ids come first, so the operation id moves from ...001 to ...004).

* fix(mobile): don't say an agent may not have started while the user is looking at it

When a launch's reply was lost after its tab had already landed, the phone said "Couldn't confirm the agent started", although the listed tab proves it did. Now a listed tab narrows the doubt to the prompt or notes ("The agent started, but couldn't confirm the notes were sent."), the notes stay unsent, and a bare launch says nothing. Only the nested-function parity pin moves, for handleCreateTerminal passing the tab list to the launch.

* test(mobile): read the launch's sent reservation through the host's params schema

The anti-slop audit rejects Reflect.get; parsing with AgentLaunchReplay also
asserts the host accepts the params the phone sent.

* test(mobile): check the launch reservation against the host without importing its schema

Mobile code may import the params contract only as types. The phone's tests
now read the sent reservation by narrowing, a chat reservation is checked
through the real host dispatcher, and the older-host drop is pinned host-side.

* fix(mobile): don't send the same review notes to a second new agent

The "+" lock is now freed when the launched tab lands, but review notes are
only cleared when the launch's reply confirms delivery, which for a prompted
launch can take up to a minute. In that window "Send review notes to AI" still
offered the same notes, and choosing a new agent session started a second
agent with them.

The notes a new agent session is being started with are now held from the tap
until that launch settles: the Send button no longer counts them, the sheet no
longer offers them, and a stale tap on the old sheet starts nothing. Notes the
host did not deliver become sendable again once the reply arrives.

* test(mobile): record the + menu and diff-note launch goldens

4 added (+ as a terminal, + as a chat, notes delivered, notes not delivered). 10 existing create-terminal goldens move only because the recorded state now shows one pending selection instead of two refs; their requests are unchanged.
…ield (stablyai#24052)

stablyai#24010's spec read them with Reflect.get, which the low-evidence lint rejects, so every PR's static analysis now fails on main.
* fix(chat): separate text paste from attachments and route by pane

Keep composer text independent of image checks and saving, and route pastes
caught underneath chat to the originating pane's mounted input. Preserve
native event data, selection replacement, undo, and target lifetime checks.

Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: lurunzi <lurunzi@gmail.com>

* fix(chat): keep focus and quiet text paste after routing it to chat

- A paste inserted into the composer now moves focus there, as the old
  menu-paste insert did; otherwise a paste routed from the hidden terminal
  left the next keystrokes going to that terminal.
- With a remote-server or not-ready workspace, pasted text no longer shows
  the "Local attachments are not available" refusal because the clipboard
  also held an image rendition (common for Office copies). The menu path
  probes for an image only when the text read is empty, so a paired browser
  does one permission-gated clipboard read for a text paste, not two.
- Latest-value refs update in a layout effect instead of during render.

* perf(clipboard): answer "is there an image?" from the format list

The chat composer asks the main process whether the clipboard holds an
image before explaining an image-only paste on a remote-server workspace.
That probe decoded the whole image (readImage().isEmpty()) on the main
thread just to return a boolean. Read clipboard.availableFormats() instead,
and share the MIME check with the paired-web probe.

* fix(chat): a chat cover owns focus, so input never reaches the hidden terminal

When a Chat UI tab opened over its terminal, the terminal's xterm kept
keyboard focus until the composer claimed it a frame later, and forever if
the composer never became ready (still starting, a question card, a phone
holding input). The previous commits rerouted paste from that hidden
terminal to the chat, but typing and Enter still went to the terminal, an
image-only or refused paste left focus there, and about twenty
terminal.focus() call sites could put it back.

Make "a covered terminal cannot hold focus" structural instead:
- The chat cover takes focus in the commit that mounts it and marks the
  covered xterm inert, so every terminal.focus() path is refused by the
  browser. Split siblings are untouched. When the chat goes away the xterm is
  un-inerted, and gets focus back only if focus was inside that chat.
- Terminal paste listeners skip anything inside a chat cover (previously only
  inside a mounted chat root). The reroute from terminal to chat is gone;
  terminal-only paste is back to main's code.
- A paste that finds no chat input (before the chat mounts, or an approval
  card with no text field) gets a visible refusal from the cover. A disabled
  composer shows the same notice inline instead of dropping the paste. New
  copy: "Can't paste — this chat isn't accepting input right now." (the old
  "Worktree not ready" toast was wrong for a chat that is still starting).
- The terminal context menu, which names its pane, keeps a small request
  event to that pane's chat, now without a clipboard payload and using the
  existing covered-pane check.
- Cmd/Ctrl+V or Shift+Insert on a non-input part of the chat focuses the
  composer (or question answer) first, so the paste lands there.
- The composer-scope check used to decide whether a text field inside the
  chat keeps its own paste matched the whole pane (the file-drop surface
  carries the same attribute). It now asks the composer whether the target
  is inside its input.

* fix(chat): don't paste a copied file's name next to the file

Copying a file in Finder or another file manager puts its name on the
clipboard as text/plain beside the file itself. Since text and images are
now pasted independently, pasting such a copy into a local or SSH chat
inserted the file name into the prompt as well as attaching the image.

On the paste-event path, text/plain that is exactly the names of the pasted
files (one per line) is the file's label, not prompt text, so it is dropped
when an image from that paste is being attached. Rich-text copies (text plus
an image rendition) still insert their text, and a copied non-image file,
which is not attached, still pastes its name as before.

* fix(chat): don't type a Finder file's name on Cmd+V either

On macOS, Cmd+V in the chat goes through the app-menu paste, which reads the
clipboard text and saves the clipboard image separately. A file copied in
Finder also puts its name on the clipboard as text, so the composer typed
the name next to the attachment. On main the menu path never read text once
an image saved.

The main process now reports the paths of the files a file manager copied
(macOS filenames plist or file URL, Explorer's FileNameW, a Linux uri-list).
Text that only labels those files waits for the image outcome: dropped when
an image is attached (or refused on a remote owner), typed when none came.
The same label rule now also accepts a path or file URL per line, which is
how Linux file managers label copied files on the paste-event path.

* fix(chat): pane focus aimed at a chat lands on the chat

Since the covered terminal became inert, focusing a pane that shows a chat
(keyboard pane navigation, focus-follows-mouse, split activation) was refused
and focus stayed on the pane the user left, so typing went to that visible
sibling terminal. The one place a pane's focus is requested now puts it on the
pane's chat cover, which hands it to the composer when the pane is revealed.
Focus already inside the chat is left alone.

* test(terminal): give fake panes the container pane focus now reads

Pane focus checks the pane's container for a chat cover, and these two
fixtures built panes with only a terminal, so four tests threw.

* refactor(native-chat): move composer paste handle and chat-root key routing into their own modules

Brings NativeChatComposer.tsx and NativeChatResolvedView.tsx back under the
400-line limit after merging main. No behavior change.

---------

Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: lurunzi <lurunzi@gmail.com>
…nty (stablyai#23788)

* fix(chat): decode Claude pastes and track terminal delivery uncertainty

Keep queued prompts pending while the existing agent status reports work, and check fresh history after a later idle fact. Preserve draft text and distinguish write rejection from unconfirmed delivery.

Co-authored-by: Harshul Rathod <harshulrathod1640@gmail.com>

* fix(native-chat): break the observed-send import cycle and keep renderer tests out of main

The observed-send path imported the clear helpers from native-chat-runtime-send,
which imports it back. Move the input-clear layer into its own module both use.
The Claude paste decoder test imported the renderer pending module from
src/main, which the node typecheck project cannot see; the echo-retirement
assertions now live in a renderer test.

* fix(native-chat): still submit a Claude chat send whose write acknowledgment was lost

A remote write whose acknowledgment is lost (timeout, dropped link) is not a
refusal, but the observed path stopped there and never sent Enter, leaving a
body that did land sitting unsubmitted in Claude's input line until the next
send's clear wiped it. Continue to the next write without re-sending the bytes,
as the unobserved path always did.

* perf(native-chat): keep terminal Chat pending delivery from re-rendering every row

The delivery notices were merged into a new Map on every render, which
invalidated the transcript row context and re-rendered every memoized row on
each stream update. The pending hook also wrote a fresh array on every status
ping and prune pass even when nothing changed, and the phone mapped its pending
list on every render, rebuilding the chat list data. Memoize the merged
notices, skip no-op pending writes, and memoize the phone's rendered pending
list. The phone also skips a transcript read when no send is due.

* fix(native-chat): never flag a queued Claude send, and flag one an idle Claude never starts

Two gaps in when terminal Chat calls a Claude send "Delivery unconfirmed":

A prompt sent while Claude is mid-turn is queued, and Claude folds it into the
running turn as a queued-command record. The transcript reader drops those
records, so once the turn ended the prompt Claude did run read as unconfirmed,
inviting a duplicate resend. A send made while the agent is busy is now never
checked; it keeps the pending behaviour it had before.

A prompt sent to an idle Claude that never starts a turn (Claude exited to the
shell, or the paste went nowhere) left the status at the same idle fact
forever, so the check never ran and the bubble stayed pending. An idle agent
starts a turn on a delivered prompt at once, so a send whose idle status is
unchanged after the existing 20 s bound is now checked against a fresh
transcript read.

* fix(native-chat): add the delivery notice strings to the English catalog

The Dismiss action's translate key was missing from en.json, which fails the
localization catalog and extraction gates. The desktop "Message not sent" and
"Delivery unconfirmed" notices were hard-coded English; route them through
translate with the same wording.

* fix(mobile): sync the held-send refs after commit instead of during render

Moving the acknowledgment-loss hold into its own hook made its render-time ref
writes new lines, which the React Doctor changed-lines gate blocks. Held sends
report after commit, so syncing those refs in a layout effect keeps them
current where they are read.

* fix(native-chat): report only definite terminal Chat send outcomes

The delivery rule inferred "Delivery unconfirmed" from "the turn ended and the
transcript has no matching row". Claude records a prompt sent mid-turn only as a
queued-command attachment, which the transcript reader drops, so that rule
flagged prompts Claude had answered. It also never fired for an idle Claude that
lost the write, because no newer turn arrives.

Keep only facts the transport reports:
- a refused write reads "Message not sent", keeps its text, and can be dismissed;
- a lost write acknowledgment holds the echo for 20 s, the phone's existing
  rule, then reads "Delivery unconfirmed" unless its row has landed.
An ordinary send, including one Claude queues mid-turn, stays pending as before.

Remove the agent-status subscription, the status-epoch origin, the fresh
500-row transcript read, the confirmed state and the no-status clock. The phone
already implements this rule, so its changes revert to main; only a test for
old-host paste envelopes remains.

* fix(i18n): translate the terminal Chat delivery notices

Add the Dismiss, "Message not sent" and "Delivery unconfirmed" strings to the
es, fr, ja, ko and zh catalogs, reusing each catalog's existing Dismiss wording.

* fix(native-chat): let a resend replace its failed terminal Chat echo

A "Message not sent" or "Delivery unconfirmed" echo kept its transcript
occurrence, so resending the same text numbered the resend as the second
copy: the one landed row retired the failed echo and pinned the resend
below the reply forever. Appending a send now drops a failed echo with the
same content first.

* fix(native-chat): unwrap a Claude paste that quotes pasted_content tags

The envelope parser refused any body containing a pasted_content tag, so a
pasted prompt that itself quotes one (a transcript excerpt, or code that
handles these tags) kept its wrapper and its echo stayed pinned below the
reply. Claude's per-paste id exists to disambiguate exactly that; only a
same-id tag inside the body is now ambiguous. Wrappers without an id keep
the strict rule.

* test(native-chat): pin which terminal Chat sends observe write outcomes

Only a Claude chat send (text or images) reports a refused or unacknowledged
write to its pending echo; other agents and slash commands keep the
unobserved write path exactly as before.

* fix(native-chat): keep failed terminal Chat sends through Stop

Stop cleared every optimistic echo, including a "Message not sent" or
"Delivery unconfirmed" bubble whose send had already settled. Stop cannot
affect that send, and the bubble is the only place its text stays copyable,
so it now survives until the user dismisses or resends it. Also moves the
observed-send import below the file header comment.

---------

Co-authored-by: Harshul Rathod <harshulrathod1640@gmail.com>
…lyai#23806)

* fix(rate-limits): stop driving a hidden Codex TUI to read usage

When the headless Codex usage call failed, Orca opened a hidden
interactive Codex, typed /status and pressed Enter without reading the
screen, then killed it after 15 s. If Codex showed its "Update available"
prompt, that Enter picked "Update now", Codex started its installer, and
the 15 s kill interrupted it, leaving the global install broken.

Drop the hidden-terminal fallback. When the headless call fails with a
non-sign-in error, read the same usage from the HTTP endpoint Orca
already calls on WSL and for the 5-hour window, so a usage check can no
longer answer any Codex startup screen.

Fixes stablyai#17415

* test(rate-limits): pin the RPC error when the Codex HTTP fallback fails

Also drop comments that still described the removed hidden-terminal fallback.

* test(rate-limits): use real Response objects in Codex fetcher tests

The changed-code gate rejects the new type assertions this PR added.
… so phone swipes never type escape text (stablyai#23943)

* fix(mobile): preserve host mouse modes for terminal scrolling

* revert(terminal): drop the out-of-band mouse-modes channel

The earlier commit sent the host's mouse tracking and encoding beside
each snapshot and stream frame, and made the phone trust that over what
its replayed bytes say. Where the snapshot text already carries the
encoding this is redundant, and where the text is wrong (a desktop pane
snapshot never writes ?1006h) the host's own mirror is seeded from that
same text, so the side channel asserts the wrong encoding too.

Remove the shared type, the host publication, the wire fields and the
phone's host-modes authority. The phone keeps its guard: while mouse
tracking is on but no replayed byte proved the encoding, a wheel scrolls
locally instead of sending a guessed legacy report.

* fix(terminal): carry the mouse encoding in desktop pane snapshots

Swiping a phone terminal running Codex typed `[M`-style mouse bytes into
Codex's prompt instead of scrolling (stablyai#23818). Codex turns on mouse
tracking with the SGR encoding (?1006h). A desktop pane snapshot is made
by xterm's serialize addon, which writes the tracking modes but never the
encoding, so the phone replayed "tracking on, legacy encoding" and sent
legacy `ESC[M` reports that Codex does not parse. The host's own terminal
model is seeded from the same snapshot, so it lost the encoding too.

The pane now mirrors its xterm's mouse encoding from the parser (1006,
1016 and a full reset, with xterm's own set/reset rules) and its snapshot
ends with the matching DECSET, after the alternate-screen switch that
readers keep. Programs on the default encoding get nothing appended.

The phone keeps a guard for hosts without this fix: while a wheel-
reporting tracking mode is on but no replayed byte proved the encoding,
a wheel or swipe scrolls locally and a tap or drag sends no mouse report
(the tap still focuses the keyboard). Click-only tracking (?9h) keeps
its arrow-key scrolling on the alternate screen, and an explicit ?1006l
still sends legacy reports.

* fix(terminal): state the default mouse encoding so legacy mouse apps keep phone input

Snapshots now say ?1006l while a program tracks the mouse with the default
encoding (daemon rehydrate and desktop pane serializer), and the phone
treats a tracking enable seen in live output as proof of its encoding.
Legacy-encoding programs such as vim with mouse=a keep phone scrolling,
taps and drags; replay-only tracking with no stated encoding still sends
nothing. An unproven drag falls back to local selection, and an empty pane
snapshot stays empty.

* chore(terminal): type the mouse-encoding tracker inputs for the low-evidence audit
* fix(claude): run Windows hooks without shell operators

Keep neutral replies inside the managed entry and payload scripts, repair missing files from managed registrations, and stop using Git Bash discovery to guess Claude's hook shell.

Co-authored-by: latte271 <junghyeyun27@gmail.com>
Co-authored-by: Bing.Z <zzb@gxsmjx.com>

* fix(claude): keep the Windows hook refresh async and its scripts after uninstall

- List windows-hook-files.ts in the CLI project so typecheck passes.
- Refresh the entry/payload pair only from a surviving entry, with an async
  existence check, so startup refresh stays off the main thread on Windows.
- Keep both scripts on uninstall like every other agent; a Claude session
  still holding old settings keeps answering instead of erroring per event.
- A payload that exists but cannot start falls through to the neutral reply,
  and the missing-payload branch exits early for background jobs.
- Update the EDR posture reference for the operator-free command.

* test(claude): run the Windows hook host legs for real

The live Windows host legs never ran: runProcessSync cannot take a string
stdin (it forces encoding 'buffer'), so every leg threw before starting a
host. Use async runProcess, pass PATHEXT (without it Windows PowerShell 5.1
prints nothing and exits 0 for a .cmd path), and name the host in each
assertion. Drop the POSIX pwsh leg: its drive-mapping shim proved nothing
about Windows, and the Windows legs cover both PowerShell hosts.

---------

Co-authored-by: latte271 <junghyeyun27@gmail.com>
Co-authored-by: Bing.Z <zzb@gxsmjx.com>
…y name (stablyai#24065)

Audit sweep over `src/renderer` (3,206 test files in scope). 58 case declarations
removed across 37 files, 3 test files deleted, 1,033 lines gone. Nine dead
production symbols removed with them.

The dominant pattern this wave was a case whose input cannot reach the behavior its
title claims:

- `getRemoteBrowserFrameStyle` takes `_metadata` unused and returns a constant
  object. Four cases fed it 958x609, a uniform high-DPI 1998x1218, an uneven
  high-DPI frame, and malformed metadata, all asserting the same constant. The
  function cannot branch on any of them; the runtime rejects server-sized frames, so
  the renderer ignores bitmap dimensions by design. Kept the one case whose name
  admits that.
- `shouldIgnoreTerminalMenuPointerDownOutside` takes `{openedAtMs, nowMs}` and has no
  button or modifier parameter. Two cases named for secondary-button and macOS
  control-click passed byte-identical timestamps. Their titles came from the
  docblock's prose, which describes behavior the CALLER delivers.
- A retry case titled "then Retry re-arms" proved that half by deleting a key from
  its own ref object and asserting it was undefined. The `stablyai#6648` budget assertion
  stays; the title now matches what the case proves.

Also removed:

- Cases that cannot fail: three asserting `toBeInstanceOf(Map)` and `size === 0`
  immediately after a `beforeEach` that clears those maps.
- Assertion-free coverage probes, one of which writes a JSON report to tmpdir and
  asserts nothing while its own comment says the sibling case is the gate.
- Copied inventories and export lists, each checked per entry: a 35-name renderer
  git-client export list (every name has 7-40 production references) and a 16-path
  caller census that never verified the callers pass the argument it exists to track.
- Provider-local replays of one factory (`createUsageProviderSlice`), keeping a
  single representative; and a case re-running a shared window-shortcut policy the
  shared suite owns.
- Duplicate invocations, including a `buildNativeChatSendBytes` block whose function
  is `buildNativeChatPasteBytes(text) + '\r'`, with three titles verbatim from the
  paste-bytes cases.
- A slice case whose only assertion reads back `antigravity: null`, a literal in
  `createEmptyRateLimitState()`, under a title claiming a "stable pending key".

Nine production symbols went, each verified to have exactly its own declaration plus
one test reference and zero production callers: `buildNativeChatSendBytes` (its
docblock said "kept for callers/tests" and had none), `isCommandMarkerId`,
`buildNativeChatRenderItems`, `collectToolResults`, `scrapeNativeChatSession`, three
now-orphaned types, and an eagerly-evaluated `WORKTREE_CARD_PROPERTY_OPTIONS`.

One deletion was reverted. An i18n JSX-spacing guard reads `.tsx` source text and
would break under a behavior-preserving rewrite, both true. But all three components
it guards have zero referencing test files, so nothing else catches
`translate('…','Detected from')}` followed by `<code>` rendering as "Detected
fromorca.yaml". The junk patterns describe a test's shape; the ratchet rules describe
consequence, and consequence wins.

Kept deliberately: a randomized parity test measured at 483 of 517 comparisons being
`identity === identity`, because 32 seeds and 2 named cases are genuine differentials
and the fix is a rewrite, not a deletion; a title-tracker parity test whose two sides
are separately-authored implementations that a sibling case proves can diverge; a
Monaco upstream-drift detector that reads the installed package, which is the shipped
contract; and nine bound guards, all confirmed to have live production callers.

Coverage is partial and stated as such: roughly 810 of 3,206 files read case-by-case,
the rest triaged by mechanical scan. Every auditor disclosed its own unread set and
those paths are tracked rather than assumed clean.

Verified: `pnpm test` over all nine renderer chunks — 2,834 files, 25,540 cases, 0
failures (a first run showed 5 failures that did not reproduce and were concurrent
load); `pnpm tc` after clearing `.tsbuildinfo`; `check-reliability-gates.mjs` 140
gates; `check:code-quality:changed` 0 new findings. All three deleted files confirmed
absent from the gate manifest, `cloud/package.json` and
`mobile/tests-typecheck-baseline.txt`.
…tablyai#24064)

Hermes fires `on_session_start` when a session is opened, switched or reset;
its payload carries no turn. Orca mapped it to `working`, so a freshly launched
Hermes held a fresh first-party `working` row for the whole 30-minute staleness
window: `terminal wait --for tui-idle` never settled against an idle composer,
and the sidebar span a phantom spinner.

Map it to `done` with `sessionBoundary`, matching what Claude and the
compatible-lifecycle providers already do for SessionStart, and let the tui-idle
evidence lane settle on a fresh boundary row. A boundary row claims a new
session owns the pane and awaits its first input, which cannot arrive mid-turn,
so it carries none of the stablyai#6011 risk that scoped that lane to DSH; a turn-end
`done` still does not settle.

Fixes stablyai#13653.

Co-authored-by: Brian Grablin <5216789+bgrablin@users.noreply.github.com>
…each (stablyai#24077)

Completes the first pass over every test area in the repository. Sweep over
`mobile/src`, `config/scripts`, `cloud/`, and `tests/` (1,494 files in scope, with
the 24 files under `mobile/src/test-support/rpc-recording/` deliberately excluded).
31 case declarations removed across 17 files, 2 test files deleted, 356 lines gone.

What went, by pattern:

- Cross-boundary replays of a shared helper. A whole mobile file re-ran
  `extractPendingAsk`/`parseAskFromStatus`/`formatAskAnswer`, all owned by
  `src/shared/native-chat-ask.test.ts`, `native-chat-ask-fifo.test.ts` and the
  renderer's interactive-prompt suite — one case title was verbatim identical to the
  owner's, and the owners' inputs are supersets. The mobile file imported the shared
  module directly and exercised no mobile transport, lifecycle or rendering.
- A case whose input cannot reach the behavior its title names: "arms it on Android
  while the drawer is open", where `use-back-claim.ts` has zero
  Platform/OS references, so flipping the mocked OS changes only shadow styles.
- Identity copiers, including one asserting `prSidebarRenderBranch(state) ===
  state.kind` against a production body that is `return state.kind`. The function
  stays; it has three live callers.
- A test of the runtime rather than the product: a case asserting Node's own
  `EventEmitter` crash contract on a bare emitter, with zero production code in the
  path. The guard it documents is exercised behaviourally by the case after it.
- Duplicate invocations, one of them provable rather than eyeballed: with
  `MODULE_SCOPE_ENV_WRITER_PIN = 0`, `files.size <= 0` is strictly implied by the
  sibling's `expect(offenders).toEqual([])`, since a non-empty `offenders` forces
  `files.size >= 1`. The pin's own doc says it may only ever be decreased from 0, so
  it could never become a meaningful bound either. Its policy guidance survives as a
  comment; the file's real ratchet and its regex self-test both stay.
- Expected values produced by the test's own arithmetic, and a p95 case strictly
  implied by a sibling that already pins exact p95 and exact max over a wider range.

One production line goes: the `export` keyword on `assignmentCleanupSteps` in
`cloud/apps/relay/src/assignment-cleanup-steps.ts`. The function itself stays and is
still called internally; only the test-only export was orphaned.

Kept deliberately: everything a gate cites, checked by case title and not only by
file path; a gate-cited case that does not deliver its claim (reported instead — see
below); a cross-version wire cell whose ledger is never invoked, left under the
raised bar for wire coverage; and every limit, bound, quota and provenance guard.

Nothing under `mobile/src/test-support/rpc-recording/` or
`mobile/rpc-foundation/goldens/` was touched — those bytes feed a `recorderSha256`
digest pinning 398 golden recordings.

Verified: `mobile` vitest over the modified mobile files (8 files, 50 cases);
`mobile/scripts/check-tests-typecheck-ratchet.mjs` OK (898 files in program, 125
grandfathered, none @ts-nocheck); relay suite 799 passed; `check-reliability-gates.mjs`
140 gates; both deleted files confirmed absent from the gate manifest,
`cloud/package.json` and `mobile/tests-typecheck-baseline.txt`.

Seven local failures were investigated and none is caused by this change: five
`mobile-web-app-*-render` tests drive `playwright-core` chromium/webkit and need
browsers this machine lacks, `release-checkout.unit.test.ts` needs cross-version git
refs, and `e2e-worker-env-isolation.unit.test.ts` fails identically with its HEAD
content restored — it recurses `tests/e2e` with symlink-following `statSync` and no
depth guard.
…ed (stablyai#24048)

* feat(secrets): warn in Settings when a credential is stored unencrypted

When no OS keyring is usable, the MiniMax stores write the credential as a
plaintext envelope and say so with a console.warn nobody reads. The users this
affects are exactly the ones who never see a main-process log, so in practice
they were told nothing (stablyai#21827).

Report it where the credential is managed instead. Each store gains a
protection reader, the status IPC carries it, and Settings renders a warning
next to the credential it applies to.

Keyed on the stored bytes, not isEncryptionAvailable(): a credential saved
before a keyring existed stays plaintext until it is saved again, so reporting
current capability would call it protected while the file says otherwise. The
readers parse the envelope kind without decrypting, so opening Settings cannot
provoke a keychain prompt.

The console.warn stays. It carries no secret material, and it is still the only
signal on a headless host with no Settings window.

* feat(secrets): extend the unsealed-credential warning to every affected store

The speech key, Linear tokens, Jira tokens and the Bitbucket credential have
the same plaintext fallback the MiniMax stores do, and the same console-only
warning nobody reads.

Add a shared `readCredentialFileProtection` for the four stores that write bare
ciphertext with no envelope, classifying with the same printable-UTF-8 test
`readStoredCredentialToken` already uses — so the reporter cannot drift into
disagreeing with the reader about the same bytes.

Linear and Jira report across every stored workspace/site rather than the
active one: sealing is a host-wide property, so a second workspace stored while
the keyring was missing is exposed even when the active one is sealed. Both
fields are optional, so an older remote host that omits them reads as unknown
rather than as sealed. Bitbucket reports null for env-supplied auth, where Orca
stores nothing and has no claim to make.

Also fixes the credential-connection test double, whose identity-function
`encryptString` wrote a readable token — faithful enough for a round-trip
assertion, but it made the suite assert that a sealed credential was exposed.

* chore(i18n): extract the unsealed-credential notice strings

CI's localization-extraction gate requires every translate() key to exist in
the primary catalog. Inserted in place rather than re-sorting the file, which
is not fully sorted and would have produced a 17k-line diff.

* test(web): pin the null protection fields on the desktop-only MiniMax bridge

The web bridge reports no protection because it stores nothing; the shape
assertions had to move with it.
…ct (stablyai#24035)

* fix(secrets): stop telling Linux users to install a keyring they already run

On a desktop Chromium does not recognise (Hyprland, sway, river, niri) the
selected backend is basic_text and sealing is unavailable, so the at-rest
protection report told the user to install and unlock gnome-keyring — which is
usually already running and serving org.freedesktop.secrets. Chromium simply
never looked, because it picks the backend from XDG_CURRENT_DESKTOP.

Split the Linux unavailable path on the selected backend: an unrecognised
desktop now says so, names XDG_CURRENT_DESKTOP, and points at
--password-store. A backend that did resolve but cannot seal keeps the
install-and-unlock text, which is correct there.

The backend read stays after isEncryptionAvailable(), the call that performs
the D-Bus probe, so nothing new blocks (STA-5765 timing unchanged).

* fix(secrets): seal credentials on Linux desktops Chromium cannot detect

Chromium picks its os_crypt backend from the desktop-environment env vars and
recognises none of the tiling compositors (Hyprland, sway, river, niri). On
those it selects basic_text, whose key Electron only exposes after an explicit
setUsePlainTextEncryption() this app never calls — so isEncryptionAvailable()
is false and every credential store takes its plaintext fallback, while
gnome-keyring sits on the session bus unasked (stablyai#21827).

Name gnome-libsecret ourselves, but only where it provably cannot hurt:

- Never for a desktop Chromium does resolve. Overriding a working selection is
  the one change that could strand already-sealed credentials, and a KDE
  session identified only by KDE_FULL_SESSION is the case that matters.
- Only when the secret service's default collection is present AND unlocked,
  measured out of process with a killable 1.5s deadline. A locked collection
  with no unlock prompter is what made isEncryptionAvailable() block for 76s to
  first window (STA-5765); selecting libsecret there would trade silent
  plaintext for a frozen app. The probe reads the Locked property, which cannot
  itself trigger a prompt.

Anything unexpected — no session bus, no gdbus, no name owner, timeout — leaves
Chromium's own choice alone, so the worst case is today's behaviour.

The desktop-presence env vars come from a review finding by @Raajik on stablyai#21831,
confirmed there against base/nix/xdg_util.cc.
…ey name (stablyai#24101)

Sweeps the triage-only backlog: 2,269 files that earlier waves saw and skipped for
size, reconstructed from the unread lists five waves of auditors disclosed. 35 case
declarations removed across 15 files, 2 test files deleted, 487 lines gone. No
production file touched.

These are large integration suites, so the junk here is individual cases buried among
real coverage rather than whole bad files. The dominant defect was again a case whose
input cannot reach the behavior its title names:

- `resume-sleeping-agent-session-remote-compat.test.ts` (deleted) — two cases titled
  for "transport-level host authority on a capable host" and "host authority is not
  known". `resume-sleeping-agent-session.ts` has no host-authority or capability
  concept at all, and its only read of `origin` is
  `if (!record.origin && record.state === 'done')`, unreachable for both rows. Both
  executed one identical path. The surviving contract is owned by
  `resume-sleeping-agent-session-execution-host-scope.test.ts`, which drives a real
  host catalog.
- `project-group-header-drag.test.ts` (deleted) — four cases setting
  `data-project-group-header-id`, which the predicate never reads. Its subject,
  `isProjectGroupHeaderActionTarget`, is byte-identical to `isRepoHeaderActionTarget`
  apart from the function name and imports the same `REPO_HEADER_ACTION_SELECTOR`, so
  all four cases were a strict subset of `project-header-drag.test.ts` using identical
  `data-repo-header-*` fixtures.
- `remote-worktree-history-cleanup.test.ts` — "repeats idempotent cleanup through the
  PTY owner" against a six-line best-effort forward with zero dedupe state. The case
  called it twice and asserted the mock recorded two calls, which is arithmetic over
  the test's own loop; nothing about idempotence was established.

Also removed:

- Runtime assertions of type-level facts, where production already makes the check at
  a stronger boundary: `const adapterSatisfiesPort: AdapterIsPort = true` followed by
  `expect(...).toBe(true)` — unconditionally true, while
  `createExpoGenerationFileSystem(): GenerationFileSystem` is explicitly annotated and
  passed into `createGenerationStore` at a typed call site. And a case named "does not
  typecheck" whose runtime assertion is a length check on its own literal, declaring
  its own local annotation so it could never notice the production annotation
  weakening.
- Private predicate tests duplicated at a real boundary: four `repo-slug-cache` cases
  delivered by `repo-slug-index.test.ts`, which drives the same resolution through the
  hook, the real store and the preload bridge, while the cache-level versions hand-seed
  the internal map and break on a cache-key format change.
- Duplicate invocations owned at the shared boundary, including commit and push
  recovery cases owned by `src/shared/source-control-recovery-agent-command.test.ts`.

Kept deliberately, verified rather than assumed: the production duplication behind the
deleted drag test was left alone, because `REPO_HEADER_ACTION_SELECTOR` ends in generic
`button, a, input, textarea, select`, so genuine action targets inside a group header
still match — it is an unspecialised copy-paste, not a live bug, and collapsing two
functions is a refactor. Reported instead.

Auditors' probes produced 20, 11 and 13 candidate hits for the signature-versus-title
shape across their chunks; every one was inspected and every one was genuine coverage.
No deletion in this wave rests on a probe alone.

Coverage is partial and stated as such: of 2,269 files, roughly 100 were read
case-by-case and the remainder reviewed at title-plus-import level. Each auditor listed
its own unread set. The largest remaining surfaces are `src/main/agent-hooks` (95),
`src/main/claude` (100), `src/renderer/src/lib/pane-manager` (62) and the 20 largest
sidebar suites.

Verified: 2,583 desktop test files / 25,636 cases pass, plus one pre-existing
`it.fails` marker; the two modified mobile files pass (57 cases);
`check-reliability-gates.mjs` 140 gates; `check:code-quality:changed` 0 new findings.
Both deleted files confirmed absent from the gate manifest, `cloud/package.json` and
`mobile/tests-typecheck-baseline.txt`.
nwparker and others added 29 commits October 2, 2026 04:26
* fix(opencode): skip absent session stores without warning

* fix(opencode): retain warnings for inaccessible session stores

---------

Co-authored-by: drakeo338 <paranoyouz@gmail.com>
…lyai#24582)

* fix(terminal): fill DOM block glyphs only in repainted rows

Adapt the block-fill approach from PR stablyai#15955 and bound painting to xterm render ranges without observer or animation-frame rescans.

Co-authored-by: mmarabel <mmarabel@users.noreply.github.com>

* fix(terminal): preserve block fills across DOM row replacements

---------

Co-authored-by: mmarabel <mmarabel@users.noreply.github.com>
* ci: bound unit jobs to one hour of execution

* docs: keep CI budget notes clear of the headless follow-up

* docs: keep CI deadline evidence in the pull request
…tion (stablyai#24612)

* fix(opencode): retain failed and stopped TUI turn outcomes

Adapt the root verdict proposal from brennanb2025 in PR stablyai#23105 to the current TUI-owned lifecycle, keeping hook-store authority and existing mainAgent semantics.

* fix(opencode): let auto-approved permissions settle before attention

* fix(opencode): reconcile cached outcomes with completed session turns

* fix(opencode): bind terminal verdicts to ending event timestamps

* fix(opencode): publish approval cards for permission requests
* chore(deps): update reviewed desktop dependencies and tooling

* chore(deps): update compatible mobile packages and Fastlane

* chore(deps): update cloud transports and enforce release age

* chore(deps): patch documentation dependencies and record review

* chore: remove dependency review reports

* test(linear): smoke-load resolved SDK through CommonJS loader

* fix(deps): keep native rebuilds from reinstalling addon dependencies

* fix(native): invoke installed node-gyp directly for Node rebuilds

* test(cloud): exclude observer probes from row-lock timing budget

* test(mobile): preserve the CSS writer receiver in viewport spy

* test(native): remove obsolete batch-shim fixture exception

* Stream native rebuild output through the process wrapper
…Auth (stablyai#24593)

Consolidates the reviewed version and visibility work from stablyai#24283 with the probe gating and structured error classification from stablyai#24296. Reject unsuccessful version probes, preserve diagnostic precedence, and assert that real quota reads start no model turn.

Co-authored-by: Pablo Werlang <19828711+werlang@users.noreply.github.com>
…in (stablyai#24596)

* fix(antigravity): make POSIX hooks executable with bounded JSON stdin

* test(antigravity): decode SSH hook shell command before assertions
)

* fix(files): ignore nested generated directories on macOS

* fix(files): filter generated filenames containing newlines
* test: check plugin fixture worktree cleanup

* test: check E2E installer package environment
* fix(terminal): let wide panes use up to 1024 columns

Adapt the wider viewport limit proposed in stablyai#16578 to the current runtime, shared RPC schemas, and preview sizing.

Co-authored-by: innocarpe <innocarpe@users.noreply.github.com>

* test(terminal): wait for probe output after command echo

* test(terminal): align RPC boundary with wider viewport limit

---------

Co-authored-by: innocarpe <innocarpe@users.noreply.github.com>
…its (stablyai#24533)

OpenCode's question tool (and Pi/OMP ask tools and custom modals) put the
hook row in a waiting state but paint no dialog wording the blocked-text
layer knows, so the hook lane read the wait as pending and the tui-idle
wait timed out with no blocked reason. For agents whose hooks are
authoritative, a wait the hook reports with no recognised dialog text now
blocks with the existing agent-interactive-prompt reason, unless input
reached the pane after the row (it may have answered the question before
the next hook arrived).
…ablyai#24536)

After a Plan-mode turn Codex shows a menu that owns the keyboard, but its Stop
hook has already fired, so a tui-idle wait settled ready and sent text into the
menu. A Codex blocked text anchor now names it an interactive prompt while its
key row ends the tail, written against a new 0.160.0 capture that is replayed
by the readiness census.
Orca downloads a newer agent-state-rules.json from a fixed GitHub release (stable or next channel), validates it like the bundled rules, and applies it without a restart; a local override wins over the download, which wins over the bundled rules. A hand-started workflow from main is the only publisher; merging publishes nothing.
…orking rows to the rules (stablyai#24541)

A headless orca serve has no window to write the Codex ready title, so a
Codex tui-idle wait settled only after three quiet seconds. Rule files gain
profile.hooks: "turn-end": a fresh hook done settles the wait, while a
working or permission row leaves the decision to the rules, so a Codex
whose Esc posts no event (before its Interrupt hook) cannot hang the wait.
Codex moves from identity-only to turn-end.
@innocarpe innocarpe closed this Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.