Conversation
…g out its 1s timeout (stablyai#23920) * fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout xterm paints nothing while DEC mode 2026 (synchronized output) is open and only force-flushes after 1000ms. Codex wraps every draw in mode 2026, so any byte gap or chunk split that loses the closing \x1b[?2026l freezes the pane for a full second and then repaints in one burst. Orca never emitted \x1b[?2026l anywhere, and three paths could destroy a TUI's: the per-PTY pending cap drops buffered output wholesale (mode 2031 was already salvaged there, 2026 was not), main sliced pending data at a blind 16KB offset that can land inside an open frame or sever the 8-byte marker, and the renderer's backlog warnings replace a queued tail that may hold the close. - salvage the 2026 latch across dropped output, mirroring the existing 2031 salvage, and append the release on both delivery sites - ground 2026 in RESET_AFTER_BYTE_GAP and the replay baseline, and in both backlog warnings, so every drop path is self-healing - make main's 16KB flush split frame-aware instead of a blind byte offset - lift the synchronized-output scanner into shared/ so main and the renderer use one implementation Closing a frame early costs one premature repaint; leaving it open costs a second of blank screen, so the asymmetry favours always closing. Also adds the reproduction this needed: the pre-existing typing bench observes the xterm BUFFER, which the parser fills while rendering is held, so it scored these freezes as fast echoes. * fix(terminal): stop the renderer's queue drain cutting inside an open DEC 2026 frame takeQueuedChunk sliced a queued chunk at a blind byte offset to fit the 16KB coalescing budget, which can strand a frame's closing \x1b[?2026l in the residual until a later drain. Same defect as main's flush split, same fix: reuse the frame-aware split helper. Usually masked because the drain coalesces adjacent chunks and reassembles what main split, but not when the budget boundary falls inside a frame. * fix(relay): keep the SSH path's bounded slice outside an open DEC 2026 frame pty-handler split pending output at a byte offset with a surrogate-pair guard but no synchronized-output awareness, so a frame straddling the 16KB wire slice had its closing \x1b[?2026l stranded in the remainder — the same defect just fixed on the local path, on the path AGENTS.md requires us to consider. Placed before the surrogate guard so that guard keeps the final say, and floored at 2 so frame alignment can never walk a healthy slice into the guard's decrement and then into the chunkChars <= 0 pause-and-retry path. Also drops a dead `splitAt === 0` branch in takeQueuedChunk: both callers pass a positive limit and the helper never returns 0 for one. The two new split tests were each confirmed to fail without their fix. * test(terminal): sweep the DEC 2026 split helper over escape-sequence shapes and every limit Covers OSC 52, DCS, repeated open/close markers and limits 1..len+3, asserting the result never exceeds the limit, never reaches 0, and stays byte-exact. Also pins that a buffer beginning inside an open frame degrades to the blind offset rather than doing something worse, and documents that callers do not thread latch state. * fix(terminal): ground DEC 2026 on the daemon slice, the recovery replays, and the process boundary Four more sites could strand the latch, found by sweeping every path that drops, splits, or replays terminal bytes. - daemon-stream-data-batcher: the 64KB bulk-write slice used a surrogate-only clamp, and its remainder is HELD until 'drain' — "seconds for multi-MB backlogs" per the file's own note. A frame straddling that boundary parked its \x1b[?2026l behind the hold, blanking the pane past xterm's 1s timeout once per frame for as long as the backlog lasted. This is the default daemon-backed pane path, so it is the one users actually hit. The new clampToSafeBulkWriteSplitIndex frame-aligns first and surrogate-clamps last, and lives in daemon-stream-data-split alongside the policy it belongs to. - replay-data-drain and remote-runtime-terminal-binary-snapshots wrote a bare \x1b[2J\x1b[3J\x1b[H, which does not clear mode 2026 — so on the SSH/remote reconnect path, the very event most likely to sever a frame, the whole replay could paint nothing. - ipc-pty-attach: trimIncompleteTerminalControlTail can cut a half-written \x1b[?2026l while its opening marker survives in the replayed prefix. - PROCESS_BOUNDARY_GROUND: the "process that armed these modes is gone" ground omitted 2026, the last unexplained gap in that file. A disable, so it still satisfies the recovery barrier's ownership scan (only ?25h may be an enable). Recovery-path expectations updated where they pin the emitted bytes. Deliberately NOT touched: apply-reattach-payload and ssh-snapshot-prepaint already ground via buildSnapshotReplayPrologue. Still unfixed, deferred with reason: terminal-output-frame-chunks.ts splits the remote wire on accumulated UTF-8 byte width and needs a different shape than the char-index helper; desktop clients reassemble in main's pending buffer, so the exposure is mobile/web only. * fix(terminal): emit the DEC 2026 release before the mode-2031 tail, and stop claiming the drop path writes it Two corrections from adversarial review of the earlier commits. 1. Ordering bug I introduced. getDroppedMode2031RendererData ends with `state.tail`, which extractPrivateModeScanTail deliberately retains as an INCOMPLETE private-mode sequence so the next chunk can resolve it. Appending the 2026 release after it put an ESC behind a dangling CSI, aborting it and silently losing whatever mode spanned the drop boundary. The release now goes first. 2. The drop-path release does not reach xterm in the dominant case, and the comment now says so instead of implying otherwise. live-data-callback's droppedOutput branch discards `data` and salvages only queries (salvageRendererQueriesFromDiscardedRestoreData handles CPR/DA1/OSC colour; \x1b[?2026l is not a query), so for hidden panes and visible panes outside foreground-restore backpressure the synthesized release was dropped. The grounded snapshot replay releases the latch instead. I tried writing it through writePtyOutputToXterm there and reverted: it consumes the pending hidden-output snapshot and broke pty-connection-hidden-snapshot-resize-signals ("re-restores a skipped alt frame"), so the release rides the restore rather than perturbing that state machine. Residual gap, documented: a cap-dropped pane whose restore never arrives. The salvage is still load-bearing on the fall-through path, so it stays. * fix(terminal): release DEC 2026 on the reattach clears, floor the split, and correct the freeze framing Remaining findings from adversarial review. - apply-reattach-payload's three bare-clear branches (:63 daemon snapshot, :229 relay replay, :269 cold restore) had no release anywhere in their sequence: I checked all seven POST_REPLAY_* profiles reachable via chooseReattachReplayReset and none contains \x1b[?2026l. Only the buildMainModelSnapshotReplayWrites branch was grounded, so covering the streamed replay path and not the main reattach path was inconsistent. Verified no production code matches these clear strings — the three test updates are mock equality, and each was confirmed to fail without the source change. - clampToSafeBulkWriteSplitIndex could return 0 (('\u{1F600}aaaa', 1) — alignment returns 1, the surrogate clamp decrements to 0), which would leave a zero-length slice that never shifts the batcher's queue entry and spin its drain loop. Unreachable from today's only caller, but it is exported with an unstated precondition. Floored at 1. - Frame alignment could halve per-PTY flush throughput: main re-queues the remainder with eligibleRound = round + 1, so the shortfall cannot be refilled in the same round, and aligned size is floor(W/F)*F — 50% worst case in the 8-16KB band, which is exactly the full-screen redraw burst that reaches the pending cap. Alignment is now rejected below half the window, preferring throughput and letting the reset profiles release the latch. Framing corrected throughout: bufferRows records a row range and clears nothing, so the pane freezes on its last painted frame — it does not go blank. The real trade is "stale but coherent for <=1s" versus "immediate partial frame", and RESET_AFTER_BYTE_GAP (written alone, with no repaint behind it in the same write) is the one site that can newly flash a partial frame. Said so at the constant instead of implying the release is free. * fix(terminal): rename the shape-flagged symbols the anti-slop audit rejects CI's anti-slop gate rejects "shape" in symbol names as structural rather than domain language: `shapes` -> `outputSamples`, and `writeCodexShapedEchoProbeScript`/`codexShapedEchoProbeScript` -> `writeCodexEchoProbeScript`/`codexEchoProbeScript`.
Disables pre-gathering of ICE candidates to avoid timeout or flakiness during test probe initialization.
…blyai#23992) Semantic sweep of src/main/ai-vault. 48 cases removed across 11 files; no production code touched, no files deleted. The largest single removal is a 36-case block (6 agents x 6 env values) in `session-scanner-agent-root-overrides.test.ts` whose assertion was a self-comparison: `root === join(root)`. The sibling `falls back to the default root for %j` pins `roots[0]` to the exact absolute default, and the extra roots the block also scanned (`agent_logs`, `.clawdbot/agents`) are homedir-derived and unaffected by the env var it varied. The stablyai#13082 rationale comment is kept on the surviving case. Inputs the production path never reads: - `session-scanner-codex-tool-records.ts:89` reads only `change.unified_diff ?? change.content` and never `change.type`, so the `{ type: 'delete' }` row was the identical path as `add`; the `add`/`update` rows remain as the two real disjuncts. - `codexSpawnDepth` accepts any positive integer, so depth 2 was the same branch as depth 1. - `agentPath` is an independent `??` fallback with no cross-field logic, so "keeps the rest of the spawn when the naming path is null" passes either way. - `getAiVaultWslHomeDirs` reads only `platform` and a `hasCachedWslDistros()` gate, so a case varying which distros are "currently running" took the default branch; the filtering lives entirely inside a mocked async call. Cases that cannot fail for the reason they name: an all-unknown-agent response whose throw requires `malformedSessionCount > 0` when it is 0; a symlink rejection byte-identical to the directory case above it (`isFile: () => false`); a runtime restamp whose fixture already carries the `executionHostId` and `id` it asserts. Also removed: duplicates of a stronger sibling in `session-list-results`, `session-parse-cache-persistence` (same `schemaVersion !==` gate), `session-scanner-claude-title` (owned by the subagent-prune test, which also asserts `subagentTranscriptCount`), and a session-scanner listing case whose count is N-independent because `fixedChildFileSegments` does one readDir plus a direct stat per child — so a per-session-readDir regression fails at N=1. Three keeps worth recording. Spy-counting tests were kept where real code runs: `session-scan-cutoff` and `session-scanner-dedup-batches` count `Array.prototype.sort` / `RegExp.prototype.test`, but drive the real scanner over 128 fixtures and assert the limit-ordered result too, so a re-sort-per-candidate regression fails them for the right reason — unlike a bench whose assertion was arithmetic over its own constants. `session-scanner-claude-unicode-scope` keeps its locally re-spelled dir-name encoder deliberately: importing the production one would hide Orca drifting from Claude's actual naming. And `session-parse-cache-persistence.test.ts:175` stays although `keys.length > 0` cannot fail — it is the only reference to the `satisfies Record<keyof AiVaultSession, true>` table, so deleting the case would make that type-level ratchet dead.
…tion (stablyai#22720) * fix(onboarding): build the agent step around skills, not CLI registration The checklist step "Enable Orca CLI" was marked done once the agent skills were installed, while Settings -> Browser still showed CLI registration as a pending step. Orca terminals already put the bundled CLI on PATH, so registration only matters for shells Orca did not launch, plus WSL, where `orca-ide` exists only once registered. - Rename the step to "Give agents Orca skills"; setup registers the CLI only for WSL (isOrcaCliRegistrationRequired), via onboarding-cli-registration.ts. - Settings -> Browser drops the CLI step outside WSL (2 steps instead of 3). - Skills panel: "All skills installed" + "Update skills" replaces the disabled button; status pills sit top-right; no "Installed" beside "Unavailable". - Full Disk Access moves to the "Start work in multiple repos" step. Fixes STA-8306 / stablyai#22524. * refactor(onboarding): simplify agent-skill step state after review - One done rule: isAgentCapabilitiesDone in feature-wall-setup-progress.ts, reused by the skills panel instead of a mirrored copy. - 'unavailable' is an install-status tone instead of a second boolean; pill/note rendering moves to AgentCapabilityStatusBadges.tsx. - "Update skills" skips Computer Use when it can't run (no warning toast); setup takes an explicit selection. - The WSL gate lives in registerOnboardingCliIfRequired; onboarding deps drop the now-unreachable host CLI branches. - BrowserUsePane: one cliRequired/cliReady pair, no host CLI status fetch, WSL-only enable path, single "Finish the steps below." string. - Full Disk Access placement goes through a SelectedStepFooter switch. - Prune orphaned locale keys (and their boot-bundle entries); fix a stale comment. * fix(skills): require CLI registration only for WSL setup * fix(skills): retain CLI install labels for WSL * docs(skills): clarify remaining WSL registration fallback * fix(skills): skip WSL registration when the host confirms managed CLI access * fix(wsl): prepare managed shell wrappers before onboarding probes * test: update daemon capability and terminal hook expectations * refactor(onboarding): remove CLI registration checks from skill setup * refactor(setup): remove redundant state and obsolete registration scaffolding * fix(settings): stop registering the CLI before installing the CLI skill The General > Orca CLI skill panel still registered `orca` on PATH before opening the install terminal, contradicting the rest of skill setup. Orca terminals already provide the CLI, so the shell command toggle now says it is only for terminals outside Orca. * style(onboarding): polish the setup checklist and first-run steps Make onboarding monochrome: completion is a neutral check, selection a neutral outline, and color only flags real problems. Tighten the checklist rail and header, single-line agent cards with a grid that scrolls only when it runs out of room, a labeled permission switch under the grid, calmer notification and skill cards, sentence-case copy, and a labeled "Hide checklist from sidebar" action. Workspace setup leads with "Add project" when no git project exists. * fix(emulator): drop the Enable Orca CLI step from agent control setup Agents that drive the emulator run in Orca terminals, which already provide the `orca` command. Agent control setup in the emulator card and Settings is now a single step: install the Orca CLI skill. * fix(onboarding): hide the Full Disk Access card once access is granted A granted card has no remaining action and only takes space on the add projects step. It also no longer flashes a "Checking" state before the first status arrives. * fix(onboarding): address review on permission warning, hide button, and translations - Name the permission switch "Yolo mode" (matching Settings > Agents) and state the risk: agents act without asking and some bypass their sandbox. - Hide the modal's "Hide checklist from sidebar" button below sm, where the header centers its title under it; the sidebar entry keeps its own control. - Translate every string this PR adds into es, fr, ja, ko, and zh.
…ts input (stablyai#23996) Semantic sweep of src/main/codex, daemon and git. 16 cases and 4 it.each rows removed across 11 files; no production code touched, no files deleted. Inert setup — the case flips the verdict by hand and the ceremony changes nothing: - three `codex-stale-pane-accounts` cases varied `environmentHomeOverride`, which `codex-stale-pane-accounts.ts:38-42` never reads (it reads `selectionKey`, `homeRoute` and `accountId`); one also rewrote `.zshrc` and called `__resetShellStartupEnvCache()` while both verdicts came from the `activeHostHomeRoute` argument the test sets directly. Names outrunning their input: - an `it.each` row named `'leap century'` stepped `['1999','12','31']` to `['2000','01','02']` — it never touches February, so it is the `'year rollover'` row under a name promising a leap rule; - `it.each([1, 2, 3])('keeps pre-ownership baseline version %s canonical…')` collapsed to version 1: `config-settings-baseline.ts:185` only validates the version is one of the three, and the policy is driven by `parsed.mcpServers === undefined`. Version 3 without `mcpServers` is an impossible shape, since v3 is what the writer emits *with* it. A self-comparison: `codex-session-index-heal.test.ts:802` looped `CODEX_SHORT_LIVED_PROBE_APP_SERVER_ARGS` asserting `args` contains each entry, but `buildNativeHealInvocation` sets `args: [...CODEX_SHORT_LIVED_PROBE_APP_SERVER_ARGS]` (`codex-session-index-heal.ts:298`) — the constant against itself. The full `toEqual` with the shim is owned by `codex-short-lived-app-server-spawn.test.ts:33`; the `CODEX_HOME` pin stays and the case is retitled to match. Five exact source greps over `POSIX_PROVIDER_SUPERVISOR_SCRIPT` went because `codex-app-server-posix-supervisor.integration.test.ts` executes that same script and asserts the behavior for real — owner-PID refusal, group leadership, group-SIGKILL escalation after a SIGTERM-ignoring provider, and exit-code relay. Three greps in that same file are KEPT, one wave after ~84 files of that shape were deleted, because nothing else can reach them: no integration test inspects the provider's env, so a leaked `ELECTRON_RUN_AS_NODE` would silently change how codex runs; and `stdin.once('close')` matters because the integration test's `stdin.end()` would still pass if only `'end'` were registered. Also collapsed four `it.each` blocks whose callback took no parameter, so every row ran the identical body while the name advertised per-scenario coverage: `['local IPC', 'SSH remote runtime']` built one transport, and `['visible blur', 'terminal tab switch', 'split pane switch']` plus `['terminal close', 'tab unmount']` each ran one harness. Those scenarios were never constructed; the false claim is removed rather than the coverage, because there was none. A single-row `it.each` whose name renders truthfully, and one whose callback is a named function that does take the parameter, are untouched.
…tablyai#23707) Every image drop is now sent to the terminal as a bracketed paste, so agent TUIs (Claude Code, Codex, Pi) attach it. Previously, names that needed shell escaping, such as `download (1).png`, and names with spaces, which Codex's shlex splits, were typed as keystrokes or pasted raw and stayed as text. Safe names are still pasted raw. Names with spaces or shell metacharacters are backslash-escaped inside the paste on POSIX shells, which Claude Code, Codex and pi-image-paste all unescape, apostrophes included. Windows shells keep double quotes. Non-ASCII characters such as the U+202F in macOS screenshot names stay bare. Names with control bytes are still typed. Fixes stablyai#23703 Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
stablyai#24000) Audit sweep over `src/cli`. 23 cases retired and 2 `it.each` tables collapsed to the rows their parameter actually reaches. What went, by pattern: - Table rows whose varied parameter production never reads, so every row ran one identical path. - Second and third invocations of a contract already proven by the case above them, differing only in a field the assertion ignores. - Argument-shape and private-predicate checks duplicated at the real CLI boundary, where the same input is already driven end to end. - Assertions whose expected value came from the same helper under test. `src/cli/command-suggestion.ts` loses `export { levenshtein }`, a re-export no production caller used. The one test that stubs edit distance spies on `../shared/edit-distance` directly, which is the module `command-suggestion` imports, so the seam it needs is unaffected. Kept deliberately: `orchestration-lifecycle-json-rejection.test.ts` and `orchestration-migration.test.ts`, both named in `config/reliability-gates.jsonc` as sole evidence for a gate. While auditing the latter, its replay dimension turned out to be inert -- `it.each([false, true])` varies `lifecycle.duplicate`, and `hasLifecycleVerdict` (`orchestration-worker-settlement.ts:112-132`) reads only `action`, `authority` and `outcome`. The gate at `reliability-gates.jsonc:15861` nonetheless records "first and replayed legacy worker_done settlements are accepted". Left exactly as found and reported rather than collapsed, because correcting a gate's claim or adding real replay coverage is the owner's call. Verified: `pnpm test src/cli` (131 files, 1474 passed), `pnpm tc`, `check-reliability-gates.mjs` (140 gates), `check:code-quality:changed`.
…cking main on sync git (stablyai#23998) * perf(git): stop blocking main on the open-on-remote git cascade `getRemoteFileUrl` ran up to 6 sequential `gitExecFileSync` calls on the Electron main thread — `remote get-url`, then `getDefaultBaseRef`'s `symbolic-ref` plus up to four `rev-parse --verify` probes — each with its own 15s timeout and no yield between them. A complete async twin already existed (`getDefaultBaseRefAsync` -> `resolveDefaultBaseRefViaExec`, sharing DEFAULT_BASE_REF_PROBES), so the sync cascade is deleted rather than converted. `getRemoteUrl`, `getRemoteFileUrl` and `getRemoteCommitUrl` become async; all four downstream callers were already async (`filesystem-git-url-handlers` inside `ipcMain.handle`, `runtime-git-diff-commands` async methods) and the provider contract already typed both wrappers `Promise<string | null>`, so no new async plumbing was needed. Removes 3 of the 10 `gitExecFileSync` sites and the confusing name collision with the unrelated async `getDefaultBaseRef` in hosted-review-creation-git-state. The base-ref regression tests keep their coverage, repointed at the public async `getBaseRefDefault`. * perf(git): resolve the repo root in one sync spawn instead of two getGitRepoRoot ran `rev-parse --is-inside-work-tree` and then `rev-parse --show-toplevel` as separate blocking spawns. Each sync git call holds the main thread for up to its whole 15s timeout, so the spawn count is the cost — and this function is called twice per "Add Project" on a linked worktree, once directly and once through getLinkedWorktreeMainRepoRoot's self-recursion. Combined into one invocation. Safe only here: in a bare repo the combined form exits non-zero, and both that throw and the plain `false` already land on the same marker-scan fallback. probeGitRepo deliberately does NOT combine — it has to read `false` cleanly to go on and detect a bare repo, which the combined form's exit 128 would misread as indeterminate. * perf(git): rebuild only the repos whose authorized roots actually changed One worktree create called `invalidateAuthorizedRootsCache()`, which dirties every registered owner. The next authorization-requiring IPC then rebuilt by listing EVERY repo — and the rebuild never consulted `dirty` when choosing what to list, so `dirty` gated only whether a rebuild ran, not its scope. At 58 repos that is 58 `git worktree list` spawns, roughly ten seconds of git wall-clock through an admission budget of four, to rediscover roots one repo changed. Both halves were needed; scoping the invalidation alone changed nothing. - `markAuthorizedRootsOwnerDirty` dirties a single owner, reusing the per-owner primitives `registerWorktreeRootsForRepo` already used. It leaves `baseRevision` and the per-repo revision map alone — that pair is the global side-effect-token fence, and bumping it would retire in-flight tokens for untouched repos. - `rebuildAuthorizedRootsCache(store, onlyDirty)` re-lists only owners that are dirty, have no listing yet, or still hold recovered roots (those are retired by comparison against a fresh listing, so skipping them would strand them as authorized). Only `ensureAuthorizedRootsCache` passes `onlyDirty`; an explicit rebuild keeps re-listing everything because callers use it to force a refresh — `filesystem-auth.test.ts` pins that contract. `invalidateAuthorizedRootsCacheForRepo` wraps the primitive and falls back to the global form for an unknown owner or a missing store, rather than silently skipping an invalidation and leaving a stale allowlist. Applied to the worktree-create path. Changes that can alter the owner SET (store swap, host/WSL re-routing, nested-repo import, folder->git upgrade) stay global. Removal paths are not converted yet. The allowlist contents are unchanged and the failure direction is a false denial rather than a false allow. The relist predicate is split into its own module so it is testable alone and the cache file stays inside its line budget without a suppression. * test(perf): measure what git orchestration actually costs the main thread The existing churn probe (ORCA_MAIN_THREAD_DIAGNOSTICS=1) reported spawn-initiation cost for git/gh/glab only — its 7 call sites all sit inside git/command-runner — so it was blind to `spawnProcess`/`runProcess`, the repo's own mandated wrapper, and to the blocking `execFileSync('ps')` per PTY resize. That understated total churn across 115 main call sites. - `spawn-observer.ts`: a settable seam, since shared code cannot import src/main. Unregistered in the daemon/relay/CLI, where it costs one boolean check. - `spawnProcess` brackets `nodeSpawn` and reports; exec-file-capture's own report is removed because it routes through runProcess and would double-count. - `posix-pty-foreground-group` now reports its full blocking duration. Note this lands on the daemon, not main, whenever the daemon hosts the PTY. - `ORCA_UNMINIFIED_MAIN=1` build flag, because a minified main bundle cannot attribute CPU-profile self time to real function names. Defaults unchanged. - `main-thread-git-cost.spec.ts` + `analyze-main-cpuprofile.mjs`: sweeps concurrency against real registered repos, captures the churn lines and a V8 CPU profile of main per phase. What it found, which is why this is worth keeping: at the width-4 admission ceiling (~90 git:status/s) main sees ZERO event-loop gaps over 50ms and a worst gap of 23ms, and is 85% idle. Git orchestration does not stall the main thread. Of the cost it does incur, spawn-init is 58%, parse 5%, stdout drain 4%. * test(perf): name the inspector params type the anti-slop gate requires The broad `object` parameter trips anti-slop(no-object-parameters); the only Profiler call that passes params sends `{ interval }`.
…ontract (stablyai#24007) Audit sweep over `src/relay`, `src/preload` and `src/shared` (1,087 test files reviewed). 101 case declarations removed across 40 files, 6 test files deleted outright, 1,143 lines gone. Executed-case count falls further, since several removals were `it.each` tables. Dominant patterns, by frequency: - Self-comparisons that cannot fail: `expect(f(x)).toBe(f(x))`, `JSON.parse(JSON.stringify(literal))` deep-equalling the literal for a type with no codec, and `normalizeKeyToken(t) === normalizeKeyToken(t)` presented as proof of memoization. - Object literals asserting their own fields back, where the guarantee comes from the type annotation and the runtime assertion cannot fail. - Copied inventories: constants compared to their own initializers, and a function returning a copy of an exported constant checked against that constant's literal contents. - Duplicate invocations of a contract owned at a stronger boundary, including provider-local replays of a shared helper. - Table rows varying a field production never reads, so every row runs one path. - Names promising more than the input exercises: a "Windows launch" case in a module with no platform input, and a case whose named branch is never entered. Two production symbols go with them, each a test-only export whose sole caller was a deleted case: - `getGitHubProjectRefInputByteLength` — a one-line forward to `getClipboardTextByteLength`. The real bound (`GITHUB_PROJECT_REF_INPUT_MAX_BYTES`) and its guard stay. - `GRAB_STYLE_PROPERTIES` — an intended shared source of truth that nothing ever consulted; the property set is hand-enumerated at three independent sites. One case was deliberately restored and strengthened rather than dropped. The relay integration suite is the only place the real `SshChannelMultiplexer` is wired to `RelayDispatcher`, so it reaches transport behavior the handler suites cannot (they use `createMockDispatcher`). Its `fs.writeFile` roundtrip is the one case producing a void result, and `JSON.stringify` drops an absent `result` member — a shape no other surviving case exercises. Restored with an assertion pinning what the client actually observes: `null`, not `undefined`. That assertion failed on first run, so the fact was previously unasserted anywhere. One deletion was reverted mid-audit. A case asserting that optional fields stay invisible to "old attach and ready decoders" builds those decoders from `z.object` schemas declared in the test file, so it demonstrates zod's unknown-key stripping rather than anything shipped. It is nonetheless the only forward-compatibility coverage these envelopes have, and `reliability-gates.jsonc:6232` names it as evidence verbatim, so it stays. Note that `check-reliability-gates.mjs` passed both with and without it: the script resolves manifest paths and commands, and does not check that a named assertion still corresponds to a live case. Kept deliberately: everything a reliability gate cites as evidence; the three `registers all expected handlers` RPC manifests (a dropped registration is a silent wire break no type checker catches, and one carries the STA-4571 `pty.ackData` ratchet); the `child-process` direct-import ratchet; and prototype-spy cases paired with a `.repeat(10_000)` input, which assert a real memory bound rather than merely forbidding a technique. Verified: `pnpm test src/shared src/relay src/preload` (1073 files, 11996 passed, 1 pre-existing `it.fails`, 131 skipped), `pnpm tc` after clearing `.tsbuildinfo`, `check-reliability-gates.mjs` (140 gates), `check:code-quality:changed` (0 new findings).
…y proves (stablyai#24025) Audit sweep over `src/main/{native-chat,startup,daemon,skills,ssh,providers,git, persistence,claude,agent-hooks,github}`. 35 case declarations removed across 23 files, 1 test file deleted, 655 lines gone. No production file touched. What went, by pattern: - Duplicate invocations of a contract owned exhaustively elsewhere: three `publishDaemonEndpoint` cases that `daemon-endpoint-publish.test.ts` already covers in 20, and three daemon health classifications (`HEALTHY`, `DEGRADED`, `UNREACHABLE`) that `daemon-health.test.ts` owns. `WEDGED` and `WEDGED-HELLO` stayed — the never-resolving-RPC and never-answers-hello paths have no other owner. - Provider-local replays of a shared helper: five `GitStatusReadLeaseOwner` cases re-run per provider, owned by `src/main/git/git-status-read-lease-owner.test.ts`, and `returns the connectionId` replayed in three provider suites against an identity getter. - Assertion-free coverage probes, including one whose comment says "no writes should happen" while nothing checks that. - Copied inventories that restate a type: `PROVIDER_FRAME_CLASSIFICATIONS` is declared `as const satisfies Record<...>`, so a missing key is already a type error and an extra key fails the excess-property check. Those cases also pinned key order, which is not a contract. - A negative control that cannot fail: asserting a profile-state filename is not an unrelated literal, in a file whose first case already pins that filename positively. - Byte-identical duplicates across files, and a second case re-asserting the `unverifiable -> true` mapping the case above it already proves. `src/main/providers/ssh-git-provider-api.test.ts` goes: 52 method names asserted `toBeTypeOf('function')` plus `toHaveLength(52)` over its own literal. Note the reason, because the obvious one is wrong. "The `IGitProvider & SshGitProvider` annotation enforces this at compile time" does NOT hold — removing an operation from the interface and its implementing class in one commit still compiles. What makes the file redundant is that all 51 extractable names are referenced by some other test under `src`, so dropping an operation breaks a behavioral test anyway. The same check kept the three `registers all expected handlers` manifests in `src/relay` during the previous wave, where eleven methods had no behavioral caller at all. An inventory test is a ratchet if and only if at least one entry is pinned solely by it; that is verified per entry, not per file. Kept deliberately: everything a reliability gate cites, checked by case TITLE and not only by file path, because the gate script resolves paths only; bound, quota and provenance guards; the Windows MSYS job-breakaway and daemon-host relocation tests, which guard failures that pass every existing gate; SSH execution-boundary verdict vocabulary; and Git capability tests covering first fallback, cached call, concurrent probes and per-host isolation as four distinct risks. Coverage is partial and stated as such: of 1,449 files in scope, roughly 990 were read case-by-case and 452 received title-and-grep triage only. The unread paths are recorded for a later sweep rather than assumed clean. Verified: per-area suites green (`daemon`+`skills` 294 files/3016 cases; `git`+`persistence` 414 files/4476 cases; and the rest), gate manifest 140 gates, `check:code-quality:changed` 0 new findings. A combined 11-path local run put 1,563 files through one machine and surfaced three timing-sensitive failures in files this change does not touch (`history-manager`, `structured-agent-session-refusal-retry`, `ssh-remote-commands`); all three pass in isolation, and no production code changed, so CI's sharded run is the arbiter.
stablyai#24026) Delivery collapsed on 2026-09-29 once send volume doubled: the worker, the retention pruner and the request path share a two-connection pool, and the database transaction rate pinned at ~140/s regardless of how many notifications were delivered. Raise the pool to six so worker and pruner stop serialising on one connection. The budget precondition stays satisfied (2 x 6 x 3 = 36 <= 64). Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
…ecorded requests against the desktop's params rules (stablyai#23732) * test(mobile): add rpc:diff to decode what a recording change moved The RPC recording goldens are content-addressed JSON, so their raw git diff is pool hashes. `pnpm --dir mobile rpc:diff [<base>]` decodes both sides and prints, per golden, the checkpoint, field and JSON path that moved with both values, grouped across checkpoints, plus added and removed goldens. `--summary <file>` appends a Markdown report capped for GitHub's step-summary limit. It reads any pooled format, so it can prove the next commit's format change moves no recorded value. Checkpoints are matched by occurrence because an id can repeat within one golden. This commit adds files under the recorder directory, which moves the header digest every golden pins; the next commit removes that header. * test(mobile): record RPC goldens without a pinned commit or input digests Every golden carried a pinned `baseline` commit plus digests of the recorder, its mount adapter and its scenario, and the record script refused to run unless the product tree matched the pin. So every behaviour change repinned to its own branch commit and rewrote all ~790 files, the squash made that commit unreachable, and main's pin job stayed red until a hand-made repin pull request landed (22 of them in 12 days). The digests could only fail when an input moved and the recording did not, which is exactly the change that carries no information; every run already re-derives each golden from the current tree and compares it. Format 6 keeps the format version, operation, family, named deltas, the value pool and the recording. Removed: the pin and fence, the three digest modules and their test, the pin guard and its CI job, and the dead scenario `version` field (the manifest reader now refuses `baseline` and `version` with a message). - `pnpm --dir mobile rpc:record [<golden-id>...] [--prune]` records all or some goldens; orphans are listed, and deleted only with `--prune`. Every derived test title now starts with its golden id so an id selects it. - `compareGolden` reports every difference in one failure (identity fields by name, the checkpoint list, each checkpoint/field/path grouped), keeps the final byte compare, and ends with the command to re-record that golden. - `unhandled-recording.test.ts` now drives a detached rejection through `runRecording` into a checkpoint and the cleanup checkpoint; no golden carries one, and disconnecting the capture passed every suite before. - Seam rules that existed only to keep a digest honest are gone; the mutant-reachability, register-completeness and one-exposure rules stay. - CI: `Mobile tests on main` runs the whole mobile suite on every merge that touches mobile/, src/shared/, the root lockfile or the host RPC paths, since `verify` never runs on main. A new `Mobile RPC Recording Replay` workflow replays the recordings on pull requests that touch src/shared/ or the root lockfile without touching mobile/. `verify` writes the `rpc:diff` report to the job summary. Proof: `rpc:diff` against the parent reports no recorded behaviour moved; each golden only loses its ten header lines. * test(mobile): check every recorded request against the host's params contract The goldens script the host's replies, so a scenario could record a success for a request the real host would refuse, and a desktop change that tightens a params schema moved no golden at all. `recorded-request-params.test.ts` parses every distinct request the corpus puts on the wire with the host dispatcher's own `parseRpcRequestParams` and the schema `rpc-params-catalog.generated.ts` binds to that method. It fails on a method the host lacks, params it refuses, params sent to a method that takes none (the dispatcher never reads them), and keys the schema silently strips unless an inventory entry gives the reason; a stale entry fails too. Each rule is also shown firing on a made-up request, since the corpus has no instance of three of them. It imports the desktop dispatcher, so it sits beside the other Node-side tests outside the RN test program, and the params-contract boundary now exempts test files, which are never bundled. It found twelve requests the host would refuse, all from invented fixture values, not product code, fixed at their source: - git.branchDiff sent `base-oid`/`head-oid`/`merge-base` where the host needs full object ids (diff-review and source-control adapters, and the branch compare replies in the manifest that feed them); - an iOS push registration without `apnsEnvironment`, which a real iOS token always carries (`push-token.ts`); the adapter now defaults to `production`; - `settings.update` given Linear's `assigned` filter as a GitHub preset, which the product type forbids; the scenario now picks `my-issues`; - GitLab `projectRef` as a string where the host and the product type take `{ host, path }` (7 methods, 5 adapters and the manifest). 46 goldens move, and a decoded comparison of every one of them shows no change other than those substitutions; `rpc:diff` lists them. * ci(mobile): detect a mobile change without a SIGPIPE-prone grep pipe Under the runner's pipefail, grep -q exiting on its first match SIGPIPEs git diff on a long file list, so a large pull request touching mobile/ read as uncovered and replayed the recordings a second time. * test(mobile): drop comments that still describe the golden header and digests Eleven adapters justified an import rule by the header a golden no longer carries, and that rule's test is gone. The census failure now names the rpc:record and --prune commands. * test(mobile): refuse a golden that keeps a key no recording writes Decoding dropped unknown top-level keys, so an old header left behind by a hand-resolved merge conflict passed every compare unseen. * ci(mobile): summarize RPC recording changes after a failed test step too * test(mobile): stream rpc:record output instead of capturing it A captured run stayed silent for its whole duration and clipped its tail, where the failure summary sits, past 8 MB. * test(ci): let the Ruby-gate contract skip the always-run RPC summary step fef088d gave the summary step an `if: ${{ !cancelled() }}`, and this test lists every gated step in `verify` and expects each to be gated on the Ruby scope. * test(mobile): replay only a golden file that is exactly what rpc:record writes Replay compared two re-encodings of decoded values, so anything decoding drops (a leftover header key, a hand edit) sat in the committed file uncompared; a key allow-list covered one case of that. Replay now passes only if the file text equals the formatted golden for the run, sharing one formatter with writeGolden, and keeps the field-level report as the failure message. The allow-list goes; the value-based compareGolden stays for the bridged run, which has no file. * test(mobile): end a corrupt or hand-edited golden's failure with the re-record command A hand edit to a pooled value failed in decode with only "Golden value <hash> does not hash to its pool key": no golden id and no command to fix it. readGolden now prefixes parse and decode failures with the golden id and ends them with the rpc:record command. The rpc:diff header also said it always exits 0; it exits non-zero when git or a golden cannot be read, and now says so. * ci(mobile): run Mobile Checks on every src/shared and root lockfile change Replaces the replay-only workflow: mobile imports hundreds of shared modules, so a shared edit can move a golden or break mobile's typecheck, and the full job catches both before merge. A root lockfile-only change skips the Ruby release checks, which read no root Node dependency. * test(mobile): list or prune orphaned goldens even when the recording run fails Orphans come from the manifest, not the run, so a failed or timed-out rpc:record still reports them; the exit code stays non-zero. README: say what a failed replay reports (first differing path per field, capped groups) and what rpc:diff compares with and without a base. * test(mobile): pin the RPC recording goldens to LF so a CRLF checkout still replays Replay now requires the committed golden text to equal exactly what rpc:record writes, which is LF. A Windows checkout with core.autocrlf=true converted every golden to CRLF and failed all 790 with "holds the same recording but is not the file rpc:record writes for it".
…lyai#24027) resources/relay is the only relay copy a packaged build resolves, but out/relay was also packed into app.asar — 14.2MB of unreachable duplicate. Kaspersky flagged app.asar as a compound object precisely because relay.js was inside it, so one script-heuristic verdict on relay.js gutted the whole install. Excluding it decouples app.asar from that verdict and drops the duplicate bytes.
stablyai#24033) stablyai#23900 probes codex --help before each launch; the fake counted it as a worker spawn, breaking two specs.
…test harness (stablyai#23986) * test: open, seed and read the agent-session record store through one harness Tests that open the durable agent-session record store, seed it, or read back what it persisted now go through agent-session-record-store-test-harness.ts instead of calling AgentSessionRecordStore.open or touching agent-sessions.json themselves. A later change that moves the store into the chat database then changes the harness instead of every test. No production code changes. Tests whose subject is the JSON file itself (its .bak recovery, salvage, schema versions, permissions, and what older builds read back) keep reading and writing the file directly; the storage move rewrites or deletes them. * test: address the record-store harness by the host's state directory The harness took the store's own folder, so each caller picked one (join(root, 'store'), or 'agent-sessions' where a test read the store the runtime owns). A later change that moves the store into the state directory's journal database could not tell those apart, and would have had to edit every caller again. Every harness function now takes the state directory, the one the test's journal database and recovery capsule already live in, and keeps the store in the same subfolder the runtime uses. Callers pass that directory; store-only tests pass their temp directory unchanged. Format tests that share a directory with harness calls take the file path from testAgentSessionStoreFilePath. The folder name moves from a private constant in the runtime to AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness shares it without importing the runtime. Its value and every path built from it are unchanged.
Four drains, each holding one provider round trip of ~100 ms plus its database statements, capped delivery near 30/s. Production inflow reached 35/s on 2026-09-30, so the backlog aged past the five-minute TTL and notifications expired. Twelve drains lift the ceiling to roughly 90/s; the pool is now six per instance, so the extra drains queue on connections instead of starving the request path. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
Correct typo in PR template regarding issue linking for outside contributors.
… in the parent's conversation (stablyai#23752) * refactor(native-chat): a subagent's rows live with that subagent, not in the conversation A subagent's rows were drawn in its parent's conversation, each captioned with the subagent's name. They now belong to the subagent: the transcript projection keeps the session's own rows as the conversation and each subagent's rows apart, keyed by the agent id its roster entry already carries, folded on their own. Desktop: a subagent's rows open in a section under the roster row that names it, from that agent's roster entry, and are windowed like any other rows. A subagent no loaded roster names opens where its first row happened, inside the section of the agent that spawned it or in the conversation. Its edits still count in the turn they were made, and revealing one opens the sections around it. Mobile shows the conversation, with each spawn's roster line. Worker reads and structured terminal reads serve the worker's own rows. Removes what the move makes redundant: the per-row caption and its copy, the producer check in the tool fold and the turn answer, the per-agent frontier interleaved in the conversation, worker-text subagent tags, and the agent id on worker-read messages. * refactor(native-chat): a diff target names the sections its row sits in Revealing a subagent's edit opens the sections around it from the target the rollup already holds, instead of looking the row up at click time. The section head keeps to the agent's name and dot; its state in words stays on the roster entry. The worker page test stubs the host through its module rather than a cast. * fix(native-chat): a working subagent's section is open; a worker page windows its own rows A subagent's section is open while its agent works and closes once it settles, the way the turn's own live run does; a section the reader opened or closed by hand keeps that choice. A subagent another subagent spawned opens inside that one's section, so a working grandchild shows inside its working parent. Openness is derived from the roster's state and the reader's choices; nothing stores an automatic open. A worker page is now the newest page of the worker's own rows. The host windows the read over them before the limit, so a subagent's burst can no longer crowd the worker's rows off the page, and "older" still means older worker rows. The scope is an in-process argument of the host's history read; no wire request carries it. * fix(native-chat): a subagent section head names the turn it sits in, for the outline rail * fix(native-chat): a subagent section's rows sit in the turn the section is shown in, for the outline rail A background subagent's rows written during a later turn carried that later turn onto their slots, so scrolling through its section lit the later turn's rail tick and then snapped back. The rollup still counts each edit in the turn it was made; only the slot, which the rail reads, takes the shown turn. * fix(mobile): Load earlier reads past pages that hold only a subagent's rows Mobile draws only the session's own rows, so an older page made entirely of a subagent's rows landed as nothing: the reader tapped Load earlier, saw the spinner, and got the same transcript back. One load now reads on (up to 8 pages) until a page holds a row of the session's own, then applies the pages in order. * test(mobile): stub the RPC client the way the other structured-session hook tests do * perf(native-chat): order subagent rows for the changed-files rollup once per change to them The rollup flattened and re-sorted every subagent row on each update, including every token the parent streamed. The ordering now keys on the projection's subagent rows, which keep their identity while only the conversation changes. * refactor(native-chat): order subagent rows in the sections hook, keeping the list under its line limit * fix(mobile): a transcript whose newest page is only a subagent's rows reads back on its own Opened while a subagent is busy, the newest page can hold nothing but that subagent's rows. Mobile draws none of them, so the reader saw an empty chat with a Load earlier button, and an empty list cannot be scrolled to page. The hook now reads back once from each such head, and the read runs on to the session's own rows. * fix(native-chat): count the live window in the session's own rows, so a subagent's burst keeps its roster The live window kept the newest 1,024 rows of every agent. A subagent writing more than that trimmed its own spawn's roster row and the prompt, and its section fell back to a closed, unnamed header. The window now keeps the newest 1,024 of the session's own rows and everything after, with an 8,192-row cap on every agent's rows as the memory backstop. A transcript with no subagent rows trims exactly as before. * fix(agent-session): window history pages by the session's own rows, with a subagent's rows riding along A history page held the newest 200 rows of every agent, so a subagent's burst could fill a page on its own: the phone opened on an empty chat and "Load earlier" landed nothing. A page now starts at the oldest of the newest `limit` rows of the session's own and serves every row from there, so the subagent's rows come with the conversation they happened in. The page stays contiguous, the cursor still names its first row, and the byte bound still applies. A transcript with no subagent rows gets the same pages as before. Clients already take a page larger than its limit: both reducers raise their retained window to the page's size. The mobile read-on and read-back stay for older hosts. * test(agent-session): a page reaches back to the start rather than leaving a subagent-only page * fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it The live window trimmed to just after the own row it dropped, so a subagent whose roster row went kept its rows at the top as an unnamed section until the parent wrote again. Trim to the oldest own row kept instead; it still fires only once an own row passes the limit, so a paged-in run of subagent rows at the head stays until then. With no subagent rows nothing changes. * perf(native-chat): cap the live window at 4,096 rows, bounding each delta's re-derivation Every live batch re-derives the transcript over every retained row. On the largest real window (7,374 rows) that cost 7-8 ms a delta on desktop against 0.6 ms at the old 1,024-row window, and held about 26 MB of row content. 4,096 halves both. The most rows any local journal puts between a roster and its subagent's last row, with the parent inside its own-row limit, is 3,005, so no observed subagent loses its roster to the lower cap. * fix(native-chat): a subagent section opens only while its roster is the running scope's live frontier A section used to open whenever its roster said the subagent was working, anywhere in the transcript and whether or not the session was running, so a background subagent's section stayed open and grew mid-transcript while the parent moved on. It now opens by default only while the session runs and the roster row naming the subagent is the newest thing the parent produced, user rows aside. Newer parent output closes it even while the subagent still works; the roster row keeps showing that live state. A subagent still working is a running scope of its own for the sections it spawned; a settled one closes its scope. Derived every render, no latch; the reader's own open or close still wins. * fix(native-chat): name a subagent's section from a client roster the window never trims A section took its name and state from a roster row in the loaded window. Once a burst trimmed that row, or the row sat on an older page, the section fell back to an unnamed, closed "Subagent" header. The shared reducer now keeps a roster keyed by agent id, folded from every roster row and revision the client receives: pages, older pages and live batches, including revisions of roster rows outside the window, which live batches already carry. The first roster naming an agent wins and its revisions update it; a removed roster row drops its entries; it is rebuilt on every page that replaces the window and bounded to 512 agents. Sections take their name, state and live-frontier place from it; placement stays under the loaded roster row, else at the section's first loaded row. Only a subagent no roster ever named stays unnamed. * feat(agent-session): a history page names the subagents whose roster row is older than it A page is a contiguous run of the journal whose older-page cursor is its first item, so it cannot pull an older roster row in without skipping the rows between. When a page held a subagent's rows but not the roster row naming it (about 11% of the moments a reader could open a session on local journals), that subagent drew as an unnamed "Subagent" header. History and hydration pages now carry an optional `subagentRoster`: the first roster entry naming each subagent whose rows are on the page and whose roster row is not, with the row's id, sequence and revision; bounded to 64 entries and 16 KB. Items and cursor are unchanged. The client seeds its roster from it. Rule 1 in docs/reference/remote-wire-compatibility.md: an optional field on an existing frame, no capability gate. An older client ignores it (the released reducer reads a page with it exactly as one without); against an older host the field is absent and the section falls back to an unnamed header. * Revert "fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it" This reverts commit 2207865. Its only purpose was to stop a subagent whose roster row an own-row trim had dropped from showing at the head of the window as an unnamed section. The client roster now names that section whatever the window holds, so the cut is back at just after the own row the limit passes. The retention test that pinned the unnamed-section case now asserts the section at the head keeps its name. * chore(native-chat): state the retention limits' own reasons, now that no name depends on the window Own-row retention keeps the conversation a reader sees from being crowded out by rows drawn as a one-row section on desktop and not at all on mobile; the every-agent cap bounds memory and each live delta's re-derivation. Neither is about keeping a roster row loaded any more. * fix(native-chat): hold the roster fold's draft map where type narrowing can see closure writes * fix(native-chat): a roster row's newer revision replaces it in the client roster too A revision that stops naming an agent (the host drops an entry it learns is not a subagent, or re-keys a provisional one) left the client roster holding the old entry, often still "working", with nothing to re-derive it. The section then read as working forever and could auto-open, while a fresh read of the same journal left it unnamed. The fold now drops an entry when a newer revision of the row that named it no longer does, before any roster takes it over. * fix(native-chat): a parent's spawn and wait calls keep the subagent they name open A subagent section auto-opened only while its roster row was the running session's newest row, so any later row closed it: a Codex wait on the agent, or the parent's text before its next spawn call. Now a row that is part of delegating to a subagent keeps that subagent open: - a Codex collab call (spawn, wait, resume, message, close) opens each agent its receiver thread ids name; one naming none is ordinary output; - a Claude spawn call names no agent, so it counts toward the roster announcing it; - a roster row at the frontier opens its most recently added agent, not all of them. A roster or call naming only agents one subagent spawned is that subagent's output, so a grandchild's roster, which the host journals as the session's row, no longer closes the spawner's section. * fix(native-chat): a parent's call right after the roster closes its subagent's section A parent's tool calls after a roster row fold into the tool run drawn above the roster, so the roster stayed the newest drawn row and its section stayed open while the parent was already reading or running commands. The fold now records the newest journal position among the rows it merged, and the live frontier orders rows by that newest part. The layout is unchanged. A spawn call folded there still counts as part of the roster announcing it. * fix(native-chat): a Codex call naming several subagents delegates to the first A Codex collab call that names several agents opened every one of their sections. It now counts as delegating to the first agent it names, so one section opens, the same as a call naming one agent. * fix(native-chat): closing a roster's list closes the sections under it Collapsing a roster row's list of subagents left their open sections drawn, so the section's own head became the only way to close them. And the list's open state lived in the row, so a row the window unmounted came back collapsed. The transcript now holds each roster list's open state beside the section choices. A closed list hides every section it anchors; each section keeps its own open or closed choice for when the list reopens. With no choice from the reader, a list is open while a section under it is open. Closing a section from its entry keeps the list open, and revealing a subagent's edit opens the list it sits under. * perf(native-chat): a reveal finds the roster lists it opens with one set lookup per entry * fix(native-chat): a subagent's roster entry heads its own rows An open section drew the agent's name twice: its entry in the roster's list, then a separate section head above its rows. The entry is now the head. The roster row draws its entries through the first open one, that agent's rows follow, then the entries after it, each run in its own windowed slot. A section no loaded roster row holds (an older page, a grandchild, an unnamed agent) keeps its own head. A roster list is open while the live frontier or a reader's choice is on one of its agents, unless the reader closed the list, so closing an agent from its entry no longer needs to pin the list open. The section emitter moves to its own module, and the trailing-run predicates it shares with the slot builder to theirs, to keep the slot builder under its line limit. * fix(native-chat): the entries after an open subagent's rows set in its roster's type The roster row's list inherits the system row's small muted type; the entries that follow an open section sit outside that row, so they now carry the same type. * docs(native-chat): a current host can also serve a page of only a subagent's rows A page is bounded by bytes after it is windowed by the session's own rows, so a burst that fills the bound yields a page, or an opening page, with none of the session's own rows. Mobile's read-on and read-back therefore serve current hosts too, not only older ones; the comments said otherwise. The retention comment still described a closed section as a row of its own; it now sits behind its roster entry. * fix(native-chat): a section's prose keeps its copy/timestamp controls inside the section An assistant row's hover controls (copy, scroll-to-top, timestamp) hang 20px below the row into the gap before the next one (`-mb-5`). Inside a subagent's section that put them below the section's left border, and on the section's last row they touched the parent's next row with no gap. Inside a section the controls now stay in flow, so the border covers them and the next row sits the normal gap below. The row-height estimate reserves the same 20px for a section's prose so windowing does not jump on measure. * test(agent-session): state each appended row's turn scope, as the journal now requires * refactor(native-chat): the client's journal retention policy lives in its own module
…er runs journaled (stablyai#23758) * fix(claude): a subagent resumed after a restart keeps its canonical id A backgrounded Claude subagent resumed by a message after its session's provider restarted parents its frames to its ORIGINAL spawn call, while the announcement the new provider sees names only the resuming call. The spawn call's alias lived only in the old provider's memory, so every row the resumed child wrote fell through to its raw call id: one child shown as two, a named roster entry with no rows and an unnamed section holding them. The alias table now also recalls what an earlier run of the session resolved, re-derived from the agent rows it journaled (canonical id beside the call its frames arrived under), read once per bound journal epoch. Nothing new is persisted. * fix(claude): a restarted provider continues the subagent roster earlier runs journaled The roster's state lived in one provider process while everything it writes is the session's. After a restart a resumed child was re-rostered in a second group row with its attempts restarted, its resumed frames lost the alias they still carry, and the one group row no turn owns was rewritten from empty, erasing the children an earlier run had listed there. A run now reads what earlier runs journaled, once per bound journal epoch: group rows give each child's entry and group, agent rows give the spawn-call aliases and the latest attempt. A group is inherited only when this run's own events reach it, and inheriting writes nothing. An inherited child is reopened by an announcement exactly as an in-process resume reopens it, and by nothing else. The alias recall is one facet of that read. Nothing new is persisted. * fix(claude): an earlier run's subagent takes Claude's restart verdict in one row At restart Claude reports each agent the previous session left running as stopped ("didn't finish before the previous session ended"). The roster ignored it, leaving the entry unverifiable, while the background-task lane, which never saw that agent announced, opened a second, unnamed row for the same agent. The roster now records the verdict on the inherited entry, keeping the time the earlier run lost contact, and still reopens the entry on an announcement. The background-task lane leaves any agent an earlier run rostered to the roster, from the same journal-derived reading. * fix(claude): a restart verdict on a child whose host died invents no stop time A journal reopened after its host died leaves a child unverifiable with no stop time. Stamping Claude's later restart verdict with the current time would show the whole outage as how long the child ran, so the earlier run's stamp, or its absence, is kept. The journal-liveness comment no longer describes the roster as unable to continue from the journal. * fix(claude): a child two journaled rows list stays live in one of them An older build could re-roster a resumed child in a later turn's row, so two group rows list it (67 children in 39 local sessions). Inheriting the second row re-pointed the child to that row's copy, so the copy the resume had reopened was left at working and swept to unverifiable, while the outcome landed on the stale copy. The first row reached keeps the child. * fix(native-chat): no run length for a settled child with no stop time A subagent whose host died without sweeping it has no stop time. After a restart, Claude's own verdict on it ("stopped") is now recorded, but the rule that hides the group's duration only covered `unverifiable`, so a mixed group showed the sibling's duration as the whole group's. The rule now keys on the missing stop time, whatever the settled state. * fix(claude): a twice-listed child resumes in the row the journal reading chose A child an older build listed in two rows was placed in whichever row this run's frames reached first, so a sibling's frame reaching the older row made the resumed child reopen there. Placement now follows the reading's one tie-break, and evicting a row drops only placements that point at it. * fix(native-chat): a resumed subagent's clock never counts the idle gap A reopened Claude child now starts a new run: its startedAt is reset to the reopening frame's time on every reopen path (resume after a restart and a same-process reactivation). The group clock reads the union of the children's latest runs, so an idle gap between runs is never counted; an ordinary overlapping fan-out reads the same as before.
…i#23768) * fix(editor): keep saved untitled notes when their tab closes Saving cleared the draft and dirty flag, so close-time cleanup took a saved note for an untouched placeholder and deleted it. Fixes stablyai#23688 * fix(editor): decide untitled-note cleanup from disk state, not a per-save flag Drop setUntitledFileHasSavedContent and its save-queue hook so deleteUntouchedOnClose keeps its creation-time template meaning. A clean untitled tab now counts as an empty placeholder only while its last-known disk content is empty or not yet loaded; the on-disk size check still gates every actual delete. Notes an agent filled while their tab was showing now stay reopenable with Cmd+Shift+T. Also test both Don't Save dialogs end to end, mirror the shared discard in the floating panel test mock, and remove markFileDirty selectors left without consumers. * test(editor): wait for untitled cleanup before asserting disk state Drain the stat-then-delete chain before the untitled-note tests assert the file survived, so the in-flight Don't Save case fails if the size check is removed. Type the main-window close dialog hook by the controller fields it reads, which lets its test drop a type assertion. --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
…tablyai#24010) * fix: keep active tab visible by docking to viewport edges Makes the current tab easier to locate in many-tab scenarios. The active tab now sticks to a viewport edge via sticky positioning when it would scroll out of view, with a full-foreground indicator bar for better visibility and arrow animation when a background tab opens off-screen. * fix(tab-bar): reveal offscreen tabs instead of nudge animation When a background tab opens beyond the visible area, automatically scroll to reveal it (unless hovering the tab strip). This replaces the previous arrow-nudge animation with direct visibility. revealTabStripElement now handles keeping the active tab visible alongside the revealed tab when both fit, or docks the active tab when needed. * fix(tab-bar): reveal tabs by identity, not count increase alone Detect opened tabs by comparing tab identities independently of count changes. Newly opened tabs are now revealed even when the total tab count stays the same—e.g., when a tab closes as another opens. * fix(tab-bar): track tabs by identity for reliable reveal on open/close Replace count-based tab detection with identity tracking so the strip correctly reveals tabs when they're added, replaced, or when the active tab closes and switches to a far-back history tab. Removes the tabCount parameter and simplifies overflow navigation by using identity sets. * fix(tab-bar): defer revealing tabs until pointer leaves When a background tab opens while the pointer hovers the tab strip, defer its reveal until the pointer leaves. This prevents the active tab from sliding away mid-interaction. Also support client-hosted rows taking active state while maintaining tab dock positioning.
…ai#24040) * fix(push): size the claim-attempt budget from the drain count With twelve drains, up to eleven peers can hold device heads, so a four-attempt claim budget can run out while claimable rows remain and the drain exits idle for a tick. Move the drain count into one module and derive the attempt budget from it. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * style(push): keep the worker and store in repo formatting Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * style(push): drop the stray semicolon in the concurrency constant Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
… settings (stablyai#23979) * fix(agent-trust): write per-user trust under the home the launched agent reads The Cursor, Copilot, Qoder and Antigravity writers and the local Codex config list resolved ~ with os.homedir() at write time, so any test that reached them wrote into the developer's real ~/.codex, ~/.cursor, ~/.copilot or ~/.gemini. Each writer now takes the home, derived once from the launch env (HOME, or USERPROFILE on Windows, else this host's home) by launchedAgentHome, which the SSH relay already used. * test: give tests that wrote the real agent or Orca home a temp one The structured Codex adoption replay pre-trusted /repos/workspace-1 in the real ~/.codex/config.toml; it now runs with a temp HOME and userData. The Codex session-resume and WSL hook tests created Orca's managed Codex home under the live userData, and the Claude Agent Teams tests wrote their tmux shim into ~/.orca; each now runs against a temp userData or HOME. * test: fail any unit test that writes the real agent or Orca home A vitest setup file wraps the node:fs mutating calls and refuses a target under the account's real ~/.codex, ~/.claude(.json), ~/.orca, ~/.cursor, ~/.copilot, ~/.gemini, ~/.qoder or Orca userData, found through os.userInfo() so a test that swaps HOME cannot hide it. The refusal is recorded and rethrown after the test, since trust writers swallow errors. Reads are untouched. It stands down only while an opted-in real-agent suite's own switch is set. It also unsets what an Orca terminal exports toward the live app (userData, Codex and Claude homes, and the Codex launch preflight CLI, which a shell test would otherwise run), so a local run matches CI. * test: type the guarded fs call from its narrowed original
…tablyai#24043) Completes the `src/main` audit. Sweep over `src/main/runtime` (flat, rpc, orchestration, relay, push) and flat `src/main`. 42 case declarations removed across 24 files, 783 lines gone. No file deleted whole, no production code touched. Highest-yield area by far was `runtime/orchestration` (21 cases from 151 files); the rest of runtime measured under 1%. What went, by pattern: - Tests whose subject is the test itself: a source-scanning boundary test with no production import at all, three of whose cases checked its own regex against strings it declares; and a benchmark whose own simulation contains the short-circuit it asserts, with `expect(wouldHaveBeen).toBe(60000)` comparing the test's own arithmetic. - Identity copiers: seven db cases where every asserted value is the literal input (`type: 'question'` in, `type === 'question'` out). - Permanently skipped tests for behavior that does not exist — two cases carrying `// TODO: inline restore on re-subscribe not yet implemented`. A skipped test for an unimplemented feature can never fail; it is a note in test syntax. - Duplicate invocations of a contract owned at a stronger boundary, including three reset scopes and two dependency-promotion cases owned by dedicated suites. - Registration manifests whose every entry is referenced by other tests. - Table rows and cases varying a field production never reads: a mobile tab-restore case varying `clientCapabilities`, which the mobile branch does not consult, and a Windows worker case in a module with no platform input at all. - Names promising more than the input exercises: a case titled for forged AppImage variables whose body only removes `AppRun`, byte-identical to a row of the `it.each` table twenty lines above. - Negative controls passing for an unrelated reason: a foreign-pane rejection whose fixture also differs in sender handle, so the handle guard can reject it. - Self-comparisons, including `format(m, { authority: 'current' }) === format(m)`. Two deletions were justified by the wrong argument and kept only after checking a better one. "A generated-catalog check gates registration" is false: that script prevents the catalog and dispatcher from drifting apart, so a developer who removes a method regenerates the catalog and the check passes. Like a type annotation over an interface and its implementing class, it verifies internal consistency and cannot see a declaration and its use removed together. Both inventories go on per-entry evidence instead — every name is referenced by other tests. The same rule kept a 37-entry terminal-method inventory in the same wave, because some of its entries are pinned solely by it. An inventory is a ratchet if and only if at least one entry is pinned solely by it; that is evidence per entry, not a verdict by shape. Kept deliberately: everything a gate cites, checked by case title and not only by file path; source-scanning ratchets that pair their negative assertion with a positive one against real source (the surviving boundary case asserts the pattern still matches the writer module, so a silently-broken regex goes red); a destructive-delete PTY waiver guard; `it.fails` markers, which go red if the bug is fixed; and `it.skipIf(platform)` cases, which do run on other hosts. Coverage is partial and stated as such: 1,195 files in scope, roughly 660 read case-by-case, the remainder title-scanned and mechanically triaged. Unread paths are recorded for a later sweep rather than assumed clean. Verified: `check-reliability-gates.mjs` (140 gates), `check:code-quality:changed` (0 new findings). No `pnpm tc` needed since no production file changed. A local run of `src/main` surfaced 42 failures in three files this change does not touch (`browser-manager-tab-identity`, `browser-manager-viewport-ownership`, `session-scanner-codex-workers`); all 42 reproduce on a pristine `origin/main` worktree, so they predate this wave and are environment-dependent locally — CI was green on main at `2d85fdc753e2`.
…pair Orca's duplicates (stablyai#22592) (stablyai#23958) * fix(codex): recognise Codex's quoted project-trust spellings in config.toml (stablyai#22592) Codex's settings screen writes project trust as ["projects"."/p"] and "trust_level" = "trusted". Orca's matchers only knew the bare spelling, so a trust write appended a second [projects."/p"] table (or a second trust_level line) and every codex command then failed with "duplicate key". The config mirror kept both spellings in Orca-managed homes for the same reason. - Project table headers are now read through the existing TOML key-path parser, so bare, quoted, literal-quoted, mixed and spaced spellings are the same table for trust writes and the managed-home mirror/dedupe. - trust_level is found by decoded key, in both the trust writer and the mirror's trust reader, and an existing key is rewritten, never duplicated. - On the next trust write, a table older Orca appended (exactly [projects."<p>"] holding only trust_level = "trusted") that duplicates the user's table, or the bare line it inserted under a quoted "trust_level", is removed; the user's table wins and the atomic writer keeps config.toml.bak. Any other duplicate, or a repair that would still leave one, leaves the file untouched and logs once. * build(cli): list the new Codex trust modules in the CLI project * fix(codex): recognise Codex's quoted hooks.state spellings and repair Orca's copies (stablyai#22592) Codex writes hook trust as ["hooks"."state"."<key>"] (and the parent as ["hooks"."state"]). Orca's hook-trust writer, parent-table check and mirror only knew the bare spelling, so a hook-trust write appended a bare copy and the file failed to parse with "Cannot declare ... twice". - The hooks.state header, parent-table and mirror checks now use the TOML key-path parser, like project tables. - The duplicate repair now also removes Orca's own hooks.state tables (an exact [hooks.state."<k>"] with only enabled + trusted_hash, or an empty [hooks.state]) that repeat a table in another spelling, and runs on hook trust writes too, so a file with both project and hook duplicates is fully repaired. The Orca-shaped copy is removed whichever order the two tables are in, only when exactly one other table (the user's) remains; anything else is left untouched and logged once. * fix(codex): carry plain-Codex plugin and project hook trust into Orca's Codex homes (stablyai#22592) Codex keeps hook trust in $CODEX_HOME/config.toml under hooks.state, keyed by the hook's source. Plugin keys (`id@mkt:path`) and project keys (`<repo>/.codex/...`) are the same in every home, but the mirror dropped every hooks.state table from ~/.codex, so Codex inside Orca asked users to re-trust plugin and project hooks they had already trusted in plain Codex. - classifyHookTrustKey splits keys into home-scoped (the home's own hooks.json/config.toml, re-keyed by install as before) and shared. - The mirror now carries shared hook trust from ~/.codex in every spelling. A key the managed home already holds keeps the managed copy, a key repeated in ~/.codex is carried once, and the parent [hooks.state] table is never copied, so the result never declares a table twice. - mergeSystemCodexConfigIntoRuntime moves to codex-config-mirror-merge.ts to keep codex-config-mirror.ts under the line limit. - Tests cover plugin/project carry in each spelling, user-hook keys staying out, repeated launches, managed-copy precedence, Windows key spellings, parent tables, and user-hook trust re-keying (trusted_hash and enabled) from every ~/.codex spelling. * fix(codex): carry session_end and interrupt hook trust into Orca's Codex homes (stablyai#22592) The shared-trust classifier parsed hook keys with Orca's own trust-key parser, which only knows the ten events Orca installs hooks for. Keys for Codex's session_end and interrupt events did not parse, so their plugin and project trust was treated as home-scoped and left out of the managed home. The classifier now reads the source path from Codex's key shape `{source}:{event}:{group}:{handler}` for any event label. A key without that shape is still never carried. Tests cover both events for plugin and project keys in both spellings, user-layer keys for both events, and five unattributable key shapes.
Group 10 is full; swap the README QR code and copy (all locales) to the new group 11 invite.
…tle (stablyai#23948) * fix(sidebar): keep hook-less agent rows while the agent runs, whatever its title Codex retitles its pane to the project name, so the sidebar's title-derived row (which required the title to name an agent) vanished while Codex kept running (stablyai#23767). Rows now take identity from the canonical pane resolver over the pane's foreground-process read and launch record, then the title; the title only decides idle/working/needs-input. The row still goes away when the PTY exits, the process tracker proves the shell is back, or the title is a shell or default title. * test(dashboard): justify the partial store fixture's type assertion * fix(sidebar): only a live process read keeps a plain-title agent row Review of the previous commit found ghost rows: the tab launch record is a latch nothing clears on WSL, after an SSH exit, or for a launch that never started, and a parked pane's process read went stale because only the mounted tracker re-derives it. - The launch record returns to main's role: a fallback only for titles that show activity, ranked below a title naming another agent (pane reuse), matching the tab icon's order. - A parked pane's command boundary retires its unconfirmable process read, like the mounted ladder's unavailable path; reveal re-reads it. * fix(sidebar): confirm before a parked marker retires an agent; read Git Bash prompt titles as the shell - A parked pane's end-of-command marker can be a nested shell's leak under a still-running full-screen agent, so confirm the foreground first (as the mounted ladder does) and retire the process read only on a shell or no answer. SSH/remote parked panes hold no incarnation to fence a host read with, so they still retire. - Git Bash emits no command marks; its `$MSYSTEM:$PWD` prompt title (MINGW64:/c/repo) is now shell evidence, so a stale Codex read there no longer keeps a ghost row after Codex exits. * fix(sidebar): trust only process-read agents for plain-title rows; per-worktree foreground selector groups by tab A daemon reattach seeds the pane's foreground entry with its launch agent, which can outlive the process while Orca is closed. The entry now records where its agent came from (agentEvidence), and the sidebar/dashboard title-derived rows only keep a plain-title row on an actual process read. Routing and the tab icon are unchanged. selectPaneForegroundAgentsForWorktree grouped every pane key per worktree; it now groups by tab once per map identity and skips worktrees with no tabs. * fix(sidebar): a parked pane's reattach keeps its own process read of the same agent The reattach seed marked a returning parked Codex pane as launch-record evidence, over the process read this session already took, so its row blinked out on reveal and stayed hidden if the user left the tab before the visible read landed. Keep the read when it names the same agent; the seed still drops byte-routing trust. * test(terminal): foreground confirmation publishes process-read evidence * fix(sidebar): a cleared pane title retires the agent's process read Codex clears its title when it exits, and the tab then shows its default title. A pane without shell command marks never re-reads its foreground process, so the retained read kept a "Codex · Idle" row after /quit (permanently for a hand-typed Codex; about 15 s while the marked-pane confirm ladder ran). Treat a blank title like the default title it shows. * fix(sidebar): the pane's process monitor retires an exited agent's process read A hook-less pane keeps its sidebar row from the tracker's foreground-process read, but nothing re-derived that read in a pane without OSC 133 command marks. After Codex exited there, a "Codex · Idle" row stayed: permanently when the shell titles its prompt, or when a killed Codex leaves its last title. The pane's agent-completion process monitor already confirms an agent's exit (no agent and no child processes, held past its settle window). It now reports that exit to the tracker, which retires its own process read and runs the confirmed-shell path the visible-pty read uses. A tracker read that names an agent seeds the monitor, so hidden panes and panes the monitor had not polled yet are watched too. A command read in flight still decides the pane, and launch records or other agents' reads are left alone. * fix(sidebar): a monitor-confirmed exit leaves the next agent in an unmarked pane identifiable The process-exit retire published shellForeground:true and left the one-shot visible sample settled; a pane without command marks has no command start to lift either, so a Codex typed again after quitting was never read and lost its row on retitle. Publish shellForeground:false and reopen the sample. * test(terminal): justify the pane binding cast in the process-exit relaunch test
… agent.launch (stablyai#22954) * fix(mobile): start + menu, quick command and diff-note agents through agent.launch The session screen's + menu, agent quick commands and diff notes' New agent session now ask the host to start the agent with agent.launchReplay, so the host picks chat or terminal from the desktop's default and delivers any prompt. Hosts without the launch capabilities keep today's paths. The phone's pending tab choice is one value (a tab, a terminal by handle, or a launched surface) instead of two refs, and a launched chat is found by its session id in the next snapshot rather than a predicted tab id. A launched surface waits a bounded number of snapshots for its tab. * test(mobile): add the + menu and diff-note launch scenarios to the recording corpus * test(mobile): repin bridged-parity tallies for the four launch goldens; drop test casts The corpus grows from 790 to 794 goldens; all four new ones replay identically. * fix(mobile): show a refused agent launch as a toast beside open tabs The inline create error renders only in an empty session, so a host refusal (for example a disabled agent) from the + menu in a session with tabs showed nothing. Always toast the failure: the caller's own copy when it gave one, otherwise the host's reason. * test(mobile): type the launch reply helper with the shared launch outcome types * fix(mobile): record a launched agent's tab as this device's pick on the host A launch carries no navigation, so the phone selected the new tab only locally while the host kept this device on the tab it had before. Leaving the session and coming back, or a reconnect that reset the screen, reopened that old tab. The "+" terminal path this replaced asked the host to select the tab for the caller. When a launched surface's tab lands in a snapshot, activate it for the caller exactly as a tap does. The resolver now names the landed tab in place of the unused `missed` flag. Route parity re-pinned for the new activation body, identity payload and strings. * fix(mobile): land on a launched agent's tab without a 500 ms wait or a blank pane The host publishes a launched tab before it replies, so the tab list the phone already holds usually has it by the time the reply arrives. The launch paths still waited for a refetch 500 ms later, leaving the phone on the old tab for that long after every launch. Read the tab list at once. On hosts without agent.launch, the chat path also unsubscribed the open terminal and cleared its handle before the chat's tab landed, while the old terminal tab stayed selected: a blank pane until the next tab list. Leave the open tab live until the chat lands, as the launch path does; applying that tab list tears the old terminal down. Route parity re-pinned for the two bodies. * fix(mobile): keep a tab the user picked while a prompted launch was still replying A quick command or review-notes launch now waits for the host to deliver the prompt, which can take up to a minute. The launched tab shows up in the tab row well before that, so a user who tapped another tab meanwhile was pulled back onto the launched one when the reply arrived, and that pick was recorded on the host. The launch now remembers which tab the phone was on when it started and only takes focus if the phone is still there when the reply lands. Any move made in between, by a tap or by the computer navigating this phone, wins. Session route parity re-pinned for the handleCreateTerminal body only. * fix(mobile): name a launched agent's tab before asking, and land on it when it is listed A "+" menu, quick-command or review-notes launch now reserves its tab before it asks the host: a fresh pane key (tab and leaf UUIDs) and, for an agent the host may start as a chat, a session id. Both are minted once per launch and sent unchanged on every replay, since the host's replay fingerprint covers them. The phone arms its pending selection with that reservation before sending, so it lands on the terminal (matched by pane halves) or chat (matched by session id) as soon as the tab is listed. For an agent whose prompt is pasted after start, that is long before the reply, which waits for delivery. Landing also frees the "+" lock; the lock holds the create's id, so an older launch's reply cannot free a newer one's. The reply now only adds its own handle or session id (an older host ignores the reservation), starts the fallback countdown, and reports prompt delivery. A tab the user picks mid-launch replaces the pending selection, so the launch-start tab check is gone. A reservation the host refuses as already taken reads "Couldn't start the agent. Try again." on the first send, and as unconfirmed after a replay. The mobile UUID fallback now yields a v4 UUID, because a pane key's leaf must be one. The host launch path moved to new-tab-agent-host-launch.ts; session route parity re-pinned for that move and the landing's lock release. * test(mobile): expect the launch reservation in the four launch scenarios The four launch scenarios now expect the pane key and session id the phone sends (the scripted ids come first, so the operation id moves from ...001 to ...004). * fix(mobile): don't say an agent may not have started while the user is looking at it When a launch's reply was lost after its tab had already landed, the phone said "Couldn't confirm the agent started", although the listed tab proves it did. Now a listed tab narrows the doubt to the prompt or notes ("The agent started, but couldn't confirm the notes were sent."), the notes stay unsent, and a bare launch says nothing. Only the nested-function parity pin moves, for handleCreateTerminal passing the tab list to the launch. * test(mobile): read the launch's sent reservation through the host's params schema The anti-slop audit rejects Reflect.get; parsing with AgentLaunchReplay also asserts the host accepts the params the phone sent. * test(mobile): check the launch reservation against the host without importing its schema Mobile code may import the params contract only as types. The phone's tests now read the sent reservation by narrowing, a chat reservation is checked through the real host dispatcher, and the older-host drop is pinned host-side. * fix(mobile): don't send the same review notes to a second new agent The "+" lock is now freed when the launched tab lands, but review notes are only cleared when the launch's reply confirms delivery, which for a prompted launch can take up to a minute. In that window "Send review notes to AI" still offered the same notes, and choosing a new agent session started a second agent with them. The notes a new agent session is being started with are now held from the tap until that launch settles: the Send button no longer counts them, the sheet no longer offers them, and a stale tap on the old sheet starts nothing. Notes the host did not deliver become sendable again once the reply arrives. * test(mobile): record the + menu and diff-note launch goldens 4 added (+ as a terminal, + as a chat, notes delivered, notes not delivered). 10 existing create-terminal goldens move only because the recorded state now shows one pending selection instead of two refs; their requests are unchanged.
…ield (stablyai#24052) stablyai#24010's spec read them with Reflect.get, which the low-evidence lint rejects, so every PR's static analysis now fails on main.
* Let scheduled CI warmers wait and measure WebRTC startup * Measure a smaller daemon shutdown fixture image * Counterbalance WebRTC startup and verify retained fixture files * Record CI fixture measurements and remove temporary pilots * Clarify fixture build dependency cleanup evidence * Make coalesced snapshot fixture delivery deterministic * test: type the PTY write delay observer
* ci: reuse qualified Windows server slots for SSH host tests * ci: reuse prepared relay addons after an exact native cache hit
* fix(cursor): read desktop login outside the main thread * fix(cursor): await native worker retirement before respawning
* fix(cursor): preserve Windows hook input across profile paths Adapt the reviewed direct-command approach from PR stablyai#24381, and set UTF-8 input/output encoding for profiles requiring PowerShell. Co-authored-by: Vladimir Kurgansky <vladimir.kurgansky@gmail.com> * fix(cursor): preserve missing-script permission replies Keep the established encoded launcher for every event and explicitly use UTF-8 input/output. Preserve original missing-script acceptance and verify ordinary/spaced profiles rather than adding shell-specific direct guards. --------- Co-authored-by: Vladimir Kurgansky <vladimir.kurgansky@gmail.com>
* fix(opencode): skip absent session stores without warning * fix(opencode): retain warnings for inaccessible session stores --------- Co-authored-by: drakeo338 <paranoyouz@gmail.com>
…lyai#24582) * fix(terminal): fill DOM block glyphs only in repainted rows Adapt the block-fill approach from PR stablyai#15955 and bound painting to xterm render ranges without observer or animation-frame rescans. Co-authored-by: mmarabel <mmarabel@users.noreply.github.com> * fix(terminal): preserve block fills across DOM row replacements --------- Co-authored-by: mmarabel <mmarabel@users.noreply.github.com>
* ci: bound unit jobs to one hour of execution * docs: keep CI budget notes clear of the headless follow-up * docs: keep CI deadline evidence in the pull request
…tion (stablyai#24612) * fix(opencode): retain failed and stopped TUI turn outcomes Adapt the root verdict proposal from brennanb2025 in PR stablyai#23105 to the current TUI-owned lifecycle, keeping hook-store authority and existing mainAgent semantics. * fix(opencode): let auto-approved permissions settle before attention * fix(opencode): reconcile cached outcomes with completed session turns * fix(opencode): bind terminal verdicts to ending event timestamps * fix(opencode): publish approval cards for permission requests
* chore(deps): update reviewed desktop dependencies and tooling * chore(deps): update compatible mobile packages and Fastlane * chore(deps): update cloud transports and enforce release age * chore(deps): patch documentation dependencies and record review * chore: remove dependency review reports * test(linear): smoke-load resolved SDK through CommonJS loader * fix(deps): keep native rebuilds from reinstalling addon dependencies * fix(native): invoke installed node-gyp directly for Node rebuilds * test(cloud): exclude observer probes from row-lock timing budget * test(mobile): preserve the CSS writer receiver in viewport spy * test(native): remove obsolete batch-shim fixture exception * Stream native rebuild output through the process wrapper
…Auth (stablyai#24593) Consolidates the reviewed version and visibility work from stablyai#24283 with the probe gating and structured error classification from stablyai#24296. Reject unsuccessful version probes, preserve diagnostic precedence, and assert that real quota reads start no model turn. Co-authored-by: Pablo Werlang <19828711+werlang@users.noreply.github.com>
…in (stablyai#24596) * fix(antigravity): make POSIX hooks executable with bounded JSON stdin * test(antigravity): decode SSH hook shell command before assertions
* test: check plugin fixture worktree cleanup * test: check E2E installer package environment
innocarpe
force-pushed
the
fix-windows-spaced-terminal-path-paste-20261002
branch
from
October 2, 2026 14:07
738d124 to
5559d8c
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upstream: stablyai#24809
ELI5
When terminal clipboard HTML turns the filename at the end of a Windows path into a fake link, the Rich Markdown editor now keeps the complete pasted text, including paths whose folders contain spaces.
What Changed
Why
A terminal can put part of a copied path in HTML text and wrap the filename in an anchor such as
http://README.md. The old matcher stopped at a folder space, so Rich Markdown could keep that synthetic link instead of the plain path the user copied.Linked Issue
Fixes stablyai#24810
Visual Proof
Measured with the same paste-handler fixtures before and after the change in Vitest's happy-dom environment on macOS. These are handler and editor-dispatch results; an actual Windows GUI paste session was not tested.
upstream/mainat1fc24d3)Read C:\Users\My Project\README.md before editing.with an HTML anchor tohttp://README.mdfalse;defaultPrevented=false; inserted text[]true;defaultPrevented=true; inserted text is the complete original plain-text valueReview \\build-server\Team Docs\README.md with the release notes.with an HTML anchor tohttp://README.mdfalse;defaultPrevented=false; inserted text[]true;defaultPrevented=true; inserted text is the complete original plain-text valueThe fixtures use the split
<span>...path prefix...</span><a href="http://README.md">README.md</a>clipboard shape. The output above is reproducible interaction evidence; no layout or styling changed, and no screenshot/video is attached.Testing
Verified locally on macOS:
oxfmt --checkandoxlint --deny-warningson both changed files.tsc --noEmit -p config/tsconfig.tc.web.jsonpassed.upstream/main: 0 new findings across 2 changed files.The actual Windows GUI was not tested. Upstream CI was still pending when this description was prepared.
AI Disclosure
GPT-6-Luna
Review
Reviewed the hostname regex escaping, path-segment separators, filename boundary, and UNC share-root compatibility. Regression coverage includes spaced drive and UNC paths,
$README.md,README.md.backup,README.md,backup, an unrelated same-sentence link, ordinary HTTP links, and non-Windows text. No command execution or filesystem path access is added.Agent skill upstream boundary
docs/reference/agent-skill-sharing-upstream-boundary.mdand copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.Notes
The change is limited to rich-editor clipboard handling. The matcher treats clipboard paths as text and escapes the URL hostname before building a regular expression. Cross-platform and Remote SSH/mobile impact was considered; only macOS happy-dom tests were run, with no actual Windows GUI or Linux GUI test.
Checklist
N/Awith reason (the interaction proof is the measured fixture table above; no screenshot/video was captured)pnpm lint,pnpm typecheck,pnpm test, andpnpm buildpass (or CI will cover; local preferred)