Skip to content

feat(usage): show ZCode Coding Plan quota on current main - #23520

Merged
nwparker merged 4 commits into
stablyai:mainfrom
guanbear:codex/pr-14556-local-usage
Sep 28, 2026
Merged

nwparker merged 4 commits into
stablyai:mainfrom
guanbear:codex/pr-14556-local-usage

Conversation

@guanbear

@guanbear guanbear commented Sep 28, 2026 •

Copy link
Copy Markdown

ELI5

Replaces #14556 with a fresh ZCode Coding Plan quota meter based directly on current main (the old branch was stacked on closed #13965). The meter shows the provider-reported 5-hour and weekly Coding Plan windows, plus a separately labeled MCP allowance when present.

What Changed

This replaces the stacked #14556 draft after #13965 was closed.

Implementation

  • Read the selected ZCode model provider's explicit API key from ~/.zcode/cli/config.json in the main process. An ambiguous or missing selection stays unavailable; another configured account is never substituted.
  • Call the configured Z.ai or BigModel quota host over HTTPS with a 15-second timeout. Only api.z.ai, open.bigmodel.cn, and dev.bigmodel.cn are allowed; nonstandard ports and redirects are rejected before a credential can follow them.
  • Send the selected Coding Plan key as a raw Authorization value, matching the Coding Plan path in ZCode's usage provider. ZCode also supports Bearer <ZCode JWT> for its ZCode JWT path; this integration reads the explicitly selected Coding Plan key from ZCode config and uses the raw-key path. CodexBar and Pulse use Bearer for their manually configured keys.
  • Parse TOKENS_LIMIT and newer CREDIT_LIMIT windows by unit and number. Suppress implausible five-hour reset timestamps while retaining the quota value. Prefer usage counts over a stale integer percentage, as documented by CodexBar and Pulse. Keep TIME_LIMIT labeled as MCP, separate from Coding Plan windows.
  • Integrate the provider with the current split rate-limit service, status bar, settings search, persisted default, and unavailable/error handling. A keyed, non-secret account identity prevents a failed account switch from reusing the previous account's quota; an omitted ZCode snapshot from an older main process no longer blocks the setup prompt. The roster and detail views both label the separate MCP allowance as MCP.

Linked Issue

Follow-up to #14556; quota tracking remains requested by #18506 and #21757.

Visual Proof

On macOS, I launched an isolated background Orca dev instance from this patch rebased onto main at 992bad5375. After refreshing usage, the live UI showed ZCode in the status-bar roster with 5-hour, weekly, and MCP windows; its detail panel showed all three values and reset countdowns. Local CDP screenshots were captured for verification but are not attached because they show account-specific quota values.

Testing

  • Live read with the locally configured ZCode account: the fetcher returned ok with 5-hour, weekly, and MCP windows and reset timestamps. The API key and account-specific quota values were not included in this PR.

  • Visual inspection of the live status-bar roster and ZCode detail panel in an isolated hidden Electron instance, using Playwright CDP. Both identify the separate MCP window.

  • Automated quota, service, and status-bar tests.

  • Node 24 / pnpm 12 typecheck passed.

  • 16 affected Vitest files, 245 tests passed after the main rebase and MCP roster-label fix. After the latest review fixes, 64 focused tests in 3 suites and a read-only live quota fetch passed on the current head.

  • Full repository lint, localization checks, max-lines ratchet, formatting, and git diff --check passed.

Review

Please review the ZCode-configured API-key path, host/redirect restrictions, and CREDIT_LIMIT parsing. A live fetch and the rendered UI were verified locally on the current main-based patch.

Agent skill upstream boundary

  • No upstream agent skill source was copied.

Limits and review notes

  • The quota endpoint is not documented as a stable third-party API. The fetcher fails visibly when its response changes.
  • ZCode OAuth-only sign-in does not necessarily leave a usable API key in the config file; that case remains unavailable. ZCode's local SQLite model_usage is useful for consumption breakdown but cannot infer remaining quota. A local entitlement-cache reader or explicit key setup could be a separate follow-up.
  • The request/parser paths are covered with synthetic fixtures. I also used my locally configured ZCode account for read-only fetch and hidden-renderer UI verification. No credential or account-specific quota value is included in this PR.

Checklist

  • Independent PR against current main.
  • Typecheck, relevant tests, full lint, localization, and diff checks pass.
  • Live quota fetch verified locally.
  • Current UI visually inspected with live account; roster and detail views verified by CDP.

AI disclosure

Implemented with OpenAI Codex.

@guanbear

Copy link
Copy Markdown
Author

@nwparker Thank you for the careful guidance on #14556. This fresh PR keeps the host allowlist and explicit unavailable/error states. I followed the distinction you called out: local model_usage records consumption but cannot give remaining quota, so the meter uses the quota endpoint and ZCode's raw Authorization header. The PR describes the undocumented-endpoint and OAuth-only limitations.

I have now validated one read-only fetch with my local ZCode configuration: it returned ok with five-hour, weekly, and MCP windows and reset times. No key or account-specific quota values are included here. Node 24 typecheck, 241 affected tests, and full lint passed. I marked this ready for review; live UI visual inspection is still listed as an open check in the description. I'd appreciate your review when you have a chance.

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ℹ️ No critical issues — minor suggestions inline.

Reviewed changes

  • ZCode quota fetcher (src/main/rate-limits/zcode-usage-fetcher.ts): reads the selected ZCode provider's API key from ~/.zcode/cli/config.json, hits the allowlisted Coding Plan quota host over HTTPS with a 15s timeout and redirect: 'error', and maps TOKENS_LIMIT / CREDIT_LIMIT / TIME_LIMIT windows into session/weekly/monthly. Credential selection refuses to substitute another configured account.
  • Service integration: zcode joins the polled provider set across service-full-cycle-preparation, service-full-cycle-application, service-polling, service-state, and service-types, with matching state slots, failure-streak tracking, stale-policy application, and test-harness mocks.
  • Status bar and settings: zcode provider id, letter/icon/display name, CLI-gated visibility, default-on migration flag, Appearance search entry, and en strings.
  • Shared contracts: ProviderRateLimits['provider'], RateLimitState, StatusBarItem, persisted _zcodeStatusBarDefaultAdded, RPC UiUpdateFields, and createEmptyRateLimitState all extended.

I verified the parser against the cited references: the unit multipliers ({1:1440,3:60,5:1,6:10080}), the TIME_LIMIT unit-5/number-1 MCP marker special case, and the counts-over-percentage logic all match CodexBar's zai.js / Pulse. The raw (non-Bearer) Authorization also matches ZCode's own monitor contract in bigmodelUsageQuotaProvider.ts. The fetcher tests assert exact URLs, headers, and percentages, and the security tests (port rejection, no account substitution, key not in error output) are meaningful.

ℹ️ Nitpicks

  • src/shared/rate-limit-types.ts:4 still documents windowMinutes as only 300 or 10080; zcode's MCP window now uses 43200. Worth widening the comment so the next reader doesn't treat it as exhaustive.
  • src/main/rate-limits/zcode-usage-fetcher.ts:248 sets credentialSource to the absolute config path (/Users/<name>/.zcode/cli/config.json), which rides into renderer/mobile snapshots. Other providers use opaque tokens (keychain, cli, desktop); a token like zcode-config would keep the home path out of the payload.

Pullfrog  | Fix all ➔ | Fix 👍s ➔ | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

Comment thread src/main/rate-limits/zcode-usage-fetcher.ts Outdated
@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Advanced

Run ID: 76f6cf9e-bf21-415c-a4c9-391e1a8ab901

📥 Commits

Reviewing files that changed from the base of the PR and between 543af74 and 869cb0a.

📒 Files selected for processing (6)
  • src/main/rate-limits/service-refresh-orchestration.test.ts
  • src/main/rate-limits/service/service-full-cycle-application.ts
  • src/main/rate-limits/zcode-usage-fetcher.test.ts
  • src/main/rate-limits/zcode-usage-fetcher.ts
  • src/renderer/src/components/status-bar/status-bar-provider-visibility.test.ts
  • src/renderer/src/components/status-bar/status-bar-provider-visibility.ts
🚧 Files skipped from review as they are similar to previous changes (2)
  • src/main/rate-limits/service-refresh-orchestration.test.ts
  • src/main/rate-limits/service/service-full-cycle-application.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 5 remain after this review.


📝 Walkthrough

Walkthrough

The change adds ZCode Coding Plan quota fetching and connects its rate limits to service polling and status-bar presentation. It adds credential selection and validation, quota response mapping, failure and stale-state handling, status-bar controls and labels, and a one-time default-item migration.

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to 869cb

The ZCode quota meter appears ready for normal merge checks; no concrete outstanding behavior risk is identified.

Security Architecture Review

Security architecture risk: 🟡 Moderate · up to 869cb

The quota request is restricted to approved HTTPS hosts, but a quota result may remain associated with an earlier ZCode account if the selected account changes while a request is in progress. The apparent exposure is the local usage display, not disclosure of the API key to an arbitrary host.

Retained concerns

  • Medium · security · inferred: A successful quota response can be applied after the selected ZCode credentials change, because the result is not checked against the current account at application time. This could display the previous account’s usage under the new selection unless a separate invalidation path intervenes.
Security review details

Security Blast Radius

  • inferred — The identified transition risk is account-specific quota misattribution in the local usage display. The inspected production caller supplies only a cycle abort signal, while the request destination is constrained to three approved HTTPS hosts; the evidence does not establish arbitrary renderer-origin requests or arbitrary-host credential delivery.

Security Findings and Attack Paths

  • inferred — If a selected account changes from A to B while A’s quota request is active, A’s successful response can reach provider state without comparison to B’s current credential identity. Whether an external configuration-change handler reliably prevents this transition remains unverified.

Trust Boundaries and Controls

  • observed — The local configuration crosses into an authenticated network request only after selected-provider, key, protocol, host, and port checks. The request has a 15-second timeout and rejects redirects.

Resilience and Maintainability Implications

  • observed — Unavailable results discard stale quota data, changed known identities prevent error-time stale retention, and aborts prevent result publication. These protections cover different transitions from a successful result whose account changes during the request.

Hardening Proposals

  • proposed — Before publishing a ZCode result, compare its credential provenance with the currently selected credentials or an invalidated account generation, and clear prior account data when selection changes. Cover success, failure, cancellation, and repeated refreshes across that transition.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 23.08% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 26 functions across 41 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Issue #13965 is closed and provides historical context only. Its coding requirements do not apply to this pull request. No active directly linked issue provides coding requirements.
Out of Scope Changes check ✅ Passed The changes remain within the ZCode Coding Plan meter scope. They add ZCode fetching, quota parsing, rate-limit integration, status-bar presentation, settings integration, persistence, error handling,…
Title check ✅ Passed The title clearly and concisely describes the primary change: adding ZCode Coding Plan quota display to the usage UI.
Description check ✅ Passed The description is detailed and covers the user impact, implementation, rationale, linked issue context, testing, visual verification, limitations, and AI disclosure. It does not use the template's ex…
  • Fix all pre-merge checks with AI

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Advanced

Run ID: 89aeb85f-9f24-4ced-8c30-b04732b06bc0

📥 Commits

Reviewing files that changed from the base of the PR and between 0b2dd0a and f7491d6.

📒 Files selected for processing (40)
  • src/main/rate-limits/rate-limit-service-test-harness.ts
  • src/main/rate-limits/service-account-target-selection.test.ts
  • src/main/rate-limits/service-antigravity-usage.test.ts
  • src/main/rate-limits/service-cursor-usage.test.ts
  • src/main/rate-limits/service-inactive-account-previews.test.ts
  • src/main/rate-limits/service-live-claude-usage.test.ts
  • src/main/rate-limits/service-minimax-usage.test.ts
  • src/main/rate-limits/service-refresh-orchestration.test.ts
  • src/main/rate-limits/service-window-activation.test.ts
  • src/main/rate-limits/service/service-full-cycle-application.ts
  • src/main/rate-limits/service/service-full-cycle-preparation.ts
  • src/main/rate-limits/service/service-polling.ts
  • src/main/rate-limits/service/service-state.ts
  • src/main/rate-limits/service/service-types.ts
  • src/main/rate-limits/zcode-usage-fetcher.test.ts
  • src/main/rate-limits/zcode-usage-fetcher.ts
  • src/renderer/src/components/settings/appearance-status-bar-search.test.ts
  • src/renderer/src/components/settings/appearance-status-bar-search.ts
  • src/renderer/src/components/settings/appearance-status-bar-zcode-toggle-search.ts
  • src/renderer/src/components/status-bar/StatusBarProviderSegment.tsx
  • src/renderer/src/components/status-bar/StatusBarVisibilityMenu.tsx
  • src/renderer/src/components/status-bar/status-bar-agent-gating.test.ts
  • src/renderer/src/components/status-bar/status-bar-agent-gating.ts
  • src/renderer/src/components/status-bar/status-bar-provider-visibility.test.ts
  • src/renderer/src/components/status-bar/status-bar-provider-visibility.ts
  • src/renderer/src/components/status-bar/tooltip.test.ts
  • src/renderer/src/components/status-bar/tooltip.tsx
  • src/renderer/src/components/status-bar/usage-error-copy.ts
  • src/renderer/src/components/status-bar/usage-provider-settings-target.ts
  • src/renderer/src/components/status-bar/use-status-bar-controller.ts
  • src/renderer/src/i18n/locales/en.json
  • src/renderer/src/store/slices/ui-hydration-workspace-preferences.test.ts
  • src/renderer/src/store/slices/ui/ui-slice-hydration-status-bar-items.ts
  • src/shared/persisted-ui-state-types.ts
  • src/shared/rate-limit-state-factory.ts
  • src/shared/rate-limit-types.test.ts
  • src/shared/rate-limit-types.ts
  • src/shared/rpc-contract/client-ui-params.ts
  • src/shared/status-bar-defaults.ts
  • src/shared/ui-chrome-types.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 6 remain after this review.

Comment thread src/main/rate-limits/service/service-full-cycle-application.ts Outdated
Comment thread src/main/rate-limits/zcode-usage-fetcher.ts Outdated
Comment thread src/renderer/src/components/status-bar/status-bar-provider-visibility.ts Outdated
@guanbear
guanbear force-pushed the codex/pr-14556-local-usage branch from f7491d6 to cd90c84 Compare September 28, 2026 03:13
Adds a ZCode provider meter alongside Claude, Codex, Gemini, Grok, and
the other usage providers:

- New bounded fetcher reads the Coding Plan credential from the ZCode
  CLI config and calls the quota endpoint on the allowlisted official
  hosts only, with no shell and a 15-second timeout.
- Wires zcode into the shared rate-limit state, polling/stale handling,
  status bar, provider menu, tooltip, settings search, and the one-shot
  default-on status item migration.
- Uses Z as the compact narrow-layout badge.
- Status item visibility is detection-gated on the ZCode CLI, which this
  stacks on top of the ZCode first-class agent PR.
@guanbear
guanbear force-pushed the codex/pr-14556-local-usage branch from cd90c84 to 543af74 Compare September 28, 2026 03:17
@guanbear

Copy link
Copy Markdown
Author

Follow-up on the live UI check: I rebased this PR onto main at 992bad5375 and tested the resulting build in an isolated background Orca profile with my local ZCode account. After a refresh, the status-bar roster showed the 5-hour, weekly, and MCP windows, and the ZCode detail panel showed each value with a reset countdown.

That check caught one UI detail: the roster called the MCP allowance 30d, even though the detail panel correctly called it MCP. I fixed the roster label in 543af74012 and verified the rendered menu again. The latest head passes Node 24 typecheck and 16 affected test files (245 tests); full repository lint passed before the final unrelated main rebase. I captured CDP screenshots locally but did not attach account-specific quota values to this public PR. Thanks again for the review guidance.

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ℹ️ No critical issues — one minor consistency note below.

Reviewed changes

  • Fixed the ZCode roster chip label for the monthly window: shortLabel now returns the section label (MCP) instead of the generic window-duration chip (30d) when p.provider === 'zcode', covering both verbose and compact roster modes.
  • Added a UsageRow test asserting the MCP label in both roster modes.

ℹ️ ZCode's MCP label is still missing on the status bar's verbose monthly-only path

The roster now renders the ZCode monthly window as "MCP", and compact status-bar mode inherits that label through getTightestUsageSection. But the verbose status-bar segment still falls back to the generic chip label in its monthly-only branch, so an account whose response carries only the MCP window would show a reset duration (or 30d) there while the roster and tooltip say "MCP".

Technical details
# ZCode MCP label not applied on the verbose monthly-only segment

## Affected sites
- `src/renderer/src/components/status-bar/StatusBarProviderSegment.tsx:164-168` — the `p.monthly && !p.session && !p.weekly` branch calls `formatRateLimitWindowChipLabel(p.monthly)` with no zcode/MCP handling.
- `src/renderer/src/components/status-bar/tooltip.tsx:190-197` — canonical MCP label source (`auto.components.status.bar.tooltip.zcode.mcp`).
- `src/renderer/src/components/status-bar/UsageRosterPanel.tsx:52-54` — the roster fix landed in this commit.

## Required outcome
- The verbose status-bar segment should label a zcode `monthly` window as "MCP", matching the roster and tooltip, rather than a duration/`30d`.

## Suggested approach (optional)
- Reuse the `getWindowSections(p)` label for the monthly section (or a small shared helper) in `VerboseProviderUsage` instead of `formatRateLimitWindowChipLabel`.

## Open questions for the human (optional)
- Is an MCP-only zcode response reachable in practice, or does the API always include the 5-hour/weekly windows?

Pullfrog  | Fix it ➔ | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

@guanbear

Copy link
Copy Markdown
Author

Thanks for the detailed review. I pushed 869cb0a854 with the requested follow-ups:

  • A five-hour window drops an implausible reset timestamp while retaining its quota value.
  • The fetcher requires an explicit selected model provider, and a keyed, non-secret account identity prevents a failed account switch from showing the previous account's quota. A transient failure for the same account can still use the existing stale-data policy.
  • An omitted ZCode snapshot from an older main process no longer keeps the setup prompt pending.

The current head passes 64 focused tests in 3 suites, Node 24 typecheck, full repository lint, and a read-only live quota fetch using my local configuration. The status-bar roster and detail panel were also checked in an isolated background Electron instance after the earlier MCP-label fix; no account-specific values are attached. I updated the PR description with the new behavior and verification. Thanks for catching these edge cases.

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes

This run reviewed the delta since the prior pullfrog review (543af740), i.e. commit 869cb0a (fix(usage): guard ZCode account identity and reset display).

  • ZCode account identity: added usageMetadata.authProvenance, an HMAC-SHA256 of the selected provider id, origin, and key (never the key itself). The service now drops a stale ZCode quota when a failing refresh belongs to a different or unknown account, while a same-account failure still falls back to the recent snapshot.
  • Credential selection tightened: an explicit, well-formed model.main provider is now required; the previous "infer from the only configured provider" path is gone, so a missing/ambiguous selection stays unavailable.
  • 5-hour reset plausibility guard: a session window whose nextResetTime is more than 5h1m out now drops resetsAt while keeping usedPercent, mirroring CodexBar's resetIsPlausible.
  • Setup CTA gating: isUsageEmptyState treats an omitted zcode key (older main process) as not-pending so the teaching CTA still shows, while a present null/pending snapshot still suppresses it.
  • Tests: added account-switch vs same-account stale handling, reset plausibility, explicit-provider selection, and isUsageEmptyState cases. Verified locally — zcode-usage-fetcher, service-refresh-orchestration, and status-bar-provider-visibility suites pass (64 tests).

The prior inline thread on the five-hour reset guard is addressed by this commit (replied and resolved). The earlier body nitpicks (config-path credentialSource, the rate-limit-types.ts:4 comment, and the verbose status-bar monthly-only MCP label) remain open but unchanged, so they are not re-raised here.

Pullfrog  | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

@nwparker

Copy link
Copy Markdown
Contributor

@guanbear thanks for rebasing this onto main — that was exactly the right move, and the security posture here is genuinely careful. I reviewed it closely because it reads a credential and makes a network call, and I want to say what held up before the blocker:

  • Host allowlist is an exact-match Set, not a suffix check, so api.z.ai.attacker.com can't slip through.
  • HTTPS enforced, port pinned to 443, redirect: 'error' so the credential can't follow a redirect.
  • /[\r\n]/.test(apiKey) — a header-injection guard on the key. Most people miss that one; nice.
  • The key is never logged, and account identity is derived via HMAC rather than stored.
  • Every one of those has a test, including "rejects a nonstandard HTTPS port before sending the key".

I also verified your central claim against the ZCode source rather than taking it on trust: /api/monitor/usage/quota/limit is ZCode's own BIGMODEL_QUOTA_PATH constant. That materially answers the hesitation on #21757 about calling an endpoint not documented for third parties — it's the endpoint ZCode itself calls, now visible in the open-source tree.

One correction to the PR description, though the code is right: ZCode uses both auth shapes. bigmodelUsageQuotaProvider.ts:673 adds Bearer to the zcode JWT; :674 sends the coding-plan JWT raw. Your raw Authorization matches the coding-plan path, which is the credential you're using — so the behaviour is correct, but the description states it more absolutely than the source does. Worth a tweak so a future reader doesn't "fix" it the wrong way.

Blocker: global fetch can crash the app

FAIL src/main/global-fetch-call-site-audit.test.ts
Global fetch (bare, globalThis.fetch, or global.fetch) uses undici, where an
unread response body can crash the whole process (orca#8695).

This one matters beyond CI being red. On any path where you don't read the body — an allowlist rejection, a non-2xx, a parse bail, the abort/timeout path — undici can take the whole main process down. A quota meter failing should never be able to kill the app.

Two sanctioned fixes, both already used by the sibling fetchers:

  1. net.fetch from Electron — see src/main/rate-limits/claude-oauth-usage-request.ts:72.
  2. cancelUnreadResponseBody(response) from src/main/lib/unread-response-body.ts on every path that doesn't consume the body — see src/main/rate-limits/codex-backend-usage-client.ts:70.

Then add the call site to AUDITED_GLOBAL_FETCH_LINES if you stay on global fetch. Given this is a credentialed request, I'd lean toward net.fetch.

Also red

  • consistent-type-assertions at zcode-usage-fetcher.ts:81 and :232. Repo rule is no assertions except as const; an unavoidable one needs a line-specific // SAFETY: note. For JSON-boundary narrowing, a type guard is usually cleaner — I ended up doing that in the harness PR for the same rule.
  • RPC recording pin and test vs non-test LoC are also failing; both are repin/ratchet gates rather than logic problems.

Happy to re-review as soon as the fetch path is sorted. This is good work and I'd like to land it.

…ncel unread bodies

Clears both CI failures on this branch.

Static analysis flagged two type assertions. Both sat on JSON this process does
not control — a user-edited `~/.zcode/config.json` and a response from a remote
quota endpoint — so the declared shapes were a guess rather than a fact. The
fields were already read defensively with `typeof` checks; only the outer cast
claimed the value was an object at all. `isRecord`/`readRecord` make that claim
a check, which also lets the two now-unused config and response types go.

The global-fetch audit failed because the request was a new unaudited call site.
Node's bundled undici can crash the whole process when a response body is left
unread and the peer closes the socket (orca#8695), and the `!response.ok` path
returned without touching the body. It now cancels the body first, matching the
sibling Codex clients, and the file is registered with its expected call count.

Co-authored-by: guanbear <guanbear@users.noreply.github.com>
@nwparker
nwparker force-pushed the codex/pr-14556-local-usage branch from 6e54d93 to 72a3fe1 Compare September 28, 2026 05:45
@nwparker

Copy link
Copy Markdown
Contributor

Thanks @guanbear — and thanks for rebasing #14556 onto current main instead of letting it rot. The credential handling here is careful in ways I appreciated on review: exact-match host allowlist, HTTPS and port-443 pinning, redirect: 'error', the /[\r\n]/ guard on the key before it reaches a header, and an HMAC'd account identity so the key itself is never stored or logged.

I took over the two CI failures rather than leave them to you:

  1. Two type assertions (zcode-usage-fetcher.ts:81, :232). Both sat on JSON this process doesn't control — a user-edited ~/.zcode/config.json and a remote quota response — so as ZcodeConfig / as QuotaResponse were guesses, not facts. The individual fields were already read defensively with typeof; only the outer cast claimed the value was an object at all. isRecord/readRecord turn that claim into a check, which let both now-unused types go.

  2. The global-fetch audit. This was a new unaudited fetch call site. Node's bundled undici can crash the whole process when a response body is left unread and the peer closes the socket (undici#5360, orca#8695) — and the !response.ok path returned without touching the body, which is exactly the shape that trips it. It now cancels the body first, matching codex-backend-usage-client.ts, and the file is registered in AUDITED_GLOBAL_FETCH_LINES. Worth knowing for future PRs: any new global fetch has to consume or cancel on every path, including the error paths.

pnpm tc clean, 658 rate-limit + audit tests green, changed-code quality gate passes. Merging once CI confirms. Your commits stay as-is and my fix credits you as co-author.

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes

This run reviewed the delta since the prior pullfrog review (869cb0a), i.e. commit 72a3fe1 (fix(usage): read the ZCode quota JSON instead of asserting it, and cancel unread bodies).

  • Runtime shape readers replace casts: the ZCode config and quota response are now narrowed through isRecord / readRecord / readMainProvider (src/main/rate-limits/zcode-usage-fetcher.ts:29-46), so a malformed config or a changed remote payload degrades to unavailable / parse rather than a property access on an unchecked as cast. Credential selection, the host allowlist, and the parser logic are untouched, so the well-formed paths are unchanged.
  • Unread response bodies are cancelled: the !response.ok path now awaits cancelUnreadResponseBody(response) before returning the failure (zcode-usage-fetcher.ts:217-224), closing the undici unread-body crash hazard, and the fetcher is registered in global-fetch-call-site-audit.test.ts with its expected one call site.

I ran zcode-usage-fetcher.test.ts, unread-response-body.test.ts, and global-fetch-call-site-audit.test.ts (15 tests pass) plus typecheck:node — all green. The delta is behavior-preserving hardening; the earlier body nitpicks are unchanged by it, so they are not re-raised here.

Pullfrog  | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

@guanbear

Copy link
Copy Markdown
Author

Thanks for the careful review and for fixing the JSON-boundary checks and unread response-body path. I updated the PR description to distinguish ZCode's two auth forms: Bearer for a ZCode JWT, and a raw Authorization value for the Coding Plan JWT used by this integration. That should help prevent a future change from applying the wrong form here.

The current CI has passed static analysis, typecheck, packaging, and the completed test shards; one test shard is still running. I'll check its result. I appreciate your help getting this ready to land.

@nwparker
nwparker merged commit 076759e into stablyai:main Sep 28, 2026
41 checks passed
Jinwoo-H added a commit that referenced this pull request Sep 28, 2026
…n the host and pane (#23602)

* test(native-chat): await the async history and journal snapshot in three tests (#23560)

#22835 made history() and journalSnapshot() async; tests from #23502 and #22944 still call them synchronously, so the typecheck job is red on every PR while main pushes do not run it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(usage): show ZCode Coding Plan quota on current main (#23520)

Shows the ZCode Coding Plan quota in the status bar alongside the Claude and Codex usage readouts, reading the key from the user's own ZCode config.

Credentials are scoped tightly: the host must be an exact match in the allowlist, HTTPS on port 443 only, `redirect: 'error'`, and the key is checked for CR/LF before it reaches a header. The key itself is never stored or logged — account identity is an HMAC.

Both JSON inputs (a user-edited config file and the remote quota response) are narrowed at runtime rather than asserted, and the request cancels an unread response body on the error path so it cannot trip the undici parser crash (orca#8695).

Co-authored-by: guanbear <guanbear@users.noreply.github.com>

* fix(mobile): paired clients re-derive a kept terminal after a cold restore (#23109)

* fix(mobile): paired clients re-derive a kept terminal after a cold restore

A renderer frame published before a cold-restored terminal's PTY registered
was fenced to an empty tab list and recorded as accepted, and the renderer
never resends unchanged content. When registerPty binds a surface the
accepted frame fenced out, re-merge that frame so the fence reads current
state.

* test(mobile): drive the live desktop window through the runtime's desktop seam

* test(mobile): the re-derive path never flushes the store synchronously

* test(mobile): a re-derived frame must not bring back a surface the host retired after accept

* fix(mobile): a re-derived frame changes membership only for the registering surface

The replay re-ran the whole accepted frame, so a surface the host retired
after accept (a phone close whose remote PTY is still exiting, or a closed
chat tab) came back. Every other surface now keeps the host's current
decision; the removal repair is extracted from the terminal retirement
helper so non-terminal tabs are removed the same way.

* test(mobile): a re-derived frame must not drop or disown a phone-created terminal the desktop has not published

* fix(mobile): a re-derived frame does not infer renderer retirements from its older frame

* revert(mobile): drop the replay of a fenced renderer frame

Reverts the production parts of a88e1eaa0a, f5b99d0003 and 5af1c6d97a: the kept
renderer frame, rederiveFencedRendererSurface and its registerPty call, and the
mergeRendererMobileSnapshot / removeMobileSessionSnapshotTabs extractions. The
fence will instead read the host's saved membership record. The test file stays
and is rewritten for that mechanism.

* fix(mobile): the paired-list fence admits a terminal the saved session still lists

After a cold restore the in-memory mobile snapshot and PTY table start empty,
so in a repo with host-authoritative terminal membership the fence dropped a
restored terminal whose renderer frame arrived before its PTY registered, and
paired clients never listed it. The fence now also admits a surface the host's
saved workspace session still lists (tab under the worktree, leaf in its
layout), and registerPty pushes the listing so pending-handle turns ready at
once. A restored pane whose PTY never returns is listed as pending-handle, as in
repos that are not host-authoritative.

* test(mobile): keep the desktop window stub's type assertion on its SAFETY line

* fix(mobile): coalesce the registration push for a listed restored terminal

registerPty pushed the paired list immediately on every registration that backs a
listed surface. The desktop's graph sync after a spawn already publishes the same
pending-handle to ready flip on the 50 ms coalescing window, so each restored pane
cost two pushes, and a restore of N panes cost N immediate full-list pushes per
client. The touch now rides the same coalescing window, which still covers a
registration no graph change follows.

The test's "unchanged" sync dropped the graph's tab, which is itself a change, and
its no-extra-push assertion ran before any coalesced push could fire; both are
fixed, and a restore of two panes is asserted to push once.

The fence comment no longer claims the new clause keeps pending leaves out of the
graph: once the surface is listed, its leaves pass the shared predicate through
that listing, as any listed surface's do.

* fix(mobile): read saved membership only from the worktree's own session partition

For a runtime-host workspace, emptying the owning partition re-routes session reads to a single
other partition that still lists the worktree. If that older copy lists a surface a retirement just
removed, a lagging renderer frame could re-admit it. The saved-membership check now reads only the
partition the worktree's host names, so it never trusts a fallback copy.

* test(mobile): pin the own saved partition for every workspace host kind

Also correct the immediate-emit comment: only an exit bypasses the window; a registration's ready flip coalesces.

* Fix terminal focus when Cmd+J wakes a workspace (#23546)

* fix: retain workspace terminal focus through wake restoration

* test: reset CPU throttling after wake focus assertion

* fix: require terminal textarea readiness before claiming focus

* test: configure React act environment for dialog regression

* fix(mobile): size a terminal's first subscribe from the document's reported cell box (#23080)

* fix(mobile): size a terminal's first subscribe from the document's reported cell box

#22960 sent phone dims on a terminal's first subscribe by opening a throwaway
empty terminal (init 80x24 ""), awaiting its ready and measuring, behind a
per-document first-subscribe mark whose lifetime was tied to web-ready. That
cost a second xterm/WebGL instance and ~150 ms per open, plus lifecycle state.

The document now measures the cell box without a terminal (xterm 6's
CharSizeService strategy, rounded as the renderer rounds it) for every
text-size preset and reports it with its viewport in web-ready; a table,
because the text scale only reaches the document after that notify. Each
init's ready reports the box xterm actually laid out, which replaces the
probe's entry. The controller answers fitDimensions/measureFitDimensions from
that table and the view's layout with no message; without a table it asks the
document as before.

The session seeds an unmeasured viewport synchronously in subscribeToTerminal,
so the first subscribe carries dims by construction. Deleted: the empty init,
its awaitReady gate, deferFirstSubscribeUntilViewportMeasured and the
subscribedDocuments mark. The fit pass is unchanged and still covers a
document that reports no cell box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): read the reported cell box through in-narrowing, not Reflect.get

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct the probe's cell-box guess from the box xterm lays out

The web-ready probe is a guess: building the WebGL addon creates no context,
so a context that fails on load lands on the DOM renderer, whose width is not
snapped and depends on the column count. Before, a ready box that differed was
only logged; the first subscribe had carried the wrong column count, the host
echoed it, the fit pass saw the viewport equal to the host's dims, and the grid
stayed slightly shrunk. The store also kept the WebGL width after a context loss.

The document now reports the box xterm laid out whenever it changes (from
onRender, which covers a renderer swap and a DPR change that
onDimensionsChange does not fire for, and at ready). The store replaces the
guess; when that changes the current text size's entry, the view calls
onCellBoxChange with xterm's grid and the session re-fits, running the bounded
fit pass if the dims moved (one resubscribe). Equal boxes do nothing.

The RN layout box now survives a document reload; the document's own viewport
only stands in until the view reports a layout (on the page, web-ready arrives
first). The mismatch console.log is gone, and the probe's rounding names the
xterm version it copies.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct each cell-box guess at most once, so a DOM renderer cannot loop

On the DOM renderer the cell width is the rounded canvas width divided by the
column count, so every re-init at new cols reported a new box. Each one counted
as a correction, a floor over floats could flip the fit between two sizes, and
each flip landed converged, which reset the resubscribe budget: an unbounded
series of full-snapshot resubscribes.

Only the first laid-out box for a guessed text size may be a correction; later
reports still update the store, so fits stay truthful, but never resubscribe on
their own. The fit's floor gains a 1e-6 epsilon so floating-point error at an
exact boundary cannot flip a column or row. New document tests pin the render
report after a renderer swap and the report at ready for a paused renderer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): make xterm the only terminal cell measurer

The document builds its real terminal before web-ready, at the app's
text scale, and reports the box xterm laid out; the first init reuses
that terminal. The page-side prediction, the per-scale guess table and
the once-per-document correction are gone. The app remembers the box
per text scale for its lifetime, so a later open at a known scale
subscribes with phone dims at once. A box that changes at the same grid
(renderer swap, pixel ratio) refits the open terminal in place.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep commands queued before the terminal WebView first loads

A subscribe sized from the stored cell box can queue init before the
native WebView reports its first load start, which cleared the queue
and left the terminal blank. Only a reload now drops queued commands.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): re-init a document that lacks the subscription's init, and fit one frame width

- Web-ready now says whether the document holds the terminal's latest init
  (a reload before the first ready drops a queued one); the session
  resubscribes any initialized terminal whose document lacks it.
- One grid fit, shared by the app and the document, fed the unrounded frame
  width React Native laid out; it keeps exact fits whole at fractional pixel
  ratios. The document's viewport-width fits are gone.
- The page builds every document at the scale the view mounted with, as the
  native WebView does.
- A new document's first cell box is compared against the grid the
  subscribe fitted from the stored box.
- The terminal built before ready stays hidden until its first init.
- The cell-box census matches glyph-measurement techniques, not names;
  the store's unused clear() is gone.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): build the terminal before ready only for the view shown at mount

A session mounts one terminal view per tab, and each built xterm and a
WebGL context before ready: 20 tabs made 20 contexts at load, past the
~16 a page (or Android's shared WebView renderer) holds, and native logged
32 context losses. Only the view shown when it mounts builds early now;
the rest build at their first init as before. Deferring the WebGL addon
instead would change the reported box: the DOM renderer lays out 7.8x15
where WebGL lays out 7.667x15 at the same font.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): write a WebView document's start values into its page, not an injected script

Android ran the pre-content injected script after the document's own in
1 of 22 documents on the emulator; that document started with no text
scale or shown flag and built a terminal it should not have. The values
now sit in the page ahead of the document script, one source object per
start pair so a render never reloads the WebView.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the pre-ready terminal measures and reports while hidden

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): measure only the laid-out frame, and refit on a new grid, not a new width

- A measure needs both of the frame's dimensions from React Native; the
  document's viewport-height fallback is gone, and before the first layout
  the handle answers no fit without asking the document.
- A frame width change that still fits the PTY's grid from the stored box is
  a no-op, so sub-pixel layout jitter no longer re-measures. The width ref is
  written in that effect rather than during render (react-doctor).
- One "last grid" ref: the last reported grid, or the one a subscribe fitted
  from the stored box.
- The page render rig measures through the frame it laid out, as the session
  does, and lets the replay's fit settle before its resize-refit witness.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let only the current terminal document's ready flush

A reload kept the WebView and its onMessage, so the old document's late
web-ready flushed the queue into the reloading view and the new document
got a second init. Each document now gets its own view (keyed on a
generation the controller owns), every notify carries the generation of
the view that received it, and a web-ready from a replaced document
flushes nothing and stamps nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop every notify from a replaced terminal document

One rule at the receive boundary: a notify from any generation but the
current one is dropped, whatever its type, not only web-ready.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make fitDimensions a pure question; name each generation counter

- fitDimensions no longer records the grid. A width change to a new grid
  asked it first, so the DOM renderer's report of that grid's box read as
  "same grid, new box" and refit again. Only the first-subscribe seed
  (seedFitDimensions) records the grid the document's first report is
  checked against.
- viewGeneration counts the views, readyGeneration counts web-readies.
- replaceDocument no longer resets the load flag; the load-start reset
  stays as the guard for a view that reloads itself.
- The name-based lifecycle census is replaced by a behavioural test.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): typecheck the handle mocks, drop the unused cell-box get

- The two handle mocks carry both fitDimensions and seedFitDimensions,
  and the fake-timer acts return nothing, so the three test files check
  under tsconfig.test.json again.
- terminalCellBoxes.get had no product caller; the store's tests assert
  through fit.
- The load-start comment says what the controller does now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold the grid the document has, ignore a replaced view's load start, dispose a failed pre-ready terminal

- The document reports a new grid even with an unchanged box, so an
  in-place reflow on WebGL is held before a later renderer swap at that
  grid; the swap then refits. The app's apply paths do not hold the grid
  themselves: the DOM renderer's box follows cols, and a grid held on
  apply would read its own box as a renderer change and loop. One
  writer (holdGrid) holds the seeded or reported grid.
- A load start from a view a replacement unmounted is ignored, as its
  notifies already are.
- A terminal whose open throws before ready is disposed, not only
  unreferenced.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): ignore every native event from a replaced terminal view

One wrapper binds each WebView lifecycle event (load start, error, HTTP
error, render process gone, content process terminated) to the view's
generation, so a replaced view's late event cannot reset, replace or
put an error over the current document.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): a DOM seed refits once on its first report, not on the refit's own

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a terminal only after its document is ready

The document still builds its terminal before ready and reports the cell
box xterm laid out in web-ready; the app now subscribes after that ready
and fits from that box, so nothing is sent to a document before it is
ready. Everything that made a pre-ready subscribe safe goes: the
app-lifetime box store, the seed fit, the per-document view generations
and their event filtering, the init tracker and the hasInit resubscribe.
The native view reloads in place again and web-ready keeps main's reload
rule. Boxes are kept per view; the grid a document last reported still
guards the in-place refit against the DOM renderer's cols-dependent box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold one reported cell box and the grid the subscribe fitted

The controller keeps only the box the current document last reported,
not a per-text-size store: the document re-reports on a scale change.
The subscribe after ready fits from that box and holds the grid it
fitted, so the DOM renderer's first report at that grid (a new box)
refits once in place and converges; refit and apply paths hold nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): fit only a ready box at the app's scale; forget a reloaded document's box and grid

A reload keeps the document's mount scale, so a ready after a text-size
change reports a box at the old scale; that box no longer sizes the first
subscribe, which then takes the no-box path. A readiness reset drops the
old document's box and held grid, so the new document's first DOM report
at the same grid does not refit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give the terminal document its frame at init, and fit text scale over it only

A subscribe sized from the ready box sends no measure, so the document
had no frame when the text size changed and reported the pre-refit row
pitch. The app's init now carries the frame it laid out, in the fields a
measure uses; the router takes it from either. The text-scale fit reads
only that frame, with no viewport fallback.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say why a frameless text-scale change skips the resize

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one cell box per terminal notify, not an array

web-ready and cell-metrics carry `cellBox: {fontScale, cellWidth,
cellHeight} | null`; the document's `laidOutCellBox` returns one or
null and the parser validates one object. The text-scale match moves
from web-ready into `handle.fitDimensions`, the one place a box is fitted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): fit terminals in the app from the reported box; drop the measure round trip

The app already holds the box the document reported, so the refit and
the fit pass await the init's ready and call `handle.fitDimensions`
instead of posting `measure` and waiting on `measure-result`. The
document's measure, its retries, and the measure promise and timeout go.
The document still resizes locally on a text-size change, so every grid
the app sends (init, resize, reflow) carries the laid-out frame it was
fitted to. `holdSubscribedGrid` replaces `subscribeFitDimensions`, so
the only fits are `fitDimensionsFromCell` and `handle.fitDimensions`.
The render rig reads its fit from the ready box. The recorder adapter
mounts the new handle with the same recorded effects; the goldens it
mounts move on their adapterSha256 header only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): keep the terminal frame in one ref, and notify a new width imperatively

The session held the frame in a height ref, a width ref, a width state
and the refit's own width ref. It now holds one `terminalFrameRef`
({width, height} | null until the first layout; a hidden 0x0 layout
keeps the last box). onLayout notifies a new width imperatively, as it
does height, and the refit's notify skips a width whose fit is the grid
the PTY has. `terminal-frame-width-refit.ts`, the width state and its
effect go. The subscribe's layout gate reads "no frame yet" directly.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a held-back terminal on the frame's first layout only

`handleTerminalFrameLayout` ran on every onLayout; it now runs once, when
the frame first has a size. Later layouts only notify a new width.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): size the first subscribe inline in subscribeToTerminal

`sizeTerminalViewportFromCellBox` wrapped five lines in a 37-line
module; the subscribe now fits the ready box against the frame, holds
that grid and records the diagnostic itself. The helper's tests fold
into the subscription tests, which move to the subscription's name.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): drop the unreachable font-size guard on the reported cell box

xterm 6.1.0-beta.303 updates the render service's cell box in the same
task that sets `options.fontSize`: CharSizeService.measure fires
onCharSizeChange, and RenderService.handleCharSizeChanged runs the
renderer's `_updateDimensions` (DomRenderer.ts:359, WebglRenderer.ts:229).
`term.onRender` fires from RenderService._renderRows after the rows
are drawn (RenderService.ts:213, CoreBrowserTerminal.ts:538), and the
document writes its text scale and the font size in one task
(text-scaling.ts applyTextScale, terminal-init.ts init). So no report
can read a box between the font and the scale; the guard and its test
go. A new test pins the real order: no report when the font is set,
the new box at the new scale on the next render.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one start seam, no source cache, the reported box as an object

- `useState` already pins each view's WebView source at mount (a new
  test re-renders at another text scale and gets the same object), so
  the module-level `webViewSources` Map goes.
- `initialTextScale` and `buildsTerminalBeforeReady` become one
  `start(): { textScale, shown }` seam.
- `reportedCellBox` holds the last reported box and grid, not a string key.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to this branch and re-record

The terminal refit now fits in the app from the reported box and reads
one frame ref, so the recorder's terminal adapter mounts the new handle
(`awaitReady` + `fitDimensions`) and options (`terminalFrameRef`),
keeping its recorded effects. `baseline` is repinned to 21954dbd2f1, the
last commit to touch a fenced path, and every golden is re-recorded.
Proof by class against HEAD: 787 header-only, 0 body moved, 0 added,
0 deleted. Header keys moved: `baseline` on all 787, and `adapterSha256`
on the 14 goldens `terminal-mount-adapters.ts` mounts (query-reply 3,
accessory-raw-send 4, takeover-report 4, viewport-refit 3). No recorded
traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold the reported cell box and its grid in one ref

The controller kept the box in `cellBoxRef`, the grid in a string
`lastGridRef` and wrote it through `terminal-held-grid.ts`. One
`heldRef` now holds `{ cellBox, grid }`, as the document's own
`reportedCellBox` does: web-ready writes the box, every cell-metrics
report writes both, `holdSubscribedGrid` writes the grid, and a
readiness reset clears it. Same write points, so the one-refit bound
holds; the DOM-loop and refit-once tests pass unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the hold-rule commit and re-record

H (f00bebba483) touched a fenced path after the last repin, so
`baseline` moves to it and every golden is re-recorded. Against the
corpus before this branch's refreshes (21954dbd2f1): 787 header-only,
0 body moved, 0 added, 0 deleted; `baseline` on all 787 and
`adapterSha256` on the 14 goldens `terminal-mount-adapters.ts` mounts.
Against the previous refresh: `baseline` only. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the main merge and re-record

The merge (b3b1b0def22) is the last commit to touch a fenced path, so
`baseline` moves to it and every golden is re-recorded. Against
97b5bb2b9ac: 787 header-only, 0 body moved, 0 added, 0 deleted;
`baseline` on all 787, and `adapterSha256` on the 14
session.diff-review-actions goldens whose adapter #22951 edited. Against
origin/main: 787 header-only, 0 body moved/added/deleted; `baseline` on
all 787 and `adapterSha256` on this branch's 14 terminal goldens.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): the terminal document holds the grid and decides each refit

The document already kept the last reported box and grid; the app kept
a mirror of both to decide the refit. Now the document decides: its
`cell-box` notify carries `{ cellBox, refit }`, sent only when the box
changes, with `refit` a box that changed at a kept grid. web-ready
records the pre-ready terminal's box at its 80x24 grid, and the first
init that reuses that terminal holds the init's grid, so the DOM
renderer's first report refits once, as the subscribe's hold did. A
re-init no longer clears the record, so a new renderer at the same grid
still refits. The app keeps one `cellBoxRef` and `holdSubscribedGrid`,
`heldRef` and the grid on the notify go. The one-refit, DOM-loop and
renderer-swap tests move to the document with the same scenarios.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one init options object, and a frame on every grid

`init` takes `{ cols, rows, data, preserveScroll, oscLinks, frame }`
instead of six positionals, and `init`, `resize` and `reflow` (handle
and messages) require `frame: TerminalFrame | null`. The refit's reflow
check reads `!dims` alone, and the controller's test file is named for
the `cell-box` notify it now covers.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one notifyTerminalFrame for the frame's layout

The frame's onLayout made four calls and held the classification
itself. It now calls `notifyTerminalFrame({ width, height })`, and the
session's terminal-webview hook keeps the one frame ref, notifies the
height, subscribes the document held back for the first layout, and
notifies a later width change. `handleTerminalFrameLayout` is named for
what it does: `subscribeIntendedActiveTerminal`. The layout tests move
to that hook.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-8 head and re-record

85d421963c6 is the last commit to touch a fenced path. Against
a676c1b65a7: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): name the init option initialData, as the message does

The init option `data` becomes `initialData`, the message field's name,
so the controller passes it through unrenamed. The `preserveScroll` why
stays on the message type only, and the document test's title names the
three grids that carry the frame.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-9 head and re-record

486566c82b4 is the last commit to touch a fenced path. Against
3371c397150: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: rerun checks against main with #23560 landed

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(ai-vault): show ZCode CLI session history (#23513)

Surfaces ZCode CLI session history in AI Vault, so past ZCode sessions show up next to the other agents' instead of being invisible.

ZCode stores sessions in the same SQLite shape OpenCode uses, so this reuses the existing OpenCode lister and parser rather than adding a second scanner — the worker only varies the agent it stamps on each row. Discovery covers the native home and any WSL homes.

SQLite rows are narrowed at runtime rather than asserted: the statement API returns untyped column values, so the declared row shape is only a claim until something checks it, and a drifted schema or a database written by another tool reaches the same code.

Co-authored-by: guanbear <guanbear@users.noreply.github.com>

* test(mobile): repin the RPC recording corpus to main after #23080 (#23565)

#23080 squash-merged a corpus pinned to its branch commit 486566c82b4, which
the squash left unreachable from main. Repin baseline to main's tip and
re-record; every golden moves only its baseline header.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* perf(ci): diff against the merge commit's first parent so PR checkouts can be shallow (#23562)

Every changed-path gate asked git for `--merge-base "$BASE_SHA" "$HEAD_SHA"`,
which needs the event payload's base SHA to be in the local graph. That is the
only reason two jobs cloned all 8127 refs' history. On a pull_request checkout
HEAD is already the merge commit, so its first parent is the base side and no
merge base has to be computed. config/scripts/git-pull-request-diff-base.mjs
resolved that for the two Node gates; the workflow's inline gates now use the
same helper through a small CLI rather than open-coding it.

code_paths gates all 22 jobs, so its checkout is charged to the start of every
one of them: measured 20.7s to 1.6s, keeping blob:none because its sparse tree
is ~7 files and leaves no blobs to refetch. Static analysis drops the filter
instead, since populating all 30,226 files makes blob:none force a second
promisor fetch: 23s to ~11s.

Verified on a real merge ref. At depth 50 the old and new forms produce
identical changed-file sets. At depth 2 the new form still works and the old one
fails with `fatal: bad object`, which is the failure a stale base would have
caused once the checkout stopped being complete.

Also drops the dead resolveBase + merge-base prelude in the changed-code gate,
whose result resolvePullRequestDiffBase already discarded on every PR.

* feat(mobile): the keyboard covers the page like a native screen, and the shell says its height (#23110)

The shell no longer shortens the WebView for the keyboard; it publishes the keyboard height like the safe-area insets, so native's keyboard lift, refit hold and dismiss key run on the page unchanged. Keyboard and inset arithmetic read the shell's OS through a host-os seam. One page-version floor (manifest pageVersion, shell floor 1) replaces per-feature accept negotiation; a page below the floor gets the existing update wall, a desktop with no bundle keeps native screens. iOS shell drops the form accessory bar and its own keyboard observers. Native session screens untouched.

* fix(agent-session): wait for in-flight session-store writes before teardown returns (#23545)

* fix(agent-session): stop lease renewal before the renewal's write lands

Clearing the renewal interval only cancelled the next tick. A tick already
past its guard still had a whole-file store transaction to commit, and the
store's transaction lock re-creates the store directory before it writes, so
that commit could land after host teardown had finished releasing everything
it touches.

`stop()` now resolves once the tick in flight has finished writing, and host
teardown's stop-lease-renewal phase waits for it. The three test harnesses
that model a host vanishing without a clean quit shared a copy of the same
incomplete shutdown; they now share one helper that waits.

The symptom was a CI flake: the refusal-oracle spec removes its temp
directory in `afterEach`, and a renewal landing mid-removal put the store
directory back, so the removal failed with ENOTEMPTY on the temp root.

* fix(agent-session): wait for the delivery loop's restart when abandoning a host

The abandon helper disposed the delivery loop and moved on. Disposing only
stops the loop's NEXT step: a step already past that check keeps going, and
the restart it runs for an accepted send reserves an owner, which is a store
commit. The store re-creates its own directory before every commit, so that
commit put the directory back under the temp-directory removal the test does
next, and the removal failed with ENOTEMPTY.

Quit already waits for exactly this work, in its drain-attaches phase — every
attach is registered with the task queue from enqueue. The helper now runs
the same drain, in quit's order, so it waits for both producers that reach
the store after the last awaited call returns.

Adds a regression test that holds the loop's restart inside its provider
acquisition and asserts abandoning does not return until it lands.

* chore: re-trigger PR checks

The push to 2ecc9c6e emitted no pull_request event, so the matrix never ran.

* fix(sidebar): clip worktree card content to its border (#23566)

* fix(runtime): answer terminal.subscribe at once for a pane the desktop already has mounted (#23512)

* fix(runtime): answer terminal.subscribe at once for a pane the desktop already has mounted

A mobile subscribe to a PTY with no headless model asked the renderer to
mount its tab and waited for a newer serializer settle. The renderer drops
mount requests for tabs it already has mounted, so a reattached daemon PTY
whose restored provider snapshot outranked the live renderer held the reply
for the full 3 s deadline. A live renderer screen now proves attachment and
is adopted directly; an unmounted pane still requests the mount and waits.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): treat any renderer answer, even a blank screen, as an attached pane

A fresh shell that has printed nothing has a registered serializer and an
empty screen; requiring non-empty data sent it back through the dropped
mount request and the 3 s wait. A blank screen skips the wait but does not
replace the chosen snapshot, so a parked pane cannot erase provider history.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): treat a mounted pane with unsettled output as attached

The stable renderer snapshot returned null both when no renderer answered
and when output advanced under every retry, so a desktop pane printing
continuously still took the dropped mount request and the 3 s wait. It now
returns a typed outcome (settled, moving, absent); moving skips the mount
and publishes the chosen snapshot, and late recovery still requires settled.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): adopt only a renderer-ordered screen when probing a mounted pane

The attachment probe read the terminal before knowing it would adopt, which
can reach the provider snapshot on the unmounted path; the read now follows
the decision. A seq-less renderer screen would replay every buffered chunk
on top of itself, so the probe keeps the chosen snapshot for it. The probe,
adopt and mount wait move into their own module to stay under the line cap.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): adopt a seq-less renderer screen when no output is pending

The seq gate only prevents a double replay of buffered output, so a settled
non-blank screen without a seq is safe when nothing is pending. That keeps
the better screen for a pane right after a deferred cold restore, before it
is renderer-ordered. The rule now applies after the mount wait as well.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(runtime): decide renderer attachment from the host's serializer flag

The mounted-pane probe re-derived attachment by serializing the renderer up to six
times, which cost ~7.5 s for a registered but unresponsive renderer on a busy PTY.
The host already holds that fact in the serializer readiness flag. The flag is never
cleared when a pane closes over a live PTY, so one null serializer answer falls back
to the mount wait: worst case is the old 3 s plus one 750 ms serialize.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(runtime): let the serializer's answer alone prove renderer attachment

serializeRendererTerminalBuffer already answers null when the host's serializer flag is
unset, so the separate flag accessor was redundant. The numeric-seq adopt test now
replays only the byte past the seam.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* Revert "perf(ci): shard the anti-slop audit across processes instead of one JS runtime (#23543)" (#23575)

This reverts commit fae0ae7a464c2991bd25885e7e61689c7f471a0f.

* perf(ci): cache pnpm verification records on Linux (#23568)

* perf(ci): pilot pnpm verification record caching on Linux

* test(ci): review pnpm verification record in mobile cache audit

* fix(native-chat): the host writes chat failures for a person, with a typed fact beside them (#23116)

* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* feat(native-chat): a typed failure fact beside every failure sentence

Adds the shared vocabulary the host writes a failure with: a closed failure kind, a
provider diagnostic that says who it is for (a person, or a log), and a refusal cause
beside the refusal code. Status rows gain an optional failure fact and rejected
submissions an optional rejection fact; the dispatch row carries it, the reducer reads
it field by field, and the projection forwards it. Older rows and older readers are
untouched: every field is optional and the schemas stay open.

* fix(native-chat): durable failure rows and rejection reasons are written for a person

Every host writer that records a failure now writes a sentence for a person beside a
typed fact, instead of embedding a refusal's message, an exception or a composed exit
string. A provider's own words travel as a separate diagnostic from the places Orca
composes them - the Claude and Codex exit stderr (a log), Codex's JSON-RPC message,
Claude's compact_error and Codex's turn error (for a person) - and are never inferred
from a string afterwards. Not signed in and oversized history are typed at the
adapter that detects them.

Covers start and restart failures, the delivery loop, dispatch rejections (content,
queue-full, write failures, provider refusals), cancel and answer confirmation rows,
compaction, the rewind fallback, and not_delivered, which released clients printed
as it was. Two leaks close on the way: a settlement retry no longer writes Orca's
probe evidence into the exit row, and an attach or journal-sink failure is recorded
as Orca's fault rather than as the provider stopping. The legacy rejection markers
and the reasons on sends in doubt stay byte-identical.

* feat(native-chat): refusals name their cause, and a failed start is worded in one place

A refusal now carries an optional cause beside its code: one closed enum of the situations
a chat write can meet, set at every emitter a structured-chat write reaches. Returned
refusals build it with refuse(code, cause, message). Store and host paths that raised a
bare Error(code) now throw AgentSessionRefusalError, whose message is still the code and
which has no code property; the RPC error mapper handles it before any other passthrough,
keeps today's wire code and message byte-identical, and adds { refusal: { code, cause } }
to the error's data. The hold throws it, and restart-resume files the cause beside the
unchanged reason. The operation ledger stores the cause beside the code, so a replay names
the same situation as the first answer. The sto…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants