Skip to content

Beyond VS Code: host API gate, 1.99 floor, the ACP agent and its editors (M60–M63) - #32

Draft
RandyNorthrup wants to merge 44 commits into
mainfrom
claude/vigilant-euler-75cvbd
Draft

RandyNorthrup wants to merge 44 commits into
mainfrom
claude/vigilant-euler-75cvbd

Conversation

@RandyNorthrup

@RandyNorthrup RandyNorthrup commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

What and why

The IDE compatibility programme (docs/ide-compatibility.md, PLAN.md D60–D62, M60–M66), built beside M45–M56 so Muse Spark reaches as many editors as it can:

  • M60/M61: npm run check:host-api records every VS Code API, Node built-in and theme variable the extension uses (docs/ide-compatibility/host-api.md). It fails when the record goes stale or portable code reaches vscode. The webview now talks to VS Code through a single host bridge.
  • M62: engines.vscode goes from ^1.125.0 to ^1.99.0, based on an API audit (the host needs nothing newer than 1.85) and a Node audit (1.99 runs Node 20.18). Host bundles now target node20.18. Tested in:
    • VSCodium 1.99.3 and 1.135 (integration tests);
    • code-server 4.99.4 and Eclipse Theia 1.75. Theia's sidebar stays blank on first open; the workaround is under Troubleshooting;
    • in CI, the latest Cursor 3.22.7 (VS Code 1.128), Devin Desktop 3.10.35 (formerly Windsurf, VS Code 1.126), Kiro 1.1.70 (VS Code 1.131) and Positron 2026.09.1 (VS Code 1.130). Each installs the VSIX and passes the 9 integration tests.
  • M63: the agent package muse-spark-code-acp, an Agent Client Protocol agent over both backends:
    • The Model API key lives in the operating system's credential store (D61), with no plaintext fallback.
    • The package is attached to each GitHub Release. The release workflow publishes it to npm, and the VSIX to Open VSX, when their tokens are set.
    • Tested in Zed 1.20, Emacs (agent-shell), Neovim (CodeCompanion) and JupyterLab (Jupyter AI). Setups are in docs/acp.md, including a Windows form (node <script>) for editors that cannot start npm's .cmd launcher.
    • The editor's stdio and HTTP MCP servers are passed to Muse Code.
    • Paid features: --web-search and --image-generation turn on web search and image generation, Model API backend only. At the first prompt the agent asks for each in the editor, naming the price, and only "Turn on" turns it on. Paid rows and approvals show their price, and the agent log counts each billed use.
  • CI for the hosts (test/hosts/):
    • hosts.yml runs on every PR that touches the product and every Monday:
      • the agent's npm package on Linux, macOS and Windows, including the key's round trip through each credential store;
      • VSCodium and code-server, at the floor and the latest release;
      • Theia;
      • JupyterLab, Emacs and Neovim.
    • forks.yml runs every Monday and by hand, covering Cursor, Devin Desktop, Kiro and Positron.
    • Both have passed on every run so far.
  • Bugs found and fixed along the way:
    • The keyring binding returns null for a missing key, so auth status reported a key that wasn't there.
    • The Model API backend's rejected status showed as completed in the editor.
    • Loading a held session again leaked the old hold.
    • docs/acp.md's Zed example was missing "type": "custom".

Merge order. main is merged in through M47: session goals, background work and the user shell, and workflows. Each was merged with a merge commit, never a rebase. This PR stays a draft until M48–M56 land, and each one is merged the same way. Notes from the merges so far:

  • M46's new Promise.withResolvers in ModelApiHost was replaced, because VS Code 1.99 and 1.100 run Node 20. The 1.99 floor's typecheck caught it.
  • "Edit automatically" now lives in core/agent/approvalRules.ts with M46's replayed-request guard, shared by the panel and the ACP agent. The panel keeps M46's one-panel condition.
  • Goals, background tasks and workflows have no ACP counterpart yet.
  • File access through the ACP client (fs/*, so the Model API backend sees unsaved buffers) waits for M48–M56, because it needs the session threaded through the Model API backend's tools.

Records with red drills:

  • docs/certification/m60.md;
  • m62.md: fork drills F1–F2;
  • m63.md: host drills H1–H5, MCP drills S1–S5, paid-feature drills P1–P6.

Checklist

  • npm run quality is green locally, with two exceptions:

    • fsAtomic.test.ts › "refuses a read-only file at once" fails when run as root (the development container). M60's record shows it fails the same way on main, and every gate before it passes.
    • Semgrep did not run: the container's proxy refuses semgrep.dev. CI's semgrep job runs it.

    After the M47 merge, the unit suite as an unprivileged user gives 1,784 passed, with coverage of 96.39 % statements and 91.71 % branches. build, security:audit, security:secrets and test:a11y (280 pages) all exit 0.

  • Tests added or changed with the code; each new check was seen to fail once on purpose (drills in the certification records).

  • CHANGELOG.md updated under Unreleased, along with the README, docs/acp.md and docs/PRIVACY.md (which now covers the ACP agent).

  • New dependencies, each with its reason in PLAN.md: @agentclientprotocol/sdk 1.4.0 and @napi-rs/keyring 2.1.0 (D61, D62), ovsx 1.2.0 (Open VSX publishing), playwright-core 1.63.0 (dev, host checks). No suppressed rule or any added.

  • Nothing in the diff contains a credential (gitleaks staged scan on every commit; security:secrets clean).

> muse-spark-code@0.8.0 check:l10n
l10n: 14 tables, 81 manifest strings, 205 source files; 0 problems
> muse-spark-code@0.8.0 check:host-api
host-api: 198 VS Code APIs, 11 files importing vscode, 14 Node built-ins, 57 theme variables; 0 problem(s)
> muse-spark-code@0.8.0 duplication
Found 0 clones.
> vitest run --coverage   (as nobody)
      Tests  1784 passed | 7 skipped (1791)
All files          |   96.39 |    91.71 |   95.96 |   96.39 |

🤖 Generated with Claude Code

https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg

claude and others added 29 commits September 26, 2026 01:46
… bridge (PLAN.md D60)

The owner's IDE compatibility plan, built beside M45–M56 and numbered
from 60 so both keep their numbers.

- PLAN.md: D60 (one product, four families; the rulings carried into
  every adapter: the key, paid features, approvals, editing claims),
  Q60–Q64, M60–M66 and the gate row. The plan itself is filed in
  docs/ide-compatibility.md.
- M60: scripts/check-host-api.mjs (npm run check:host-api, in
  quality:gates; --write regenerates) keeps
  docs/ide-compatibility/host-api.md: the manifest and build targets,
  the 198 VS Code APIs used at run time and where (found with
  TypeScript's checker, including provider members, options fields and
  VS Code objects handed to code that takes them by shape), the 11 files
  that import vscode, the Node built-ins, acquireVsCodeApi and the 57
  theme variables. It fails when the record is stale, and whenever
  src/core, src/shared, src/webview or a listed portable host module
  reaches vscode through any import, type-only ones included.
- M61 steps 1–2: src/webview/hostBridge.ts is the webview's one way to
  its host; ChatSurface and ConversationMessage move to a vscode-free
  module, createLogger takes the channel by shape, and DictationSetup
  moves to the core, so the conversation controller, both backend
  managers and the tool harness are portable. No behaviour change.

Drills A–K in docs/certification/m60.md. In this container the unit
suite passes as an unprivileged user (1,484 tests, coverage thresholds
met); as root the read-only test in fsAtomic.test.ts fails, as it does
on main. gitleaks is clean (history and staged); semgrep and the
integration tests were not run here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
… npm publishing (PLAN.md D61, D62)

muse-spark-code-acp runs the panel's two backends as an Agent Client
Protocol agent for Zed, JetBrains IDEs, Neovim, Emacs and the other ACP
editors (src/acp, src/runtime; @agentclientprotocol/sdk 1.4.0). Tool
calls with diffs, plans, permission prompts (a cancelled or unknown
answer rejects), questions as forms, modes, model and effort, skills as
commands, and sessions listed, loaded and resumed. The backend is chosen
at launch; paid features stay off.

D61: `auth set|status|clear` keeps the Model API key in the OS credential
store through @napi-rs/keyring 2.1.0 (Secret Service only on Linux, no
plaintext fallback). Checked live against GNOME Keyring 46.1, which found
that the binding returns null for a missing entry despite its typings.

The package is built and checked in CI, attached to each release, and
published to npm when NPM_TOKEN is set; the VSIX is published to Open VSX
when OVSX_PAT is set (ovsx 1.2.0).

The new tests also found that the Model API backend's `rejected` status
was shown as completed, and that a session loaded twice kept the old
hold. Both are fixed. Drills A-P and the gate run are in
docs/certification/m63.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
The owner signed the Eclipse Publisher Agreement and created the Open VSX
namespace, so the next tag publishes there. npm has held the owner's
account for suspicious activity, so docs/acp.md now gives the install from
the GitHub Release's URL and says the package is not on npm yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
engines.vscode goes from ^1.125.0 to ^1.99.0, so editors built on VS Code
1.99 or later can install the extension. The host typechecks against
every published @types/vscode from 1.85 on. VS Code 1.99 and 1.100 run
Node 20.18/20.19, so the host library is ES2023, its bundles target
node20.18 (the ACP agent keeps node22), and the one Node 22 API in the
host, Promise.withResolvers, is replaced. The unit tests' panel fake is
typed from the interface so it holds at every version.

Tested in VSCodium 1.99.3 and 1.135 (the integration tests, 9 passing in
each) and in code-server 4.99.4 (VS Code 1.99.3, Node 20.18.3: a
conversation and an approval against the fake CLI), which refused the
same VSIX with the 1.125 floor. Record: docs/certification/m62.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Theia 1.75, built from npm as a browser app, runs the panel, a
conversation and an approval against the fake CLI. It claims VS Code API
1.134. Found: Theia never fires onView: for a webview view, so the
sidebar opened first stays blank until a command or the tab starts the
extension. README's Troubleshooting gives the shortcut (Ctrl+Esc); the
fix belongs in Theia, since onStartupFinished would start the extension,
and read SecretStorage, in every VS Code window.

The owner set NPM_TOKEN; the package name was free on 2026-09-26, so the
next tag publishes muse-spark-code-acp to npm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Emacs 29.3 with acp.el 0.15.2, shell-maker 0.97.3 and agent-shell 0.79.2,
and Neovim 0.11.4 with CodeCompanion v19.25.0, each drove the packaged
agent against the fake Muse Code CLI: the modes, the model and effort, a
streamed reply, a tool call allowed and one rejected through each
client's own prompt (agent-shell's y and C-c C-c, CodeCompanion's g2 and
g3). Nothing in the agent changed. docs/acp.md now carries the tested
configuration for both; hosts.md marks them Preview.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Zed 1.20.2 (its Linux release, software Vulkan on Xvfb, driven with
xdotool) listed Muse Spark under External Agents; its thread showed the
model and effort selectors, streamed the reply, and ran or skipped a
command from Zed's permission card. docs/acp.md's Zed example gains the
"type": "custom" field Zed now requires; hosts.md marks Zed Preview.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Jupyter AI 3.2.0 runs ACP agents as chat personas, so JupyterLab 4 is
reached through the agent now rather than waiting for M65's native
extension. With JupyterLab 4.6.3 and a local persona file (its name must
contain "persona"), the chat showed the agent's model, mode and effort
pickers and its context gauge, and allowed and rejected a command from
its buttons. Found: Jupyter AI passes its notebook tools as MCP servers,
which the agent does not forward yet (M63c, next).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
The engine's per-session MCP servers gain a stdio kind beside HTTP (MSP
takes both), and the ACP agent passes the editor's stdio and HTTP servers
to Muse Code on session/new, load and resume when the host granted
sessionMcp, each optional; it advertises mcpCapabilities.http on that
backend. SSE, the unstable ACP transport and repeated names are left out,
the Model API backend runs none, and only server names are logged.

JupyterLab's notebook tools (Jupyter AI's HTTP MCP server) now reach the
agent. Drills S1-S5 in docs/certification/m63.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Every host run by hand in M62 and M63 is now a script in test/hosts/,
and the Hosts workflow runs each against the fake Muse Code CLI:

- the agent's npm package on Linux, macOS and Windows: the stdio suite
  against the installed package (MUSE_ACP_PACKAGE_DIR), then the key's
  round trip through each OS credential store (keystore.sh)
- VSCodium 1.99.32846 and latest: the integration tests
- code-server 4.99.4 and latest, Eclipse Theia 1.75.0: the VSIX, a
  reply, a command allowed and one rejected, driven in Chrome
  (playwright-core 1.63.0, a dev dependency)
- JupyterLab 4.6.3 with Jupyter AI 3.2.0, Emacs (acp.el v0.15.1,
  agent-shell v0.77.4), Neovim 0.11.4 with CodeCompanion v19.25.0

Pull requests and pushes to main that touch the product, every Monday,
and by hand. Drills H1-H5 in docs/certification/m63.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Cursor, Devin Desktop (Windsurf's new name), Kiro and Positron, each at
its latest release from its own update feed, the one its nixpkgs update
script or Homebrew cask reads (test/hosts/fork-release.mjs). run-fork.sh
unpacks the package without installing it (AppImage, .deb, tar), prints
the fork's version and VS Code base into the job summary, installs the
VSIX with the fork's CLI and checks it is listed, then runs the
integration tests in the fork. FORK_URL tests a given package instead.

The feeds are refused in the container: the three package formats were
exercised with VSCodium 1.99's AppImage, .deb and tarball (9 passing
each); drills F1 (the old 1.125 floor refused) and F2 (no product.json)
in docs/certification/m62.md. The VSCodium config becomes
installed.vscode-test.mjs, shared by both workflows.

Weekly, by hand, and on pull requests that change the check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Conflicts: MuseCodeHost.test.ts takes main's snapshot resume with this
branch's stdio MCP server in its config; conversationController.test.ts
keeps DictationSetup from core/voice/dictation (M61) with main's
GoalCommandVerb. The ACP agent's test fake gains AgentSession's new
controlGoal. The agent passes goal events over (no ACP counterpart yet).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Forks run 36222382949 passed in all four: Cursor 3.22.7 (VS Code
1.128.0), Devin Desktop 3.10.35 (1.126.0), Kiro 1.1.70 and Positron
2026.09.1 (1.130.0), each installing the VSIX and passing the 9
integration tests; Hosts run 36222382972 passed its 12 jobs. hosts.md
moves the four forks to Preview and fixes the stale Planned rows for Zed
and Neovim.

fork-release.mjs took Kiro's own 1.1.70 for its VS Code base: describe
now takes only a VS Code release number (1.NN.N) and otherwise lists
product.json's version fields.

PRIVACY.md gains the ACP agent: what it sends, the editor's MCP servers,
the key in the OS credential store, where Model API sessions live, the
redacted stderr log. PLAN.md records the playwright-core pin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
npm installs muse-spark-code-acp on Windows as a .cmd launcher, which
some editors cannot start (Node refuses to spawn a .cmd without a shell).
The guide now gives `node "<npm root -g>\muse-spark-code-acp\dist\acp.js"`
as the command that works everywhere; it is the form the Hosts workflow's
packaged suite already runs on Linux, macOS and Windows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Web search and image generation on the Model API backend, behind
--web-search and --image-generation (refused with the Muse Code
backend). At the first prompt the agent asks for each in the editor: a
row "Turn on Web search?" with what it does and its price, and a
session/request_permission offering "Turn on" and "Keep off". Only
"Turn on" turns it on, for the life of the process; a refusal, a
cancelled question or a client that cannot answer leaves it off and it
is not asked again. Two prompts at once share one question, and a
cancel while it is asked ends the prompt without a turn.

The Model API backend reads isPaidFeatureOn from src/acp/paid.ts and
logs every billed use with a running total. Paid rows and paid
approvals name their price in the title. New strings in all 15
languages; docs/acp.md, PRIVACY.md, PLAN.md (D62's condition met) and
the changelog updated. Drills P1-P6 in docs/certification/m63.md.

File access through the client (fs/*) waits for M46-M56.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
The second Forks run (36223213890) listed Kiro's product.json version
fields: its base is vsCodeVersion, 1.131.0. describe now reads that key
too; hosts.md, PLAN.md and the M62 record carry the version. All four
forks passed again, 9 integration tests each.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
…patibility branch

Conflicts:
- conversationController.ts: M46 guarded "Edit automatically" with one
  panel holding the session and no replayed request. The rule stays in
  core/agent/approvalRules.ts, shared with the ACP agent, which now also
  refuses a replayed request; the one-panel condition stays in the
  controller.
- PLAN.md: M46's D39 and this branch's D60-D62, both kept.

After the merge:
- ModelApiHost's new Promise.withResolvers (M46) is a new Promise:
  VS Code 1.99 and 1.100 run Node 20 (M62). The host's ES2023 library
  refused it at typecheck, as M62's check intends.
- The ACP agent's test fake gains M46's five AgentSession methods.
- docs/ide-compatibility/host-api.md regenerated: 23 commands and 7
  keybindings; no new VS Code API.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Conflict: README.md's Troubleshooting, where M47's two workflow entries
and this branch's Theia entry were added at the same place; all three
kept. Nothing else needed: no new AgentSession member, no VS Code API
newer than the 1.99 floor, no Node 20 gap; host-api.md unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Joins origin/main 890e371 (PRs #35-#44) onto PR #32's head 68d450e.

Conflicts, each kept on both sides:
- agentBackend.ts: the PR's stdio MCP server union with main's side-chat
  options on startSession, resumeSession and forkSession.
- webviewSetup.ts / chatSurface.ts: ChatSurface stays in the PR's
  vscode-free chatSurface.ts and gains main's isSideChat.
- conversationController.test.ts: DictationSetup from core (the PR's move)
  plus main's new constants.
- PLAN.md: D41-D47 then D60-D62; main's gate table and notes with the
  PR's Cycles entry and its two rows. README, AGENTS, CHANGELOG and the
  certification index: both sides in milestone order, the PR's CHANGELOG
  entries under [Unreleased].

Integration fixes:
- The ACP runtime takes M48-M56's new manager inputs: Muse Code's default
  sandbox network, the in-memory prompt cache, M49's memory store, no paid
  subagents, and the shell job's C# (now shipped in the agent's package).
- MUSE_LOGIN_ARGS is back for the agent's `login` (M55 removed it because
  the panel signs in by device code).
- The fake session's steer returns a TurnSubmission, as main's
  interface does.
- The 1.99 floor: M55's two Promise.withResolvers became settleOnAbort
  (Node 20). The host typechecks at @types/vscode 1.99.0 and
  @types/node 20.19. The host API record was regenerated.
- M56's premise at older VS Code: fetch is routed from the floor on, but
  WebSocket only from 1.112.0. Diagnostics now says per global whether the
  editor routes it. The README, the D43 amendment and the certification
  record in m62.md cover it.
- acpRuntime.test.ts no longer finds a Muse Code installed on the
  machine running it.

Gates (Windows 11, Node 24.20.0): format:check, lint, typecheck,
check:l10n, check:host-api, deadcode, cycles, duplication and test:unit
(2,591 passed, 5 skipped) exit 0. The ACP stdio suite passes in-repo and
against the packed agent. build exits 1 at the size check only:
extension.js 600.5/600 KiB, acp.js 874.1/800 KiB. No budget was raised;
the choice is recorded in PLAN.md section 7 for the owner.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Open VSX refuses a publish into a namespace that does not exist, and RandyNorthrup had none (GET /api/RandyNorthrup answered 404 on 2026-09-27). The Open VSX job now creates it with OVSX_PAT when it is missing and skips when it exists; any other answer stops the job.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Main gained the 0.9.0 release fixes, 0.9.1, M57 (the Model API backend
as its own bundle, dist/modelApi.js) and M58 (a popup before every paid
use). Seven files conflicted, all prose or lists (AGENTS, CHANGELOG,
PLAN, README, the certification index, package.json's cycles entries);
both sides are kept. PR #32's changelog entries move out of [0.9.0],
where git placed them, into [Unreleased].

Adapted where the two sides met:

- The ACP agent loads the Model API backend from dist/modelApi.js beside
  acp.js, the file the .vsix ships, and its package now ships it: one
  build of the backend for both packages (PLAN.md D6 amendment).
  check-bundle-split fails when acp.js carries the backend; the host API
  gate lists modelApiEntry.ts as portable; the agent's notices take
  modelApi.js's packages. acp.js: 874.1 -> 713.2 KiB.
- M58 in the agent (D62 amendment): a paid feature is on only with its
  flag, and each use asks over session/request_permission in the
  conversation it is for, with the panel's words (paidUseQuestion, moved
  into paidConsent.ts) and allow_once / allow_always / reject_once.
  "Always" only with --trust-workspace, kept per folder in the agent's
  data folder (acp/paid-uses.json) and forgotten when the agent starts
  without the flag. No session or client to ask, a cancel, or an unknown
  option denies; subagents are always denied. The first prompt's
  "Turn on" question and its three strings are gone.
- ModelApiHostDeps.allowsPaidUse names the conversation (sessionId, a
  child's being its parent's), since one host serves every conversation
  of a folder.
- build.mjs: modelApi.js targets HOST_NODE_TARGET (node20.18); the test
  helper builds for the same target. fakeModelApi no longer uses
  Promise.withResolvers, which M57's integration test would call on
  VS Code 1.99's Node 20.18. src/acp uses isPromptSettledError (M57's
  lint rule). The approval title drops M58's removed paidFeature.
- The host API record regenerated (200 APIs, 17 Node built-ins).

Tests, drills P1-P9 and B1-B3, the packed agent's stdio suite and the
integration run (10 passing on 1.139.1 and on 1.99.0) are in
docs/certification/pr32-integration.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
dist/acp.js is 713.2 KiB now that the agent loads the Model API backend
from dist/modelApi.js (874.1 KiB, over its 800 KiB budget, on the tree
joined before M57). The budget is the measured size plus about 15 %,
rounded up to a multiple of 50 KiB. Zod is 445.2 KiB of it, 257.6 of
that the locales of the classic API the ACP SDK imports. The agent is
installed once and never loaded by VS Code, so its size is a download
(175.1 KiB gzipped), not a start-up cost. No other budget changes; the
extension stays at 432.5 of 600 KiB.

PLAN.md's open budget question (section 7) is recorded as resolved.
Drill: a 700 KiB budget fails the size check.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The agent runs outside VS Code, so VS Code's http.* settings and
museSpark.environmentVariables do not reach it. Measured against a local
proxy (nothing left the machine) on Node 22.0.0 to 24.20.0: the agent's
own fetch (the Model API backend) ignores HTTPS_PROXY unless
NODE_USE_ENV_PROXY=1 is also set, which Node honours from 22.21 and 24.0
(NODE_OPTIONS=--use-env-proxy from 22.21 and 24.5), reading HTTPS_PROXY,
falling back to HTTP_PROXY, and NO_PROXY, not ALL_PROXY. NODE_EXTRA_CA_CERTS
works on every version; --use-system-ca is accepted from 22.15. Muse Code,
started by the agent, reads the proxy variables itself (M56) with loopback
kept off the proxy.

The agent does not turn Node's switch on by itself, and its network-failure
advice still names VS Code's settings: recorded as PLAN.md Q66 for the
owner rather than changed here.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The key's round trip through the OS credential store (test/hosts/
keystore.sh, a made-up key) had run on GitHub's runners and by hand only
through gnome-keyring. With the agent packed from the merge and a
portable Node 22.23.3 it passed on the Windows 11 VM (Credential Manager,
in the owner's console session, through a one-off scheduled task) and on
the Mac mini (a keychain of the run's own, the owner's keychain settings
put back). Nothing is left on either rig.

Found: over SSH with a key, Windows refuses Credential Manager
(ERROR_NO_SUCH_LOGON_SESSION); the agent says the store cannot be used and
stores nothing. docs/acp.md now says to store the key from a desktop
session.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The agent now keeps the paid features allowed always in a folder in
acp/paid-uses.json beside its sessions (folder hashes and feature names
only), asks before each paid use rather than once for its price, and
reaches api.meta.ai through Node's fetch, which uses a proxy only when
its environment asks for one. Say so where the privacy notice describes
the agent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The first local semgrep run over PR #32's code (its container could not
fetch semgrep's rules) flags detect-child-process on the spawn behind
muse-spark-code-acp login. The command is the CLI resolveLaunch found
(install layout, PATH or an absolute --muse-binary), the arguments its
launcher's fixed prefix and MUSE_LOGIN_ARGS, as an array with no shell:
the same launch as muse serve. Suppressed with its reason inline and a
row in the escape-hatch register, as the other spawn sites are.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
npm run quality on 87383c7: exit 0 (2,652 unit tests, every bundle under
budget, a11y 336 pages clean, SAST 0 findings). The first full run failed
at SAST on the agent's login spawn, fixed in 87383c7; that run is the
suppression's drill. The integration run on the merge: 10 passing on
VS Code 1.139.1 and on 1.99.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The owner's ruling: the agent does not re-route through a proxy or add
undici; it says so, and its advice names what works outside VS Code.

- At start, on the Model API backend, one log warning when HTTPS_PROXY or
  HTTP_PROXY (either case) is set and Node's switch is off (only
  NODE_USE_ENV_PROXY=1, or --use-env-proxy in execArgv or NODE_OPTIONS,
  turns it on), or when this Node has no switch at all (below 22.21, 23):
  the requests go to Meta directly. Names only, never an address; the two
  spellings are one variable on Windows. Nothing on the Muse Code backend.
- A request that never reaches Meta: M56's classifier takes a
  NetworkAdvice, and the runtime's manager passes 'agent' down to the
  bundle's client, so certificate, proxy-credential and unreachable
  failures name NODE_EXTRA_CA_CERTS, --use-system-ca, HTTPS_PROXY and
  NODE_USE_ENV_PROXY instead of VS Code's http.* settings. Three new
  strings in en.ts and the 14 tables; the extension's advice is unchanged.
- The runtime takes sleep from main.ts, so the retries run instantly in
  tests; the fake Model API can fail a fetch with a socket code.

PLAN.md Q66 resolved and D62 amended; docs/acp.md and CHANGELOG updated;
tests and drills Q1-Q10 in docs/certification/pr32-integration.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
npm run quality on 83a4530: exit 0 (2,674 unit tests, 252 source files
through the l10n gate, acp.js 715.7 of 850 KiB, SAST 0 findings).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
RandyNorthrup added a commit that referenced this pull request Sep 28, 2026
Codex's second round on PR #50 found three issues; three class reviewers
then swept D49, D50 and M67-M85 for their siblings, and everything they
found is fixed here in one commit.

- TypeSafe is billed to its own key: AGENTS.md rule 12 names that one
  exception, and D50/M85 keep every other part of the paid invariant
  (PaidFeatureGate, badge, paid row, PaidUsage, the D48 popup).
- M75 is built first in wave 3; M73 and M74 land with a passing run and
  have no reachable path or setting before it.
- M72 keeps ignored files the turn changed, in a shadow repository only,
  and says exactly what it could not restore.
- Rules conflicts: D4's network list amended, Restricted Mode (no git, no
  shell) stated, the user's own turn defined, paid best-of-N on the Model
  API only, the CLI's settings never written by M83, no invented file
  names (M76, M79), GitLab moved to "not taken".
- Order and promises: verify loop first, M80 last (waits for PR #32),
  table rows match their sections, every milestone has Acceptance and
  Tests, the log reducer lives in M73, ported-code notices go through the
  generator.
- Security: tools on the ide server, untrusted content, automatic actions
  following the mode, M44b's fetch rules restored, check-command option
  injection, PR worktrees with project config off, M78's parser-based
  allow rules, M80's trigger and secret handling, M81's pipe and
  loopback-only browser, M83/M84 redaction and import limits.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
RandyNorthrup added a commit that referenced this pull request Sep 28, 2026
…e CI key's one way in

Codex's fourth round on PR #50:
- The extension cannot ask VS Code about a folder's trust before it
  opens, and a trusted parent (the home folder) can trust its storage, so
  an unauthored PR's conversation is held in Plan mode with its project
  configuration off until the user confirms trust in the extension's own
  card, whatever VS Code's trust says.
- M80's key invariant now says what is true: its one way in is auth
  set's standard input; in CI the Action's step shell is the only
  environment it is ever in, which PR #32's rule 8 amendment names; no
  process the agent or exec starts gets it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
RandyNorthrup and others added 8 commits September 28, 2026 07:58
Main's plan for the next milestones (D49, D50, M67-M85) joins PR #32's
D60-D62 and M60-M66. PLAN.md: D49 and D50 before D60; AGENTS.md rule 12
keeps both the ACP agent's paid-use sentence and D50's planned TypeSafe
exception.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
fix/cli-sign-in-detection at 2a324d4: the CLI's sign-in read from the
credential file's structure, account/read where only the CLI can say,
sign-out through account/logout. Conflicts in the certification index,
report.ts's imports and deviceSignIn.ts, both sides kept.

- Node 20: PR #49's deviceSignIn.ts and accountHost.ts used
  Promise.withResolvers, which VS Code 1.99/1.100 (Node 20.18, M62) lack;
  the host typecheck at ES2023 refused it. accountHost races the handshake
  with unlessAborted, the device sign-in stops through its own controller
  and settleOnAbort. PR #49's 220 sign-in tests pass unchanged.
- The ACP agent's Muse Code readiness moves off credentialFileExists()
  (gone, and wrong after muse logout, which leaves the file emptied) to
  PR #49's CliAccount: META_API_KEY, else the file's verdict, else
  account/read on a short-lived host. unsupportedHere is 'cannot run'
  with the panel's sentence; authenticate forgets earlier answers.
  Tested over stdio (sign-in asked after muse logout) and in process
  against the fake CLI for every verdict.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Outside VS Code (the ACP agent) the operating system's credential store
stands in for SecretStorage, the store SecretStorage itself rests on: the
key goes in only through auth set's standard input, never from an
environment variable, an argument or a file, and never reaches a child
process. The one named exception is M80's planned CI bootstrap: the
Action's step shell is the one environment the key is ever in, piped to
auth set and unset before exec. D61's Never list and M80's note say the
same; CHANGELOG.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Three read-only reviews of the diff since 2559a9b, one per class, fixed
in one round:

- Grants: 'always' adds only its feature, merged inside the write queue,
  so a stale set is never written back; a change fails on a file that is
  there but unreadable instead of writing over it.
- PaidUseConsent (both hosts): an 'always' that cannot be kept lets the
  use go ahead once, logged as such, instead of failing it.
- Grants lapse at start only on the Model API agent; a Muse Code agent
  beside it no longer clears them.
- The client's permission answers are parsed with zod (rule 7).
- Paid-use row ids are UUIDs; a prompt is busy, and cancellable, while
  the skills are first announced; a failed form request is declined.
- A missing dist/modelApi.js says to reinstall the agent (new string in
  15 languages); the agent's log names its own key store.
- Docs: D62 sign-in, M63 status, the command name, acp.md's links (npm
  README) and authenticate sentence, CHANGELOG counts, PRIVACY.

Drills R1-R9 and K1 in docs/certification/pr32-integration.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
PR #49 now asks the CLI again on every macOS user action unless the
answer is under 5 s old. The agent's user action is authenticate (the
user saying they signed in): a session or a list reads the file, or on
macOS takes the panel's passive estimate, instead of starting muse serve
(and perhaps a Keychain prompt) on every call. Tested in process with a
runtime reading the file as macOS does, on any runner; drill R10.

The host API record regenerated for PR #49's code: 201 APIs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A forced Model API backend asks the CLI nothing, and a same-account
Keychain re-sign-in counts. No conflict; nothing the ACP agent uses
changes (it chooses its backend by flag and never runs the device
sign-in).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
npm run quality: exit 0 (2,873 unit tests, acp.js 724.8 of 850 KiB, SAST
0 findings); test:integration 10 passing on VS Code 1.139.1 and 1.99.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The same tree as PR #49's head 5184f26, merged here before; no file changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review 🔄 Running since 2026-09-28T16:30:14.505556Z a209130 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-28T18:07:33.046116Z 4eb0156 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a209130aae

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/runtime/backends.ts
Comment thread src/acp/agent.ts Outdated
Comment thread src/acp/questions.ts Outdated
Comment thread src/acp/agent.ts Outdated
PR #49 (merged in a209130) asks the CLI for the sign-in (account/read),
and the fake answers from the credential file's contents: a `meta` entry
holding `access_token` is a browser sign-in, anything else is signed out.
test/hosts/fake-muse.sh still wrote the old placeholder {"fake":true}, so
the extension and the agent saw Muse Code signed out and the Hosts jobs
for code-server, JupyterLab ("Authentication required") and Emacs failed.

It now writes the browser sign-in's shape as captured
(test/unit/helpers/credentialShapes.ts), placeholders only. Locally on
a209130: code-server 4.99.4 failed before and passes after; JupyterLab
and Emacs pass with it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg

RandyNorthrup commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner Author

Hosts is red on a209130 in four jobs: code-server 4.99.4, code-server latest, ACP client jupyter and ACP client emacs. In all four, Muse Code is treated as signed out: Jupyter reports "Authentication required: Muse Code is not signed in", and code-server's log says account/read: loggedOut.

Cause. Since PR #49 (merged here in a209130), the extension and the agent ask the CLI for the sign-in state (account/read). The fake CLI answers from what the credential file contains: a meta entry holding access_token counts as signed in. test/hosts/fake-muse.sh still writes the old placeholder {"fake":true}, so the fake answers signed out.

Fix, pushed as 7c0ff0c. fake-muse.sh now writes the browser sign-in's captured shape from test/unit/helpers/credentialShapes.ts, with placeholders only. Local runs on a209130:

Check Before the fix After the fix
code-server 4.99.4 fails, the same as CI passes
JupyterLab not run passes, including the MCP server line
Emacs passes locally twice passes, 10 of 10 checks

Emacs fails in CI but passes locally with the old placeholder, and I haven't yet explained that difference. CI's Jupyter error still names the same cause.

The Codex review requested on a209130 may need requesting again for the new head.


Generated by Claude Code

- Form answers are parsed with zod and must fit their question: an
  offered option, none twice, within the question's bounds; otherwise the
  questions are declined, as each field was required (AGENTS.md rule 7).
- Backend failures in the agent's log go through failureForLog: an MSP
  error by its kind and code, never the CLI's message (rule 8).
- A loaded or resumed session is set to the mode the editor is told
  before anything is replayed; a backend that refuses it fails the load.
- META_API_KEY left as it is: the user's own variable is the CLI's
  documented credential, inherited as the extension does (PLAN.md §1).

Drills and verdicts in docs/certification/m63.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg

RandyNorthrup commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner Author

build / quality (windows-latest) failed on 7c0ff0c in test:a11y: one violation of the target-size rule. In the light/question-explain scenario, the composer's model .pill measured 139.7×21 px, although .pill has min-height: var(--ms-control-size) (26 px). All other gates in that job passed.

Why I'm not changing anything for it yet:

Update: the run for 83833fe, which served as the one re-run, passed Windows a11y along with every other check (24 of 24). I have not found a cause. The 21 px is the pill's content height with min-height not applied, which points to a timing issue in the Windows headless Chrome scan rather than to the stylesheet. If it fails again, I'll pin it down in the harness scenario.


Generated by Claude Code

On top of 83833fe, another session's answer to the same review: its
tests, decline log line and mode-on-load stay; this goes further where
they differ.

- Credential variables: the agent takes every credential variable
  (*_API_KEY and the names hooks never get) out of its own environment
  at start and hands them to Muse Code's processes only, as the
  extension's muse serve inherits the user's META_API_KEY (D1), where it
  still counts as the CLI's credential. The shell tool, hooks, git and
  the Windows helpers never see one. Rule 8, D61, docs/acp.md.
- Logs: every backend, CLI or client failure in src/acp is named by
  failureForLog (PR #49), never by its message; describe() is gone.
- Client answers: the elicitation answer is parsed with zod, and each
  answer against its question (a single choice only an offered option,
  as the form's oneOf; free text only where there are none; distinct
  choices within bounds, at least one when unset, as the panel's card);
  anything else declines. With the permission answers, every response
  the agent asks for is checked.
- Load and resume: the session is set to the mode shown, a listed
  model and the effort shown, or the load fails.

Drills C1-C4, L1, E1-E3, M1-M3 in docs/certification/pr32-integration.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4eb0156c72

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/acp/agent.ts Outdated
Comment thread src/acp/agent.ts Outdated
RandyNorthrup and others added 2 commits September 28, 2026 12:23
- What the agent shows comes from the backend's answer or an explicit
  set, never from a value the handle merely holds (P1). Muse Code's
  resumed handle holds the model the agent asked for, since
  session/resume takes none, while the CLI keeps the one the session last
  ran on. matchAdvertised now asks listModels(sessionId) for the active
  model, as the panel's adopt does, and keeps it only where the agent
  lists it; otherwise it sets the default. The mode and the effort are
  set explicitly on every path; the commands, history and plan are the
  backend's own answers.
- A session that fails setup is let go (P2). adopt replaces register:
  a new, loaded or resumed session is set up before the agent holds it,
  so no request finds it and no id reaches the client until then. Any
  failure disposes it, including a handle whose events cannot be
  followed.
- Grok's review of the first cut (1 P1, 2 P2, all held): both hosts hand
  a held session back retained, so a reload now lets the held wrapper go
  before anything runs on the shared session, and a failed reload holds
  nothing for that id. A wrapper let go decides nothing more on the
  editor's late answers (permission, form, a failed form; a paid use is
  denied), and its running prompt ends cancelled instead of never
  answering. The reload tests hand the held session back, as the hosts do.
- Grok's second look (2 P1, both held): a session let go now stops its
  running turn on the backend once started (release replaces dispose),
  and a load still being set up is let go by a newer load or a close,
  failing rather than being held; a session let go writes nothing more
  to the backend. Closing an id neither held nor being set up is refused.

Drills N1-N2, R1-R6, H1, C1-C3, D1-D6, X1 in docs/certification/pr32-integration.md;
PLAN.md D62, docs/acp.md and CHANGELOG updated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ows it

Grok's pre-Codex look at ca263c5: 1 P1, 6 P2; the P1 and three P2s held.

- A reload waits until nothing holds or sets up that id, each turn
  stopped, before the replacement follows the session: Muse Code hands a
  new listener the prompts still open, which reached it for the turn
  being stopped. It looks again after each wait, as another load may
  have started (P1).
- A released session cancels its turn once the start is answered, even
  a failed start: past its deadline Muse Code's turn/start still runs.
- A released session delivers nothing more to the editor.
- The backend stopping lets go of loads being set up too; they fail.
- The fake session hands a new listener its open prompts, as MuseSession
  does, and the reload test checks turn/cancel goes out first.

Not changed, with reasons in docs/certification/pr32-integration.md: a
failed turn/cancel (no stronger stop through AgentSession), and
cancelQuestions after a release (no await between the check and it).

Drills G1-G4 added; R4 and C1 retargeted; C2 retired with hasStarted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

RandyNorthrup and others added 2 commits September 28, 2026 13:32
…cels

session/cancel sent while sendTurn was still waiting for its answer
reached the backend before the turn existed, so the turn ran to its end,
editing and billing, while the editor was told it was cancelled. cancel()
now waits for the start to be answered, as release() already did. A
cancel notification for a session not held no longer throws inside the
SDK's notification handler.

Found by a read-only Muse Code review (contributor model) while Codex and
Grok Build were at their limits. Drills C1 and C2 recorded.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

46ba540: a read-only Muse Code review (contributor model; Codex and Grok were at their limits) of 4eb0156..496fdee found a P1 (session/cancel while a turn is starting reached the backend before the turn existed, so the turn ran on while the editor was told cancelled) and a P2 (a cancel for a session not held threw inside the SDK's notification handler). Both fixed with drills C1/C2 (docs/certification/pr32-integration.md); full Win11 npm run quality exit 0.

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants