Skip to content

M72: turn checkpoints — restore files, the conversation or both, and redo - #55

Merged
RandyNorthrup merged 177 commits into
mainfrom
feature/m72-checkpoints
Oct 2, 2026
Merged

RandyNorthrup merged 177 commits into
mainfrom
feature/m72-checkpoints

Conversation

@RandyNorthrup

@RandyNorthrup RandyNorthrup commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

M72 turn checkpoints (PLAN.md D51) and the release preparation for 0.10.0.

What

  • Turn checkpoints: restore files, the conversation or both to before a turn, then redo. The records are git refs updated by compare-and-swap, each window has its own index and presence file, and the shadow repository never touches the workspace's .git.
  • The correction batch before the release (docs/certification/m72.md, "Correction batch before 0.10.0"):
    • memory notes, the Memory view and exports keep restore copies and hold the restore lease (an export in either spelling of the workspace folder);
    • the user's ! command is re-checked at its real start (trust, session, Host closing, the user's own Stop);
    • a shell that could not start no longer leaves a false sticky "native uncertainty", and the model is told nothing about a command that never entered;
    • Windows paths over 260 characters: the repository is made in a short folder and published by one rename, and a path git cannot use is refused with a clear message in all 14 languages;
    • the integration tests get a short user-data folder under a long checkout path (macOS socket limit), and the macOS worker cap is typed so typecheck:host passes.
  • Version 0.10.0 with the release notes restructured: a Highlights block first, and one Added, one Changed and one Fixed heading (73 entries, none lost).

History

The head carries the pull request's original commits (the merge a424e52 has e73e854, which holds the previous head 9be57e5, as a parent), so the review threads and authorship stay. The tree is exactly the verified candidate: the content of that history was redone on main 3270944. Both earlier Codex threads are answered and resolved.

Proof

  • Independent verifiers (read-only, three classes) accepted the memory and the shell lanes; their findings are in the table "Verifier findings and what was done" (one declined, one right and fixed: the export lease). The path-limits lane's verifier did not run (the account's session limit ended it); the lead read that diff and found no defect, and the four-machine gate is its evidence.
  • Red drills for every new guard, named in the record with the test each one fails: memory D01-D18, shell S01-S20, path limits N01-N15 (an equivalence drill besides), the test-infrastructure helper, and the export lease. Restoration hashes are given where the lanes recorded them.
  • A literal npm run quality on the code tree of 129099c8 on Kubuntu, the Mac mini, the Windows 11 VM and the Windows host: 246 test files, 3,830 to 3,850 tests passed, statements 94.4 to 94.5 %, a11y 380 pages with 0 violations, no secrets, 0 SAST findings, 0 clones, every size gate ok (extension 591.8 of 600 KiB, Model API 398.5 of 400, checkpoint store 188.5 of 225). Two earlier rounds stopped at cheap steps and were fixed (check:host-api, jscpd). Only documentation differs from the head that is merged here.
  • Hosted CI on this head runs the three operating systems with the integration tests on VS Code stable and on the 1.99 floor.

After the merge

Tag v0.10.0: the release workflow publishes the GitHub Release (VSIX and the ACP package), the Marketplace, Open VSX and npm (muse-spark-code-acp); each channel is checked for the real version and bytes afterwards.

🤖 Generated with Claude Code

claude and others added 30 commits September 26, 2026 01:46
… bridge (PLAN.md D60)

The owner's IDE compatibility plan, built beside M45–M56 and numbered
from 60 so both keep their numbers.

- PLAN.md: D60 (one product, four families; the rulings carried into
  every adapter: the key, paid features, approvals, editing claims),
  Q60–Q64, M60–M66 and the gate row. The plan itself is filed in
  docs/ide-compatibility.md.
- M60: scripts/check-host-api.mjs (npm run check:host-api, in
  quality:gates; --write regenerates) keeps
  docs/ide-compatibility/host-api.md: the manifest and build targets,
  the 198 VS Code APIs used at run time and where (found with
  TypeScript's checker, including provider members, options fields and
  VS Code objects handed to code that takes them by shape), the 11 files
  that import vscode, the Node built-ins, acquireVsCodeApi and the 57
  theme variables. It fails when the record is stale, and whenever
  src/core, src/shared, src/webview or a listed portable host module
  reaches vscode through any import, type-only ones included.
- M61 steps 1–2: src/webview/hostBridge.ts is the webview's one way to
  its host; ChatSurface and ConversationMessage move to a vscode-free
  module, createLogger takes the channel by shape, and DictationSetup
  moves to the core, so the conversation controller, both backend
  managers and the tool harness are portable. No behaviour change.

Drills A–K in docs/certification/m60.md. In this container the unit
suite passes as an unprivileged user (1,484 tests, coverage thresholds
met); as root the read-only test in fsAtomic.test.ts fails, as it does
on main. gitleaks is clean (history and staged); semgrep and the
integration tests were not run here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
… npm publishing (PLAN.md D61, D62)

muse-spark-code-acp runs the panel's two backends as an Agent Client
Protocol agent for Zed, JetBrains IDEs, Neovim, Emacs and the other ACP
editors (src/acp, src/runtime; @agentclientprotocol/sdk 1.4.0). Tool
calls with diffs, plans, permission prompts (a cancelled or unknown
answer rejects), questions as forms, modes, model and effort, skills as
commands, and sessions listed, loaded and resumed. The backend is chosen
at launch; paid features stay off.

D61: `auth set|status|clear` keeps the Model API key in the OS credential
store through @napi-rs/keyring 2.1.0 (Secret Service only on Linux, no
plaintext fallback). Checked live against GNOME Keyring 46.1, which found
that the binding returns null for a missing entry despite its typings.

The package is built and checked in CI, attached to each release, and
published to npm when NPM_TOKEN is set; the VSIX is published to Open VSX
when OVSX_PAT is set (ovsx 1.2.0).

The new tests also found that the Model API backend's `rejected` status
was shown as completed, and that a session loaded twice kept the old
hold. Both are fixed. Drills A-P and the gate run are in
docs/certification/m63.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
The owner signed the Eclipse Publisher Agreement and created the Open VSX
namespace, so the next tag publishes there. npm has held the owner's
account for suspicious activity, so docs/acp.md now gives the install from
the GitHub Release's URL and says the package is not on npm yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
engines.vscode goes from ^1.125.0 to ^1.99.0, so editors built on VS Code
1.99 or later can install the extension. The host typechecks against
every published @types/vscode from 1.85 on. VS Code 1.99 and 1.100 run
Node 20.18/20.19, so the host library is ES2023, its bundles target
node20.18 (the ACP agent keeps node22), and the one Node 22 API in the
host, Promise.withResolvers, is replaced. The unit tests' panel fake is
typed from the interface so it holds at every version.

Tested in VSCodium 1.99.3 and 1.135 (the integration tests, 9 passing in
each) and in code-server 4.99.4 (VS Code 1.99.3, Node 20.18.3: a
conversation and an approval against the fake CLI), which refused the
same VSIX with the 1.125 floor. Record: docs/certification/m62.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Theia 1.75, built from npm as a browser app, runs the panel, a
conversation and an approval against the fake CLI. It claims VS Code API
1.134. Found: Theia never fires onView: for a webview view, so the
sidebar opened first stays blank until a command or the tab starts the
extension. README's Troubleshooting gives the shortcut (Ctrl+Esc); the
fix belongs in Theia, since onStartupFinished would start the extension,
and read SecretStorage, in every VS Code window.

The owner set NPM_TOKEN; the package name was free on 2026-09-26, so the
next tag publishes muse-spark-code-acp to npm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Emacs 29.3 with acp.el 0.15.2, shell-maker 0.97.3 and agent-shell 0.79.2,
and Neovim 0.11.4 with CodeCompanion v19.25.0, each drove the packaged
agent against the fake Muse Code CLI: the modes, the model and effort, a
streamed reply, a tool call allowed and one rejected through each
client's own prompt (agent-shell's y and C-c C-c, CodeCompanion's g2 and
g3). Nothing in the agent changed. docs/acp.md now carries the tested
configuration for both; hosts.md marks them Preview.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Zed 1.20.2 (its Linux release, software Vulkan on Xvfb, driven with
xdotool) listed Muse Spark under External Agents; its thread showed the
model and effort selectors, streamed the reply, and ran or skipped a
command from Zed's permission card. docs/acp.md's Zed example gains the
"type": "custom" field Zed now requires; hosts.md marks Zed Preview.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Jupyter AI 3.2.0 runs ACP agents as chat personas, so JupyterLab 4 is
reached through the agent now rather than waiting for M65's native
extension. With JupyterLab 4.6.3 and a local persona file (its name must
contain "persona"), the chat showed the agent's model, mode and effort
pickers and its context gauge, and allowed and rejected a command from
its buttons. Found: Jupyter AI passes its notebook tools as MCP servers,
which the agent does not forward yet (M63c, next).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
The engine's per-session MCP servers gain a stdio kind beside HTTP (MSP
takes both), and the ACP agent passes the editor's stdio and HTTP servers
to Muse Code on session/new, load and resume when the host granted
sessionMcp, each optional; it advertises mcpCapabilities.http on that
backend. SSE, the unstable ACP transport and repeated names are left out,
the Model API backend runs none, and only server names are logged.

JupyterLab's notebook tools (Jupyter AI's HTTP MCP server) now reach the
agent. Drills S1-S5 in docs/certification/m63.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Every host run by hand in M62 and M63 is now a script in test/hosts/,
and the Hosts workflow runs each against the fake Muse Code CLI:

- the agent's npm package on Linux, macOS and Windows: the stdio suite
  against the installed package (MUSE_ACP_PACKAGE_DIR), then the key's
  round trip through each OS credential store (keystore.sh)
- VSCodium 1.99.32846 and latest: the integration tests
- code-server 4.99.4 and latest, Eclipse Theia 1.75.0: the VSIX, a
  reply, a command allowed and one rejected, driven in Chrome
  (playwright-core 1.63.0, a dev dependency)
- JupyterLab 4.6.3 with Jupyter AI 3.2.0, Emacs (acp.el v0.15.1,
  agent-shell v0.77.4), Neovim 0.11.4 with CodeCompanion v19.25.0

Pull requests and pushes to main that touch the product, every Monday,
and by hand. Drills H1-H5 in docs/certification/m63.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Cursor, Devin Desktop (Windsurf's new name), Kiro and Positron, each at
its latest release from its own update feed, the one its nixpkgs update
script or Homebrew cask reads (test/hosts/fork-release.mjs). run-fork.sh
unpacks the package without installing it (AppImage, .deb, tar), prints
the fork's version and VS Code base into the job summary, installs the
VSIX with the fork's CLI and checks it is listed, then runs the
integration tests in the fork. FORK_URL tests a given package instead.

The feeds are refused in the container: the three package formats were
exercised with VSCodium 1.99's AppImage, .deb and tarball (9 passing
each); drills F1 (the old 1.125 floor refused) and F2 (no product.json)
in docs/certification/m62.md. The VSCodium config becomes
installed.vscode-test.mjs, shared by both workflows.

Weekly, by hand, and on pull requests that change the check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Conflicts: MuseCodeHost.test.ts takes main's snapshot resume with this
branch's stdio MCP server in its config; conversationController.test.ts
keeps DictationSetup from core/voice/dictation (M61) with main's
GoalCommandVerb. The ACP agent's test fake gains AgentSession's new
controlGoal. The agent passes goal events over (no ACP counterpart yet).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Forks run 36222382949 passed in all four: Cursor 3.22.7 (VS Code
1.128.0), Devin Desktop 3.10.35 (1.126.0), Kiro 1.1.70 and Positron
2026.09.1 (1.130.0), each installing the VSIX and passing the 9
integration tests; Hosts run 36222382972 passed its 12 jobs. hosts.md
moves the four forks to Preview and fixes the stale Planned rows for Zed
and Neovim.

fork-release.mjs took Kiro's own 1.1.70 for its VS Code base: describe
now takes only a VS Code release number (1.NN.N) and otherwise lists
product.json's version fields.

PRIVACY.md gains the ACP agent: what it sends, the editor's MCP servers,
the key in the OS credential store, where Model API sessions live, the
redacted stderr log. PLAN.md records the playwright-core pin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
npm installs muse-spark-code-acp on Windows as a .cmd launcher, which
some editors cannot start (Node refuses to spawn a .cmd without a shell).
The guide now gives `node "<npm root -g>\muse-spark-code-acp\dist\acp.js"`
as the command that works everywhere; it is the form the Hosts workflow's
packaged suite already runs on Linux, macOS and Windows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Web search and image generation on the Model API backend, behind
--web-search and --image-generation (refused with the Muse Code
backend). At the first prompt the agent asks for each in the editor: a
row "Turn on Web search?" with what it does and its price, and a
session/request_permission offering "Turn on" and "Keep off". Only
"Turn on" turns it on, for the life of the process; a refusal, a
cancelled question or a client that cannot answer leaves it off and it
is not asked again. Two prompts at once share one question, and a
cancel while it is asked ends the prompt without a turn.

The Model API backend reads isPaidFeatureOn from src/acp/paid.ts and
logs every billed use with a running total. Paid rows and paid
approvals name their price in the title. New strings in all 15
languages; docs/acp.md, PRIVACY.md, PLAN.md (D62's condition met) and
the changelog updated. Drills P1-P6 in docs/certification/m63.md.

File access through the client (fs/*) waits for M46-M56.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
The second Forks run (36223213890) listed Kiro's product.json version
fields: its base is vsCodeVersion, 1.131.0. describe now reads that key
too; hosts.md, PLAN.md and the M62 record carry the version. All four
forks passed again, 9 integration tests each.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
…patibility branch

Conflicts:
- conversationController.ts: M46 guarded "Edit automatically" with one
  panel holding the session and no replayed request. The rule stays in
  core/agent/approvalRules.ts, shared with the ACP agent, which now also
  refuses a replayed request; the one-panel condition stays in the
  controller.
- PLAN.md: M46's D39 and this branch's D60-D62, both kept.

After the merge:
- ModelApiHost's new Promise.withResolvers (M46) is a new Promise:
  VS Code 1.99 and 1.100 run Node 20 (M62). The host's ES2023 library
  refused it at typecheck, as M62's check intends.
- The ACP agent's test fake gains M46's five AgentSession methods.
- docs/ide-compatibility/host-api.md regenerated: 23 commands and 7
  keybindings; no new VS Code API.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Conflict: README.md's Troubleshooting, where M47's two workflow entries
and this branch's Theia entry were added at the same place; all three
kept. Nothing else needed: no new AgentSession member, no VS Code API
newer than the 1.99 floor, no Node 20 gap; host-api.md unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9UjDkubgiZqw9y8c3hBDg
Joins origin/main 890e371 (PRs #35-#44) onto PR #32's head 68d450e.

Conflicts, each kept on both sides:
- agentBackend.ts: the PR's stdio MCP server union with main's side-chat
  options on startSession, resumeSession and forkSession.
- webviewSetup.ts / chatSurface.ts: ChatSurface stays in the PR's
  vscode-free chatSurface.ts and gains main's isSideChat.
- conversationController.test.ts: DictationSetup from core (the PR's move)
  plus main's new constants.
- PLAN.md: D41-D47 then D60-D62; main's gate table and notes with the
  PR's Cycles entry and its two rows. README, AGENTS, CHANGELOG and the
  certification index: both sides in milestone order, the PR's CHANGELOG
  entries under [Unreleased].

Integration fixes:
- The ACP runtime takes M48-M56's new manager inputs: Muse Code's default
  sandbox network, the in-memory prompt cache, M49's memory store, no paid
  subagents, and the shell job's C# (now shipped in the agent's package).
- MUSE_LOGIN_ARGS is back for the agent's `login` (M55 removed it because
  the panel signs in by device code).
- The fake session's steer returns a TurnSubmission, as main's
  interface does.
- The 1.99 floor: M55's two Promise.withResolvers became settleOnAbort
  (Node 20). The host typechecks at @types/vscode 1.99.0 and
  @types/node 20.19. The host API record was regenerated.
- M56's premise at older VS Code: fetch is routed from the floor on, but
  WebSocket only from 1.112.0. Diagnostics now says per global whether the
  editor routes it. The README, the D43 amendment and the certification
  record in m62.md cover it.
- acpRuntime.test.ts no longer finds a Muse Code installed on the
  machine running it.

Gates (Windows 11, Node 24.20.0): format:check, lint, typecheck,
check:l10n, check:host-api, deadcode, cycles, duplication and test:unit
(2,591 passed, 5 skipped) exit 0. The ACP stdio suite passes in-repo and
against the packed agent. build exits 1 at the size check only:
extension.js 600.5/600 KiB, acp.js 874.1/800 KiB. No budget was raised;
the choice is recorded in PLAN.md section 7 for the owner.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Open VSX refuses a publish into a namespace that does not exist, and RandyNorthrup had none (GET /api/RandyNorthrup answered 404 on 2026-09-27). The Open VSX job now creates it with OVSX_PAT when it is missing and skips when it exists; any other answer stops the job.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Main gained the 0.9.0 release fixes, 0.9.1, M57 (the Model API backend
as its own bundle, dist/modelApi.js) and M58 (a popup before every paid
use). Seven files conflicted, all prose or lists (AGENTS, CHANGELOG,
PLAN, README, the certification index, package.json's cycles entries);
both sides are kept. PR #32's changelog entries move out of [0.9.0],
where git placed them, into [Unreleased].

Adapted where the two sides met:

- The ACP agent loads the Model API backend from dist/modelApi.js beside
  acp.js, the file the .vsix ships, and its package now ships it: one
  build of the backend for both packages (PLAN.md D6 amendment).
  check-bundle-split fails when acp.js carries the backend; the host API
  gate lists modelApiEntry.ts as portable; the agent's notices take
  modelApi.js's packages. acp.js: 874.1 -> 713.2 KiB.
- M58 in the agent (D62 amendment): a paid feature is on only with its
  flag, and each use asks over session/request_permission in the
  conversation it is for, with the panel's words (paidUseQuestion, moved
  into paidConsent.ts) and allow_once / allow_always / reject_once.
  "Always" only with --trust-workspace, kept per folder in the agent's
  data folder (acp/paid-uses.json) and forgotten when the agent starts
  without the flag. No session or client to ask, a cancel, or an unknown
  option denies; subagents are always denied. The first prompt's
  "Turn on" question and its three strings are gone.
- ModelApiHostDeps.allowsPaidUse names the conversation (sessionId, a
  child's being its parent's), since one host serves every conversation
  of a folder.
- build.mjs: modelApi.js targets HOST_NODE_TARGET (node20.18); the test
  helper builds for the same target. fakeModelApi no longer uses
  Promise.withResolvers, which M57's integration test would call on
  VS Code 1.99's Node 20.18. src/acp uses isPromptSettledError (M57's
  lint rule). The approval title drops M58's removed paidFeature.
- The host API record regenerated (200 APIs, 17 Node built-ins).

Tests, drills P1-P9 and B1-B3, the packed agent's stdio suite and the
integration run (10 passing on 1.139.1 and on 1.99.0) are in
docs/certification/pr32-integration.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
dist/acp.js is 713.2 KiB now that the agent loads the Model API backend
from dist/modelApi.js (874.1 KiB, over its 800 KiB budget, on the tree
joined before M57). The budget is the measured size plus about 15 %,
rounded up to a multiple of 50 KiB. Zod is 445.2 KiB of it, 257.6 of
that the locales of the classic API the ACP SDK imports. The agent is
installed once and never loaded by VS Code, so its size is a download
(175.1 KiB gzipped), not a start-up cost. No other budget changes; the
extension stays at 432.5 of 600 KiB.

PLAN.md's open budget question (section 7) is recorded as resolved.
Drill: a 700 KiB budget fails the size check.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The agent runs outside VS Code, so VS Code's http.* settings and
museSpark.environmentVariables do not reach it. Measured against a local
proxy (nothing left the machine) on Node 22.0.0 to 24.20.0: the agent's
own fetch (the Model API backend) ignores HTTPS_PROXY unless
NODE_USE_ENV_PROXY=1 is also set, which Node honours from 22.21 and 24.0
(NODE_OPTIONS=--use-env-proxy from 22.21 and 24.5), reading HTTPS_PROXY,
falling back to HTTP_PROXY, and NO_PROXY, not ALL_PROXY. NODE_EXTRA_CA_CERTS
works on every version; --use-system-ca is accepted from 22.15. Muse Code,
started by the agent, reads the proxy variables itself (M56) with loopback
kept off the proxy.

The agent does not turn Node's switch on by itself, and its network-failure
advice still names VS Code's settings: recorded as PLAN.md Q66 for the
owner rather than changed here.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The key's round trip through the OS credential store (test/hosts/
keystore.sh, a made-up key) had run on GitHub's runners and by hand only
through gnome-keyring. With the agent packed from the merge and a
portable Node 22.23.3 it passed on the Windows 11 VM (Credential Manager,
in the owner's console session, through a one-off scheduled task) and on
the Mac mini (a keychain of the run's own, the owner's keychain settings
put back). Nothing is left on either rig.

Found: over SSH with a key, Windows refuses Credential Manager
(ERROR_NO_SUCH_LOGON_SESSION); the agent says the store cannot be used and
stores nothing. docs/acp.md now says to store the key from a desktop
session.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The agent now keeps the paid features allowed always in a folder in
acp/paid-uses.json beside its sessions (folder hashes and feature names
only), asks before each paid use rather than once for its price, and
reaches api.meta.ai through Node's fetch, which uses a proxy only when
its environment asks for one. Say so where the privacy notice describes
the agent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The first local semgrep run over PR #32's code (its container could not
fetch semgrep's rules) flags detect-child-process on the spawn behind
muse-spark-code-acp login. The command is the CLI resolveLaunch found
(install layout, PATH or an absolute --muse-binary), the arguments its
launcher's fixed prefix and MUSE_LOGIN_ARGS, as an array with no shell:
the same launch as muse serve. Suppressed with its reason inline and a
row in the escape-hatch register, as the other spawn sites are.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
npm run quality on 87383c7: exit 0 (2,652 unit tests, every bundle under
budget, a11y 336 pages clean, SAST 0 findings). The first full run failed
at SAST on the agent's login spawn, fixed in 87383c7; that run is the
suppression's drill. The integration run on the merge: 10 passing on
VS Code 1.139.1 and on 1.99.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The owner's ruling: the agent does not re-route through a proxy or add
undici; it says so, and its advice names what works outside VS Code.

- At start, on the Model API backend, one log warning when HTTPS_PROXY or
  HTTP_PROXY (either case) is set and Node's switch is off (only
  NODE_USE_ENV_PROXY=1, or --use-env-proxy in execArgv or NODE_OPTIONS,
  turns it on), or when this Node has no switch at all (below 22.21, 23):
  the requests go to Meta directly. Names only, never an address; the two
  spellings are one variable on Windows. Nothing on the Muse Code backend.
- A request that never reaches Meta: M56's classifier takes a
  NetworkAdvice, and the runtime's manager passes 'agent' down to the
  bundle's client, so certificate, proxy-credential and unreachable
  failures name NODE_EXTRA_CA_CERTS, --use-system-ca, HTTPS_PROXY and
  NODE_USE_ENV_PROXY instead of VS Code's http.* settings. Three new
  strings in en.ts and the 14 tables; the extension's advice is unchanged.
- The runtime takes sleep from main.ts, so the retries run instantly in
  tests; the fake Model API can fail a fetch with a socket code.

PLAN.md Q66 resolved and D62 amended; docs/acp.md and CHANGELOG updated;
tests and drills Q1-Q10 in docs/certification/pr32-integration.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
npm run quality on 83a4530: exit 0 (2,674 unit tests, 252 source files
through the l10n gate, acp.js 715.7 of 850 KiB, SAST 0 findings).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
RandyNorthrup and others added 6 commits October 1, 2026 10:04
…store

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex round 4 (PR #55): an archive hides what was made at or before its
wall-clock time and stays for a day or more. With the clock set back
after an archive, an unarchived conversation's new captures read as
taken before it, and record() dropped them. Unarchiving now removes the
conversation's archive files (plain files, no git, so Restricted Mode
too), through the port's new unforgetSession.

Red drill: unforgetSession made a no-op fails "keeps the new checkpoints
of a conversation unarchived after the clock went back" (Kubuntu);
restored byte-exact (b2e3e59b).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex round 4 (PR #55): when a tool's write failed after its copy was
taken (refused at its last check), the turn's end still made a change
entry from the copy, so a restore replaced the untouched file (new inode
and times, a broken hard link) and Redo recorded a change of nothing.
A copied file that ends the turn at its copy's size and bytes (the same
blob) is now no change. Size first; only same-size files are hashed.

Red drill: endsAsCopied returning an empty set fails "leaves alone an
ignored file a tool copied but never changed, and restores one rewritten
at its size" (Kubuntu); restored byte-exact (3019c186).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex security review, round 4 (PR #55): the separation check compared
the storage and workspace paths as written. When VS Code's storage path
runs through a link or junction into a trusted workspace, the check
passed, and a malicious repository could plant a hook in the shadow
repository that git would run on the next capture. The check now also
compares the paths as resolved (canonicalPath) and refuses before any
file is made. Tool writes were already confined by resolved paths.

Red drill: the resolved comparison disabled fails "refuses storage that
a link on its way puts inside the workspace" on Kubuntu (symlink) and
the Windows VM (junction); restored byte-exact (9036a5a7).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex round 4, P1 (PR #55): a save the user made in the editor while a
turn ran ended up in the turn's end capture, and changedOutside() looked
only between turns, so a restore attributed the user's bytes to the
model and overwrote them once the editor was clean.

VS Code's save events are the user's own (the extension's tools write
files directly and save no document). While a turn runs in the window,
the store notes each saved file against the open turns, and against a
turn whose record comes next (saves since its start capture); the turn's
end record keeps them as an optional `userSaves` list, and a restore
refuses those files as changed outside the turns. Records without the
field (0.10.0 candidates) restore as before. Limit: writers VS Code does
not see (another editor, a terminal) are not told apart.

Red drills (Kubuntu): ignoring userSaves in changedOutside fails both new
tests; not counting saves made before the turn's record fails the second;
restored byte-exact (b8e9e235).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

@codex review

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 357216b253

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 357216b253

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/host/checkpoints/checkpointStore.ts Outdated
Comment thread src/host/checkpoints/checkpointStore.ts
Comment thread src/host/backend/mcpServers.ts
Comment thread src/host/checkpoints/checkpointRetention.ts
Comment thread src/host/conversation/conversationController.ts
Comment thread src/extension.ts Outdated
RandyNorthrup and others added 7 commits October 1, 2026 12:20
The Windows run of b244132 passed every gate, unit tests included, and
was cancelled at the 30-minute job limit during the accessibility gate
(the run before took 29 min 11 s). This is the job's budget, not a test
timeout; every test deadline is unchanged. Kubuntu, the Mac mini and the
Windows VM passed the full gate of b244132.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Found by the M83 lane: `\b[a-z][\w+.-]*://` restarts at every word start
of a long run of dotted or dashed words with no `://`, so redacting a
40,000-character line took 3.5 s, and every log line passes through
redactSecrets. The scheme is now at most 32 characters, which keeps the
scan linear; a 32-character scheme still redacts.

Red drill: the unbounded pattern back fails "reads a long line of dotted
words with no URL in linear time" (28.6 s against a 1 s bound); restored
byte-exact (f27df521).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ts rewind

P1: a file the extension writes or deletes in the user's name (Create
AGENTS.md, an export, Revert's write and delete, a plan's publication or
stale-stage removal, the Memory view's notes, index lines and trash) goes
through workspace.fs or node fs and fires no save event, so a turn running
meanwhile took it for its own and its restore undid it (deleting the user's
AGENTS.md). Each is now noted with noteUserSave once it is done: asUserEdit
for the lease callers, PlanIoOptions.noteUserWrite for plans, and a view
memory store whose io notes its writes plus afterDelete for the trash. The
model's memory tools keep the unnoted store and stay restorable; the
environment probe's git lease writes nothing and is not noted.

P2: restoreFiles refuses a rewind whose session or turn differs from the
files' before any confirmation, restore or rewind.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nseen ends

Codex round on PR #55, two P1 findings in the restore's outside-change scan:

- An earlier turn of the restored conversation that was still running when
  the restored turn started (the conversation open in two windows) was cut
  off by the "selected turn onward" slice and skipped by the other-
  conversation scan, so its edits after that start were reverted as the
  restored turn's. The undo range stays the selected turn onward; the
  conversation's earlier turns now reach changedOutside(), and one whose end
  number is at or past the selected start number (equal: numbered at once
  in two windows), or whose end no window saw, is blamed like another
  conversation's overlapping turn: its start-to-end paths (to the current
  capture when its end was not seen) are refused. Earlier turns that ended
  before the selected start add nothing.
- Another conversation's turn whose window went before its end was seen
  collapsed to a zero-length interval at its start, so it missed a later
  restored turn it overlapped. An unseen end is now bounded by the start of
  its conversation's next turn (its paths read up to that turn's start
  tree), else by now (up to the current capture).

Cost: an unfinished turn whose window went keeps blocking the paths changed
since it began, until retention drops it: for another conversation only
until that conversation continues; for the restored conversation's own
earlier turn, for every later turn's restore.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A window reloaded mid-turn leaves that turn's end unseen. Counting such a
turn as overlapping every later turn of its conversation refused every
path changed since it began on each later restore, until retention
dropped the record. Its window is gone and the conversation went on in
the next turn, so its end came no later than that turn's start, the
same bound the other conversations' unseen ends take. A turn still
running elsewhere still overlaps. Limit: the conversation open in two
windows, one gone mid-turn while the other went on.

Red drill: the unbounded rule back fails "bounds an earlier turn whose
window went before its end was seen at the next turn (a reload)"
(Kubuntu); restored byte-exact (512f3588).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0ca990740c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/host/checkpoints/checkpointStore.ts
Comment thread src/host/checkpoints/checkpointStore.ts
Comment thread src/host/checkpoints/checkpointStore.ts
Comment thread src/host/checkpoints/checkpointStore.ts
Comment thread src/host/checkpoints/checkpointStore.ts
@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 0ca990740c

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

RandyNorthrup and others added 4 commits October 1, 2026 15:02
…ed files a turn with no end may have changed

Codex round 6 (PR #55), two P1s:
- Two windows can give turns of one conversation the same `sequence`;
  their order was then the refs' hash order, and a tied earlier turn
  could land in the undo range and be reverted. A turn tied with the
  restored one is never undone with it; it is blamed, as an earlier
  turn still running is.
- A blamed turn whose end no window recorded has no list of the ignored
  files it changed, which git's trees cannot show, so a restore could
  write this turn's ignored-file copy over its later bytes. With such a
  turn blamed, every ignored file the restore would put back is refused.

Red drills (Kubuntu): tie handling off fails "never undoes a turn
numbered at once with the restored one, whichever the refs list first";
the ignored refusal off fails "refuses an ignored file it would put back
when a blamed turn left no list of its ignored changes"; restored
byte-exact (e48a5ead before the lint wrap).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A save that landed while a turn's start capture ran (after the index had
hashed the file, before the capture finished) was dropped from the turn's
user saves, because the capture was timed at its end; a restore could then
overwrite it. A tool's copy of an ignored file taken after the wall clock
was set back fell before its turn's start and was discarded, so a restore
reported no earlier copy.

The store now orders its own in-window events by a monotonic clock
(performance.now() unless a test injects one): a capture is stamped before
it reads any file, and journal entries, running marks, recent saves, open
turns and pending ends use that clock. The wall-clock createdAt/endedAt that
are persisted and compared across windows are unchanged; the monotonic
stamps are never recorded.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex round 6, P1 (PR #55): a save in an idle window on the folder
during another window's turn was recorded only against the idle
window's own (absent) turns, so the turn's end took the saved bytes for
its own and a restore could overwrite them.

Every window keeps a saves file beside its presence ({pid, saves,
droppedThrough}), written whole, no lock and no git, bounded by
CHECKPOINT_PEER_SAVES_MAX and CHECKPOINT_PEER_SAVE_KEEP_MS; a turn's end
reads every other window's (closed ones too) after its end capture and
adds the files saved since its start. Saves it cannot know (an unread
file written since, a save the cap dropped, a turn past the keep time)
fail closed: every file the turn changed counts as saved.

Integrated by the lead onto the monotonic in-window clock: the other
windows' saves count from the wall-clock time the start capture began
(Snapshot/OpenTurn `startedWallAt`), so a peer's save during the start
capture counts too; the capture-save test runs for this window and for
an idle one. Drill: counting from the capture's completion
(record.createdAt) fails the idle-window case; restored byte-exact
(e75dd983). The lane's twelve drills are in its report.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…w test setup

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 1fb31bc685

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1fb31bc685

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/host/checkpoints/checkpointHost.ts
Comment thread src/host/checkpoints/ignoredScan.ts
Comment thread src/host/checkpoints/checkpointStore.ts
Comment thread src/host/checkpoints/checkpointStore.ts
Comment thread src/host/conversation/conversationController.ts
Comment thread src/host/checkpoints/checkpointStore.ts
PR #55 reached a seventh Codex round of findings in one family: the
restore infers which changes were the turn's from whole-workspace
captures, and multiple windows, subagents, saves and clocks keep giving
new cases. Owner's decision: 0.10.0 ships with museSpark.turnCheckpoints
off by default and marked Preview (all 15 setting descriptions say so),
and M86 rebuilds the restore on the model tools' own writes. PLAN D63
and M86, README, PRIVACY, CHANGELOG and the certification record it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@RandyNorthrup
RandyNorthrup merged commit bdfb651 into main Oct 2, 2026
19 checks passed
@RandyNorthrup
RandyNorthrup deleted the feature/m72-checkpoints branch October 2, 2026 00:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants