Work on the D49 program (PLAN.md M67–M85) paused on 2026-09-28 when the Claude, Codex and Grok Build usage limits were all reached. This issue is the resume point for a cloud agent or the next session. Follow AGENTS.md and CLAUDE.md as always: every change is reviewed in one pass (three classes: concurrency/lifecycle; wire/validation/security; failure paths/honesty/docs), red-drilled, and gated with the complete npm run quality before it is pushed. A third review round on one module means a redesign, not another patch.
Open pull requests (pushed, gate passed; each waits only for its final Codex review)
| PR |
Milestone |
Head |
State |
| #32 |
Other editors (ACP agent, M60–M63) |
46ba5406 |
Muse Code's review of the last two rounds found a P1 (cancel while a turn is starting) and a P2; both fixed, drilled, full Win11 gate exit 0. Merge after review; then the M63c cloud agent resumes; then release 0.10.0 (never 1.0). |
| #52 |
M69 web fetch |
4c4d14c7 |
Redesign: page text as served, all untrusted. Muse Code's review: one P2 open — picture in OMITTED (src/core/web/htmlToMarkdown.ts) drops the fallback <img> and its alt text; fix by removing 'picture' from the list. |
| #53 |
M79 plans as files |
b6f24e79 |
Codex threads answered and resolved. |
| #54 |
M68 verify loop |
51c80e83 |
Newer work is on wip/m68-verify-loop (below); thread PRRT_kwDOUkzj5M6m0_2F still to answer. |
| #55 |
M72 checkpoints |
9be57e55 |
Superseded by the redesign on wip/m72-checkpoints (below). |
| #57 |
M67 code intelligence |
107510a4 |
Waiting for review. |
| #51 |
External (Android tests) |
— |
The owner's call. |
Work in progress (branches wip/*; not reviewed or gated as a whole)
Each branch's last commit message states where it stopped. Resume each by reading its docs/certification/m<nn>.md and PLAN section, running npm run quality, fixing, reviewing in one pass, then opening the PR from the branch (rename it feature/... if you like).
| Branch |
Milestone |
Where it stopped |
wip/m68-verify-loop |
M68 |
Round-4 fix plus reviewer fixes: re-check changesWhatRunsNow after the guard; check-card grants keyed apart from the shell tool. Then push to PR #54 and answer its thread. |
wip/m70-review |
M70 /review |
Built, Kubuntu gate failed only on jscpd (fixed in cefa9587); three-class review findings not yet collected. Cert Gate section says "to be recorded". |
wip/m71-git-prs |
M71 git and PRs |
Built and gated (3af80e40); review fixes in progress (ConversationGit rewrite). |
wip/m72-checkpoints |
M72 checkpoints |
Mid-redesign: records as git refs updated by compare-and-swap, per-window files, prune grace, no lock. Findings the design must remove: below. |
wip/m73-observation-packing |
M73 |
Muse Code's draft, not checked. Must stay an M75 evaluation arm with no user setting until a live M75 run shows the floors held. |
wip/m74-handoff |
M74 /handoff |
Muse Code's draft, not checked. |
wip/m75-eval |
M75 evaluation |
Rebuilt harness; reviewer fixes to the runner in progress; full 10-task live baseline not run yet. |
wip/m76-agents |
M76 custom agents |
Muse Code's draft plus its own review fix; not checked by a finisher. |
wip/m77-board |
M77 session board, best-of-N |
Muse Code's draft with translations; one jscpd clone; no certification record. |
wip/m78-auto-safe |
M78 Auto made safe |
Builder stopped mid-build. |
wip/m81-browser |
M81 browser check |
Muse Code's draft; its security claims (loopback-only, pipe transport, temp profile) not yet verified. |
wip/m82-awareness |
M82 budget and notices |
Fixing the review findings below. |
wip/m83-import |
M83 import from other agents |
Fixing the review findings below. |
wip/m84-export |
M84 export and share |
Fixing an independent review's findings; transferInvalidField exists in en.ts only. |
Not started: M80 (headless and CI; needs PR #32 merged), M85 (needs the owner's TypeSafe key).
Lesson from this round: Muse Code drafts had the right shape and good translations, but every one had real gaps in security or correctness logic, and several claimed checks that had not run. Treat each Muse draft as unverified and check every claim against the code.
Review findings still to fix (Grok Build, before the limits)
M82 (wip/m82-awareness)
I'll review the diff between that base commit and HEAD against the working rules, starting with AGENTS.md and the change stat.The change is M82 (notices, reply cost, session cap). I'll read the plan sections it cites and the source diffs next.
The last diff read was cut off. I'll finish the reservation, notice, and docs paths, then write the full report.
The spend ledger and the model-switch path still need a line check before the report.## Class 1 — concurrency, ordering, stale state, lifecycle
P1. src/core/backends/modelapi/ModelApiHost.ts:2450 and :6062, shutdown at :5669 and :6446. A capped request that reports usage adds that cost to budgetSpentUsd only after adoptOutput has already pushed the function_call with no output. persist then returns without writing while hasUnansweredSessionCall is true, so for the whole approval or tool wait the file still has the previous spend. Reload or a crash in that window restores budgetSpentUsd ?? 0, and the next reservation is computed against that remainder, so a request that does not fit the real balance is sent. dispose also fires cancel() and does not wait for it; close() then awaits the saving promise captured before closeReservation runs, so a started request aborted by window close is charged only in memory. Record the settled spend, and an in-flight reservation once a stream event has arrived, in a write that is not skipped with the replay, and await that write in close() before it returns.
P1. src/core/backends/modelapi/ModelApiHost.ts:1899 and :1926, with setModel at :5232. The controller applies a model change as soon as it is confirmed (conversationController.ts:2982) and does not wait for the open request. The body already sent still names the old model, but noteUsage prices the reported tokens with this.modelId at settlement time. Switching muse-spark-1.3 to an already-confirmed muse-spark-1.3-contributor while the response is in flight records the bill at the contributor rates ($0.10 / $0.20 per million instead of $1.25 / $4.25), so later requests are reserved against a ledger below what Meta charged and are sent past the cap. The same settlement writes budgetBase back from that open reservation, undoing the clear in setModel, and the next estimate starts from the previous model's reported input count. Stamp the model id onto the reservation when the body is built, settle the dollars with that id, and install the base only when this.modelId is still that id.
P2. src/host/conversation/turnNotifications.ts:64 and :110, src/core/backends/musecode/MuseCodeHost.ts:511 and :1096. A question replayed to a later surface is delivered without isReplayed (onEvent flags only approvalRequested; approval/listPending does the same for userInput/requested). attentionNotice therefore always returns a question notice. The window's key set suppresses the second toast only until BACKGROUND_NOTICE_KEYS_MAX (200) drops the oldest key; the notifier test shows a forgotten key is raised again. After 200 later notices, opening a second surface on a session that still has that question shows the waiting toast again. Mark replayed questions isReplayed and ignore them in attentionNotice, as replayed approvals already are.
Class 2 — wire evidence, validation, security
P2. src/core/backends/modelapi/schemas.ts:165 and ModelApiHost.ts:1900; stored ledger at src/core/backends/modelapi/sessionStore.ts:222. Response usage is z.number() with no finite, non-negative check, and budgetSpentUsd is restored the same way. estimateCostUsd is linear, so a schema-valid output_tokens or input_tokens below zero, or a stored budgetSpentUsd below zero, decreases the ledger. The next reserveRequest then sees more room than the cap has and sends a request that does not fit. Reject non-finite and negative usage before it is added, and reject a negative stored spend.
Class 3 — failure paths, honesty, docs
P2. src/core/backends/modelapi/ModelApiHost.ts:4873. A failed countInputTokens is logged with describe(error), and ModelApiError's message is the envelope's error.message (client.ts:111). That string is free text from the network and is written to the log as sent. Log a status and a fixed protocol word.
P2. package.nls.json:97, README.md session-budget row, and docs/certification/m82.md:98 and :174. Those say spend is kept with the session and that the cap can be passed only by the input estimate or a paid subagent. The file omits settled spend for the whole time a call is unanswered, and a model change during a request both under-prices that request and puts the old base back. H10 (test/unit/modelApiHost.test.ts:2116) reloads only after a finished turn whose save succeeded, and N4 (test/unit/turnNotifications.test.ts:96) only passes a replayed approval, so both drills stay green on those paths. Persist and settle as in the two P1s, mark replayed questions, and assert a reload during an open tool call and a second surface's replayed question.
P1: 2. P2: 4.
M83 (wip/m83-import)
I'll review the diff between that base commit and HEAD against the working rules and the plan sections this change touches. Starting with AGENTS.md and the diff stat.The change is M83, the agent-import feature. I'll read the plan sections it cites and the new source before judging defects.The plan section is M83. Next I’ll read the import implementation and the host paths that write it.The import path is large enough that the confinement, masking, and apply logic need to be checked against the helpers they call.
A few confinement and apply paths still need exact line checks before the report.## Class 1 — concurrency, ordering, stale state, lifecycle
P2. src/core/import/agentImport.ts:1141 (appendSections), reached from a second museSpark.importFromAgents while the first is still between that read and appendText at line 1154. Confirm is a non-modal notification (src/host/agentImportHost.ts:73) and registerLoggedCommand (src/extension.ts:437) does not serialize the command. Two accepted imports of the same rules section both read AGENTS.md before either appends, both treat the heading as absent, and both append. The file gains the section twice. The re-read only observes an append that has already finished.
Fix: keep a single in-flight import and append under an exclusive lock that covers the read and the append.
Class 2 — wire evidence, validation, security
P1. src/host/importIo.ts:111 (appendFile), after isWithinRoot at src/core/import/agentImport.ts:1127. canonicalPath (src/host/canonicalPath.ts:35) treats realpath ENOENT as a missing file and returns the workspace path. A dangling symlink is ENOENT, so AGENTS.md → /home/user/.bashrc (target absent, parent present) is judged inside the workspace. readText also gets ENOENT and the append proceeds with an empty current file. appendFile follows the symlink and creates the outside file. The body is the repository's rules text. The same destination is read with no confine before confirm: readPlanText (src/host/commands/agentImportCommands.ts:341) uses stat/readFile, which follow a symlink whose target exists, so a link to an existing key file is loaded (up to 16 MiB) while building the preview; apply then refuses that write. memoryImportIo.realPath rewrites a link even when the target is missing, so the apply test does not exercise this.
Fix: lstat AGENTS.md before the plan read and the append, refuse any symlink, and append with O_NOFOLLOW.
P1. src/host/agentImportHost.ts:93 (openTarget), for the hooks copy built at src/core/import/agentImport.ts:1046. Project hooks are opened with no confineWorkspacePath check. isPathPresent is lstat, so a symlink counts as the file being present, and the prompt shows the lexical path .muse/hooks.json (importDisplayPath). A trusted repository whose .muse is a symlink to ~/.config/muse, or whose .muse/hooks.json is a symlink to ~/.bashrc, opens that outside file when it exists. Paste-and-save, which the prompt asks for, overwrites it.
Fix: confine the hooks path to the workspace root and do not open it when the canonical path leaves the workspace or any component is a symlink.
P2. src/core/import/importConvert.ts:411. command goes through maskText; statusMessage is stored raw and then copied into the preview and the clipboard (src/host/commands/agentImportCommands.ts:252 fences copy.text with no further mask). A handler { "type": "command", "command": "echo ok", "statusMessage": "<GitHub-shaped token or password=…>" } shows and copies that value.
Fix: run statusMessage through maskText before it is stored on the hook.
Class 3 — failure paths, honesty, docs
P2. src/host/commands/agentImportCommands.ts:276. readPlanText returns undefined for tooLarge and for a failed decode, and logs nothing in those cases. planImportApply (src/core/import/agentImport.ts:1074) only treats status 'read' as a real settings file, so a settings file over AGENT_IMPORT_FILE_MAX_BYTES (64 KiB) — still under the 1 MiB hook-config limit — or a file that is not valid text, yields no known server names and hasLegacyMcpKey: false while hasMuseSettingsFile stays true. The preview says to merge, omits the legacy-key warning, and offers servers the file already has. Merging mcpServers into a file that already has mcp_servers is the case Muse Code loads neither key.
Fix: when the settings file is present but not read, skip server and hook copies and say the file could not be read.
P2. src/core/import/agentImport.ts:1154. A rules append is not checked against RULES_FILE_MAX_BYTES (64 KiB). Source files may be that large, and the section adds a heading, so one full-size rules file makes a new AGENTS.md exceed the limit. An existing AGENTS.md of 40 KiB plus a 30 KiB section does the same. loadRuleFile (src/core/context/rules.ts:82) then skips the whole file, including rules that were already loading. The import still reports the section as written. A skill over the same limit is refused (src/core/import/agentImport.ts:527).
Fix: refuse the append when the resulting AGENTS.md would exceed RULES_FILE_MAX_BYTES, and say so in the preview.
P2. src/core/import/agentImport.ts:666 and :690. Both Claude agent walks pass depth: 1, so collectMarkdown never enters a subdirectory. agents/review/security.md is never imported. Command walks use AGENT_IMPORT_COMMAND_MAX_DEPTH (3). PLAN.md M83 "As built" (around line 7128) says agents and commands are both read from namespaced subfolders three deep.
Fix: walk Claude agents/ to the same depth as commands/, and slug the directory prefixes the way commands already do.
P1: 2. P2: 5.
M72 lock, third round (wip/m72-checkpoints; the redesign should remove these by construction)
I'll review the diff between that base commit and HEAD against the working rules, starting with AGENTS.md and the change summary.The change is M72: checkpoint-store locking across windows. I'll read the plan sections it cites and the new lock code next.The lock lives in storeLock.ts. Next I’ll trace every caller for races around acquire, release, and presence.
The lock and turn-end paths are the risk. I'll confirm the races and the docs against the code, then write the report.## Class 1 — concurrency, ordering, stale state, lifecycle
P1. src/host/conversation/conversationCheckpoints.ts:239, src/host/checkpoints/checkpointStore.ts:1838, src/host/checkpoints/checkpointStore.ts:1847. Ending a turn unpublishes it and records its end together. markTurn(false) drops the turn from runningTurns and, once the presence write finishes, trimJournal deletes every staged copy not still covered by an older running turn. endTurn has already removed the turn from openTurns before copiesSince, and held is filled only if the lock comes back busy. On a normal end the recording is already done, so endTurn starts before that trim: copiesSince keeps the staged path, trimJournal unlinks the file, then finishTurn runs hash-object on it (checkpointStore.ts:1085) and throws. That error is not StoreBusyError, so it is not retried, and endAfterRecording only logs it (conversationCheckpoints.ts:281). The record never gets endedAt or the tool pre-image, so an ignored file the tools copied (.env) cannot be restored, and every change until the next capture is treated as unsure. If the turn ends before record has set openTurns, the trim finishes first and the end is saved with the pre-image already gone. Hold those staged paths for the in-flight end, and trim only after finishTurn has imported them.
P1. src/host/checkpoints/checkpointStore.ts:2068, src/host/checkpoints/checkpointStore.ts:2100, src/host/checkpoints/checkpointStore.ts:2165. forgetSession and queueForget set the archive in memory and, when the other window holds the lock, defer the records write and return. dispose aborts the store and clears the retry timer; scheduleRetry (checkpointStore.ts:491) then does nothing because the signal is already aborted. ConversationCheckpoints.forget therefore resolves with no error (conversationCheckpoints.ts:610) while records.json never gains forgotten. The other window keeps that conversation's checkpoints, including copied file bytes. CHANGELOG.md:39 says an archive happens later on its own. Write the archive before forget resolves, and on close flush it or report that it was not written.
P2. src/host/conversation/conversationCheckpoints.ts:514. A turn this panel did not start (turnStarted: queued message, scheduled run) is marked with publish, which is not awaited and whose failure is only logged (conversationCheckpoints.ts:183). The backend is already editing. Until that presence write lands, or for up to one heartbeat if it fails, the other window does not see the turn and will restore over it. SECURITY.md:84 and README.md:824 say the turn is marked before it can change a file. Await the mark, and do not leave a failed publish until the next beat.
P2. src/host/checkpoints/storeLock.ts:219, src/host/checkpoints/storeLock.ts:417. Takeover deletes the moved lock when the owner id still matches, without calling isGone on the moved file. A heartbeat (utimes, or a presence rewrite) between the stale check and the rename leaves a live window's lock deleted and the taker free to create its own, while the first window still has isHeld and keeps writing. liveWindows has the same gap: it rms a presence file it judged gone, then dropOrphanRefs (checkpointStore.ts:649) drops that window's pins and staging/ even if the file was replaced with a fresh one after the read. After the move or before the delete, put the file back unless that same bytes are still gone.
P2. src/host/checkpoints/storeLock.ts:303, src/host/checkpoints/storeLock.ts:427. writePresence checks isPresent only before mkdir / writeFileAtomically. leave (from dispose) deletes the presence file in that gap, and the in-flight write creates it again with the old running list and a fresh mtime. beat will not remove it (isPresent is already false). While that process is still alive, the other window's turnBlocking (checkpointStore.ts:776) refuses every restore. Re-check isPresent after the write and delete the file if the window has left.
P2. src/host/checkpoints/checkpointStore.ts:2166, src/host/checkpoints/checkpointStore.ts:1526. dispose clears the heartbeat and aborts git, but runSteps never reads the signal, so a restore already writing files keeps going with isHeld true and a lock file that is no longer refreshed. After CHECKPOINT_OWNER_STALE_MS the other window takes the lock over. Those file writes, and the update-refs for the index, pins and keep refs (checkpointStore.ts:871, checkpointStore.ts:1381, checkpointStore.ts:1656, checkpointStore.ts:1744, checkpointStore.ts:1784), never call assertHeld, so both windows write. SECURITY.md:81 says a write is refused once the lock is taken over. Keep beating until the in-flight operation finishes, and check the lock before each ref update and each file write.
Class 2 — wire evidence, validation, security
none
Class 3 — failure paths, honesty, docs
P2. src/host/checkpoints/checkpointStore.ts:1784, src/host/checkpoints/checkpointStore.ts:1807. A record that throws for any reason other than StoreBusyError does not schedule deletion of the pin, and ConversationCheckpoints.record does not call release (conversationCheckpoints.ts:252). Pins of a live window are kept by dropOrphanRefs, so the capture stays for the life of the window. A busy failure after update-ref of the keep ref schedules only the pin delete, so the keep ref stays with no record until the next maintain. docs/certification/m72.md:401 says a failed record leaves nothing behind. On every failed record, delete the pin and any keep ref that was created.
P2. src/host/conversation/conversationCheckpoints.ts:471. If the running mark cannot be written, beforeTurn throws a hardcoded English sentence that embeds error.message. sendMessage posts that string as sendFailed (conversationController.ts:2893). An EACCES/ENOENT from the presence write includes the absolute path under the user profile, in the panel and in the log, and the sentence is not in en.ts or the 14 l10n/ui.*.json tables. README.md:826 says the panel says why. Show one localized reason and log a fixed phrase.
P2. docs/certification/m72.md:372, test/unit/storeLock.test.ts:139. Drill L9 claims "writes allowed after the lock was taken over" and points at a test that only calls StoreLock.assertHeld after overwriting the lock file. Removing assertHeld from save, deleteRefs and prune, or the update-refs that never call it, still leaves that test green. Point the drill at a store write that must fail once the lock file is no longer this window's.
P1: 2. P2: 7.
M68 round 4 (wip/m68-verify-loop)
I'll review the diff between that base commit and HEAD against the working rules, starting with AGENTS.md and the plan sections this change touches.The change is one commit on file judging for verify-loop commands. I'll read the plan section and the full diff next.The diff is the conservative command-word split and moving the rule check into authorizeCommand. I'll read that path and the M68 plan next.
The session-rule invalidation is sampled once, then a later await can change the ledger, and the same rule still allows the shell tool.
Class 1 — concurrency, ordering, stale state, lifecycle
P1 src/core/backends/modelapi/ModelApiHost.ts:3766
authorizeCommand decides at line 3673. When changesWhatRunsNow(ruleCommand) is false and a session rule exists, it returns allow with no further read of the ledger (line 3680). runVerifyCommand then awaits the guard (line 3762) and only then runs the command (line 3766). then_run always installs that guard; a changedFiles check does too when it has paths. A child turn is started with void child.session.sendTurn (line 3238), so it runs while the parent is in that await. The child's edit calls parentSession.ledger.noteOutsideEdit (line 4316).
Inputs: a session allow for npm run lint (or node scripts/t.js); the parent is in then_run of that command after editing src/a.ts only; during isAsEdited's read, the child writes package.json or scripts/t.js. The guard only compares the then_run file's hash, so it returns true. The command runs under the session rule. A re-check of changesWhatRunsNow would now be true and would ask.
Fix: after the guard returns, call changesWhatRunsNow(ruleCommand) again and, if it is true, take the mode shell verdict instead of running.
Class 2 — wire evidence, validation, security
P1 src/core/backends/modelapi/ModelApiHost.ts:4057
The check card stores the allow on the shell tool's own key: allowForSession(shell.name, ruleCommand) (line 2465), and authorizeCommand builds that same key (lines 3667–3671). After the edit, only authorizeCommand ignores the rule (line 3673). The shell tool does not. decideAndRun puts the model's command string on the query (lines 4183–4186) and verdictWithHook calls permissions.verdict (line 4057), which returns allow when that key is present.
Inputs: the user has already chosen "Always allow in this session" for node scripts/t.js. The model edits scripts/t.js, then calls the shell tool with that same command (same round or the next). The automatic check would ask; the shell call does not, and the modified script runs.
Fix: in verdictWithHook / decideAndRun, if the call is the shell tool and changesWhatRunsNow(command) is true, use verdictFor(mode, 'shell') instead of the session rule.
Class 3 — failure paths, honesty, docs
P2 PLAN.md:6677
That sentence still defines canChangeWhatRuns as command-defining files, code-loading files, or a file the command line names. canChangeWhatRuns('src/a.ts', 'node "scripts/my check.js"') is now true: any edited file counts when a word is not plain, and a workspace-relative path matches as a substring. README and CHANGELOG in this diff say that; PLAN does not.
Fix: update the M68 permission bullet so it matches canChangeWhatRuns in src/core/verify/codeFiles.ts.
P1: 2. P2: 1.
Written by Claude Code at the pause, 2026-09-28.
Work on the D49 program (PLAN.md M67–M85) paused on 2026-09-28 when the Claude, Codex and Grok Build usage limits were all reached. This issue is the resume point for a cloud agent or the next session. Follow
AGENTS.mdandCLAUDE.mdas always: every change is reviewed in one pass (three classes: concurrency/lifecycle; wire/validation/security; failure paths/honesty/docs), red-drilled, and gated with the completenpm run qualitybefore it is pushed. A third review round on one module means a redesign, not another patch.Open pull requests (pushed, gate passed; each waits only for its final Codex review)
46ba54064c4d14c7pictureinOMITTED(src/core/web/htmlToMarkdown.ts) drops the fallback<img>and its alt text; fix by removing'picture'from the list.b6f24e7951c80e83wip/m68-verify-loop(below); threadPRRT_kwDOUkzj5M6m0_2Fstill to answer.9be57e55wip/m72-checkpoints(below).107510a4Work in progress (branches
wip/*; not reviewed or gated as a whole)Each branch's last commit message states where it stopped. Resume each by reading its
docs/certification/m<nn>.mdand PLAN section, runningnpm run quality, fixing, reviewing in one pass, then opening the PR from the branch (rename itfeature/...if you like).wip/m68-verify-loopchangesWhatRunsNowafter the guard; check-card grants keyed apart from the shell tool. Then push to PR #54 and answer its thread.wip/m70-review/reviewcefa9587); three-class review findings not yet collected. Cert Gate section says "to be recorded".wip/m71-git-prs3af80e40); review fixes in progress (ConversationGit rewrite).wip/m72-checkpointswip/m73-observation-packingwip/m74-handoff/handoffwip/m75-evalwip/m76-agentswip/m77-boardwip/m78-auto-safewip/m81-browserwip/m82-awarenesswip/m83-importwip/m84-exporttransferInvalidFieldexists inen.tsonly.Not started: M80 (headless and CI; needs PR #32 merged), M85 (needs the owner's TypeSafe key).
Lesson from this round: Muse Code drafts had the right shape and good translations, but every one had real gaps in security or correctness logic, and several claimed checks that had not run. Treat each Muse draft as unverified and check every claim against the code.
Review findings still to fix (Grok Build, before the limits)
M82 (wip/m82-awareness)
I'll review the diff between that base commit and HEAD against the working rules, starting with
AGENTS.mdand the change stat.The change is M82 (notices, reply cost, session cap). I'll read the plan sections it cites and the source diffs next.The last diff read was cut off. I'll finish the reservation, notice, and docs paths, then write the full report.
The spend ledger and the model-switch path still need a line check before the report.## Class 1 — concurrency, ordering, stale state, lifecycle
P1.
src/core/backends/modelapi/ModelApiHost.ts:2450and:6062, shutdown at:5669and:6446. A capped request that reports usage adds that cost tobudgetSpentUsdonly afteradoptOutputhas already pushed thefunction_callwith no output.persistthen returns without writing whilehasUnansweredSessionCallis true, so for the whole approval or tool wait the file still has the previous spend. Reload or a crash in that window restoresbudgetSpentUsd ?? 0, and the next reservation is computed against that remainder, so a request that does not fit the real balance is sent.disposealso firescancel()and does not wait for it;close()then awaits the saving promise captured beforecloseReservationruns, so a started request aborted by window close is charged only in memory. Record the settled spend, and an in-flight reservation once a stream event has arrived, in a write that is not skipped with the replay, and await that write inclose()before it returns.P1.
src/core/backends/modelapi/ModelApiHost.ts:1899and:1926, withsetModelat:5232. The controller applies a model change as soon as it is confirmed (conversationController.ts:2982) and does not wait for the open request. The body already sent still names the old model, butnoteUsageprices the reported tokens withthis.modelIdat settlement time. Switchingmuse-spark-1.3to an already-confirmedmuse-spark-1.3-contributorwhile the response is in flight records the bill at the contributor rates ($0.10 / $0.20 per million instead of $1.25 / $4.25), so later requests are reserved against a ledger below what Meta charged and are sent past the cap. The same settlement writesbudgetBaseback from that open reservation, undoing the clear insetModel, and the next estimate starts from the previous model's reported input count. Stamp the model id onto the reservation when the body is built, settle the dollars with that id, and install the base only whenthis.modelIdis still that id.P2.
src/host/conversation/turnNotifications.ts:64and:110,src/core/backends/musecode/MuseCodeHost.ts:511and:1096. A question replayed to a later surface is delivered withoutisReplayed(onEventflags onlyapprovalRequested;approval/listPendingdoes the same foruserInput/requested).attentionNoticetherefore always returns a question notice. The window's key set suppresses the second toast only untilBACKGROUND_NOTICE_KEYS_MAX(200) drops the oldest key; the notifier test shows a forgotten key is raised again. After 200 later notices, opening a second surface on a session that still has that question shows the waiting toast again. Mark replayed questionsisReplayedand ignore them inattentionNotice, as replayed approvals already are.Class 2 — wire evidence, validation, security
P2.
src/core/backends/modelapi/schemas.ts:165andModelApiHost.ts:1900; stored ledger atsrc/core/backends/modelapi/sessionStore.ts:222. Responseusageisz.number()with no finite, non-negative check, andbudgetSpentUsdis restored the same way.estimateCostUsdis linear, so a schema-validoutput_tokensorinput_tokensbelow zero, or a storedbudgetSpentUsdbelow zero, decreases the ledger. The nextreserveRequestthen sees more room than the cap has and sends a request that does not fit. Reject non-finite and negative usage before it is added, and reject a negative stored spend.Class 3 — failure paths, honesty, docs
P2.
src/core/backends/modelapi/ModelApiHost.ts:4873. A failedcountInputTokensis logged withdescribe(error), andModelApiError's message is the envelope'serror.message(client.ts:111). That string is free text from the network and is written to the log as sent. Log a status and a fixed protocol word.P2.
package.nls.json:97,README.mdsession-budget row, anddocs/certification/m82.md:98and:174. Those say spend is kept with the session and that the cap can be passed only by the input estimate or a paid subagent. The file omits settled spend for the whole time a call is unanswered, and a model change during a request both under-prices that request and puts the old base back. H10 (test/unit/modelApiHost.test.ts:2116) reloads only after a finished turn whose save succeeded, and N4 (test/unit/turnNotifications.test.ts:96) only passes a replayed approval, so both drills stay green on those paths. Persist and settle as in the two P1s, mark replayed questions, and assert a reload during an open tool call and a second surface's replayed question.P1: 2. P2: 4.
M83 (wip/m83-import)
I'll review the diff between that base commit and HEAD against the working rules and the plan sections this change touches. Starting with
AGENTS.mdand the diff stat.The change is M83, the agent-import feature. I'll read the plan sections it cites and the new source before judging defects.The plan section is M83. Next I’ll read the import implementation and the host paths that write it.The import path is large enough that the confinement, masking, and apply logic need to be checked against the helpers they call.A few confinement and apply paths still need exact line checks before the report.## Class 1 — concurrency, ordering, stale state, lifecycle
P2.
src/core/import/agentImport.ts:1141(appendSections), reached from a secondmuseSpark.importFromAgentswhile the first is still between that read andappendTextat line 1154. Confirm is a non-modal notification (src/host/agentImportHost.ts:73) andregisterLoggedCommand(src/extension.ts:437) does not serialize the command. Two accepted imports of the same rules section both readAGENTS.mdbefore either appends, both treat the heading as absent, and both append. The file gains the section twice. The re-read only observes an append that has already finished.Fix: keep a single in-flight import and append under an exclusive lock that covers the read and the append.
Class 2 — wire evidence, validation, security
P1.
src/host/importIo.ts:111(appendFile), afterisWithinRootatsrc/core/import/agentImport.ts:1127.canonicalPath(src/host/canonicalPath.ts:35) treatsrealpathENOENTas a missing file and returns the workspace path. A dangling symlink isENOENT, soAGENTS.md→/home/user/.bashrc(target absent, parent present) is judged inside the workspace.readTextalso getsENOENTand the append proceeds with an empty current file.appendFilefollows the symlink and creates the outside file. The body is the repository's rules text. The same destination is read with no confine before confirm:readPlanText(src/host/commands/agentImportCommands.ts:341) usesstat/readFile, which follow a symlink whose target exists, so a link to an existing key file is loaded (up to 16 MiB) while building the preview; apply then refuses that write.memoryImportIo.realPathrewrites a link even when the target is missing, so the apply test does not exercise this.Fix:
lstatAGENTS.mdbefore the plan read and the append, refuse any symlink, and append withO_NOFOLLOW.P1.
src/host/agentImportHost.ts:93(openTarget), for the hooks copy built atsrc/core/import/agentImport.ts:1046. Project hooks are opened with noconfineWorkspacePathcheck.isPathPresentislstat, so a symlink counts as the file being present, and the prompt shows the lexical path.muse/hooks.json(importDisplayPath). A trusted repository whose.museis a symlink to~/.config/muse, or whose.muse/hooks.jsonis a symlink to~/.bashrc, opens that outside file when it exists. Paste-and-save, which the prompt asks for, overwrites it.Fix: confine the hooks path to the workspace root and do not open it when the canonical path leaves the workspace or any component is a symlink.
P2.
src/core/import/importConvert.ts:411.commandgoes throughmaskText;statusMessageis stored raw and then copied into the preview and the clipboard (src/host/commands/agentImportCommands.ts:252fencescopy.textwith no further mask). A handler{ "type": "command", "command": "echo ok", "statusMessage": "<GitHub-shaped token or password=…>" }shows and copies that value.Fix: run
statusMessagethroughmaskTextbefore it is stored on the hook.Class 3 — failure paths, honesty, docs
P2.
src/host/commands/agentImportCommands.ts:276.readPlanTextreturnsundefinedfortooLargeand for a failed decode, and logs nothing in those cases.planImportApply(src/core/import/agentImport.ts:1074) only treats status'read'as a real settings file, so a settings file overAGENT_IMPORT_FILE_MAX_BYTES(64 KiB) — still under the 1 MiB hook-config limit — or a file that is not valid text, yields no known server names andhasLegacyMcpKey: falsewhilehasMuseSettingsFilestays true. The preview says to merge, omits the legacy-key warning, and offers servers the file already has. MergingmcpServersinto a file that already hasmcp_serversis the case Muse Code loads neither key.Fix: when the settings file is present but not read, skip server and hook copies and say the file could not be read.
P2.
src/core/import/agentImport.ts:1154. A rules append is not checked againstRULES_FILE_MAX_BYTES(64 KiB). Source files may be that large, and the section adds a heading, so one full-size rules file makes a newAGENTS.mdexceed the limit. An existingAGENTS.mdof 40 KiB plus a 30 KiB section does the same.loadRuleFile(src/core/context/rules.ts:82) then skips the whole file, including rules that were already loading. The import still reports the section as written. A skill over the same limit is refused (src/core/import/agentImport.ts:527).Fix: refuse the append when the resulting
AGENTS.mdwould exceedRULES_FILE_MAX_BYTES, and say so in the preview.P2.
src/core/import/agentImport.ts:666and:690. Both Claude agent walks passdepth: 1, socollectMarkdownnever enters a subdirectory.agents/review/security.mdis never imported. Command walks useAGENT_IMPORT_COMMAND_MAX_DEPTH(3). PLAN.md M83 "As built" (around line 7128) says agents and commands are both read from namespaced subfolders three deep.Fix: walk Claude
agents/to the same depth ascommands/, and slug the directory prefixes the way commands already do.P1: 2. P2: 5.
M72 lock, third round (wip/m72-checkpoints; the redesign should remove these by construction)
I'll review the diff between that base commit and HEAD against the working rules, starting with
AGENTS.mdand the change summary.The change is M72: checkpoint-store locking across windows. I'll read the plan sections it cites and the new lock code next.The lock lives instoreLock.ts. Next I’ll trace every caller for races around acquire, release, and presence.The lock and turn-end paths are the risk. I'll confirm the races and the docs against the code, then write the report.## Class 1 — concurrency, ordering, stale state, lifecycle
P1.
src/host/conversation/conversationCheckpoints.ts:239,src/host/checkpoints/checkpointStore.ts:1838,src/host/checkpoints/checkpointStore.ts:1847. Ending a turn unpublishes it and records its end together.markTurn(false)drops the turn fromrunningTurnsand, once the presence write finishes,trimJournaldeletes every staged copy not still covered by an older running turn.endTurnhas already removed the turn fromopenTurnsbeforecopiesSince, andheldis filled only if the lock comes back busy. On a normal end the recording is already done, soendTurnstarts before that trim:copiesSincekeeps the staged path,trimJournalunlinks the file, thenfinishTurnrunshash-objecton it (checkpointStore.ts:1085) and throws. That error is notStoreBusyError, so it is not retried, andendAfterRecordingonly logs it (conversationCheckpoints.ts:281). The record never getsendedAtor the tool pre-image, so an ignored file the tools copied (.env) cannot be restored, and every change until the next capture is treated as unsure. If the turn ends beforerecordhas setopenTurns, the trim finishes first and the end is saved with the pre-image already gone. Hold those staged paths for the in-flight end, and trim only afterfinishTurnhas imported them.P1.
src/host/checkpoints/checkpointStore.ts:2068,src/host/checkpoints/checkpointStore.ts:2100,src/host/checkpoints/checkpointStore.ts:2165.forgetSessionandqueueForgetset the archive in memory and, when the other window holds the lock, defer the records write and return.disposeaborts the store and clears the retry timer;scheduleRetry(checkpointStore.ts:491) then does nothing because the signal is already aborted.ConversationCheckpoints.forgettherefore resolves with no error (conversationCheckpoints.ts:610) whilerecords.jsonnever gainsforgotten. The other window keeps that conversation's checkpoints, including copied file bytes.CHANGELOG.md:39says an archive happens later on its own. Write the archive beforeforgetresolves, and on close flush it or report that it was not written.P2.
src/host/conversation/conversationCheckpoints.ts:514. A turn this panel did not start (turnStarted: queued message, scheduled run) is marked withpublish, which is not awaited and whose failure is only logged (conversationCheckpoints.ts:183). The backend is already editing. Until that presence write lands, or for up to one heartbeat if it fails, the other window does not see the turn and will restore over it.SECURITY.md:84andREADME.md:824say the turn is marked before it can change a file. Await the mark, and do not leave a failed publish until the next beat.P2.
src/host/checkpoints/storeLock.ts:219,src/host/checkpoints/storeLock.ts:417. Takeover deletes the moved lock when the owner id still matches, without callingisGoneon the moved file. A heartbeat (utimes, or a presence rewrite) between the stale check and therenameleaves a live window's lock deleted and the taker free to create its own, while the first window still hasisHeldand keeps writing.liveWindowshas the same gap: itrms a presence file it judged gone, thendropOrphanRefs(checkpointStore.ts:649) drops that window's pins andstaging/even if the file was replaced with a fresh one after the read. After the move or before the delete, put the file back unless that same bytes are still gone.P2.
src/host/checkpoints/storeLock.ts:303,src/host/checkpoints/storeLock.ts:427.writePresencechecksisPresentonly beforemkdir/writeFileAtomically.leave(fromdispose) deletes the presence file in that gap, and the in-flight write creates it again with the oldrunninglist and a fresh mtime.beatwill not remove it (isPresentis already false). While that process is still alive, the other window'sturnBlocking(checkpointStore.ts:776) refuses every restore. Re-checkisPresentafter the write and delete the file if the window has left.P2.
src/host/checkpoints/checkpointStore.ts:2166,src/host/checkpoints/checkpointStore.ts:1526.disposeclears the heartbeat and aborts git, butrunStepsnever reads the signal, so a restore already writing files keeps going withisHeldtrue and a lock file that is no longer refreshed. AfterCHECKPOINT_OWNER_STALE_MSthe other window takes the lock over. Those file writes, and theupdate-refs for the index, pins and keep refs (checkpointStore.ts:871,checkpointStore.ts:1381,checkpointStore.ts:1656,checkpointStore.ts:1744,checkpointStore.ts:1784), never callassertHeld, so both windows write.SECURITY.md:81says a write is refused once the lock is taken over. Keep beating until the in-flight operation finishes, and check the lock before each ref update and each file write.Class 2 — wire evidence, validation, security
none
Class 3 — failure paths, honesty, docs
P2.
src/host/checkpoints/checkpointStore.ts:1784,src/host/checkpoints/checkpointStore.ts:1807. Arecordthat throws for any reason other thanStoreBusyErrordoes not schedule deletion of the pin, andConversationCheckpoints.recorddoes not callrelease(conversationCheckpoints.ts:252). Pins of a live window are kept bydropOrphanRefs, so the capture stays for the life of the window. A busy failure afterupdate-refof the keep ref schedules only the pin delete, so the keep ref stays with no record until the nextmaintain.docs/certification/m72.md:401says a failed record leaves nothing behind. On every failedrecord, delete the pin and any keep ref that was created.P2.
src/host/conversation/conversationCheckpoints.ts:471. If the running mark cannot be written,beforeTurnthrows a hardcoded English sentence that embedserror.message.sendMessageposts that string assendFailed(conversationController.ts:2893). AnEACCES/ENOENTfrom the presence write includes the absolute path under the user profile, in the panel and in the log, and the sentence is not inen.tsor the 14l10n/ui.*.jsontables.README.md:826says the panel says why. Show one localized reason and log a fixed phrase.P2.
docs/certification/m72.md:372,test/unit/storeLock.test.ts:139. Drill L9 claims "writes allowed after the lock was taken over" and points at a test that only callsStoreLock.assertHeldafter overwriting the lock file. RemovingassertHeldfromsave,deleteRefsandprune, or theupdate-refs that never call it, still leaves that test green. Point the drill at a store write that must fail once the lock file is no longer this window's.P1: 2. P2: 7.
M68 round 4 (wip/m68-verify-loop)
I'll review the diff between that base commit and HEAD against the working rules, starting with
AGENTS.mdand the plan sections this change touches.The change is one commit on file judging for verify-loop commands. I'll read the plan section and the full diff next.The diff is the conservative command-word split and moving the rule check intoauthorizeCommand. I'll read that path and the M68 plan next.The session-rule invalidation is sampled once, then a later await can change the ledger, and the same rule still allows the shell tool.
Class 1 — concurrency, ordering, stale state, lifecycle
P1
src/core/backends/modelapi/ModelApiHost.ts:3766authorizeCommanddecides at line 3673. WhenchangesWhatRunsNow(ruleCommand)is false and a session rule exists, it returns allow with no further read of the ledger (line 3680).runVerifyCommandthen awaits the guard (line 3762) and only then runs the command (line 3766).then_runalways installs that guard; achangedFilescheck does too when it has paths. A child turn is started withvoid child.session.sendTurn(line 3238), so it runs while the parent is in that await. The child's edit callsparentSession.ledger.noteOutsideEdit(line 4316).Inputs: a session allow for
npm run lint(ornode scripts/t.js); the parent is inthen_runof that command after editingsrc/a.tsonly; duringisAsEdited's read, the child writespackage.jsonorscripts/t.js. The guard only compares thethen_runfile's hash, so it returns true. The command runs under the session rule. A re-check ofchangesWhatRunsNowwould now be true and would ask.Fix: after the guard returns, call
changesWhatRunsNow(ruleCommand)again and, if it is true, take the mode shell verdict instead of running.Class 2 — wire evidence, validation, security
P1
src/core/backends/modelapi/ModelApiHost.ts:4057The check card stores the allow on the shell tool's own key:
allowForSession(shell.name, ruleCommand)(line 2465), andauthorizeCommandbuilds that same key (lines 3667–3671). After the edit, onlyauthorizeCommandignores the rule (line 3673). The shell tool does not.decideAndRunputs the model'scommandstring on the query (lines 4183–4186) andverdictWithHookcallspermissions.verdict(line 4057), which returns allow when that key is present.Inputs: the user has already chosen "Always allow in this session" for
node scripts/t.js. The model editsscripts/t.js, then calls the shell tool with that same command (same round or the next). The automatic check would ask; the shell call does not, and the modified script runs.Fix: in
verdictWithHook/decideAndRun, if the call is the shell tool andchangesWhatRunsNow(command)is true, useverdictFor(mode, 'shell')instead of the session rule.Class 3 — failure paths, honesty, docs
P2
PLAN.md:6677That sentence still defines
canChangeWhatRunsas command-defining files, code-loading files, or a file the command line names.canChangeWhatRuns('src/a.ts', 'node "scripts/my check.js"')is now true: any edited file counts when a word is not plain, and a workspace-relative path matches as a substring. README and CHANGELOG in this diff say that; PLAN does not.Fix: update the M68 permission bullet so it matches
canChangeWhatRunsinsrc/core/verify/codeFiles.ts.P1: 2. P2: 1.
Written by Claude Code at the pause, 2026-09-28.