Skip to content

Resume point: D49 work paused 2026-09-28 (usage limits) #58

Description

@RandyNorthrup

Work on the D49 program (PLAN.md M67–M85) paused on 2026-09-28 when the Claude, Codex and Grok Build usage limits were all reached. This issue is the resume point for a cloud agent or the next session. Follow AGENTS.md and CLAUDE.md as always: every change is reviewed in one pass (three classes: concurrency/lifecycle; wire/validation/security; failure paths/honesty/docs), red-drilled, and gated with the complete npm run quality before it is pushed. A third review round on one module means a redesign, not another patch.

Open pull requests (pushed, gate passed; each waits only for its final Codex review)

PR Milestone Head State
#32 Other editors (ACP agent, M60–M63) 46ba5406 Muse Code's review of the last two rounds found a P1 (cancel while a turn is starting) and a P2; both fixed, drilled, full Win11 gate exit 0. Merge after review; then the M63c cloud agent resumes; then release 0.10.0 (never 1.0).
#52 M69 web fetch 4c4d14c7 Redesign: page text as served, all untrusted. Muse Code's review: one P2 open — picture in OMITTED (src/core/web/htmlToMarkdown.ts) drops the fallback <img> and its alt text; fix by removing 'picture' from the list.
#53 M79 plans as files b6f24e79 Codex threads answered and resolved.
#54 M68 verify loop 51c80e83 Newer work is on wip/m68-verify-loop (below); thread PRRT_kwDOUkzj5M6m0_2F still to answer.
#55 M72 checkpoints 9be57e55 Superseded by the redesign on wip/m72-checkpoints (below).
#57 M67 code intelligence 107510a4 Waiting for review.
#51 External (Android tests) — The owner's call.

Work in progress (branches wip/*; not reviewed or gated as a whole)

Each branch's last commit message states where it stopped. Resume each by reading its docs/certification/m<nn>.md and PLAN section, running npm run quality, fixing, reviewing in one pass, then opening the PR from the branch (rename it feature/... if you like).

Branch Milestone Where it stopped
wip/m68-verify-loop M68 Round-4 fix plus reviewer fixes: re-check changesWhatRunsNow after the guard; check-card grants keyed apart from the shell tool. Then push to PR #54 and answer its thread.
wip/m70-review M70 /review Built, Kubuntu gate failed only on jscpd (fixed in cefa9587); three-class review findings not yet collected. Cert Gate section says "to be recorded".
wip/m71-git-prs M71 git and PRs Built and gated (3af80e40); review fixes in progress (ConversationGit rewrite).
wip/m72-checkpoints M72 checkpoints Mid-redesign: records as git refs updated by compare-and-swap, per-window files, prune grace, no lock. Findings the design must remove: below.
wip/m73-observation-packing M73 Muse Code's draft, not checked. Must stay an M75 evaluation arm with no user setting until a live M75 run shows the floors held.
wip/m74-handoff M74 /handoff Muse Code's draft, not checked.
wip/m75-eval M75 evaluation Rebuilt harness; reviewer fixes to the runner in progress; full 10-task live baseline not run yet.
wip/m76-agents M76 custom agents Muse Code's draft plus its own review fix; not checked by a finisher.
wip/m77-board M77 session board, best-of-N Muse Code's draft with translations; one jscpd clone; no certification record.
wip/m78-auto-safe M78 Auto made safe Builder stopped mid-build.
wip/m81-browser M81 browser check Muse Code's draft; its security claims (loopback-only, pipe transport, temp profile) not yet verified.
wip/m82-awareness M82 budget and notices Fixing the review findings below.
wip/m83-import M83 import from other agents Fixing the review findings below.
wip/m84-export M84 export and share Fixing an independent review's findings; transferInvalidField exists in en.ts only.

Not started: M80 (headless and CI; needs PR #32 merged), M85 (needs the owner's TypeSafe key).

Lesson from this round: Muse Code drafts had the right shape and good translations, but every one had real gaps in security or correctness logic, and several claimed checks that had not run. Treat each Muse draft as unverified and check every claim against the code.

Review findings still to fix (Grok Build, before the limits)

M82 (wip/m82-awareness)

I'll review the diff between that base commit and HEAD against the working rules, starting with AGENTS.md and the change stat.The change is M82 (notices, reply cost, session cap). I'll read the plan sections it cites and the source diffs next.
The last diff read was cut off. I'll finish the reservation, notice, and docs paths, then write the full report.
The spend ledger and the model-switch path still need a line check before the report.## Class 1 — concurrency, ordering, stale state, lifecycle

P1. src/core/backends/modelapi/ModelApiHost.ts:2450 and :6062, shutdown at :5669 and :6446. A capped request that reports usage adds that cost to budgetSpentUsd only after adoptOutput has already pushed the function_call with no output. persist then returns without writing while hasUnansweredSessionCall is true, so for the whole approval or tool wait the file still has the previous spend. Reload or a crash in that window restores budgetSpentUsd ?? 0, and the next reservation is computed against that remainder, so a request that does not fit the real balance is sent. dispose also fires cancel() and does not wait for it; close() then awaits the saving promise captured before closeReservation runs, so a started request aborted by window close is charged only in memory. Record the settled spend, and an in-flight reservation once a stream event has arrived, in a write that is not skipped with the replay, and await that write in close() before it returns.

P1. src/core/backends/modelapi/ModelApiHost.ts:1899 and :1926, with setModel at :5232. The controller applies a model change as soon as it is confirmed (conversationController.ts:2982) and does not wait for the open request. The body already sent still names the old model, but noteUsage prices the reported tokens with this.modelId at settlement time. Switching muse-spark-1.3 to an already-confirmed muse-spark-1.3-contributor while the response is in flight records the bill at the contributor rates ($0.10 / $0.20 per million instead of $1.25 / $4.25), so later requests are reserved against a ledger below what Meta charged and are sent past the cap. The same settlement writes budgetBase back from that open reservation, undoing the clear in setModel, and the next estimate starts from the previous model's reported input count. Stamp the model id onto the reservation when the body is built, settle the dollars with that id, and install the base only when this.modelId is still that id.

P2. src/host/conversation/turnNotifications.ts:64 and :110, src/core/backends/musecode/MuseCodeHost.ts:511 and :1096. A question replayed to a later surface is delivered without isReplayed (onEvent flags only approvalRequested; approval/listPending does the same for userInput/requested). attentionNotice therefore always returns a question notice. The window's key set suppresses the second toast only until BACKGROUND_NOTICE_KEYS_MAX (200) drops the oldest key; the notifier test shows a forgotten key is raised again. After 200 later notices, opening a second surface on a session that still has that question shows the waiting toast again. Mark replayed questions isReplayed and ignore them in attentionNotice, as replayed approvals already are.

Class 2 — wire evidence, validation, security

P2. src/core/backends/modelapi/schemas.ts:165 and ModelApiHost.ts:1900; stored ledger at src/core/backends/modelapi/sessionStore.ts:222. Response usage is z.number() with no finite, non-negative check, and budgetSpentUsd is restored the same way. estimateCostUsd is linear, so a schema-valid output_tokens or input_tokens below zero, or a stored budgetSpentUsd below zero, decreases the ledger. The next reserveRequest then sees more room than the cap has and sends a request that does not fit. Reject non-finite and negative usage before it is added, and reject a negative stored spend.

Class 3 — failure paths, honesty, docs

P2. src/core/backends/modelapi/ModelApiHost.ts:4873. A failed countInputTokens is logged with describe(error), and ModelApiError's message is the envelope's error.message (client.ts:111). That string is free text from the network and is written to the log as sent. Log a status and a fixed protocol word.

P2. package.nls.json:97, README.md session-budget row, and docs/certification/m82.md:98 and :174. Those say spend is kept with the session and that the cap can be passed only by the input estimate or a paid subagent. The file omits settled spend for the whole time a call is unanswered, and a model change during a request both under-prices that request and puts the old base back. H10 (test/unit/modelApiHost.test.ts:2116) reloads only after a finished turn whose save succeeded, and N4 (test/unit/turnNotifications.test.ts:96) only passes a replayed approval, so both drills stay green on those paths. Persist and settle as in the two P1s, mark replayed questions, and assert a reload during an open tool call and a second surface's replayed question.

P1: 2. P2: 4.

M83 (wip/m83-import)

I'll review the diff between that base commit and HEAD against the working rules and the plan sections this change touches. Starting with AGENTS.md and the diff stat.The change is M83, the agent-import feature. I'll read the plan sections it cites and the new source before judging defects.The plan section is M83. Next I’ll read the import implementation and the host paths that write it.The import path is large enough that the confinement, masking, and apply logic need to be checked against the helpers they call.
A few confinement and apply paths still need exact line checks before the report.## Class 1 — concurrency, ordering, stale state, lifecycle

P2. src/core/import/agentImport.ts:1141 (appendSections), reached from a second museSpark.importFromAgents while the first is still between that read and appendText at line 1154. Confirm is a non-modal notification (src/host/agentImportHost.ts:73) and registerLoggedCommand (src/extension.ts:437) does not serialize the command. Two accepted imports of the same rules section both read AGENTS.md before either appends, both treat the heading as absent, and both append. The file gains the section twice. The re-read only observes an append that has already finished.

Fix: keep a single in-flight import and append under an exclusive lock that covers the read and the append.

Class 2 — wire evidence, validation, security

P1. src/host/importIo.ts:111 (appendFile), after isWithinRoot at src/core/import/agentImport.ts:1127. canonicalPath (src/host/canonicalPath.ts:35) treats realpath ENOENT as a missing file and returns the workspace path. A dangling symlink is ENOENT, so AGENTS.md → /home/user/.bashrc (target absent, parent present) is judged inside the workspace. readText also gets ENOENT and the append proceeds with an empty current file. appendFile follows the symlink and creates the outside file. The body is the repository's rules text. The same destination is read with no confine before confirm: readPlanText (src/host/commands/agentImportCommands.ts:341) uses stat/readFile, which follow a symlink whose target exists, so a link to an existing key file is loaded (up to 16 MiB) while building the preview; apply then refuses that write. memoryImportIo.realPath rewrites a link even when the target is missing, so the apply test does not exercise this.

Fix: lstat AGENTS.md before the plan read and the append, refuse any symlink, and append with O_NOFOLLOW.

P1. src/host/agentImportHost.ts:93 (openTarget), for the hooks copy built at src/core/import/agentImport.ts:1046. Project hooks are opened with no confineWorkspacePath check. isPathPresent is lstat, so a symlink counts as the file being present, and the prompt shows the lexical path .muse/hooks.json (importDisplayPath). A trusted repository whose .muse is a symlink to ~/.config/muse, or whose .muse/hooks.json is a symlink to ~/.bashrc, opens that outside file when it exists. Paste-and-save, which the prompt asks for, overwrites it.

Fix: confine the hooks path to the workspace root and do not open it when the canonical path leaves the workspace or any component is a symlink.

P2. src/core/import/importConvert.ts:411. command goes through maskText; statusMessage is stored raw and then copied into the preview and the clipboard (src/host/commands/agentImportCommands.ts:252 fences copy.text with no further mask). A handler { "type": "command", "command": "echo ok", "statusMessage": "<GitHub-shaped token or password=…>" } shows and copies that value.

Fix: run statusMessage through maskText before it is stored on the hook.

Class 3 — failure paths, honesty, docs

P2. src/host/commands/agentImportCommands.ts:276. readPlanText returns undefined for tooLarge and for a failed decode, and logs nothing in those cases. planImportApply (src/core/import/agentImport.ts:1074) only treats status 'read' as a real settings file, so a settings file over AGENT_IMPORT_FILE_MAX_BYTES (64 KiB) — still under the 1 MiB hook-config limit — or a file that is not valid text, yields no known server names and hasLegacyMcpKey: false while hasMuseSettingsFile stays true. The preview says to merge, omits the legacy-key warning, and offers servers the file already has. Merging mcpServers into a file that already has mcp_servers is the case Muse Code loads neither key.

Fix: when the settings file is present but not read, skip server and hook copies and say the file could not be read.

P2. src/core/import/agentImport.ts:1154. A rules append is not checked against RULES_FILE_MAX_BYTES (64 KiB). Source files may be that large, and the section adds a heading, so one full-size rules file makes a new AGENTS.md exceed the limit. An existing AGENTS.md of 40 KiB plus a 30 KiB section does the same. loadRuleFile (src/core/context/rules.ts:82) then skips the whole file, including rules that were already loading. The import still reports the section as written. A skill over the same limit is refused (src/core/import/agentImport.ts:527).

Fix: refuse the append when the resulting AGENTS.md would exceed RULES_FILE_MAX_BYTES, and say so in the preview.

P2. src/core/import/agentImport.ts:666 and :690. Both Claude agent walks pass depth: 1, so collectMarkdown never enters a subdirectory. agents/review/security.md is never imported. Command walks use AGENT_IMPORT_COMMAND_MAX_DEPTH (3). PLAN.md M83 "As built" (around line 7128) says agents and commands are both read from namespaced subfolders three deep.

Fix: walk Claude agents/ to the same depth as commands/, and slug the directory prefixes the way commands already do.

P1: 2. P2: 5.

M72 lock, third round (wip/m72-checkpoints; the redesign should remove these by construction)

I'll review the diff between that base commit and HEAD against the working rules, starting with AGENTS.md and the change summary.The change is M72: checkpoint-store locking across windows. I'll read the plan sections it cites and the new lock code next.The lock lives in storeLock.ts. Next I’ll trace every caller for races around acquire, release, and presence.
The lock and turn-end paths are the risk. I'll confirm the races and the docs against the code, then write the report.## Class 1 — concurrency, ordering, stale state, lifecycle

P1. src/host/conversation/conversationCheckpoints.ts:239, src/host/checkpoints/checkpointStore.ts:1838, src/host/checkpoints/checkpointStore.ts:1847. Ending a turn unpublishes it and records its end together. markTurn(false) drops the turn from runningTurns and, once the presence write finishes, trimJournal deletes every staged copy not still covered by an older running turn. endTurn has already removed the turn from openTurns before copiesSince, and held is filled only if the lock comes back busy. On a normal end the recording is already done, so endTurn starts before that trim: copiesSince keeps the staged path, trimJournal unlinks the file, then finishTurn runs hash-object on it (checkpointStore.ts:1085) and throws. That error is not StoreBusyError, so it is not retried, and endAfterRecording only logs it (conversationCheckpoints.ts:281). The record never gets endedAt or the tool pre-image, so an ignored file the tools copied (.env) cannot be restored, and every change until the next capture is treated as unsure. If the turn ends before record has set openTurns, the trim finishes first and the end is saved with the pre-image already gone. Hold those staged paths for the in-flight end, and trim only after finishTurn has imported them.

P1. src/host/checkpoints/checkpointStore.ts:2068, src/host/checkpoints/checkpointStore.ts:2100, src/host/checkpoints/checkpointStore.ts:2165. forgetSession and queueForget set the archive in memory and, when the other window holds the lock, defer the records write and return. dispose aborts the store and clears the retry timer; scheduleRetry (checkpointStore.ts:491) then does nothing because the signal is already aborted. ConversationCheckpoints.forget therefore resolves with no error (conversationCheckpoints.ts:610) while records.json never gains forgotten. The other window keeps that conversation's checkpoints, including copied file bytes. CHANGELOG.md:39 says an archive happens later on its own. Write the archive before forget resolves, and on close flush it or report that it was not written.

P2. src/host/conversation/conversationCheckpoints.ts:514. A turn this panel did not start (turnStarted: queued message, scheduled run) is marked with publish, which is not awaited and whose failure is only logged (conversationCheckpoints.ts:183). The backend is already editing. Until that presence write lands, or for up to one heartbeat if it fails, the other window does not see the turn and will restore over it. SECURITY.md:84 and README.md:824 say the turn is marked before it can change a file. Await the mark, and do not leave a failed publish until the next beat.

P2. src/host/checkpoints/storeLock.ts:219, src/host/checkpoints/storeLock.ts:417. Takeover deletes the moved lock when the owner id still matches, without calling isGone on the moved file. A heartbeat (utimes, or a presence rewrite) between the stale check and the rename leaves a live window's lock deleted and the taker free to create its own, while the first window still has isHeld and keeps writing. liveWindows has the same gap: it rms a presence file it judged gone, then dropOrphanRefs (checkpointStore.ts:649) drops that window's pins and staging/ even if the file was replaced with a fresh one after the read. After the move or before the delete, put the file back unless that same bytes are still gone.

P2. src/host/checkpoints/storeLock.ts:303, src/host/checkpoints/storeLock.ts:427. writePresence checks isPresent only before mkdir / writeFileAtomically. leave (from dispose) deletes the presence file in that gap, and the in-flight write creates it again with the old running list and a fresh mtime. beat will not remove it (isPresent is already false). While that process is still alive, the other window's turnBlocking (checkpointStore.ts:776) refuses every restore. Re-check isPresent after the write and delete the file if the window has left.

P2. src/host/checkpoints/checkpointStore.ts:2166, src/host/checkpoints/checkpointStore.ts:1526. dispose clears the heartbeat and aborts git, but runSteps never reads the signal, so a restore already writing files keeps going with isHeld true and a lock file that is no longer refreshed. After CHECKPOINT_OWNER_STALE_MS the other window takes the lock over. Those file writes, and the update-refs for the index, pins and keep refs (checkpointStore.ts:871, checkpointStore.ts:1381, checkpointStore.ts:1656, checkpointStore.ts:1744, checkpointStore.ts:1784), never call assertHeld, so both windows write. SECURITY.md:81 says a write is refused once the lock is taken over. Keep beating until the in-flight operation finishes, and check the lock before each ref update and each file write.

Class 2 — wire evidence, validation, security

none

Class 3 — failure paths, honesty, docs

P2. src/host/checkpoints/checkpointStore.ts:1784, src/host/checkpoints/checkpointStore.ts:1807. A record that throws for any reason other than StoreBusyError does not schedule deletion of the pin, and ConversationCheckpoints.record does not call release (conversationCheckpoints.ts:252). Pins of a live window are kept by dropOrphanRefs, so the capture stays for the life of the window. A busy failure after update-ref of the keep ref schedules only the pin delete, so the keep ref stays with no record until the next maintain. docs/certification/m72.md:401 says a failed record leaves nothing behind. On every failed record, delete the pin and any keep ref that was created.

P2. src/host/conversation/conversationCheckpoints.ts:471. If the running mark cannot be written, beforeTurn throws a hardcoded English sentence that embeds error.message. sendMessage posts that string as sendFailed (conversationController.ts:2893). An EACCES/ENOENT from the presence write includes the absolute path under the user profile, in the panel and in the log, and the sentence is not in en.ts or the 14 l10n/ui.*.json tables. README.md:826 says the panel says why. Show one localized reason and log a fixed phrase.

P2. docs/certification/m72.md:372, test/unit/storeLock.test.ts:139. Drill L9 claims "writes allowed after the lock was taken over" and points at a test that only calls StoreLock.assertHeld after overwriting the lock file. Removing assertHeld from save, deleteRefs and prune, or the update-refs that never call it, still leaves that test green. Point the drill at a store write that must fail once the lock file is no longer this window's.

P1: 2. P2: 7.

M68 round 4 (wip/m68-verify-loop)

I'll review the diff between that base commit and HEAD against the working rules, starting with AGENTS.md and the plan sections this change touches.The change is one commit on file judging for verify-loop commands. I'll read the plan section and the full diff next.The diff is the conservative command-word split and moving the rule check into authorizeCommand. I'll read that path and the M68 plan next.
The session-rule invalidation is sampled once, then a later await can change the ledger, and the same rule still allows the shell tool.

Class 1 — concurrency, ordering, stale state, lifecycle

P1 src/core/backends/modelapi/ModelApiHost.ts:3766

authorizeCommand decides at line 3673. When changesWhatRunsNow(ruleCommand) is false and a session rule exists, it returns allow with no further read of the ledger (line 3680). runVerifyCommand then awaits the guard (line 3762) and only then runs the command (line 3766). then_run always installs that guard; a changedFiles check does too when it has paths. A child turn is started with void child.session.sendTurn (line 3238), so it runs while the parent is in that await. The child's edit calls parentSession.ledger.noteOutsideEdit (line 4316).

Inputs: a session allow for npm run lint (or node scripts/t.js); the parent is in then_run of that command after editing src/a.ts only; during isAsEdited's read, the child writes package.json or scripts/t.js. The guard only compares the then_run file's hash, so it returns true. The command runs under the session rule. A re-check of changesWhatRunsNow would now be true and would ask.

Fix: after the guard returns, call changesWhatRunsNow(ruleCommand) again and, if it is true, take the mode shell verdict instead of running.

Class 2 — wire evidence, validation, security

P1 src/core/backends/modelapi/ModelApiHost.ts:4057

The check card stores the allow on the shell tool's own key: allowForSession(shell.name, ruleCommand) (line 2465), and authorizeCommand builds that same key (lines 3667–3671). After the edit, only authorizeCommand ignores the rule (line 3673). The shell tool does not. decideAndRun puts the model's command string on the query (lines 4183–4186) and verdictWithHook calls permissions.verdict (line 4057), which returns allow when that key is present.

Inputs: the user has already chosen "Always allow in this session" for node scripts/t.js. The model edits scripts/t.js, then calls the shell tool with that same command (same round or the next). The automatic check would ask; the shell call does not, and the modified script runs.

Fix: in verdictWithHook / decideAndRun, if the call is the shell tool and changesWhatRunsNow(command) is true, use verdictFor(mode, 'shell') instead of the session rule.

Class 3 — failure paths, honesty, docs

P2 PLAN.md:6677

That sentence still defines canChangeWhatRuns as command-defining files, code-loading files, or a file the command line names. canChangeWhatRuns('src/a.ts', 'node "scripts/my check.js"') is now true: any edited file counts when a word is not plain, and a workspace-relative path matches as a substring. README and CHANGELOG in this diff say that; PLAN does not.

Fix: update the M68 permission bullet so it matches canChangeWhatRuns in src/core/verify/codeFiles.ts.

P1: 2. P2: 1.


Written by Claude Code at the pause, 2026-09-28.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions