M68: verify loop — diagnostics and checks after every edit - #54
Conversation
PLAN.md D49, milestone M68. - Model API: after a round that edited files, before the next request, the edited files' errors and warnings from VS Code's language servers (new and fixed since each file's last check, 50 at most) and the user's check commands, in a Check edits row and one note the model reads as tool data. museSpark.diagnosticsAfterEdits (on), checkCommands (none) and formatOnEdit (off), all machine-scoped. - Checks, run_checks and then_run take the shell tool's permission path: asked wherever a shell command asks (every mode but Bypass), never run in Plan or Restricted Mode, "always allow in this session" keyed on the configured command; changed paths quoted after --, a path starting with - refused; the check's time cap. Three failing rounds in a row stop the checks for the turn, and the model and the panel are told. - then_run on write_file and edit_file (Action Fusion, reimplemented, not ported): asked like any command, run only if the file's SHA-256 still matches what the edit (and its format) left; the row keeps the diff and adds Then ran. Format on edit through the document's formatter. - Language servers report only on files an editor shows (measured in VS Code 1.139.1 and 1.125.0), so edited files, and a file the getDiagnostics tool is asked about, open beside the editor as a preview without focus while their server reports. - Muse Code: a MODEL_TEXT note with each message to check edited files with mcp__ide__getDiagnostics and run the named checks. - 21 strings and 7 manifest strings in fifteen languages; harness scenario verify; 64 new unit tests, an integration test in both VS Code versions, 26 red drills; live: case19, 3 requests a run, contributor model. - The bundle-split gate lists verifyLoop.ts and verifyTools.ts as lazy. - Docs: README (Checking edits, settings, diagnostics), CHANGELOG, PLAN (M68 status), SECURITY, PRIVACY, AGENTS layout, docs/certification/m68.md. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
P1 (Windows paths): Windows PowerShell 5.1's argument passing and cmd.exe
re-parsing let a crafted path add an option (`a"` + `b --inject` reached
node.exe as `a b`, `--inject`) or run a command through a .cmd
(`x&echo.INJECTED`, `%OS%` expanded, `^` dropped), measured on this machine.
A path with `"` & | < > ^ % ! now refuses the check on Windows, a leading
`@` refuses it everywhere, and only files that exist reach a check
(run_checks refuses a path that names nothing). Round trips through
PowerShell 5.1 into node.exe and into a .cmd are tests.
P1 (hooks): checks, run_checks and then_run go through the user's
PreToolUse, PostToolUse and PostToolUseFailure hooks as calls of the shell
tool (block, rewrite, ask, context, reason, stop), and a hook's denial is
told apart from the user's Reject, whose feedback is kept.
Also: a session rule no longer answers after the conversation edited a
file that decides what the command runs (package.json, Makefile, configs, a
file the command names); one VerifyState per session, reset only by a
user's message (rejections and the fix loop's stop hold for run_checks and
across a goal's wake); no check run twice for one edit, none after the last
round; one 64,000-character budget for the note; code the editor runs
(eslint.config.js, package.json, node_modules) is never shown or
formatted, and nothing more after one is written; a file with no report,
unsaved, or past the first 8 is "not checked", never clean, and the
baseline moves only once the note is delivered; one editor queue with
Stop honoured, timers started after the show, own tabs closed after the
read; getDiagnostics confined by real path, stopped with its MCP caller,
and answered from what was read before the tab closed (the integration
test found the JSON server clearing diagnostics on close); a failed format
write-back keeps the edit; honest rows and exports ("Failed without an
exit code", a skipped then_run exported as not run, the summary exported);
the Muse Code note names getDiagnostics only with the ide server; the
harness scenario fixed and reshot.
Tests, 40 red drills with SHA-256 restore, the integration test on VS Code
1.139.1 and 1.125.0, docs (README, CHANGELOG, SECURITY, PRIVACY, AGENTS,
PLAN M68/§8/§9, docs/certification/m68.md). Owner decisions stay open: the
side editor group, the upstream MSP ask, the diagnosticsAfterEdits default.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c61c8fac63
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
- A PostToolUse or PostToolUseFailure hook that ends the turn now stops the remaining checks: runChecks, the one loop behind the automatic checks and run_checks, breaks as soon as a hook's stop is recorded. - A caller queued for the editor (the loop or getDiagnostics) races its wait with its own stop and leaves at once; the next caller still waits for every one ahead of it. - A file an editor already shows is read when its server reported on it since the write (a report log kept for the extension's life, compared with the file's modification time), or when, with no new report, the shown document holds exactly what is on disk. - run_checks rounds without edits count toward the fix loop's bound. Tests for each, drills R41 to R46 with SHA-256 restore, the integration test on VS Code 1.139.1 and 1.125.0, PLAN and the certification updated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1b29fca0ec
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…nd 2) Codex found three on 1b29fca; each fix closes its class. Every act M68 performs on a file after an await now uses the path confinement found at the edit (the real path, links and junctions resolved, and the canonical workspace-relative name) and checks the file just before acting: - format on edit: the formatter opens only an unchanged real path, and the write-back happens only while the file still holds what the edit wrote (a change made while the formatter ran stands, logged); - then_run: the SHA-256 guard right before the command (unchanged); - the automatic diagnostics: by the real path (was the lexical one), the real path and the edit's fingerprint checked before the open and again before the read ("not checked: changed" otherwise); - getDiagnostics: its confined real path re-checked before open and read; - check path arguments: the canonical names, each re-resolved to the same real path just before the command ("not run: the file changed"). Showing a file reuses a tab it has outside the user's group as it is (its preview state kept), and afterwards the tab that was in front of each group comes back. The fingerprint moves to src/core/verify/fingerprint.ts, shared by the tools and the editor. Promise.withResolvers is gone from the editor queue (VS Code 1.99 runs Node 20). Tests, drills R47 to R56, the integration test on 1.139.1 and 1.125.0, and the cert's list of every guarded site. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 45f936cd1d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
# Conflicts: # README.md # docs/certification/README.md
Codex's third round found four more; by the owner's rule the two areas they come from are redesigned, not patched. One conditional write: fsAtomic.writeFileIfUnchanged does every atomic write's steps and, immediately before each rename attempt, reads and hashes the target's bytes; a target that no longer holds the expected text (or is gone) is left alone. ToolIo carries it, and format on edit's write-back goes through it. The residual window between that read and the rename itself is documented (PLAN §9, SECURITY, the certification). One VerifyLedger per session replaces VerifyState and the turn's editedInRound, checksSinceEdit and roundCheckRuns: every edit advances a version; every check that finished (automatic, run_checks, or a then_run of a configured check's command) is recorded against the versions of what it covered; a check is not run again while a run of it is current; a round's verdict reads only the runs since the previous verdict that are still on the latest state; the fix loop advances only on a failed verdict. reset() is called from every path that admits user input (a queued turn, each drained steer); a goal's wake keeps it. Tests for the four cases and the ledger's rules, drills R57 to R68, the integration tests on VS Code 1.139.1 and 1.125.0, and docs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- P1: a steer no longer clears what the conversation wrote. The ledger has two resets: resetForMessage (everything) and resetForSteer (the fix loop, the rejections, the runs), so a steer never lets a session rule answer again for a check whose script the model rewrote, nor resumes showing or formatting after the model wrote a config the editor runs. - A round passes only when no current run of any check failed; the latest run of each check over the same files counts. - then_run is recorded as a check only for a check that takes no files, and a time-out only when the check's cap is the shell's. - A subagent's edits advance its parent's version of the file. - The ledger resets only once a user's message is admitted (after the start and UserPromptSubmit hooks), and never in a subagent. - The write-back re-checks unsaved changes; the conditional write compares before it starts and never makes a folder again; the refusal is internal (MODEL_TEXT.fileChangedBeforeWrite removed); the log names every case. - The residual window is described per platform (POSIX, Windows); hard links are documented; stale comments and docs brought up to date; the cert lists R57 to R75 with what each broke, and no longer claims a gate result that was not recorded. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- P1: a subagent's edit now counts for its parent as the parent's own: noteOutsideEdit takes the file and its names, so after a subagent writes package.json or eslint.config.js the parent's session rule stops answering for the check and its editor stops showing and formatting. - A check is recorded against the state it started on (ledger snapshot taken before the command), so an edit made while it ran leaves it behind. - A then_run's pass counts for a check only when the check's cap is no shorter than the shell's; a time-out only when it is no longer. - The format write-back asks about unsaved text by the name given and the real path, first and again right before the rename (the conditional write's isReplaceable). - CHANGELOG: the then_run rule holds only for a check that takes no files. Tests for each, drills R76 to R80, and R57 to R80 rerun on the final code. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 51c80e8367
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…ound 4) canChangeWhatRuns split a command on whitespace, so a quoted script path (`node "scripts/my check.js"`) never matched the edited script and an always-allowed check ran the model-modified script without asking. A command is now split into words only when every word is plain (letters, digits, `_ . / : + -`, and `=` between an option and its value); anything a shell could read otherwise (quotes, escapes, variables, substitutions, globs, braces, operators, PowerShell's `@` and `,`) makes the words uncertain, and then any edited file counts as changing what the command runs. An edited file whose path (either slash form) occurs anywhere in the command's text counts too. Sibling: authorizeCommand now judges the rule on the command it is keyed on, a PreToolUse hook's rewrite included, instead of on the command before the rewrite. Tests, drills R81 to R83, README, CHANGELOG and the certification. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md
The owner's option B for Grok's two P1s on a9f665b: - The verify loop's "Always allow in this session" grants are keyed apart from the shell tool (VERIFY_COMMAND_RULE_KEY): a check's or then_run's grant never answers for the model's own shell call of the same command, nor the shell's for them. The card and the hooks still show the shell tool; the approval path records a grant under the query's key. - authorizeThenGuard judges the rule again after the guard: if a file that decides what the command runs was edited meanwhile (a subagent's edit during the guard's read), the command is authorized again and guarded again before it runs (at most twice). The shell tool's own session rules keep their pre-M68 behaviour; whether they should lapse when the model edits a file the command names is PLAN §3 Q10. Tests, drills R84 and R85, README, CHANGELOG and the certification. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Round-4 fix (f2d20a5) plus reviewer fixes in progress: re-check changesWhatRunsNow after the guard; key check-card grants apart from the shell tool (not option b). PR #54 thread PRRT_kwDOUkzj5M6m0_2F still unanswered. Not reviewed or gated as a whole. Resume: read the milestone's docs/certification record and PLAN.md section, run npm run quality, fix, review in one pass, then open the PR. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… candidate The staged tree of the 2026-09-30 pause (index 0a70cc1), on top of main 3270944 (PRs #32, #52, #53, #54, #57 merged): checkpoint records as git refs updated by compare-and-swap, per-window index and presence files, checkpointed ordinary and conditional writes, host close before SessionEnd, the checkpoint store as its own bundle, and version 0.10.0 with the release notes promoted. Not in this commit, and still open (docs/certification/m72.md, issue #58): the memory writes through checkpoints, the memory GUI activity lease, the final admission callback of user `!` commands, the unavailable-interpreter proof, the Windows long-path initializer, and the macOS short profile for the integration tests. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Summary
Resume the existing M68 verify loop against current main. Preserve M67 code intelligence, M69 web fetch, M79 plan files and PR32's portable/ACP interfaces.
Exact candidate and history
Reviewed/pushed head:
b718db33fcfa6c6b92351935af2bcedc77cc8dce.Committed tree:
0c30f1bffebd51e5e1430ba00137bfed0626ef4a.Original PR head
51c80e83and local feature headbf0d2305remain ancestors. Current mainf7db5715is a normal merge parent. The final ancestry merge retained the verified tree byte for byte; no force push or gate bypass. Original dirty feature/resume worktrees remain untouched.Local evidence
0c30f1bf; all source checks, covered unit tests, build/budgets, audit, accessibility, history secrets and SASTd10c6b42; exact before/after bindings18a147bc1204fcb36c027065680a593a255b64eae3b34d89e2f9c70ac0552ee4The final tree differs from the three-rig product tree only in the milestone certification document. Runtime, tests, manifests, dependencies and shipped package inputs are identical. Final host quality and staged secrets were repeated after that documentation-only change.
Windows full runs pass 3,436 tests with 27 skips; Mac/Linux pass 3,426 with 37 platform skips. Windows host coverage is 94.97% statements / 90.27% branches / 96.56% functions / 94.95% lines. Thresholds remain unchanged.
Retained earlier failures: fixture formatting, obsolete optional-signal lint condition, a short-path fixture hold mismatch, Mac disk exhaustion before clone, and a transient Windows generation-file
EPERM. Each is recorded separately; the actual-short-path runtime suite passed 34 tests after native canonicalization, and the paid-grant owning suite passed all 22 under the same TEMP before the unchanged full retry passed. No suppressed rules, timeout extensions or ignores were used. The staged scan's source-digest false positive was fixed by naming SHA-256 accurately, without an ignore.Certification:
docs/certification/m68.md. Complete source-bound receipts/logs and platform handoff are retained in the maintainer's externalmuse-goal-evidence-20260929directory. Historical live captures retain their original source/date; this resume made zero new model/paid calls. Grok review is explicitly skipped at the owner's instruction because credits are unavailable.Hosted proof
Exact head
b718db33fcfa6c6b92351935af2bcedc77cc8dce: CI 36686125768 passed all seven jobs (Windows/macOS/Ubuntu quality; macOS dictation; package; gitleaks; semgrep). Hosts 36686125195 passed all twelve jobs. Total: 19 successful checks. Both headSha values were independently matched; review threads are resolved. Forks was not triggered.