Skip to content

M68: verify loop — diagnostics and checks after every edit - #54

Merged
RandyNorthrup merged 17 commits into
mainfrom
feature/m68-verify-loop
Sep 30, 2026
Merged

RandyNorthrup merged 17 commits into
mainfrom
feature/m68-verify-loop

Conversation

@RandyNorthrup

@RandyNorthrup RandyNorthrup commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Summary

Resume the existing M68 verify loop against current main. Preserve M67 code intelligence, M69 web fetch, M79 plan files and PR32's portable/ACP interfaces.

  • Share pending workspace-write notices across live Model API conversations, children and native workspace aliases; keep notices through input resets and invalidate prior check versions at write start and completion.
  • Pin runtime execution to the captured canonical directory and native identity; refuse retargeted/replaced owners at mutations, commands and final publication.
  • Record successful/partial rename writes in the owning verification round. Preserve Stop before the first write.
  • Notify plan publication without loading the lazy backend; capture the exact original session object before lookup/confirmation, so a same-id replacement cannot inherit its write.
  • Preserve conditional formatter replacement, current dirty-buffer checks and owned-stage cleanup. Update docs, translations, host API inventory and certification.

Exact candidate and history

Reviewed/pushed head: b718db33fcfa6c6b92351935af2bcedc77cc8dce.
Committed tree: 0c30f1bffebd51e5e1430ba00137bfed0626ef4a.

Original PR head 51c80e83 and local feature head bf0d2305 remain ancestors. Current main f7db5715 is a normal merge parent. The final ancestry merge retained the verified tree byte for byte; no force push or gate bypass. Original dirty feature/resume worktrees remain untouched.

Local evidence

Check Observed result
Windows host, fresh install + literal quality exit 0 on final tree 0c30f1bf; all source checks, covered unit tests, build/budgets, audit, accessibility, history secrets and SAST
Mac mini / Kubuntu VM / Windows 11 VM fresh install + literal quality exit 0 each on product tree d10c6b42; exact before/after bindings
Installed ACP package 9 tests passed on each rig; shared agent tar SHA-256 18a147bc1204fcb36c027065680a593a255b64eae3b34d89e2f9c70ac0552ee4
Actual VS Code 1.99.0 / Node 20.18.3 19 tests passed per rig, including three real diagnostics/JSON formatting cases, activation and code intelligence; packaged plan module also passed
Accessibility 360 pages per full run, no violations/undecided/missing results; eight existing exemptions
Current bounded behavior proof 52 tests pass before/after 16 production mutants: 15 intended assertions and one explicitly labeled exact contract guard; all nine production hashes restored
Final staged secrets exit 0; ordinary commit hooks also passed
Independent review Root and independent lifecycle reviewer accepted current source, ownership/cancellation seams, fixture corrections and inspected evidence

The final tree differs from the three-rig product tree only in the milestone certification document. Runtime, tests, manifests, dependencies and shipped package inputs are identical. Final host quality and staged secrets were repeated after that documentation-only change.

Windows full runs pass 3,436 tests with 27 skips; Mac/Linux pass 3,426 with 37 platform skips. Windows host coverage is 94.97% statements / 90.27% branches / 96.56% functions / 94.95% lines. Thresholds remain unchanged.

Retained earlier failures: fixture formatting, obsolete optional-signal lint condition, a short-path fixture hold mismatch, Mac disk exhaustion before clone, and a transient Windows generation-file EPERM. Each is recorded separately; the actual-short-path runtime suite passed 34 tests after native canonicalization, and the paid-grant owning suite passed all 22 under the same TEMP before the unchanged full retry passed. No suppressed rules, timeout extensions or ignores were used. The staged scan's source-digest false positive was fixed by naming SHA-256 accurately, without an ignore.

Certification: docs/certification/m68.md. Complete source-bound receipts/logs and platform handoff are retained in the maintainer's external muse-goal-evidence-20260929 directory. Historical live captures retain their original source/date; this resume made zero new model/paid calls. Grok review is explicitly skipped at the owner's instruction because credits are unavailable.

Hosted proof

Exact head b718db33fcfa6c6b92351935af2bcedc77cc8dce: CI 36686125768 passed all seven jobs (Windows/macOS/Ubuntu quality; macOS dictation; package; gitleaks; semgrep). Hosts 36686125195 passed all twelve jobs. Total: 19 successful checks. Both headSha values were independently matched; review threads are resolved. Forks was not triggered.

RandyNorthrup and others added 3 commits September 28, 2026 05:35
PLAN.md D49, milestone M68.

- Model API: after a round that edited files, before the next request,
  the edited files' errors and warnings from VS Code's language servers
  (new and fixed since each file's last check, 50 at most) and the user's
  check commands, in a Check edits row and one note the model reads as
  tool data. museSpark.diagnosticsAfterEdits (on), checkCommands (none)
  and formatOnEdit (off), all machine-scoped.
- Checks, run_checks and then_run take the shell tool's permission path:
  asked wherever a shell command asks (every mode but Bypass), never run
  in Plan or Restricted Mode, "always allow in this session" keyed on the
  configured command; changed paths quoted after --, a path starting with
  - refused; the check's time cap. Three failing rounds in a row stop the
  checks for the turn, and the model and the panel are told.
- then_run on write_file and edit_file (Action Fusion, reimplemented, not
  ported): asked like any command, run only if the file's SHA-256 still
  matches what the edit (and its format) left; the row keeps the diff and
  adds Then ran. Format on edit through the document's formatter.
- Language servers report only on files an editor shows (measured in VS
  Code 1.139.1 and 1.125.0), so edited files, and a file the getDiagnostics
  tool is asked about, open beside the editor as a preview without focus
  while their server reports.
- Muse Code: a MODEL_TEXT note with each message to check edited files
  with mcp__ide__getDiagnostics and run the named checks.
- 21 strings and 7 manifest strings in fifteen languages; harness scenario
  verify; 64 new unit tests, an integration test in both VS Code versions,
  26 red drills; live: case19, 3 requests a run, contributor model.
- The bundle-split gate lists verifyLoop.ts and verifyTools.ts as lazy.
- Docs: README (Checking edits, settings, diagnostics), CHANGELOG, PLAN
  (M68 status), SECURITY, PRIVACY, AGENTS layout, docs/certification/m68.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
P1 (Windows paths): Windows PowerShell 5.1's argument passing and cmd.exe
re-parsing let a crafted path add an option (`a"` + `b --inject` reached
node.exe as `a b`, `--inject`) or run a command through a .cmd
(`x&echo.INJECTED`, `%OS%` expanded, `^` dropped), measured on this machine.
A path with `"` & | < > ^ % ! now refuses the check on Windows, a leading
`@` refuses it everywhere, and only files that exist reach a check
(run_checks refuses a path that names nothing). Round trips through
PowerShell 5.1 into node.exe and into a .cmd are tests.

P1 (hooks): checks, run_checks and then_run go through the user's
PreToolUse, PostToolUse and PostToolUseFailure hooks as calls of the shell
tool (block, rewrite, ask, context, reason, stop), and a hook's denial is
told apart from the user's Reject, whose feedback is kept.

Also: a session rule no longer answers after the conversation edited a
file that decides what the command runs (package.json, Makefile, configs, a
file the command names); one VerifyState per session, reset only by a
user's message (rejections and the fix loop's stop hold for run_checks and
across a goal's wake); no check run twice for one edit, none after the last
round; one 64,000-character budget for the note; code the editor runs
(eslint.config.js, package.json, node_modules) is never shown or
formatted, and nothing more after one is written; a file with no report,
unsaved, or past the first 8 is "not checked", never clean, and the
baseline moves only once the note is delivered; one editor queue with
Stop honoured, timers started after the show, own tabs closed after the
read; getDiagnostics confined by real path, stopped with its MCP caller,
and answered from what was read before the tab closed (the integration
test found the JSON server clearing diagnostics on close); a failed format
write-back keeps the edit; honest rows and exports ("Failed without an
exit code", a skipped then_run exported as not run, the summary exported);
the Muse Code note names getDiagnostics only with the ide server; the
harness scenario fixed and reshot.

Tests, 40 red drills with SHA-256 restore, the integration test on VS Code
1.139.1 and 1.125.0, docs (README, CHANGELOG, SECURITY, PRIVACY, AGENTS,
PLAN M68/§8/§9, docs/certification/m68.md). Owner decisions stay open: the
side editor group, the upstream MSP ask, the diagnosticsAfterEdits default.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-28T19:18:31.765128Z 51c80e8 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c61c8fac63

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/core/backends/modelapi/ModelApiHost.ts Outdated
Comment thread src/host/editor/verifyEditor.ts Outdated
Comment thread src/host/editor/verifyEditor.ts
Comment thread src/core/backends/modelapi/ModelApiHost.ts Outdated
- A PostToolUse or PostToolUseFailure hook that ends the turn now stops the
  remaining checks: runChecks, the one loop behind the automatic checks and
  run_checks, breaks as soon as a hook's stop is recorded.
- A caller queued for the editor (the loop or getDiagnostics) races its wait
  with its own stop and leaves at once; the next caller still waits for
  every one ahead of it.
- A file an editor already shows is read when its server reported on it
  since the write (a report log kept for the extension's life, compared
  with the file's modification time), or when, with no new report, the
  shown document holds exactly what is on disk.
- run_checks rounds without edits count toward the fix loop's bound.

Tests for each, drills R41 to R46 with SHA-256 restore, the integration
test on VS Code 1.139.1 and 1.125.0, PLAN and the certification updated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1b29fca0ec

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/core/backends/modelapi/tools.ts Outdated
Comment thread src/core/backends/modelapi/ModelApiHost.ts Outdated
Comment thread src/host/editor/verifyEditor.ts Outdated
…nd 2)

Codex found three on 1b29fca; each fix closes its class. Every act M68
performs on a file after an await now uses the path confinement found at
the edit (the real path, links and junctions resolved, and the canonical
workspace-relative name) and checks the file just before acting:

- format on edit: the formatter opens only an unchanged real path, and the
  write-back happens only while the file still holds what the edit wrote
  (a change made while the formatter ran stands, logged);
- then_run: the SHA-256 guard right before the command (unchanged);
- the automatic diagnostics: by the real path (was the lexical one), the
  real path and the edit's fingerprint checked before the open and again
  before the read ("not checked: changed" otherwise);
- getDiagnostics: its confined real path re-checked before open and read;
- check path arguments: the canonical names, each re-resolved to the same
  real path just before the command ("not run: the file changed").

Showing a file reuses a tab it has outside the user's group as it is (its
preview state kept), and afterwards the tab that was in front of each
group comes back.

The fingerprint moves to src/core/verify/fingerprint.ts, shared by the
tools and the editor. Promise.withResolvers is gone from the editor queue
(VS Code 1.99 runs Node 20). Tests, drills R47 to R56, the integration
test on 1.139.1 and 1.125.0, and the cert's list of every guarded site.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 45f936cd1d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/core/backends/modelapi/tools.ts Outdated
Comment thread src/core/backends/modelapi/ModelApiHost.ts
Comment thread src/core/backends/modelapi/ModelApiHost.ts Outdated
Comment thread src/core/backends/modelapi/ModelApiHost.ts Outdated
RandyNorthrup and others added 5 commits September 28, 2026 10:07
# Conflicts:
#	README.md
#	docs/certification/README.md
Codex's third round found four more; by the owner's rule the two areas they
come from are redesigned, not patched.

One conditional write: fsAtomic.writeFileIfUnchanged does every atomic
write's steps and, immediately before each rename attempt, reads and
hashes the target's bytes; a target that no longer holds the expected text
(or is gone) is left alone. ToolIo carries it, and format on edit's
write-back goes through it. The residual window between that read and the
rename itself is documented (PLAN §9, SECURITY, the certification).

One VerifyLedger per session replaces VerifyState and the turn's
editedInRound, checksSinceEdit and roundCheckRuns: every edit advances a
version; every check that finished (automatic, run_checks, or a then_run
of a configured check's command) is recorded against the versions of what
it covered; a check is not run again while a run of it is current; a
round's verdict reads only the runs since the previous verdict that are
still on the latest state; the fix loop advances only on a failed verdict.
reset() is called from every path that admits user input (a queued turn,
each drained steer); a goal's wake keeps it.

Tests for the four cases and the ledger's rules, drills R57 to R68, the
integration tests on VS Code 1.139.1 and 1.125.0, and docs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- P1: a steer no longer clears what the conversation wrote. The ledger
  has two resets: resetForMessage (everything) and resetForSteer (the fix
  loop, the rejections, the runs), so a steer never lets a session rule
  answer again for a check whose script the model rewrote, nor resumes
  showing or formatting after the model wrote a config the editor runs.
- A round passes only when no current run of any check failed; the latest
  run of each check over the same files counts.
- then_run is recorded as a check only for a check that takes no files,
  and a time-out only when the check's cap is the shell's.
- A subagent's edits advance its parent's version of the file.
- The ledger resets only once a user's message is admitted (after the
  start and UserPromptSubmit hooks), and never in a subagent.
- The write-back re-checks unsaved changes; the conditional write compares
  before it starts and never makes a folder again; the refusal is
  internal (MODEL_TEXT.fileChangedBeforeWrite removed); the log names every
  case.
- The residual window is described per platform (POSIX, Windows); hard
  links are documented; stale comments and docs brought up to date; the
  cert lists R57 to R75 with what each broke, and no longer claims a gate
  result that was not recorded.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- P1: a subagent's edit now counts for its parent as the parent's own:
  noteOutsideEdit takes the file and its names, so after a subagent writes
  package.json or eslint.config.js the parent's session rule stops
  answering for the check and its editor stops showing and formatting.
- A check is recorded against the state it started on (ledger snapshot
  taken before the command), so an edit made while it ran leaves it behind.
- A then_run's pass counts for a check only when the check's cap is no
  shorter than the shell's; a time-out only when it is no longer.
- The format write-back asks about unsaved text by the name given and the
  real path, first and again right before the rename (the conditional
  write's isReplaceable).
- CHANGELOG: the then_run rule holds only for a check that takes no files.

Tests for each, drills R76 to R80, and R57 to R80 rerun on the final code.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RandyNorthrup

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 51c80e8367

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/core/verify/codeFiles.ts Outdated
RandyNorthrup and others added 5 commits September 28, 2026 13:27
…ound 4)

canChangeWhatRuns split a command on whitespace, so a quoted script path
(`node "scripts/my check.js"`) never matched the edited script and an
always-allowed check ran the model-modified script without asking.

A command is now split into words only when every word is plain (letters,
digits, `_ . / : + -`, and `=` between an option and its value); anything
a shell could read otherwise (quotes, escapes, variables, substitutions,
globs, braces, operators, PowerShell's `@` and `,`) makes the words
uncertain, and then any edited file counts as changing what the command
runs. An edited file whose path (either slash form) occurs anywhere in the
command's text counts too.

Sibling: authorizeCommand now judges the rule on the command it is keyed
on, a PreToolUse hook's rewrite included, instead of on the command before
the rewrite.

Tests, drills R81 to R83, README, CHANGELOG and the certification.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The owner's option B for Grok's two P1s on a9f665b:

- The verify loop's "Always allow in this session" grants are keyed apart
  from the shell tool (VERIFY_COMMAND_RULE_KEY): a check's or then_run's
  grant never answers for the model's own shell call of the same command,
  nor the shell's for them. The card and the hooks still show the shell
  tool; the approval path records a grant under the query's key.
- authorizeThenGuard judges the rule again after the guard: if a file that
  decides what the command runs was edited meanwhile (a subagent's edit
  during the guard's read), the command is authorized again and guarded
  again before it runs (at most twice).

The shell tool's own session rules keep their pre-M68 behaviour; whether
they should lapse when the model edits a file the command names is PLAN
§3 Q10. Tests, drills R84 and R85, README, CHANGELOG and the certification.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Round-4 fix (f2d20a5) plus reviewer fixes in progress: re-check changesWhatRunsNow after the guard; key check-card grants apart from the shell tool (not option b). PR #54 thread PRRT_kwDOUkzj5M6m0_2F still unanswered.

Not reviewed or gated as a whole. Resume: read the milestone's
docs/certification record and PLAN.md section, run npm run quality,
fix, review in one pass, then open the PR.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RandyNorthrup
RandyNorthrup merged commit 3270944 into main Sep 30, 2026
19 checks passed
@RandyNorthrup
RandyNorthrup deleted the feature/m68-verify-loop branch September 30, 2026 22:18
RandyNorthrup added a commit that referenced this pull request Oct 1, 2026
… candidate

The staged tree of the 2026-09-30 pause (index 0a70cc1), on top of main
3270944 (PRs #32, #52, #53, #54, #57 merged): checkpoint records as git refs
updated by compare-and-swap, per-window index and presence files, checkpointed
ordinary and conditional writes, host close before SessionEnd, the checkpoint
store as its own bundle, and version 0.10.0 with the release notes promoted.

Not in this commit, and still open (docs/certification/m72.md, issue #58): the
memory writes through checkpoints, the memory GUI activity lease, the final
admission callback of user `!` commands, the unavailable-interpreter proof, the
Windows long-path initializer, and the macOS short profile for the integration
tests.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant