From 06028225f4341fa09f5ef8b3bfcb091ee53d2a4a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Micha=C5=82=20Pierzcha=C5=82a?= Date: Fri, 25 Sep 2026 09:10:00 +0200 Subject: [PATCH 1/3] docs(skill): add a test-audit skill with the mutation discipline that worked MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adapts the upstream test-audit card to this repository. The authoring gate, junk patterns, retention bar, candidate-evidence fields, and subagent dispatch plan are its structure; what changes is everything this repo actually enforces. Junk patterns are the shapes that turned up as real findings here rather than a generic list: type pins that survive erasure (`Equal = true`, `X ? true : never`, `satisfies` compared after it stops existing at runtime), a lone `typeof x === 'function'` sitting directly above the test that calls `x`, an ordering claim with no ordering assertion, and a `switch` written inside a test body and compared to its own arms. The retention bar is likewise earned: registry and gate-manifest completeness, cross-language source inspection, `@ts-expect-error` directives (an unused one is a compile error, so they do bite), exact-array facets where a membership scan cannot see a duplicate, identity pins where deep-equality does not subsume `toBe`, and vocabulary pins that survive a production-and-fixture co-update. The mutation rule is the part worth keeping: plant the production edit that represents the bug, show the old body green and the new body red with both counts, revert, and show `git status` clean. Arguing detection is not evidence; measuring it is. Validation commands are this repo's — `--project unit-core`/`apple-runner`, `oxlint`, `tsc -b` per package plus the root project, `check:affected --run`, `check:fallow` — and record that the upstream `run-vitest.mjs` and `check-changed.mjs` helpers do not exist here, so nobody goes looking. It also notes the trap found the hard way: `scripts/**` tests are in no `tsconfig` include, so a helper added there needs an explicit type-check. Routed from the AGENTS.md task table, which stays inside its 10 kB budget at 8.0 kB. --- AGENTS.md | 1 + skills/test-audit/SKILL.md | 78 ++++++++++++++++++++++++++++++++++++++ 2 files changed, 79 insertions(+) create mode 100644 skills/test-audit/SKILL.md diff --git a/AGENTS.md b/AGENTS.md index 9e592a8a93..97b8c6630d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -12,6 +12,7 @@ Load only the procedures relevant to the task: | Domain vocabulary | `CONTEXT.md`, `docs/agents/domain.md` | | Architecture decisions | `docs/adr/README.md` | | Tests or gate selection | `docs/agents/testing.md` | +| Writing, auditing, or sweeping tests; dispatching test-audit subagents | `skills/test-audit/SKILL.md` | | Selector capture, polling, or interaction fast paths | `docs/agents/selector-capture.md` | | Adding or changing a CLI flag | `docs/agents/cli-flags.md` | | Opening or reviewing a PR | `docs/agents/pull-requests.md` | diff --git a/skills/test-audit/SKILL.md b/skills/test-audit/SKILL.md new file mode 100644 index 0000000000..7f08963bf5 --- /dev/null +++ b/skills/test-audit/SKILL.md @@ -0,0 +1,78 @@ +--- +name: test-audit +description: Write, change, review, or sweep tests. Authoring gate for new tests, plus an audit workflow and parallel subagent dispatch plan for finding tests that cannot fail, duplicating proof, or test-only production seams. +--- + +# Test audit + +Three modes, one value criterion: a test earns its cost by guarding behavior, credible regressions, or independently meaningful contracts. Optimize for confidence, not deletion count. + +## Authoring gate + +Before adding a test, answer; if you can't, don't add it yet: + +1. Which observable behavior, invariant, or independent contract does it guard? +2. Which credible regression makes it fail? Name the production edit. +3. Why doesn't existing coverage catch it? Each contract has one primary owner at the strongest boundary. +4. Does it need a production seam no production caller needs? Move the test to the real boundary instead. + +Bug regressions must fail on pre-fix code for the intended reason. A regression test that never visibly failed is proving the mock. + +## Junk patterns + +Authoring rejects matches; audits hunt for existing matches: + +- assertion-free probes; lone `typeof x === 'function'` followed by a sibling that calls `x`; +- self-comparison (`f(x)` vs `f(x)`, `map.get(k)` vs itself, fixture vs re-spelled fixture); +- expected value from the helper under test (oracle-as-expectation); a `switch` written in the test body; +- byte-identical or renamed `test()` blocks; the same golden replayed through a second entry point; +- type pins posing as runtime assertions: `Equal = true`, `X ? true : never`, `satisfies` compared after erasure, `toHaveLength(n)` over test literals. Erase them mentally first; if nothing executable remains, convert to the `void [fixtures]` idiom; +- mocks asserting their own inputs; negative controls passing for the wrong reason; +- names promising more than the input exercises (an ordering claim with no ordering assertion is a repair candidate, not a deletion); +- dead code called only by tests; tests preserving test-only exports. + +## Retention criteria + +Keep (and refuse to delete) tests enforcing: command-registry/gate-manifest/layering/DI-seam completeness; golden tables and cross-language source inspection (`protocol.test.ts` grepping Obj-C is a real detector); `@ts-expect-error` pins (unused directives fail `tsc`, so they bite); deliberate identity pins where deep-equal doesn't subsume `toBe`; exact-array facets where a membership scan can't see duplicates; vocabulary pins that survive production+fixture co-updates. Static-or-slow isn't deletion-worthy; similarity to implementation isn't either — prove you can't break it. + +## Candidate evidence + +Record before editing; missing field ⇒ not ready: + +- location + test name; detectable failure (or why none exists); +- non-test callers of covered seams; stronger remaining owner-boundary proof (quote both assertion sets, file:line, show superset relation); +- skip guards in candidate and sibling (`skipIf`, `.skip`, `AGENT_DEVICE_*` env, whether any workflow or script sets it); +- which Vitest project includes the file (`unit-core`, `apple-runner`, `provider-integration`, `fuzz-worker`, `interaction-contract`, `output-economy`) — unmatched = never-run finding; note: `macos.yml` names its darwin files explicitly; +- unlock deletion; focused validation command; risk. + +## Mutation discipline + +The audit's core proof: plant a production edit that represents the bug the test claims to catch; the pre-edit body must stay green while the post-edit body goes red. Report both counts (`8 passed` → `1 failed | 7 passed`). Revert and show `git status` clean before committing. For deletions, prove the sibling catches the mutant. Never argue detection — measure it. + +## Validation + +This repo's commands (upstream `run-vitest.mjs`/`check-changed.mjs` don't exist): + +1. Focused: `npx vitest run --project unit-core ` (or `--project apple-runner` for `packages/platform-apple/src/runner/**`). +2. `pnpm format` (whole repo), `npx oxlint . --deny-warnings`, `npx tsc -b packages/...` for package edits, `npx tsc -p tsconfig.json` for `src/`. +3. `pnpm check:affected --run` at the pushed commit; `pnpm check:fallow --base origin/main` when touching production exports. +4. `git diff --numstat`; report production/test deltas separately. +5. Don't edit during Vitest runs; `scripts/**` tests are outside `tsc` scope — type-check new helpers there explicitly. + +## Dispatching subagents + +For broad scope, run parallel read-only lanes and keep editing ownership exclusive: + +- subsystem lanes (platform packages; `src/daemon`; contracts/kernel/registry/capture-kit/selectors; `test/integration` + maestro/ad-replay/session-journal + provider packages; `scripts/**`); +- one cross-cutting pattern sweep (junk-pattern greps across all tests, each hit read and confirmed); +- one never-run inventory (project includes, gate ownership via `scripts/gate` and `.github/workflows`, unsatisfiable guards). + +Mandate in every prompt: READ-ONLY; read the production owner and name the bug; verify skip guards and project selection; retention list above verbatim; "a false positive costs more than a miss"; output capped, highest-confidence first, with validation commands. + +## Landing + +One coherent owner-boundary batch per PR; prefer net-negative production LOC; delete the test-only seam with the test. Push only when authorized; rebase on review feedback and answer findings by fixing the rule, not the cited site. After landing, rebase to `main` and rediscover; don't carry stale candidate lists. + +## Handoff + +Report: categories deleted; owner simplifications; retained false positives with reasons; mutations run with before/after counts; production vs test LOC; PR state; named follow-ups needing owner judgment (never-run env gates, tsconfig coverage gaps, suspected product bugs). From 4791a65d5d15119d6bb08e12360c8262e57e7296 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Micha=C5=82=20Pierzcha=C5=82a?= Date: Fri, 25 Sep 2026 12:51:53 +0200 Subject: [PATCH 2/3] docs(agents): route the test-audit procedure through docs/agents and repair its commands --- AGENTS.md | 2 +- docs/agents/test-audit.md | 57 +++++++++++++++++++++++++++++ skills/test-audit/SKILL.md | 74 +++----------------------------------- 3 files changed, 62 insertions(+), 71 deletions(-) create mode 100644 docs/agents/test-audit.md diff --git a/AGENTS.md b/AGENTS.md index 97b8c6630d..61672c50b9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -12,7 +12,7 @@ Load only the procedures relevant to the task: | Domain vocabulary | `CONTEXT.md`, `docs/agents/domain.md` | | Architecture decisions | `docs/adr/README.md` | | Tests or gate selection | `docs/agents/testing.md` | -| Writing, auditing, or sweeping tests; dispatching test-audit subagents | `skills/test-audit/SKILL.md` | +| Writing, auditing, or sweeping tests; dispatching test-audit subagents | `docs/agents/test-audit.md`, `skills/test-audit/SKILL.md` | | Selector capture, polling, or interaction fast paths | `docs/agents/selector-capture.md` | | Adding or changing a CLI flag | `docs/agents/cli-flags.md` | | Opening or reviewing a PR | `docs/agents/pull-requests.md` | diff --git a/docs/agents/test-audit.md b/docs/agents/test-audit.md new file mode 100644 index 0000000000..ebfcbb9421 --- /dev/null +++ b/docs/agents/test-audit.md @@ -0,0 +1,57 @@ +# Test Audit + +A test earns its cost by guarding behavior, credible regressions, or independent contracts. Optimize for confidence, not deletions. Gate selection stays in `docs/agents/testing.md`. + +## Authoring gate + +Before adding a test, answer all: + +1. Which observable behavior, invariant, or independent contract does it guard? +2. Which credible regression makes it fail? Name the production edit. +3. Why doesn't existing coverage catch it? Each contract has one owner at its strongest boundary. +4. Does it need a production seam no production caller needs? Move the test to the real boundary. + +A bug regression must fail on pre-fix code for the intended reason; one that never failed visibly proves the mock. + +## Junk patterns + +Authoring rejects; audits hunt: + +- assertion-free probes; a lone `typeof x === 'function'` above the sibling calling `x`; +- self-comparison (`f(x)` vs `f(x)`, `map.get(k)` vs itself, fixture vs re-spelled fixture); +- expectation drawn from the helper under test; a `switch` in the test body compared to its arms; +- byte-identical or renamed `test()` blocks; one golden replayed via a second entry; +- type pins posing as runtime assertions (`Equal = true`, `X ? true : never`, `satisfies` after erasure, `toHaveLength(n)` over literals); if erasing leaves nothing executable, use `void [fixtures]`; +- mocks asserting their own inputs; negative controls passing for the wrong reason; +- names promising more than the input exercises (an ordering claim with no ordering assertion); +- dead code called only by tests; tests preserving test-only exports. + +## Retention criteria + +Refuse to delete tests enforcing: registry, gate-manifest, layering, and DI-seam completeness; golden tables and cross-language source inspection; `@ts-expect-error` pins (an unused directive fails `tsc`); identity pins deep-equal can't subsume; exact-array facets a membership scan can't see duplicates in; vocabulary pins surviving a production+fixture co-update. Static-or-slow and similarity to the implementation are not deletion-worthy: prove you can't break it. + +## Candidate evidence + +Record before editing; a missing field blocks the edit: + +- location and test name; the detectable failure, or why none; +- non-test callers of the covered seams; the stronger remaining owner-boundary proof (both assertion sets, `file:line`); +- skip guards in candidate and sibling (`skipIf`, `.skip`, env flags, whether a workflow sets them); +- which runner executes the file: a Vitest project (`vitest.config.ts`), a `node --test` script (`test:smoke`, `check:*:test`), native XCTest, or a `CHECK_CATALOG` gate id. Only a file no runner reaches is a never-run finding; +- what deletion unlocks; risk. + +## Mutation discipline + +Plant the production edit representing the claimed bug: the pre-edit body must stay green while the post-edit body goes red. Report both counts (`8 passed` → `1 failed | 7 passed`). Revert and show `git status` clean before committing. For a deletion, prove the sibling catches the mutant. Never argue detection: measure it. + +## Focused validation + +`pnpm exec vitest run --project unit-core `, or `--project apple-runner` for `packages/platform-apple/src/runner/**`. Then `pnpm check:quick` and `pnpm check:affected --run` on the pushed commit; report production vs test LOC from `git diff --numstat origin/main...HEAD`. Most `scripts/**` tests are in no `tsconfig` include, so `pnpm typecheck` skips helpers there. + +## Dispatching subagents + +Run parallel read-only lanes with exclusive edit ownership: subsystem lanes, one cross-cutting junk-pattern sweep, one never-run inventory (Vitest includes, `node --test` globs, gate ids). Mandate in every prompt: READ-ONLY; read the production owner and name the bug; verify skip guards and runner/gate ownership; the retention list verbatim; false positives cost more than misses. + +## Landing + +One owner-boundary batch per PR; prefer net-negative production LOC. Report retained false positives, then rebase and rediscover. diff --git a/skills/test-audit/SKILL.md b/skills/test-audit/SKILL.md index 7f08963bf5..1b16a01b99 100644 --- a/skills/test-audit/SKILL.md +++ b/skills/test-audit/SKILL.md @@ -1,78 +1,12 @@ --- name: test-audit -description: Write, change, review, or sweep tests. Authoring gate for new tests, plus an audit workflow and parallel subagent dispatch plan for finding tests that cannot fail, duplicating proof, or test-only production seams. +description: Write, change, review, or sweep tests. Use when authoring new tests, auditing existing tests that cannot fail or duplicate proof, or dispatching parallel read-only test-audit subagents. --- # Test audit -Three modes, one value criterion: a test earns its cost by guarding behavior, credible regressions, or independently meaningful contracts. Optimize for confidence, not deletion count. +Read `docs/agents/test-audit.md` before adding or deleting a test. It holds the authoring gate, the junk-pattern catalog, the retention criteria that refuse a deletion, the candidate evidence fields, and the focused validation commands. -## Authoring gate +The one non-negotiable: never argue that a test already detects something or can be deleted — measure it. Plant the production edit representing the claimed bug, show the pre-edit body green and the post-edit body red with both counts, revert, and show `git status` clean. -Before adding a test, answer; if you can't, don't add it yet: - -1. Which observable behavior, invariant, or independent contract does it guard? -2. Which credible regression makes it fail? Name the production edit. -3. Why doesn't existing coverage catch it? Each contract has one primary owner at the strongest boundary. -4. Does it need a production seam no production caller needs? Move the test to the real boundary instead. - -Bug regressions must fail on pre-fix code for the intended reason. A regression test that never visibly failed is proving the mock. - -## Junk patterns - -Authoring rejects matches; audits hunt for existing matches: - -- assertion-free probes; lone `typeof x === 'function'` followed by a sibling that calls `x`; -- self-comparison (`f(x)` vs `f(x)`, `map.get(k)` vs itself, fixture vs re-spelled fixture); -- expected value from the helper under test (oracle-as-expectation); a `switch` written in the test body; -- byte-identical or renamed `test()` blocks; the same golden replayed through a second entry point; -- type pins posing as runtime assertions: `Equal = true`, `X ? true : never`, `satisfies` compared after erasure, `toHaveLength(n)` over test literals. Erase them mentally first; if nothing executable remains, convert to the `void [fixtures]` idiom; -- mocks asserting their own inputs; negative controls passing for the wrong reason; -- names promising more than the input exercises (an ordering claim with no ordering assertion is a repair candidate, not a deletion); -- dead code called only by tests; tests preserving test-only exports. - -## Retention criteria - -Keep (and refuse to delete) tests enforcing: command-registry/gate-manifest/layering/DI-seam completeness; golden tables and cross-language source inspection (`protocol.test.ts` grepping Obj-C is a real detector); `@ts-expect-error` pins (unused directives fail `tsc`, so they bite); deliberate identity pins where deep-equal doesn't subsume `toBe`; exact-array facets where a membership scan can't see duplicates; vocabulary pins that survive production+fixture co-updates. Static-or-slow isn't deletion-worthy; similarity to implementation isn't either — prove you can't break it. - -## Candidate evidence - -Record before editing; missing field ⇒ not ready: - -- location + test name; detectable failure (or why none exists); -- non-test callers of covered seams; stronger remaining owner-boundary proof (quote both assertion sets, file:line, show superset relation); -- skip guards in candidate and sibling (`skipIf`, `.skip`, `AGENT_DEVICE_*` env, whether any workflow or script sets it); -- which Vitest project includes the file (`unit-core`, `apple-runner`, `provider-integration`, `fuzz-worker`, `interaction-contract`, `output-economy`) — unmatched = never-run finding; note: `macos.yml` names its darwin files explicitly; -- unlock deletion; focused validation command; risk. - -## Mutation discipline - -The audit's core proof: plant a production edit that represents the bug the test claims to catch; the pre-edit body must stay green while the post-edit body goes red. Report both counts (`8 passed` → `1 failed | 7 passed`). Revert and show `git status` clean before committing. For deletions, prove the sibling catches the mutant. Never argue detection — measure it. - -## Validation - -This repo's commands (upstream `run-vitest.mjs`/`check-changed.mjs` don't exist): - -1. Focused: `npx vitest run --project unit-core ` (or `--project apple-runner` for `packages/platform-apple/src/runner/**`). -2. `pnpm format` (whole repo), `npx oxlint . --deny-warnings`, `npx tsc -b packages/...` for package edits, `npx tsc -p tsconfig.json` for `src/`. -3. `pnpm check:affected --run` at the pushed commit; `pnpm check:fallow --base origin/main` when touching production exports. -4. `git diff --numstat`; report production/test deltas separately. -5. Don't edit during Vitest runs; `scripts/**` tests are outside `tsc` scope — type-check new helpers there explicitly. - -## Dispatching subagents - -For broad scope, run parallel read-only lanes and keep editing ownership exclusive: - -- subsystem lanes (platform packages; `src/daemon`; contracts/kernel/registry/capture-kit/selectors; `test/integration` + maestro/ad-replay/session-journal + provider packages; `scripts/**`); -- one cross-cutting pattern sweep (junk-pattern greps across all tests, each hit read and confirmed); -- one never-run inventory (project includes, gate ownership via `scripts/gate` and `.github/workflows`, unsatisfiable guards). - -Mandate in every prompt: READ-ONLY; read the production owner and name the bug; verify skip guards and project selection; retention list above verbatim; "a false positive costs more than a miss"; output capped, highest-confidence first, with validation commands. - -## Landing - -One coherent owner-boundary batch per PR; prefer net-negative production LOC; delete the test-only seam with the test. Push only when authorized; rebase on review feedback and answer findings by fixing the rule, not the cited site. After landing, rebase to `main` and rediscover; don't carry stale candidate lists. - -## Handoff - -Report: categories deleted; owner simplifications; retained false positives with reasons; mutations run with before/after counts; production vs test LOC; PR state; named follow-ups needing owner judgment (never-run env gates, tsconfig coverage gaps, suspected product bugs). +Dispatch parallel lanes read-only with exclusive edit ownership, paste the retention list into every prompt, and treat a false positive as costing more than a miss. From 5559b2386039d4fc982ce97018ccfff0ec0e7d82 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Micha=C5=82=20Pierzcha=C5=82a?= Date: Fri, 25 Sep 2026 13:46:08 +0200 Subject: [PATCH 3/3] docs(agents): fold the test-audit skill card into its procedure doc --- AGENTS.md | 2 +- docs/agents/test-audit.md | 22 +++++++++++----------- skills/test-audit/SKILL.md | 12 ------------ 3 files changed, 12 insertions(+), 24 deletions(-) delete mode 100644 skills/test-audit/SKILL.md diff --git a/AGENTS.md b/AGENTS.md index 61672c50b9..00ac5f7f8d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -12,7 +12,7 @@ Load only the procedures relevant to the task: | Domain vocabulary | `CONTEXT.md`, `docs/agents/domain.md` | | Architecture decisions | `docs/adr/README.md` | | Tests or gate selection | `docs/agents/testing.md` | -| Writing, auditing, or sweeping tests; dispatching test-audit subagents | `docs/agents/test-audit.md`, `skills/test-audit/SKILL.md` | +| Writing, auditing, or sweeping tests; dispatching test-audit subagents | `docs/agents/test-audit.md` | | Selector capture, polling, or interaction fast paths | `docs/agents/selector-capture.md` | | Adding or changing a CLI flag | `docs/agents/cli-flags.md` | | Opening or reviewing a PR | `docs/agents/pull-requests.md` | diff --git a/docs/agents/test-audit.md b/docs/agents/test-audit.md index ebfcbb9421..c9e9af4c29 100644 --- a/docs/agents/test-audit.md +++ b/docs/agents/test-audit.md @@ -4,7 +4,7 @@ A test earns its cost by guarding behavior, credible regressions, or independent ## Authoring gate -Before adding a test, answer all: +Answer all before adding a test: 1. Which observable behavior, invariant, or independent contract does it guard? 2. Which credible regression makes it fail? Name the production edit. @@ -20,24 +20,24 @@ Authoring rejects; audits hunt: - assertion-free probes; a lone `typeof x === 'function'` above the sibling calling `x`; - self-comparison (`f(x)` vs `f(x)`, `map.get(k)` vs itself, fixture vs re-spelled fixture); - expectation drawn from the helper under test; a `switch` in the test body compared to its arms; -- byte-identical or renamed `test()` blocks; one golden replayed via a second entry; +- near-identical or renamed `test()` blocks; one golden replayed twice; - type pins posing as runtime assertions (`Equal = true`, `X ? true : never`, `satisfies` after erasure, `toHaveLength(n)` over literals); if erasing leaves nothing executable, use `void [fixtures]`; - mocks asserting their own inputs; negative controls passing for the wrong reason; -- names promising more than the input exercises (an ordering claim with no ordering assertion); +- names promising more than the input exercises; - dead code called only by tests; tests preserving test-only exports. ## Retention criteria -Refuse to delete tests enforcing: registry, gate-manifest, layering, and DI-seam completeness; golden tables and cross-language source inspection; `@ts-expect-error` pins (an unused directive fails `tsc`); identity pins deep-equal can't subsume; exact-array facets a membership scan can't see duplicates in; vocabulary pins surviving a production+fixture co-update. Static-or-slow and similarity to the implementation are not deletion-worthy: prove you can't break it. +Refuse to delete tests enforcing: registry, gate-manifest, layering, and DI-seam completeness; golden tables and cross-language source inspection; `@ts-expect-error` pins (an unused directive fails `tsc`); identity pins deep-equal can't subsume; exact-array facets a membership scan can't see duplicates in; vocabulary pins surviving a production+fixture co-update. Static-or-slow and implementation-similarity are not deletion-worthy: prove you can't break it. ## Candidate evidence Record before editing; a missing field blocks the edit: -- location and test name; the detectable failure, or why none; -- non-test callers of the covered seams; the stronger remaining owner-boundary proof (both assertion sets, `file:line`); -- skip guards in candidate and sibling (`skipIf`, `.skip`, env flags, whether a workflow sets them); -- which runner executes the file: a Vitest project (`vitest.config.ts`), a `node --test` script (`test:smoke`, `check:*:test`), native XCTest, or a `CHECK_CATALOG` gate id. Only a file no runner reaches is a never-run finding; +- location, test name, detectable failure (or why none); +- non-test callers of the covered seams; the stronger remaining owner proof (both assertion sets, `file:line`); +- skip guards in candidate and sibling (`skipIf`, `.skip`, env flags a workflow sets); +- the runner that executes the file: a Vitest project (`vitest.config.ts`), a `node --test` script (`test:smoke`, `check:*:test`), native XCTest, or a `CHECK_CATALOG` gate id. Only a file no runner reaches is a never-run finding; - what deletion unlocks; risk. ## Mutation discipline @@ -46,12 +46,12 @@ Plant the production edit representing the claimed bug: the pre-edit body must s ## Focused validation -`pnpm exec vitest run --project unit-core `, or `--project apple-runner` for `packages/platform-apple/src/runner/**`. Then `pnpm check:quick` and `pnpm check:affected --run` on the pushed commit; report production vs test LOC from `git diff --numstat origin/main...HEAD`. Most `scripts/**` tests are in no `tsconfig` include, so `pnpm typecheck` skips helpers there. +`pnpm exec vitest run --project unit-core `, or `--project apple-runner` for `packages/platform-apple/src/runner/**`; then `pnpm check:quick` and `pnpm check:affected --run` on the pushed commit. `scripts/**` tests are in no `tsconfig` include, so `pnpm typecheck` skips helpers there. ## Dispatching subagents -Run parallel read-only lanes with exclusive edit ownership: subsystem lanes, one cross-cutting junk-pattern sweep, one never-run inventory (Vitest includes, `node --test` globs, gate ids). Mandate in every prompt: READ-ONLY; read the production owner and name the bug; verify skip guards and runner/gate ownership; the retention list verbatim; false positives cost more than misses. +Run parallel read-only lanes with exclusive edit ownership: subsystem lanes (per platform package, `src/daemon`, core packages, `scripts/**`), one cross-cutting junk-pattern sweep, one never-run inventory (Vitest includes, `node --test` globs, gate ids). Mandate in every prompt: READ-ONLY; read the production owner and name the bug; verify skip guards and runner/gate ownership; paste the retention list; false positives cost more than misses. ## Landing -One owner-boundary batch per PR; prefer net-negative production LOC. Report retained false positives, then rebase and rediscover. +One owner-boundary batch per PR; delete the test-only seam with the test; prefer net-negative production LOC. Report retained false positives, then rebase to `main` and rediscover. diff --git a/skills/test-audit/SKILL.md b/skills/test-audit/SKILL.md deleted file mode 100644 index 1b16a01b99..0000000000 --- a/skills/test-audit/SKILL.md +++ /dev/null @@ -1,12 +0,0 @@ ---- -name: test-audit -description: Write, change, review, or sweep tests. Use when authoring new tests, auditing existing tests that cannot fail or duplicate proof, or dispatching parallel read-only test-audit subagents. ---- - -# Test audit - -Read `docs/agents/test-audit.md` before adding or deleting a test. It holds the authoring gate, the junk-pattern catalog, the retention criteria that refuse a deletion, the candidate evidence fields, and the focused validation commands. - -The one non-negotiable: never argue that a test already detects something or can be deleted — measure it. Plant the production edit representing the claimed bug, show the pre-edit body green and the post-edit body red with both counts, revert, and show `git status` clean. - -Dispatch parallel lanes read-only with exclusive edit ownership, paste the retention list into every prompt, and treat a false positive as costing more than a miss.