From d61c128e26988f573ead510d1d1cb4930d56c1df Mon Sep 17 00:00:00 2001 From: 3metaJun <251347867+3metaJun@users.noreply.github.com> Date: Fri, 11 Sep 2026 11:44:51 +0800 Subject: [PATCH 1/3] fix: restore portable workflow guidance --- skills/meta-mode/SKILL.md | 213 ++++++++++++++++++++++++-------------- skills/reflect/SKILL.md | 4 +- skills/swarm/SKILL.md | 4 +- skills/why/SKILL.md | 4 +- 4 files changed, 144 insertions(+), 81 deletions(-) diff --git a/skills/meta-mode/SKILL.md b/skills/meta-mode/SKILL.md index e4df274..beade6a 100644 --- a/skills/meta-mode/SKILL.md +++ b/skills/meta-mode/SKILL.md @@ -5,79 +5,140 @@ description: "Route a non-trivial engineering task through a verifiable mstack w # Meta mode -Use this skill for a task that needs more than one focused action. It is the -portable replacement for pstack's `poteto-mode`. It keeps the workflow intact -while leaving execution details to the current harness. - -## Apply the mode - -1. State the requested result in one sentence. -2. Choose the smallest matching playbook from `playbooks/`. -3. Read that playbook and every principle it names before you act. -4. Write a short todo list whose first entries are the playbook steps. -5. Choose a local or configured remote execution environment. Use a remote - environment only when the task needs another machine. -6. Choose a model for each role from the user's mstack model configuration. - Use `inherit-parent` when no override exists. -7. End each step with evidence from the real artifact. - -The mode stays active for the current task. Do not apply it to a casual turn or -after the user opts out. - -## Choose a playbook - -| Request | Playbook | -| --- | --- | -| Understand a subsystem | `investigation.md` | -| Fix a defect | `bug-fix.md` | -| Improve measured performance | `perf-issue.md` or `hillclimb.md` | -| Add behavior | `feature.md` | -| Reshape code without changing behavior | `refactoring.md` | -| Compare a design or prototype | `prototype.md` or `visual-parity.md` | -| Write or change a skill | `authoring-a-skill.md` | -| Review or evaluate agent behavior | `eval.md` or `interrogate` | -| Drive a pull request | `babysit.md`, `shipping.md`, or `opening-a-pr.md` | -| Run work while the user is away | `autonomous-run.md` or `orchestrate.md` | -| Resume or suspend work | `session-pickup.md` or `pause-safely.md` | -| Plan several phases | `multi-phase-plan.md` | - -Use `figure-it-out` when no listed playbook fits. Use `swarm` for independent -coverage and `arena` for competing proposals. Keep each worker in its own -worktree or output directory when it writes files. - -## Map capabilities - -Read [the capability map](references/capability-matrix.md) when a playbook -mentions delegation, background work, remote execution, model roles, or session -history. The map distinguishes a native implementation from a documented -fallback. Report the distinction when it changes the result. - -The canonical skill tree uses capability names. Adapters map those names to a -Harness. Do not put vendor-specific paths, commands, transcript formats, or -model names in this skill. - -## Resolve installed paths - -Playbooks use `` for the directory that contains the installed -mstack skills. They use `` for the optional `meta-mode-tools` -artifact directory. Resolve both paths from the active `meta-mode` skill before -you run a command. In this repository, the paths are `/skills` and -`/tools/meta-mode`. The default installer places them at -`/skills` and `/tools/meta-mode`, so the tools are -at `../../tools/meta-mode` from the `meta-mode` skill directory. If the install -uses an artifact destination override, use the destination that the installer -reports. - -## Configure models - -Copy `profiles/models.example.json` to `~/.config/mstack/models.json` and replace -`inherit-parent` with models available in that Harness. Run -`node scripts/model-config.mjs --file ~/.config/mstack/models.json` to validate -the file. Keep the role names stable so a workflow can move between Harnesses. - -## Finish with proof - -Run the narrowest meaningful verification and read its output. For code, run -the relevant test and inspect the diff. For an installation, inspect the target -directory and adapted frontmatter. For a remote run, record the environment and -the command that produced the evidence. +## Non-negotiables + +The Principles section below grounds every trigger. In your reply, name each principle that shaped a decision and the specific choice it changed. Cite only principles whose leaf SKILL.md you read this session. + +Remaining triggers: + +- Nontrivial change, architecture decision, or "are we sure?" → the **how** skill. +- About to ask a human to choose between approaches → classify it before asking. If the answer is a fact you could observe by running something (behavior, timing, layout, output, perf, even whether an eval separates), it is not the human's to answer. Sketch it via the Prototype playbook (`playbooks/prototype.md`) and let the result decide. If the task is a read-only Investigation whose deliverable is a cited answer, stay in it and answer from the evidence rather than building a sketch. Reserve the question for a genuine product or preference call no experiment can settle. +- Any code → name the data shape first, and choose its organizing structure per **principle-model-the-domain**. +- Code crossing a function boundary → the **architect** skill, parallel design exploration before implementing. +- Parallel fan-out → the **swarm** skill for coverage matrices, races, gauntlets, and exploration partitions. Use **arena** for design or code bakeoffs with base selection and grafting. +- Contested design → the **interrogate** skill (multi-model adversarial) before shipping. +- Nontrivial multi-step → write the throughput checkpoint (Feature step 3). +- Any prose surface → the **unslop** skill. Your reply is a prose surface. Write it per **Writing the reply**. Agent-facing prose also follows the **create-skill** skill (the active Harness skill-authoring flow). +- Docs, RFCs, readmes, PR descriptions, or commit messages → the **technical-writing** skill (`/technical-writing`). +- Before commit → the active Harness cleanup command. +- Before review → the **no-comments** skill (`/no-comments`). +- Shipping UI / IDE / CLI → use the matching control skill exposed by the active Harness. For bug fixes, reproduce first on the same surface yourself. Hand to the user only under the narrow Bug fix step 1 exception. +- Any PR-status request → the **Babysit** playbook (`playbooks/babysit.md`). That includes "babysit this", "get it green", "address the bugbot comments", and the commonest phrasing, "check on PR X" / "anything outstanding on X". Never triggered by merely opening a PR. Declare its mode before polling. The playbook's step 1 owns the request-to-mode mapping. Reaching for `drive` inside a phase agent stops that agent finishing its turn. +- Asked to land or ship a green stack → the **Shipping** playbook (`playbooks/shipping.md`). Green is not safe. Nothing gets armed before an independent per-PR verdict, and only the contiguous verified run from the root lands. +- Bugbot or the agentic security review commented → skeptical posture. They catch real bugs and also file non-issues and nitpicks, so assess each on its merits and dismiss noise with a concrete reason instead of churning code. Triage fix / dismiss / ask per `references/bugbot-triage.md`. +- Broken skill mid-task → fix it in its own PR. Don't block. Don't silently work around it. +- Long, autonomous, or multi-phase work, or any task the user steps away from to review later ("going to bed", "trust it when i'm back", "/loop until X") → a decision trail via the **show-me-your-work** skill. Commit it when stakes need an auditable record. Keep it local otherwise. + +## Principles + +Read the leaf skill in full for any principle you apply. Each entry names when it applies. + +**Core** + +- **Laziness Protocol** (**principle-laziness-protocol**). Refactoring, sizing a diff, or tempted to add abstractions, layers, or signal threading. Bias to deletion and the smallest change that solves the problem. +- **Foundational Thinking** (**principle-foundational-thinking**). Before writing logic: core types and data structures, scaffold-vs-feature sequencing, what concurrent actors share. +- **Redesign from First Principles** (**principle-redesign-from-first-principles**). Integrating a new requirement into an existing design. Redesign as if it had been foundational from day one. +- **Attack the Premise** (**principle-attack-the-premise**). Two or more fixes that share one premise have failed the same gate. Take a census of which actors hold the imbalance before the next fix, then question the premise instead of writing another fix that assumes it. +- **Subtract Before You Add** (**principle-subtract-before-you-add**). Sequencing an addition, refactor, or rewrite. Remove dead weight first, then build on the simpler base. +- **Minimize Reader Load** (**principle-minimize-reader-load**). Reviewing or shaping code that's hard to trace. Count layers and hidden state, collapse one-caller wrappers, shrink mutable scope. +- **Outcome-Oriented Execution** (**principle-outcome-oriented-execution**). Planned rewrites and migrations with explicit phase boundaries. Converge on the target architecture, don't preserve throwaway compatibility states. +- **Experience First** (**principle-experience-first**). Product, UX, or feature-scope tradeoffs. Choose user delight over implementation convenience. +- **Exhaust the Design Space** (**principle-exhaust-the-design-space**). A novel interaction or architectural decision with no precedent. Build 2-3 competing prototypes and compare before committing. +- **Build the Lever** (**principle-build-the-lever**). Any non-trivial work. Build the tool that does or proves it (codemod, script, generator), not by hand. The tool is the artifact a reviewer reruns. + +**Architecture** + +- **Model the Domain** (**principle-model-the-domain**). Writing stateful logic, or code that branches a lot or repeats a shape assumption across files. Encode the domain in a structure (state machine, typed model, table or registry, reducer, boundary, the right collection) instead of scattered conditionals. +- **Boundary Discipline** (**principle-boundary-discipline**). Wiring validation, error handling, or framework adapters. Guards at system boundaries, trust internal types, keep business logic pure. +- **Type System Discipline** (**principle-type-system-discipline**). Designing types or a signature in any typed language. Make illegal states unrepresentable, brand primitives, parse external data at boundaries. +- **Make Operations Idempotent** (**principle-make-operations-idempotent**). Designing commands, lifecycle steps, or loops that run amid crashes and retries. Converge to the same end state. +- **Migrate Callers Then Delete Legacy APIs** (**principle-migrate-callers-then-delete-legacy-apis**). Introducing a new internal API while old callers exist. Migrate and delete in one wave. +- **Separate Before Serializing Shared State** (**principle-separate-before-serializing-shared-state**). Concurrent actors might write the same file, branch, key, or object. Eliminate the sharing first. + +**Verification** + +- **Prove It Works** (**principle-prove-it-works**). After a task, before declaring done. Verify against the real artifact, not a proxy or "it compiles". +- **Fix Root Causes** (**principle-fix-root-causes**). Debugging. Trace each symptom to its root cause, reproduce first, ask why until you reach it. +- **Sequence Work into Verifiable Units** (**principle-sequence-verifiable-units**). Multi-step work (sweeps, migrations, runs of similar edits) and how you stack commits and PRs. Break work into small units that each end in a check, verify each before the next, and order delivery so the sequence proves itself. +- **Test Behavior, Not Implementation** (**principle-test-behavior-not-implementation**). Writing, changing, or keeping a test. Call the code the way its users do and assert the result against a literal expected value. If the test would still pass when every imported function returns `undefined`, rewrite the assertion or delete the test. + +**Delegation** + +- **Guard the Context Window** (**principle-guard-the-context-window**). Context fills up: large outputs, long files, repeated reads, fan-out planning. Route bulk to subagents, keep summaries in the main thread. +- **Never Block on the Human** (**principle-never-block-on-the-human**). Tempted to ask "should I do X?" on reversible work. Proceed, present the result, let the human course-correct. + +**Meta** + +- **Encode Lessons in Structure** (**principle-encode-lessons-in-structure**). You catch yourself writing the same instruction a second time. Encode it as a lint, metadata flag, runtime check, or script instead of more text. + +## Autonomy + +**Just do it.** Use any MCP tool. Reversible work and external actions (team chat, ticket updates, kicking off evals) proceed without asking. + +**Always pause** for irreversible writes: force-push to shared branches, deploys, data deletion, customer messages. + +**Session overrides:** "Don't stop" / "going to bed" / "run until done" / "be fully autonomous" → keep going. + +**No is an acceptable answer.** Asked whether to do something, invited to add scope, or shown an approach, reply with your real judgment. Decline, push back, or say "this doesn't earn its place" when true. A recommendation is a judgment, not a validation. Agreement is not the default, candor over sycophancy. + +## Subagents + +Use the active Harness's documented delegation mechanism for workers spawned by a +playbook. Routed skills such as `how`, `why`, `interrogate`, `reflect`, and +`swarm` define their own worker roles and review rules. Keep worker prompts and +results in the format that the active Harness supports. + +Run independent workers concurrently when the Harness supports it. If it does +not, run them one at a time and record the fallback. Select each worker's model +from `/setup-mstack`; use `inherit-parent` when no override exists. Give each +worker a focused file or question and read its complete result before deciding. + +You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. A second opinion is the same prompt against a different model. Agreement is high-signal. + +## Writing the reply + +Write the reply clean as you draft it. A cleanup pass after drafting does not remove these patterns. + +- **Short declarative sentences.** One thought per sentence, ended with a period. +- **No long-dash character anywhere.** Write a file-list bullet as a sentence ("`main.js` owns persistence and the IPC handlers") and a bold section header as its own sentence ("**Verification.** End to end via CDP"). +- **A colon as a mid-sentence connector is also out** (unslop rule 14). A colon before a list is fine. +- **Terse is not an excuse to drop content.** Short sentences, but every section the playbook's reply names stays: details, tradeoffs, choices, open decisions. +- **Frame impact for the consumer and the maintainer.** Name who the work is for (an end user, a colleague importing the library) and what changes for them before any implementation detail. Then what the next engineer who owns this code inherits. If you can't say what either would notice, the work or the explanation is off. +- **Never fabricate a link, citation, or transcript reference.** Link only artifacts you produced or read this session. +- **Every claim carries its evidence or its label in the same sentence.** Measured, inferred, or guess. A prediction or an unseen cause is a guess. Never hand the human a check you could run. + +Every playbook ends with a reply written this way, PR link as `https://github.com///pull/`. The per-playbook lines below name only the content unique to that playbook. + +## Comments + +Comments follow the same rule as the reply. Write them clean as you go. Keep a comment only for a non-obvious *why* the code can't show. A verify or test script gets no phase-narrating comments such as `// Phase 1: add cards`. The assertion or log string documents the step, as in `assert(ok, 'persisted across restart')`. This applies to every file you produce, including the delegate's diff. + +## Playbooks + +Open a todolist whose first items are the matched playbook's steps, copied in verbatim, before any task-specific todos. A step you choose not to do stays in the list with a one-line `skip: `. Match the task to a playbook below, open its file, and copy its steps in verbatim. + +A large or cross-cutting effort (a migration across many call sites, an ambitious multi-part change), or work the user steps away from to trust later, routes to the **figure-it-out** skill even when a narrower playbook like Feature fits. Use **figure-it-out** whenever no bundled playbook fits. It designs a bespoke, rigorous playbook for the task. A standing project-scale program (multi-day, many stacked PRs, a fleet of subagents under one coordinator) routes to **Orchestrate** instead. figure-it-out designs one bespoke run, orchestrate runs the program. + +- **Investigation.** Read-only question: how does X work, why was Y built this way, are we sure about Z, should we do X or Y. `playbooks/investigation.md`. +- **Bug fix.** A reported defect to reproduce, root-cause, and fix with runtime evidence. `playbooks/bug-fix.md`. +- **Perf issue.** A measured slowness to trace and improve against a baseline. `playbooks/perf-issue.md`. +- **Hillclimb.** Sustained, scientific improvement of one metric against a target: loop hypotheses with before/after measurement, a decision log, and one commit per accepted win. Distinct from Perf issue, which is a one-off fix. `playbooks/hillclimb.md`. +- **Runtime forensics.** Diagnose a runtime symptom (leak, idle-CPU spin, glitch) from live instrumentation. The deliverable is a diagnosis, not a fix. `playbooks/runtime-forensics.md`. +- **Trace forensics.** Diagnose a captured profiling artifact (cpuprofile, trace, spindump, heap snapshot) handed to you after the fact. The deliverable is a diagnosis, not a fix. `playbooks/trace-forensics.md`. +- **Feature.** New or changed behavior, built from a named data shape. `playbooks/feature.md`. +- **Refactoring.** A behavior-preserving change to structure or shape (rename, extract, inline, dedupe, move). `playbooks/refactoring.md`. +- **Prototype.** A throwaway sketch to make a design or behavioral decision cheaply, or to settle an empirical fork by observing it instead of asking the human ("prototype", "mock it up", "try this layout", "sketch it to decide"). `playbooks/prototype.md`. +- **Visual parity.** Pixel-exact UI equivalence: matching two implementations or migrating a styling system. `playbooks/visual-parity.md`. +- **Authoring or modifying a skill.** Writing or editing a SKILL.md. `playbooks/authoring-a-skill.md`. +- **Eval.** Testing how a skill, structure, or prompt change affects agent behavior before promoting it. `playbooks/eval.md`. +- **Babysit.** Driving a PR or a stack to merge-ready: conflicts, review threads, CI. `playbooks/babysit.md`. +- **Shipping.** The half after Babysit. Independently verifying a green stack, then landing the contiguous verified run bottom-up through `gh` by default or Origin when its CLI is available. `playbooks/shipping.md`. +- **Autonomous run.** A long task to drive to completion without stopping ("run until done", "/loop until X"). `playbooks/autonomous-run.md`. +- **Orchestrate.** A standing project handed to one coordinator chat: multi-day, many stacked PRs, dozens to hundreds of subagents, minimal human turns ("run this whole project", "own this migration until it lands"). Distinct from Autonomous run, which drives one task to a predicate. Work one agent could finish inside the session's budget routes there, not here, however program-shaped the phrasing sounds. `playbooks/orchestrate.md`. +- **Autopilot-full.** A queue of independent PRs run to merged with full autonomy. One owner per PR carries build through merge, and the root swarm-verifies each merge-ready head before its owner merges ("autopilot this queue", "full autopilot", one-owner-per-PR programs). `playbooks/autopilot-full.md`. +- **Autopilot-stack.** A queue of changes built and verified with full autonomy, delivered as one linear reviewed base-branch stack the operator lands herself ("autopilot-stack", "stack them, don't ship", "build the stack, I'll land it"). `playbooks/autopilot-stack.md`. +- **Session pickup.** Resuming or taking over a prior agent's in-flight work from a transcript, cloud-agent URL, or pushed branch. `playbooks/session-pickup.md`. +- **Pause safely.** Suspending in-flight work cleanly so it can be resumed, on an explicit pause, going offline, a Harness restart, or imminent context compaction. The complement to Session pickup. Full steps: `playbooks/pause-safely.md`. +- **Multi-phase or multi-PR plan.** Work that spans phases or stacked PRs. `playbooks/multi-phase-plan.md`. +- **Worktree and simulator cleanup.** Reclaiming local disk by pruning merged or abandoned git worktrees and stale iOS simulators ("what's using my disk", "clean up worktrees", "prune safe-to-prune worktrees", "free up space", "delete old simulators"). `playbooks/worktree-cleanup.md`. +- **Opening a PR.** Invoked at the end of every other playbook. `playbooks/opening-a-pr.md`. diff --git a/skills/reflect/SKILL.md b/skills/reflect/SKILL.md index 28cb0ca..0c4e6ff 100644 --- a/skills/reflect/SKILL.md +++ b/skills/reflect/SKILL.md @@ -27,7 +27,7 @@ For each candidate, read the first JSONL line and check that `message.content[0] ### 2. Spawn three reviewers in parallel -One message, three `Task` calls, `worker type: generalPurpose`, explicit `model:` on each, agent mode (`readonly: false`). Reviewers need MCP access for context lookups (tickets, chat threads, observability traces referenced in the transcript). Readonly strips MCPs. +One delegation request with three reviewers, an explicit model on each, and agent mode with the access needed for context lookups (tickets, chat threads, and observability traces referenced in the transcript). Reviewers need the active Harness's connected tools. | Lens | `model` | Prompt template | |---|---|---| @@ -39,7 +39,7 @@ Pass each template verbatim, substituting the transcript path or digest where ma ### 3. Synthesize -One `Task` call, `worker type: generalPurpose`, using the active mstack `synthesizer` role (default `inherit-parent`), agent mode (`readonly: false`). The synthesizer's quality check includes spot-verifying citations, which can require MCP access. Readonly strips MCPs. Use `references/synthesizer.md` verbatim, with each reviewer's full output inlined where marked. The synthesizer returns a structured Accepted / Rejected / Backlog list. +One delegation request using the active mstack `synthesizer` role (default `inherit-parent`) and the access needed for citation checks. Use `references/synthesizer.md` verbatim, with each reviewer's full output inlined where marked. The synthesizer returns a structured Accepted / Rejected / Backlog list. ### 4. Structural enforcement check diff --git a/skills/swarm/SKILL.md b/skills/swarm/SKILL.md index a9b4803..ad058fd 100644 --- a/skills/swarm/SKILL.md +++ b/skills/swarm/SKILL.md @@ -5,7 +5,7 @@ description: "Fan out N parallel workers, drain them, and return one report. Use # Swarm -Fan out N parallel cloud workers. They may cover separate slices, race the same brief, or mix both. The parent waits, aggregates, and returns one report. +Fan out N parallel workers. They may cover separate slices, race the same brief, or mix both. The parent waits, aggregates, and returns one report. ## Start @@ -26,7 +26,7 @@ Open a todolist with one entry per phase before launching anything. ## Phase B: Fan out -Spawn all N workers in one message with `worker type: generalPurpose`, `environment: "cloud"`, `background execution: true`, and the configured model. Use `environment: "local"` only when the worker needs access to something on the user's computer. +Spawn all N workers through the active Harness's delegation API, using the configured model and background execution when supported. Use a local worker when the task needs access to something on the user's computer. When a worker must start from a non-default pushed branch, pass `cloud_base_branch`. diff --git a/skills/why/SKILL.md b/skills/why/SKILL.md index 6ba1216..b50c7c7 100644 --- a/skills/why/SKILL.md +++ b/skills/why/SKILL.md @@ -140,7 +140,9 @@ Take the synthesizer's output and present it to the user. You may lightly edit f The output structure is the one in `references/synthesizer-prompt.md`: The Question, The Code in Question, What We Found, What We Can Reasonably Infer, Competing Hypotheses, What We Don't Know, Sources Consulted, Confidence Summary. Adapt as needed, but keep the confidence separation intact, and keep Sources Consulted as one line per investigator, including the ones that returned nothing or were skipped, with the reason. -After the Sources Consulted block, if the user's `why` question is a prethe current harness to actually changing this code, convert the lineage findings into a Preserve / Change / Avoid / Risk constraint set suitable for planning the change. +After the Sources Consulted block, if the user's `why` question is a prelude to +actually changing this code, convert the lineage findings into a Preserve / +Change / Avoid / Risk constraint set suitable for planning the change. ## Common Failure Modes to Avoid From ef8a6d8899d05077671fd01ac4b661351130779e Mon Sep 17 00:00:00 2001 From: 3metaJun <251347867+3metaJun@users.noreply.github.com> Date: Fri, 11 Sep 2026 11:45:41 +0800 Subject: [PATCH 2/3] test: guard portable skill integrity --- package.json | 2 +- scripts/skill-integrity.test.mjs | 52 ++++++++++++++++++++++++++++++++ 2 files changed, 53 insertions(+), 1 deletion(-) create mode 100644 scripts/skill-integrity.test.mjs diff --git a/package.json b/package.json index 9e8d8c9..26fa7ad 100644 --- a/package.json +++ b/package.json @@ -23,7 +23,7 @@ "install-skills": "node scripts/install.mjs", "optimize-context": "node scripts/optimize-context.mjs", "reconcile-context": "node scripts/reconcile-context.mjs", - "test": "node scripts/validate.mjs && node --test scripts/install.test.mjs scripts/context.test.mjs scripts/audit-context.test.mjs scripts/worktree-audit.test.mjs scripts/sync-upstream.test.mjs scripts/runtime.test.mjs scripts/environment.test.mjs" + "test": "node scripts/validate.mjs && node --test scripts/install.test.mjs scripts/context.test.mjs scripts/audit-context.test.mjs scripts/worktree-audit.test.mjs scripts/sync-upstream.test.mjs scripts/runtime.test.mjs scripts/environment.test.mjs scripts/skill-integrity.test.mjs" }, "bin": { "mstack": "scripts/install.mjs" diff --git a/scripts/skill-integrity.test.mjs b/scripts/skill-integrity.test.mjs new file mode 100644 index 0000000..802ce71 --- /dev/null +++ b/scripts/skill-integrity.test.mjs @@ -0,0 +1,52 @@ +import assert from "node:assert/strict"; +import { existsSync, readFileSync, readdirSync } from "node:fs"; +import { dirname, join, resolve } from "node:path"; +import { fileURLToPath } from "node:url"; +import test from "node:test"; + +const repoRoot = resolve(dirname(fileURLToPath(import.meta.url)), ".."); +const skillsRoot = join(repoRoot, "skills"); + +function readSkill(name) { + return readFileSync(join(skillsRoot, name, "SKILL.md"), "utf8"); +} + +test("meta-mode retains the workflow contract after portability adaptation", () => { + const content = readSkill("meta-mode"); + for (const heading of [ + "## Non-negotiables", + "## Principles", + "## Autonomy", + "## Subagents", + "## Writing the reply", + "## Comments", + "## Playbooks", + ]) { + assert.match(content, new RegExp(`^${heading}$`, "m")); + } + assert.doesNotMatch(content, /(?:poteto|Cursor|AskQuestion|subagent_type|run_in_background|disable-model-invocation|prethe current)/i); +}); + +test("meta-agent points to an existing complete entry skill", () => { + const agent = readFileSync(join(repoRoot, "agents", "meta-agent.md"), "utf8"); + assert.match(agent, /`meta-mode` skill's `SKILL\.md`/); + assert.match(readSkill("meta-mode"), /^## Principles$/m); +}); + +test("portable skill files contain no broken replacement artifacts", () => { + const forbidden = [ + /prethe current harness/i, + /(?:poteto-mode|setup-pstack|poteto-agent)/i, + /worker type:\s*generalPurpose/i, + ]; + const files = readdirSync(skillsRoot, { withFileTypes: true }) + .filter((entry) => entry.isDirectory()) + .map((entry) => join(skillsRoot, entry.name, "SKILL.md")) + .filter((path) => existsSync(path)); + for (const path of files) { + const content = readFileSync(path, "utf8"); + for (const pattern of forbidden) { + assert.doesNotMatch(content, pattern, `${path} contains ${pattern}`); + } + } +}); From b36ef68b5a1c7d04667f25815946f03e7c23a060 Mon Sep 17 00:00:00 2001 From: 3metaJun <251347867+3metaJun@users.noreply.github.com> Date: Fri, 11 Sep 2026 12:42:20 +0800 Subject: [PATCH 3/3] fix: harden portable skill review checks --- scripts/check-upstream.mjs | 7 ++-- scripts/skill-integrity.test.mjs | 26 ++++++++++-- scripts/sync-upstream.test.mjs | 6 +++ skills/how/SKILL.md | 6 +-- skills/interrogate/SKILL.md | 2 +- skills/meta-mode/SKILL.md | 40 +++++++++++++++++++ .../meta-mode/playbooks/worktree-cleanup.md | 2 +- .../meta-mode/references/capability-matrix.md | 8 ++-- skills/reflect/SKILL.md | 2 +- skills/swarm/SKILL.md | 2 +- skills/why/SKILL.md | 4 +- 11 files changed, 85 insertions(+), 20 deletions(-) diff --git a/scripts/check-upstream.mjs b/scripts/check-upstream.mjs index 0414be5..34e2795 100644 --- a/scripts/check-upstream.mjs +++ b/scripts/check-upstream.mjs @@ -28,6 +28,7 @@ if (args.includes("--help")) { } const targetRoot = resolve(valueAfter("--target") ?? repoRoot); +const strict = args.includes("--strict"); const syncLock = acquireUpstreamSyncLock(targetRoot, { recoverStale: false }); const releaseReadLockAtExit = () => { try { @@ -46,8 +47,9 @@ if (!pstack || typeof pstack !== "object") throw new Error("profiles/upstreams.j const source = valueAfter("--source") ?? process.env.MSTACK_PSTACK_SOURCE; if (!source) { - console.log("Pass --source or set MSTACK_PSTACK_SOURCE to check the upstream inventory."); - process.exit(0); + const message = "Pass --source or set MSTACK_PSTACK_SOURCE to check the upstream inventory."; + (strict ? console.error : console.log)(message); + process.exit(strict ? 1 : 0); } const sourceRoot = resolve(source); @@ -102,7 +104,6 @@ function resolveInside(root, path, label) { const sourceSkills = resolveInside(sourceRoot, "skills", "upstream skills directory"); if (!existsSync(sourceSkills)) throw new Error(`Upstream skills directory not found: ${sourceSkills}`); -const strict = args.includes("--strict"); function gitOutput(arguments_) { try { diff --git a/scripts/skill-integrity.test.mjs b/scripts/skill-integrity.test.mjs index 802ce71..b786cc3 100644 --- a/scripts/skill-integrity.test.mjs +++ b/scripts/skill-integrity.test.mjs @@ -11,6 +11,14 @@ function readSkill(name) { return readFileSync(join(skillsRoot, name, "SKILL.md"), "utf8"); } +function walkMarkdownFiles(directory) { + return readdirSync(directory, { withFileTypes: true }).flatMap((entry) => { + const path = join(directory, entry.name); + if (entry.isDirectory()) return walkMarkdownFiles(path); + return entry.isFile() && path.endsWith(".md") ? [path] : []; + }); +} + test("meta-mode retains the workflow contract after portability adaptation", () => { const content = readSkill("meta-mode"); for (const heading of [ @@ -38,15 +46,25 @@ test("portable skill files contain no broken replacement artifacts", () => { /prethe current harness/i, /(?:poteto-mode|setup-pstack|poteto-agent)/i, /worker type:\s*generalPurpose/i, + /cursor/i, + /application support\/the current harness/i, ]; - const files = readdirSync(skillsRoot, { withFileTypes: true }) + const skillDirectories = readdirSync(skillsRoot, { withFileTypes: true }) .filter((entry) => entry.isDirectory()) - .map((entry) => join(skillsRoot, entry.name, "SKILL.md")) - .filter((path) => existsSync(path)); + .map((entry) => join(skillsRoot, entry.name)); + const skillFiles = skillDirectories.map((directory) => join(directory, "SKILL.md")); + assert.equal( + skillFiles.filter((path) => existsSync(path)).length, + skillDirectories.length, + "every skill directory must contain SKILL.md", + ); + const files = skillDirectories.flatMap(walkMarkdownFiles); + assert.ok(files.length >= skillFiles.length, "integrity scan must include every skill markdown file"); for (const path of files) { const content = readFileSync(path, "utf8"); + const normalized = content.replaceAll("`", "").replace(/\s+/g, " "); for (const pattern of forbidden) { - assert.doesNotMatch(content, pattern, `${path} contains ${pattern}`); + assert.doesNotMatch(normalized, pattern, `${path} contains ${pattern}`); } } }); diff --git a/scripts/sync-upstream.test.mjs b/scripts/sync-upstream.test.mjs index 57d32b0..1f8564e 100644 --- a/scripts/sync-upstream.test.mjs +++ b/scripts/sync-upstream.test.mjs @@ -20,6 +20,12 @@ import test from "node:test"; const synchronizer = resolve("scripts", "sync-upstream.mjs"); const checker = resolve("scripts", "check-upstream.mjs"); +test("strict upstream checks require a source checkout", () => { + const checked = run(checker, ["--strict"]); + assert.notEqual(checked.status, 0); + assert.match(checked.stderr, /Pass --source /); +}); + function write(path, content) { mkdirSync(resolve(path, ".."), { recursive: true }); writeFileSync(path, content, "utf8"); diff --git a/skills/how/SKILL.md b/skills/how/SKILL.md index 5ab64f2..10a4446 100644 --- a/skills/how/SKILL.md +++ b/skills/how/SKILL.md @@ -20,7 +20,7 @@ When in doubt, take the simple path. Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Spawn all explorers in a single message: -- `worker type`: `generalPurpose` +- `worker role`: `explorer` - `model`: the active mstack `explorer` role, defaulting to `inherit-parent` - `readonly`: `true` @@ -30,7 +30,7 @@ Each explorer gets the prompt in `references/explorer-prompt.md` with its angle Spawn one Task subagent that explores and explains in one pass: -- `worker type`: `generalPurpose` +- `worker role`: `synthesizer` - `model`: the active mstack `synthesizer` role, defaulting to `inherit-parent` - `readonly`: `true` @@ -40,7 +40,7 @@ Build its prompt from `references/explainer-prompt.md` without the explorer-find Once all explorers have returned, spawn one Task subagent to synthesize their findings into one explanation: -- `worker type`: `generalPurpose` +- `worker role`: `synthesizer` - `model`: the active mstack `synthesizer` role, defaulting to `inherit-parent` - `readonly`: `true` diff --git a/skills/interrogate/SKILL.md b/skills/interrogate/SKILL.md index 350bed1..9b1669b 100644 --- a/skills/interrogate/SKILL.md +++ b/skills/interrogate/SKILL.md @@ -39,7 +39,7 @@ Launch all reviewers in a single message using the delegation tool. Use the conf | Reviewer | mstack `reviewer` role | For each reviewer: -- `worker type`: `generalPurpose` +- `worker role`: `reviewer` - `model`: the configured mstack `reviewer` role, or `inherit-parent` when no override exists - `readonly`: `true` diff --git a/skills/meta-mode/SKILL.md b/skills/meta-mode/SKILL.md index beade6a..4dabb22 100644 --- a/skills/meta-mode/SKILL.md +++ b/skills/meta-mode/SKILL.md @@ -5,6 +5,18 @@ description: "Route a non-trivial engineering task through a verifiable mstack w # Meta mode +## Apply the mode + +1. State the requested result in one sentence. +2. Choose the smallest matching playbook from `playbooks/`. +3. Read that playbook and every principle it names before you act. +4. Write a short todo list whose first entries are the playbook steps. +5. Choose a local or configured remote execution environment. Use a remote environment only when the task needs another machine. +6. Choose a model for each role from the user's mstack model configuration. Use `inherit-parent` when no override exists. +7. End each step with evidence from the real artifact. + +The mode stays active for the current task. Do not apply it to a casual turn or after the user opts out. + ## Non-negotiables The Principles section below grounds every trigger. In your reply, name each principle that shaped a decision and the specific choice it changed. Cite only principles whose leaf SKILL.md you read this session. @@ -29,6 +41,29 @@ Remaining triggers: - Broken skill mid-task → fix it in its own PR. Don't block. Don't silently work around it. - Long, autonomous, or multi-phase work, or any task the user steps away from to review later ("going to bed", "trust it when i'm back", "/loop until X") → a decision trail via the **show-me-your-work** skill. Commit it when stakes need an auditable record. Keep it local otherwise. +## Map capabilities + +Read [the capability map](references/capability-matrix.md) when a playbook +mentions delegation, background work, remote execution, model roles, or session +history. The map distinguishes a native implementation from a documented +fallback. Report the distinction when it changes the result. + +The canonical skill tree uses capability names. Adapters map those names to a +Harness. Keep vendor-specific paths, commands, transcript formats, and model +names in the Harness adapter or a user environment profile. + +## Resolve installed paths + +Playbooks use `` for the directory that contains the installed +mstack skills. They use `` for the optional `meta-mode-tools` +artifact directory. Resolve both paths from the active `meta-mode` skill before +you run a command. In this repository, the paths are `/skills` and +`/tools/meta-mode`. The default installer places them at +`/skills` and `/tools/meta-mode`, so the tools are at +`../../tools/meta-mode` from the `meta-mode` skill directory. If the install +uses an artifact destination override, use the destination that the installer +reports. + ## Principles Read the leaf skill in full for any principle you apply. Each entry names when it applies. @@ -93,6 +128,11 @@ not, run them one at a time and record the fallback. Select each worker's model from `/setup-mstack`; use `inherit-parent` when no override exists. Give each worker a focused file or question and read its complete result before deciding. +Choose models by difficulty. Route cross-cutting design, concurrency, and subtle +algorithms to the strongest configured judgment role. Route trivial mechanical +edits to the fast implementer role. Role-specific settings override these +defaults, and `inherit-parent` uses the parent chat model. + You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. A second opinion is the same prompt against a different model. Agreement is high-signal. ## Writing the reply diff --git a/skills/meta-mode/playbooks/worktree-cleanup.md b/skills/meta-mode/playbooks/worktree-cleanup.md index 2426a54..52ed1e9 100644 --- a/skills/meta-mode/playbooks/worktree-cleanup.md +++ b/skills/meta-mode/playbooks/worktree-cleanup.md @@ -7,7 +7,7 @@ 3. Verify usage before deleting. For every `verify-recent-chat` row, or anything you doubt, fan subagents out to read the transcripts and report whether the chat is pinned or ongoing and which worktrees it touches (principle-guard-the-context-window, transcripts are bulk). A pinned chat spawns arena and repro trees into sibling worktrees via background subagents, and those are in use even when their names never hit the sidebar. 4. Pause on irreversible loss. `wip:N` is N tracked uncommitted edits. Show the diff and get a decision first, since removing a clean worktree is recoverable from its branch but uncommitted work is gone. `scratch:N` is untracked throwaway, safe to drop, but name the files. Per Autonomy, clean and merged and not-in-use proceeds. `wip` and in-use pause. 5. Prune the confirmed set. Per path, `git worktree remove --force `. If the dir survives on ignored build artifacts, `rm -rf` it, then `git worktree prune`. Branch refs survive, so no commits are lost. Confirm with `df -h /` and re-list. -6. Simulators and other reclaimers. Simulators are usually the next-biggest win. `xcrun simctl --set testing delete all` (XCTestDevices clones), `xcrun simctl delete unavailable`, and `xcrun simctl runtime list` then `runtime delete ` for old runtimes. More when needed: Xcode `DerivedData` and `iOS DeviceSupport`, `~/Library/Application Support/the current harness` (`state.vscdb.backup`, and `snapshots/roots/` where a `` named for a folder you opened as a workspace balloons), package caches (pnpm, uv, brew, yarn). Clear only caches the user has not said to keep. +6. Simulators and other reclaimers. Simulators are usually the next-biggest win. `xcrun simctl --set testing delete all` (XCTestDevices clones), `xcrun simctl delete unavailable`, and `xcrun simctl runtime list` then `runtime delete ` for old runtimes. More when needed: Xcode `DerivedData` and `iOS DeviceSupport`, the active Harness state directory and workspace snapshot cache, package caches (pnpm, uv, brew, yarn). Clear only caches the user has not said to keep. This is the one playbook that deletes user state with no code review to catch a slip, so the gates above are the review. diff --git a/skills/meta-mode/references/capability-matrix.md b/skills/meta-mode/references/capability-matrix.md index a2599ed..9886edc 100644 --- a/skills/meta-mode/references/capability-matrix.md +++ b/skills/meta-mode/references/capability-matrix.md @@ -21,7 +21,7 @@ credentials, and run history when it is available. A local or SSH fallback must provide those lifecycle guarantees explicitly, so report the difference when it affects the result or the operator's control. -Benny remains a Cursor-specific automation pack because its Slack triggers and -run lifecycle depend on Cursor Automations. Its operational skills can reuse -mstack workflows, but installing mstack in another Harness does not recreate -those triggers. +Benny remains a vendor-specific automation pack because its Slack triggers and +run lifecycle depend on one Harness's automation service. Its operational +skills can reuse mstack workflows, but installing mstack in another Harness +does not recreate those triggers. diff --git a/skills/reflect/SKILL.md b/skills/reflect/SKILL.md index 0c4e6ff..08b8afe 100644 --- a/skills/reflect/SKILL.md +++ b/skills/reflect/SKILL.md @@ -35,7 +35,7 @@ One delegation request with three reviewers, an explicit model on each, and agen | Tooling | the active mstack `reviewer` role, default `inherit-parent` | `references/tooling-reviewer.md` | | Divergent | the active mstack `reviewer` role, default `inherit-parent` | `references/divergent-reviewer.md` | -Pass each template verbatim, substituting the transcript path or digest where marked. Reviewers return findings in the `Task` response body. +Pass each template verbatim, substituting the transcript path or digest where marked. Reviewers return findings in the delegation result body. ### 3. Synthesize diff --git a/skills/swarm/SKILL.md b/skills/swarm/SKILL.md index ad058fd..3207254 100644 --- a/skills/swarm/SKILL.md +++ b/skills/swarm/SKILL.md @@ -28,7 +28,7 @@ Open a todolist with one entry per phase before launching anything. Spawn all N workers through the active Harness's delegation API, using the configured model and background execution when supported. Use a local worker when the task needs access to something on the user's computer. -When a worker must start from a non-default pushed branch, pass `cloud_base_branch`. +When a worker must start from a non-default pushed branch, provide the active Harness's base-branch option. Every brief stands alone. Include the goal, scope, exact slice or race arm, how to verify, and what to report. Reports use `PASS`, `ISSUES`, or `BLOCKED` with evidence. diff --git a/skills/why/SKILL.md b/skills/why/SKILL.md index b50c7c7..e6a02bc 100644 --- a/skills/why/SKILL.md +++ b/skills/why/SKILL.md @@ -77,7 +77,7 @@ Aim for a complete **coverage map**, not a minimal one. Document the null, don't Launch all matching investigators in a single message so they run concurrently. Don't ask one agent to cover multiple MCPs. Subagent config (each): -- `worker type`: `generalPurpose` +- `worker role`: `explorer` - `model`: the active mstack `explorer` role, defaulting to `inherit-parent` - `readonly`: `false` (agent mode). **Do not use readonly/Ask mode.** It strips MCP access, which disables MCP-backed investigators entirely. Investigators still shouldn't write anything. @@ -121,7 +121,7 @@ If your scope assessment suggests a single-commit trivial target where the PR de Spawn one synthesizer subagent: -- `worker type`: `generalPurpose` +- `worker role`: `synthesizer` - `model`: the active mstack `synthesizer` role, defaulting to `inherit-parent` - `readonly`: `false` (agent mode). The synthesizer's quality check spot-verifies citations, which can require MCP access. Readonly/Ask mode strips MCPs and defeats that.