A narrative tour of the system: the harness engineering, the gates, the habit sensors, the skills, and the loop that lets an agent get better at a codebase over time. If you read one document, read this one.
For the rules themselves see CONTRACT.md; for the design rationale and decisions see DESIGN.md.
A CI pipeline and a set of "AI coding standards" are usually built as two separate things that slowly disagree with each other. Foundry's whole premise is that they are one rule set, run in three places:
| Placement | When it runs | Who it talks to | What runs |
|---|---|---|---|
| In-loop | while the agent is editing | the agent, mid-task | habit-hooks (a Stop hook) |
| Pre-push | before code leaves your machine | you | mise run gate (via /ship) |
| CI | on every pull request | the permanent record | the same gate + guards |
Run different rules in each and you get "passes locally, fails in CI", plus a failure mode unique to agents: the agent fixes what the hook flagged, CI complains about something else, the agent fixes that and regresses the first thing. One rule set, three placements, no drift.
The system is three moving parts, deliberately separated by who consumes them:
foundry(this repo) — the CI half. Reusable GitHub Actions workflows,misetask templates, config presets. Consumed by repositories.cmaintz-skills— the agent half. Claude Code skills and hooks (ship, the habit-hooks Stop hook). Consumed by Claude Code as an installed plugin.habit-hooks— a third-party tool (installed viauv), not written by us. It's the structural-smell sensor layer, and Foundry borrows its tool-independent smell vocabulary as a backbone.
The seam between foundry and cmaintz-skills is CONTRACT.md,
copied verbatim into both. If those two copies ever need to differ, the split
was wrong.
Every repo, whatever language, exposes the same six verbs through mise:
mise run fix # auto-fix what is mechanically fixable
mise run lint # style + structural smells (non-mutating)
mise run typecheck # static types
mise run test # tests + coverage floor
mise run audit # vulnerabilities, secrets, SAST
mise run gate # all of the above, in order. THE ORACLE.
Everything else in the system — CI, the git hooks, the ship skill — calls
verbs, never tools. A skill says mise run lint; it never says eslint or
phpstan. That indirection is the entire reason one skill library can serve a
React repo and a Kotlin repo: the smell vocabulary is language-independent, only
the tool behind each verb changes.
mise also pins the toolchain (node = "22.14.0"), which kills
"works-on-my-machine" between a laptop and the CI runner at the same time.
This is the load-bearing distinction in the whole design.
- Deterministic layer = the oracle. Formatters, linters, type checkers, tests + coverage, dependency/secret/SAST scanners. Pinned, reproducible, offline-capable. These decide pass/fail. Nothing else does.
- Probabilistic layer = a proposer, never an authority. An LLM's output is a patch. A patch is accepted only if the deterministic gate then passes from a clean tree. A model's claim that it fixed something is not evidence — the exit code is.
Concretely: an agent may run mise run fix, may act on habit-hooks' coaching,
may write code — but "done" is defined solely as a green mise run gate. The
agent proposes; the gate disposes.
Three probabilistic roles are kept deliberately separate (people mush them):
- Fixer — given a failure + coaching, produce a patch. Bounded.
- Reviewer — the things a linter genuinely can't judge: naming, domain drift, "is this the right seam", spec compliance. Advisory, never blocking.
- Triager — flaky vs real, dedupe, route to a ticket.
The CI gate (gate.yml, or foundry's reusable ts.yml + _guards.yml) is four
jobs. CI is 100% deterministic on purpose — no model, no API key, no secrets
beyond the default token.
gate(Deterministic gate) — runsmise run gate: the exact verb a developer runs locally. If this and a laptop ever disagree, that's a bug in the setup, not the code.habits(Structural smells) — runshabit-hooks. Fails on new smells beyond the snoozed baseline (see §6).secrets(Secret scan) — gitleaks over full history.ruleset-guard— the anti-gaming control (see §7).
habit-hooks is the reflex layer. It wraps detectors (eslint, knip, ts-morph,
jscpd, PMD, phpmd, ruff…) and maps their findings onto tool-independent
smells: oversized-file, oversized-function, too-many-parameters,
high-complexity, deep-nesting, duplication, dead code. A smell means the same
thing in Kotlin and PHP; only the detector differs.
What makes it more than a linter: when it fails, it emits coaching aimed at
the agent, and that coaching argues against mechanical compliance —
"splitting a file at line 200 into foo-1.ts and foo-2.ts satisfies the
threshold while leaving the real problem in place." That anti-gaming framing is
why it's used to teach an agent rather than just gate it.
It runs in two placements:
- In-loop: a global Stop hook (
~/.claude/settings.json) fireshabit-hooks-guard.ps1when the agent is about to finish — but only in repos that opted in by having a.habit-hooks/directory, so it's silent everywhere else and safe to install globally. It's a Stop hook, not PostToolUse, because habit-hooks costs ~25s cold (~6s warm) — far too slow to fire after every edit, and it matches habit-hooks' own guidance: "run before considering work complete." - CI: the
habitsjob.
Config lives in two places, and it matters which:
- Global (
~/.claude/settings.json): the Stop-hook wiring only. - Per-repo (
.habit-hooks/config.toml,.habit-hooks/snooze.json): which plugins are on, and the accepted-baseline of existing smells. Committed with the code.
Retrofitting linters onto a real codebase makes everything red on day one, and you abandon it in week two. So Foundry never gates on absolute cleanliness; it ratchets:
- ESLint's native bulk suppression (
--suppress-all/--prune-suppressions) records existing violations ineslint-suppressions.json. - habit-hooks'
habit-snoozerecords existing smells insnooze.json. - Coverage floors are set to today's real number, not an aspiration.
Each baseline is committed and may only ever shrink. Fix an any, and
--prune-suppressions removes it from the baseline permanently.
The obvious attack on any ratchet is to weaken the rule instead of fixing the
code — the cheapest fix for high-complexity is // eslint-disable. So
ruleset-guard enforces mechanically: a PR that changes a ruleset file
(configs, thresholds, suppression baselines) and production source is blocked
unless a human applies the ruleset-change label.
The guard is tightening-aware — the risk is one-directional. Loosening
needs a human; tightening, or a change with no semantic effect, never does. So
a bundled ruleset+source PR passes without a label when every ruleset file it
touches is proven not to loosen the gate: a suppression baseline whose count only
shrank, a snooze list that only got shorter, or a pure CRLF/LF flip. This is what
lets mise run fix prune the baseline and ship that in the same PR as the fix.
Governance is split across three enforcement points by when they run — a
client-side hook (cmaintz-skills), the CI guard (_guards.yml), and the regenerator
(bootstrap.yml). This table is the single map of what's protected and how:
| File | Controls | Protected by |
|---|---|---|
snooze.json |
structural-smell baseline | hook blocks hand-edits (incl. Bash writes) · ruleset-guard blocks growth without the label · bootstrap regenerates (prune only shrinks) |
eslint-suppressions.json |
lint baseline | same three |
.jscpd.json |
which paths the duplication sensor scans | ruleset-guard — widening the ignore list alongside source needs the label |
pmd/ruleset.xml, thresholds, mise.toml, workflows |
rule definitions / gate wiring | ruleset-guard — direction can't be proven, so any change alongside source needs the label |
ruleset_paths regex |
which files ruleset-guard treats as gate-defining (this table's first column) | the guards input default, or a consumer's override — keep it in sync when you add a guarded file |
The recurring failure mode: a new gate-defining file (like .jscpd.json) isn't
added to ruleset_paths, so widening it slips past the guard. When you introduce
one, add it to the regex in the same change.
Foundry's own skill layer is deliberately thin, because most of the practice layer is already written well by others. They're installed as plugins (not vendored), so upstream fixes flow and attribution stays put.
| Source | Role | How it touches the gate |
|---|---|---|
| habit-hooks | reflex / enforcement | is the habits placement; coaching feeds the fixer |
| mattpocock/skills | breadth of practice: tdd, diagnosing-bugs, research, codebase-design, code-review (Standards+Spec)… |
procedures the agent runs; ship calls code-review |
| devill/ivetts-skills | the flywheel: learn, hotspot-rec, build-project-review |
learn routes lessons into the gate; hotspot-rec reads git history for what to improve |
| cmaintz-skills (ours) | ship (pre-PR orchestration) + the habit-hooks hook |
drives the whole gate locally |
Three layers, cleanly divided:
- Reflex → hooks (habit-hooks, pre-commit). Involuntary.
- Practice → skills (
tdd,ship,code-review…). Invoked, procedural. - Standing context →
AGENTS.md+CONTEXT.md. Ambient shared vocabulary. (CLAUDE.mdis a one-line include ofAGENTS.md, so Codex/Gemini/Cursor read the same source.)
/ship is where a change goes from working tree to PR, running the gate before
the PR exists (which is what keeps CI free of API keys):
mise run fix— mechanical fixes land silentlyhabit-hooks— coaching → the agent fixes the smellsmise run gate— must be green. The oracle, not the agent's opinion.- Review in a fresh context — a sub-agent that sees only the diff and the
spec, never the conversation that wrote the code (an agent reviewing its own
work reviews its intent, not its diff). Uses a repo-local
build-project-reviewskill if present, elsemattpocock-skills:code-review. - Conventional commit →
gh pr create.
Installing everything leaves three things called "code review". They layer,
most-specific first: a repo-local build-project-review skill → Matt's
code-review (Standards + Spec) → the built-in /code-review (fast manual bug
hunt). They namespace as plugin:skill, so nothing actually collides.
This is what makes the system get better rather than just stay clean. Ivett's
learn skill reflects on a session and routes each lesson to the store that will
actually enforce it, in priority order:
- a deterministic hook / check — enforcement, so the mistake becomes impossible
AGENTS.md/CLAUDE.md— standing context- a new skill — a reusable procedure
- auto-memory — last resort
That ordering is the project's thesis expressed as a skill: prefer the
placement that makes a mistake impossible over the one that merely reminds you
not to make it. A one-off fix in this session becomes a rule the next session
can't skip. Unlike habit-hooks, learn is a model-invoked skill (not a hook) —
it can't auto-fire, so the practice is to run /learn at session boundaries,
before anything gets written to memory.
The loop, end to end:
work → habit-hooks flags a smell → agent fixes it → /learn decides:
"this class of mistake should be impossible" → new deterministic check
→ next time, the gate catches it before a human ever sees it
Proven and built out: TypeScript / Node (mise/ts.toml, ts.yml, eslint +
tsc + vitest + habit-hooks-typescript).
Reachable cheaply (habit-hooks already has detectors): Java (PMD), PHP
(phpmd), Python (ruff), Ruby (rubocop). Each needs a mise template +
a workflow.
No habit-hooks support, would use native tooling: C#/.NET (Roslyn analyzers / editorconfig), Kotlin (detekt), Swift (SwiftLint), HTML/CSS (a line-count / max-file gate, since the full smell suite doesn't apply).
The architecture is additive: a new stack is a new mise template and a new
reusable workflow, mapping that language's tools onto the same six verbs and the
same smell names. Nothing already built changes.
~/.claude/settings.json # global: the habit-hooks Stop hook, no-attribution config
~/.claude/hooks/ # the hook script (habit-hooks-guard.ps1)
<repo>/mise.toml # the six verbs for this repo
<repo>/.habit-hooks/ # config.toml + snooze.json (the smell baseline)
<repo>/eslint-suppressions.json # the lint baseline (ratchet)
<repo>/.github/workflows/ # the gate (inline, or calling foundry)
<repo>/AGENTS.md # standing context (CLAUDE.md @-includes it)
foundry/ # reusable workflows, mise templates, this doc
cmaintz-skills/ # ship skill + the hook, as an installable plugin
Applying Foundry to a real polyglot monorepo (a Spring Boot backend + Angular frontend + a browser extension) surfaced patterns worth codifying:
- Verbs are namespaced per package. One root
mise.tomlexposesbackend:gate,frontend:gate, etc., plus an aggregategate. CI runs each package's gate as its own job. - habit-hooks is per package, not per repo. Its TypeScript detectors
(eslint/knip/ts-morph/jscpd) resolve from the package's own
node_modules, so the config and snooze baseline live infrontend/.habit-hooks/, and the sensor runs fromfrontend/. Backend Java/PMD lives inbackend/.habit-hooks/. - Generate the smell baseline in CI, not locally. On Windows the full-repo
file list blows the ~8191-char command-line limit, so a manual
workflow_dispatchjob generates eachsnooze.jsonon Linux and commits it. The per-package smell jobs soft-pass until their baseline exists. - In-loop scope is
--branch, not--all. The Stop hook checks only the changeset — correct for "flag what you touched", and it dodges the Windows limit. - Dependency audit needs a lockfile, generated on demand. osv-scanner reads a
gradle.lockfile; enable Gradle dependency locking but keep the lockfile gitignored and regenerate it in the audit job, so lock drift never breaks the gate. - Ratchet by diff where a baseline file is overkill. Spotless
(
ratchetFrom origin/main), Semgrep (--baseline-commit) and the no-varPMD rule (CI checks only changed files) all enforce on new code while leaving legacy untouched — the ratchet, without a committed baseline. - Integrate, don't replace, existing CI. The pilot kept its Docker image-publish job verbatim and only added Foundry's gate + guard jobs alongside.
A consumer doesn't put everything in one ci.yml. Split by cadence and blast
radius into separate workflow files, each composing foundry's reusable pieces:
| File | Calls (facade) | Trigger |
|---|---|---|
gate.yml |
gate.yml facade — with: { stack, working_directory }, once per package; includes structural smells |
PR + push |
security.yml |
security.yml facade — secrets + ruleset-guard + SAST in one call |
PR + push |
bootstrap.yml |
bootstrap.yml — baseline refresh |
workflow_dispatch |
deploy.yml |
image publish / release — push-to-main only | push |
# .github/workflows/security.yml — pin the facade; it runs secrets + ruleset-guard + SAST
name: security
on: { pull_request: {}, push: { branches: [main] } }
jobs:
security:
uses: CMaintz/foundry/.github/workflows/security.yml@v2
permissions: { contents: read, pull-requests: read } # the secret scan lists the PR's commits
with:
ruleset_paths: '^(mise\.toml|backend/\.habit-hooks/|frontend/\.habit-hooks/|\.github/workflows/)'Why split, not one file: needs: can't cross workflow files, so unrelated jobs
don't serialise; a slow quality run doesn't gate a fast security result; and
each file has one obvious trigger. Keep deploy push-only — a deploy job
under pull_request shows up as a permanently skipped check, which is noise and
can wedge branch protection (see below).
-
concurrencylives in the caller, not the reusable (a reusable workflow can't set it for you). Give each caller fileconcurrency: { group: <name>-${{ github.ref }}, cancel-in-progress: true }so a new push cancels the in-flight run for that branch instead of paying for both. -
Path-filter what genuinely can't be affected — but mind the required-check trap. A docs-only PR doesn't need the TS gate. The naive fix (
on: pull_request: paths:) backfires: a required check that's path-filtered out is reported asExpectedand never arrives, so the PR can't merge. The working pattern is a single always-running aggregate job that the branch protection requires, whichneeds:the real jobs and passes when they either succeed or are legitimately skipped:jobs: gate: { if: ..., uses: ... } # heavy, may be skipped by an inner filter gate-ok: # THIS is the required check needs: [gate] if: always() runs-on: ubuntu-latest steps: - run: '[ "${{ needs.gate.result }}" != "failure" ] || exit 1'
Require
gate-ok, never the heavy job directly. Skipped ≠ failed, so a docs PR goes green without running the gate, and a real failure still blocks. -
Consider dropping the
push:mainrun entirely — it's usually redundant. The tempting rationale for keeping it is a post-merge integration check: with strict mode off, a PR can merge behind main, so only a main run sees the merged result. But that gap is self-healing — the next PR's gate runsmise run gateagainst current main (its merge-base includes the integrated result), so any real break surfaces there, one PR later. With rebase-merge the merged commit is also the same code the PR gated. So on a minute-constrained repo, run the gate suite onpull_requestonly (keepworkflow_dispatchas a manual full-run for after a genuinely risky merge); leavepush:mainfordeploy/bootstrap. Measure before assuming value:gh run list --branch main --event push— if no main run has ever caught what a PR missed, it's pure cost. If you do keeppush:main(e.g. you can't rely on rebase-merge), at least scope it: the naive classifier emits "run everything" off-PR, re-running the whole suite unfiltered on every merge. Give thechangesjob an event-aware range — PR →merge-base..head;push→github.event.before..github.sha; dispatch/first/ force-push → run everything (fail-safe) — so a backend-only merge skips frontend.
- Required check names must match the job's
name:string exactly. Protection matches on the rendered check name, not the job id. Rename a job (e.g.SpotBugs (backend)→SpotBugs (backend, report-only)) and the old required check never reports — PRs hang "Expected" forever. Update protection and the job name in the same change. - A
workflow_dispatch-only job never appears as a PR check — so don't addbootstrapto required checks; it would block every PR waiting on a run that isn't coming. - A PR opened with
GITHUB_TOKENwaits for approval. GitHub treatsgithub-actions[bot]as a first-time contributor, so the prune PR's checks sit at "awaiting approval" and auto-merge never fires. Givebootstrap.ymla GitHub App (app_client_id+app_private_key) and the PR is opened as the App instead. - A job gated
if: github.event_name == 'pull_request'is skipped on push. That's fine for a required check (skipped ≠ failed on the branch it doesn't run on), but a job skipped on the PR itself (wrongif) counts as neither pass nor fail and can stall the merge. Gate on the event, not by accident. - Enable "require branches to be up to date" (strict mode) so a PR is re-tested against the latest main before merge — this, plus the guard diffing from the merge-base, is what stops a stale base from either sneaking a regression in or false-flagging an untouched file.