Conversation
…j2 partial (M1) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… loader (M2) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… --project (M3) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… (M4) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…k skill + docs (M5) Kill criterion failed (+10% vs -30% target); old /groom stays canonical, new experimental /groom-playbook skill drives the core groom playbook. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… dirs Each jinja body now renders with its own ChoiceLoader chain: the body itself (by basename) -> step dir -> playbook dir -> local/global/core playbook roots -> src/templates. A PlaybookEnvironment.join_path override resolves ./ and ../ against the including file's own directory at any nesting depth; bare names fall through to the chain unchanged. src/templates stays last so _partials/ keeps resolving plugin-root relative, at the cost of a playbook-local _partials/ shadowing it. The graph and step chrome templates render through a separate plugin-only env so a playbook cannot shadow them. Also drop the "no playbook.md, skipping" warning: a directory without a manifest is now skipped silently. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…se, validation, per-scope waves (M1)
…reference docs (M4)
…amed states in loader (M1) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… transition report (M2) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…or (M3) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ll-run integration test (M4) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…iving protocol, docs (M5) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…lenient inputs/outputs parsing (M1)
…ch + inputs/outputs bullets (M2)
…iving protocol (M3)
…uts reference + bootstrap driving (M4)
…ybook scan + name clash (M1) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ons.j2 + step append + clash STOP + --no-lessons (M2) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rker + driving lessons binding (M3) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s reference, stale shadowing refs removed (M4) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tch file target + {instance} interpolation, plan-transition rejection, report prefix (M1)
…vocabulary docs across playbook.md, CLAUDE.md, driving partials (M2)
…ites Rebuild core groom playbook: playbook.yaml with states/graph, new steps (draft-plan, decompose-work, verify-references), _scripts status hooks, _specs design docs. Each step ships promptfoo eval suite (tests.yaml, fixtures, model variants). Add eval tooling: justfile recipes (eval, eval-md, smoke, regress), bin/eval-md.sh + bin/report-md.jq markdown reporter. Ignore repo-root _booping/ runtime log. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- frontmatter-update on a file without a frontmatter block now prepends one instead of aborting; existing content becomes the body unchanged - groom steps stop authoring `reviewed_at: null` — run hooks own the artifact frontmatter, files open at their H1 - run slugs take their timestamp from the local clock (`date +%Y%m%d-%H-%M`), never from model memory Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ansition An absolute --target ending in the machine's declared artifact: now implies the workdir, so hooks run beside the artifact without a separate --workdir. Relative targets, non-matching paths, and machines without artifact: keep cwd. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…script mining core
…capture and wiring
A workdir inside home_dir/{project}/ now assembles the same context as a
repo workdir; repo_directory becomes optional and hook resolution falls
back to the process cwd when the containment-resolved project has none.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…eview to awaiting-approval Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…or release — M1 snapshots + mdcheck as Python scripts Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…or release — M2 eval tooling moves to scripts/ Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…or release — M3 dead code deleted Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…or release — M4 CLAUDE.md layout + commands sweep Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…nting — M1 extraction engine
…nting — M2 session-stats subcommand, session-time retired
…nting — M3 metrics_ key rename, migration 006, config surfacing and docs
…nting — CHANGELOG entry for session-stats
…agnitudes in units
…ncached remainder
The vault lived under home_dir at ~/Dev/@A/notes/projects/claude-booping, tracked in the private notes repo. Move it to vault/ here so plans, retros and lessons version alongside the code they describe. - .booping gains vault_path: vault - _core_playbooks becomes a relative symlink (../playbooks), keeping the absolute home path out of a public repo - vault/notes/ and .booping.log are gitignored: local scratchpad and operational log, not part of the published vault - drop _booping/.booping.log, dead since migration 004
Migration 004 converted _booping/agent_*.md into targeted _lessons/ files; nothing has read that path since. The developer agent description still pointed at it. Lessons are discovered automatically, so the sentence is removed rather than repointed.
…ss cases Both judged a freshly generated artifact whose shape varies run to run — the fixtures corpus rubrics and the llm-tests row-quality rubrics went red on roughly half the runs without any source change. The deterministic smoke cases stay. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The suites judge freshly generated artifacts, so a full run lands on a different red set each time — a per-sha gate blocked merges on judge variance rather than on source drift. The sticky PR comment stays as the advisory record of a run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Six incremental migrations described a history no 1.0 vault will ever walk step by step. Fold them into three that each cover a coherent layout change: vault layout, tracks and lessons, session metrics. - fixture marker watermark drops 6 -> 3, and the migrate snapshot follows - README gains a 'What it's for' section, links the checked-in vault/, and states the beta-grade maturity of 1.0
Owner
Author
|
Superseded by #20, which releases the same commits from |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #18.
Overview
The playbook track toward 1.0 — eighteen sprints on the
v1.0branch:Groom-as-Playbook pilot —
_playbook_driving.j2shared driving protocol, core discovery scope, dir-form steps, Jinja-rendered bodies with--step/--project, the coregroomplaybook + experimental/groom-playbookskill (parallel to/groom, which stays canonical).Jinja include resolution — playbook/step-dir-relative
{% include %}via per-bodyChoiceLoaderchains.Subgraphs with repeat — one level of nested graphs (
dependencies+ innergraph+ driver-judgedrepeatprose), per-scope wave resolution, subgraph render sections; step bodies no longer embedded in the composed render.Run state machines — opt-in
states:inplaybook.yaml(artifact-backed,{instance}-aware),playbook-statefrontier reader +playbook-transitionas the single writer of run state withfrontmatter-update/scripthooks, resumable runs. File-targetfrontmatter-updatehooks bootstrap a frontmatter block on files that lack one.Metadata-driven runner — uniform
--stepfetch for every section, bootstrap-spawn protocol: the driver never pastes step bodies, delegated agents fetch their own instructions.Playbook lessons —
_lessons/dirs at four scopes (global root, project root, playbook, step viastep:frontmatter), merged across discovery roots most-specific-wins, injected byrender-playbook: composed## Lessonssection for the run, step lessons appended to--stepoutput,--no-lessonsescape hatch. Playbook name shadowing replaced by a blocking name-clash STOP.Playbook-authoring into core + delegation-level model — the
playbook-authoringplaybook and the shared_lib/eval engine (promptfoo assert modules, subscription-auth claude provider, grader/rubric prompts, return-contract convention) move into the plugin as core assets; step fieldagent:renamed todetached:framework-wide (legacy key → blocking STOP); advisoryinputs:/outputs:step frontmatter dropped from the framework (superseding the earlier metadata-runner shape); delegation levels + review-gate design taught to the authoring playbook.Groom playbook respec (v3) — plan-as-directory:
plans/{slug}/doubles as plan directory and run workdir (index.md+plan.md), retiring_runs/; directory-plan support in the plan model;research_agentconfig key + driving-partial protocol updates; regenerated manifest, preamble, and_scripts/; step prompts rewritten to the new step contracts; per-step eval suites and fixtures realigned. Then the graph collapsed to six steps (intake → research-codebase / research-web → draft-plan → cross-review → present) with a new detached second-modelcross-reviewstep, intake reworked around context tables + an on-disk brief, and gate confirmations / step questions routed throughAskUserQuestion(review gates stay prose asks).Develop-as-Playbook port — the core
developplaybook, parallel to/develop(which stays canonical): flat chainintake → provision → develop-loop → verify → wrap-upwithinline_steps; one run machine writing the shared plan-lifecycle statuses straight onto the plan's ownindex.md(plan directory = run workdir, no_scripts/, nosprints.mdre-render);verifyis the only detached step (sonnet:medium), a guardrails-only validator with no user-review gate — the run's only stops are the branch-name confirmation and intake approval; step bodies stay near-verbatim to the skill prose with the referenced docs and partials wired (_git_guideshared atplaybooks/_partials/); authoring_specs/+ decision log committed alongside.Committed playbook reports —
booping render-playbookgains a repeatable--set <key>=<value>config-override flag and a config-pinnablenow()(--set now=…freezes every rendered timestamp), making renders byte-reproducible; a minimal fixture vault atplaybooks/_fixtures/vault/makes them project-independent;just playbook-reportsrenders every core playbook into a committedplaybooks/<name>/_reports/output.md, so source drift shows up as a diff.Retro-as-Playbook port — the core
retroplaybook, parallel to/retro(which stays canonical): flat chainintake → prepare → gather-feedback → research-issues → synthesize → save, step bodies built verbatim on the skill's phases (intake = Phase 0, prepare = Phase 1 with the two parallel mining briefs, gather-feedback = Phase 2 whole — open-ended round then issue triage + per-plan goal verdicts, research-issues = Phase 3); run machineawaiting-retro → awaiting-learningon the primary plan's ownindex.md, exit edge stampingreviewed_atand closing the multi-plan working set via aclose-working-setscript; the save closing report tables the issues per plan with root causes + action items; the retrospective template moved out of the eagerly-embedded partial into lazy-loadeddocs/retrospective_template.md, read at draft time by both the skill and the playbook; authoring_specs/+ decision log committed alongside. Plan discovery hardening rode along:plans/{slug}/index.mdrecognized besideplan.md(plan.mdwins in legacy dual-file dirs), unloadable plan files skipped with a warning instead of crashing, and the develop-loop rule that every milestone group gets a fresh worker agent.Config
corenamespacing + lifecycle retirement — every setting the shipped playbook set owns moves underconfig.core(playbook-owned atcore.{name}_playbook, loop-wide or multi-playbook directly undercore), leavinghome_diras the only other top-level key; theskills:block is deleted. Config validation goes with it —validate_skills,SkillConfigand theUNTRUSTED_PROJECT_KEYStier filter are removed, so a user's own playbook namespace needs no plugin change to be legal and no tier drops a key (a project-tiercore.macrosentry now executes). Thechatandhelpskills are deleted, leavingcode-reviewandplaybook; every delegation table renders from the singleplaybooks/_partials/playbook_agents.mdand bothavailable_agentssurfaces are gone. The shared plan lifecycle is retired wholesale:config.plan, thetransitionandvault-commitcommands and theplan_status:frontmatter mirror are all deleted, each playbook's ownstates:block becoming the whole status vocabulary — groom's three status-mirror scripts give way to an explicitcommit-planhook, retro's skipped-sibling close to a localdrop-planscript, and develop's twovault-commitcallers to plain staged git.marker-setlearnslatest_migration=@latest,booping-create-projectseeds it on both branches, and migration002_config_core_namespacingships to carry existing vaults across the rename (global tier called out as a manual step). A shared_partials/timestamps.mdgives any playbook the run's clock.Plan date + summary hygiene —
createdis stamped to the minute (YYYY-MM-DD HH:MM) and rendered into the plan frontmatter shape from the run's own clock rather than documented as a bare date; groom'sintakenow carries that shape explicitly, having previously createdindex.mdwithout ever being shown it.completedis normalised to the same shape across the vault, and the surrounding date keys are documented as sharing it.Hook clock through the macro system — the bespoke
@now/@todayinterpolation tokens are retired in favour of the same call the templates use,{{ macro('core.macros.date', '+%Y-%m-%d %H:%M') }}, written directly in the hook value and rendered as Jinja with themacroglobal; hook tokenising moves toshlexso a quoted expression survives the split. One clock source for the CLI and the rendered bodies, and--stub-macronow pins a transition exactly as it pins a render.@headstays a token — it resolves against the repo directory, where a macro would run in the process cwd.Split-boundary cleanup — the three residue families the PR-19 architecture review catalogued, cleared in one sprint.
cancelledbecomes a real terminal on both the groom and develop machines, declared once per superstate (in-spec/awaiting-plan-reviewon groom, a newcancellablesuperstate on develop) and inherited byresolve_edgeswith zero engine change, so thelatest_plansquery'scancelledfilter finally has a writer. The two near-identicalclose-working-setscripts collapse into one parameterized script in a new sharedplaybooks/_scripts/root:_dispatch_scriptresolves hooks playbook-dir first then each discovery root, and passes trailing hook tokens to the script as argv (script close-working-set --status done --prefix learn), so per-playbook differences live in eachplaybook.yamlhook line. Playbook bodies stop restating the machine — retro/learn intake validate against "the status the## Statesection names as this run's entry" instead of spelled literals, four transition-restating blocks reduce to one-line pointers, and three manifest graph narrations go (the step table is the contract). Fossil sweep: every retired-skill handoff rewritten to/playbook {name}form, the "parallel to /X" trigger clauses and develop's "stays canonical" rule dropped, the deadsrc/templates/_partials/_learn_targets.j2folded into the live learn copy,docs/cross_validation.mddeleted; dead engine code out (Edge.skill,DIR_PLAN_NAME,config["plan"]docstrings,macros.now-era help strings),specs_dirdropped from the bootstrap prompt, code-review prose repointed at the delegation table, and thechore/branch rule reworded out of the task-type namespace.Flat lessons with
targets:— lessons collapse to two flat_lessons/roots (global home + Project Vault), each file routed by atargets:frontmatter list ({playbook},{playbook}/{step},agent:{id},skill:{name}); an untargeted lesson injects nowhere. The per-playbook/per-step_lessons/dirs and thestep:key are retired (legacy content surfaces a migration note), and learn now shows the playbook table of contents, fetches target spaces on demand, and pins exacttargets:in its review table before writing.Scaffold CLI —
booping scaffold <dotted.path> <dest>materialises any config-declared file tree (playbook skeletons, the vault layout, a project's own trees) with Jinja-seeded file content,--force/--set/--stub-macrocontrols, and the setup flow reuses it for vault creation.Frontmatter query surface — any markdown directory in the vault becomes queryable:
booping queryon the CLI,query/as_tableJinja filters in bodies, config-declared query specs addressable at any dotted path (mirroringscaffold/config-get), with--wherefiltering, sorting, column projection and table/JSON/YAML/paths output.macro()over config-declaredcore.macrosreplaces the fixednow()global,--stub-macropins any macro for reproducible renders, plan discovery becomes the user-editable orderedcore.plans.globkey, andsprints.mdis seeded once as a live Obsidian Bases view —render-sprintsand/chat's re-render are gone.Query surface: migration support — the
migrateplaybook goes from listed to functional: every render surface gates on the.boopingmarker'slatest_migrationwatermark and stops with a/playbook migratenotice when behind; migration authoring gets numeric--whereordering (:gt/:lt), query specs over the plugin's own files (root: core),{{ booping.latest_migration }}in templates, and format-preservingbooping marker-setwrites.Testing infrastructure — committed render snapshots grow teeth:
just snapshotsdiffs the committed playbook reports against a fresh hermetic fixture render (writing nothing),just snapshots-acceptis the only baseline writer,mdcheckenforces structural rules over every report, andjust cichains lint → typecheck → pytest → snapshots → mdcheck, run by CI on every push and PR.core.macrosentries gaincommand:+cwd: repo|vaultscoping (thegit_commitmacro),--projectpins the config merge as well as discovery for byte-identical renders, and the@headhook shorthand is retired in favour of thegit_commitmacro.Retro track split — the plan track now ends at develop: plans close at
done|fail|cancelledwith no awaiting-retro tail. Retro and learn run as their own track over a standaloneretrospectives/{slug}.mdartifact (awaiting-retro → awaiting-learning → done), joined to plans only through theretro:frontmatter key (null= the retro queue);playbook-transition/playbook-stategain--targetaddressing and a machine may declare noartifact:at all; an idempotent migration relocates oldplans/{slug}/retro.mdfiles and normalises terminal statuses.Code review track split — reviews persist as vault artifacts
codereviews/{plan-dirname}/{ts}.md(ad-hoc scopescodereviews/{target-slug}/{ts}.md,plan: null) with their own machinein-agent-review → human-review → done; plans accumulate every closed review incode_reviews:as history rather than a queue flag, so everydoneplan stays reviewable and re-review is first class; review templates gain the global tier, layering core → global → project with by-name override.Docs brought current against the delivered set in a spec-driven docs run: README, the full MkDocs site (
documentation/), andCLAUDE.mdrewritten and source-verified page by page;CHANGELOG.mdseeded at v1.0.0 with the release's user-visible changes.🤖 Generated with Claude Code