Repository navigation
feat: add skills and cross-run memory to the Loop language - #1
Merged
Merged
Conversation
Add two optional, first-class knobs to the Loop language: - `use skills: a, b` — named execution skills the loop may invoke during plan/act (coordinate proven skills instead of one mega-prompt). - `done when the skill "x" approves` / `scores N or more` — a review skill as the verifier, bridging an abstract goal to a verifiable check. - `remember in "<file.md>"` — cross-run memory: read past lessons into the first plan, append a dated outcome entry on stop. Wired end-to-end: loop-spec schema + TS mirror, parser, runtime engine (memory read/write, skill-predicate routing to runSkill, reflection capture), Claude Code runner (runSkill + skills/memory in prompts, grants the Skill tool, negation-aware parseSkillVerdict), MockRunner, CLI events, viz label, VSCode highlight/hover/completions, docs, examples, and parser+runtime tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X74Z3fZXh7hiSZ8LmVzpGs
Add a self-contained animated SVG hero at the top of the README — a terminal trace running examples/fix_test.loop: plan → act → observe (FAIL) → reflect → plan → act → observe (PASS) → done. Lines default opaque, so it degrades to the full static trace where CSS animation isn't honored. No build tooling or fonts. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X74Z3fZXh7hiSZ8LmVzpGs
tickets-forge-dev
force-pushed
the
claude/loop-engineering-guide-fwm4iu
branch
from
June 24, 2026 04:41
68dd417 to
ebbeda2
Compare
tickets-forge-dev
added a commit
that referenced
this pull request
Jun 29, 2026
* Plan: agentic-engineering constructs for Loop, broken into stories Translate "The New SDLC with Vibe Coding" into a Loop roadmap, in the deck's own vocabulary. Adds a story-by-story implementation plan and a dogfooded pipeline (today's grammar building tomorrow's constructs). - docs/agentic-engineering-plan.md: the epic broken into 13 stories across 4 waves (eval core, the dials, harness completeness, niche/ standards) plus a cross-cutting docs reframe; each story names files, a runnable `done when`, dependencies, and gates. - agentic-engineering.loop: the same epic as a pipeline (epic->stages = stories), validated via `loop-run parse` with its flow printed. Headline: make EVALS first-class (tests vs evals; output vs trajectory), the gap every analysis lens ranked #1 — "without both, it is vibe coding". Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Config defaults: config-tier `each cycle:` + project `loop.config` file Make the language easier to use by removing repeated config. Two tiers, both following the existing lowest-wins cascade (git/models): 1. Config-tier `each cycle:` — a file-level default cycle that every loop inherits; a per-loop `each cycle:` still overrides it. Resolved at parse time and threaded through pipelines/stages. 2. `loop.config` (or `.looprc`) — a per-project defaults file in the same config-tier syntax (each cycle / models / git). The runner walks up from the .loop file and folds it in as the lowest tier, so a file's own config and per-loop directives override it. `parse()` gains an optional `defaultCycle` so the project default seeds the cycle cascade. Also: schema + IR (`Config.cycle`), `show` renders the file-level default, AGENTS.md documents both, and agentic-engineering.loop dogfoods the cycle hoist (13 repeated `each cycle:` lines collapsed to one config-tier line). Tests: parser 32, runtime 75, vscode 14, stdlib/viz green (129 total). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Wave 1 Story 1: evals — tests vs evals, output vs trajectory, the bar Make verification a conjunction and evals first-class, using the deck's `done when … on the output / on the trajectory` syntax. - IR + schema: `Loop.doneWhen` is now a `Predicate[]` (all must pass). The skill predicate gains `subject` ("output" default | "trajectory") and `bar` (an inline `the bar:` rubric). - Parser: collects multiple `done when` lines; parses the `on the output`/`on the trajectory` qualifier; attaches an indented `the bar:` line to a skill eval (error if applied to a non-skill). - Engine: observe evaluates the whole conjunction, short-circuiting on the first miss; routes skill predicates (evals) to runSkill, the rest to the shell verifier. - Renderers: show + viz render multiple predicates and label evals/ trajectory; show now renders skill evals (previously dropped). - vscode linter, schema, and AGENTS.md updated; tests added. Trajectory predicates currently receive the act summary as context; Story 2 wires the real captured trajectory. Tests: parser 36, runtime 75, vscode 14, viz 5, stdlib 3. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Wave 1 Story 2: capture the trajectory and route it to trajectory evals The runner already parsed every tool call (with inputs) only to discard it after the live display. Retain it and feed it to trajectory evals. - ActResult gains `trajectory`; ClaudeCodeRunner.act() captures the streamed tool-call activity (now streams whenever capturing, so it works without a live `onActivity` display). - SkillVerifyInput gains `subject` + `bar`; runSkill frames the prompt by subject (judge the path vs the output) and states the rubric. - Engine keeps `lastTrajectory` from each act and hands trajectory evals the captured path (output evals still get the act summary). - Tests: a trajectory eval receives the trajectory + bar; an output eval receives the act summary. Runtime 77. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Wave 1 Stories 3-4: reflect sees the trajectory; trajectory-eval bar nudge Story 3 (feedback wiring): a failure's trajectory now flows into reflect, so `reflect on the path it took` can see the tool calls and fix the path (e.g. "you edited the test") rather than only the output. - ReflectInput gains `trajectory`; engine threads `lastTrajectory` through applyActions; the runner includes it in the reflect prompt. Story 4 (anti-thrash, safe part): a soft vscode lint nudges any trajectory eval that lacks `the bar:` — gating "done" on an LM judging a non-deterministic path without explicit pass conditions invites its own thrash. The full rigor-gated "without both [test and eval]" lint lands in Wave 2 Story 5, where `rigor:` supplies the severity (a blanket version would wrongly fire on every test-only loop). Tests: runtime 78, vscode 16. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Friendly surface: plain-English sugar + `loop-run explain` Make Loop approachable for non-experts while keeping all power — these are pure sugar/views over the same IR; the full grammar is unchanged. - Aliases (parser): `check:` / `verify:` = a `done when` (a bare value is a shell command, a predicate phrase is parsed as-is); `in:` / `look in:` / `files:` / `context:` = `look at:`; `when it breaks` = `when it fails`; `when it gets stuck` = `when blocked`. - `loop-run explain <file>` — describes a loop/pipeline/flow in plain English ("works toward… each round it… it's done when… it gives up after N tries"), so non-experts can trust what they wrote. A 6-line friendly loop now reads and explains cleanly. Tests: parser 38, runtime 79. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Docs: update the tutorial for evals, config defaults, and friendly syntax Bring the manual and grammar docs in line with the new features: - MANUAL.md: CLI table gains `show` / `explain` / `ls`; a "Tests vs evals" subsection (output vs trajectory, `the bar:`); a "Friendly shorthands" subsection (`check:`/`in:`/`when it breaks`) with a 6-line example; config tier documents the `each cycle:` default and the `loop.config` project file + the lowest-wins cascade. - AGENTS.md: `loop-run explain` and the friendly shorthands. Docs-only. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Wave 2: rigor dial (+ sensible defaults) and conductor/orchestrator mode Story 5 (rigor): a config-tier `rigor: vibe coding | structured ai-assisted | agentic engineering` dial. Under structured/agentic, every loop is born with a reflect-on-fail back-edge and a thrash guard unless it sets its own — the "sensible defaults" that remove boilerplate. `vibe coding` (and no rigor) injects nothing, so existing files are unchanged. The rigor-gated "without both, it is vibe coding" lint now fires (tests without an eval, or vice-versa) only when rigor opts in. Story 6 (mode): a config-tier `mode: conductor | orchestrator` naming the in-session vs async posture, with a lint warning on the costliest `vibe coding` + `orchestrator` quadrant. IR + schema + parser (threaded ParseDefaults) + show badges + cli project defaults + lints. Tests: parser 41, vscode 19. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Wave 3: hooks at lifecycle points; observe / OpEx report Story 7 (hooks): a `hooks:` block binds a deterministic check (a command or test) to a lifecycle point — before each cycle / after plan|act|observe / on commit / on push / on stop. A failing hook blocks the loop ("hooks block unsafe commits"). Parser + IR + schema + engine enforcement (before-cycle, per-step, on-commit, on-stop) + show rendering + tests. Story 8 (observe): a config-tier `observe:` block (trace every cycle / meter tokens and cost / stop and warn if cost exceeds "$N"). The CLI prints a stop-time OpEx report — cycles, reflects (back-edges), first-pass success, outcome — making "token burn from unverified loops" visible. summarizeOpex + formatOpexSummary + show + schema + tests. Tests: parser 44, runtime 82, vscode 19. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Wave 4: sandbox, knowledge/examples, parallel stages, standards/identity Story 9 (sandbox): a config-tier `sandbox:` block — no network / egress allowlist / cpu-memory-time caps — declaring run isolation as config. Story 10 (context): `examples:` (patterns to imitate) and `knowledge:` (read-only reference the agent must not edit) complete context engineering's six parts; both flow into the plan context. Story 11 (parallel stages): `stages in parallel:` groups stages that run concurrently (Promise.all, barrier-join, fail-fast). Real worktree isolation for file-safe parallel edits is the operator's setup; the orchestration is in the engine. Story 12 (standards/identity): `use tools from the "<server>"` names MCP servers; `runs as: <identity>` gives unattended runs an auditable principal. IR + schema + parser + engine + show + tests across all four. Tests: parser 47, runtime 83. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Examples in the new syntax + run .loop files as tests - examples/agentic/: 8 new examples showcasing the new features — evals (tests + output/trajectory + the bar), the friendly surface, the rigor dial (auto-injected defaults), hooks, observe/OpEx, sandbox, parallel stages, and context engineering's six parts. - examples/forge-sandbox.loop: isolation now declared via the `sandbox:` block + `runs as:` (was prose-only). - examples/feature/csv-export.loop: the build stage now pairs a test with a trajectory eval — "without both, it is vibe coding" in a real example. - packages/runtime/test/examples.test.js: parses every .loop in the repo and runs each standalone loop/pipeline through the engine — the framework's own "do all example shapes run?" check. All 57 .loop files parse/show/explain cleanly; full suite 159 tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Docs: tutorial coverage for the agentic-engineering features AGENTS.md vocabulary + a MANUAL.md "Agentic-engineering features" table covering rigor, mode, hooks, observe, sandbox, runs as, examples/knowledge, MCP tools, and parallel stages — pointing at examples/agentic/. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * Docs: remove references to the source article Scrub named citations and "the deck" references from docs, comments, example files, and the lint message — keeping the generic concept descriptions and the grammar keywords (rigor levels) intact. No behavior change; tests unchanged in count (159 green). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py * docs(site): expand the tutorial to cover the whole system Add a "The full system" section group to docs/index.html (the tutorial website), in its existing design language, covering every new construct with runnable examples: - Tests & evals (output/trajectory subjects, the bar) - Context & knowledge (look at/in, knowledge, examples, MCP tools) - Config tier & loop.config (with a concrete project-config file and the lowest-wins cascade) - Rigor, sensible defaults & mode - Hooks (lifecycle checkpoints) - Observability (OpEx report) & sandbox/identity - Friendly syntax & `loop-run explain` Extends the in-page syntax highlighter's keyword list for the new constructs and adds the nav links. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py --------- Co-authored-by: Claude <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Loop already covered triggers, verification, and human gates well, but two
building blocks of loop engineering were missing: coordinating named skills,
and remembering lessons across runs. Add both as first-class, optional knobs.
Language:
use skills: a, b— execution skills the loop may invoke during plan/actdone when the skill "x" approves/scores N or more— a review skill asverifier, bridging an abstract goal to a verifiable check
remember in "<file.md>"— read past lessons into the first plan, append adated outcome entry on stop (the across-run counterpart to
reflect)Wired end-to-end: loop-spec schema + TS mirror, parser, runtime engine (memory
read/write, skill-predicate routing, reflection capture), Claude Code runner
(runSkill + skills/memory in prompts), CLI events, VSCode highlight/hover/
completions, viz label, docs, examples, and tests (parser + runtime, all green).
Also documents the four-condition "when to build a loop" test and skill-driven
development in AGENTS.md.
Co-Authored-By: Claude Opus 4.8 (1M context) noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_01X74Z3fZXh7hiSZ8LmVzpGs