Skip to content

feat: add skills and cross-run memory to the Loop language - #1

Merged
tickets-forge-dev merged 2 commits into
masterfrom
claude/loop-engineering-guide-fwm4iu
Jun 25, 2026
Merged

tickets-forge-dev merged 2 commits into
masterfrom
claude/loop-engineering-guide-fwm4iu

Conversation

@tickets-forge-dev

Copy link
Copy Markdown
Owner

Loop already covered triggers, verification, and human gates well, but two
building blocks of loop engineering were missing: coordinating named skills,
and remembering lessons across runs. Add both as first-class, optional knobs.

Language:

  • use skills: a, b — execution skills the loop may invoke during plan/act
  • done when the skill "x" approves / scores N or more — a review skill as
    verifier, bridging an abstract goal to a verifiable check
  • remember in "<file.md>" — read past lessons into the first plan, append a
    dated outcome entry on stop (the across-run counterpart to reflect)

Wired end-to-end: loop-spec schema + TS mirror, parser, runtime engine (memory
read/write, skill-predicate routing, reflection capture), Claude Code runner
(runSkill + skills/memory in prompts), CLI events, VSCode highlight/hover/
completions, viz label, docs, examples, and tests (parser + runtime, all green).

Also documents the four-condition "when to build a loop" test and skill-driven
development in AGENTS.md.

Co-Authored-By: Claude Opus 4.8 (1M context) noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_01X74Z3fZXh7hiSZ8LmVzpGs

claude added 2 commits June 24, 2026 04:38
Add two optional, first-class knobs to the Loop language:

- `use skills: a, b` — named execution skills the loop may invoke during
  plan/act (coordinate proven skills instead of one mega-prompt).
- `done when the skill "x" approves` / `scores N or more` — a review skill
  as the verifier, bridging an abstract goal to a verifiable check.
- `remember in "<file.md>"` — cross-run memory: read past lessons into the
  first plan, append a dated outcome entry on stop.

Wired end-to-end: loop-spec schema + TS mirror, parser, runtime engine
(memory read/write, skill-predicate routing to runSkill, reflection capture),
Claude Code runner (runSkill + skills/memory in prompts, grants the Skill
tool, negation-aware parseSkillVerdict), MockRunner, CLI events, viz label,
VSCode highlight/hover/completions, docs, examples, and parser+runtime tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X74Z3fZXh7hiSZ8LmVzpGs
Add a self-contained animated SVG hero at the top of the README — a terminal
trace running examples/fix_test.loop: plan → act → observe (FAIL) → reflect →
plan → act → observe (PASS) → done. Lines default opaque, so it degrades to the
full static trace where CSS animation isn't honored. No build tooling or fonts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X74Z3fZXh7hiSZ8LmVzpGs
@tickets-forge-dev
tickets-forge-dev force-pushed the claude/loop-engineering-guide-fwm4iu branch from 68dd417 to ebbeda2 Compare June 24, 2026 04:41
@tickets-forge-dev
tickets-forge-dev merged commit 449c8e5 into master Jun 25, 2026
2 checks passed
tickets-forge-dev added a commit that referenced this pull request Jun 29, 2026
* Plan: agentic-engineering constructs for Loop, broken into stories

Translate "The New SDLC with Vibe Coding" into a Loop roadmap, in the
deck's own vocabulary. Adds a story-by-story implementation plan and a
dogfooded pipeline (today's grammar building tomorrow's constructs).

- docs/agentic-engineering-plan.md: the epic broken into 13 stories
  across 4 waves (eval core, the dials, harness completeness, niche/
  standards) plus a cross-cutting docs reframe; each story names files,
  a runnable `done when`, dependencies, and gates.
- agentic-engineering.loop: the same epic as a pipeline (epic->stages =
  stories), validated via `loop-run parse` with its flow printed.

Headline: make EVALS first-class (tests vs evals; output vs trajectory),
the gap every analysis lens ranked #1 — "without both, it is vibe coding".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Config defaults: config-tier `each cycle:` + project `loop.config` file

Make the language easier to use by removing repeated config. Two tiers,
both following the existing lowest-wins cascade (git/models):

1. Config-tier `each cycle:` — a file-level default cycle that every loop
   inherits; a per-loop `each cycle:` still overrides it. Resolved at parse
   time and threaded through pipelines/stages.

2. `loop.config` (or `.looprc`) — a per-project defaults file in the same
   config-tier syntax (each cycle / models / git). The runner walks up from
   the .loop file and folds it in as the lowest tier, so a file's own config
   and per-loop directives override it. `parse()` gains an optional
   `defaultCycle` so the project default seeds the cycle cascade.

Also: schema + IR (`Config.cycle`), `show` renders the file-level default,
AGENTS.md documents both, and agentic-engineering.loop dogfoods the cycle
hoist (13 repeated `each cycle:` lines collapsed to one config-tier line).

Tests: parser 32, runtime 75, vscode 14, stdlib/viz green (129 total).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Wave 1 Story 1: evals — tests vs evals, output vs trajectory, the bar

Make verification a conjunction and evals first-class, using the
deck's `done when … on the output / on the trajectory` syntax.

- IR + schema: `Loop.doneWhen` is now a `Predicate[]` (all must pass).
  The skill predicate gains `subject` ("output" default | "trajectory")
  and `bar` (an inline `the bar:` rubric).
- Parser: collects multiple `done when` lines; parses the
  `on the output`/`on the trajectory` qualifier; attaches an indented
  `the bar:` line to a skill eval (error if applied to a non-skill).
- Engine: observe evaluates the whole conjunction, short-circuiting on
  the first miss; routes skill predicates (evals) to runSkill, the rest
  to the shell verifier.
- Renderers: show + viz render multiple predicates and label evals/
  trajectory; show now renders skill evals (previously dropped).
- vscode linter, schema, and AGENTS.md updated; tests added.

Trajectory predicates currently receive the act summary as context;
Story 2 wires the real captured trajectory. Tests: parser 36, runtime
75, vscode 14, viz 5, stdlib 3.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Wave 1 Story 2: capture the trajectory and route it to trajectory evals

The runner already parsed every tool call (with inputs) only to discard
it after the live display. Retain it and feed it to trajectory evals.

- ActResult gains `trajectory`; ClaudeCodeRunner.act() captures the
  streamed tool-call activity (now streams whenever capturing, so it
  works without a live `onActivity` display).
- SkillVerifyInput gains `subject` + `bar`; runSkill frames the prompt
  by subject (judge the path vs the output) and states the rubric.
- Engine keeps `lastTrajectory` from each act and hands trajectory evals
  the captured path (output evals still get the act summary).
- Tests: a trajectory eval receives the trajectory + bar; an output eval
  receives the act summary. Runtime 77.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Wave 1 Stories 3-4: reflect sees the trajectory; trajectory-eval bar nudge

Story 3 (feedback wiring): a failure's trajectory now flows into reflect,
so `reflect on the path it took` can see the tool calls and fix the path
(e.g. "you edited the test") rather than only the output.
- ReflectInput gains `trajectory`; engine threads `lastTrajectory` through
  applyActions; the runner includes it in the reflect prompt.

Story 4 (anti-thrash, safe part): a soft vscode lint nudges any trajectory
eval that lacks `the bar:` — gating "done" on an LM judging a
non-deterministic path without explicit pass conditions invites its own
thrash. The full rigor-gated "without both [test and eval]" lint lands in
Wave 2 Story 5, where `rigor:` supplies the severity (a blanket version
would wrongly fire on every test-only loop).

Tests: runtime 78, vscode 16.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Friendly surface: plain-English sugar + `loop-run explain`

Make Loop approachable for non-experts while keeping all power — these
are pure sugar/views over the same IR; the full grammar is unchanged.

- Aliases (parser): `check:` / `verify:` = a `done when` (a bare value is
  a shell command, a predicate phrase is parsed as-is); `in:` / `look in:`
  / `files:` / `context:` = `look at:`; `when it breaks` = `when it fails`;
  `when it gets stuck` = `when blocked`.
- `loop-run explain <file>` — describes a loop/pipeline/flow in plain
  English ("works toward… each round it… it's done when… it gives up
  after N tries"), so non-experts can trust what they wrote.

A 6-line friendly loop now reads and explains cleanly. Tests: parser 38,
runtime 79.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Docs: update the tutorial for evals, config defaults, and friendly syntax

Bring the manual and grammar docs in line with the new features:

- MANUAL.md: CLI table gains `show` / `explain` / `ls`; a "Tests vs
  evals" subsection (output vs trajectory, `the bar:`); a "Friendly
  shorthands" subsection (`check:`/`in:`/`when it breaks`) with a 6-line
  example; config tier documents the `each cycle:` default and the
  `loop.config` project file + the lowest-wins cascade.
- AGENTS.md: `loop-run explain` and the friendly shorthands.

Docs-only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Wave 2: rigor dial (+ sensible defaults) and conductor/orchestrator mode

Story 5 (rigor): a config-tier `rigor: vibe coding | structured
ai-assisted | agentic engineering` dial. Under structured/agentic, every
loop is born with a reflect-on-fail back-edge and a thrash guard unless it
sets its own — the "sensible defaults" that remove boilerplate. `vibe
coding` (and no rigor) injects nothing, so existing files are unchanged.
The rigor-gated "without both, it is vibe coding" lint now fires (tests
without an eval, or vice-versa) only when rigor opts in.

Story 6 (mode): a config-tier `mode: conductor | orchestrator` naming the
in-session vs async posture, with a lint warning on the costliest
`vibe coding` + `orchestrator` quadrant.

IR + schema + parser (threaded ParseDefaults) + show badges + cli project
defaults + lints. Tests: parser 41, vscode 19.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Wave 3: hooks at lifecycle points; observe / OpEx report

Story 7 (hooks): a `hooks:` block binds a deterministic check (a command
or test) to a lifecycle point — before each cycle / after plan|act|observe
/ on commit / on push / on stop. A failing hook blocks the loop ("hooks
block unsafe commits"). Parser + IR + schema + engine enforcement
(before-cycle, per-step, on-commit, on-stop) + show rendering + tests.

Story 8 (observe): a config-tier `observe:` block (trace every cycle /
meter tokens and cost / stop and warn if cost exceeds "$N"). The CLI prints
a stop-time OpEx report — cycles, reflects (back-edges), first-pass
success, outcome — making "token burn from unverified loops" visible.
summarizeOpex + formatOpexSummary + show + schema + tests.

Tests: parser 44, runtime 82, vscode 19.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Wave 4: sandbox, knowledge/examples, parallel stages, standards/identity

Story 9 (sandbox): a config-tier `sandbox:` block — no network / egress
allowlist / cpu-memory-time caps — declaring run isolation as config.

Story 10 (context): `examples:` (patterns to imitate) and `knowledge:`
(read-only reference the agent must not edit) complete context
engineering's six parts; both flow into the plan context.

Story 11 (parallel stages): `stages in parallel:` groups stages that run
concurrently (Promise.all, barrier-join, fail-fast). Real worktree
isolation for file-safe parallel edits is the operator's setup; the
orchestration is in the engine.

Story 12 (standards/identity): `use tools from the "<server>"` names MCP
servers; `runs as: <identity>` gives unattended runs an auditable
principal.

IR + schema + parser + engine + show + tests across all four. Tests:
parser 47, runtime 83.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Examples in the new syntax + run .loop files as tests

- examples/agentic/: 8 new examples showcasing the new features — evals
  (tests + output/trajectory + the bar), the friendly surface, the rigor
  dial (auto-injected defaults), hooks, observe/OpEx, sandbox, parallel
  stages, and context engineering's six parts.
- examples/forge-sandbox.loop: isolation now declared via the `sandbox:`
  block + `runs as:` (was prose-only).
- examples/feature/csv-export.loop: the build stage now pairs a test with
  a trajectory eval — "without both, it is vibe coding" in a real example.
- packages/runtime/test/examples.test.js: parses every .loop in the repo
  and runs each standalone loop/pipeline through the engine — the
  framework's own "do all example shapes run?" check.

All 57 .loop files parse/show/explain cleanly; full suite 159 tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Docs: tutorial coverage for the agentic-engineering features

AGENTS.md vocabulary + a MANUAL.md "Agentic-engineering features" table
covering rigor, mode, hooks, observe, sandbox, runs as, examples/knowledge,
MCP tools, and parallel stages — pointing at examples/agentic/.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* Docs: remove references to the source article

Scrub named citations and "the deck" references from docs, comments,
example files, and the lint message — keeping the generic concept
descriptions and the grammar keywords (rigor levels) intact. No behavior
change; tests unchanged in count (159 green).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

* docs(site): expand the tutorial to cover the whole system

Add a "The full system" section group to docs/index.html (the tutorial
website), in its existing design language, covering every new construct
with runnable examples:

- Tests & evals (output/trajectory subjects, the bar)
- Context & knowledge (look at/in, knowledge, examples, MCP tools)
- Config tier & loop.config (with a concrete project-config file and the
  lowest-wins cascade)
- Rigor, sensible defaults & mode
- Hooks (lifecycle checkpoints)
- Observability (OpEx report) & sandbox/identity
- Friendly syntax & `loop-run explain`

Extends the in-page syntax highlighter's keyword list for the new
constructs and adds the nav links.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants