Repository navigation
Plan: agentic-engineering constructs for Loop, broken into stories - #2
Merged
Merged
Conversation
Translate "The New SDLC with Vibe Coding" into a Loop roadmap, in the deck's own vocabulary. Adds a story-by-story implementation plan and a dogfooded pipeline (today's grammar building tomorrow's constructs). - docs/agentic-engineering-plan.md: the epic broken into 13 stories across 4 waves (eval core, the dials, harness completeness, niche/ standards) plus a cross-cutting docs reframe; each story names files, a runnable `done when`, dependencies, and gates. - agentic-engineering.loop: the same epic as a pipeline (epic->stages = stories), validated via `loop-run parse` with its flow printed. Headline: make EVALS first-class (tests vs evals; output vs trajectory), the gap every analysis lens ranked #1 — "without both, it is vibe coding". Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Make the language easier to use by removing repeated config. Two tiers, both following the existing lowest-wins cascade (git/models): 1. Config-tier `each cycle:` — a file-level default cycle that every loop inherits; a per-loop `each cycle:` still overrides it. Resolved at parse time and threaded through pipelines/stages. 2. `loop.config` (or `.looprc`) — a per-project defaults file in the same config-tier syntax (each cycle / models / git). The runner walks up from the .loop file and folds it in as the lowest tier, so a file's own config and per-loop directives override it. `parse()` gains an optional `defaultCycle` so the project default seeds the cycle cascade. Also: schema + IR (`Config.cycle`), `show` renders the file-level default, AGENTS.md documents both, and agentic-engineering.loop dogfoods the cycle hoist (13 repeated `each cycle:` lines collapsed to one config-tier line). Tests: parser 32, runtime 75, vscode 14, stdlib/viz green (129 total). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Make verification a conjunction and evals first-class, using the
deck's `done when … on the output / on the trajectory` syntax.
- IR + schema: `Loop.doneWhen` is now a `Predicate[]` (all must pass).
The skill predicate gains `subject` ("output" default | "trajectory")
and `bar` (an inline `the bar:` rubric).
- Parser: collects multiple `done when` lines; parses the
`on the output`/`on the trajectory` qualifier; attaches an indented
`the bar:` line to a skill eval (error if applied to a non-skill).
- Engine: observe evaluates the whole conjunction, short-circuiting on
the first miss; routes skill predicates (evals) to runSkill, the rest
to the shell verifier.
- Renderers: show + viz render multiple predicates and label evals/
trajectory; show now renders skill evals (previously dropped).
- vscode linter, schema, and AGENTS.md updated; tests added.
Trajectory predicates currently receive the act summary as context;
Story 2 wires the real captured trajectory. Tests: parser 36, runtime
75, vscode 14, viz 5, stdlib 3.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
The runner already parsed every tool call (with inputs) only to discard it after the live display. Retain it and feed it to trajectory evals. - ActResult gains `trajectory`; ClaudeCodeRunner.act() captures the streamed tool-call activity (now streams whenever capturing, so it works without a live `onActivity` display). - SkillVerifyInput gains `subject` + `bar`; runSkill frames the prompt by subject (judge the path vs the output) and states the rubric. - Engine keeps `lastTrajectory` from each act and hands trajectory evals the captured path (output evals still get the act summary). - Tests: a trajectory eval receives the trajectory + bar; an output eval receives the act summary. Runtime 77. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
…nudge Story 3 (feedback wiring): a failure's trajectory now flows into reflect, so `reflect on the path it took` can see the tool calls and fix the path (e.g. "you edited the test") rather than only the output. - ReflectInput gains `trajectory`; engine threads `lastTrajectory` through applyActions; the runner includes it in the reflect prompt. Story 4 (anti-thrash, safe part): a soft vscode lint nudges any trajectory eval that lacks `the bar:` — gating "done" on an LM judging a non-deterministic path without explicit pass conditions invites its own thrash. The full rigor-gated "without both [test and eval]" lint lands in Wave 2 Story 5, where `rigor:` supplies the severity (a blanket version would wrongly fire on every test-only loop). Tests: runtime 78, vscode 16. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Make Loop approachable for non-experts while keeping all power — these
are pure sugar/views over the same IR; the full grammar is unchanged.
- Aliases (parser): `check:` / `verify:` = a `done when` (a bare value is
a shell command, a predicate phrase is parsed as-is); `in:` / `look in:`
/ `files:` / `context:` = `look at:`; `when it breaks` = `when it fails`;
`when it gets stuck` = `when blocked`.
- `loop-run explain <file>` — describes a loop/pipeline/flow in plain
English ("works toward… each round it… it's done when… it gives up
after N tries"), so non-experts can trust what they wrote.
A 6-line friendly loop now reads and explains cleanly. Tests: parser 38,
runtime 79.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
…ntax Bring the manual and grammar docs in line with the new features: - MANUAL.md: CLI table gains `show` / `explain` / `ls`; a "Tests vs evals" subsection (output vs trajectory, `the bar:`); a "Friendly shorthands" subsection (`check:`/`in:`/`when it breaks`) with a 6-line example; config tier documents the `each cycle:` default and the `loop.config` project file + the lowest-wins cascade. - AGENTS.md: `loop-run explain` and the friendly shorthands. Docs-only. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Story 5 (rigor): a config-tier `rigor: vibe coding | structured ai-assisted | agentic engineering` dial. Under structured/agentic, every loop is born with a reflect-on-fail back-edge and a thrash guard unless it sets its own — the "sensible defaults" that remove boilerplate. `vibe coding` (and no rigor) injects nothing, so existing files are unchanged. The rigor-gated "without both, it is vibe coding" lint now fires (tests without an eval, or vice-versa) only when rigor opts in. Story 6 (mode): a config-tier `mode: conductor | orchestrator` naming the in-session vs async posture, with a lint warning on the costliest `vibe coding` + `orchestrator` quadrant. IR + schema + parser (threaded ParseDefaults) + show badges + cli project defaults + lints. Tests: parser 41, vscode 19. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Story 7 (hooks): a `hooks:` block binds a deterministic check (a command
or test) to a lifecycle point — before each cycle / after plan|act|observe
/ on commit / on push / on stop. A failing hook blocks the loop ("hooks
block unsafe commits"). Parser + IR + schema + engine enforcement
(before-cycle, per-step, on-commit, on-stop) + show rendering + tests.
Story 8 (observe): a config-tier `observe:` block (trace every cycle /
meter tokens and cost / stop and warn if cost exceeds "$N"). The CLI prints
a stop-time OpEx report — cycles, reflects (back-edges), first-pass
success, outcome — making "token burn from unverified loops" visible.
summarizeOpex + formatOpexSummary + show + schema + tests.
Tests: parser 44, runtime 82, vscode 19.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Story 9 (sandbox): a config-tier `sandbox:` block — no network / egress allowlist / cpu-memory-time caps — declaring run isolation as config. Story 10 (context): `examples:` (patterns to imitate) and `knowledge:` (read-only reference the agent must not edit) complete context engineering's six parts; both flow into the plan context. Story 11 (parallel stages): `stages in parallel:` groups stages that run concurrently (Promise.all, barrier-join, fail-fast). Real worktree isolation for file-safe parallel edits is the operator's setup; the orchestration is in the engine. Story 12 (standards/identity): `use tools from the "<server>"` names MCP servers; `runs as: <identity>` gives unattended runs an auditable principal. IR + schema + parser + engine + show + tests across all four. Tests: parser 47, runtime 83. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
- examples/agentic/: 8 new examples showcasing the new features — evals (tests + output/trajectory + the bar), the friendly surface, the rigor dial (auto-injected defaults), hooks, observe/OpEx, sandbox, parallel stages, and context engineering's six parts. - examples/forge-sandbox.loop: isolation now declared via the `sandbox:` block + `runs as:` (was prose-only). - examples/feature/csv-export.loop: the build stage now pairs a test with a trajectory eval — "without both, it is vibe coding" in a real example. - packages/runtime/test/examples.test.js: parses every .loop in the repo and runs each standalone loop/pipeline through the engine — the framework's own "do all example shapes run?" check. All 57 .loop files parse/show/explain cleanly; full suite 159 tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
AGENTS.md vocabulary + a MANUAL.md "Agentic-engineering features" table covering rigor, mode, hooks, observe, sandbox, runs as, examples/knowledge, MCP tools, and parallel stages — pointing at examples/agentic/. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Scrub named citations and "the deck" references from docs, comments, example files, and the lint message — keeping the generic concept descriptions and the grammar keywords (rigor levels) intact. No behavior change; tests unchanged in count (159 green). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Add a "The full system" section group to docs/index.html (the tutorial website), in its existing design language, covering every new construct with runnable examples: - Tests & evals (output/trajectory subjects, the bar) - Context & knowledge (look at/in, knowledge, examples, MCP tools) - Config tier & loop.config (with a concrete project-config file and the lowest-wins cascade) - Rigor, sensible defaults & mode - Hooks (lifecycle checkpoints) - Observability (OpEx report) & sandbox/identity - Friendly syntax & `loop-run explain` Extends the in-page syntax highlighter's keyword list for the new constructs and adds the nav links. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
# Conflicts: # docs/MANUAL.md # packages/runtime/src/cli.ts
tickets-forge-dev
added a commit
that referenced
this pull request
Jun 29, 2026
Bump for the npm release that ships the agentic-engineering constructs (#2): - @loop-lang/{parser,runtime,stdlib,viz} 0.2.0 → 0.3.0 (parser/runtime/viz carry the new constructs; stdlib lockstep so the workspace publish has no duplicate). - @loop-lang/loop (installer) 0.5.0 → 0.6.0. - Sync inter-package dep pins + loop-vscode's parser devDep to 0.3.0 so the extension links/bundles the current parser (knows the new constructs). - CHANGELOG 0.6.0 entry. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Translate "The New SDLC with Vibe Coding" into a Loop roadmap, in the
deck's own vocabulary. Adds a story-by-story implementation plan and a
dogfooded pipeline (today's grammar building tomorrow's constructs).
across 4 waves (eval core, the dials, harness completeness, niche/
standards) plus a cross-cutting docs reframe; each story names files,
a runnable
done when, dependencies, and gates.stories), validated via
loop-run parsewith its flow printed.Headline: make EVALS first-class (tests vs evals; output vs trajectory),
the gap every analysis lens ranked #1 — "without both, it is vibe coding".
Co-Authored-By: Claude Opus 4.8 noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py