Skip to content

Plan: agentic-engineering constructs for Loop, broken into stories - #2

Merged
tickets-forge-dev merged 15 commits into
masterfrom
claude/loop-lang-concepts-3p5tz1
Jun 29, 2026
Merged

tickets-forge-dev merged 15 commits into
masterfrom
claude/loop-lang-concepts-3p5tz1

Conversation

@tickets-forge-dev

Copy link
Copy Markdown
Owner

Translate "The New SDLC with Vibe Coding" into a Loop roadmap, in the
deck's own vocabulary. Adds a story-by-story implementation plan and a
dogfooded pipeline (today's grammar building tomorrow's constructs).

  • docs/agentic-engineering-plan.md: the epic broken into 13 stories
    across 4 waves (eval core, the dials, harness completeness, niche/
    standards) plus a cross-cutting docs reframe; each story names files,
    a runnable done when, dependencies, and gates.
  • agentic-engineering.loop: the same epic as a pipeline (epic->stages =
    stories), validated via loop-run parse with its flow printed.

Headline: make EVALS first-class (tests vs evals; output vs trajectory),
the gap every analysis lens ranked #1 — "without both, it is vibe coding".

Co-Authored-By: Claude Opus 4.8 noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py

claude and others added 15 commits June 26, 2026 16:34
Translate "The New SDLC with Vibe Coding" into a Loop roadmap, in the
deck's own vocabulary. Adds a story-by-story implementation plan and a
dogfooded pipeline (today's grammar building tomorrow's constructs).

- docs/agentic-engineering-plan.md: the epic broken into 13 stories
  across 4 waves (eval core, the dials, harness completeness, niche/
  standards) plus a cross-cutting docs reframe; each story names files,
  a runnable `done when`, dependencies, and gates.
- agentic-engineering.loop: the same epic as a pipeline (epic->stages =
  stories), validated via `loop-run parse` with its flow printed.

Headline: make EVALS first-class (tests vs evals; output vs trajectory),
the gap every analysis lens ranked #1 — "without both, it is vibe coding".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Make the language easier to use by removing repeated config. Two tiers,
both following the existing lowest-wins cascade (git/models):

1. Config-tier `each cycle:` — a file-level default cycle that every loop
   inherits; a per-loop `each cycle:` still overrides it. Resolved at parse
   time and threaded through pipelines/stages.

2. `loop.config` (or `.looprc`) — a per-project defaults file in the same
   config-tier syntax (each cycle / models / git). The runner walks up from
   the .loop file and folds it in as the lowest tier, so a file's own config
   and per-loop directives override it. `parse()` gains an optional
   `defaultCycle` so the project default seeds the cycle cascade.

Also: schema + IR (`Config.cycle`), `show` renders the file-level default,
AGENTS.md documents both, and agentic-engineering.loop dogfoods the cycle
hoist (13 repeated `each cycle:` lines collapsed to one config-tier line).

Tests: parser 32, runtime 75, vscode 14, stdlib/viz green (129 total).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Make verification a conjunction and evals first-class, using the
deck's `done when … on the output / on the trajectory` syntax.

- IR + schema: `Loop.doneWhen` is now a `Predicate[]` (all must pass).
  The skill predicate gains `subject` ("output" default | "trajectory")
  and `bar` (an inline `the bar:` rubric).
- Parser: collects multiple `done when` lines; parses the
  `on the output`/`on the trajectory` qualifier; attaches an indented
  `the bar:` line to a skill eval (error if applied to a non-skill).
- Engine: observe evaluates the whole conjunction, short-circuiting on
  the first miss; routes skill predicates (evals) to runSkill, the rest
  to the shell verifier.
- Renderers: show + viz render multiple predicates and label evals/
  trajectory; show now renders skill evals (previously dropped).
- vscode linter, schema, and AGENTS.md updated; tests added.

Trajectory predicates currently receive the act summary as context;
Story 2 wires the real captured trajectory. Tests: parser 36, runtime
75, vscode 14, viz 5, stdlib 3.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
The runner already parsed every tool call (with inputs) only to discard
it after the live display. Retain it and feed it to trajectory evals.

- ActResult gains `trajectory`; ClaudeCodeRunner.act() captures the
  streamed tool-call activity (now streams whenever capturing, so it
  works without a live `onActivity` display).
- SkillVerifyInput gains `subject` + `bar`; runSkill frames the prompt
  by subject (judge the path vs the output) and states the rubric.
- Engine keeps `lastTrajectory` from each act and hands trajectory evals
  the captured path (output evals still get the act summary).
- Tests: a trajectory eval receives the trajectory + bar; an output eval
  receives the act summary. Runtime 77.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
…nudge

Story 3 (feedback wiring): a failure's trajectory now flows into reflect,
so `reflect on the path it took` can see the tool calls and fix the path
(e.g. "you edited the test") rather than only the output.
- ReflectInput gains `trajectory`; engine threads `lastTrajectory` through
  applyActions; the runner includes it in the reflect prompt.

Story 4 (anti-thrash, safe part): a soft vscode lint nudges any trajectory
eval that lacks `the bar:` — gating "done" on an LM judging a
non-deterministic path without explicit pass conditions invites its own
thrash. The full rigor-gated "without both [test and eval]" lint lands in
Wave 2 Story 5, where `rigor:` supplies the severity (a blanket version
would wrongly fire on every test-only loop).

Tests: runtime 78, vscode 16.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Make Loop approachable for non-experts while keeping all power — these
are pure sugar/views over the same IR; the full grammar is unchanged.

- Aliases (parser): `check:` / `verify:` = a `done when` (a bare value is
  a shell command, a predicate phrase is parsed as-is); `in:` / `look in:`
  / `files:` / `context:` = `look at:`; `when it breaks` = `when it fails`;
  `when it gets stuck` = `when blocked`.
- `loop-run explain <file>` — describes a loop/pipeline/flow in plain
  English ("works toward… each round it… it's done when… it gives up
  after N tries"), so non-experts can trust what they wrote.

A 6-line friendly loop now reads and explains cleanly. Tests: parser 38,
runtime 79.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
…ntax

Bring the manual and grammar docs in line with the new features:

- MANUAL.md: CLI table gains `show` / `explain` / `ls`; a "Tests vs
  evals" subsection (output vs trajectory, `the bar:`); a "Friendly
  shorthands" subsection (`check:`/`in:`/`when it breaks`) with a 6-line
  example; config tier documents the `each cycle:` default and the
  `loop.config` project file + the lowest-wins cascade.
- AGENTS.md: `loop-run explain` and the friendly shorthands.

Docs-only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Story 5 (rigor): a config-tier `rigor: vibe coding | structured
ai-assisted | agentic engineering` dial. Under structured/agentic, every
loop is born with a reflect-on-fail back-edge and a thrash guard unless it
sets its own — the "sensible defaults" that remove boilerplate. `vibe
coding` (and no rigor) injects nothing, so existing files are unchanged.
The rigor-gated "without both, it is vibe coding" lint now fires (tests
without an eval, or vice-versa) only when rigor opts in.

Story 6 (mode): a config-tier `mode: conductor | orchestrator` naming the
in-session vs async posture, with a lint warning on the costliest
`vibe coding` + `orchestrator` quadrant.

IR + schema + parser (threaded ParseDefaults) + show badges + cli project
defaults + lints. Tests: parser 41, vscode 19.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Story 7 (hooks): a `hooks:` block binds a deterministic check (a command
or test) to a lifecycle point — before each cycle / after plan|act|observe
/ on commit / on push / on stop. A failing hook blocks the loop ("hooks
block unsafe commits"). Parser + IR + schema + engine enforcement
(before-cycle, per-step, on-commit, on-stop) + show rendering + tests.

Story 8 (observe): a config-tier `observe:` block (trace every cycle /
meter tokens and cost / stop and warn if cost exceeds "$N"). The CLI prints
a stop-time OpEx report — cycles, reflects (back-edges), first-pass
success, outcome — making "token burn from unverified loops" visible.
summarizeOpex + formatOpexSummary + show + schema + tests.

Tests: parser 44, runtime 82, vscode 19.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Story 9 (sandbox): a config-tier `sandbox:` block — no network / egress
allowlist / cpu-memory-time caps — declaring run isolation as config.

Story 10 (context): `examples:` (patterns to imitate) and `knowledge:`
(read-only reference the agent must not edit) complete context
engineering's six parts; both flow into the plan context.

Story 11 (parallel stages): `stages in parallel:` groups stages that run
concurrently (Promise.all, barrier-join, fail-fast). Real worktree
isolation for file-safe parallel edits is the operator's setup; the
orchestration is in the engine.

Story 12 (standards/identity): `use tools from the "<server>"` names MCP
servers; `runs as: <identity>` gives unattended runs an auditable
principal.

IR + schema + parser + engine + show + tests across all four. Tests:
parser 47, runtime 83.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
- examples/agentic/: 8 new examples showcasing the new features — evals
  (tests + output/trajectory + the bar), the friendly surface, the rigor
  dial (auto-injected defaults), hooks, observe/OpEx, sandbox, parallel
  stages, and context engineering's six parts.
- examples/forge-sandbox.loop: isolation now declared via the `sandbox:`
  block + `runs as:` (was prose-only).
- examples/feature/csv-export.loop: the build stage now pairs a test with
  a trajectory eval — "without both, it is vibe coding" in a real example.
- packages/runtime/test/examples.test.js: parses every .loop in the repo
  and runs each standalone loop/pipeline through the engine — the
  framework's own "do all example shapes run?" check.

All 57 .loop files parse/show/explain cleanly; full suite 159 tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
AGENTS.md vocabulary + a MANUAL.md "Agentic-engineering features" table
covering rigor, mode, hooks, observe, sandbox, runs as, examples/knowledge,
MCP tools, and parallel stages — pointing at examples/agentic/.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Scrub named citations and "the deck" references from docs, comments,
example files, and the lint message — keeping the generic concept
descriptions and the grammar keywords (rigor levels) intact. No behavior
change; tests unchanged in count (159 green).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
Add a "The full system" section group to docs/index.html (the tutorial
website), in its existing design language, covering every new construct
with runnable examples:

- Tests & evals (output/trajectory subjects, the bar)
- Context & knowledge (look at/in, knowledge, examples, MCP tools)
- Config tier & loop.config (with a concrete project-config file and the
  lowest-wins cascade)
- Rigor, sensible defaults & mode
- Hooks (lifecycle checkpoints)
- Observability (OpEx report) & sandbox/identity
- Friendly syntax & `loop-run explain`

Extends the in-page syntax highlighter's keyword list for the new
constructs and adds the nav links.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WX3xSGKczGFix6RYrQA7Py
# Conflicts:
#	docs/MANUAL.md
#	packages/runtime/src/cli.ts
@tickets-forge-dev
tickets-forge-dev merged commit 6b52431 into master Jun 29, 2026
2 checks passed
@tickets-forge-dev
tickets-forge-dev deleted the claude/loop-lang-concepts-3p5tz1 branch June 29, 2026 14:32
tickets-forge-dev added a commit that referenced this pull request Jun 29, 2026
Bump for the npm release that ships the agentic-engineering constructs (#2):
- @loop-lang/{parser,runtime,stdlib,viz} 0.2.0 → 0.3.0 (parser/runtime/viz carry
  the new constructs; stdlib lockstep so the workspace publish has no duplicate).
- @loop-lang/loop (installer) 0.5.0 → 0.6.0.
- Sync inter-package dep pins + loop-vscode's parser devDep to 0.3.0 so the
  extension links/bundles the current parser (knows the new constructs).
- CHANGELOG 0.6.0 entry.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants