Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
14 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions .claude/skills/loopflow/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,8 @@ schedule: <when> run unattended on a cadence
runner: <agent> which agent executes the loop
target: <dir> operate on another directory/repo
recommend skills with ctx ctx is this file's skill source — recommends + installs skills per loop goal
grant ctx: skills, agents, mcps, harnesses capability groups ctx may recommend (fail-closed; default skills+agents; mcps/harnesses are recommend-only)
ctx may use my own model "<provider>/<model>" declares a user-owned model — unlocks harness recommendations (dry-run only)
```

Predicates:
Expand Down Expand Up @@ -178,6 +180,13 @@ When ctx's tools are available (`ctx__loop_provision`, `ctx__recommend_bundle`):
first keeps the `.loop` self-contained and reproducible; the second lets a headless
`loop run` re-resolve the bundle from ctx.
3. Offer `top up skills from ctx` if the loop should pull more skills when a cycle fails.
4. **Beyond skills** — if the goal needs more than skills, add a `grant ctx: skills, agents,
mcps, harnesses` line for the groups that apply (fail-closed; omit it for skills-only).
`mcps` and `harnesses` are **recommend-only** — ctx surfaces them with an install command
the user runs; the loop never auto-installs them. Harnesses additionally need a
`ctx may use my own model "<provider>/<model>"` line, and always come as a `--dry-run`
command. Pass the granted groups (and own-model) to `ctx__loop_provision` as `permissions` /
`own_llm` / `model_provider` / `model`.

When ctx is **not** attached, skip this silently and author `use skills:` by hand as usual —
the loop runs the same either way.
Expand Down
46 changes: 45 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -99,6 +99,8 @@ rigor: vibe coding | structured ai-assisted | agentic engineering (the spectru
mode: conductor | orchestrator (supervision posture: in-session/sync vs async/opens-a-PR)
runs as: <identity> (an auditable principal for unattended runs)
recommend skills with ctx (config tier: ctx is this file's skill source — recommends + installs skills per loop goal; see "Skill source: ctx" below)
grant ctx: skills, agents, mcps, harnesses (config tier: capability groups the file lets ctx recommend; fails closed, default skills+agents; mcps/harnesses are recommend-only)
ctx may use my own model "<provider>/<model>" (config tier: declares a user-owned/local/API model — unlocks ctx harness recommendations, always dry-run)
observe: (block) trace every cycle / meter tokens and cost / stop and warn if cost exceeds "$N"
sandbox: (block) no network access / allow egress to "host" only / cap cpu at … memory at … time at …
hooks: (loop body block) before each cycle | after act | on commit | on stop : "<cmd>" passes|finds nothing (a failing hook blocks)
Expand All @@ -114,9 +116,11 @@ stages in parallel: (inside a pipeline: the indented stages run concurrently)
done when the test "billing.spec.ts::apostrophe" passes # a named test
done when "pnpm test" passes # a shell command, exit 0
done when "semgrep --severity=high" finds nothing # a shell command, empty output
done when "pnpm test flaky" passes 3 times # flake guard: re-run the check, EVERY run must pass
done when a human confirms "looks right at 375px" # a human check
done when the skill "email-review" approves # an eval: approved / not
done when the skill "email-review" scores 8 or more # an eval: numeric threshold
done when the skill "code-review" approves by 3 judges # consensus: N independent verdicts, majority wins
done when the skill "api-review" scores 8 or more on the output # an eval of WHAT was produced
done when the skill "path-review" approves on the trajectory # an eval of HOW the agent got there
the bar: didn't weaken a test to go green; no writes outside api/ # the rubric the judge scores against
Expand All @@ -125,6 +129,16 @@ done when the skill "path-review" approves on the trajectory # an eval
The command in a predicate runs in the user's shell with their privileges (like an npm
script). It IS meant to be a real command. Prefer a fast, deterministic check.

**Flake guard — `passes N times`.** Append `N times` to a `test` or command predicate to re-run
it `N` times and require every run to pass (the first failure short-circuits). Reach for it when a
green can pass by luck — a timing- or order-dependent test — so "done" means "passes *reliably*",
not "passed *once*".

**Judge panel — `by N judges`.** Append `by N judges` to a skill predicate to collect `N`
independent verdicts and take the majority (early-exit once decided). A single LM judge wobbles
near the bar; consensus smooths the noise. The deterministic counterpart of the flake guard:
flake guard for tests, judge panel for evals.

### Tests vs evals — list as many `done when` as you need

A loop may have **multiple `done when` lines, and ALL must pass** (a conjunction). Use this to
Expand Down Expand Up @@ -186,8 +200,33 @@ loop "harden the stripe webhook handler":
MCP server before the first plan, and `top up skills from ctx` after a failed cycle reflects.
- **No ctx attached?** The lines are inert — the loop runs exactly as it would without them.

**Beyond skills — the full capability set.** By default ctx provisions only `skills`
(and the agents Loop loads the same way). A `grant ctx:` line widens what ctx may recommend to
any of `skills, agents, mcps, harnesses`, **failing closed** — only listed groups are returned:

```loop
recommend skills with ctx
grant ctx: skills, agents, mcps, harnesses # capability grants (fail-closed)
ctx may use my own model "ollama/llama3.1" # unlocks harness recs (dry-run only)

loop "stand up a local agent loop":
goal: an MCP agent loop running on local ollama with filesystem access
use skills recommended by ctx
done when "pytest tests/agent_loop" passes
```

- **skills / agents** install into `~/.claude/skills` (as before) and merge into the cycle's
skill set.
- **mcps** are **recommend-only**: ctx surfaces fitting MCP servers + a suggested
`ctx-mcp-install <name>`; the runtime emits them on a `ctx` event, it never auto-registers one.
- **harnesses** (autogen, langfuse, …) recommend **only** when you declare a user-owned model
(`ctx may use my own model "…"`), and ship as an explicit `ctx-harness-install <name> --dry-run`
command — never an automatic install. This is the one capability that pulls real software, so it
stays human-gated by design.

Setup: `claude mcp add ctx -- ctx-mcp-server` (needs `pip install claude-ctx`). See
`examples/ctx_skills.loop` and `docs/ctx-skill-source.md`.
`examples/ctx_skills.loop`, `examples/ctx_capabilities.loop`, and `docs/ctx-skill-source.md`.
Full customer-facing walkthrough (setup, own-model, the capability model): `docs/ctx-integration-guide.md`.

### `remember in` — cross-run memory

Expand Down Expand Up @@ -411,6 +450,11 @@ flow, show the file chain. `loop-run ls` lists every loop in the repo.

- `loop-run run file.loop` — execute it on Claude Code (plan/act/observe, reflect on failure,
verify with `done when`, pause at human gates).
- `loop-run run file.loop --log run.log` — also append every event to a local NDJSON log
(secrets are scrubbed before anything is persisted).
- `loop-run run file.loop --resume run.log` — resume an interrupted run from its log: satisfied
stages / flow steps / for-each items are skipped, the first incomplete unit picks up (flow
carry-forward summaries restored from the log).
- `loop-run show file.loop` — print the loop's flow as compact ASCII (and `loop-run ls` to list them).
- `loop-run explain file.loop` — describe the loop in plain English (a friendly check of what it will do).
- `loop-run viz file.loop` — open a visual HTML schematic of the flow.
Expand Down
36 changes: 36 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,42 @@ Versions track the `@loop-lang/loop` installer package.

## [Unreleased]

## [0.7.0] — 2026-07-02

> `@loop-lang/loop` 0.7.0 · `@loop-lang/{parser,runtime,stdlib,viz}` 0.4.0 · `loopflow` (vscode) 0.5.0

### Added
- **ctx as a skill source** — `recommend skills with ctx` / `use skills recommended by ctx`
/ `top up skills from ctx`: a loop equips itself via the ctx MCP server before the first
plan and re-equips after a failed cycle reflects. Capability grants
(`grant ctx: skills, agents, mcps, harnesses`, fail-closed) and own-model gating
(`ctx may use my own model "…"`, dry-run-only harness recommendations).
- **Verification reliability** — the flake guard (`done when "…" passes 3 times`: every
run must pass, first failure short-circuits) and judge panels
(`the skill "…" approves by 3 judges`: majority of independent verdicts, early-exit
once decided). Rendered in `show`/`explain` (`×3`, `· 3 judges`).
- **Event log & telemetry** — `--log <path>` / `LOOP_LOG_FILE` appends every runtime
event as durable NDJSON (header + seq'd lines); `LOOP_EVENTS_URL` streams the same
events to a control-plane collector (shared run id, fan-out). **Secret redaction on by
default**: env-derived values and well-known credential shapes are scrubbed before any
sink persists an event (`LOOP_REDACT=off` to disable).
- **Resume** — `loop-run run <file> --resume run.log` skips every unit the log proves
satisfied (definitions, stages, flow steps, for-each items), restores flow carry-forward
summaries, and warns when the `.loop` source changed since the logged run.
- **Browser playground** — `docs/playground.html`: the parser, ASCII shape view, explain,
and soft linter bundled to 27 kB of client-side JS; parse-on-type with inline errors and
example loops. Linked from the tutorial.
- **Docs** — "How verification works: what 'done' actually depends on" in the manual
(verdict factors: working dir, shell env, exit codes, `finds nothing` semantics, flake /
judge hardening); event-log, resume, and redaction sections; README aligned with the
tutorial (all `.loop` blocks verified against the parser).

### Changed
- **New logo** — the gap ring (one ring, one gap: the loop still iterating), with a
solid-tile variant as the favicon / app icon across the site and README.
- `loopflow` (vscode) 0.5.0 — bundles the 0.4.0 parser (judge panels, flake guard, ctx
lines all recognized); output panel renders `⏩ resumed` events.

## [0.6.0] — 2026-06-29

> `@loop-lang/loop` 0.6.0 · `@loop-lang/{parser,runtime,stdlib,viz}` 0.3.0 · `loop-vscode` 0.4.0
Expand Down
75 changes: 63 additions & 12 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@

<p align="center"><b>An open, natural-language DSL for loop engineering.</b><br/>Describe a staged, self-correcting, human-gated agent workflow in plain English, press ▶, and it runs on Claude Code.</p>

<p align="center"><i>Stop tuning prompts. Start editing the loop.</i></p>
<p align="center"><i>Stop babysitting the agent. Write the goal once — the loop plans, acts, reflects on red,<br/>and stops only when the check is green, at the gates you set.</i></p>

<p align="center"><img src="docs/assets/loop-demo.svg" alt="A Loop turning a failing test green: plan → act → observe (FAIL) → reflect → plan → act → observe (PASS) → done" width="760"></p>

Expand All @@ -24,9 +24,12 @@

<p align="center">
<a href="https://loopflow.live">Tutorial</a> ·
<a href="https://loopflow.live/workshop.html">Workshop</a> ·
<a href="https://loopflow.live/playground.html">⚡ Playground</a> ·
<a href="https://loopflow.live/workshop.html">🛠️ Workshop</a> ·
<a href="https://loopflow.live/game.html">🎮 Lab</a> ·
<a href="https://github.com/tickets-forge-dev/loop-lang/blob/master/docs/MANUAL.md">Manual</a> ·
<a href="https://loopflow.live/keywords/index.html">Keyword reference</a>
<a href="https://loopflow.live/keywords/index.html">Keywords</a> ·
<a href="docs/FAQ.md">FAQ</a>
</p>

## Quickstart
Expand Down Expand Up @@ -88,24 +91,48 @@ Compose loops into **stages** and **pipelines**, with humans wired in where judg

```loop
pipeline "ship feature":
stage security:

stage "security":
goal: no high or critical vulnerabilities
done when "semgrep --severity=high" finds nothing
each cycle: plan, act, observe
each cycle: plan, then act, then observe
when it fails: reflect, then plan again

stage build:
stage "build":
goal: feature works and tests pass
a human approves the plan first
then each cycle: act, observe
done when "pnpm test" passes
each cycle: act, then observe

stage ui:
stage "ui":
goal: matches design, responsive at 375px
each cycle: plan, act, observe
each cycle: plan, then act, then observe
a human reviews before stopping
```

## Verify like you mean it

`done when` is the loop's definition of reality — so LoopFlow gives verification real teeth.
List several checks (**all must pass**), mix deterministic tests with LM-judged evals, and
harden both sides against false greens:

```loop
loop "harden checkout":
goal: checkout works, reliably, and was built the right way
done when "pnpm test checkout" passes 3 times # flake guard: every run must pass
done when the skill "code-review" approves by 3 judges # judge panel: majority of 3 verdicts
done when the skill "path-review" approves on the trajectory # judges HOW it got there
the bar: didn't weaken a test to go green; no writes outside src/checkout
```

- **Flake guard** — `passes N times` re-runs a test/command; one lucky green isn't "done".
- **Judge panel** — `by N judges` takes a majority of independent verdicts; one wobbly LM
judgment isn't "done" either.
- **Trajectory evals** — catch what a green test can't: an agent that gamed the check.

Full mechanics — what a verdict is actually affected by (working dir, shell env, exit codes) —
in [How verification works](docs/MANUAL.md#how-verification-works--what-done-actually-depends-on).

## Skills and memory

Two knobs make a loop coordinate proven work and learn over time:
Expand All @@ -127,6 +154,12 @@ loop "decide whether to cancel the morning run":
first plan and appends an outcome entry when it stops. `reflect` is within-run memory;
`remember` is its across-run counterpart. See [`examples/skills_memory.loop`](examples/skills_memory.loop).

And a loop can **equip itself**: with [ctx](https://github.com/stevesolun/ctx) attached as the
skill source, `use skills recommended by ctx` resolves + installs the right skill bundle for the
goal before the first plan, and `top up skills from ctx` pulls more after a failed cycle
reflects. Opt-in, fail-closed, inert without ctx — see
[the integration guide](docs/MANUAL.md) and [`examples/ctx_capabilities.loop`](examples/ctx_capabilities.loop).

## Compose loops

Compose loops into **pipelines** (stages in order, fail-fast), chain whole files with **`flow`**, and fan out over a plan with **`for each`** — humans wired in where judgment lives. Full grammar with worked examples: the [tutorial](https://loopflow.live) and the [manual](docs/MANUAL.md).
Expand Down Expand Up @@ -192,6 +225,19 @@ When you run a loop via `/loopflow`, the skill asks if you want the dashboard an
opens it and updates it as each step happens — pipeline stages, flow steps, and sprint stories
filling in as the loop progresses.

Prefer a file you can grep later? Persist the same event stream as NDJSON — and use it to
**resume** an interrupted run:

```
loop-run run file.loop --log run.log # append every event to a local log (secrets scrubbed)
loop-run run file.loop --resume run.log # skip what the log proves done; pick up where it died
```

See **Event log & telemetry** in [`docs/MANUAL.md`](docs/MANUAL.md) for the format, redaction,
resume semantics, and the `LOOP_EVENTS_URL` remote collector. Want to *feel* the language first?
Open the [**browser playground**](https://loopflow.live/playground.html) — type a `.loop`, see its
shape live, no install.

## Project layout

| Package | Purpose |
Expand All @@ -210,12 +256,17 @@ filling in as the loop progresses.

## Status

Early. v1 in progress: parser, runtime, VSCode extension, BMAD preset. See the [roadmap](#roadmap) and [open issues](../../issues).
Active. Shipped: parser + runtime (pipelines, flows, for-each, evals, judge panels, flake
guard), event log + `--resume`, secret-scrubbed telemetry, live dashboard, VSCode extension,
template library, browser playground, ctx skill provisioning. See the [roadmap](#roadmap)
and [open issues](../../issues).

## Roadmap

- **v1** — parser, single-loop + sequential pipeline runtime on Claude Code, blocking human nodes, VSCode extension, BMAD preset.
- **v2** — visual graph editor (the `loop-spec` IR is built for it), async human nodes, reactive stages, scheduling, a community preset registry (`use someone/their-method`).
- **Next** — runner abstraction (run loops on your own local/API model), a GitHub Action
(loops as CI quality gates), a community template registry (`use someone/their-method`).
- **Later** — visual graph editor (the `loop-spec` IR is built for it), async human nodes,
reactive stages, scheduling.

## Built with LoopFlow

Expand Down
Loading
Loading