diff --git a/.claude/skills/loopflow/SKILL.md b/.claude/skills/loopflow/SKILL.md index 99ab4e5..740eea0 100644 --- a/.claude/skills/loopflow/SKILL.md +++ b/.claude/skills/loopflow/SKILL.md @@ -60,6 +60,8 @@ schedule: run unattended on a cadence runner: which agent executes the loop target: operate on another directory/repo recommend skills with ctx ctx is this file's skill source — recommends + installs skills per loop goal +grant ctx: skills, agents, mcps, harnesses capability groups ctx may recommend (fail-closed; default skills+agents; mcps/harnesses are recommend-only) +ctx may use my own model "/" declares a user-owned model — unlocks harness recommendations (dry-run only) ``` Predicates: @@ -178,6 +180,13 @@ When ctx's tools are available (`ctx__loop_provision`, `ctx__recommend_bundle`): first keeps the `.loop` self-contained and reproducible; the second lets a headless `loop run` re-resolve the bundle from ctx. 3. Offer `top up skills from ctx` if the loop should pull more skills when a cycle fails. +4. **Beyond skills** — if the goal needs more than skills, add a `grant ctx: skills, agents, + mcps, harnesses` line for the groups that apply (fail-closed; omit it for skills-only). + `mcps` and `harnesses` are **recommend-only** — ctx surfaces them with an install command + the user runs; the loop never auto-installs them. Harnesses additionally need a + `ctx may use my own model "/"` line, and always come as a `--dry-run` + command. Pass the granted groups (and own-model) to `ctx__loop_provision` as `permissions` / + `own_llm` / `model_provider` / `model`. When ctx is **not** attached, skip this silently and author `use skills:` by hand as usual — the loop runs the same either way. diff --git a/AGENTS.md b/AGENTS.md index 59c7c89..b09a28c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -99,6 +99,8 @@ rigor: vibe coding | structured ai-assisted | agentic engineering (the spectru mode: conductor | orchestrator (supervision posture: in-session/sync vs async/opens-a-PR) runs as: (an auditable principal for unattended runs) recommend skills with ctx (config tier: ctx is this file's skill source — recommends + installs skills per loop goal; see "Skill source: ctx" below) +grant ctx: skills, agents, mcps, harnesses (config tier: capability groups the file lets ctx recommend; fails closed, default skills+agents; mcps/harnesses are recommend-only) +ctx may use my own model "/" (config tier: declares a user-owned/local/API model — unlocks ctx harness recommendations, always dry-run) observe: (block) trace every cycle / meter tokens and cost / stop and warn if cost exceeds "$N" sandbox: (block) no network access / allow egress to "host" only / cap cpu at … memory at … time at … hooks: (loop body block) before each cycle | after act | on commit | on stop : "" passes|finds nothing (a failing hook blocks) @@ -114,9 +116,11 @@ stages in parallel: (inside a pipeline: the indented stages run concurrently) done when the test "billing.spec.ts::apostrophe" passes # a named test done when "pnpm test" passes # a shell command, exit 0 done when "semgrep --severity=high" finds nothing # a shell command, empty output +done when "pnpm test flaky" passes 3 times # flake guard: re-run the check, EVERY run must pass done when a human confirms "looks right at 375px" # a human check done when the skill "email-review" approves # an eval: approved / not done when the skill "email-review" scores 8 or more # an eval: numeric threshold +done when the skill "code-review" approves by 3 judges # consensus: N independent verdicts, majority wins done when the skill "api-review" scores 8 or more on the output # an eval of WHAT was produced done when the skill "path-review" approves on the trajectory # an eval of HOW the agent got there the bar: didn't weaken a test to go green; no writes outside api/ # the rubric the judge scores against @@ -125,6 +129,16 @@ done when the skill "path-review" approves on the trajectory # an eval The command in a predicate runs in the user's shell with their privileges (like an npm script). It IS meant to be a real command. Prefer a fast, deterministic check. +**Flake guard — `passes N times`.** Append `N times` to a `test` or command predicate to re-run +it `N` times and require every run to pass (the first failure short-circuits). Reach for it when a +green can pass by luck — a timing- or order-dependent test — so "done" means "passes *reliably*", +not "passed *once*". + +**Judge panel — `by N judges`.** Append `by N judges` to a skill predicate to collect `N` +independent verdicts and take the majority (early-exit once decided). A single LM judge wobbles +near the bar; consensus smooths the noise. The deterministic counterpart of the flake guard: +flake guard for tests, judge panel for evals. + ### Tests vs evals — list as many `done when` as you need A loop may have **multiple `done when` lines, and ALL must pass** (a conjunction). Use this to @@ -186,8 +200,33 @@ loop "harden the stripe webhook handler": MCP server before the first plan, and `top up skills from ctx` after a failed cycle reflects. - **No ctx attached?** The lines are inert — the loop runs exactly as it would without them. +**Beyond skills — the full capability set.** By default ctx provisions only `skills` +(and the agents Loop loads the same way). A `grant ctx:` line widens what ctx may recommend to +any of `skills, agents, mcps, harnesses`, **failing closed** — only listed groups are returned: + +```loop +recommend skills with ctx +grant ctx: skills, agents, mcps, harnesses # capability grants (fail-closed) +ctx may use my own model "ollama/llama3.1" # unlocks harness recs (dry-run only) + +loop "stand up a local agent loop": + goal: an MCP agent loop running on local ollama with filesystem access + use skills recommended by ctx + done when "pytest tests/agent_loop" passes +``` + +- **skills / agents** install into `~/.claude/skills` (as before) and merge into the cycle's + skill set. +- **mcps** are **recommend-only**: ctx surfaces fitting MCP servers + a suggested + `ctx-mcp-install `; the runtime emits them on a `ctx` event, it never auto-registers one. +- **harnesses** (autogen, langfuse, …) recommend **only** when you declare a user-owned model + (`ctx may use my own model "…"`), and ship as an explicit `ctx-harness-install --dry-run` + command — never an automatic install. This is the one capability that pulls real software, so it + stays human-gated by design. + Setup: `claude mcp add ctx -- ctx-mcp-server` (needs `pip install claude-ctx`). See -`examples/ctx_skills.loop` and `docs/ctx-skill-source.md`. +`examples/ctx_skills.loop`, `examples/ctx_capabilities.loop`, and `docs/ctx-skill-source.md`. +Full customer-facing walkthrough (setup, own-model, the capability model): `docs/ctx-integration-guide.md`. ### `remember in` — cross-run memory @@ -411,6 +450,11 @@ flow, show the file chain. `loop-run ls` lists every loop in the repo. - `loop-run run file.loop` — execute it on Claude Code (plan/act/observe, reflect on failure, verify with `done when`, pause at human gates). +- `loop-run run file.loop --log run.log` — also append every event to a local NDJSON log + (secrets are scrubbed before anything is persisted). +- `loop-run run file.loop --resume run.log` — resume an interrupted run from its log: satisfied + stages / flow steps / for-each items are skipped, the first incomplete unit picks up (flow + carry-forward summaries restored from the log). - `loop-run show file.loop` — print the loop's flow as compact ASCII (and `loop-run ls` to list them). - `loop-run explain file.loop` — describe the loop in plain English (a friendly check of what it will do). - `loop-run viz file.loop` — open a visual HTML schematic of the flow. diff --git a/CHANGELOG.md b/CHANGELOG.md index 7407fa5..ed5003a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,42 @@ Versions track the `@loop-lang/loop` installer package. ## [Unreleased] +## [0.7.0] — 2026-07-02 + +> `@loop-lang/loop` 0.7.0 · `@loop-lang/{parser,runtime,stdlib,viz}` 0.4.0 · `loopflow` (vscode) 0.5.0 + +### Added +- **ctx as a skill source** — `recommend skills with ctx` / `use skills recommended by ctx` + / `top up skills from ctx`: a loop equips itself via the ctx MCP server before the first + plan and re-equips after a failed cycle reflects. Capability grants + (`grant ctx: skills, agents, mcps, harnesses`, fail-closed) and own-model gating + (`ctx may use my own model "…"`, dry-run-only harness recommendations). +- **Verification reliability** — the flake guard (`done when "…" passes 3 times`: every + run must pass, first failure short-circuits) and judge panels + (`the skill "…" approves by 3 judges`: majority of independent verdicts, early-exit + once decided). Rendered in `show`/`explain` (`×3`, `· 3 judges`). +- **Event log & telemetry** — `--log ` / `LOOP_LOG_FILE` appends every runtime + event as durable NDJSON (header + seq'd lines); `LOOP_EVENTS_URL` streams the same + events to a control-plane collector (shared run id, fan-out). **Secret redaction on by + default**: env-derived values and well-known credential shapes are scrubbed before any + sink persists an event (`LOOP_REDACT=off` to disable). +- **Resume** — `loop-run run --resume run.log` skips every unit the log proves + satisfied (definitions, stages, flow steps, for-each items), restores flow carry-forward + summaries, and warns when the `.loop` source changed since the logged run. +- **Browser playground** — `docs/playground.html`: the parser, ASCII shape view, explain, + and soft linter bundled to 27 kB of client-side JS; parse-on-type with inline errors and + example loops. Linked from the tutorial. +- **Docs** — "How verification works: what 'done' actually depends on" in the manual + (verdict factors: working dir, shell env, exit codes, `finds nothing` semantics, flake / + judge hardening); event-log, resume, and redaction sections; README aligned with the + tutorial (all `.loop` blocks verified against the parser). + +### Changed +- **New logo** — the gap ring (one ring, one gap: the loop still iterating), with a + solid-tile variant as the favicon / app icon across the site and README. +- `loopflow` (vscode) 0.5.0 — bundles the 0.4.0 parser (judge panels, flake guard, ctx + lines all recognized); output panel renders `⏩ resumed` events. + ## [0.6.0] — 2026-06-29 > `@loop-lang/loop` 0.6.0 · `@loop-lang/{parser,runtime,stdlib,viz}` 0.3.0 · `loop-vscode` 0.4.0 diff --git a/README.md b/README.md index 4e2f3d8..5f344f4 100644 --- a/README.md +++ b/README.md @@ -6,7 +6,7 @@

An open, natural-language DSL for loop engineering.
Describe a staged, self-correcting, human-gated agent workflow in plain English, press ▶, and it runs on Claude Code.

-

Stop tuning prompts. Start editing the loop.

+

Stop babysitting the agent. Write the goal once — the loop plans, acts, reflects on red,
and stops only when the check is green, at the gates you set.

A Loop turning a failing test green: plan → act → observe (FAIL) → reflect → plan → act → observe (PASS) → done

@@ -24,9 +24,12 @@

Tutorial · - Workshop · + ⚡ Playground · + 🛠️ Workshop · + 🎮 Lab · Manual · - Keyword reference + Keywords · + FAQ

## Quickstart @@ -88,24 +91,48 @@ Compose loops into **stages** and **pipelines**, with humans wired in where judg ```loop pipeline "ship feature": - stage security: + + stage "security": goal: no high or critical vulnerabilities done when "semgrep --severity=high" finds nothing - each cycle: plan, act, observe + each cycle: plan, then act, then observe when it fails: reflect, then plan again - stage build: + stage "build": goal: feature works and tests pass a human approves the plan first - then each cycle: act, observe done when "pnpm test" passes + each cycle: act, then observe - stage ui: + stage "ui": goal: matches design, responsive at 375px - each cycle: plan, act, observe + each cycle: plan, then act, then observe a human reviews before stopping ``` +## Verify like you mean it + +`done when` is the loop's definition of reality — so LoopFlow gives verification real teeth. +List several checks (**all must pass**), mix deterministic tests with LM-judged evals, and +harden both sides against false greens: + +```loop +loop "harden checkout": + goal: checkout works, reliably, and was built the right way + done when "pnpm test checkout" passes 3 times # flake guard: every run must pass + done when the skill "code-review" approves by 3 judges # judge panel: majority of 3 verdicts + done when the skill "path-review" approves on the trajectory # judges HOW it got there + the bar: didn't weaken a test to go green; no writes outside src/checkout +``` + +- **Flake guard** — `passes N times` re-runs a test/command; one lucky green isn't "done". +- **Judge panel** — `by N judges` takes a majority of independent verdicts; one wobbly LM + judgment isn't "done" either. +- **Trajectory evals** — catch what a green test can't: an agent that gamed the check. + +Full mechanics — what a verdict is actually affected by (working dir, shell env, exit codes) — +in [How verification works](docs/MANUAL.md#how-verification-works--what-done-actually-depends-on). + ## Skills and memory Two knobs make a loop coordinate proven work and learn over time: @@ -127,6 +154,12 @@ loop "decide whether to cancel the morning run": first plan and appends an outcome entry when it stops. `reflect` is within-run memory; `remember` is its across-run counterpart. See [`examples/skills_memory.loop`](examples/skills_memory.loop). +And a loop can **equip itself**: with [ctx](https://github.com/stevesolun/ctx) attached as the +skill source, `use skills recommended by ctx` resolves + installs the right skill bundle for the +goal before the first plan, and `top up skills from ctx` pulls more after a failed cycle +reflects. Opt-in, fail-closed, inert without ctx — see +[the integration guide](docs/MANUAL.md) and [`examples/ctx_capabilities.loop`](examples/ctx_capabilities.loop). + ## Compose loops Compose loops into **pipelines** (stages in order, fail-fast), chain whole files with **`flow`**, and fan out over a plan with **`for each`** — humans wired in where judgment lives. Full grammar with worked examples: the [tutorial](https://loopflow.live) and the [manual](docs/MANUAL.md). @@ -192,6 +225,19 @@ When you run a loop via `/loopflow`, the skill asks if you want the dashboard an opens it and updates it as each step happens — pipeline stages, flow steps, and sprint stories filling in as the loop progresses. +Prefer a file you can grep later? Persist the same event stream as NDJSON — and use it to +**resume** an interrupted run: + +``` +loop-run run file.loop --log run.log # append every event to a local log (secrets scrubbed) +loop-run run file.loop --resume run.log # skip what the log proves done; pick up where it died +``` + +See **Event log & telemetry** in [`docs/MANUAL.md`](docs/MANUAL.md) for the format, redaction, +resume semantics, and the `LOOP_EVENTS_URL` remote collector. Want to *feel* the language first? +Open the [**browser playground**](https://loopflow.live/playground.html) — type a `.loop`, see its +shape live, no install. + ## Project layout | Package | Purpose | @@ -210,12 +256,17 @@ filling in as the loop progresses. ## Status -Early. v1 in progress: parser, runtime, VSCode extension, BMAD preset. See the [roadmap](#roadmap) and [open issues](../../issues). +Active. Shipped: parser + runtime (pipelines, flows, for-each, evals, judge panels, flake +guard), event log + `--resume`, secret-scrubbed telemetry, live dashboard, VSCode extension, +template library, browser playground, ctx skill provisioning. See the [roadmap](#roadmap) +and [open issues](../../issues). ## Roadmap -- **v1** — parser, single-loop + sequential pipeline runtime on Claude Code, blocking human nodes, VSCode extension, BMAD preset. -- **v2** — visual graph editor (the `loop-spec` IR is built for it), async human nodes, reactive stages, scheduling, a community preset registry (`use someone/their-method`). +- **Next** — runner abstraction (run loops on your own local/API model), a GitHub Action + (loops as CI quality gates), a community template registry (`use someone/their-method`). +- **Later** — visual graph editor (the `loop-spec` IR is built for it), async human nodes, + reactive stages, scheduling. ## Built with LoopFlow diff --git a/docs/MANUAL.md b/docs/MANUAL.md index ba69a5a..ec0c5fc 100644 --- a/docs/MANUAL.md +++ b/docs/MANUAL.md @@ -116,7 +116,7 @@ the thrash guard). ## 4. The CLI ``` -loop-run [--model ] [--live] [--out ] +loop-run [--model ] [--live] [--log ] [--resume ] [--out ] ``` | Command | What it does | @@ -136,6 +136,10 @@ loop-run [--model ] Code. Omit to use the CLI default. - `--live` — for `run`, open a live browser dashboard and stream every step to it as the loop executes (see below). +- `--log ` — for `run`, append the full event stream to a local NDJSON log + (see **Event log & telemetry** below). Overrides `LOOP_LOG_FILE`. +- `--resume ` — for `run`, skip everything a prior run's event log proves already + satisfied and pick up at the first incomplete unit (see **Resuming an interrupted run**). - `--out ` — for `viz`, the HTML output file (default: the `.loop` name with an `.html` extension). - `--json` — print the parsed loop-spec JSON before the command's normal work (handy with @@ -168,6 +172,109 @@ on connect (and dedupes on reconnect via `Last-Event-ID`), so events fired befor connects are not lost and a transient drop doesn't double-deliver. The dashboard is self-contained (no external assets) and binds to `127.0.0.1` only. +### Event log & telemetry + +Every meaningful thing a run does is a **structured event** — `loop-start`, each +`node-enter` / `node-exit` (with the attempt number), `observe` (pass/fail + output), +`transition`, `reflect`, `loop-back`, human gates, `ctx` provision/top-up, `git` actions, +`hook` results, the `model` tier per phase, `stop` (with the reason), `loop-end`, and the +pipeline / flow / for-each envelopes. The live dashboard renders this stream; you can also +**persist it** — to a local file and/or a remote collector — for auditing, debugging a +thrashing loop, or metering cost. + +**Off by default.** With no flag and no env var, a run persists nothing and behaves exactly +as before. Persistence is **best-effort**: a failing log or an unreachable collector never +throws and never blocks a run — telemetry can't break a loop. + +#### Local log — `--log` / `LOOP_LOG_FILE` + +```bash +loop-run run test.loop --log run.log # one-off (overrides LOOP_LOG_FILE) +LOOP_LOG_FILE=run.log loop-run run test.loop # env — same effect +``` + +The log is **NDJSON**: one JSON object per line, so it streams and greps cleanly. The first +line is a `loop.log.v1` header (the run id + metadata); every line after is one event with a +monotonic `seq` and an ISO timestamp: + +```jsonc +{"v":"loop.log.v1","runId":"…","ts":"2026-07-01T20:04:43.443Z","meta":{"loop_path":"test.loop","principal":"idan"}} +{"seq":0,"ts":"…","event":{"type":"loop-start","name":"test loop"}} +{"seq":1,"ts":"…","event":{"type":"node-enter","node":"plan","attempt":1}} +{"seq":2,"ts":"…","event":{"type":"observe","passed":false,"output":"1 failing test"}} +{"seq":3,"ts":"…","event":{"type":"reflect","focus":"which layer broke","text":"the API returned 500"}} +{"seq":4,"ts":"…","event":{"type":"stop","reason":"done"}} +``` + +Each event is written with a synchronous append, so it's on disk the instant it fires — the +log survives a `Ctrl-C` or a crash with **no lost tail**. The parent directory is created if +missing. It works the same across all run modes (default, `--live`, `--events`). + +**Secrets are scrubbed before anything is persisted.** Command output can echo credentials +(a failing `git push` printing a token, a dumped env). Every event passes through a redactor +before it reaches the file log or the HTTP collector: values of env vars whose *names* look +secret-bearing (`*_TOKEN`, `*_SECRET`, `*_PASSWORD`, `*_KEY`, …) are replaced with +`[redacted:]`, and well-known credential shapes (GitHub / Slack / AWS / `sk-…` API +keys, JWTs, PEM private keys, `Bearer …` headers, `password=…` assignments) are masked by +pattern. Best-effort by design — treat it as a seatbelt, not a licence to log freely. Disable +with `LOOP_REDACT=off` (e.g. when debugging the redactor itself). + +Read it back with any NDJSON tool — e.g. with `jq`: + +```bash +jq 'select(.event.type == "observe")' run.log # every verification result +jq -r 'select(.event.type=="reflect") | .event.text' run.log # what each failure taught it +jq 'select(.event.type == "stop") | .event.reason' run.log # how it ended +``` + +#### Remote collector — `LOOP_EVENTS_URL` + +To stream the same events to a control plane over HTTP, set `LOOP_EVENTS_URL` (and, if the +collector needs it, `LOOP_EVENTS_TOKEN`). Events POST to +`/api/v1/runs//events`, each carrying the monotonic `seq` so the server is +idempotent on retries and out-of-order delivery. + +| Env var | Meaning | +|---|---| +| `LOOP_LOG_FILE` | Local NDJSON log path (the `--log` flag overrides it). | +| `LOOP_EVENTS_URL` | Control-plane collector base URL — enables the HTTP sink. | +| `LOOP_EVENTS_TOKEN` | Shared API token, sent as `x-api-token` (optional). | +| `LOOP_RUN_ID` | Correlate a run's events across sinks (optional; a UUID is generated if unset). | + +The file log and the HTTP collector can run **together** — the same event fans out to both, +sharing one `runId`, so a local trace and the control-plane record line up. Set `LOOP_RUN_ID` +yourself when you want a run's id to match something you already track (a CI job, a ticket). + +### Resuming an interrupted run — `--resume` + +The event log is also a **journal**. A long pipeline that dies at stage 4 — crash, `Ctrl-C`, +laptop lid — doesn't have to start over: + +```bash +loop-run run epic.loop --log run.log # …dies at stage 4 of 6 +loop-run run epic.loop --resume run.log --log run.log # stages 1–3 skip, stage 4 picks up +``` + +What resume does: + +- **Skips what's proven done.** A unit whose end event says `satisfied: true` in the log — + a whole definition, a pipeline stage, a flow step, a `for each` item — is skipped, with a + `⏩ resumed` line in the trace. Everything else (failed, or interrupted mid-flight with no + end event) re-runs from scratch. +- **Restores flow context.** A flow step's handoff summary is recorded on its end event, so + a resumed flow hands the *next* step the same carry-forward text the original run produced. +- **Detects drift.** The log header carries a hash of the `.loop` source; if the file changed + since the logged run you get a warning (units are matched by name/position). Editing the + file to *fix* the failing stage and then resuming is the normal workflow — the warning is + informational, not an error. +- **Composes with `--log`.** Point `--log` at the same file to keep one continuous journal + (headers separate the runs), or at a new file for a clean second record. Nested work needs + no bookkeeping: a sub-file a flow step runs is summarised by that step's own end event. + +The unit of resume is deliberately the *stage / step / item*, not the mid-loop cycle — a +half-finished loop re-verifies from its own `done when`, which is exactly what makes skipping +safe: nothing is trusted that a check didn't prove. + ## 5. Language reference A `.loop` file is indentation-structured. `loop` / `pipeline` sit at column 0; their body @@ -296,9 +403,11 @@ what you intend. done when the test "billing.spec.ts::apostrophe" passes # a named test done when "pnpm test" passes # shell command, exit 0 (`succeeds` also works) done when "semgrep --severity=high" finds nothing # shell command, empty stdout +done when "pnpm test flaky" passes 3 times # flake guard: re-run, every run must pass done when a human confirms "looks right at 375px" # a human check done when the skill "email-review" approves # an eval: approved / not done when the skill "email-review" scores 8 or more # an eval: numeric threshold +done when the skill "code-review" approves by 3 judges # judge panel: N verdicts, majority wins done when the skill "api-review" scores 8 or more on the output # an eval of WHAT was produced done when the skill "path-review" approves on the trajectory # an eval of HOW it got there the bar: didn't weaken a test to go green; no writes outside api/ # the rubric the judge scores against @@ -307,6 +416,23 @@ done when the skill "path-review" approves on the trajectory # an eval The command runs in your shell with your privileges (like an npm script). Keep it fast and deterministic. +**Flake guard — `passes N times`.** Append `N times` to a `test` or command predicate +(`passes` / `succeeds` / `finds nothing`) to re-run the check `N` times and require **every** +run to pass. The first failing run short-circuits (the rest don't execute). Use it when a +green can hold by luck — a timing-dependent or order-dependent test — so "done" means +"passes *reliably*", not "passed *once*". `show` renders it as `×N`; a plain check (or +`1 time`) is the usual single run. + +**Judge panel — `by N judges`.** Append `by N judges` to a skill predicate to collect `N` +independent verdicts and take the **majority**. A single LM judge wobbles run-to-run near the +bar; independent samples average the noise out. The panel early-exits once the vote is +mathematically decided (2 approvals out of 3 → the third judge never runs), each verdict is +emitted as its own `skill-verify` event (`judge 1/3: …`), and the observe output reports the +tally (`judges: 2/2 approved (majority of 3 reached)`). Composes with everything else: +`scores 8 or more on the trajectory by 5 judges`. This is the eval-side counterpart of the +flake guard — **flake guard for tests, judge panel for evals**; costs N× the eval, so reserve +it for checks where a wrong "done" is expensive. + #### Tests vs evals A loop can list **several `done when` lines, and all must pass.** Use this to combine the @@ -323,6 +449,56 @@ weakening it. Pair a test with an eval when "done" means both *it works* and *it the right way*. Build the review skill manually first and confirm it judges well, then wire it in. See `examples/skills_memory.loop` and `examples/email_review.loop`. +#### How verification works — what "done" actually depends on + +The most common question about Loop: *when the loop says "done", what exactly decided that?* +The full mechanics, end to end: + +**When it runs.** Verification happens at the **observe** node of every cycle — after act, +before any transition. Nothing is verified mid-act; a cycle without `observe` in its +`each cycle:` never checks at all (it relies on a human gate to stop). + +**The conjunction.** Every `done when` line must pass, in the order written. The **first +failing check short-circuits** — later predicates don't run that cycle. Their combined output +becomes the observe result; on failure that text is *"the last failure"* your `look at:` +context refers to, and it feeds the `reflect` step. Verification isn't just a gate — it's the +loop's sensory input. + +**Where a command runs (this is what surprises people).** A `test` / command predicate runs +as a real shell command: + +| Factor | Effect on the verdict | +|---|---| +| **Working directory** | The loop's effective `baseDir` — your repo dir by default, the `target:` dir if the config sets one, and **the worktree** when the git policy is `work in a worktree`. A check that passes in-repo can fail in a fresh worktree (untracked files, unbuilt artifacts). | +| **Your shell + your env** | The command runs with *your* privileges and environment, like an npm script. A `PATH`, `NODE_ENV`, or missing env var difference between your machine and CI changes the verdict. | +| **Exit code** | `passes` / `succeeds` = **exit 0**. Nothing else is inspected. A test runner configured to exit 0 on failures will make the loop lie to you — verify your runner's exit-code behavior first. | +| **Output emptiness** | `finds nothing` = exit 0 **and** empty stdout+stderr. A scanner that prints a benign banner never "finds nothing" — silence it or wrap it. | +| **The `test` shorthand** | `the test "x::y" passes` desugars to `npm test -- x::y` by default. If your runner isn't npm-style, use an explicit command predicate instead. | +| **Flakiness** | One lucky green = done, unless you add `passes N times` (every run must pass, first failure short-circuits). | +| **Output size** | Captured output is truncated (~4 kB into the trace) — the verdict uses the exit code, not the text, so truncation never changes pass/fail. | + +**How an eval decides.** A skill predicate never touches the shell — it's routed to the +runner, which invokes the named review skill as a judge. What the judge sees is the +**subject**: the act summary (`on the output`, default) or the captured path and tool calls +(`on the trajectory`). It scores against the `the bar:` rubric if present; `scores N or more` +compares its numeric score to the threshold; `by N judges` repeats the judgment independently +and takes the majority. An eval's verdict is an LM's judgment — it can wobble; the bar and the +panel are the two tools that stabilize it. + +**What never auto-passes.** `a human confirms "…"` always waits for you. A loop with **no** +`done when` at all can never be machine-satisfied — it needs `a human reviews before stopping`, +or it runs until the `after N tries` guard / the hard cap (25 cycles) stops it unsatisfied. + +**What the verdict triggers.** Pass + goal met → the loop stops (then `also:` finishing passes, +then a human review if declared — and an `on stop` hook can still **veto** the stop). Fail → +the `when it fails:` transition (reflect → plan again). So the checks you write are literally +the loop's definition of reality: a wrong predicate doesn't make the loop fail — it makes the +loop *stop caring about the right thing*. + +**Rules of thumb.** Fast, deterministic, loud-on-failure commands; `finds nothing` for +scanners; `passes N times` for anything with timing in it; a test **and** an eval when "done" +has a quality dimension; a trajectory eval when you're worried the agent will game the test. + ### Config tier (top of file) ```loop diff --git a/docs/agentic-engineering-plan.md b/docs/agentic-engineering-plan.md index 6d74141..f38e6fb 100644 --- a/docs/agentic-engineering-plan.md +++ b/docs/agentic-engineering-plan.md @@ -1,7 +1,7 @@ # Plan — Agentic Engineering constructs for Loop > Bringing agentic-engineering discipline into Loop. -> Branch: `claude/loop-lang-concepts-3p5tz1`. Companion pipeline: [`agentic-engineering.loop`](../agentic-engineering.loop). +> Branch: `claude/loop-lang-concepts-3p5tz1`. Companion pipeline: [`agentic-engineering.loop`](../examples/agentic-engineering.loop). ## Context diff --git a/docs/ctx-integration-guide.md b/docs/ctx-integration-guide.md new file mode 100644 index 0000000..86b1269 --- /dev/null +++ b/docs/ctx-integration-guide.md @@ -0,0 +1,388 @@ +# Loop × ctx — the self-equipping coding loop + +> Your loop already knows *what* to build and *how to check it's done*. +> ctx makes it know *what to bring* — the skills, agents, MCP servers, and model +> harnesses the job needs — and loads them before the first plan. + +This is the complete guide to the Loop ⇄ ctx integration: what it is, why it +matters, how to set it up, and how to drive the full capability set — including +running on **your own local or API model**. + +--- + +## 1. The 60-second pitch + +A `.loop` file is a plain-English, self-correcting workflow: a goal, a way to +verify "done", human gates, and a retry edge. It already runs your agent in a +tight plan → act → observe → reflect cycle until the tests pass. + +The one thing a loop *couldn't* do was **equip itself**. `use skills: a, b` +assumes `a` and `b` already exist on disk. Someone had to know the right skills, +find them, and install them by hand. + +**ctx closes that gap.** Point a loop at a goal and ctx recommends the smallest +useful bundle of capabilities for it and provisions them — so the loop walks in +already holding the right tools: + +- **Skills & agents** — installed straight into `~/.claude/skills`, ready for the + loop's very first plan. +- **MCP servers** — recommended with a one-line install command (e.g. a + filesystem or database server the goal implies). +- **Model harnesses** — when you bring your own model (local Ollama, an API + model), ctx recommends a fitting agent harness (AutoGen, Langfuse, …) as a + ready-to-run, **dry-run** install command. + +It is **opt-in, fail-closed, and human-gated by design.** A loop with no ctx +attached runs exactly as before. Nothing heavier than a skill is ever installed +without you asking. + +**The outcome you're buying:** stop hand-curating tooling for every workflow. +Describe the goal; the loop arrives equipped. + +--- + +## 2. The problem it solves + +Teams writing agentic workflows hit the same wall: + +| Without ctx | With ctx | +|---|---| +| You must already know which skills a task needs. | Describe the goal; ctx recommends the bundle. | +| Skills are installed by hand, per machine, per person. | The loop installs them at run time, reproducibly. | +| MCP servers and model harnesses are wired up manually. | Recommended for the goal, with the exact install command. | +| "Bring your own model" means assembling a harness yourself. | Declare your model; ctx recommends a fitting harness. | +| Tooling drift between author's box and CI. | The `.loop` re-resolves its bundle on every headless run. | + +ctx is the **provisioning layer beneath Loop**. Loop stays the driver; ctx is +the quartermaster. + +--- + +## 3. What you get — the capability set + +ctx recommends across four capability groups. A `.loop` *grants* which ones +apply (see §6). Each group behaves differently, on purpose: + +| Group | Installed automatically? | What happens | +|-------|--------------------------|--------------| +| **skills** | ✅ into `~/.claude/skills` | Merged into the loop's skill set for plan/act. | +| **agents** | ✅ into `~/.claude` | Sub-agents the loop can invoke, loaded the same way. | +| **mcps** | ❌ **recommend-only** | Fitting MCP servers surfaced with a `ctx-mcp-install ` command. The loop never auto-registers one. | +| **harnesses** | ❌ **recommend-only, gated** | Recommended only when you declare your own model; shipped as a `ctx-harness-install --dry-run` command you run. Never auto-installed. | + +**Why the split?** Skills and agents are small, sandboxed, and the loop needs +them in hand to work. MCP servers and harnesses pull real software and touch your +machine's configuration — so ctx *recommends* them and hands you the exact +command, but the decision to install stays yours. That's the trust boundary that +makes this safe to run unattended. + +--- + +## 4. How it works + +``` + ┌────────────┐ grant + goal + own-model ┌─────────────────┐ + You → │ .loop │ ────────────────────────────► │ ctx-mcp-server │ + │ (Loop) │ ◄──────────────────────────── │ (recommender) │ + └─────┬──────┘ ctx.loop_adapter.v1 contract └────────┬────────┘ + │ │ + │ skills/agents → installed │ recommend_bundle + │ mcps/harnesses → surfaced (recommend-only) │ + harness recommender + ▼ ▼ + plan → act → observe → reflect ↺ ~/.claude/skills + the graph +``` + +1. A loop that opts into ctx calls `ctx__loop_provision` **once before the first + plan**, passing its goal, the capability grants, and (optionally) your model. +2. ctx returns a single read-only JSON contract (`ctx.loop_adapter.v1`): the + skills/agents it installed, and the MCP servers / harnesses it recommends. +3. The loop merges skills + agents into its working set, and surfaces the + recommend-only items on its event stream for you (or your host) to act on. +4. On a failed cycle, `top up skills from ctx` asks for *more* — the loop learns + what it was missing from the failure and re-equips before the next plan. + +If ctx isn't attached, every ctx line is inert and the loop runs unchanged. A +ctx call that fails emits one "skipped" event and the loop continues. **A loop +never fails because ctx is missing.** + +--- + +## 5. Setup + +### Prerequisites +- [Loop](https://github.com/tickets-forge-dev/loop-lang) (`.loop` runtime / the + `/loopflow` skill in Claude Code). +- Python 3.11+ for ctx. + +### Install & attach ctx + +```bash +# 1. Install ctx and seed its recommendation graph +pip install claude-ctx +ctx-init --graph --model-mode skip # extracts the recommendation graph into ~/.claude/skill-wiki + +# 2. Attach ctx's tools to Claude Code over MCP +claude mcp add ctx -- ctx-mcp-server +claude mcp list # → ctx: ✔ Connected +``` + +That exposes the tools the Loop bridge uses: + +- `ctx__recommend_bundle` — read-only preview of what ctx would recommend. +- `ctx__loop_provision` — recommend + install skills/agents, recommend mcps/harnesses, return the contract. +- `ctx__loop_topup` — the same, for *additional* capabilities after a failed cycle. + +> **No graph yet?** ctx will return an empty (but valid) contract — the loop runs +> on whatever it already names. Re-run `ctx-init --graph` to seed or refresh. + +--- + +## 6. The grammar + +Five lines, all additive, all inert without ctx attached. + +```loop +recommend skills with ctx # config: ctx is this file's capability source +grant ctx: skills, agents, mcps, harnesses # config: which groups ctx may recommend (fail-closed) +ctx may use my own model "ollama/llama3.1" # config: declare your model → unlocks harnesses + +loop "stand up a local agent loop": + goal: an MCP agent loop on local ollama with filesystem access, with passing tests + use skills recommended by ctx for "local ollama agent loop with filesystem MCP" # loop body + top up skills from ctx when a step needs more # loop body + done when "pytest agent/tests/test_loop.py" passes +``` + +| Line | Tier | Effect | +|------|------|--------| +| `recommend skills with ctx` | config | Declares ctx as the file's capability source. | +| `grant ctx: ` | config | Capability groups ctx may recommend. **Fails closed** — omit it and ctx defaults to `skills + agents`; list only what you want. | +| `ctx may use my own model "/"` | config | Declares a user-owned/local/API model. Required to unlock **harness** recommendations. | +| `use skills recommended by ctx [for ""]` | loop body | Provision the bundle for the goal (or an explicit intent) before the first plan. | +| `top up skills from ctx when a step needs more` | loop body | After a failed cycle reflects, pull additional capabilities before re-planning. | + +### Fail-closed permissions — what it means + +`grant ctx:` is an allow-list, not a wish-list. ctx returns **only** the groups +you name: + +- No `grant ctx:` line → `skills + agents` (the original, safe default). +- `grant ctx: skills` → skills only; agents/mcps/harnesses are never returned. +- `grant ctx: skills, mcps` → skills installed, MCP servers recommended; no agents, no harnesses. + +A typo in a group name grants nothing for that token — it can never accidentally +widen access. + +--- + +## 7. Using your own model (the harness story) + +This is the feature that turns Loop × ctx from "skill installer" into "bring your +own model agent platform". + +If you run on a **local model** (Ollama, llama.cpp) or **your own API model**, +you usually need a *harness* — an agent framework like AutoGen or an +observability layer like Langfuse — wired to that model. ctx recommends one for +your goal and model, and hands you the command to install it. + +### Step 1 — declare your model + +```loop +ctx may use my own model "ollama/llama3.1" +``` + +The string is `"/"`. The provider (before the first `/`) and the +full model id are both passed to ctx so it can score harnesses for your exact +setup. + +### Step 2 — grant the harness group + +```loop +grant ctx: skills, harnesses +``` + +Harnesses are **double-gated**: they're returned only when *both* `harnesses` is +granted *and* a model is declared. Grant `harnesses` without a model and ctx +fails closed with a clear warning instead of recommending something it can't fit: + +```json +"warnings": ["harnesses granted but no user-owned model declared + (set own_llm / model_provider / model) — skipping harness recs."] +``` + +### Step 3 — run, review, install + +ctx returns the recommended harnesses with fit scores and a **dry-run** install +command: + +```json +"capabilities": { + "harnesses": [ + { "name": "autogen", "type": "harness", "fit_score": 1.0, + "install_command": "ctx-harness-install autogen --dry-run" }, + { "name": "langfuse", "type": "harness", "fit_score": 0.9, + "install_command": "ctx-harness-install langfuse --dry-run" } + ] +}, +"harness_install": "ctx-harness-install autogen --dry-run" +``` + +The loop **never installs a harness for you.** It surfaces the command; you run +it. `--dry-run` shows exactly what would be installed before anything touches +your machine. Drop `--dry-run` when you're ready. + +> **Why gated and dry-run?** A harness is the one capability that pulls a full +> framework and runs code against your model. Keeping it an explicit, previewable +> step is what lets you grant `harnesses` in a workflow that otherwise runs +> unattended. + +--- + +## 8. A full worked example + +`examples/ctx_capabilities.loop`: + +```loop +recommend skills with ctx +grant ctx: skills, agents, mcps, harnesses +ctx may use my own model "ollama/llama3.1" + +loop "stand up a local agent loop": + goal: an MCP agent loop running on local ollama with filesystem access, with passing tests + look at: agent/loop.py, agent/tests/test_loop.py + use skills recommended by ctx for "local ollama agent loop with filesystem MCP" + top up skills from ctx when a step needs more + each cycle: plan, then act, then observe + done when "pytest agent/tests/test_loop.py" passes + when it fails: reflect on the failing assertion, then plan again + after 6 tries: stop and warn "local agent loop still red — needs a human" +``` + +Print its shape: + +```bash +loop show examples/ctx_capabilities.loop +``` +``` +loop "stand up a local agent loop" + ↻ plan → act → observe (each cycle) + ↺ on fail: reflect → plan (the back-edge) + ✓ done when: "pytest agent/tests/test_loop.py" passes + ⛔ guard: after 6 tries → stop & warn "local agent loop still red — needs a human" +``` + +Run it: + +```bash +loop run examples/ctx_capabilities.loop --events +``` + +What happens on the first cycle: +1. ctx provisions skills + agents for the goal → installed, merged into the plan. +2. The filesystem MCP server is **recommended** (with its install command) on the + `ctx` event — you decide whether to register it. +3. Because a model is declared, a fitting **harness** is recommended as a dry-run + command. +4. plan → act → observe runs. If the tests fail, `top up skills from ctx` pulls + more before the next plan. + +--- + +## 9. The contract (for integrators) + +Every provision/top-up call returns one stable, versioned JSON object. Build +against it directly if you're embedding Loop or driving ctx from another host: + +```jsonc +{ + "version": "ctx.loop_adapter.v1", + "permissions": { "skills": true, "agents": true, "mcps": true, "harnesses": true }, + "use_skills": ["..."], // skill + agent names now resolvable on disk + "installed": ["..."], // freshly installed this call + "skipped": ["..."], // already present + "unavailable": [{ "name": "...", "status": "not-in-wiki" }], + "recommended": [{ "name": "...", "type": "skill", "score": 146.9 }], + "capabilities": { + "skills": [{ "name": "...", "type": "skill", "status": "installed" }], + "agents": [{ "name": "...", "type": "agent", "status": "installed" }], + "mcps": [{ "name": "...", "type": "mcp-server", "status": "available", + "install_command": "ctx-mcp-install ..." }], + "harnesses": [{ "name": "...", "type": "harness", "fit_score": 1.0, + "install_command": "ctx-harness-install ... --dry-run" }] + }, + "harness_install": "ctx-harness-install ... --dry-run", // or null + "warnings": [] +} +``` + +The contract is **additive and back-compatible**: the original +`use_skills`/`installed`/`skipped` keys are unchanged, so existing skills-only +integrations keep working untouched. + +MCP tool parameters (`ctx__loop_provision` / `ctx__loop_topup`): + +| Param | Type | Meaning | +|-------|------|---------| +| `goal` | string | What the capabilities are for. | +| `intent` | string | Optional query override. | +| `permissions` | string[] | Granted groups. Omit → `skills + agents`. | +| `own_llm` / `model_provider` / `model` | bool / string / string | Your model — unlocks harnesses. | +| `top_k` | int | Recommendations per group (≤ 5). | +| `dry_run` | bool | Recommend without installing skills/agents. | + +--- + +## 10. Safety & trust + +Designed to be safe to grant in unattended workflows: + +- **Opt-in.** No ctx attached → every ctx line is a no-op. Existing loops are unaffected. +- **Fail-closed.** Capabilities are an allow-list. Nothing outside the grant is ever returned. +- **Recommend-only for heavy capabilities.** MCP servers and harnesses are never + auto-installed — ctx hands you the command; you run it. +- **Dry-run by default for harnesses.** See exactly what would be installed first. +- **Human-gated, double-gated for harnesses.** They require both the grant *and* a declared model. +- **Degrades quietly.** A failed ctx call emits one event and the loop continues + with whatever it already names. +- **Reproducible.** Author-time names are baked into a literal `use skills:` line, + while the directive re-resolves on headless runs — so CI matches the author's box. + +--- + +## 11. FAQ + +**Do I have to use ctx?** No. It's entirely optional and opt-in. Loops without +ctx lines behave identically. + +**Will it install things I didn't approve?** Only skills and agents are installed +automatically, and only from groups you granted. MCP servers and harnesses are +never auto-installed. + +**Can I preview before anything changes?** Yes — `ctx__recommend_bundle` is a +read-only preview, and harness/MCP recommendations are always commands you choose +to run. Use `dry_run: true` to recommend skills/agents without installing them. + +**Does it work headless / in CI?** Yes. `loop run ` re-resolves the bundle +through the ctx MCP server, so an unattended run equips itself the same way an +author's session did. + +**What if my goal needs a tool ctx doesn't know?** ctx recommends from its graph; +unknown items simply don't appear. The loop still runs with whatever it names. +Re-seed or extend the graph to teach ctx new capabilities. + +**Local model or API model?** Both. Declare it with +`ctx may use my own model "/"`. That's what unlocks harness +recommendations tuned to your setup. + +--- + +## 12. Reference + +- Worked examples: `examples/ctx_skills.loop` (skills only), + `examples/ctx_capabilities.loop` (full capability set). +- Grammar in context: `AGENTS.md` → *Skill source: ctx*. +- Mechanics & contract: `docs/ctx-skill-source.md`. +- ctx itself: (`pip install claude-ctx`). + +**One line to remember:** *ctx provisions; Loop drives.* You describe the goal — +the loop arrives equipped. diff --git a/docs/ctx-skill-source.md b/docs/ctx-skill-source.md index 8ab0132..d40eae0 100644 --- a/docs/ctx-skill-source.md +++ b/docs/ctx-skill-source.md @@ -19,9 +19,12 @@ ctx-init --graph --model-mode skip # seed the recommendation graph claude mcp add ctx -- ctx-mcp-server # attach ctx's MCP tools ``` -That exposes four tools the Loop bridge uses: `ctx__recommend_bundle` (preview), -`ctx__loop_provision` (recommend + install + return names), and -`ctx__loop_topup` (add more on a failing cycle). +That exposes the tools the Loop bridge uses: `ctx__recommend_bundle` (preview), +`ctx__loop_provision` (recommend + install + return names), and `ctx__loop_topup` +(add more on a failing cycle). `loop_provision`/`loop_topup` accept an optional +`permissions` array (`skills, agents, mcps, harnesses`) plus `own_llm` / +`model_provider` / `model`, and return the versioned `ctx.loop_adapter.v1` +contract (see *Capability groups* below). ## Grammar @@ -43,6 +46,32 @@ loop "harden the stripe webhook handler": | `recommend skills with ctx` | config | Declares ctx as the file's skill source. | | `use skills recommended by ctx [for ""]` | loop body | Author-time: bake resolved names into `use skills:`. Run-time: re-resolve before the first plan. | | `top up skills from ctx when a step needs more` | loop body | Run-time: after a cycle fails and reflects, pull additional skills before re-planning. | +| `grant ctx: skills, agents, mcps, harnesses` | config | Capability groups the file lets ctx recommend. Fails closed; default (no line) = skills+agents. | +| `ctx may use my own model "/"` | config | Declares a user-owned/local/API model — unlocks harness recommendations (dry-run only). | + +## Capability groups (beyond skills) + +ctx recommends across four entity types; a `.loop` grants which ones apply. The +model **fails closed** — with no `grant ctx:` line the grant defaults to +`skills + agents` (the original behaviour), and only listed groups are ever +returned. + +| Group | Installed? | Behaviour | +|-------|-----------|-----------| +| `skills` | yes → `~/.claude/skills` | Merged into the cycle's skill set, as before. | +| `agents` | yes → `~/.claude` | Loaded the same way Loop loads named (sub)agents. | +| `mcps` | **no — recommend-only** | Fitting MCP servers surfaced with a suggested `ctx-mcp-install `; emitted on the `ctx` event. The runtime never auto-registers one. | +| `harnesses` | **no — recommend-only, gated** | Recommended only when the loop declares a user-owned model (`ctx may use my own model …`); shipped as an explicit `ctx-harness-install --dry-run` command. Never an automatic install. | + +The provision/top-up calls return the `ctx.loop_adapter.v1` contract: +`{ version, permissions, use_skills, installed, skipped, unavailable, +recommended, capabilities{skills,agents,mcps,harnesses}, harness_install, +warnings }`. The runtime merges `use_skills` (skills + agents) into the loop and +surfaces `capabilities.mcps` / `capabilities.harnesses` / `harness_install` on +the `ctx` event for the host or a human to act on — it never installs an MCP +server or a harness on its own. + +See `examples/ctx_capabilities.loop` for the full-capability example. ## How it works diff --git a/docs/game.html b/docs/game.html index 76baa38..fe47b7e 100644 --- a/docs/game.html +++ b/docs/game.html @@ -225,7 +225,7 @@
-
LoopFlow Studio
+
LoopFlow Studio
← tutorial 🛠️ workshop @@ -271,7 +271,7 @@

Run