ctx capability convergence + reliability & recovery + playground — release 0.7.0 - #9
Merged
tickets-forge-dev merged 14 commits intoJul 2, 2026
Conversation
…odel gating Extend the ctx skill-source integration beyond skills to the full capability set, behind a fail-closed permission model. - DSL (additive, config tier): `grant ctx: skills, agents, mcps, harnesses` and `ctx may use my own model "<provider>/<model>"`. - Runtime: CtxProvisionResult carries capability groups + harnessInstall + warnings; McpCtxAdapter threads permissions/own-model and parses the ctx.loop_adapter.v1 contract; the engine merges skills+agents and surfaces mcps/harnesses on the ctx event (never auto-installs); the cli passes grants. - Schema, AGENTS.md, docs/ctx-skill-source.md, docs/ctx-integration-guide.md, examples/ctx_capabilities.loop. - Back-compat: no grant => skills+agents, exactly as before. Tests: parser 54, runtime 101 (4 new), full build green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`ctx may use my own model "ollama/…"` only unlocks ctx harness recommendations; it never makes the loop run on that model. Warn at run time when the provider's local binary (e.g. ollama) isn't on PATH, so the author isn't surprised that running a recommended harness later would fail. API providers (no local binary) never warn. New ownModel.ts (ownModelBinaryWarning + commandOnPath), wired into the cli ctx path, +1 runtime test (injected PATH lookup). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Stream every LoopEvent to a collector when LOOP_EVENTS_URL is set (+ LOOP_EVENTS_TOKEN, LOOP_RUN_ID). Best-effort, fire-and-forget, flushes on exit; silent when unset/unreachable. +3 tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Moved agentic-engineering.loop from root to examples/ - Moved demo.loop from root to examples/ - Updated references in test and docs to reflect new paths Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
Records the full loop event stream to a local file, the durable sibling of
the LOOP_EVENTS_URL control-plane sink. Every event the engine already emits
(loop/node/observe/reflect/transition/human/ctx/git/hook/model/stop) appends
as one NDJSON line: a `loop.log.v1` header then `{seq, ts, event}` per event.
- makeFileEventSink: synchronous appendFileSync per event, so the log is
durable the instant it fires — survives Ctrl-C / crash with no lost tail
(flush() is a no-op). Best-effort like the HTTP sink: creates the parent
dir once, disables itself quietly on any write error, never throws.
- combineSinks: fan one event stream out to HTTP + file, sharing one runId so
the collector and the local log correlate. eventSinkFromEnv wires both, so
all three run paths (--events, --live, default) log with no cli.ts churn.
- --log <path> flag overrides LOOP_LOG_FILE for one-off runs.
Off by default: no env, no flag → no sink, behaves exactly as before.
Tests: 5 sink + override/precedence cases; runtime suite 111/111.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…P_EVENTS_URL) MANUAL.md §4 gains an "Event log & telemetry" section: what the event stream records, the NDJSON format (header + seq'd event lines) with jq read-back examples, durability + best-effort guarantees, and the env-var table for the local log and the remote HTTP collector. Usage line + flag list pick up --log. README points at it from the live-dashboard section. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…es N times` Re-runs a test/command `done when` check N times and requires EVERY run to pass, so "done" means "passes reliably" not "passed once" — a guard against a green that only holds by luck (timing- or order-dependent tests). - parser: an optional `N times` suffix on `passes` / `succeeds` / `finds nothing` and `the test "…" passes`, surfaced as `runs` on the predicate. Only set when > 1, so a plain check (and "1 time") keep the existing single-run shape. The `check:`/`verify:` sugar accepts it too. - verify: ShellVerifier loops `runs` times; the first failing run short-circuits (labelled `run i/N failed —`); a clean pass reports `(passed N/N runs)`. - show/explain: rendered as `×N` (ASCII) and "N times in a row" (prose). Docs: AGENTS.md + MANUAL.md predicate sections. Parser 58/58, runtime 115/115 (4 new verifier tests prove the exact run count + short-circuit). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ge panels, secret redaction, --resume Three features that turn a demo loop into one you can trust unattended: • Judge panels — `done when the skill "x" approves by N judges`. N independent verdicts, majority wins, early-exit once mathematically decided. Each vote is its own skill-verify event (`judge i/N: …`); observe reports the tally. The eval-side counterpart of `passes N times`: flake guard for tests, judge panel for evals. Composes with scores/subject/the bar. • Secret redaction — every event is scrubbed before any sink persists it: values of secret-named env vars (*_TOKEN/_SECRET/_PASSWORD/_KEY/…) become [redacted:<VAR>], and well-known shapes (GitHub/Slack/AWS/sk- keys, JWTs, PEM blocks, Bearer headers, password= assignments) are masked. On by default; LOOP_REDACT=off to disable. Best-effort: a redactor fault falls back to the raw event rather than breaking telemetry. • Resume — `loop run file.loop --resume run.log`. The NDJSON log doubles as a journal: buildResumePlan() replays it (container stack + step depth keep nested sub-runs honest) and the engine skips every unit whose end event says satisfied — whole definitions, pipeline stages, flow steps, foreach items — emitting ⏩ resumed events. Flow handoff summaries now ride on flow-step-end so a resumed flow restores carry-forward context; the log header carries a sha256 of the .loop source and the CLI warns on drift. Interrupted or failed units re-run from scratch: nothing is trusted that a check didn't prove. Tests: parser 61/61 (judge grammar), runtime 132/132 — 6 redaction, 4 panel (majority/early-exit/single-judge parity), 7 resume (log-scan unit tests + end-to-end crash-resume for pipeline, flow w/ summary restore, foreach). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The parser, the ASCII "show" renderer, explain, and the soft linter are all pure TypeScript — one esbuild pass (npm run build:playground) bundles the whole language toolchain to 27 kB of client-side JS. docs/playground.html is a split-pane editor: parse-on-type with inline errors, the compact flow view, lint nudges, and the plain-English explain, plus a picker of examples that showcase the grammar (flake guard, judge panels, pipelines, flows). Linked from the tutorial nav and the hands-on cards. No install, no server. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MANUAL.md gains the section users actually ask for — "How verification works: what 'done' actually depends on": when checks run (the observe node), the short-circuiting conjunction, and the factors that flip a verdict (working dir under target:/worktree isolation, shell env, exit-code semantics, output emptiness for `finds nothing`, the npm-style `test` desugar, flakiness, truncation), how evals decide (subject, the bar, scores, panels), what never auto-passes, and what the verdict triggers. Plus: "Resuming an interrupted run" under telemetry, the redaction guarantees, and the `by N judges` panel in the predicate reference. AGENTS.md and README track the same features. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One ring, one gap at 1–2 o'clock: the loop that hasn't closed yet, still iterating. Replaces the gimbal-detail gyro mark, which turned to mud below 32px. Single flat teal stroke, round caps, crisp at 16px, Vercel-simple. Applied everywhere the mark lives: the favicon SVG (README inherits it) and the inline header lockups in the tutorial, workshop, and game pages. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rial
The icon file (gyro-icon.svg — name kept so all 30 favicon links and the
README hero update with zero churn) is now the solid tile: filled rounded
square, gap-ring knocked out — app-icon treatment of the same mark. The flat
gap ring lives on as docs/logo.svg; page headers keep it inline.
README brought in line with the tutorial and the current feature set:
- hero voice matches the tutorial ("Stop babysitting the agent…")
- links row gains Playground / Workshop / Lab / FAQ
- the pipeline example now actually parses (quoted stage names, real cycle
lines — the old block used pre-v1 syntax); all 5 loop blocks in the README
are verified against the parser
- new "Verify like you mean it" section: done-when conjunction, flake guard,
judge panels, trajectory evals, link to the verification mechanics
- ctx self-equipping paragraph under Skills and memory
- Status/Roadmap reflect what actually shipped (resume, telemetry, playground)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…print.yaml New "You already have a spec (the BMAD flow)" walkthrough in Real Workflows, between the Jira ticket and the Forge case study: copy the load-spec + story-template kit, the three-line for-each driver over YOUR sprint.yaml (backlog stays the source of truth — entry text becomes the story's context), one checklist per story with the pause-and-ask on failure, and the --log/--resume journal so a long sprint survives an interruption. Notes the greenfield path (discover → design produce the spec first) and that the flow is method-neutral — BMAD is one example. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bump for the npm release that ships ctx skill provisioning, verification
reliability (flake guard + judge panels), the event log + secret redaction +
--resume, and the browser playground:
- @loop-lang/{parser,runtime,stdlib,viz} 0.3.0 → 0.4.0 (lockstep)
- @loop-lang/loop (installer) 0.6.0 → 0.7.0
- loopflow (vscode) 0.4.0 → 0.5.0; inter-package dep pins + the extension's
parser devDep synced to 0.4.0 so the bundle knows the new grammar
- CHANGELOG 0.7.0 entry
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Everything since 0.6.0, released as 0.7.0:
ctx integration (the branch's original scope)
recommend skills with ctx, run-time provision + top-up)grant ctx: skills, agents, mcps, harnesses, fail-closed) + own-model gating (dry-run harness recs)LOOP_EVENTS_URLcontrol-plane telemetry sinkVerification reliability
done when "…" passes 3 times(every run must pass, short-circuits)the skill "…" approves by 3 judges(majority, early-exit once decided)Telemetry, redaction, resume
--log/LOOP_LOG_FILE: durable NDJSON event journalLOOP_REDACT=off)--resume run.log: skip units the log proves satisfied; flow carry-forward restored; source-hash drift warningSite & docs
Release
Tests: parser 61 · runtime 132 · stdlib 3 · viz 10 · vscode 19 — all green.
🤖 Generated with Claude Code