Repository navigation
feat(ctx): integrate ctx as an opt-in skill source for loops - #8
Merged
Merged
Conversation
Add the Loop side of the Loop↔ctx skill-provisioning bridge: a `.loop` can let
ctx recommend + install the skills its loops need instead of assuming they
already exist in ~/.claude/skills.
Grammar (parser + schema + AGENTS.md + loopflow skill):
- config tier: `recommend skills with ctx`
- loop: `use skills recommended by ctx [for "<intent>"]`, `top up skills from ctx`
Runtime:
- McpCtxAdapter (packages/runtime/src/ctx.ts) speaks MCP to a `ctx-mcp-server`
and calls ctx__loop_provision / ctx__loop_topup.
- engine provisions skills before the first plan and tops up after a failed
cycle; cli attaches the adapter only when a file opts in (dynamic import
keeps the MCP SDK off every other run). A missing server/SDK degrades to a
warning — ctx is always optional, the loop runs the same without it.
Adds @modelcontextprotocol/sdk, the ctx event type, docs/ctx-skill-source.md,
examples/ctx_skills.loop, and parser/engine tests.
Also bundles vscode packaging fixes (rename loop-vscode→loopflow, publisher
Loop-Lang, drop --minify, ignore *.vsix) per request to ship as one PR.
Server side: stevesolun/ctx#209.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…odel gating Extend the ctx skill-source integration beyond skills to the full capability set, behind a fail-closed permission model. - DSL (additive, config tier): `grant ctx: skills, agents, mcps, harnesses` and `ctx may use my own model "<provider>/<model>"`. - Runtime: CtxProvisionResult carries capability groups + harnessInstall + warnings; McpCtxAdapter threads permissions/own-model and parses the ctx.loop_adapter.v1 contract; the engine merges skills+agents and surfaces mcps/harnesses on the ctx event (never auto-installs); the cli passes grants. - Schema, AGENTS.md, docs/ctx-skill-source.md, docs/ctx-integration-guide.md, examples/ctx_capabilities.loop. - Back-compat: no grant => skills+agents, exactly as before. Tests: parser 54, runtime 101 (4 new), full build green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`ctx may use my own model "ollama/…"` only unlocks ctx harness recommendations; it never makes the loop run on that model. Warn at run time when the provider's local binary (e.g. ollama) isn't on PATH, so the author isn't surprised that running a recommended harness later would fail. API providers (no local binary) never warn. New ownModel.ts (ownModelBinaryWarning + commandOnPath), wired into the cli ctx path, +1 runtime test (injected PATH lookup). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Stream every LoopEvent to a collector when LOOP_EVENTS_URL is set (+ LOOP_EVENTS_TOKEN, LOOP_RUN_ID). Best-effort, fire-and-forget, flushes on exit; silent when unset/unreachable. +3 tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Moved agentic-engineering.loop from root to examples/ - Moved demo.loop from root to examples/ - Updated references in test and docs to reflect new paths Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
Records the full loop event stream to a local file, the durable sibling of
the LOOP_EVENTS_URL control-plane sink. Every event the engine already emits
(loop/node/observe/reflect/transition/human/ctx/git/hook/model/stop) appends
as one NDJSON line: a `loop.log.v1` header then `{seq, ts, event}` per event.
- makeFileEventSink: synchronous appendFileSync per event, so the log is
durable the instant it fires — survives Ctrl-C / crash with no lost tail
(flush() is a no-op). Best-effort like the HTTP sink: creates the parent
dir once, disables itself quietly on any write error, never throws.
- combineSinks: fan one event stream out to HTTP + file, sharing one runId so
the collector and the local log correlate. eventSinkFromEnv wires both, so
all three run paths (--events, --live, default) log with no cli.ts churn.
- --log <path> flag overrides LOOP_LOG_FILE for one-off runs.
Off by default: no env, no flag → no sink, behaves exactly as before.
Tests: 5 sink + override/precedence cases; runtime suite 111/111.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…P_EVENTS_URL) MANUAL.md §4 gains an "Event log & telemetry" section: what the event stream records, the NDJSON format (header + seq'd event lines) with jq read-back examples, durability + best-effort guarantees, and the env-var table for the local log and the remote HTTP collector. Usage line + flag list pick up --log. README points at it from the live-dashboard section. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…es N times` Re-runs a test/command `done when` check N times and requires EVERY run to pass, so "done" means "passes reliably" not "passed once" — a guard against a green that only holds by luck (timing- or order-dependent tests). - parser: an optional `N times` suffix on `passes` / `succeeds` / `finds nothing` and `the test "…" passes`, surfaced as `runs` on the predicate. Only set when > 1, so a plain check (and "1 time") keep the existing single-run shape. The `check:`/`verify:` sugar accepts it too. - verify: ShellVerifier loops `runs` times; the first failing run short-circuits (labelled `run i/N failed —`); a clean pass reports `(passed N/N runs)`. - show/explain: rendered as `×N` (ASCII) and "N times in a row" (prose). Docs: AGENTS.md + MANUAL.md predicate sections. Parser 58/58, runtime 115/115 (4 new verifier tests prove the exact run count + short-circuit). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ge panels, secret redaction, --resume Three features that turn a demo loop into one you can trust unattended: • Judge panels — `done when the skill "x" approves by N judges`. N independent verdicts, majority wins, early-exit once mathematically decided. Each vote is its own skill-verify event (`judge i/N: …`); observe reports the tally. The eval-side counterpart of `passes N times`: flake guard for tests, judge panel for evals. Composes with scores/subject/the bar. • Secret redaction — every event is scrubbed before any sink persists it: values of secret-named env vars (*_TOKEN/_SECRET/_PASSWORD/_KEY/…) become [redacted:<VAR>], and well-known shapes (GitHub/Slack/AWS/sk- keys, JWTs, PEM blocks, Bearer headers, password= assignments) are masked. On by default; LOOP_REDACT=off to disable. Best-effort: a redactor fault falls back to the raw event rather than breaking telemetry. • Resume — `loop run file.loop --resume run.log`. The NDJSON log doubles as a journal: buildResumePlan() replays it (container stack + step depth keep nested sub-runs honest) and the engine skips every unit whose end event says satisfied — whole definitions, pipeline stages, flow steps, foreach items — emitting ⏩ resumed events. Flow handoff summaries now ride on flow-step-end so a resumed flow restores carry-forward context; the log header carries a sha256 of the .loop source and the CLI warns on drift. Interrupted or failed units re-run from scratch: nothing is trusted that a check didn't prove. Tests: parser 61/61 (judge grammar), runtime 132/132 — 6 redaction, 4 panel (majority/early-exit/single-judge parity), 7 resume (log-scan unit tests + end-to-end crash-resume for pipeline, flow w/ summary restore, foreach). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The parser, the ASCII "show" renderer, explain, and the soft linter are all pure TypeScript — one esbuild pass (npm run build:playground) bundles the whole language toolchain to 27 kB of client-side JS. docs/playground.html is a split-pane editor: parse-on-type with inline errors, the compact flow view, lint nudges, and the plain-English explain, plus a picker of examples that showcase the grammar (flake guard, judge panels, pipelines, flows). Linked from the tutorial nav and the hands-on cards. No install, no server. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MANUAL.md gains the section users actually ask for — "How verification works: what 'done' actually depends on": when checks run (the observe node), the short-circuiting conjunction, and the factors that flip a verdict (working dir under target:/worktree isolation, shell env, exit-code semantics, output emptiness for `finds nothing`, the npm-style `test` desugar, flakiness, truncation), how evals decide (subject, the bar, scores, panels), what never auto-passes, and what the verdict triggers. Plus: "Resuming an interrupted run" under telemetry, the redaction guarantees, and the `by N judges` panel in the predicate reference. AGENTS.md and README track the same features. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One ring, one gap at 1–2 o'clock: the loop that hasn't closed yet, still iterating. Replaces the gimbal-detail gyro mark, which turned to mud below 32px. Single flat teal stroke, round caps, crisp at 16px, Vercel-simple. Applied everywhere the mark lives: the favicon SVG (README inherits it) and the inline header lockups in the tutorial, workshop, and game pages. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rial
The icon file (gyro-icon.svg — name kept so all 30 favicon links and the
README hero update with zero churn) is now the solid tile: filled rounded
square, gap-ring knocked out — app-icon treatment of the same mark. The flat
gap ring lives on as docs/logo.svg; page headers keep it inline.
README brought in line with the tutorial and the current feature set:
- hero voice matches the tutorial ("Stop babysitting the agent…")
- links row gains Playground / Workshop / Lab / FAQ
- the pipeline example now actually parses (quoted stage names, real cycle
lines — the old block used pre-v1 syntax); all 5 loop blocks in the README
are verified against the parser
- new "Verify like you mean it" section: done-when conjunction, flake guard,
judge panels, trajectory evals, link to the verification mechanics
- ctx self-equipping paragraph under Skills and memory
- Status/Roadmap reflect what actually shipped (resume, telemetry, playground)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…print.yaml New "You already have a spec (the BMAD flow)" walkthrough in Real Workflows, between the Jira ticket and the Forge case study: copy the load-spec + story-template kit, the three-line for-each driver over YOUR sprint.yaml (backlog stays the source of truth — entry text becomes the story's context), one checklist per story with the pause-and-ask on failure, and the --log/--resume journal so a long sprint survives an interruption. Notes the greenfield path (discover → design produce the spec first) and that the flow is method-neutral — BMAD is one example. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bump for the npm release that ships ctx skill provisioning, verification
reliability (flake guard + judge panels), the event log + secret redaction +
--resume, and the browser playground:
- @loop-lang/{parser,runtime,stdlib,viz} 0.3.0 → 0.4.0 (lockstep)
- @loop-lang/loop (installer) 0.6.0 → 0.7.0
- loopflow (vscode) 0.4.0 → 0.5.0; inter-package dep pins + the extension's
parser devDep synced to 0.4.0 so the bundle knows the new grammar
- CHANGELOG 0.7.0 entry
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ergence ctx capability convergence + reliability & recovery + playground — release 0.7.0
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What & why
The Loop side of the Loop↔ctx skill-provisioning bridge. Today a
.loopnames theskills its loops need (
use skills: a, b) and assumes they already exist in~/.claude/skills— it can't discover or install one. This makes ctx anopt-in skill source: ctx recommends + installs the right skills for a loop's
goal, so the names resolve.
Pairs with the ctx server side: stevesolun/ctx#209 (adds the
ctx__loop_provision/ctx__loop_topupMCP tools this client calls).Grammar
recommend skills with ctxuse skills recommended by ctx [for "<intent>"]top up skills from ctxParser + JSON schema +
AGENTS.md+ the loopflow skill doc all updated;docs/ctx-skill-source.mdandexamples/ctx_skills.loopadded.Runtime
packages/runtime/src/ctx.ts—McpCtxAdapter, an MCP stdio client thatspawns
ctx-mcp-serverand callsctx__loop_provision/ctx__loop_topup.engine.ts— provisions skills before the first plan; tops up after afailed cycle reflects. Merges resolved names into the loop's working skill set.
cli.ts— attaches the adapter only when a file opts in. The MCP SDK is adynamic import, so it stays off the path for every non-ctx run; a missing
server or SDK degrades to a warning. ctx is always optional — the loop runs
the same without it.
Adds
@modelcontextprotocol/sdkand actxLoopEventtype.Also in this PR (bundled per request)
vscode packaging fixes — rename
loop-vscode→loopflow, publisherLoop-Lang,drop
--minify, ignore*.vsix. Unrelated to ctx; folded in to ship as one PR.Tests
Not included
docs/loop-ctx-integration-report.md(a design write-up) was left out — itdescribes a ctx module
ctx.adapters.loopthat doesn't match what shipped in#209 (
ctx.adapters.generic.loop_tools). Can update + add it if wanted.