Skip to content

feat(ctx): integrate ctx as an opt-in skill source for loops - #8

Merged
tickets-forge-dev merged 16 commits into
masterfrom
feat/ctx-skill-integration
Jul 2, 2026
Merged

tickets-forge-dev merged 16 commits into
masterfrom
feat/ctx-skill-integration

Conversation

@tickets-forge-dev

Copy link
Copy Markdown
Owner

What & why

The Loop side of the Loop↔ctx skill-provisioning bridge. Today a .loop names the
skills its loops need (use skills: a, b) and assumes they already exist in
~/.claude/skills — it can't discover or install one. This makes ctx an
opt-in skill source: ctx recommends + installs the right skills for a loop's
goal, so the names resolve.

Pairs with the ctx server side: stevesolun/ctx#209 (adds the
ctx__loop_provision / ctx__loop_topup MCP tools this client calls).

Grammar

tier syntax
config recommend skills with ctx
loop use skills recommended by ctx [for "<intent>"]
loop top up skills from ctx

Parser + JSON schema + AGENTS.md + the loopflow skill doc all updated;
docs/ctx-skill-source.md and examples/ctx_skills.loop added.

Runtime

  • packages/runtime/src/ctx.ts — McpCtxAdapter, an MCP stdio client that
    spawns ctx-mcp-server and calls ctx__loop_provision / ctx__loop_topup.
  • engine.ts — provisions skills before the first plan; tops up after a
    failed cycle reflects. Merges resolved names into the loop's working skill set.
  • cli.ts — attaches the adapter only when a file opts in. The MCP SDK is a
    dynamic import, so it stays off the path for every non-ctx run; a missing
    server or SDK degrades to a warning. ctx is always optional — the loop runs
    the same without it.

Adds @modelcontextprotocol/sdk and a ctx LoopEvent type.

Also in this PR (bundled per request)

vscode packaging fixes — rename loop-vscode→loopflow, publisher Loop-Lang,
drop --minify, ignore *.vsix. Unrelated to ctx; folded in to ship as one PR.

Tests

  • parser: 51 pass / 0 fail (incl. 4 new ctx grammar tests)
  • runtime: 100 pass / 0 fail (incl. new ctx provision/top-up engine tests)
  • both packages build clean.

Not included

docs/loop-ctx-integration-report.md (a design write-up) was left out — it
describes a ctx module ctx.adapters.loop that doesn't match what shipped in
#209 (ctx.adapters.generic.loop_tools). Can update + add it if wanted.

tickets-forge-dev and others added 16 commits June 30, 2026 12:20
Add the Loop side of the Loop↔ctx skill-provisioning bridge: a `.loop` can let
ctx recommend + install the skills its loops need instead of assuming they
already exist in ~/.claude/skills.

Grammar (parser + schema + AGENTS.md + loopflow skill):
  - config tier: `recommend skills with ctx`
  - loop: `use skills recommended by ctx [for "<intent>"]`, `top up skills from ctx`

Runtime:
  - McpCtxAdapter (packages/runtime/src/ctx.ts) speaks MCP to a `ctx-mcp-server`
    and calls ctx__loop_provision / ctx__loop_topup.
  - engine provisions skills before the first plan and tops up after a failed
    cycle; cli attaches the adapter only when a file opts in (dynamic import
    keeps the MCP SDK off every other run). A missing server/SDK degrades to a
    warning — ctx is always optional, the loop runs the same without it.

Adds @modelcontextprotocol/sdk, the ctx event type, docs/ctx-skill-source.md,
examples/ctx_skills.loop, and parser/engine tests.

Also bundles vscode packaging fixes (rename loop-vscode→loopflow, publisher
Loop-Lang, drop --minify, ignore *.vsix) per request to ship as one PR.

Server side: stevesolun/ctx#209.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…odel gating

Extend the ctx skill-source integration beyond skills to the full capability
set, behind a fail-closed permission model.

- DSL (additive, config tier): `grant ctx: skills, agents, mcps, harnesses`
  and `ctx may use my own model "<provider>/<model>"`.
- Runtime: CtxProvisionResult carries capability groups + harnessInstall +
  warnings; McpCtxAdapter threads permissions/own-model and parses the
  ctx.loop_adapter.v1 contract; the engine merges skills+agents and surfaces
  mcps/harnesses on the ctx event (never auto-installs); the cli passes grants.
- Schema, AGENTS.md, docs/ctx-skill-source.md, docs/ctx-integration-guide.md,
  examples/ctx_capabilities.loop.
- Back-compat: no grant => skills+agents, exactly as before.

Tests: parser 54, runtime 101 (4 new), full build green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`ctx may use my own model "ollama/…"` only unlocks ctx harness recommendations;
it never makes the loop run on that model. Warn at run time when the provider's
local binary (e.g. ollama) isn't on PATH, so the author isn't surprised that
running a recommended harness later would fail. API providers (no local binary)
never warn.

New ownModel.ts (ownModelBinaryWarning + commandOnPath), wired into the cli ctx
path, +1 runtime test (injected PATH lookup).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Stream every LoopEvent to a collector when LOOP_EVENTS_URL is set (+ LOOP_EVENTS_TOKEN, LOOP_RUN_ID). Best-effort, fire-and-forget, flushes on exit; silent when unset/unreachable. +3 tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Moved agentic-engineering.loop from root to examples/
- Moved demo.loop from root to examples/
- Updated references in test and docs to reflect new paths

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
Records the full loop event stream to a local file, the durable sibling of
the LOOP_EVENTS_URL control-plane sink. Every event the engine already emits
(loop/node/observe/reflect/transition/human/ctx/git/hook/model/stop) appends
as one NDJSON line: a `loop.log.v1` header then `{seq, ts, event}` per event.

- makeFileEventSink: synchronous appendFileSync per event, so the log is
  durable the instant it fires — survives Ctrl-C / crash with no lost tail
  (flush() is a no-op). Best-effort like the HTTP sink: creates the parent
  dir once, disables itself quietly on any write error, never throws.
- combineSinks: fan one event stream out to HTTP + file, sharing one runId so
  the collector and the local log correlate. eventSinkFromEnv wires both, so
  all three run paths (--events, --live, default) log with no cli.ts churn.
- --log <path> flag overrides LOOP_LOG_FILE for one-off runs.

Off by default: no env, no flag → no sink, behaves exactly as before.
Tests: 5 sink + override/precedence cases; runtime suite 111/111.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…P_EVENTS_URL)

MANUAL.md §4 gains an "Event log & telemetry" section: what the event stream
records, the NDJSON format (header + seq'd event lines) with jq read-back
examples, durability + best-effort guarantees, and the env-var table for the
local log and the remote HTTP collector. Usage line + flag list pick up --log.
README points at it from the live-dashboard section.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…es N times`

Re-runs a test/command `done when` check N times and requires EVERY run to pass,
so "done" means "passes reliably" not "passed once" — a guard against a green
that only holds by luck (timing- or order-dependent tests).

- parser: an optional `N times` suffix on `passes` / `succeeds` / `finds nothing`
  and `the test "…" passes`, surfaced as `runs` on the predicate. Only set when
  > 1, so a plain check (and "1 time") keep the existing single-run shape. The
  `check:`/`verify:` sugar accepts it too.
- verify: ShellVerifier loops `runs` times; the first failing run short-circuits
  (labelled `run i/N failed —`); a clean pass reports `(passed N/N runs)`.
- show/explain: rendered as `×N` (ASCII) and "N times in a row" (prose).

Docs: AGENTS.md + MANUAL.md predicate sections. Parser 58/58, runtime 115/115
(4 new verifier tests prove the exact run count + short-circuit).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ge panels, secret redaction, --resume

Three features that turn a demo loop into one you can trust unattended:

• Judge panels — `done when the skill "x" approves by N judges`. N independent
  verdicts, majority wins, early-exit once mathematically decided. Each vote is
  its own skill-verify event (`judge i/N: …`); observe reports the tally. The
  eval-side counterpart of `passes N times`: flake guard for tests, judge panel
  for evals. Composes with scores/subject/the bar.

• Secret redaction — every event is scrubbed before any sink persists it:
  values of secret-named env vars (*_TOKEN/_SECRET/_PASSWORD/_KEY/…) become
  [redacted:<VAR>], and well-known shapes (GitHub/Slack/AWS/sk- keys, JWTs,
  PEM blocks, Bearer headers, password= assignments) are masked. On by
  default; LOOP_REDACT=off to disable. Best-effort: a redactor fault falls
  back to the raw event rather than breaking telemetry.

• Resume — `loop run file.loop --resume run.log`. The NDJSON log doubles as a
  journal: buildResumePlan() replays it (container stack + step depth keep
  nested sub-runs honest) and the engine skips every unit whose end event says
  satisfied — whole definitions, pipeline stages, flow steps, foreach items —
  emitting ⏩ resumed events. Flow handoff summaries now ride on flow-step-end
  so a resumed flow restores carry-forward context; the log header carries a
  sha256 of the .loop source and the CLI warns on drift. Interrupted or failed
  units re-run from scratch: nothing is trusted that a check didn't prove.

Tests: parser 61/61 (judge grammar), runtime 132/132 — 6 redaction, 4 panel
(majority/early-exit/single-judge parity), 7 resume (log-scan unit tests +
end-to-end crash-resume for pipeline, flow w/ summary restore, foreach).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The parser, the ASCII "show" renderer, explain, and the soft linter are all
pure TypeScript — one esbuild pass (npm run build:playground) bundles the
whole language toolchain to 27 kB of client-side JS. docs/playground.html is
a split-pane editor: parse-on-type with inline errors, the compact flow view,
lint nudges, and the plain-English explain, plus a picker of examples that
showcase the grammar (flake guard, judge panels, pipelines, flows). Linked
from the tutorial nav and the hands-on cards. No install, no server.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MANUAL.md gains the section users actually ask for — "How verification works:
what 'done' actually depends on": when checks run (the observe node), the
short-circuiting conjunction, and the factors that flip a verdict (working
dir under target:/worktree isolation, shell env, exit-code semantics, output
emptiness for `finds nothing`, the npm-style `test` desugar, flakiness,
truncation), how evals decide (subject, the bar, scores, panels), what never
auto-passes, and what the verdict triggers. Plus: "Resuming an interrupted
run" under telemetry, the redaction guarantees, and the `by N judges` panel
in the predicate reference. AGENTS.md and README track the same features.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One ring, one gap at 1–2 o'clock: the loop that hasn't closed yet, still
iterating. Replaces the gimbal-detail gyro mark, which turned to mud below
32px. Single flat teal stroke, round caps, crisp at 16px, Vercel-simple.
Applied everywhere the mark lives: the favicon SVG (README inherits it) and
the inline header lockups in the tutorial, workshop, and game pages.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rial

The icon file (gyro-icon.svg — name kept so all 30 favicon links and the
README hero update with zero churn) is now the solid tile: filled rounded
square, gap-ring knocked out — app-icon treatment of the same mark. The flat
gap ring lives on as docs/logo.svg; page headers keep it inline.

README brought in line with the tutorial and the current feature set:
- hero voice matches the tutorial ("Stop babysitting the agent…")
- links row gains Playground / Workshop / Lab / FAQ
- the pipeline example now actually parses (quoted stage names, real cycle
  lines — the old block used pre-v1 syntax); all 5 loop blocks in the README
  are verified against the parser
- new "Verify like you mean it" section: done-when conjunction, flake guard,
  judge panels, trajectory evals, link to the verification mechanics
- ctx self-equipping paragraph under Skills and memory
- Status/Roadmap reflect what actually shipped (resume, telemetry, playground)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…print.yaml

New "You already have a spec (the BMAD flow)" walkthrough in Real Workflows,
between the Jira ticket and the Forge case study: copy the load-spec +
story-template kit, the three-line for-each driver over YOUR sprint.yaml
(backlog stays the source of truth — entry text becomes the story's context),
one checklist per story with the pause-and-ask on failure, and the
--log/--resume journal so a long sprint survives an interruption. Notes the
greenfield path (discover → design produce the spec first) and that the flow
is method-neutral — BMAD is one example.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bump for the npm release that ships ctx skill provisioning, verification
reliability (flake guard + judge panels), the event log + secret redaction +
--resume, and the browser playground:
- @loop-lang/{parser,runtime,stdlib,viz} 0.3.0 → 0.4.0 (lockstep)
- @loop-lang/loop (installer) 0.6.0 → 0.7.0
- loopflow (vscode) 0.4.0 → 0.5.0; inter-package dep pins + the extension's
  parser devDep synced to 0.4.0 so the bundle knows the new grammar
- CHANGELOG 0.7.0 entry

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ergence

ctx capability convergence + reliability & recovery + playground — release 0.7.0
@tickets-forge-dev
tickets-forge-dev merged commit 668944f into master Jul 2, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant