Skip to content

ctx capability convergence + reliability & recovery + playground — release 0.7.0 - #9

Merged
tickets-forge-dev merged 14 commits into
feat/ctx-skill-integrationfrom
feat/ctx-capability-convergence
Jul 2, 2026
Merged

tickets-forge-dev merged 14 commits into
feat/ctx-skill-integrationfrom
feat/ctx-capability-convergence

Conversation

@tickets-forge-dev

@tickets-forge-dev tickets-forge-dev commented Jun 30, 2026 •

Copy link
Copy Markdown
Owner

Everything since 0.6.0, released as 0.7.0:

ctx integration (the branch's original scope)

  • ctx as an opt-in skill source (recommend skills with ctx, run-time provision + top-up)
  • capability grants (grant ctx: skills, agents, mcps, harnesses, fail-closed) + own-model gating (dry-run harness recs)
  • LOOP_EVENTS_URL control-plane telemetry sink

Verification reliability

  • flake guard: done when "…" passes 3 times (every run must pass, short-circuits)
  • judge panels: the skill "…" approves by 3 judges (majority, early-exit once decided)

Telemetry, redaction, resume

  • --log / LOOP_LOG_FILE: durable NDJSON event journal
  • secret redaction on by default before any sink persists (env-derived + pattern-based; LOOP_REDACT=off)
  • --resume run.log: skip units the log proves satisfied; flow carry-forward restored; source-hash drift warning

Site & docs

  • browser playground (parser+show+explain+lint client-side, 27 kB)
  • "How verification works" mechanics section; event-log/resume/redaction docs
  • spec-driven (PRD + sprint.yaml) walkthrough in the tutorial
  • new logo: gap ring mark + solid-tile icon; README aligned with the tutorial (all loop blocks parser-verified)

Release

  • libs 0.4.0 · @loop-lang/loop 0.7.0 · loopflow (vscode) 0.5.0 · CHANGELOG entry

Tests: parser 61 · runtime 132 · stdlib 3 · viz 10 · vscode 19 — all green.

🤖 Generated with Claude Code

tickets-forge-dev and others added 14 commits June 30, 2026 14:27
…odel gating

Extend the ctx skill-source integration beyond skills to the full capability
set, behind a fail-closed permission model.

- DSL (additive, config tier): `grant ctx: skills, agents, mcps, harnesses`
  and `ctx may use my own model "<provider>/<model>"`.
- Runtime: CtxProvisionResult carries capability groups + harnessInstall +
  warnings; McpCtxAdapter threads permissions/own-model and parses the
  ctx.loop_adapter.v1 contract; the engine merges skills+agents and surfaces
  mcps/harnesses on the ctx event (never auto-installs); the cli passes grants.
- Schema, AGENTS.md, docs/ctx-skill-source.md, docs/ctx-integration-guide.md,
  examples/ctx_capabilities.loop.
- Back-compat: no grant => skills+agents, exactly as before.

Tests: parser 54, runtime 101 (4 new), full build green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`ctx may use my own model "ollama/…"` only unlocks ctx harness recommendations;
it never makes the loop run on that model. Warn at run time when the provider's
local binary (e.g. ollama) isn't on PATH, so the author isn't surprised that
running a recommended harness later would fail. API providers (no local binary)
never warn.

New ownModel.ts (ownModelBinaryWarning + commandOnPath), wired into the cli ctx
path, +1 runtime test (injected PATH lookup).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Stream every LoopEvent to a collector when LOOP_EVENTS_URL is set (+ LOOP_EVENTS_TOKEN, LOOP_RUN_ID). Best-effort, fire-and-forget, flushes on exit; silent when unset/unreachable. +3 tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Moved agentic-engineering.loop from root to examples/
- Moved demo.loop from root to examples/
- Updated references in test and docs to reflect new paths

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
Records the full loop event stream to a local file, the durable sibling of
the LOOP_EVENTS_URL control-plane sink. Every event the engine already emits
(loop/node/observe/reflect/transition/human/ctx/git/hook/model/stop) appends
as one NDJSON line: a `loop.log.v1` header then `{seq, ts, event}` per event.

- makeFileEventSink: synchronous appendFileSync per event, so the log is
  durable the instant it fires — survives Ctrl-C / crash with no lost tail
  (flush() is a no-op). Best-effort like the HTTP sink: creates the parent
  dir once, disables itself quietly on any write error, never throws.
- combineSinks: fan one event stream out to HTTP + file, sharing one runId so
  the collector and the local log correlate. eventSinkFromEnv wires both, so
  all three run paths (--events, --live, default) log with no cli.ts churn.
- --log <path> flag overrides LOOP_LOG_FILE for one-off runs.

Off by default: no env, no flag → no sink, behaves exactly as before.
Tests: 5 sink + override/precedence cases; runtime suite 111/111.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…P_EVENTS_URL)

MANUAL.md §4 gains an "Event log & telemetry" section: what the event stream
records, the NDJSON format (header + seq'd event lines) with jq read-back
examples, durability + best-effort guarantees, and the env-var table for the
local log and the remote HTTP collector. Usage line + flag list pick up --log.
README points at it from the live-dashboard section.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…es N times`

Re-runs a test/command `done when` check N times and requires EVERY run to pass,
so "done" means "passes reliably" not "passed once" — a guard against a green
that only holds by luck (timing- or order-dependent tests).

- parser: an optional `N times` suffix on `passes` / `succeeds` / `finds nothing`
  and `the test "…" passes`, surfaced as `runs` on the predicate. Only set when
  > 1, so a plain check (and "1 time") keep the existing single-run shape. The
  `check:`/`verify:` sugar accepts it too.
- verify: ShellVerifier loops `runs` times; the first failing run short-circuits
  (labelled `run i/N failed —`); a clean pass reports `(passed N/N runs)`.
- show/explain: rendered as `×N` (ASCII) and "N times in a row" (prose).

Docs: AGENTS.md + MANUAL.md predicate sections. Parser 58/58, runtime 115/115
(4 new verifier tests prove the exact run count + short-circuit).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ge panels, secret redaction, --resume

Three features that turn a demo loop into one you can trust unattended:

• Judge panels — `done when the skill "x" approves by N judges`. N independent
  verdicts, majority wins, early-exit once mathematically decided. Each vote is
  its own skill-verify event (`judge i/N: …`); observe reports the tally. The
  eval-side counterpart of `passes N times`: flake guard for tests, judge panel
  for evals. Composes with scores/subject/the bar.

• Secret redaction — every event is scrubbed before any sink persists it:
  values of secret-named env vars (*_TOKEN/_SECRET/_PASSWORD/_KEY/…) become
  [redacted:<VAR>], and well-known shapes (GitHub/Slack/AWS/sk- keys, JWTs,
  PEM blocks, Bearer headers, password= assignments) are masked. On by
  default; LOOP_REDACT=off to disable. Best-effort: a redactor fault falls
  back to the raw event rather than breaking telemetry.

• Resume — `loop run file.loop --resume run.log`. The NDJSON log doubles as a
  journal: buildResumePlan() replays it (container stack + step depth keep
  nested sub-runs honest) and the engine skips every unit whose end event says
  satisfied — whole definitions, pipeline stages, flow steps, foreach items —
  emitting ⏩ resumed events. Flow handoff summaries now ride on flow-step-end
  so a resumed flow restores carry-forward context; the log header carries a
  sha256 of the .loop source and the CLI warns on drift. Interrupted or failed
  units re-run from scratch: nothing is trusted that a check didn't prove.

Tests: parser 61/61 (judge grammar), runtime 132/132 — 6 redaction, 4 panel
(majority/early-exit/single-judge parity), 7 resume (log-scan unit tests +
end-to-end crash-resume for pipeline, flow w/ summary restore, foreach).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The parser, the ASCII "show" renderer, explain, and the soft linter are all
pure TypeScript — one esbuild pass (npm run build:playground) bundles the
whole language toolchain to 27 kB of client-side JS. docs/playground.html is
a split-pane editor: parse-on-type with inline errors, the compact flow view,
lint nudges, and the plain-English explain, plus a picker of examples that
showcase the grammar (flake guard, judge panels, pipelines, flows). Linked
from the tutorial nav and the hands-on cards. No install, no server.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MANUAL.md gains the section users actually ask for — "How verification works:
what 'done' actually depends on": when checks run (the observe node), the
short-circuiting conjunction, and the factors that flip a verdict (working
dir under target:/worktree isolation, shell env, exit-code semantics, output
emptiness for `finds nothing`, the npm-style `test` desugar, flakiness,
truncation), how evals decide (subject, the bar, scores, panels), what never
auto-passes, and what the verdict triggers. Plus: "Resuming an interrupted
run" under telemetry, the redaction guarantees, and the `by N judges` panel
in the predicate reference. AGENTS.md and README track the same features.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One ring, one gap at 1–2 o'clock: the loop that hasn't closed yet, still
iterating. Replaces the gimbal-detail gyro mark, which turned to mud below
32px. Single flat teal stroke, round caps, crisp at 16px, Vercel-simple.
Applied everywhere the mark lives: the favicon SVG (README inherits it) and
the inline header lockups in the tutorial, workshop, and game pages.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rial

The icon file (gyro-icon.svg — name kept so all 30 favicon links and the
README hero update with zero churn) is now the solid tile: filled rounded
square, gap-ring knocked out — app-icon treatment of the same mark. The flat
gap ring lives on as docs/logo.svg; page headers keep it inline.

README brought in line with the tutorial and the current feature set:
- hero voice matches the tutorial ("Stop babysitting the agent…")
- links row gains Playground / Workshop / Lab / FAQ
- the pipeline example now actually parses (quoted stage names, real cycle
  lines — the old block used pre-v1 syntax); all 5 loop blocks in the README
  are verified against the parser
- new "Verify like you mean it" section: done-when conjunction, flake guard,
  judge panels, trajectory evals, link to the verification mechanics
- ctx self-equipping paragraph under Skills and memory
- Status/Roadmap reflect what actually shipped (resume, telemetry, playground)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…print.yaml

New "You already have a spec (the BMAD flow)" walkthrough in Real Workflows,
between the Jira ticket and the Forge case study: copy the load-spec +
story-template kit, the three-line for-each driver over YOUR sprint.yaml
(backlog stays the source of truth — entry text becomes the story's context),
one checklist per story with the pause-and-ask on failure, and the
--log/--resume journal so a long sprint survives an interruption. Notes the
greenfield path (discover → design produce the spec first) and that the flow
is method-neutral — BMAD is one example.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bump for the npm release that ships ctx skill provisioning, verification
reliability (flake guard + judge panels), the event log + secret redaction +
--resume, and the browser playground:
- @loop-lang/{parser,runtime,stdlib,viz} 0.3.0 → 0.4.0 (lockstep)
- @loop-lang/loop (installer) 0.6.0 → 0.7.0
- loopflow (vscode) 0.4.0 → 0.5.0; inter-package dep pins + the extension's
  parser devDep synced to 0.4.0 so the bundle knows the new grammar
- CHANGELOG 0.7.0 entry

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@tickets-forge-dev tickets-forge-dev changed the title feat(ctx): capability grants — agents, MCP servers, harnesses + own-model gating ctx capability convergence + reliability & recovery + playground — release 0.7.0 Jul 2, 2026
@tickets-forge-dev
tickets-forge-dev merged commit 67b5d52 into feat/ctx-skill-integration Jul 2, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant