Skip to content

Enforce every hard stop outside the agent's reach; correct and condense the knowledge base - #26

Open
espi wants to merge 3 commits into
mainfrom
claude/twitter-post-requirements-eq7h5r
Open

espi wants to merge 3 commits into
mainfrom
claude/twitter-post-requirements-eq7h5r

Conversation

@espi

@espi espi commented Sep 29, 2026

Copy link
Copy Markdown
Owner

This is the follow-up from reviewing and merging #21–#25. It has three commits, and each can be reviewed or reverted on its own.

1. 46d7214: the loops this repo ships could fail to stop

Two findings from #25 apply to this repo's own templates. I checked both against the primary sources before changing anything.

  • Claude Code < v2.1.281 could retry a turn indefinitely while ignoring --max-turns (changelog, Sep 23). templates/ralph/run.sh passed no per-call cap and had no timeout. Its cost meter returned 0 unless you wired one in. Its counter only ticks between calls, so a runaway call would never return to it.
  • The ralph-wiggum plugin's iteration cap is in a file the agent can write. It's stored in .claude/ralph-loop.local.md, and stop-hook.sh:51 only enforces it when max_iterations is greater than 0. An agent can remove its own cap. new-loop, first-loop.md, overnight-build.md and ui-audit-overnight.md all recommended the plugin as the cap.

What run.sh does now:

  • It refuses to run on Claude Code < 2.1.281, or if there's no timeout/gtimeout.
  • Every call gets --max-turns, --max-budget-usd (set to the remaining budget) and a wall-clock timeout.
  • Spend is read from the CLI's own --output-format json total_cost_usd. If a call times out or returns no cost, the loop stops instead of guessing.
  • PROTECTED_PATHS fingerprints the prompt, harness and tests. If the agent edits its own checks, the loop stops.
  • Stall detection hashes the whole worktree, so empty commits don't count as progress.
  • It is deny-by-default (--permission-prompts none) with an --allowedTools allowlist.

Tested: every exit path against a stub CLI (old version, done, bad check, stall, tampering, budget, iteration cap, no JSON, timeout), plus one real run on v2.1.284 (it metered $0.043 and exited 0).

Documentation updated to match:

  • guardrails/ (README, checklist, budget.env, which gains per-call caps and CLAUDE_CODE_MAX_TURNS)
  • the loop-guardrails and new-loop skills
  • CLAUDE.md, with the principle "the meter must not live inside the thing it meters"
  • three runbooks. The UI-audit runbook now runs headless through run.sh; the plugin is a supervised fallback only.

self-edit-guard now runs on PRs to any base, and on edited. Stacked PRs from the routine (#23, #25) and PRs retargeted to main previously skipped it. You approved this change.

2. 9637d3c: the primer gets a size budget

The primer grew from 559 lines (Jul 21) to 3,303 (Sep 28). Nothing capped it, so a one-off trim would have grown back.

  • Step 5 of the update-knowledge skill: the primer changes only when the state of the practice changes, and detail goes to sources.md. This is an unprotected region, and a human-authored change, not a self-edit.
  • The new kb-size.yml fails any PR that leaves the primer over 750 lines.

3. 09e3980: knowledge-base corrections and condensed primer

Corrections (verified against primaries):

  • claude plugin eval was not "the three hard stops as a CLI flag set". Its caps are per-case frontmatter fields, it has no stall detection, and it wasn't the first turn-plus-dollar cap.
  • Three version numbers are fixed.
  • SWE-Proof is pinned to v1.
  • SaltBench's rule is filed under hard stop knowledge: 2026-06-22 update pass (7-day delta from Jun 15) #3.
  • The METR URL is added.
  • arXiv:2609.19844 is re-characterised.
  • A stale note about the weekly whats-new digest is corrected.

Backlog:

Primer: 3,303 → 661 lines.

  • The §1–§7 structure is kept, and every load-bearing fact still carries a confidence tag.
  • All removed prose is kept verbatim in archive/primer-detail-2026-09.md, and the per-claim records stay in sources.md.
  • A line-level check shows that the only old primer lines not present verbatim somewhere in knowledge/ are the ones these corrections replaced.

Worth a closer look

  • Behaviour changes in templates/ralph/run.sh.
    • It now stops when a call times out, because the call's spend is unknown.
    • It needs timeout. On macOS, that means brew install coreutils.
    • The default AGENT_CMD is deny-by-default, so a build loop needs an --allowedTools allowlist.
  • Open question: review-tool spend caps. The condensing agent flagged a contradiction in the old §5A about which review tools ship enforced spend caps. The new text avoids the contradictory blanket claim; you may want to settle it.
  • What CI checks. Both workflows are checked locally (YAML valid; self-edit-guard OK). kb-size will run for the first time on this PR.

🤖 Generated with Claude Code

https://claude.ai/code/session_01YAFQxQa6sakZp18SoDhMP7


Generated by Claude Code

PR #25's research showed two ways this repo's own loops could fail to stop,
both verified against primary sources before this change:

- Claude Code < v2.1.281 could retry a turn indefinitely, ignoring
  --max-turns (changelog, Sep 23 2026). templates/ralph/run.sh passed no
  per-call cap, no timeout, and its cost meter returned 0 unless wired, so a
  single runaway `claude -p` call would never return to the counter.
- The ralph-wiggum plugin keeps its iteration count in
  .claude/ralph-loop.local.md inside the agent's worktree, and its stop hook
  only enforces it while max_iterations > 0 (plugins/ralph-wiggum/hooks/
  stop-hook.sh:51). An agent can remove its own cap.

run.sh now:
- refuses to run on Claude Code < 2.1.281, or without timeout/gtimeout;
- caps each call from outside: --max-turns, --max-budget-usd (the remaining
  budget), and a wall-clock timeout;
- meters spend from the CLI's own JSON result (total_cost_usd) instead of a
  stub; stops rather than guesses when a call times out or returns no cost;
- fingerprints PROTECTED_PATHS (prompt, harness, tests) and stops if the
  agent edits its own checks;
- detects stalls from the full worktree, so empty commits aren't progress;
- defaults to deny-by-default permissions (--permission-prompts none).
Tested: every exit path against a stub CLI, plus one real run (v2.1.284).

Docs: guardrails/README + checklist + budget.env, loop-guardrails and
new-loop skills, CLAUDE.md, first-loop / overnight-build / ui-audit runbooks
now require per-call caps and v2.1.281+, and stop recommending the plugin as
the only cap on unattended runs.

self-edit-guard now runs on PRs to any base and on `edited`, so stacked and
retargeted routine PRs are checked (human-approved change).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YAFQxQa6sakZp18SoDhMP7
The primer grew from 559 lines (Jul 21) to 3,303 (Sep 28) — each weekly pass
appended per-release changelog detail, paper digests and correction
narratives, and the last pass alone added ~940 lines. Step 5 of the
update-knowledge skill told the routine to update the primer for every new
fact and nothing capped it, so a one-off trim would regrow within weeks.

- update-knowledge SKILL.md step 5 (unprotected region; human-authored, not
  a self-edit): the primer changes only when the state of the practice
  changes; detail goes to sources.md; prefer replacing to appending; stay
  <= 750 lines, condensing in the same PR if needed.
- .github/workflows/kb-size.yml: fails any PR that leaves the primer over 750
  lines. A budget the agent grades itself on is no budget.
- CLAUDE.md: records the budget under "Knowledge base discipline".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YAFQxQa6sakZp18SoDhMP7
…ecided calls

Review of PRs #22, #23 and #25 (merged 2026-09-29), each finding verified
against primary sources: the raw changelog, plugin-evals docs, arXiv API,
metr.org and the ralph-wiggum stop-hook.sh.

Corrections:
- `claude plugin eval` was framed as "the three hard stops shipped as a CLI
  flag set". Its turn and time caps are per-case frontmatter fields, not
  flags; it has no stall detection; and it wasn't the first turn+dollar cap.
- Three version numbers were wrong (Monitor deadline and modelPricing
  multiplier are v2.1.271; the tool_use_id fix is v2.1.274).
- SWE-Proof is pinned to v1, SaltBench's budget-stop rule moves to hard
  stop #3, the METR URL is added, and arXiv:2609.19844 is re-characterised.
- The stale "whats-new digest discontinued" note is corrected.
- Added (High, source read): the ralph-wiggum plugin's cap sits in a
  worktree file, and `max_iterations: 0` disables it.

Backlog:
- The CSA/METR "~700 agents" item was restored; the correction that
  dropped it was wrong, since METR does say 700.
- The OpenAI research-spend caveat was archived; both items had been
  dropped without archiving.
- Five human calls decided by this PR's guardrail changes (and #21) move
  to the archive.

Primer: 3,303 -> 661 lines, back to a briefing. All removed prose is
verbatim in archive/primer-detail-2026-09.md, and per-claim records stay in
sources.md. A line-level check confirms the only primer text not present
verbatim elsewhere is the text these corrections replaced.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YAFQxQa6sakZp18SoDhMP7
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants