Skip to content

guardrails: containment as a required companion, verify-the-agent rule, and retiring a subagent cap that no longer exists - #21

Merged
espi merged 1 commit into
mainfrom
claude/focused-clarke-x07iz3
Sep 29, 2026
Merged

espi merged 1 commit into
mainfrom
claude/focused-clarke-x07iz3

Conversation

@espi

@espi espi commented Sep 7, 2026

Copy link
Copy Markdown
Owner

Implements the five suggestions you confirmed from PR #20's "Routine self-improvements" section — recommendations 1B, 2A, 3A, 4B, 5B.

All of these sit outside the routine's auto-apply envelope by design, which is exactly why they couldn't ride the knowledge PR. Every commit here is human-authored on your instruction; there is no self-edit: commit, and self-edit-guard passes with 0 checked.


What changed, by decision

S1 → B: containment as a required companion, not a fourth hard stop

guardrails/README.md, guardrails/checklist.md

The three hard stops bound what a loop spends and how long it runs. Neither bounds what an attacker-supplied input makes the harness do — and the known mechanisms fire before any stop condition or verification step runs:

  • GitSpawn executes attacker code during the agent's startup git status, before the workspace-trust prompt and outside the command sandbox. Triggered by cloning the target repo.
  • Instruction privilege escalation re-labels tool-level content as a genuine user message; automatic permission review does not stop it.

This repo's model is "loops run from here against other repositories" — we clone untrusted code by design, which is precisely the premise both attacks need.

Written as a required section rather than a fourth numbered rule, per your call: the three hard stops are single checkable numbers, containment is a posture made of several settings. Keeping the list at three preserves its force; skipping the section is still incomplete. Defaults cover scoped egress, unreachable credentials, deny-by-default unattended runs, and real isolation — with a note that a network proxy alone isn't sufficient, since agents have been observed editing /etc/hosts to route around one.

S5 → B: --permission-prompts none, and an honest note on --restricted

guardrails/README.md

Documents the v2.1.259 flag as the cheapest containment win available — including that it removes tools needing a human answer and surfaces denials in stream-json output, so a harness can see what it blocked. ant apply was not included, per your pick.

Also records what the research pass actually established about --restricted: the primary docs describe only tool removal and settings scoping, say nothing about OS-level isolation, and the flag is absent from Anthropic's own sandbox-environments comparison page. So it's documented as a permission gate, not a sandbox — the honest reading, and one that matters because it's tempting to reach for.

S2 → A: this routine's threat model, written down

runbooks/staying-current.md

The research these routines do now describes an attack on the shape these routines are. Previously that reasoning lived only in reviewers' heads.

The section states the exposure plainly (scheduled, unattended, reads the open web, no permission prompts, IPE reproduces through /schedule-shaped features, neither paper proposes a mitigation) and then why it's still safe to run: the defenses are structural and sit outside the agent's judgement — never merges, blast radius limited to knowledge/, CI-enforced fence, claude/-only pushes. None of them depends on a compromised agent correctly noticing it's been manipulated.

It ends with what would invalidate this — enabling unrestricted pushes, widening the auto-edit envelope, removing self-edit-guard, adding connectors that reach beyond this repo. The load-bearing property is that the routine's autonomy is bounded by things it cannot edit, and each of those is a way of accidentally trading it away.

S3 → A: verify the agent, don't relay it

.claude/skills/update-knowledge/SKILL.md (step 4)

Step 4 now says a subagent's summary is not a primary source — it's a claim about one — and mandates two checks: every arXiv citation machine-verified against the arXiv API (ID, exact title, v1 date, category), and every version attribution checked against the feature's own doc page rather than the nearest changelog heading.

Both earned their place this week: the arXiv check has caught a misdated ID in an earlier pass, and the version rule is what surfaced that modelPricing is v2.1.242, not v2.1.243.

Note for review: step 4 is a protected region under self-edit-guard. This is a normal human-authored commit with no self-edit: prefix — the exact route the guard's own failure message names ("make it a normal human-authored commit"). The gate did its job: the routine surfaced the idea and could not apply it.

S4 → declined

No change. sources.md already carries the whats-new digest status and is read every pass at step 1, so editing the skill would be churn — and wrong text in the least-reviewed file if the digest resumes.


Two things I did beyond the five — please review these specifically

Both are consequences of your decisions rather than new ideas, but you should see them called out rather than buried in the diff.

1. loop-guardrails gains containment as check 6. That skill declares "the standard is guardrails/checklist.md" and then enumerates five items. Adding a required section to the checklist without teaching the skill to audit it would mean loop-guardrails reports a loop clean while skipping it — required in name only. It's reported as ✅/⚠️/❌ rather than pass-fail, since unlike caps 1–3 it isn't a single number.

2. A removed env var that four files still recommended. CLAUDE_CODE_MAX_SUBAGENTS_PER_SESSION was removed in v2.1.224 (Aug 7) and is gone from the docs entirely — the sub-agents page now says "There's no limit on the total number of subagents Claude can spawn over a session." It was still being recommended in guardrails/README.md, runbooks/staying-current.md (three places, including both Routine config tables), and artifact-audit's own guardrails section. Setting it today does nothing.

That last one is the uncomfortable find: our guardrails files were telling readers to rely on a cap that doesn't exist. Replaced with CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS and CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH, with an explicit note that neither bounds a session's total lifetime spawns — that ceiling is now yours to enforce. It's also a neat illustration of the repo's own thesis: a vendor default was deleted underneath us, which is why the hard stops belong in your harness.

This is strictly artifact-audit's job, not update-knowledge's. I fixed it here because the knowledge to do so had just landed in the same session and the files were already open — but the monthly audit should have caught it, and didn't. Worth asking why.


Verification

🤖 Generated with Claude Code

https://claude.ai/code/session_01Ca7z6e9N5SAhNjqM8gFM5R


Generated by Claude Code

…e; retire a removed subagent cap

Implements the five suggestions confirmed from PR #20's "Routine
self-improvements" section (recommendations 1B, 2A, 3A, 4B, 5B). All are
human-authored changes outside the routine's auto-apply envelope, which is why
they could not ride the knowledge PR.

Containment (S1, option B) — guardrails/README.md, guardrails/checklist.md
  The three hard stops bound what a loop spends and how long it runs. Neither
  bounds what an attacker-supplied input makes the harness do, and the known
  mechanisms fire before any stop condition or verification step runs: GitSpawn
  executes on the agent's startup git status, before the workspace-trust
  prompt; instruction privilege escalation lands before automatic permission
  review sees the content. Added as a required companion section rather than a
  fourth numbered hard stop, deliberately: the three are single checkable
  numbers, containment is a posture. Defaults cover scoped network egress,
  unreachable credentials, deny-by-default unattended runs, and real isolation.

Deny-by-default for headless loops (S5, option B) — guardrails/README.md
  Documents --permission-prompts none (v2.1.259+) as the cheapest containment
  win available, including that it removes tools needing a human answer and
  surfaces denials in stream-json output. Also records that --restricted is a
  permission gate, not a sandbox: the primary docs say nothing about OS-level
  isolation and it is absent from Anthropic's own sandbox-environments
  comparison page.

Routine threat model (S2, option A) — runbooks/staying-current.md
  Writes down what update-knowledge can reach and why it is still safe to run.
  The research it does now describes an attack on the shape it is: instruction
  privilege escalation reproduces through /goal- and /schedule-shaped features,
  and neither paper proposes a mitigation. The defenses are structural — never
  merges, blast radius limited to knowledge/, CI-enforced fence, claude/-only
  pushes — so they do not depend on the agent noticing it was manipulated. Also
  lists what would invalidate that reasoning.

Verify the agent, don't relay it (S3, option A) — update-knowledge SKILL.md
  Step 4 now requires machine-verifying every arXiv citation against the arXiv
  API, and checking version attributions against the feature's own doc page
  rather than the nearest changelog heading. Step 4 is a protected region, so
  this is a normal human-authored commit, not a self-edit: one — the route the
  guard's own error message names.

Consistency and drift fixes
  loop-guardrails gains containment as check 6, since a required section that
  nothing audits is not required. And guardrails/README.md, staying-current.md
  (three places) and artifact-audit SKILL.md all still recommended
  CLAUDE_CODE_MAX_SUBAGENTS_PER_SESSION, which was removed in v2.1.224 and now
  does nothing — replaced with the concurrency and depth caps that remain, and
  noting neither bounds a session's total lifetime spawns.

S4 (recording the whats-new digest as unreliable in the skill) was declined:
sources.md already carries it and is read every pass at step 1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ca7z6e9N5SAhNjqM8gFM5R
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants