Repository navigation
knowledge: Sep 7–14 pass — Anthropic ships the three hard stops as a CLI flag set, plus two corrections to this KB - #22
Merged
Conversation
…CLI flag set, and two KB corrections Five-agent fan-out across tooling, ecosystem, key voices, guardrails/cost and verification/skills. No update-knowledge PR was open (PR #21 touches guardrails/, runbooks/ and the skill, not knowledge/), so the merged baseline was clean. Every load-bearing claim was re-read against its primary rather than relayed from an agent report. Headline: `claude plugin eval` (v2.1.269) is the first first-party Claude Code command carrying all three hard stops plus a deterministic gate — max_turns, timeout_seconds, --max-cost-usd (checked before each run starts, exit 2 + partial:true on a budget stop) and --threshold. Its no-plugin ablation arm and hidden case definitions are four of PROCTOR's five deterministic guardrails, arrived at independently. In the same release, CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS (1-256) is changelog-only and absent from both env-vars and workflows, while workflows still presents 16 concurrent / 1,000 total as the runaway-loop ceiling. Two corrections to this KB: - "No review tool ships an enforced budget cap" was wrong — an error of coverage across four passes. Anthropic Code Review and Greptile both ship work-stopping spend caps. The merge-gate "no" stands and is now an explicit vendor design commitment rather than an absence. - The LiteLLM backlog quote does not exist. The real primitive is budget reservation (pre-flight admission control, on by default); fail_closed_budget_enforcement is a separate counter-degradation backstop. Also: arXiv:2609.01222 read in full — X-CPE's mechanism plausibly reaches this repo's own knowledge/ directory via CLAUDE.md's read instruction, making the human PR review the load-bearing control rather than a formality (inference, Medium; flagged for a human, nothing applied to guardrails/). Anthropic's Sep 10 threat-intel report documents adversaries running scheduled agent swarms that iterate until success. OpenAI Agents API and Cursor Projects both shipped unbounded-horizon loops the same week with no documented ceilings. Sept 14 weekly-limit change re-checked on its own effective date: still uncorroborated on every Anthropic-controlled surface; downgraded, kept open. Four backlog items resolved and archived, five narrowed or corrected, eleven opened — including two structural blind spots (x.com returns 402; the arXiv Atom API rate-limited the whole pass, so verification used OAI-PMH plus abs submission history). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VWoWKBnpy6V1mVHGtu6Z1g
espi
pushed a commit
that referenced
this pull request
Sep 21, 2026
…and two corrections this routine owes itself Window Sep 14 -> Sep 21, 2026. Branched from PR #22's head (open and unmerged) per the skill's step 1, so this PR is stacked on it. Headline: Mandiant's September 2026 report carries Case study 6, "Denial-of-Wallet" using rogue reasoning loop — a ledger agent hit a null value, looped, and made "over 15,000 high-frequency, high-cost reasoning API calls" for a "~$50,000 cloud-billing spike" plus database locking that halted live transactions. Its recommended controls independently reproduce all three hard stops. METR's stolen-key disclosure is the companion case where the ceiling did not exist at all ("no natural token spend ceiling"). Two corrections are this routine's own inference failures, not changes in the world: the whats-new digest was lagging, not discontinued (w35-w37 now 200), and Anthropic's Sept 14 limit change was never contradicted — the page simply had not been updated yet. Both were produced by repetition creating false confidence across passes. Verification: VP-Control (read in full) measures the separate-checker rule and finds evidence-source independence worth ~2x model independence (40.9pp vs 11.3pp), which shifts the doctrine's emphasis. OverclaimBench finds agents' final responses "are not reliable accounts of their actions." SaltBench reframes hard stop #1 — "a budget stop is a halt, never a failure" — and self-reports probes certifying a fence that did not exist. Two findings aim at this repo's self-improvement envelope: poisoned-benchmark contamination persists through clean re-evaluation (2609.17817), and GuardrailLoop pins policy+evaluation hashes before every stage. Flagged for a human, not applied. Eight backlog items resolved and archived, five narrowed, eleven opened. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SQ341kwv5YJKY8Yd2tAsmk
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Weekly
update-knowledgepass. Window Sep 7 → Sep 14, 2026; last pass 2026-09-07.Routine self-improvements
Applied (auto-gated): none this pass.
The gate found no may-auto-apply candidate. I checked the two things that qualify — every intra-repo path referenced by
SKILL.md(runbooks/staying-current.md,knowledge/archive/resolved-caveats.md, the threeknowledge/files,.github/workflows/self-edit-guard.yml,guardrails/budget.env) resolves on disk, and every internal step-number cross-reference (steps 1, 2, 4, 7, 8) points at the right step. No typos or markdown glitches found. Default-deny applies, so nothing was written.Independently: PR #21 is open and modifies this same
SKILL.md. Even a qualifying self-edit should have been suppressed this pass to avoid conflicting with it.Suggested (needs your confirm):
general-purpose)". Subagents are now spawned with the Agent tool;TaskCreate/Get/Update/Listwere removed on newer models in v2.1.233. Routed to human by gate check 5 (repo-internal evidence) — this is justified by harness behaviour, not by anything checkable on disk — and by the always-suggest "changes what a step does" trigger.claude/knowledge-update-YYYY-MM-DD; this session was instructed to develop onclaude/focused-clarke-5kjv2q. Both areclaude/-prefixed, so the push restriction (and the never-touch-mainfloor) held either way, and I used the designated branch. Worth reconciling the wording so the two don't appear to conflict. Routed to human: it touches the branch rule, which is always-suggest.export.arxiv.org/api/queryreturnedRate exceededfor this entire pass (15/15 probes over ~30 min, plus manual attempts). Verification fell back to arXiv's OAI-PMHGetRecord(ID / exact title / primary category) and the abs submission history (v1 date) — both arXiv-operated, and I re-confirmed both endpoints myself. A caution belongs with it: OAI's<created>is the announcement date, not the v1 date, so substituting it would reintroduce exactly the misdating the step guards against. Step 4 is a protected region — human-authored only.Knowledge changes
The headline: the doctrine shipped as a command
claude plugin eval(Claude Code v2.1.269, Sep 11) is the first first-party Claude Code command carrying all three hard stops plus a deterministic gate in one interface —max_turns(default 10),timeout_seconds(default 300), and--max-cost-usd: "Checked before each run starts. Once spent, nothing further starts; runs already in flight finish, so spend can pass the ceiling by those runs." Same bounded-overshoot semantics this KB recorded for Managed Agents session budgets, now as a flag, with exit 2 +partial: truedistinguishing a budget stop from a quality failure. Gate:--threshold(default1.0).Its anti-self-grading design is the interesting part: three runs per case, a no-plugin ablation arm ("If a case scores 1.0 both with and without the plugin, the plugin isn't what made it pass"), case definitions hidden from the agent, and the explicit instruction to "suspect the judge before the plugin." That is four of PROCTOR's five deterministic guardrails, arrived at independently.
One trap recorded: a usage-limit hit mid-suite scores ~0 and "isn't marked
partial, so the result can look like a regression" — budget exhaustion masquerading as quality loss.Pulling the other way in the same release:
CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS(1–256) is changelog-only — absent from bothenv-varsandworkflows(I fetched both in full and grepped) — whileworkflowsstill presents "Up to 16 concurrent agents" and "1,000 agents total per run | Prevents runaway loops" as the ceiling. A documented runaway-loop cap with an undocumented 16× escape hatch.Also:
maxEffortLevel(v2.1.267), a genuinely non-raisable ceiling ("the lowest applies, so a cap set in one scope can't be raised from another"), and a/goalsilent-stall fix (v2.1.269) — a stop-condition loop that stalls without saying so defeats stall detection.Two corrections to this knowledge base
fail_closed_budget_enforcement" appears nowhere in LiteLLM's docs or repo. The real primitive is budget reservation — "If the reservation would exceed the budget, LiteLLM rejects the request before sending it to the provider" — genuine pre-flight admission control, on by default, with a documented hole where cost can't be estimated.fail_closed_budget_enforcementis a separate counter-degradation backstop.CLAUDE_CODE_MAX_SUBAGENTS_PER_SESSIONis not "gone from the docs" as recorded 2026-08-17 —env-varsnow carries an explicit tombstone. Guidance unchanged; the old text was corrected in place rather than left beside the correction.arXiv:2609.01222 was read in full (a standing backlog item). X-CPE's definition is not file-specific — it covers any "context source that is more persistent for the agent" — and its Attack Vector A-7 exploits
@-import chains fromCLAUDE.md. This repo reaches the same edge by prose:CLAUDE.mdtells every arriving agent to readknowledge/00-primer.md, and this routine writes web-sourced research there for a later run to read back (σ_session → σ_project).The paper doesn't test this configuration, so the transfer is inference, not a result (Medium). But the honest consequence is worth stating plainly: the human PR review is not a formality here — it is the only control standing between web-sourced text and a persistent context source. That sharpens the existing rule that a routine self-edit is never justified by web-sourced research.
Flagged, not applied. Whether
guardrails/should say more, and whether this routine should mark web-sourced text as tainted in some machine-visible way, is a human call — it's on the backlog, not in this diff.The same paper half-resolves the "neither IPE paper proposes a mitigation" caveat: it carries a disclosure section claiming OpenAI and Anthropic acknowledged the findings and that Codex, Gemini CLI and Cline shipped fixes — but names no versions, dates or CVEs, so finding the vendor-side artifact is now its own open item. arXiv:2608.27299 is unchanged.
Other new facts
max_concurrent_subagents: 4— no turn cap, stop condition or spend ceiling. Cursor "Projects" (Sep 10) "delegates tasks to thousands of subagents", runs laptop-closed on schedule/PR/Slack triggers, and documents none of the three. A clean natural experiment againstclaude plugin eval: ceilings are not yet a market norm.rubyhack.ai(Sep 11, same authors) — RubyGems swarm from May 5 2026, 2,000+ malicious packages,.yardoptsRCE in doc builds, and the vendor-confirmed link: "The June agents were accessing 49 of the same files as the wiki agents, which OpenAI has confirmed were theirs."Archived (4)
Moved to
knowledge/archive/resolved-caveats.md: AI Engineer World's Fair 2026 sessions (resolved — and the item was wrong on its own premise: Yegge's WF26 talk is "Agentic Security", not "Harness Engineering," and the "Harness Engineering fireside" was a Tessl side event with Dru Knox, not Guy Podjarny, with no recording; rewritten before archiving so the error isn't preserved); LiteLLM fail-closed pre-flight rejection (resolved as the correction above); "read arXiv:2609.01222v2 in full" and "read PROCTOR in full" (both read, each leaving a narrower successor question); AgentGuard identity (disambiguated tobmdhodl/agent47).Every currently-open re-verify item
Per step 2, the full standing backlog, not just what's new:
Security / highest priority
claude ultrareviewpath — re-checked; still no evidence of a patch through v2.1.270. Three negatives: no Anthropic advisory published in Sept 2026, no security-shaped ultrareview entry in v2.1.265–270, no CVE assigned (CVE-2026-19592 is Codex's, and an earlier note here conflated them). Absence of evidence, not proof — Manifold withheld the config key.knowledge/in practice? (reframed) — inference, not a tested result. Human call on whetherguardrails/needs more.--restrictedis not an OS-level sandbox — no new information; docs still neither claim nor deny OS isolation;CLAUDE_CODE_RESTRICTED=1still unfindable in primary docs.Tooling / docs
5.
CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTSundocumented (new) — which cap does it lift, and does the 1,000 total still bind?6. Spend-limit-bar version conflict (new) — gateway docs say v2.1.251, changelog says v2.1.259; both first-party.
7. 1 GB tool-results cap is changelog-only (new).
8. Agent Plugins 1.0 has no canonical repo URL in this KB (new) —
github.com/agent-plugins/spec404s.9.
SKILL.mdcross-tool execution — unchanged; portable-with-testing, not drop-in.10. Graphite is a Cursor product (new) — the KB still implicitly treats it as independent.
Cost / limits
11. Anthropic's Sept 14 weekly-limit change — re-checked on the effective date; still uncorroborated. Help center unchanged ("Updated over a week ago"), still says limits "return to their standard levels" after Sep 13, and mentions neither Sept 14 nor +25% nor any reduction. No Anthropic surface carries it; the only primary is an X post returning 402. Downgraded further.
12. OpenAI "Research acceleration" spend figures — openai.com still 403s (re-verified twice). Willison's chart axis remains the best independently-opened corroboration.
13. Promote "no review tool ships a merge gate" to a primer finding? (corrected) — budget-cap half was wrong; merge-gate half now rests on an explicit vendor commitment.
14.
$47K / 11-day,$500M in one month, $6,000 overnight / $4,200 refactor — still self-reported or unattributed; the SEO cluster recycled them again this pass. Do not cite.15. June 15 2026 billing split — no primary Anthropic announcement of the pause; re-check when a revised plan lands.
16. Microsoft dropping Claude Code — no primary Microsoft announcement read.
Research / infrastructure
17.
export.arxiv.org/api/queryreturned "Rate exceeded" for the whole pass (new) — fallback recorded; check whether the Atom API recovers.18. The x.com blind spot is structural (new) — HTTP 402 on every status URL, and Steinberger, Osmani and Anthropic's limit announcement all publish there. Worth a human deciding whether a workaround is warranted.
19. Adopt PROCTOR's canary-case guardrail? (reframed) — cheap, and this repo has no equivalent.
20. Read in full next pass — arXiv:2609.12216, 2609.10969, 2609.11076.
21. Willison's Sep 4 rogue-wikis report (narrowed) — the follow-up primary is read; only the
/etc/hostsproxy-bypass detail remains relayed.22. roborev.io/changelog 403s (and
releases.atom403s while the HTML releases page is readable); Greptile's changelog renders via JS — use a raw fetch.23. This query space is dominated by marketing — named do-not-cite list carried forward.
24. The "costliest thing is managing the agent loop" slogan — community paraphrase, not a Cherny quote.
25.
/goal"Codex invented it, Claude copied in 11 days" — single secondary.26. Huntley's Loom — self-described;
ghuntley/loomcommit activity was unverifiable this pass (GitHub access is scoped toespi/loops).27. EvoAgentBench / SkillCheck — still too thin to promote.
Notes for the reviewer
guardrails/,runbooks/and.claude/skills/update-knowledge/SKILL.md. This PR touches onlyknowledge/, so the two should not conflict — but they are thematically adjacent (PR guardrails: containment as a required companion, verify-the-agent rule, and retiring a subagent cap that no longer exists #21's containment section and this PR's X-CPE finding are about the same threat class), and reading them together is worth the few extra minutes.main, nothing outsideknowledge/was modified, and noself-edit:commit exists in this branch.🤖 Generated with Claude Code
https://claude.ai/code/session_01VWoWKBnpy6V1mVHGtu6Z1g
Generated by Claude Code