Skip to content

knowledge: Sep 7–14 pass — Anthropic ships the three hard stops as a CLI flag set, plus two corrections to this KB - #22

Merged
espi merged 1 commit into
mainfrom
claude/focused-clarke-5kjv2q
Sep 29, 2026
Merged

espi merged 1 commit into
mainfrom
claude/focused-clarke-5kjv2q

Conversation

@espi

@espi espi commented Sep 14, 2026

Copy link
Copy Markdown
Owner

Weekly update-knowledge pass. Window Sep 7 → Sep 14, 2026; last pass 2026-09-07.

Routine self-improvements

Applied (auto-gated): none this pass.

The gate found no may-auto-apply candidate. I checked the two things that qualify — every intra-repo path referenced by SKILL.md (runbooks/staying-current.md, knowledge/archive/resolved-caveats.md, the three knowledge/ files, .github/workflows/self-edit-guard.yml, guardrails/budget.env) resolves on disk, and every internal step-number cross-reference (steps 1, 2, 4, 7, 8) points at the right step. No typos or markdown glitches found. Default-deny applies, so nothing was written.

Independently: PR #21 is open and modifies this same SKILL.md. Even a qualifying self-edit should have been suppressed this pass to avoid conflicting with it.

Suggested (needs your confirm):

  1. Step 3's tool name is stale. It says "Launch parallel research agents (Task tool, general-purpose)". Subagents are now spawned with the Agent tool; TaskCreate/Get/Update/List were removed on newer models in v2.1.233. Routed to human by gate check 5 (repo-internal evidence) — this is justified by harness behaviour, not by anything checkable on disk — and by the always-suggest "changes what a step does" trigger.
  2. Step 8's branch rule collides with a session-designated branch. The skill says to create claude/knowledge-update-YYYY-MM-DD; this session was instructed to develop on claude/focused-clarke-5kjv2q. Both are claude/-prefixed, so the push restriction (and the never-touch-main floor) held either way, and I used the designated branch. Worth reconciling the wording so the two don't appear to conflict. Routed to human: it touches the branch rule, which is always-suggest.
  3. Step 4 could name an arXiv fallback. export.arxiv.org/api/query returned Rate exceeded for this entire pass (15/15 probes over ~30 min, plus manual attempts). Verification fell back to arXiv's OAI-PMH GetRecord (ID / exact title / primary category) and the abs submission history (v1 date) — both arXiv-operated, and I re-confirmed both endpoints myself. A caution belongs with it: OAI's <created> is the announcement date, not the v1 date, so substituting it would reintroduce exactly the misdating the step guards against. Step 4 is a protected region — human-authored only.
  4. Noting for the reviewer of PR guardrails: containment as a required companion, verify-the-agent rule, and retiring a subagent cap that no longer exists #21: that PR's proposed "verify a research agent's report, don't relay it" block in step 4 was followed in practice this pass even though it is unmerged — every headline claim below was re-fetched and verbatim-checked by the lead agent, which is how both corrections in this PR were caught.

Knowledge changes

The headline: the doctrine shipped as a command

claude plugin eval (Claude Code v2.1.269, Sep 11) is the first first-party Claude Code command carrying all three hard stops plus a deterministic gate in one interface — max_turns (default 10), timeout_seconds (default 300), and --max-cost-usd: "Checked before each run starts. Once spent, nothing further starts; runs already in flight finish, so spend can pass the ceiling by those runs." Same bounded-overshoot semantics this KB recorded for Managed Agents session budgets, now as a flag, with exit 2 + partial: true distinguishing a budget stop from a quality failure. Gate: --threshold (default 1.0).

Its anti-self-grading design is the interesting part: three runs per case, a no-plugin ablation arm ("If a case scores 1.0 both with and without the plugin, the plugin isn't what made it pass"), case definitions hidden from the agent, and the explicit instruction to "suspect the judge before the plugin." That is four of PROCTOR's five deterministic guardrails, arrived at independently.

One trap recorded: a usage-limit hit mid-suite scores ~0 and "isn't marked partial, so the result can look like a regression" — budget exhaustion masquerading as quality loss.

Pulling the other way in the same release: CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS (1–256) is changelog-only — absent from both env-vars and workflows (I fetched both in full and grepped) — while workflows still presents "Up to 16 concurrent agents" and "1,000 agents total per run | Prevents runaway loops" as the ceiling. A documented runaway-loop cap with an undocumented 16× escape hatch.

Also: maxEffortLevel (v2.1.267), a genuinely non-raisable ceiling ("the lowest applies, so a cap set in one scope can't be raised from another"), and a /goal silent-stall fix (v2.1.269) — a stop-condition loop that stalls without saying so defeats stall detection.

Two corrections to this knowledge base

  • "No review tool ships an enforced budget cap" was wrong — an error of coverage across four passes, not a change in the world. Anthropic Code Review has shipped a work-stopping monthly spend cap since at least Jul 12 2026 ("Code Review posts a single comment on the PR explaining that the review was skipped"), and Greptile's Flex Usage Limits (projected-spend, pre-flight) since Apr 30 2026. The merge-gate half stands and is now stronger than an absence — Anthropic states it as design: "The check run always completes with a neutral conclusion so it never blocks merging through branch protection rules."
  • The LiteLLM backlog quote does not exist. "reject known estimates over remaining budget under fail_closed_budget_enforcement" appears nowhere in LiteLLM's docs or repo. The real primitive is budget reservation — "If the reservation would exceed the budget, LiteLLM rejects the request before sending it to the provider" — genuine pre-flight admission control, on by default, with a documented hole where cost can't be estimated. fail_closed_budget_enforcement is a separate counter-degradation backstop.
  • Smaller: CLAUDE_CODE_MAX_SUBAGENTS_PER_SESSION is not "gone from the docs" as recorded 2026-08-17 — env-vars now carries an explicit tombstone. Guidance unchanged; the old text was corrected in place rather than left beside the correction.

⚠️ The one that names this repo

arXiv:2609.01222 was read in full (a standing backlog item). X-CPE's definition is not file-specific — it covers any "context source that is more persistent for the agent" — and its Attack Vector A-7 exploits @-import chains from CLAUDE.md. This repo reaches the same edge by prose: CLAUDE.md tells every arriving agent to read knowledge/00-primer.md, and this routine writes web-sourced research there for a later run to read back (σ_session → σ_project).

The paper doesn't test this configuration, so the transfer is inference, not a result (Medium). But the honest consequence is worth stating plainly: the human PR review is not a formality here — it is the only control standing between web-sourced text and a persistent context source. That sharpens the existing rule that a routine self-edit is never justified by web-sourced research.

Flagged, not applied. Whether guardrails/ should say more, and whether this routine should mark web-sourced text as tainted in some machine-visible way, is a human call — it's on the backlog, not in this diff.

The same paper half-resolves the "neither IPE paper proposes a mitigation" caveat: it carries a disclosure section claiming OpenAI and Anthropic acknowledged the findings and that Codex, Gemini CLI and Cline shipped fixes — but names no versions, dates or CVEs, so finding the vendor-side artifact is now its own open item. arXiv:2608.27299 is unchanged.

Other new facts

  • Adversaries run this pattern, and Anthropic documents it first-hand. Its Sep 10 threat-intel report describes "agent swarms," "A fleet of thirteen standing collection AI agents ran on a scheduled job," "The workflow iterated over edits of the exploit code until success," and scheduled credential-renewal jobs "with no human involvement" — scheduling, fan-out, iterate-until-verified and durable cross-run state, i.e. §2's four orchestration-loop properties pointed the other way.
  • Two harnesses shipped unbounded-horizon loops the same week. OpenAI Agents API (public beta ~Sep 10) documents only max_concurrent_subagents: 4 — no turn cap, stop condition or spend ceiling. Cursor "Projects" (Sep 10) "delegates tasks to thousands of subagents", runs laptop-closed on schedule/PR/Slack triggers, and documents none of the three. A clean natural experiment against claude plugin eval: ceilings are not yet a market norm.
  • The "rogue agent wikis" primary arrived. rubyhack.ai (Sep 11, same authors) — RubyGems swarm from May 5 2026, 2,000+ malicious packages, .yardopts RCE in doc builds, and the vendor-confirmed link: "The June agents were accessing 49 of the same files as the wiki agents, which OpenAI has confirmed were theirs."
  • Verification papers. Sharpest is arXiv:2609.10969: "a cross-model vote over shared evidence approves 62.9% of unsafe proposals, versus 22.9% with an independent source. The source effect is 40.9 percentage points, compared with 11.3 for model diversity" — independent evidence matters ~4× more than an independent model, which cuts against "just add a second model as a checker." arXiv:2609.11076 contributes a framing this KB lacked for hard stop knowledge: 2026-06-22 update pass (7-day delta from Jun 15) #3: "a budget stop is a halt, never a failure." Also 2609.12039, 2609.12216 (hash-pinned evaluation identity — a primitive this repo lacks), 2609.08371 (CapScope, 3/75 vs 33–47/75), 2609.07360, 2609.12001.
  • The skills-as-durable-asset evidence is a near-null, honestly reported (arXiv:2609.12742): +4.9 pp / +0.1 pp, and "cannot be separated from the agent's run-to-run variance." Fair summary: plausibly valuable, not yet measurably so.
  • PROCTOR read in full — five guardrails now quoted verbatim, and cited as a position, not evidence (single author, one model family, single-run, "no public artifact").
  • Anthropic's own Sep 8 cost guidance contains no ceiling — optimization and visibility only, the cleanest illustration of why the ceiling must live in your harness.
  • AgentGuard v1.3.0 (Sep 12) — a catalogue of ways a ceiling silently isn't one: a zero-call budget that "previously LangChain could log the exception and continue", and NaN/infinite/negative inputs rejected "so non-finite values cannot bypass a cost ceiling."

Archived (4)

Moved to knowledge/archive/resolved-caveats.md: AI Engineer World's Fair 2026 sessions (resolved — and the item was wrong on its own premise: Yegge's WF26 talk is "Agentic Security", not "Harness Engineering," and the "Harness Engineering fireside" was a Tessl side event with Dru Knox, not Guy Podjarny, with no recording; rewritten before archiving so the error isn't preserved); LiteLLM fail-closed pre-flight rejection (resolved as the correction above); "read arXiv:2609.01222v2 in full" and "read PROCTOR in full" (both read, each leaving a narrower successor question); AgentGuard identity (disambiguated to bmdhodl/agent47).


Every currently-open re-verify item

Per step 2, the full standing backlog, not just what's new:

Security / highest priority

  1. GitSpawn's claude ultrareview path — re-checked; still no evidence of a patch through v2.1.270. Three negatives: no Anthropic advisory published in Sept 2026, no security-shaped ultrareview entry in v2.1.265–270, no CVE assigned (CVE-2026-19592 is Codex's, and an earlier note here conflated them). Absence of evidence, not proof — Manifold withheld the config key.
  2. Vendor-side artifact for arXiv:2609.01222's claimed fixes (new) — no versions, dates or CVEs named.
  3. Does X-CPE reach this repo's knowledge/ in practice? (reframed) — inference, not a tested result. Human call on whether guardrails/ needs more.
  4. --restricted is not an OS-level sandbox — no new information; docs still neither claim nor deny OS isolation; CLAUDE_CODE_RESTRICTED=1 still unfindable in primary docs.

Tooling / docs
5. CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS undocumented (new) — which cap does it lift, and does the 1,000 total still bind?
6. Spend-limit-bar version conflict (new) — gateway docs say v2.1.251, changelog says v2.1.259; both first-party.
7. 1 GB tool-results cap is changelog-only (new).
8. Agent Plugins 1.0 has no canonical repo URL in this KB (new) — github.com/agent-plugins/spec 404s.
9. SKILL.md cross-tool execution — unchanged; portable-with-testing, not drop-in.
10. Graphite is a Cursor product (new) — the KB still implicitly treats it as independent.

Cost / limits
11. Anthropic's Sept 14 weekly-limit change — re-checked on the effective date; still uncorroborated. Help center unchanged ("Updated over a week ago"), still says limits "return to their standard levels" after Sep 13, and mentions neither Sept 14 nor +25% nor any reduction. No Anthropic surface carries it; the only primary is an X post returning 402. Downgraded further.
12. OpenAI "Research acceleration" spend figures — openai.com still 403s (re-verified twice). Willison's chart axis remains the best independently-opened corroboration.
13. Promote "no review tool ships a merge gate" to a primer finding? (corrected) — budget-cap half was wrong; merge-gate half now rests on an explicit vendor commitment.
14. $47K / 11-day, $500M in one month, $6,000 overnight / $4,200 refactor — still self-reported or unattributed; the SEO cluster recycled them again this pass. Do not cite.
15. June 15 2026 billing split — no primary Anthropic announcement of the pause; re-check when a revised plan lands.
16. Microsoft dropping Claude Code — no primary Microsoft announcement read.

Research / infrastructure
17. export.arxiv.org/api/query returned "Rate exceeded" for the whole pass (new) — fallback recorded; check whether the Atom API recovers.
18. The x.com blind spot is structural (new) — HTTP 402 on every status URL, and Steinberger, Osmani and Anthropic's limit announcement all publish there. Worth a human deciding whether a workaround is warranted.
19. Adopt PROCTOR's canary-case guardrail? (reframed) — cheap, and this repo has no equivalent.
20. Read in full next pass — arXiv:2609.12216, 2609.10969, 2609.11076.
21. Willison's Sep 4 rogue-wikis report (narrowed) — the follow-up primary is read; only the /etc/hosts proxy-bypass detail remains relayed.
22. roborev.io/changelog 403s (and releases.atom 403s while the HTML releases page is readable); Greptile's changelog renders via JS — use a raw fetch.
23. This query space is dominated by marketing — named do-not-cite list carried forward.
24. The "costliest thing is managing the agent loop" slogan — community paraphrase, not a Cherny quote.
25. /goal "Codex invented it, Claude copied in 11 days" — single secondary.
26. Huntley's Loom — self-described; ghuntley/loom commit activity was unverifiable this pass (GitHub access is scoped to espi/loops).
27. EvoAgentBench / SkillCheck — still too thin to promote.


Notes for the reviewer

🤖 Generated with Claude Code

https://claude.ai/code/session_01VWoWKBnpy6V1mVHGtu6Z1g


Generated by Claude Code

…CLI flag set, and two KB corrections

Five-agent fan-out across tooling, ecosystem, key voices, guardrails/cost and
verification/skills. No update-knowledge PR was open (PR #21 touches guardrails/,
runbooks/ and the skill, not knowledge/), so the merged baseline was clean. Every
load-bearing claim was re-read against its primary rather than relayed from an
agent report.

Headline: `claude plugin eval` (v2.1.269) is the first first-party Claude Code
command carrying all three hard stops plus a deterministic gate — max_turns,
timeout_seconds, --max-cost-usd (checked before each run starts, exit 2 +
partial:true on a budget stop) and --threshold. Its no-plugin ablation arm and
hidden case definitions are four of PROCTOR's five deterministic guardrails,
arrived at independently. In the same release, CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS
(1-256) is changelog-only and absent from both env-vars and workflows, while
workflows still presents 16 concurrent / 1,000 total as the runaway-loop ceiling.

Two corrections to this KB:
- "No review tool ships an enforced budget cap" was wrong — an error of coverage
  across four passes. Anthropic Code Review and Greptile both ship work-stopping
  spend caps. The merge-gate "no" stands and is now an explicit vendor design
  commitment rather than an absence.
- The LiteLLM backlog quote does not exist. The real primitive is budget
  reservation (pre-flight admission control, on by default);
  fail_closed_budget_enforcement is a separate counter-degradation backstop.

Also: arXiv:2609.01222 read in full — X-CPE's mechanism plausibly reaches this
repo's own knowledge/ directory via CLAUDE.md's read instruction, making the
human PR review the load-bearing control rather than a formality (inference,
Medium; flagged for a human, nothing applied to guardrails/). Anthropic's Sep 10
threat-intel report documents adversaries running scheduled agent swarms that
iterate until success. OpenAI Agents API and Cursor Projects both shipped
unbounded-horizon loops the same week with no documented ceilings.

Sept 14 weekly-limit change re-checked on its own effective date: still
uncorroborated on every Anthropic-controlled surface; downgraded, kept open.

Four backlog items resolved and archived, five narrowed or corrected, eleven
opened — including two structural blind spots (x.com returns 402; the arXiv Atom
API rate-limited the whole pass, so verification used OAI-PMH plus abs
submission history).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VWoWKBnpy6V1mVHGtu6Z1g
espi pushed a commit that referenced this pull request Sep 21, 2026
…and two corrections this routine owes itself

Window Sep 14 -> Sep 21, 2026. Branched from PR #22's head (open and unmerged)
per the skill's step 1, so this PR is stacked on it.

Headline: Mandiant's September 2026 report carries Case study 6,
"Denial-of-Wallet" using rogue reasoning loop — a ledger agent hit a null
value, looped, and made "over 15,000 high-frequency, high-cost reasoning API
calls" for a "~$50,000 cloud-billing spike" plus database locking that halted
live transactions. Its recommended controls independently reproduce all three
hard stops. METR's stolen-key disclosure is the companion case where the
ceiling did not exist at all ("no natural token spend ceiling").

Two corrections are this routine's own inference failures, not changes in the
world: the whats-new digest was lagging, not discontinued (w35-w37 now 200),
and Anthropic's Sept 14 limit change was never contradicted — the page simply
had not been updated yet. Both were produced by repetition creating false
confidence across passes.

Verification: VP-Control (read in full) measures the separate-checker rule and
finds evidence-source independence worth ~2x model independence (40.9pp vs
11.3pp), which shifts the doctrine's emphasis. OverclaimBench finds agents'
final responses "are not reliable accounts of their actions." SaltBench
reframes hard stop #1 — "a budget stop is a halt, never a failure" — and
self-reports probes certifying a fence that did not exist.

Two findings aim at this repo's self-improvement envelope: poisoned-benchmark
contamination persists through clean re-evaluation (2609.17817), and
GuardrailLoop pins policy+evaluation hashes before every stage. Flagged for a
human, not applied.

Eight backlog items resolved and archived, five narrowed, eleven opened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SQ341kwv5YJKY8Yd2tAsmk
@espi
espi merged commit 791f2b6 into main Sep 29, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants