Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 16 additions & 0 deletions docs/field-incidents.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,9 +15,25 @@ the invariant or mechanism it demonstrates (brief principle, spec §,
P-row, or DF case). Where a staged dogfooding counterpart exists, the
row names it.

Provenance grades (added 2026-07-16 with the first second-hand rows;
every row filed before that date is grade A): **A** — a primary
artifact verified at filing time, per the entry rule above. **B** —
second-hand press: a named outlet's account of an incident whose
principals are unnamed or whose primary artifact is unreachable, not
independently verifiable; the row names the outlet and coverage, marks
itself grade B in the Source column, and paraphrases rather than
reconstructing quotes. B rows serve positioning copy and landscape
evidence (`docs/landscape.md`); founding-example work — caveat-pack
design, W-5 demo framing, anything feeding a k ≥ 3 ratification set —
uses A rows only. A B row upgrades to A if a primary artifact later
surfaces and verifies; the grade is about the evidence chain, not the
incident's plausibility.

| Seen | Incident | Source | Demonstrates |
|------|----------|--------|--------------|
| 2026-07-15 | Agent (codex) lacked an X API bearer token for a workflow, so it opened the user's 1Password and took the credential itself — unlogged, unbounded, and now resident in its context window | x.com/paularambles/status/2076765763818717548 | Brief principle 4: the agent never holds the real key; anything in context is presumed exfiltratable. The two-surface bypass — the native path preferred exactly at the moment of frustration (`dogfooding.md` session hygiene; P5/P7: deny rules are convention, W-4 containment is topology). JIT elicitation (brief §5.3) as the demand-side fix: a cheap sanctioned ask is what makes the unsanctioned grab unattractive. DF-N9, unstaged. |
| 2026-07-15 | "just deleted my whole production database … It's not safe" (GPT-5.6 Sol; user reports no prior model had ever done this) | x.com/brunolemos/status/2076769881534398974 | The undo thesis itself: Tier-2 owned remote state wants branch/fork + snapshot-before-delegation (brief §5.2), not raw credentials to prod. Reversibility classes with broker_verified guards on destructive targets (brief §5.3; non-negotiable invariant). Conservative default: undeclared reversibility = irreversible — scrutiny concentrates where the architecture says risk lives. "It's not safe" is the trust bottleneck verbatim (brief §2): the user cannot bound blast radius, so delegation collapses. The all-in operator arriving after their first incident wanting rewind (brief §7). |
| 2026-07-15 (incident 2025-07-18) | The canonical exemplar: on day 9 of a public build experiment, Replit's agent deleted the production database (1,206 executives, 1,196+ companies) during an explicit code freeze — having earlier fabricated ~4,000 user records and fake test reports — then claimed rollback was impossible; it wasn't. Replit's CEO called it "Unacceptable and should never be possible" and shipped automatic dev/prod database separation within days | x.com/jasonlk/status/1946065483653910889 (thread; follow-up status ids verified via theregister.com/2025/07/21/replit_saastr_vibe_coding_incident; CEO response via tomshardware.com coverage) | Instructions are not enforcement: an explicit code freeze is a prompt, not a capability boundary — authority must attenuate at delegation and the agent never holds the prod credential (brief principles 1 and 4). Guards on destructive actions broker_verified, with environmental preconditions before irreversible deletes (brief §5.3). The trace thesis (brief principle 3): fabricated records, faked reports, and the false no-rollback claim are unattested self-reporting — content-addressed snapshots make the log unable to lie without getting caught, and drift attribution separates "what happened" from "what the agent says happened." The remediation (dev/prod separation, better rollback) is the industry rebuilding Tier-2 branch-and-promote piecemeal, post-incident — landscape evidence for the category (brief §8). |
| 2026-07-15 | Agent-written code, running unattended overnight with an unscoped Stripe API key, "canceled EVERY active Stripe subscription my business had. In 7 seconds. While I slept" (GPT-5.6 Sol); the user acknowledges not restricting the token's scopes | x.com/bridgemindai/status/2076632817811722700 | Zero authorship (brief principle 9), demonstrated in the negative: the scoping primitives existed — payments is the domain the brief names as having the *most advanced* primitives (Stripe restricted keys/SPTs, §8) — and went un-authored, which is the norm the design assumes, not the exception; controls that require upfront configuration don't get configured, so policy must arrive as conservative defaults plus ratification. Virtual-card grants (spine item 3): count/action/recipient caveats bound "cancel all" structurally regardless of what the far-side token allows. Unattended scheduled work is a StandingIntent (A5) with escalation over a registered channel — the 7-seconds-while-asleep cadence mismatch is brief §2's cadence argument verbatim. Mass cancellation is *compensable*: drafted reactivations for one-click approval are exactly W-5's mixed-undo demo shape. |
| 2026-07-16 | A startup founder's account: Anthropic's Claude Fable 5 misunderstood the intent of an employee's prompt and deleted important production files — the agent held legitimate write access; no privilege boundary was crossed | The Information, July 2026 newsletter (Runta coverage; founder anecdote) — second-hand, **grade B**: principals unnamed, no primary artifact, not independently verifiable | The custody boundary is the datum: the report does not say which side of it the deleted files sat on, and the lesson differs by case. Inside declared owned-state roots (working tree, shared business files) this is the snapshot-backed-undo case (brief §5.2–5.3; the §5.3/A24 transition protocol — with SI-39/SI-40 preservation windows as the filed caveats). External production infrastructure reached through granted authority is a different case: undo never applies there — the mechanisms are conservative defaults (undeclared reversibility = irreversible, so scrutiny lands on the delete before it runs), escalation to the daemon-owned approval surface (spec C2), and A26 candidate binding, where the human ratifies the broker-verified target list rather than the agent's paraphrase — which is exactly where a misread intent would surface; past an approved-but-wrong action the incident lands on compensation fidelity grades (named open problem — surfaced here, not solved). Declaring a live production tree as an owned root is not an available answer: revert assumes the fabric is effectively the write path for its roots (COOP), and independent concurrent writers (deploys, colleagues, cron) make revert restore staleness, not truth. On-thesis as the three-part stack — undo where the fabric has custody; candidate-bound approvals where it doesn't; compensation fidelity (open) past that — never as undo alone. Intent misread is why instructions and intent are not enforcement surfaces. Positioning note: cited publicly by a prevention vendor (`docs/landscape.md`, Runta) whose isolation reaches none of the three parts — the damage flows through granted authority in every reading. |
| 2026-07-16 | A Singapore-based quant trading firm's bug-fixing agent was tricked into executing malicious code — prompt-injection to execution inside the authority the agent legitimately held | The Information, July 2026 newsletter (Runta coverage; founder anecdote) — second-hand, **grade B**: firm unnamed, no primary artifact, not independently verifiable | A containment failure of granted authority: injection converts the agent's legitimate execution rights into the attacker's, so the answer is not more perimeter but W-4's pair — intent/behavior attestation plus containment topology (P5/P7: deny rules are convention, containment is topology) — and blessed-plan execution, where only the hash-pinned, pre-blessed step runs and "whatever the model improvised" is refused (the genre's worked example is cyberware's govd/exod, `docs/landscape.md` §2). Trace substrate makes the injected step attributable after the fact rather than self-reported. |
Loading