Skip to content

feat(v1.10): budget guardrail - cost confirm before each AI phase - #76

Merged
jameskhair-code merged 3 commits into
mainfrom
feat/v1.10-budget-guardrail
Jun 11, 2026
Merged

jameskhair-code merged 3 commits into
mainfrom
feat/v1.10-budget-guardrail

Conversation

@jameskhair-code

@jameskhair-code jameskhair-code commented Jun 11, 2026 •

Copy link
Copy Markdown
Owner

Item 4 of the v1.10 preflight (docs/planning/v1.10-charter.md) — last code item before the release chore. Consumes item 2's verified prices.

What changed

Projection helper — usage.py:project_step_cost. The simplest honest projection, per the charter: when the step has ≥ 10 calls of usage history, per-book cost = the step's historical totals priced at the requested model, divided by (calls × the step's default batch size — the log doesn't record per-call book counts, so the default batch is the stated assumption). Below that sample, a static conservative fallback (~2× the measured per-book token averages from the PR #74 cost table). The basis is labelled in the prompt line ("usage history, 158 calls" / "static estimate, no usage history"). Unpriced models return None — no dollar estimate exists anywhere in the toolkit for them, so the gate cannot fire. replay_usage_log gained an optional step filter (same parser, no second reader).

Prompt — _common.py:budget_guardrail. Before a step's AI phase: This batch: N books ≈ $X (basis) — proceed? when the projection exceeds usage.confirm_above_usd (default $1.00). At or below the threshold: silent. Declining prints "Aborted — no AI calls made" and exits cleanly. --auto-apply-high does not bypass — it's a review-tier flag, not a yes-flag, and the charter wants the guardrail fronting every campaign run.

Charter correction (review feedback, commit 6109c50): the charter's "--dry-run must bypass" wording conflated write-safety with spend-safety — dry-runs make real AI calls, so the spend gate now fires on them too. The dry_run parameter and its projection-only branch were removed; the gate has no bypass path at all.

Wiring. The four AI commands gate at the start of their AI block — clean-titles, comments-enrich, tags-enrich, and lcc-enrich (placed after the catalog cascade, before AI paths 3c and 3d, with n = the full batch since the per-book history basis amortizes the catalog/AI mix). regrade passes the threshold through its dispatch, so re-grade runs are gated too. Helper sits entirely outside the review/apply flow — nothing downstream of the AI call changed.

Charter verification note: enrich-identifiers constructs no AI client at all — nothing to gate. tags-cleanup does make AI calls but wasn't in the charter's touch-point list; captured in docs/planning/inbox.md (low exposure: 150-book batches, few calls).

Config. usage.confirm_above_usd added to config.example.json (with comment) and config_schema.py:UsageConfig. The real config.json needs no change — the default applies.

Verification

  • python -m pytest -q — 659 passed (12 new: history/static/unknown-step/unpriced projection cases, replay step filter, prompt-over-threshold / silent-under-threshold / decline-exits / no-dry-run-bypass / unpriced-never-prompts guardrail cases, config type validation, knob default).
  • Real library, the charter checks (run before the dry-run correction; the correction only removes a bypass, it does not touch the projection or prompt paths exercised here):
    • Over threshold prompts with a plausible estimate: clean-titles 'title:"~."' --limit 1000 → This batch: 1000 books ≈ $1.84 (usage history, 158 calls) — proceed? [y/N] — matches the hand-computed per-book figure from the same history.
    • Decline writes nothing: answered n → "Aborted — no AI calls made"; audit.log and usage.jsonl line counts unchanged (6 / 230 before and after).
    • Under threshold doesn't prompt: 5-book dry-run proceeded straight to AI with no guardrail output (still true post-correction — under-threshold is silent for all runs).

Second commit removes three local scratch logs (claude-smoke.log, enrich.log, an old lcc-probe report) accidentally swept into the first commit by git add -A — files remain on disk, just untracked.

🤖 Generated with Claude Code

jameskhair-code and others added 3 commits June 11, 2026 08:13
Before a step run's AI phase: "This batch: N books ~= $X - proceed?"
above usage.confirm_above_usd (default $1.00). Projection basis is
per-step usage history (>=10 calls, converted per-book via the step's
default batch size) with a static conservative fallback; the basis is
labelled in the prompt line. Declining exits before any AI call;
--dry-run shows the projection without prompting; unpriced models
cannot fire the gate (no estimate exists for them anywhere).

Wired into the four AI commands (clean-titles, lcc-enrich,
comments-enrich, tags-enrich) and regrade dispatch. enrich-identifiers
verified to make no AI calls (charter note); tags-cleanup gap captured
in the inbox.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
claude-smoke.log, enrich.log, and the lcc-probe report are local
session debris swept up by git add -A; files stay on disk.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Dry-runs make real AI calls, so the spend gate must not bypass them -
the charter wording conflated write-safety with spend-safety. Removes
the dry_run parameter and the dim-line branch; the gate now has no
bypass path at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@jameskhair-code
jameskhair-code merged commit b51b37b into main Jun 11, 2026
2 checks passed
@jameskhair-code
jameskhair-code deleted the feat/v1.10-budget-guardrail branch June 11, 2026 12:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant