Skip to content

fix(sdk): an authored budget refusal names the limit it crossed - #594

Merged
khaliqgant merged 3 commits into
mainfrom
fix/authored-budget-error-names-limit
Oct 1, 2026
Merged

khaliqgant merged 3 commits into
mainfrom
fix/authored-budget-error-names-limit

Conversation

@khaliqgant

@khaliqgant khaliqgant commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Problem

When the kernel refuses an authored flow's next step for budget, the authored runner reports only:

FAILED [step_failed] Flow budget exceeded before the next step.

It says nothing about which limit tripped or by how much. Cloud run f28314ed (software-factory, cloud#4033) stopped with exactly this message after 26 successful steps. From the message alone, the operator couldn't tell whether the flow header's wallclock: "2h", its dollars: 25, or Cloud's per-launch run budget (180m) had fired. It took summing the step durations to show that the header's wallclock was the limit: 132.9 min of summed step time against 2h.

Change

packages/sdk/src/authored-budget.ts: the accumulator already holds the exact totals it sends to the kernel as prior_spend. On a budget_exceeded admission, the error now names each declared limit the carried spend crossed, using the kernel's strict > comparisons from machine/budget.rs:

step_failed: Flow budget exceeded before step "run-27": wallclock 132.9m used of 2h declared in the flow's budget header.
  • It covers wallclock, dollars, total tokens, input tokens and output tokens. Several crossed limits are joined with ; .
  • Dollars say metered (some steps unmetered) when any charge was unmetered, because the figure is then a lower bound.
  • A /day budget says in today's window of the flow's budget header.
  • The refused step is named when the child spec has exactly one step. Otherwise the message says "the next step".
  • When the carried totals don't explain the refusal, the original wording stays, so the SDK never invents a reason.
  • The error code (step_failed) and completion reason (budget_exceeded) are unchanged. The kernel is untouched.

Tests

packages/sdk/tests/budget-attribution.test.ts adds three tests:

  • An end-to-end AuthoredBudget.execute refusal with the f28314ed totals, checking the full message, code and completion reason.
  • Several crossed dimensions, unmetered dollars, and a day window.
  • Fallback to the generic wording when a total equals its limit, since equal is not over.

Mutation check: I reverted src/authored-budget.ts to origin/main and stubbed the new export so the file compiled.

$ npx vitest run tests/budget-attribution.test.ts   # with the fix reverted
   × authored budget refusal names the limit it crossed > reports wallclock used against the declared header on refusal
   × authored budget refusal names the limit it crossed > names every crossed dimension, marks unmetered dollars, and scopes a day window
   × authored budget refusal names the limit it crossed > keeps the generic wording when the carried totals do not explain the kernel refusal
      Tests  3 failed | 5 passed (8)

Then I restored the file (cmp identical) and re-ran:

      Tests  8 passed (8)

Full SDK suite (npm test: kernel build, typecheck, build, test typecheck, vitest), run locally:

 Test Files  8 failed | 220 passed | 3 skipped (231)
      Tests  13 failed | 3563 passed | 27 skipped (3603)

The 13 failures are live-daemon, webhook and trigger tests (webhook-live, provider-trigger-executor, live-kernel, event-await-cli, communication-mixed-resume, mcp, cloud-run, authored-node-runtime). None of them touch the budget code. Re-running those 8 files with src/authored-budget.ts swapped back to origin/main (rebuilt) gives the identical 13 failed | 115 passed | 20 skipped. They fail on this machine regardless of the change, and CI is the authority for them.

Related finding (not changed here)

The same run spent $39.34 by Claude's total_cost_usd under dollars: 25 without the dollar cap tripping. The kernel's metered dollars come from pricedUsage(model, input_tokens, output_tokens), using worker-usage.ts usageResult and model-pricing.ts. Claude's usage.input_tokens excludes cache_read_input_tokens and cache_creation_input_tokens. Recomputing the run's Claude steps with the published workerSpend gives $6.45 metered against $39.34 reported, so a dollar cap on a cache-heavy Claude flow cannot trip near its declared value. That is reported separately and not fixed in this PR.

🤖 Generated with Claude Code


Note

Cursor Bugbot is generating a summary for commit ee2383c. Configure here.


Summary by cubic

The kernel refuses an authored flow's next step for budget with only budget_exceeded, so the SDK now names which declared limit the carried spend crossed and by how much, instead of the generic "Flow budget exceeded before the next step."

  • Reports wallclock, dollars, and token limits as used vs declared, joined with "; " when several limits were crossed.
  • Names the refused step when the child spec has exactly one step; falls back to the original wording when the carried totals don't explain the refusal, so the SDK never invents a reason.
  • Compares durations as quantities rather than rendered text so rounding never hides a strict overrun, printing exact milliseconds when needed, and compares legacy dollar limits with more than six decimals at their own precision.
  • The error code and completion reason are unchanged.

Written for commit 2847595. Summary will update on new commits.

Review in cubic

The kernel refuses admission of an authored flow's next step with only
`budget_exceeded`, and the authored runner reported "Flow budget exceeded
before the next step." — no limit, no numbers. Cloud run f28314ed stopped
on it after 26 successful steps, and the operator had to reconstruct from
step durations that the flow header's `wallclock: "2h"` (summed step time,
132.9 min) had tripped, not the dollars cap or the Cloud run budget.

The accumulator already carries the exact totals it sends as prior_spend,
so the refusal now names each declared limit the carried spend crossed,
with used vs declared, using the kernel's strict comparisons, and names the
refused step when the child spec has exactly one:

  Flow budget exceeded before step "run-27": wallclock 132.9m used of 2h
  declared in the flow's budget header.

Dollars say when some charges were unmetered, and a day-window budget says
so. When the carried totals cannot explain the refusal, the original
wording stays, so the SDK never invents a reason. Error code and
completion reason are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 6ec5730b-95f2-416e-af7b-a52cb4c2ada2

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-30T15:26:48.881937Z 2847595 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

Devin Review

Comment thread packages/sdk/src/authored-budget.ts Outdated
Comment thread packages/sdk/src/authored-budget.ts Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ee2383c567

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/sdk/src/authored-budget.ts Outdated
Review of #594: rounding both durations to one decimal minute/second could
render a strict overrun as "2m used of 2m declared" (120001 vs 120000 ms)
or a sub-second one as "0s used of 0s"; both now fall back to exact
milliseconds when the rounded forms coincide. A legacy maxDollars with more
than six decimals was dropped from the message although the kernel compares
it as written; the carried microdollars are now compared at the limit's own
precision and the limit is printed as declared.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@khaliqgant

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: dca79253a9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/sdk/src/authored-budget.ts Outdated
Review of #594 (Codex): an exact-hour limit renders in hours and the spend
in rounded minutes, so 7200001 ms against 2h printed "120m used of 2h" and
the string-equality fallback missed it. The fallback now compares the
quantities the texts denote and drops to exact milliseconds unless the
displayed spend itself reads as over the displayed limit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@khaliqgant

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Chef's kiss.

Reviewed commit: 2847595bb5

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@khaliqgant
khaliqgant merged commit b11eea7 into main Oct 1, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant