Repository navigation
feat: journal each step's actual cost beside its metered charge (display only) - #595
Conversation
…lay only) A step's metered `budget.dollars` prices only plain input and output tokens, so it undercounts cache-heavy Claude agents about sixfold (cloud#4033's run metered $6.45 against $39.34 reported). This records what the step actually cost, for display, without changing what counts toward `maxDollars`. - SDK `reportedCost()` (reported-cost.ts): the CLI's own `total_cost_usd` (source `cli`), else a full-usage estimate at the existing frozen rates with cache writes at 1.25x and cache reads at 0.1x input (source `priced`), else nothing: an unknown cost stays unknown, never $0. - Workers send it as `reported_cost` on `step.complete`. The kernel validates it as a decimal and journals it on `step.completed` beside `budget`. It never enters AttemptResult budget folding, so admission and `budget_exceeded` are unchanged. Optional and omitted when absent, so older journals replay and serialize byte-identically. - Read API: run-state (`flows status --json`) adds per-step metered `spend` and `reported_cost`, and a run-level `reported_cost` total. The total is marked incomplete when a model attempt reported nothing or compaction dropped attempts; deterministic steps count as $0. - MODEL_PRICING is unchanged: no model ids were added, so enforcement prices exactly what it did before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Devin Review found 2 potential issues.
2 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)
… once - Segment rollover keeps earlier completions in the journal and the fold has already read them, so an epoch summary no longer resets per-step spend/reported_cost or marks the run's reported total incomplete. - Codex's input_tokens already includes cached_input_tokens: the priced estimate now splits them out instead of charging cached input twice. Claude, which reports cache tokens beside input, is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ebe7bd69b9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…no model - Memoized reuse copied the source completion's reported_cost into the new run, double-counting a model that did not execute. The reuse path now clears it alongside `budget` (Codex P2). - The fold treated only deterministic steps as free, so agent-typed steps that never call a model (authored root, helpers) marked the total incomplete. A completion is free when it is a memoized reuse, or metered zero tokens and zero dollars without flagging them unknown. Every model attempt meters its tokens, so a missing reported_cost there is still unknown (Cursor). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 2426fd0. Configure here.
Crash recovery and cancellation journal a default zero budget for agent and llm attempts whose model did run, so the zero-token rule presented a lower bound as a complete total (Cursor). A completion without reported_cost is now known-free only when it is a deterministic step, a memoized reuse, or the authored root (its model calls are child runs). Everything else stays unknown and the total reads as a lower bound. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a02a7d68db
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…red-root The free-completion check keyed on the step id alone, so an ordinary flow with an agent step named `authored-root` would read as known $0 (Codex). It now requires the spawned spec's `relayflows.authored-root.v1` discriminator, via the same `declaresAuthoredRoot` the authored-verdict projection uses. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 27c25758cc
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
… are model-free - The reported-cost types and fold move to reported-cost-total.ts, and exact decimal addition to decimal-dollars.ts. run-state.ts drops to 456 lines, below its 468 on main (Codex P1, small-module rule). `addDollars` and `ReportedCostTotal` are re-exported, so run-state's public API is unchanged. - Model-free steps are derived once from the spawned spec: deterministic steps, provider helpers (`relayflows:helper:v1` instruction, now a shared constant in helper-instruction.ts), and a discriminated authored root. Helper-only flows now report a measured $0 rather than a lower bound (Codex P2). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ea59824e92
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…kers report no_model
- Version skew (Codex P1): a newer SDK attached to a pre-upgrade daemon would
send `reported_cost`, which that daemon's `deny_unknown_fields` rejects,
failing the completion. The daemon now lists `features: ["reported_cost"]`
in its protocol-0 hello reply, and JournalClient sends the field only when
advertised, dropping it otherwise. Metered usage is always sent. An older
SDK against a newer daemon never sends it. The protocol stays 0: this is
additive and needs no lockstep upgrade.
- Internal effect workers (authored provider helper, plugin, MCP, YAML
helper) run no model. They now complete with an explicit
`reported_cost: { dollars: "0.000000", source: "no_model" }`, which the fold
counts as a known $0 without naming a cost source. This relies on the worker
that completed the step, not on matching instruction text (Codex P2).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fbf18f3e6f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
- The kernel refuses a `reported_cost` whose source is `no_model` but carries a non-zero amount, so a malformed completion fails closed instead of being silently read as a complete $0 (Codex P2). - `completeChannelPost` is a broker effect that runs no model; it now sends NO_MODEL_COST like the other internal effect workers (Codex P2). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
|
Codex Review: Didn't find any major issues. Nice work! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |

Why
A step's metered
budget.dollars, the figure the kernel charges againstmaxDollars, prices only plain input and output tokens (workerSpend→pricedUsage). Claude's cache-read and cache-creation tokens are left out, so cache-heavy agents are undercounted roughly sixfold. cloud#4033's software-factory run metered $6.45 against $39.34 the Claude CLI reported, which is why itsdollars: 25cap never tripped.This PR shows the real cost. It does not change what counts toward a cap. Whether to enforce on real cost is a separate decision, and the read-only audit of deployed caps feeds it.
What
reported_costonstep.completed(display only), besidebudget:source: "cli": the CLI's owntotal_cost_usd.source: "priced": an estimate from the digest's full token usage at the existing frozenMODEL_PRICINGrates. Cache writes count at 1.25× and cache reads at 0.1× the input rate. Only used for a model that already has a price.source: "no_model": sent by the internal effect workers (provider helper, plugin, MCP), which run no model. It counts as a known $0.features: ["reported_cost"]in its protocol-0helloreply, andJournalClientsends the field only when that's advertised. A newer SDK against an older daemon drops it and never trips itsdeny_unknown_fields. No protocol bump.step.completeaccepts an optionalreported_cost, validates it as a non-negative decimal, and journals it verbatim. It is carried onAttemptResultonly to reach the journal entry. It is never folded into budget state, so admission,dollars_unmeteredandbudget_exceededare unchanged. The field isOptionwithskip_serializing_if, so journals written before it replay and re-serialize byte-identically (tested).flows status --json, adds:spend(the step's metered charges across attempts) andreported_cost: { dollars, complete, source } | null;reported_cost: { dollars, complete, source }.complete: falsewhen a model attempt reported no cost, sodollarsis then a lower bound.no_modelreports, memoized reuse (reused_from), and a discriminated authored root. A zero metered budget alone is not proof no model ran, since crash recovery journals one.reported_cost(kernel).reported-cost-total.ts, andrun-state.tsis shorter than on main.costUsd(the last attempt's transcript total) is unchanged.Enforcement is unchanged
MODEL_PRICINGis untouched. No model ids were added, so metered pricing and everymaxDollarsverdict are exactly what they were.reported_costis display-only by construction: the kernel never reads it when charging.Pinned by tests:
budget_gate.rs→reported_cost_is_journaled_but_never_charged_to_the_budget: a $7.17 reported cost under a $0.01 cap still admits the next step. The journal'sbudget/spendstay the metered $0.001.budget-unmetered-live.test.ts→ an unpriced Claude model with a reported $12.50 under a $0.000001 cap journalsreported_costand still succeeds, withbudgetexactly{dollars: '0', dollars_unmetered: true}as before.reported-cost.test.ts→workerSpendstill prices input and output only ($0.972355) while the reported estimate includes cache ($2.597355).Validation
deny_unknown_fieldswould refuse it.cargo test --workspace: all pass.cargo clippy: no findings in changed lines; the existing findings are all present on main too.typecheck,typecheck:testsandbuildpass.vitest run: 3573 passed, 6 failed. The 6 failures are in 5 files (surfacedistnot built in this worktree, an installed package at 1.4.2 vs a pinned 1.4.0, a Cloud transport test, a daemon socket, and the event-await CLI). They fail identically with this PR's SDK sources reverted to main, so they are environmental.cargo fmt: not applied repo-wide. main itself isn't rustfmt-clean under the current toolchain, so only the changed hunks were formatted.🤖 Generated with Claude Code
Note
Medium Risk
Touches step completion validation and journal schema with backward-compat gating, but budget enforcement paths are unchanged and covered by new tests; main risk is wire/client skew or misread totals in status views.
Overview
Adds display-only actual step cost alongside the existing metered
budget, without changing dollar-cap enforcement.Journal & wire: Optional
reported_costonstep.completed(dollars+source:cli,priced, orno_model). The kernel validates and journals it but never folds it into budget state. Memoized reuse clearsreported_costso the reusing run does not inherit the source run’s figure.SDK workers: LLM/agent workers compute cost from CLI
total_cost_usdor full-token pricing (including cache); helper/plugin/MCP/channel paths sendno_modelat $0.step.completeincludesreported_costonly whenhelloadvertisesfeatures: ["reported_cost"], so older daemons are not broken by unknown fields.Status UI:
flows status --jsongains per-step and run-levelreported_costtotals (completeflags missing costs) plus per-stepspend, while runspendstays metered-only.addDollarsmoves todecimal-dollars.tsfor shared exact summation.Reviewed by Cursor Bugbot for commit 13e2c9f. Bugbot is set up for automated code reviews on this repo. Configure here.