Skip to content

feat: journal each step's actual cost beside its metered charge (display only) - #595

Merged
khaliqgant merged 8 commits into
mainfrom
feat/journal-reported-cost
Sep 30, 2026
Merged

khaliqgant merged 8 commits into
mainfrom
feat/journal-reported-cost

Conversation

@khaliqgant

@khaliqgant khaliqgant commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Why

A step's metered budget.dollars, the figure the kernel charges against maxDollars, prices only plain input and output tokens (workerSpend → pricedUsage). Claude's cache-read and cache-creation tokens are left out, so cache-heavy agents are undercounted roughly sixfold. cloud#4033's software-factory run metered $6.45 against $39.34 the Claude CLI reported, which is why its dollars: 25 cap never tripped.

This PR shows the real cost. It does not change what counts toward a cap. Whether to enforce on real cost is a separate decision, and the read-only audit of deployed caps feeds it.

What

  • reported_cost on step.completed (display only), beside budget:
    "reported_cost": { "dollars": "7.169405", "source": "cli" }
    • source: "cli": the CLI's own total_cost_usd.
    • source: "priced": an estimate from the digest's full token usage at the existing frozen MODEL_PRICING rates. Cache writes count at 1.25× and cache reads at 0.1× the input rate. Only used for a model that already has a price.
    • source: "no_model": sent by the internal effect workers (provider helper, plugin, MCP), which run no model. It counts as a known $0.
    • Otherwise absent: an unknown cost is never journaled as $0 (e.g. Codex).
  • Version skew: the daemon lists features: ["reported_cost"] in its protocol-0 hello reply, and JournalClient sends the field only when that's advertised. A newer SDK against an older daemon drops it and never trips its deny_unknown_fields. No protocol bump.
  • Kernel: step.complete accepts an optional reported_cost, validates it as a non-negative decimal, and journals it verbatim. It is carried on AttemptResult only to reach the journal entry. It is never folded into budget state, so admission, dollars_unmetered and budget_exceeded are unchanged. The field is Option with skip_serializing_if, so journals written before it replay and re-serialize byte-identically (tested).
  • Read API: run-state, i.e. flows status --json, adds:
    • per step: spend (the step's metered charges across attempts) and reported_cost: { dollars, complete, source } | null;
    • per run: reported_cost: { dollars, complete, source }.
    • complete: false when a model attempt reported no cost, so dollars is then a lower bound.
    • Known $0, without making the total incomplete: deterministic steps, no_model reports, memoized reuse (reused_from), and a discriminated authored root. A zero metered budget alone is not proof no model ran, since crash recovery journals one.
    • Memoized reuse no longer copies the source run's reported_cost (kernel).
    • The fold lives in reported-cost-total.ts, and run-state.ts is shorter than on main.
    • The Cloud mirror's existing costUsd (the last attempt's transcript total) is unchanged.

Enforcement is unchanged

MODEL_PRICING is untouched. No model ids were added, so metered pricing and every maxDollars verdict are exactly what they were. reported_cost is display-only by construction: the kernel never reads it when charging.

Pinned by tests:

  • budget_gate.rs → reported_cost_is_journaled_but_never_charged_to_the_budget: a $7.17 reported cost under a $0.01 cap still admits the next step. The journal's budget/spend stay the metered $0.001.
  • budget-unmetered-live.test.ts → an unpriced Claude model with a reported $12.50 under a $0.000001 cap journals reported_cost and still succeeds, with budget exactly {dollars: '0', dollars_unmetered: true} as before.
  • reported-cost.test.ts → workerSpend still prices input and output only ($0.972355) while the reported estimate includes cache ($2.597355).

Validation

  • Red→green: the 4 new SDK tests fail against origin/main's sources and pass here. The kernel tests don't compile without the field, and the wire's deny_unknown_fields would refuse it.
  • cargo test --workspace: all pass. cargo clippy: no findings in changed lines; the existing findings are all present on main too.
  • SDK: typecheck, typecheck:tests and build pass. vitest run: 3573 passed, 6 failed. The 6 failures are in 5 files (surface dist not built in this worktree, an installed package at 1.4.2 vs a pinned 1.4.0, a Cloud transport test, a daemon socket, and the event-await CLI). They fail identically with this PR's SDK sources reverted to main, so they are environmental.
  • cargo fmt: not applied repo-wide. main itself isn't rustfmt-clean under the current toolchain, so only the changed hunks were formatted.

🤖 Generated with Claude Code


Note

Medium Risk
Touches step completion validation and journal schema with backward-compat gating, but budget enforcement paths are unchanged and covered by new tests; main risk is wire/client skew or misread totals in status views.

Overview
Adds display-only actual step cost alongside the existing metered budget, without changing dollar-cap enforcement.

Journal & wire: Optional reported_cost on step.completed (dollars + source: cli, priced, or no_model). The kernel validates and journals it but never folds it into budget state. Memoized reuse clears reported_cost so the reusing run does not inherit the source run’s figure.

SDK workers: LLM/agent workers compute cost from CLI total_cost_usd or full-token pricing (including cache); helper/plugin/MCP/channel paths send no_model at $0. step.complete includes reported_cost only when hello advertises features: ["reported_cost"], so older daemons are not broken by unknown fields.

Status UI: flows status --json gains per-step and run-level reported_cost totals (complete flags missing costs) plus per-step spend, while run spend stays metered-only. addDollars moves to decimal-dollars.ts for shared exact summation.

Reviewed by Cursor Bugbot for commit 13e2c9f. Bugbot is set up for automated code reviews on this repo. Configure here.

…lay only)

A step's metered `budget.dollars` prices only plain input and output tokens,
so it undercounts cache-heavy Claude agents about sixfold (cloud#4033's run
metered $6.45 against $39.34 reported). This records what the step actually
cost, for display, without changing what counts toward `maxDollars`.

- SDK `reportedCost()` (reported-cost.ts): the CLI's own `total_cost_usd`
  (source `cli`), else a full-usage estimate at the existing frozen rates with
  cache writes at 1.25x and cache reads at 0.1x input (source `priced`), else
  nothing: an unknown cost stays unknown, never $0.
- Workers send it as `reported_cost` on `step.complete`. The kernel validates
  it as a decimal and journals it on `step.completed` beside `budget`. It never
  enters AttemptResult budget folding, so admission and `budget_exceeded` are
  unchanged. Optional and omitted when absent, so older journals replay and
  serialize byte-identically.
- Read API: run-state (`flows status --json`) adds per-step metered `spend`
  and `reported_cost`, and a run-level `reported_cost` total. The total is
  marked incomplete when a model attempt reported nothing or compaction
  dropped attempts; deterministic steps count as $0.
- MODEL_PRICING is unchanged: no model ids were added, so enforcement prices
  exactly what it did before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-30T17:34:10.398174Z 13e2c9f Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 74187314-f14f-40e1-83b6-a4f55d58cac8

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

2 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)

Devin Review

Comment thread packages/sdk/src/reported-cost.ts
Comment thread packages/sdk/src/run-state.ts Outdated
… once

- Segment rollover keeps earlier completions in the journal and the fold has
  already read them, so an epoch summary no longer resets per-step
  spend/reported_cost or marks the run's reported total incomplete.
- Codex's input_tokens already includes cached_input_tokens: the priced
  estimate now splits them out instead of charging cached input twice. Claude,
  which reports cache tokens beside input, is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@khaliqgant

Copy link
Copy Markdown
Member Author

@codex review

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/sdk/src/run-state.ts Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ebe7bd69b9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/sdk/src/run-state.ts Outdated
…no model

- Memoized reuse copied the source completion's reported_cost into the new
  run, double-counting a model that did not execute. The reuse path now clears
  it alongside `budget` (Codex P2).
- The fold treated only deterministic steps as free, so agent-typed steps that
  never call a model (authored root, helpers) marked the total incomplete. A
  completion is free when it is a memoized reuse, or metered zero tokens and
  zero dollars without flagging them unknown. Every model attempt meters its
  tokens, so a missing reported_cost there is still unknown (Cursor).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@khaliqgant

Copy link
Copy Markdown
Member Author

@codex review

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 2426fd0. Configure here.

Comment thread packages/sdk/src/run-state.ts Outdated
Crash recovery and cancellation journal a default zero budget for agent and
llm attempts whose model did run, so the zero-token rule presented a lower
bound as a complete total (Cursor). A completion without reported_cost is now
known-free only when it is a deterministic step, a memoized reuse, or the
authored root (its model calls are child runs). Everything else stays unknown
and the total reads as a lower bound.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@khaliqgant

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a02a7d68db

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/sdk/src/run-state.ts Outdated
…red-root

The free-completion check keyed on the step id alone, so an ordinary flow
with an agent step named `authored-root` would read as known $0 (Codex). It
now requires the spawned spec's `relayflows.authored-root.v1` discriminator,
via the same `declaresAuthoredRoot` the authored-verdict projection uses.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@khaliqgant

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 27c25758cc

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/sdk/src/run-state.ts Outdated
Comment thread packages/sdk/src/run-state.ts Outdated
… are model-free

- The reported-cost types and fold move to reported-cost-total.ts, and exact
  decimal addition to decimal-dollars.ts. run-state.ts drops to 456 lines,
  below its 468 on main (Codex P1, small-module rule). `addDollars` and
  `ReportedCostTotal` are re-exported, so run-state's public API is unchanged.
- Model-free steps are derived once from the spawned spec: deterministic
  steps, provider helpers (`relayflows:helper:v1` instruction, now a shared
  constant in helper-instruction.ts), and a discriminated authored root.
  Helper-only flows now report a measured $0 rather than a lower bound (Codex P2).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@khaliqgant

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ea59824e92

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/sdk/src/protocol.ts
Comment thread packages/sdk/src/reported-cost-total.ts
…kers report no_model

- Version skew (Codex P1): a newer SDK attached to a pre-upgrade daemon would
  send `reported_cost`, which that daemon's `deny_unknown_fields` rejects,
  failing the completion. The daemon now lists `features: ["reported_cost"]`
  in its protocol-0 hello reply, and JournalClient sends the field only when
  advertised, dropping it otherwise. Metered usage is always sent. An older
  SDK against a newer daemon never sends it. The protocol stays 0: this is
  additive and needs no lockstep upgrade.
- Internal effect workers (authored provider helper, plugin, MCP, YAML
  helper) run no model. They now complete with an explicit
  `reported_cost: { dollars: "0.000000", source: "no_model" }`, which the fold
  counts as a known $0 without naming a cost source. This relies on the worker
  that completed the step, not on matching instruction text (Codex P2).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@khaliqgant

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: fbf18f3e6f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread kernel/relayflowd/src/engine/remote.rs
Comment thread packages/sdk/src/reported-cost-total.ts
- The kernel refuses a `reported_cost` whose source is `no_model` but carries
  a non-zero amount, so a malformed completion fails closed instead of being
  silently read as a complete $0 (Codex P2).
- `completeChannelPost` is a broker effect that runs no model; it now sends
  NO_MODEL_COST like the other internal effect workers (Codex P2).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@khaliqgant

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Nice work!

Reviewed commit: 13e2c9fe60

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@khaliqgant
khaliqgant merged commit 129aabf into main Sep 30, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant