Skip to content

Add GPT-6.1 Sol and Haiku 5.5 pricing, and correct Sonnet 5 and 5.5 rates - #43

Merged
starkdmi merged 4 commits into
mainfrom
update-model-pricing
Oct 8, 2026
Merged

starkdmi merged 4 commits into
mainfrom
update-model-pricing

Conversation

@starkdmi

@starkdmi starkdmi commented Oct 8, 2026

Copy link
Copy Markdown
Owner

Summary

  • Add GPT-6.1 Sol at $2 input, $0.10 cached input, $2.50 cache writes and $10 output per million tokens. That is GPT-6 Sol's rates with cached input halved. Fast mode is 2x, and the >272K whole-request surcharge applies.

  • Add Claude Haiku 5.5 at $0.10 input, $0.125 cache writes, $0.01 cache reads and $0.50 output. A single request whose prompt exceeds 100K tokens is repriced at 5x on every rate.

  • Give Claude Sonnet 5.5 its own catalog entry. The normalizer was folding it into Sonnet 5. Its cache reads drop from $0.20/M to $0.10/M from 2026-10-07, and aggregates spanning that date stay unpriced.

  • Remove the Sonnet 5 increase to $3/$15 on 2026-09-01. Anthropic cancelled it and made the $2/$10 launch price standard. Sonnet 5 and 5.5 usage since that date was overpriced by about 50%.

  • Stop two aggregate records from claiming to be a single request, which let per-request context tiers price combined usage:

    • Claude stats-cache.json per-model totals.
    • Loose Codex usage lines, such as the thread running total that codex exec --json writes on turn.completed.

    Both are parsed without a request count, and repricing clears the stale count on stored records. Request totals are unchanged, because stats-cache summaries are excluded from rollups and events without a count already roll up as one request.

  • Bump the pricing ruleset from 4 to 5 so stored estimates are repriced.

Cursor CSV rows are deliberately left without a request count: each row is a session meter, not a request.

Verification

  • cargo test -p statsai-pricing -p statsai-adapters -p statsai-store
  • cargo clippy -p statsai-pricing -p statsai-adapters -p statsai-store --all-targets -- -D warnings
  • ./scripts/rust-ci.sh full (passed locally and in the pre-push hook)
  • The new aggregate-request regression tests fail with the parser or repricing fixes reverted.

Pricing sources: OpenAI GPT-6.1 Sol, Anthropic pricing.

GPT-6.1 Sol keeps GPT-6 Sol's input, cache-write and output rates and
halves cached input to $0.10/M. Fast mode is 2x standard, and prompts
above 272K tokens reprice the whole request at 2x input and cache and
1.5x output. Bump the pricing ruleset so stored usage is repriced.
… rise

Haiku 5.5 is $0.10/$0.50 with $0.125 cache writes and $0.01 cache reads
for prompts up to 100K tokens. A single request above that reprices
every rate at 5x.

Sonnet 5.5 was folded into Sonnet 5 by the normalizer. It now has its
own entry, and its cache reads drop from $0.20/M to $0.10/M from
2026-10-07, with aggregates spanning that date left unpriced.

Sonnet 5's $2/$10 launch price became standard and the increase to
$3/$15 scheduled for 2026-09-01 was cancelled, so that boundary is gone.
The stats-cache parser reused the per-message usage reader, which stamps
`requests: 1`. Its per-model totals span every cached session, so Haiku
5.5's above-100K tier repriced combined usage at 5x. Parse the totals
without a request count, and clear the stale count on stored summaries
when they are repriced.
The Codex adapter's fallback for loose `usage` lines stamped
`requests: 1`, but `codex exec --json` writes the thread's running total
on `turn.completed`, so a combined total above 272K was billed at the
GPT long-context rates. Parse these lines without a request count, and
clear the stale count on stored events when they are repriced.
@starkdmi
starkdmi merged commit 2deee99 into main Oct 8, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant