Repository navigation
Add GPT-6.1 Sol and Haiku 5.5 pricing, and correct Sonnet 5 and 5.5 rates - #43
Merged
Merged
Conversation
GPT-6.1 Sol keeps GPT-6 Sol's input, cache-write and output rates and halves cached input to $0.10/M. Fast mode is 2x standard, and prompts above 272K tokens reprice the whole request at 2x input and cache and 1.5x output. Bump the pricing ruleset so stored usage is repriced.
… rise Haiku 5.5 is $0.10/$0.50 with $0.125 cache writes and $0.01 cache reads for prompts up to 100K tokens. A single request above that reprices every rate at 5x. Sonnet 5.5 was folded into Sonnet 5 by the normalizer. It now has its own entry, and its cache reads drop from $0.20/M to $0.10/M from 2026-10-07, with aggregates spanning that date left unpriced. Sonnet 5's $2/$10 launch price became standard and the increase to $3/$15 scheduled for 2026-09-01 was cancelled, so that boundary is gone.
The stats-cache parser reused the per-message usage reader, which stamps `requests: 1`. Its per-model totals span every cached session, so Haiku 5.5's above-100K tier repriced combined usage at 5x. Parse the totals without a request count, and clear the stale count on stored summaries when they are repriced.
The Codex adapter's fallback for loose `usage` lines stamped `requests: 1`, but `codex exec --json` writes the thread's running total on `turn.completed`, so a combined total above 272K was billed at the GPT long-context rates. Parse these lines without a request count, and clear the stale count on stored events when they are repriced.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add GPT-6.1 Sol at $2 input, $0.10 cached input, $2.50 cache writes and $10 output per million tokens. That is GPT-6 Sol's rates with cached input halved. Fast mode is 2x, and the >272K whole-request surcharge applies.
Add Claude Haiku 5.5 at $0.10 input, $0.125 cache writes, $0.01 cache reads and $0.50 output. A single request whose prompt exceeds 100K tokens is repriced at 5x on every rate.
Give Claude Sonnet 5.5 its own catalog entry. The normalizer was folding it into Sonnet 5. Its cache reads drop from $0.20/M to $0.10/M from 2026-10-07, and aggregates spanning that date stay unpriced.
Remove the Sonnet 5 increase to $3/$15 on 2026-09-01. Anthropic cancelled it and made the $2/$10 launch price standard. Sonnet 5 and 5.5 usage since that date was overpriced by about 50%.
Stop two aggregate records from claiming to be a single request, which let per-request context tiers price combined usage:
stats-cache.jsonper-model totals.usagelines, such as the thread running total thatcodex exec --jsonwrites onturn.completed.Both are parsed without a request count, and repricing clears the stale count on stored records. Request totals are unchanged, because stats-cache summaries are excluded from rollups and events without a count already roll up as one request.
Bump the pricing ruleset from 4 to 5 so stored estimates are repriced.
Cursor CSV rows are deliberately left without a request count: each row is a session meter, not a request.
Verification
cargo test -p statsai-pricing -p statsai-adapters -p statsai-storecargo clippy -p statsai-pricing -p statsai-adapters -p statsai-store --all-targets -- -D warnings./scripts/rust-ci.sh full(passed locally and in the pre-push hook)Pricing sources: OpenAI GPT-6.1 Sol, Anthropic pricing.