Repository navigation
Add a prompt-cache report and per-day cache diagnostics - #42
Merged
Merged
Conversation
Add `statsai report cache`, which shows how much logical input the prompt cache served, how much context each model call processed, and when reuse dropped. For Claude Code and Codex it walks each agent's calls in order and flags a suspected cache loss when cached reads fall at least 50% and 4,096 tokens below what the previous call left reusable. Each loss records the reusable context it missed and its cost: the call's price minus its price had that context been read from the cache, priced as one request by the call's own model. Losses are grouped by the gap since the previous call. Model changes, compactions, and calls that cannot be ordered or split are boundaries, not losses. A Claude Code call that neither read nor wrote the cache, as when a custom endpoint serves it, is counted apart and not judged. The report takes --from/--to, --all, --provider, --account, --session, --timeline for local 10-minute buckets, --details for per-call evidence, and --json. The daemon serves the same report at /reports/cache, and `statsai schema cache-report` prints its schema. - Codex turn events list each model call with its request start, model, and compaction boundary. Rows too large to read still mark compaction. - Claude Code events record their request start, sub-agent, and compaction boundary, following file order because compaction rewrites history. - Daily summaries carry an optional, compact `cache_health` object. It is sent only to receivers that advertise support, and an existing store gains it once, rewriting only summaries whose content changes. Fields that are zero on most days are left out then. - The cache hit ratio now counts cache writes in its denominator. - Migration 30 orders the session index by start time, so sessions read in order and the calls after a changed one are an index range.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add
statsai report cache, which shows how much logical input the promptcache served, how much context each model call processed, and when reuse
dropped. For Claude Code and Codex it walks each agent's calls in order and
flags a suspected cache loss when cached reads fall at least 50% and 4,096
tokens below what the previous call left reusable. Each loss records the
reusable context it missed and its cost: the call's price minus its price
had that context been read from the cache, priced as one request by the
call's own model. Losses are grouped by the gap since the previous call.
Model changes, compactions, and calls that cannot be ordered or split are
boundaries, not losses. A Claude Code call that neither read nor wrote the
cache, as when a custom endpoint serves it, is counted apart and not judged.
The report takes --from/--to, --all, --provider, --account, --session,
--timeline for local 10-minute buckets, --details for per-call evidence,
and --json. The daemon serves the same report at /reports/cache, and
statsai schema cache-reportprints its schema.and compaction boundary. Rows too large to read still mark compaction.
boundary, following file order because compaction rewrites history.
cache_healthobject. It issent only to receivers that advertise support, and an existing store
gains it once, rewriting only summaries whose content changes. Fields
that are zero on most days are left out then.
in order and the calls after a changed one are an index range.