Skip to content

Count Codex token_usage_record rows and skip repeated token totals - #39

Merged
starkdmi merged 8 commits into
mainfrom
ai/codex-usage-records-19f4
Sep 23, 2026
Merged

starkdmi merged 8 commits into
mainfrom
ai/codex-usage-records-19f4

Conversation

@starkdmi

Copy link
Copy Markdown
Owner

Behaviour

Codex rollout parsing now matches CLI 0.147–0.155 usage lines.

  • Repeated token_count snapshots no longer inflate totals. When total_token_usage is present and equals the previous total in the same file, that line contributes no usage. A second copy for another limit_id (for example premium) is the same case. The line still produces a quota observation, with no usage sample. A missing total_token_usage still uses last_token_usage. A smaller total (fork or resume reset) still counts as new usage.
  • The phantom token_count after compacted is not billed. The compaction fixture repeats total_tokens 11081643 with a different last_token_usage (total_tokens 17983). That 17983 is no longer added. Advancing last-usage totals in that fixture sum to 390224.
  • token_usage_record is a usage source, including compaction. The line is classified before the headless-usage fallback, including when "usage" sits past the 256-byte header. Before a file’s first record, usage still comes from token_count. From the first record on, only payload.usage on records is counted; later token_count lines supply quota and context only. Compaction inference that never appears on a token_count is therefore included. turn_token_usage and thread_token_usage are not summed.
  • cache_write_input_tokens / cacheWriteInputTokens are an inclusive subset of input, clamped the same way as cached input.
  • Record ids are kept on the parsed record only: turn_id, root_turn_id, and response_id. For a sub-agent file, session_id (the parent) and root_turn_id are available so a later change can roll cost up to the parent turn. No store columns were added.

Historical data

Stored history is corrected only when those rollout files are rescanned. CODEX_SCAN_CACHE_PARSER_REVISION is unchanged, so this does not force a full-history resync.

Tests

Synthetic fixtures under tests/fixtures/codex/usage-record/ cover a mixed-era file (legacy token_count, then records, including a compaction record) and a sub-agent file. cargo fmt --check, cargo clippy --workspace --all-targets -- -D warnings, and cargo test --workspace passed.

Open in Web Open in Cursor 

cursoragent and others added 8 commits September 23, 2026 11:57
Repeated token_count snapshots and the post-compaction phantom were added
as new usage. token_usage_record lines, including compaction inference,
were dropped because usage sits past the header window.

Historical sessions keep their stored totals until those files are rescanned.
The Codex scan-cache parser revision is unchanged.

Co-authored-by: Dmitry Starkov <21260939+starkdmi@users.noreply.github.com>
…count

Records are keyed by the same session rule as token_count, so a fork's
copied parent records still pair. A compacted line ends the pairing
window, pairing compares normalized counts, and a paired token_count's
quota sample links to a record emitted outside any turn. A counter
reset without last_token_usage counts the new total, standalone records
get their own event kind, and the archive quota scan follows session_meta.
@starkdmi
starkdmi marked this pull request as ready for review September 23, 2026 18:15
@starkdmi
starkdmi merged commit b182235 into main Sep 23, 2026
4 checks passed
@starkdmi
starkdmi deleted the ai/codex-usage-records-19f4 branch September 23, 2026 18:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants