Skip to content

Account plan evidence - #14

Merged
starkdmi merged 13 commits into
mainfrom
account-plan-evidence
Aug 26, 2026
Merged

starkdmi merged 13 commits into
mainfrom
account-plan-evidence

Conversation

@starkdmi

@starkdmi starkdmi commented Aug 25, 2026 •

Copy link
Copy Markdown
Owner

Adds Codex account-plan and identity evidence collection, and fixes the scan
performance and reconciliation defects found while building it.

What this adds

Codex writes identity and plan claims into several local artifacts. This reads
them into four new local tables — account_identity_observations,
account_plan_observations, conversation_account_bindings, and
account_evidence_checkpoints — and syncs a privacy-safe projection of the plan
observations plus aggregate coverage summaries.

Evidence sources: auth.json and swapped-out auth-<label>.json snapshots,
codex_otel.* telemetry rows, Reloading auth for account lines, rate-limit
reset history, and login logs. Conversation bindings can re-attribute individual
events away from their source's directory-level account assignment.

Sync moves to sync_batch.v5 / sync_ack.v5, which adds counters and
authoritative snapshot IDs for the two new collections. Only the canonical
account reference, provider, plan names, observation time, provider-reported
bounds, evidence kind, confidence, and aggregate counts leave the machine —
email, provider user ID, conversation and turn IDs, artifact paths, and raw
provenance stay local.

What this does not do

It does not yet re-attribute historical usage to the account that actually
incurred it. On a real 11-month store it produced 26 identity observations and
113 conversation bindings, every one of which agreed with the directory-level
assignment already in place — zero events changed account. Codex only retains
identity-bearing telemetry for a short window (two disjoint ranges totalling
about five days in the store this was validated against), and the current
generation stopped emitting those attributes altogether. Treat this as
groundwork, not a shipped attribution capability.

The known next step is reading last_refresh on swapped-out auth-<label>.json
files as an interval end, paired with auth_time as the start. That is what
would let a directory's ownership actually change hands over time.

Performance and correctness fixes

  • A full reparse deleted every quota row a source owned, then reinserted them.
    On a multi-gigabyte store that ran for tens of minutes. Cache invalidation now
    reconciles by file, and a full reread (--no-cache) retires only rows no
    rescanned file explains. Measured on a 3.8 GB store: scan --no-cache went
    from >10 minutes to 69s, with every table byte-identical afterwards.
  • Source reconciliation reads were scoped, and unchanged quota rows are no
    longer rewritten.
  • A historical auth-<label>.json was recorded as an AuthSnapshot, which closed
    a source's account interval that nothing could reopen — permanently orphaning
    every later event. It is recorded as the LoginSuccess it evidences instead.
  • Identity is dated by auth_time rather than
    chatgpt_subscription_last_checked, which moves on plan revalidation and
    claimed an account switch weeks before the login.
  • Repeated same-identity telemetry collapses to run endpoints across
    incremental scans, so the ledger stops growing on every daemon pass.
  • Usage events and conversation bindings hashed session identity from disjoint
    spaces, so no binding could ever match an event.
  • Quota cycles no longer count as rollup metadata, which double-charged them
    against the D1 budget and could retry a batch forever.
  • The Release workflow no longer runs on pull requests; releases come from
    version tags only.

Schema

Local store schema 18 → 22, additive only (new tables and indexes, no drops or
data rewrites). The store refuses to open a database migrated by a newer binary,
so a machine that runs this cannot be downgraded without restoring its store.

Requires the hosted API to accept sync_batch.v5 before this ships — the
collector emits v5 with no downgrade path.

Verification

Rust: 210 adapter, 237 store, 193 CLI tests, fmt and clippy clean. Validated
against a clone of a real 3.8 GB store: full reparse and re-reparse both
lossless, evidence ledger and attribution counts unchanged.

Uncommitted collector/store/sync work for account-plan evidence, shelved so
the quota-cycle stack (PR #13) can be tested locally. Not reviewed for merge:
carries a sync_batch.v4 definition that conflicts with PR #13.
Attribution:
- A conversation bound to two accounts no longer discards a manual
  assignment. Clearing the account left `account_identity_source` still
  naming `UserConfigured` against an account that was gone.
- Reset history no longer truncates a source-wide assignment. The
  truncating loop accepted any strong observation while only an auth
  reload could reopen an interval, so one conversation-scoped entry
  naming another account unattributed the rest of the source for good.

Evidence integrity:
- Telemetry identity is read only from the structured attribute prefix,
  and only when the body names its event and its account exactly once.
  A prompt quoting `event.name` and `user.account_id` could otherwise
  mint a binding for an account this device never used.
- A login is identified by what it says and when, not by its file and
  line, so rotating `codex-login.log` no longer duplicates the login
  history that these append-only rows keep forever.

Durability:
- The legacy plan conversion runs inside `Store::migrate`, so one
  unparsable subscription or account payload rolled the transaction back
  before the completion flag was written and failed every later open. It
  now skips rows it cannot read; they could not have been converted.

Cost:
- Quota rows carrying the same plan collapse into one observation per
  run instead of one per raw reading. Ten thousand rows saying "Pro"
  were ten thousand ledger rows, as many sync queries, and a snapshot
  split across fifty fragments.
- Conversation reattribution reads only its own source's events and
  returns early when the source has no bindings, instead of
  deserializing every row of `usage_events` on every scan.

Docs:
- The account-plan privacy paragraph no longer splits the stripped-field
  list into two disconnected lists.
The truncating loop already ignored it, but the loop that reopens an
interval still searched all strong evidence for its confirmation and its
end. A turn-scoped reset-history entry naming another account could
therefore close the interval an auth reload had just opened, leaving the
rest of the source unattributed with nothing able to reopen it.

Also names sync_batch.v5 as the current schema in the README and the task
collection notes.
`chatgpt_subscription_last_checked` moves when the embedded plan claims
are revalidated, which is not when the account changed: signing out of
one account and into another rewrites the account id and leaves that
stamp where it was. Using it for the identity observation therefore
claimed the source was already on the new account weeks before the
login, and an `AuthSnapshot` ends a source's account interval, so the
previous account lost every event in between. The identity now carries
`auth_time`; the plan observation keeps the subscription check.

Quota cycles also stop counting as rollup metadata. They have their own
splitting path and the backend upserts the whole collection in a single
statement, so counting them charged them twice against the D1 budget,
made `has_non_quota_cycle_payload` true for a batch holding nothing
else -- which returned the same chunk from every split and retried it
forever -- and sent each contribution through both splitters.
The provider names the plan while serving a request, so a quota reading
is evidence of the plan as of that moment -- and a fresher one than
`auth.json`, which can sit unchanged on disk for weeks. It was emitted
with `is_current_snapshot: false`, so an account whose logs report "plus"
every day still read as `last_detected` the moment its last declared
provider period ran out, and it dropped off every current-plan surface.

Still no invented billing window: the reading declares no period, so
`active_from` and `active_until` stay empty and the timeline keeps
showing it as an observation rather than a provider period.
Stripping them left every account in the dashboard showing as a bare
`acct_` hash with no way to tell one of your own accounts from another,
which is most of the reason accounts sync at all. The account's own
identity is not what this feature had to make private: the new evidence
types are, and those carry hashes under their own contracts. Server-side
canonicalisation already prefers the hashes and only falls back to these,
so it was unaffected either way.

`plan_name` still goes, because a plan is evidence now rather than an
account attribute.
The account attribution shipped inert: every gate in the collector was
tuned to artifacts the current Codex build no longer produces, so the
evidence tables stayed empty and every source kept the one-account-per-
directory attribution this feature was meant to replace.

- Telemetry identity was only accepted from `codex.conversation_starts`
  events, which this generation emits a handful of times in total. The
  account id rides on ordinary events (`codex.turn_ttft`,
  `codex.tool_decision`, ...), so any single declared `codex.*` event in
  the trusted structured prefix now qualifies. The free-text cut, the
  sole-attribute rule, and the double-event rejection are unchanged.
- Auth reloads moved to the `codex_login::auth::manager` target, and the
  message now closes a tracing span (`app_server.request{...}: Reloading
  auth for account <id>`). The line is accepted at start of body or
  right after a span close, nothing may follow the account id, and a
  match at or past the first free-text field stays inert.
- Multi-account setups keep swapped-out logins as `auth-<label>.json`
  beside the live file and switch by swapping. Each is a genuine auth
  artifact for a real past login, so it is read as a dated historical
  snapshot: identity and plan claims observed at that moment, never
  current, and skipped entirely when no timestamp can be derived.
- Usage events hashed their session identity from the file path, while
  conversation bindings hash the id Codex reports as `conversation.id`.
  The two spaces are disjoint, so no binding could ever attribute an
  event. The session_meta `id` is that same UUID and now becomes the
  session identity as soon as the file declares it; the scan cache
  namespace moves so already-parsed files pick it up.
- A conversation restates the same account hundreds of times, each row
  went to the ledger, and the ledger is re-read on every reconcile.
  Consecutive same-identity telemetry and reload observations now
  collapse to their run endpoints. The run key is conversation-blind so
  parallel conversations cannot split runs; every alternation - the
  actual switch signal - survives exactly.

The evidence parser version moves to v2 so checkpointed telemetry
databases are re-read under the new rules.
Two review findings against the evidence collector, plus release-CI
noise.

- A swapped-out `auth-<label>.json` was recorded as an AuthSnapshot, and
  `ends_source_attribution` turns that into a source-wide boundary. But
  a historical file never becomes current again, so nothing could reopen
  the interval it closed: every event after it was permanently
  unattributed. It is now recorded as the LoginSuccess it evidences -
  dated identity and plan claims that create the account and its
  history without ever acting as an auth-state boundary.
- The in-scan run collapse never crossed a telemetry checkpoint: each
  incremental scan sees only rows past its high-water mark, so every
  daemon pass appended a fresh pair of endpoints and the ledger kept
  growing - the exact cost the collapse exists to remove. The store now
  slides the persisted endpoint instead: when the two newest rows for
  the source already continue the incoming row's run, the incoming row
  replaces the newest one. A run's first point is never touched, any
  differing identity in between breaks the match, and a replayed
  endpoint from a full rescan is recognized by id and left alone, so
  every alternation still survives.

The Release workflow also ran on every pull request - plan-only and
unable to publish, but a transient network failure inside it still
marks the PR red. Releases come from version tags alone now
(pr-run-mode = "skip").
Quota-plan compaction grouped by account before time, so an A-B-A source
timeline could erase the first A run. Keep source time order, make migration
21 retry-safe, and reserve the backend ownership queries in HTTP chunk sizing.
Preserve authentication boundaries and make the quota index migration retry-safe.
Scope source reconciliation reads and skip unchanged record rewrites.
Reconcile cache invalidations by file instead of deleting all quota history.
Select the underscored `user_account_id` alias in the log row filter so
telemetry carrying it without an email reaches the parser.
Retire the previous evidence checkpoints so devices replay rows they
already skipped.
A `--no-cache` scan rereads every file, so file-level reconciliation already
rewrites everything it produces. Deleting the source first walked every
observation, window, and payload row, which stalled for tens of minutes on a
multi-gigabyte store. Rows no rescanned file explains are retired directly
instead. `--replace` keeps the blanket delete it documents.
The widened telemetry filter accepts the `user_account_id` spelling, but no
observed Codex build writes it on an allowlisted target, so retiring every
checkpoint to replay for it cost about 34 seconds per device and recovered
nothing. The version stays at v2 and the filter fix stands on its own.

The auth fixture also carried a real ChatGPT account id. It is now a
placeholder, matching the synthetic emails already in that test.
@starkdmi
starkdmi merged commit e46e6a6 into main Aug 26, 2026
4 checks passed
@starkdmi
starkdmi deleted the account-plan-evidence branch September 1, 2026 13:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant