Skip to content

Repair Claude archive collection and cut import work to what changed - #11

Merged
starkdmi merged 1 commit into
mainfrom
perf/conversation-collect
Aug 21, 2026
Merged

starkdmi merged 1 commit into
mainfrom
perf/conversation-collect

Conversation

@starkdmi

@starkdmi starkdmi commented Aug 21, 2026 •

Copy link
Copy Markdown
Owner

What this fixes

statsai conversation collect --provider claude_code failed outright with FOREIGN KEY constraint failed (SQLite error 787) — on a fresh store, on an incremental run, and with --no-cache alike.

Claude records the conversation identity per line, and a resumed session records its parent's sessionId before its own. A trace edit reconstructed mid-file was bound to whichever identity was current at that moment, while the conversation was written under the last one seen — so the edit referenced a row that never existed. Where the parent conversation happened to exist from another file the FK was satisfied instead, and the edit was silently filed against the parent.

Edits are now re-bound to the conversation their file is finally written as (TraceEdit::rebind_conversation), deriving a deterministic identifier so a re-import replaces an edit rather than duplicating it. Codex was never affected — it derives its identity once from the file header.

⚠️ The import revision moves to archive.v8, so the first collect after this change re-imports every source. That is the point: it is how already-cached files carrying mis-attributed edits get corrected. With the changes below that is roughly 12 minutes for ~7.7 GB rather than the hours it would previously have taken.

Why imports were slow

Profiling showed 100% of wall time inside upsert_archive_conversations, with pread/pwrite and FTS5 flush/merge dominating. The store was doing work proportional to the number of statements it issued, not to the content that actually changed:

  • Byte-identical rows were rewritten. The retention check treated equal content as "replace", so every re-import deleted and reinserted every content part and dragged the full-text index through both. A row is now kept when every persisted field matches, and base64 is decoded only once a write is known to be needed.
  • Everything was one statement at a time. Once a transaction touches FTS5, SQLite flushes that index at every statement boundary. Item and part writes are now batched, and reconciling stale source records went from two statements per item (~680k) to two per conversation against a staged table — joined with CROSS JOIN to pin the order, since the planner otherwise walked every item recorded for the source and made imports quadratic.
  • Parsing was single-threaded. Files are now reconstructed on several threads and written in file order (two files can describe the same conversation, so the winner must not depend on scheduling), bounded by how much reconstructed content is held at once.

Also: prepare_cached throughout, a larger page cache, temp_store = MEMORY, and two adapter fixes — content JSON was rendered twice per item (once to hash, once to store) and tool blocks were deep-cloned purely to serialize them.

Durability

Commit durability is relaxed for the imports alone, via a guard that restores the previous value on drop. Imports rebuild from the provider's own files, and each file's rows commit together with the cache entry recording it, so a commit lost to a power cut costs a re-collect. The code-change refresh that follows carries metrics forward for commits Git no longer rescans, across several separate transactions, so it keeps the store's normal durability — as does every other write path.

Results

Measured back-to-back against the previous build on the same archives (absolute timings on this machine vary several-fold with load, so only same-session A/B pairs are meaningful):

  before after
OpenCode (1.6 GB), fresh 384 s 24 s
OpenCode, forced re-import 51 s 20 s — 0 parts rewritten (was 63,886)
Codex 120-file fixture, fresh 103 s 6 s
Codex fixture, forced re-import 16 s 3 s — 0 parts rewritten (was 34,196)
claude_code, full crashed 11 s

Full fresh collect of all four providers (7.7 GB, 2,026 conversations, 318k items): 12.3 min. Steady-state incremental run: ~5 s.

Verification

Counters aren't trusted here — the resulting databases were compared directly. Importing the same fixture with both builds yields identical conversation/item/part counts, text and binary byte totals, trace-edit IDs and FTS row counts, both on a fresh import and after forced re-imports, and with parallel reconstruction enabled. integrity_check passes on the result.

910 tests pass; cargo clippy --workspace --all-targets -- -D warnings and cargo fmt --all --check are clean. New coverage: re-binding on a resumed session, metadata corrections landing on unchanged bytes, unchanged content not being rewritten, entries cached under an earlier revision being re-read, durability relaxation staying scoped, parallel results preserving file order, and the reconstruction budget gating concurrent claims. The pre-existing per-file durability guarantee (archive_collection_commits_each_candidate_before_the_next) still holds: a file that fails leaves every file before it stored.

Deliberately not done

  • Per-session incremental import for OpenCode. Its cache key covers the whole opencode.db, so any use of OpenCode still re-parses 1.6 GB (~20 s, writing nothing). Fixing that changes cache invalidation semantics.
  • One transaction per group of files. Worth ~16%, but it breaks the per-file durability that is deliberately tested.
  • Claude files containing two sessionIds (2 of 119 here, from resumed sessions) still collapse into one conversation keyed by the last identity seen, with item IDs derived from interim ones. Not a crash, and arguably wrong, but restructuring conversation identity is a separate decision.

`conversation collect --provider claude_code` failed outright with a
foreign key violation. Claude records the conversation identity per
line, and a resumed session records its parent's identifier before its
own, so trace edits reconstructed mid-file referenced a conversation
that was never written. Edits are now re-bound to the conversation
their file is finally written as. The import revision moves to
archive.v8 so files imported under the old reconstruction are read
again, which means the first run after this change re-imports every
source.

The store did work proportional to the statements it issued rather
than to the content that changed:

- Re-importing replaced byte-identical rows, deleting and reinserting
  each one along with its full-text index entry. A row is kept when
  every persisted field matches, and base64 is decoded only once a
  write is known to be needed.
- Items, content parts, and their reconciliation were written one
  statement at a time. Once a transaction has touched FTS5, SQLite
  flushes that index at every statement boundary, so these are now
  batched, and reconciling stale source records is two statements per
  conversation against a staged table, joined in an order that keeps
  it from walking every item recorded for the source.
- Archive files are reconstructed on several threads and written in
  file order, bounded by how much reconstructed content is held at
  once. A file that fails still leaves every file before it stored.

Commit durability is relaxed for the imports alone, which rebuild from
the provider's own files. The code-change refresh that follows carries
metrics forward for commits Git no longer rescans, so it keeps the
store's normal durability.

Measured against the previous build on the same archives, back to
back: OpenCode 384s to 24s, a 120-file Codex fixture 103s to 6s, and
re-imports write no rows where they previously rewrote every one.
Imported content is byte-identical to the previous build's.
@starkdmi
starkdmi merged commit 8da110c into main Aug 21, 2026
11 checks passed
@starkdmi
starkdmi deleted the perf/conversation-collect branch September 1, 2026 13:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant