Repair Claude archive collection and cut import work to what changed - #11
Merged
Merged
Conversation
`conversation collect --provider claude_code` failed outright with a foreign key violation. Claude records the conversation identity per line, and a resumed session records its parent's identifier before its own, so trace edits reconstructed mid-file referenced a conversation that was never written. Edits are now re-bound to the conversation their file is finally written as. The import revision moves to archive.v8 so files imported under the old reconstruction are read again, which means the first run after this change re-imports every source. The store did work proportional to the statements it issued rather than to the content that changed: - Re-importing replaced byte-identical rows, deleting and reinserting each one along with its full-text index entry. A row is kept when every persisted field matches, and base64 is decoded only once a write is known to be needed. - Items, content parts, and their reconciliation were written one statement at a time. Once a transaction has touched FTS5, SQLite flushes that index at every statement boundary, so these are now batched, and reconciling stale source records is two statements per conversation against a staged table, joined in an order that keeps it from walking every item recorded for the source. - Archive files are reconstructed on several threads and written in file order, bounded by how much reconstructed content is held at once. A file that fails still leaves every file before it stored. Commit durability is relaxed for the imports alone, which rebuild from the provider's own files. The code-change refresh that follows carries metrics forward for commits Git no longer rescans, so it keeps the store's normal durability. Measured against the previous build on the same archives, back to back: OpenCode 384s to 24s, a 120-file Codex fixture 103s to 6s, and re-imports write no rows where they previously rewrote every one. Imported content is byte-identical to the previous build's.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this fixes
statsai conversation collect --provider claude_codefailed outright withFOREIGN KEY constraint failed(SQLite error 787) — on a fresh store, on an incremental run, and with--no-cachealike.Claude records the conversation identity per line, and a resumed session records its parent's
sessionIdbefore its own. A trace edit reconstructed mid-file was bound to whichever identity was current at that moment, while the conversation was written under the last one seen — so the edit referenced a row that never existed. Where the parent conversation happened to exist from another file the FK was satisfied instead, and the edit was silently filed against the parent.Edits are now re-bound to the conversation their file is finally written as (
TraceEdit::rebind_conversation), deriving a deterministic identifier so a re-import replaces an edit rather than duplicating it. Codex was never affected — it derives its identity once from the file header.archive.v8, so the first collect after this change re-imports every source. That is the point: it is how already-cached files carrying mis-attributed edits get corrected. With the changes below that is roughly 12 minutes for ~7.7 GB rather than the hours it would previously have taken.Why imports were slow
Profiling showed 100% of wall time inside
upsert_archive_conversations, withpread/pwriteand FTS5 flush/merge dominating. The store was doing work proportional to the number of statements it issued, not to the content that actually changed:CROSS JOINto pin the order, since the planner otherwise walked every item recorded for the source and made imports quadratic.Also:
prepare_cachedthroughout, a larger page cache,temp_store = MEMORY, and two adapter fixes — content JSON was rendered twice per item (once to hash, once to store) and tool blocks were deep-cloned purely to serialize them.Durability
Commit durability is relaxed for the imports alone, via a guard that restores the previous value on drop. Imports rebuild from the provider's own files, and each file's rows commit together with the cache entry recording it, so a commit lost to a power cut costs a re-collect. The code-change refresh that follows carries metrics forward for commits Git no longer rescans, across several separate transactions, so it keeps the store's normal durability — as does every other write path.
Results
Measured back-to-back against the previous build on the same archives (absolute timings on this machine vary several-fold with load, so only same-session A/B pairs are meaningful):
Full fresh collect of all four providers (7.7 GB, 2,026 conversations, 318k items): 12.3 min. Steady-state incremental run: ~5 s.
Verification
Counters aren't trusted here — the resulting databases were compared directly. Importing the same fixture with both builds yields identical conversation/item/part counts, text and binary byte totals, trace-edit IDs and FTS row counts, both on a fresh import and after forced re-imports, and with parallel reconstruction enabled.
integrity_checkpasses on the result.910 tests pass;
cargo clippy --workspace --all-targets -- -D warningsandcargo fmt --all --checkare clean. New coverage: re-binding on a resumed session, metadata corrections landing on unchanged bytes, unchanged content not being rewritten, entries cached under an earlier revision being re-read, durability relaxation staying scoped, parallel results preserving file order, and the reconstruction budget gating concurrent claims. The pre-existing per-file durability guarantee (archive_collection_commits_each_candidate_before_the_next) still holds: a file that fails leaves every file before it stored.Deliberately not done
opencode.db, so any use of OpenCode still re-parses 1.6 GB (~20 s, writing nothing). Fixing that changes cache invalidation semantics.sessionIds (2 of 119 here, from resumed sessions) still collapse into one conversation keyed by the last identity seen, with item IDs derived from interim ones. Not a crash, and arguably wrong, but restructuring conversation identity is a separate decision.