fix: bind the packed matrix to its sig (the #81 torn-read fix missed the matrix) - #83
Merged
Merged
Conversation
The content_sig added in #81 hashes ids/metas only. The interleaving "writer A saves matrix.npy, writer B saves a whole set, A then writes ids/metas/sig" leaves B's same-shape matrix under A's ids and a sig that vouches for them, so rows were still served under the wrong ids. Each save now writes matrix/ids/metas under a fresh generation token that only that writer touches, and publishes sig.json last via os.replace; a reader follows the sig to exactly one writer's complete set. Older generations are swept after publishing. Legacy flat-layout caches still load. Regression test drives the real interleaving through _save_packed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This was referenced Sep 22, 2026
thorwhalen
added a commit
that referenced
this pull request
Sep 22, 2026
…#84) Post-merge review of #83. - A writer whose disk-clear was skipped by the _packed_stale guard loaded a packed set another process published before its latest writes, so it did not see records it had just written. _load_packed now treats the disk cache as a miss while this process has unpublished writes; the rebuild republishes, which also heals the cache for other readers. - np.load of an empty/truncated .npy raises EOFError, which escaped _load_packed and made every matrix() call raise while that sig stayed. - Data files and the sig tmp are fsynced before os.replace, and the directory after (POSIX), so a power loss cannot leave a durable sig naming data blocks that never reached disk. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Member
Author
|
Post-merge review: three defects fixed in #84. (1) A writer could load another process's older packed set and miss records it had just written, because the |
This was referenced Sep 22, 2026
CorpusStore.matrix() raises KeyError/EOFError when read while another process is writing records
#85
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Post-merge refute review of #81.
Defect
#81's
content_sighashesids.json+metas.jsononly; the matrix is checked by shape alone. The interleaving the PR itself describes still passes every check:np.save(matrix.npy)Disk: B's matrix beside A's ids/metas and a sig that vouches for A's ids/metas. Shapes agree, the sig agrees, so
_load_packedserves rows under the wrong ids. The #81 tests only mutate ids/metas without rewriting the sig, so they never exercise this.Fix
Each
_save_packedwritesmatrix-<gen>.npy,ids-<gen>.json,metas-<gen>.jsonunder a fresh uuid generation that only that writer touches, then publishessig.json(naming the generation, pluscontent_sigandshape) last via an atomicos.replace. A reader follows the sig to exactly one writer's complete set, so a mixed set can't exist on disk. This is option 1 of #77 done with a file-level pointer rather than a directory swap. It also stops a writer truncating amatrix.npythat another process has memory-mapped, because files are never rewritten in place.sig-*.tmp) are swept after publishing, keeping this writer's set and whatever setsig.jsonnames by then. A sweep that races another writer can only cause a cache miss, never wrong rows.matrix.npy+ sig with nogeneration) still load, so no forced rebuild. Olderirreading a new-layout cache finds nomatrix.npy, treats it as a miss and rebuilds.raglab(never setspacked_dir) andtruffle(not on box) can't reach this path through anything public.Tests
test_interleaved_packed_writers_never_serve_mismatched_rowsruns the real interleaving through_save_packedby patchingnp.save. It fails on master (row forr1served asr2's) and passes here.test_legacy_flat_packed_layout_still_loadsandtest_republishing_sweeps_older_generationsare new.Not addressed (still open under #77): a staleness race. If a writer builds a matrix, then another process does
put_record(clearing the cache), then the first writer publishes, the cache misses the new record until the next write. The rows are stale but still correct for their ids.Self-reviewed only (the worker was told not to spawn a sub-agent reviewer). This needs a post-merge review.
🤖 Generated with Claude Code