Skip to content

fix(store): matrix()/metas() no longer raise while another process writes - #87

Merged
thorwhalen merged 2 commits into
masterfrom
fix-matrix-concurrent-reads
Sep 22, 2026
Merged

thorwhalen merged 2 commits into
masterfrom
fix-matrix-concurrent-reads

Conversation

@thorwhalen

@thorwhalen thorwhalen commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Closes #85.

Reproduction

A stress script (6 processes on one file-backed corpus, 200 random put_record / delete_record / matrix() / metas() each, readers forced to rebuild) failed in 10 of 10 runs on master with KeyError (listed id without a file), EOFError (torn .npy) and JSONDecodeError (torn meta). With puts and reads only (no deletes, the #77 scenario), 6 of 6 runs failed. With this branch: 0 of 10 and 0 of 10.

Change

Builds on #83/#84's publish step (write elsewhere, then os.replace) rather than changing the packed-cache design.

  • Vector before meta. put_record writes the vector first. Ids are listed from meta, so a listed id always has a vector. delete_record already removed the meta first, and now also tolerates a concurrent delete of the same record.
  • Atomic per-record writes. _json_store / _ndarray_store write through _AtomicFiles: a hidden temp file in the same directory, then os.replace. Hidden, so the dol.Files listing never shows it as a key. Reads, listing and deletes still go through dol.Files unchanged. No fsync per record (atomic against other processes, not durable against power loss), so a bulk build does not pay one. The on-disk format is unchanged (indent=4, UTF-8 JSON; np.save bytes), so existing corpora read as before. On Windows, os.replace refuses while another process has the file open. The write is retried with back-off, then falls back to the old in-place write rather than failing.
  • Tolerant rebuild. _build_matrix and metas() treat a record that raises KeyError / EOFError / ValueError on read as not yet written. A failed record is read once more: if its meta is gone by then it was deleted, which is not a gap. If any record still can't be read, matrix() logs a warning naming the ids and returns the partial result. It caches that result in-process, like any build, but never publishes it as the packed set.
  • Writes stay inside the store. The write path comes from dol.Files itself (root + key, the same path reads and deletes use), and a key that climbs out of the root is refused.

Tests

tests/test_store_concurrency.py: write order; meta-without-vector and torn meta/vector files are skipped, not cached, not published, and picked up once whole; deleted-mid-read is not a gap; no temp files left and temp files are not keys; the JSON format still matches dol.JsonFiles; the Windows retry-then-in-place path and the POSIX re-raise; absolute keys land inside the root and escaping keys are refused. It also runs a subprocess stress test: two writers build 40 large records (with overwrites) while the test process loops on matrix() / metas(), in three rounds. On master, 8 of the new tests fail deterministically, and the stress test fails in about 5 of 6 runs. On this branch it passed every run. Full suite: 576 passed, 5 skipped, 1 xfailed.

Found on the way, not fixed here: #86

A reader that rebuilds mid-build, and reads nothing torn, publishes a packed set that looks complete but predates the writer's later records. The writer clears the cache only once per write session, so later fresh processes are served that set. That is a correctness gap in what may be published, separate from #85 (reads raising). It is filed as #86 with options, and pinned by a strict xfail test here.

Remaining edge case

Two writers racing put_record and delete_record on the same id can still leave a meta without a vector. Reads now skip that record with a warning instead of raising. Until the record is rewritten or deleted, that corpus's matrix is not published to the packed cache, so each fresh process rebuilds it once.

Dependents

fleet_dependents lists raglab and truffle. raglab: 49 passed against this branch (it uses CorpusStore.memory()). truffle was out of scope for this run and was not tested. The change keeps every public name and signature. The only behaviour changes are that a racing read no longer raises and that per-record files are replaced atomically.

Review

An independent refute-review agent found one blocker, which is fixed in the second commit. _AtomicFiles built the write path with Path(root, key). For an absolute key (an artifact id can be one, e.g. in links), that overwrote a file outside the store, while reads still looked under the root. Its should-fix items are also applied: a permanently unreadable record no longer forces a full re-read on every search in the process, and the stress test is strengthened, since it used to catch master only about 1 run in 5. It confirmed that the on-disk bytes are identical to dol.JsonFiles output, that mapping semantics (in, get, pop, len, KeyError) are unchanged, and that no path caches or publishes an incomplete set. Not changed: temp files left by a crash between write and replace are hidden but never swept; no fsync per record (by design, see above).

🤖 Generated with Claude Code

thorwhalen and others added 2 commits September 22, 2026 15:22
…ites

Closes #85.

- put_record writes the vector before the meta (delete_record already
  removes the meta first), so a listed id always has a vector.
- Per-record files are written atomically (hidden temp file + os.replace,
  the same publish step as the packed cache's sig.json, without fsync).
  On Windows a replace blocked by a reader is retried, then falls back to
  the old in-place write. On-disk format is unchanged.
- _build_matrix/metas skip a record that vanishes or cannot be decoded
  mid-read; a failed record is re-read once (a missing meta then means
  deleted, not a gap). A read that still left records out is returned but
  neither cached nor published as the packed set, and is logged.
- delete_record tolerates a concurrent delete of the same record.

Stale-cache follow-up found on the way: #86 (xfail test added).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ocess

- _AtomicFiles built the write path with Path(root, key), which lets an
  absolute key (an artifact id can be an absolute path, e.g. links) replace
  the root and overwrite a file outside the store, while reads still looked
  under the root. The path now comes from dol.Files itself (root + key, as
  reads and deletes use), and a key climbing out of the root is refused.
- A partial matrix (records unreadable after the retry) is now cached
  in-process, still never published, so one damaged record no longer
  makes every search in that process re-read every record file. The
  warning says how to clear it.
- metas() logs skipped records at debug level.
- The subprocess stress test uses two writers and larger records, in three
  rounds; on master it now fails in about 5 of 6 runs (was 1 of 5).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@thorwhalen
thorwhalen merged commit 7c41dac into master Sep 22, 2026
10 of 12 checks passed
@thorwhalen
thorwhalen deleted the fix-matrix-concurrent-reads branch September 22, 2026 15:39
thorwhalen added a commit that referenced this pull request Sep 22, 2026
On Windows a drive-letter key (store\C:\...) is not a valid path, so the
write raises OSError (as it did before #87) instead of landing inside the
root; assert that, and that the file the key names is untouched. #87's
Windows job failed on this and was missed because the Windows job does
not fail the workflow.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CorpusStore.matrix() raises KeyError/EOFError when read while another process is writing records

1 participant