Problem
Found during the round-3 post-merge review of #84. This is not caused by #83/#84: those made the packed cache safe. The failure is one layer down, in the per-record meta/vectors stores that _build_matrix reads when the packed cache misses.
A writer clears the packed cache on its first write, so every reader that misses rebuilds from the per-record files while the writer is still adding them. Two races then escape matrix():
put_record writes meta[id] before vectors[id]. A reader that lists meta in between gets KeyError: [Errno 2] No such file or directory: '.../vectors/<id>'.
dol.Files writes are not atomic. A reader can read a half-written vector .npy (EOFError: No data left in file from np.load) or, presumably, a half-written meta JSON (JSONDecodeError).
Repro: 6 processes on one file-backed corpus, each doing 200 random put_record / matrix() calls. About 1 run in 3 has a reader that raises one of the two errors above. The stress script is small; it can be re-derived from this description.
The realistic trigger is the one #77 described: a query issued while ir maintain (scheduled) or ir build is writing to the same corpus.
Suggested fix (not applied, because it is outside the reviewed PR)
- In
put_record, write the vector before the meta, so a listed id always has a vector. delete_record already removes the meta first.
- In
_build_matrix and metas(), treat a record that disappears or cannot be decoded mid-read (KeyError, EOFError, ValueError) as not yet written: skip it, and do not publish a packed set built from a read that skipped anything.
- Longer term, write per-record files atomically (tmp +
os.replace) in the store factories.
Problem
Found during the round-3 post-merge review of #84. This is not caused by #83/#84: those made the packed cache safe. The failure is one layer down, in the per-record
meta/vectorsstores that_build_matrixreads when the packed cache misses.A writer clears the packed cache on its first write, so every reader that misses rebuilds from the per-record files while the writer is still adding them. Two races then escape
matrix():put_recordwritesmeta[id]beforevectors[id]. A reader that listsmetain between getsKeyError: [Errno 2] No such file or directory: '.../vectors/<id>'.dol.Fileswrites are not atomic. A reader can read a half-written vector.npy(EOFError: No data left in filefromnp.load) or, presumably, a half-written meta JSON (JSONDecodeError).Repro: 6 processes on one file-backed corpus, each doing 200 random
put_record/matrix()calls. About 1 run in 3 has a reader that raises one of the two errors above. The stress script is small; it can be re-derived from this description.The realistic trigger is the one #77 described: a query issued while
ir maintain(scheduled) orir buildis writing to the same corpus.Suggested fix (not applied, because it is outside the reviewed PR)
put_record, write the vector before the meta, so a listed id always has a vector.delete_recordalready removes the meta first._build_matrixandmetas(), treat a record that disappears or cannot be decoded mid-read (KeyError,EOFError,ValueError) as not yet written: skip it, and do not publish a packed set built from a read that skipped anything.os.replace) in the store factories.