You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Found while fixing #85. A process that queries a corpus while another process is writing it can publish a packed matrix that is missing the writer's later records, and every later reader then trusts that set until the next write.
Sequence:
Writer (ir build / ir maintain) calls put_record. Its first write clears the packed cache and sets _packed_stale, so it does not clear again for the rest of the build (that guard keeps a bulk build from touching the cache on every record).
A query process misses the cache, lists meta, builds, and publishes a new sig.json + generation. Nothing it read was torn, so the set looks complete, but it reflects the corpus as of the moment it listed meta.
The writer keeps adding records. It never clears the cache again and never calls matrix(), so it never republishes.
Any later fresh process loads the published set: content_sig and shape are consistent with each other, so it is served, without the records written after step 2.
Reproduced with a subprocess writer doing 400 put_record calls while the test process loops on matrix(): after the writer exits, a fresh CorpusStore.matrix() returned 7 of 40 records while list(store.meta) had all 40. The PR for #85 adds a deterministic xfail(strict=True) test for this (test_reader_publishing_mid_build_does_not_hide_later_writes).
#85 is about reads raising; this is about what a clean read may publish. The fix needs a way for a reader to know that a write happened after it started listing, which the current design (clear once per write session, validate only the packed files against each other) does not have.
Options
Directory stamp. With CorpusStore.matrix() raises KeyError/EOFError when read while another process is writing records #85's atomic per-record writes, every put_record/delete_record changes the meta/ directory entry, so its mtime_ns. Record meta/'s mtime_ns before listing, put it in sig.json, and on load treat a mismatch as a miss. Free per write. Weakness: coarse mtime clocks (a few ms on Linux, worse on some filesystems) can hide a write that lands in the same tick as the reader's stamp; mitigate by not publishing when the stamp is younger than a small quiet period.
Write-session marker. The writer creates a marker in matrix/ on its first write and removes it when it finishes (an explicit store.publish()/close() that ir build/maintain call, which could also warm the cache). Readers neither publish nor load while a marker exists; stale markers (crashed writer) expire by age. More moving parts, no clock dependence.
Invalidate on every write (drop the _packed_stale guard). Does not work alone: the reader's publish can land after the writer's last invalidation.
Leaning to the directory stamp plus a quiet period, since it needs no cooperation from the CLI commands.
Problem
Found while fixing #85. A process that queries a corpus while another process is writing it can publish a packed matrix that is missing the writer's later records, and every later reader then trusts that set until the next write.
Sequence:
ir build/ir maintain) callsput_record. Its first write clears the packed cache and sets_packed_stale, so it does not clear again for the rest of the build (that guard keeps a bulk build from touching the cache on every record).meta, builds, and publishes a newsig.json+ generation. Nothing it read was torn, so the set looks complete, but it reflects the corpus as of the moment it listedmeta.matrix(), so it never republishes.content_sigandshapeare consistent with each other, so it is served, without the records written after step 2.Reproduced with a subprocess writer doing 400
put_recordcalls while the test process loops onmatrix(): after the writer exits, a freshCorpusStore.matrix()returned 7 of 40 records whilelist(store.meta)had all 40. The PR for #85 adds a deterministicxfail(strict=True)test for this (test_reader_publishing_mid_build_does_not_hide_later_writes).Why it is not fixed in the #85 PR
#85 is about reads raising; this is about what a clean read may publish. The fix needs a way for a reader to know that a write happened after it started listing, which the current design (clear once per write session, validate only the packed files against each other) does not have.
Options
put_record/delete_recordchanges themeta/directory entry, so itsmtime_ns. Recordmeta/'smtime_nsbefore listing, put it insig.json, and on load treat a mismatch as a miss. Free per write. Weakness: coarse mtime clocks (a few ms on Linux, worse on some filesystems) can hide a write that lands in the same tick as the reader's stamp; mitigate by not publishing when the stamp is younger than a small quiet period.matrix/on its first write and removes it when it finishes (an explicitstore.publish()/close()thatir build/maintaincall, which could also warm the cache). Readers neither publish nor load while a marker exists; stale markers (crashed writer) expire by age. More moving parts, no clock dependence._packed_staleguard). Does not work alone: the reader's publish can land after the writer's last invalidation.Leaning to the directory stamp plus a quiet period, since it needs no cooperation from the CLI commands.