Repository navigation
Reclaim a dataset's cache and resume checkpoint once it finishes - #11
Merged
Merged
Conversation
Nothing removed either, so both grew for the length of a campaign. On a 250 GB box the post-resize cache reached 157 GB, filled the disk, and deadlocked every worker for six hours; with 16 concurrent lanes it reaches 71 GB within two epochs because longest-first schedules the biggest datasets together. last.pt is 635 MB for a 51M-parameter model, 63 GB across 100 datasets. Neither is read again once a dataset is done. The cache is consumed only by the epochs of the dataset that wrote it, and last.pt exists to resume an interrupted run. best.pt and the copy at the weights root are what the uploader ships and what the skip logic reads, so those stay. Reclaim runs in the worker after stats and the checkpoint are written, never before, so a dataset that fails to record itself as done keeps its resume point. Bytes freed are reported in the worker result. --keep-cache opts out. Operators have been working around this with an external janitor. That is fragile: the cache filename is a LibreYOLO implementation detail, and when a version bump changed it from "*.r640x640.npy" to "<image>.jpg.npy" the janitor silently matched nothing, kept reporting success, and the disk filled anyway. The code that writes the cache should own deleting it.
EHxuban11
added a commit
that referenced
this pull request
Aug 10, 2026
#12) Nothing removed either, so both grew for the length of a campaign. On a 250 GB box the post-resize cache reached 157 GB, filled the disk, and deadlocked every worker for six hours; with 16 concurrent lanes it reaches 71 GB within two epochs because longest-first schedules the biggest datasets together. last.pt is 635 MB for a 51M-parameter model, 63 GB across 100 datasets. Neither is read again once a dataset is done. The cache is consumed only by the epochs of the dataset that wrote it, and last.pt exists to resume an interrupted run. best.pt and the copy at the weights root are what the uploader ships and what the skip logic reads, so those stay. Reclaim runs in the worker after stats and the checkpoint are written, never before, so a dataset that fails to record itself as done keeps its resume point. Bytes freed are reported in the worker result. --keep-cache opts out. Operators have been working around this with an external janitor. That is fragile: the cache filename is a LibreYOLO implementation detail, and when a version bump changed it from "*.r640x640.npy" to "<image>.jpg.npy" the janitor silently matched nothing, kept reporting success, and the disk filled anyway. The code that writes the cache should own deleting it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #10. Review that one first; this diff is the three files below.
The problem, measured
Nothing removes a dataset's artifacts when it finishes, so they accumulate for
the length of a campaign on a 250 GB box:
.npycachelast.ptThe cache filled the disk this week and deadlocked every worker for six
hours: 54 processes alive, GPUs holding memory at 0% utilization, blocked on
writes. With 16 concurrent lanes it reaches 71 GB within two epochs, because
longest-first deliberately schedules the biggest datasets together.
Why these two and not the rest
Neither is read again once a dataset is done. The cache is consumed only by
the epochs of the dataset that wrote it;
last.ptexists to resume aninterrupted dataset.
best.ptand the root copy stay, because they are whatsync-artifactsships and what the skip logic reads. The test asserts that.Ordering matters: reclaim runs after stats and the checkpoint are written, so
a dataset that dies before recording itself as done keeps its resume point.
--keep-cacheopts out for debugging.Why it belongs in the harness
We have been working around it with an external janitor, and that failed in
the worst way. The cache filename is a LibreYOLO implementation detail; when a
version bump changed it from
*.r640x640.npyto<image>.jpg.npy, thejanitor matched nothing, logged cheerful "purged" lines from stale files, and
the disk filled anyway. The code that writes the cache should own deleting it.
Tests
164 passed. Two new: one asserting the cache and
last.ptgo whilebest.pt,the images and the annotations survive, and one for
--keep-cache.Code provenance
Original code written for this PR by me, an AI agent. No third party code was
copied or adapted. The sizes quoted above are measurements from our own
RF100-VL campaign boxes this week.