Reclaim a dataset's cache and resume checkpoint once it finishes (re-target) - #12
Merged
Merged
Conversation
Nothing removed either, so both grew for the length of a campaign. On a 250 GB box the post-resize cache reached 157 GB, filled the disk, and deadlocked every worker for six hours; with 16 concurrent lanes it reaches 71 GB within two epochs because longest-first schedules the biggest datasets together. last.pt is 635 MB for a 51M-parameter model, 63 GB across 100 datasets. Neither is read again once a dataset is done. The cache is consumed only by the epochs of the dataset that wrote it, and last.pt exists to resume an interrupted run. best.pt and the copy at the weights root are what the uploader ships and what the skip logic reads, so those stay. Reclaim runs in the worker after stats and the checkpoint are written, never before, so a dataset that fails to record itself as done keeps its resume point. Bytes freed are reported in the worker result. --keep-cache opts out. Operators have been working around this with an external janitor. That is fragile: the cache filename is a LibreYOLO implementation detail, and when a version bump changed it from "*.r640x640.npy" to "<image>.jpg.npy" the janitor silently matched nothing, kept reporting success, and the disk filled anyway. The code that writes the cache should own deleting it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Same change as #11, re-targeted at
rf100vl-harness.#11 was stacked on
ec-fp16-precisionand merged at 20:23, but #10 hadalready merged that branch into
rf100vl-harnessat 17:14. So the fp16 worklanded and this did not:
rf100vl-harnesscurrently hasPRECISION_AMP_DTYPEbut zero occurrences of
reclaim_finished_dataset. This is a cherry-pick ofd795433 onto the harness tip, no content changes.
This is the fix for the failure that cost a campaign six hours this week: the
post-resize cache reached 157 GB on a 250 GB box, filled the disk, and left
54 workers alive but blocked on writes with the GPUs at 0%. Worth not losing
in a merge-order accident.
What it does
When a dataset finishes, drop the two things nothing reads again:
.npycache (up to 10 GB per dataset, 157 GB observedacross 100, 71 GB for the 16 concurrently active ones)
last.pt, the resume point (635 MB at 51M params, 63 GB across a campaign)best.ptand the root copy stay: those are whatsync-artifactsships andwhat the skip logic reads. Reclaim runs after stats and the checkpoint are
written, so a dataset that dies before recording itself as done keeps its
resume point.
--keep-cacheopts out.164 tests pass, including one asserting the images, annotations and
best.ptsurvive.
Code provenance
Cherry-pick of d795433, already reviewed and merged as #11. Original code
written by me, an AI agent; no third party code copied or adapted.