Skip to content

Reclaim a dataset's cache and resume checkpoint once it finishes - #11

Merged
EHxuban11 merged 1 commit into
ec-fp16-precisionfrom
cache-purge-on-done
Aug 10, 2026
Merged

EHxuban11 merged 1 commit into
ec-fp16-precisionfrom
cache-purge-on-done

Conversation

@EHxuban11

Copy link
Copy Markdown
Contributor

Stacked on #10. Review that one first; this diff is the three files below.

The problem, measured

Nothing removes a dataset's artifacts when it finishes, so they accumulate for
the length of a campaign on a 250 GB box:

artifact size across 100 datasets
post-resize .npy cache ~10 GB for the largest dataset 157 GB observed
last.pt 635 MB at 51M params 63 GB

The cache filled the disk this week and deadlocked every worker for six
hours
: 54 processes alive, GPUs holding memory at 0% utilization, blocked on
writes. With 16 concurrent lanes it reaches 71 GB within two epochs, because
longest-first deliberately schedules the biggest datasets together.

Why these two and not the rest

Neither is read again once a dataset is done. The cache is consumed only by
the epochs of the dataset that wrote it; last.pt exists to resume an
interrupted dataset. best.pt and the root copy stay, because they are what
sync-artifacts ships and what the skip logic reads. The test asserts that.

Ordering matters: reclaim runs after stats and the checkpoint are written, so
a dataset that dies before recording itself as done keeps its resume point.

--keep-cache opts out for debugging.

Why it belongs in the harness

We have been working around it with an external janitor, and that failed in
the worst way. The cache filename is a LibreYOLO implementation detail; when a
version bump changed it from *.r640x640.npy to <image>.jpg.npy, the
janitor matched nothing, logged cheerful "purged" lines from stale files, and
the disk filled anyway. The code that writes the cache should own deleting it.

Tests

164 passed. Two new: one asserting the cache and last.pt go while best.pt,
the images and the annotations survive, and one for --keep-cache.

Code provenance

Original code written for this PR by me, an AI agent. No third party code was
copied or adapted. The sizes quoted above are measurements from our own
RF100-VL campaign boxes this week.

Nothing removed either, so both grew for the length of a campaign. On a
250 GB box the post-resize cache reached 157 GB, filled the disk, and
deadlocked every worker for six hours; with 16 concurrent lanes it reaches
71 GB within two epochs because longest-first schedules the biggest datasets
together. last.pt is 635 MB for a 51M-parameter model, 63 GB across 100
datasets.

Neither is read again once a dataset is done. The cache is consumed only by
the epochs of the dataset that wrote it, and last.pt exists to resume an
interrupted run. best.pt and the copy at the weights root are what the
uploader ships and what the skip logic reads, so those stay.

Reclaim runs in the worker after stats and the checkpoint are written, never
before, so a dataset that fails to record itself as done keeps its resume
point. Bytes freed are reported in the worker result. --keep-cache opts out.

Operators have been working around this with an external janitor. That is
fragile: the cache filename is a LibreYOLO implementation detail, and when a
version bump changed it from "*.r640x640.npy" to "<image>.jpg.npy" the
janitor silently matched nothing, kept reporting success, and the disk filled
anyway. The code that writes the cache should own deleting it.
@EHxuban11
EHxuban11 merged commit d795433 into ec-fp16-precision Aug 10, 2026
1 check passed
EHxuban11 added a commit that referenced this pull request Aug 10, 2026
#12)

Nothing removed either, so both grew for the length of a campaign. On a
250 GB box the post-resize cache reached 157 GB, filled the disk, and
deadlocked every worker for six hours; with 16 concurrent lanes it reaches
71 GB within two epochs because longest-first schedules the biggest datasets
together. last.pt is 635 MB for a 51M-parameter model, 63 GB across 100
datasets.

Neither is read again once a dataset is done. The cache is consumed only by
the epochs of the dataset that wrote it, and last.pt exists to resume an
interrupted run. best.pt and the copy at the weights root are what the
uploader ships and what the skip logic reads, so those stay.

Reclaim runs in the worker after stats and the checkpoint are written, never
before, so a dataset that fails to record itself as done keeps its resume
point. Bytes freed are reported in the worker result. --keep-cache opts out.

Operators have been working around this with an external janitor. That is
fragile: the cache filename is a LibreYOLO implementation detail, and when a
version bump changed it from "*.r640x640.npy" to "<image>.jpg.npy" the
janitor silently matched nothing, kept reporting success, and the disk filled
anyway. The code that writes the cache should own deleting it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant