Skip to content

Reclaim a dataset's cache and resume checkpoint once it finishes (re-target) - #12

Merged
EHxuban11 merged 1 commit into
rf100vl-harnessfrom
cache-purge-to-harness
Aug 10, 2026
Merged

EHxuban11 merged 1 commit into
rf100vl-harnessfrom
cache-purge-to-harness

Conversation

@EHxuban11

Copy link
Copy Markdown
Contributor

Same change as #11, re-targeted at rf100vl-harness.

#11 was stacked on ec-fp16-precision and merged at 20:23, but #10 had
already merged that branch into rf100vl-harness at 17:14. So the fp16 work
landed and this did not: rf100vl-harness currently has PRECISION_AMP_DTYPE
but zero occurrences of reclaim_finished_dataset. This is a cherry-pick of
d795433 onto the harness tip, no content changes.

This is the fix for the failure that cost a campaign six hours this week: the
post-resize cache reached 157 GB on a 250 GB box, filled the disk, and left
54 workers alive but blocked on writes with the GPUs at 0%. Worth not losing
in a merge-order accident.

What it does

When a dataset finishes, drop the two things nothing reads again:

  • its post-resize .npy cache (up to 10 GB per dataset, 157 GB observed
    across 100, 71 GB for the 16 concurrently active ones)
  • last.pt, the resume point (635 MB at 51M params, 63 GB across a campaign)

best.pt and the root copy stay: those are what sync-artifacts ships and
what the skip logic reads. Reclaim runs after stats and the checkpoint are
written, so a dataset that dies before recording itself as done keeps its
resume point. --keep-cache opts out.

164 tests pass, including one asserting the images, annotations and best.pt
survive.

Code provenance

Cherry-pick of d795433, already reviewed and merged as #11. Original code
written by me, an AI agent; no third party code copied or adapted.

Nothing removed either, so both grew for the length of a campaign. On a
250 GB box the post-resize cache reached 157 GB, filled the disk, and
deadlocked every worker for six hours; with 16 concurrent lanes it reaches
71 GB within two epochs because longest-first schedules the biggest datasets
together. last.pt is 635 MB for a 51M-parameter model, 63 GB across 100
datasets.

Neither is read again once a dataset is done. The cache is consumed only by
the epochs of the dataset that wrote it, and last.pt exists to resume an
interrupted run. best.pt and the copy at the weights root are what the
uploader ships and what the skip logic reads, so those stay.

Reclaim runs in the worker after stats and the checkpoint are written, never
before, so a dataset that fails to record itself as done keeps its resume
point. Bytes freed are reported in the worker result. --keep-cache opts out.

Operators have been working around this with an external janitor. That is
fragile: the cache filename is a LibreYOLO implementation detail, and when a
version bump changed it from "*.r640x640.npy" to "<image>.jpg.npy" the
janitor silently matched nothing, kept reporting success, and the disk filled
anyway. The code that writes the cache should own deleting it.
@EHxuban11
EHxuban11 merged commit 981d381 into rf100vl-harness Aug 10, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant