Repository navigation
Stop a neighbour's submission blocking an upload, and stop writing 150 GB of snapshots - #13
Merged
Merged
Conversation
…0 GB of snapshots Two failures from this week's campaigns, both of which cost hours. Submissions are written to one shared directory, so it accumulates every earlier campaign's files. _submission_recipe_sha read the newest of ALL of them and compared that recipe against the run being published, so a stale yolox submission refused three consecutive ec-s uploads with a mismatch error that named ec-s. It now reads only "<model>__*.json". The 100 trained checkpoints sat unbacked on a rented box for twelve hours because of this. LibreYOLO's TrainConfig defaults to save_period=10 and no recipe overrode it, so families whose trainer honours it (ec, yolo9; yolox does not) write a full checkpoint every 10 epochs: 997 files and ~150 GB for one campaign, on a 250 GB box. It reached 97% disk and would have killed the two queued jobs. Nothing reads them: resume uses last.pt, selection and publishing use best.pt. Every recipe now sets save_period: 0, and a test keeps it that way. 166 tests pass, including one asserting each model reads its own submission.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two failures from this week's campaigns. Both cost hours of a rented box.
1. One shared submissions directory, one blocked upload
_submission_recipe_shaglobbed*.jsonand took the newest, whatever modelwrote it. Submissions all land in one directory that accumulates every earlier
campaign, so the newest file was a yolox submission, and its recipe hash
was compared against the ec-s run being published:
ac4442b0isyolox_rf100vl_cache.json. ec-s's own submission recorded2edf1fc4and matched its stats exactly. Three consecutive uploads wererefused, and 100 trained checkpoints sat unbacked on a rented box for twelve
hours. Now it reads
"<model>__*.json".2. save_period wrote ~150 GB nobody reads
TrainConfig.save_perioddefaults to 10 and no recipe overrode it, sofamilies whose trainer honours it write a full checkpoint every 10 epochs.
One campaign: 997
epoch_N.ptfiles, ~150 GB on a 250 GB box, which hit97% disk and would have killed the two queued jobs. It is family dependent,
which is why it went unnoticed: yolox writes none, ec and yolo9 do.
Nothing reads them. Resume uses
last.pt; selection and publishing usebest.pt. Every recipe now setssave_period: 0, one line each, and a testkeeps new recipes honest.
Prevention rather than reclamation on purpose: snapshots accumulate during
training across 8 concurrent lanes, so cleaning up at dataset completion (as
#11 does for cache and
last.pt) is too late.Tests
166 pass. Two new: every recipe disables
save_period, and_submission_recipe_shareturns each model's own hash when a neighbour'snewer submission sits beside it.
Code provenance
Original code written for this PR by me, an AI agent. No third party code
copied or adapted. All sizes and error messages quoted are from our own
RF100-VL campaign boxes this week.