Skip to content

Stop a neighbour's submission blocking an upload, and stop writing 150 GB of snapshots - #13

Merged
EHxuban11 merged 1 commit into
rf100vl-harnessfrom
harness-upload-and-snapshots
Aug 12, 2026
Merged

EHxuban11 merged 1 commit into
rf100vl-harnessfrom
harness-upload-and-snapshots

Conversation

@EHxuban11

Copy link
Copy Markdown
Contributor

Two failures from this week's campaigns. Both cost hours of a rented box.

1. One shared submissions directory, one blocked upload

_submission_recipe_sha globbed *.json and took the newest, whatever model
wrote it. Submissions all land in one directory that accumulates every earlier
campaign, so the newest file was a yolox submission, and its recipe hash
was compared against the ec-s run being published:

IncompletePublish: Recipe mismatch for ec-s: ec-b16-nocache.json hashes to
2edf1fc4 but the submission records ac4442b0

ac4442b0 is yolox_rf100vl_cache.json. ec-s's own submission recorded
2edf1fc4 and matched its stats exactly. Three consecutive uploads were
refused, and 100 trained checkpoints sat unbacked on a rented box for twelve
hours. Now it reads "<model>__*.json".

2. save_period wrote ~150 GB nobody reads

TrainConfig.save_period defaults to 10 and no recipe overrode it, so
families whose trainer honours it write a full checkpoint every 10 epochs.
One campaign: 997 epoch_N.pt files, ~150 GB on a 250 GB box, which hit
97% disk and would have killed the two queued jobs. It is family dependent,
which is why it went unnoticed: yolox writes none, ec and yolo9 do.

Nothing reads them. Resume uses last.pt; selection and publishing use
best.pt. Every recipe now sets save_period: 0, one line each, and a test
keeps new recipes honest.

Prevention rather than reclamation on purpose: snapshots accumulate during
training across 8 concurrent lanes, so cleaning up at dataset completion (as
#11 does for cache and last.pt) is too late.

Tests

166 pass. Two new: every recipe disables save_period, and
_submission_recipe_sha returns each model's own hash when a neighbour's
newer submission sits beside it.

Code provenance

Original code written for this PR by me, an AI agent. No third party code
copied or adapted. All sizes and error messages quoted are from our own
RF100-VL campaign boxes this week.

…0 GB of snapshots

Two failures from this week's campaigns, both of which cost hours.

Submissions are written to one shared directory, so it accumulates every
earlier campaign's files. _submission_recipe_sha read the newest of ALL of
them and compared that recipe against the run being published, so a stale
yolox submission refused three consecutive ec-s uploads with a mismatch error
that named ec-s. It now reads only "<model>__*.json". The 100 trained
checkpoints sat unbacked on a rented box for twelve hours because of this.

LibreYOLO's TrainConfig defaults to save_period=10 and no recipe overrode it,
so families whose trainer honours it (ec, yolo9; yolox does not) write a full
checkpoint every 10 epochs: 997 files and ~150 GB for one campaign, on a
250 GB box. It reached 97% disk and would have killed the two queued jobs.
Nothing reads them: resume uses last.pt, selection and publishing use best.pt.
Every recipe now sets save_period: 0, and a test keeps it that way.

166 tests pass, including one asserting each model reads its own submission.
@EHxuban11
EHxuban11 merged commit b7cf016 into rf100vl-harness Aug 12, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant