Repository navigation
Guard the silent failures found in the RF100-VL yolox campaigns - #7
Merged
Merged
Conversation
Fix the ETA lane count. lanes counted distinct GPUs, so an 8-GPU box at --jobs-per-gpu 2 divided remaining work by 8 instead of 16 and reported roughly double the true remaining time. Count running datasets: while anything is pending the worker pool is saturated, so that is the pool size. Stop averaging pycocotools -1 sentinels into the size-stratified means. -1 means "no ground truth in this area range", not "scored zero", and averaging it in published a yolox-nano mAP_small of -0.0535: a negative average precision. Each mean now covers only the datasets where the metric is defined; a metric defined nowhere stays -1 to keep saying so. Warn when the training-time selection metric contradicts the reported test score. Selection runs inside the trainer against the in-memory model; the reported score reloads a checkpoint and evaluates it standalone. When those disagree by a lot, the two paths disagree about the model. In 2026-08 that was BatchNorm eps reverting on a class-count rebuild (libreyolo #700): yolox-nano selected on valid 0.5663, published test 0.1620, and a full 100-dataset campaign shipped a headline 0.3601 that should have been 0.4853. Nothing flagged it. Warns and records rf100vl.selection_test_divergence; never fails the run. Claude-Session: https://claude.ai/code/session_01D8uB5n3e1zYjix18gT2Y4w
Contributor
Author
|
Reviewed (Cursor agent, on Xuban's behalf). LGTM — merging.
Agreed that CI pytest is green. |
This was referenced Aug 3, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Harness half of the 2026-08 campaign postmortem. Library half is LibreYOLO/libreyolo#710.
Every problem in those campaigns was silent. Crash handling is genuinely good (~300 dataset-runs, 0 failures, dataset+epoch resume, OOM to grad-accum). What was missing was anything that notices a wrong result.
What
rf100vl_dash.py).lanes = len(gpus_active)counted GPUs, not workers, so an 8-GPU box at--jobs-per-gpu 2divided remaining work by 8 instead of 16 and reported roughly double. Now counts running datasets._aggregate_metrics). pycocotools returns -1 for an area range with no ground truth. Averaging it in published a yolox-nanomAP_smallof -0.0535, a negative average precision.valid_mAP50_95fromstats.jsonagainst its reported test AP, warns above 0.15 AP, recordsrf100vl.selection_test_divergence.Why the third one matters most
yolox-nano selected on valid 0.5663 and published test 0.1620 on
ball, while yolox-tiny agreed to 0.002 on the same datasets. Cause was BatchNorm eps reverting on a class-count rebuild (LibreYOLO#700): trained at 1e-5, evaluated at 1e-3, ruinous only for the depthwise variant. A full 100-dataset campaign published 0.3601 where the truth was 0.4853, and the harness said nothing.A 16-point gap between adjacent sizes of one family is not a result, it is a bug report. Adjacent sizes land ~3 points apart.
Check
tests/test_campaign_guards.py(new, 7 tests): sentinel excluded from the mean, published negative mAP_small cannot recur, metric defined nowhere stays -1, the real yolox-nano numbers are flagged, the real yolox-tiny numbers are not, missing stats and absent weights-root are non-errors. All passing.Not verified
sync-artifactsreportsuploaded N, skipped 0while silently shipping nothing outside the paths it knows. It dropped every checkpoint and a provenance note from a corrected run because the weights root was flat rather than.runs-shaped. It also never uploads the recipe JSON, so a run whoserecipe_sha256points at a box-local file is unreproducible once that box is gone.