Skip to content

Guard the silent failures found in the RF100-VL yolox campaigns - #7

Merged
EHxuban11 merged 1 commit into
rf100vl-harnessfrom
fix/campaign-silent-failure-guards
Aug 3, 2026
Merged

EHxuban11 merged 1 commit into
rf100vl-harnessfrom
fix/campaign-silent-failure-guards

Conversation

@EHxuban11

Copy link
Copy Markdown
Contributor

Harness half of the 2026-08 campaign postmortem. Library half is LibreYOLO/libreyolo#710.

Every problem in those campaigns was silent. Crash handling is genuinely good (~300 dataset-runs, 0 failures, dataset+epoch resume, OOM to grad-accum). What was missing was anything that notices a wrong result.

What

  • ETA lane count (rf100vl_dash.py). lanes = len(gpus_active) counted GPUs, not workers, so an 8-GPU box at --jobs-per-gpu 2 divided remaining work by 8 instead of 16 and reported roughly double. Now counts running datasets.
  • -1 sentinel exclusion (_aggregate_metrics). pycocotools returns -1 for an area range with no ground truth. Averaging it in published a yolox-nano mAP_small of -0.0535, a negative average precision.
  • Selection-vs-test divergence warning (new). Compares each dataset's valid_mAP50_95 from stats.json against its reported test AP, warns above 0.15 AP, records rf100vl.selection_test_divergence.

Why the third one matters most

yolox-nano selected on valid 0.5663 and published test 0.1620 on ball, while yolox-tiny agreed to 0.002 on the same datasets. Cause was BatchNorm eps reverting on a class-count rebuild (LibreYOLO#700): trained at 1e-5, evaluated at 1e-3, ruinous only for the depthwise variant. A full 100-dataset campaign published 0.3601 where the truth was 0.4853, and the harness said nothing.

A 16-point gap between adjacent sizes of one family is not a result, it is a bug report. Adjacent sizes land ~3 points apart.

Check

tests/test_campaign_guards.py (new, 7 tests): sentinel excluded from the mean, published negative mAP_small cannot recur, metric defined nowhere stays -1, the real yolox-nano numbers are flagged, the real yolox-tiny numbers are not, missing stats and absent weights-root are non-errors. All passing.

Not verified

  • Warn-only by design. A genuinely hard dataset can drift, and a run that refuses to emit results is worse than one that emits them loudly caveated. The 0.15 threshold is a judgement call from two campaigns, not a tuned value.
  • The ETA fix undercounts lanes once the queue drains and fewer datasets run than the pool size. Little work remains by then, but it is not exact.
  • Not fixed here, and worth a follow-up: sync-artifacts reports uploaded N, skipped 0 while silently shipping nothing outside the paths it knows. It dropped every checkpoint and a provenance note from a corrected run because the weights root was flat rather than .runs-shaped. It also never uploads the recipe JSON, so a run whose recipe_sha256 points at a box-local file is unreproducible once that box is gone.

Fix the ETA lane count. lanes counted distinct GPUs, so an 8-GPU box at
--jobs-per-gpu 2 divided remaining work by 8 instead of 16 and reported
roughly double the true remaining time. Count running datasets: while
anything is pending the worker pool is saturated, so that is the pool
size.

Stop averaging pycocotools -1 sentinels into the size-stratified means.
-1 means "no ground truth in this area range", not "scored zero", and
averaging it in published a yolox-nano mAP_small of -0.0535: a negative
average precision. Each mean now covers only the datasets where the
metric is defined; a metric defined nowhere stays -1 to keep saying so.

Warn when the training-time selection metric contradicts the reported
test score. Selection runs inside the trainer against the in-memory
model; the reported score reloads a checkpoint and evaluates it
standalone. When those disagree by a lot, the two paths disagree about
the model. In 2026-08 that was BatchNorm eps reverting on a class-count
rebuild (libreyolo #700): yolox-nano selected on valid 0.5663, published
test 0.1620, and a full 100-dataset campaign shipped a headline 0.3601
that should have been 0.4853. Nothing flagged it. Warns and records
rf100vl.selection_test_divergence; never fails the run.

Claude-Session: https://claude.ai/code/session_01D8uB5n3e1zYjix18gT2Y4w
@EHxuban11

Copy link
Copy Markdown
Contributor Author

Reviewed (Cursor agent, on Xuban's behalf). LGTM — merging.

  • Lanes fix is correct: len(running) is the pool size while anything is pending, and the residual undercount once the queue drains is honestly documented and immaterial.
  • Sentinel exclusion is right: >= 0.0 keeps genuine zero scores, exclusion is per-metric, and an all-sentinel metric stays -1 instead of fabricating a number. The published negative mAP_small cannot recur, and the test asserts exactly that.
  • Divergence gate is the highest-value change: warn-only is the correct posture for a benchmark harness, the threshold is documented as a judgement call, the result is recorded as selection_test_divergence in the output JSON, and the tests pin the real campaign numbers (nano's 0.5663/0.1620 flagged, tiny's 0.6198/0.6091 quiet). Verified the helper's imports (load_json, Path, os, json, Any) all resolve in va_bench/rf100vl.py.

Agreed that sync-artifacts (skip-report vs explicit manifest) stays a follow-up — it needs a design call and shouldn't ride in a bug-fix PR. Given it nearly published an unreproducible correction today, it's the top remaining harness item.

CI pytest is green.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant