Skip to content

DDP lanes: train one dataset across several GPUs (--gpus-per-job) - #15

Merged
EHxuban11 merged 1 commit into
rf100vl-harnessfrom
rf100vl-ddp-lanes
Aug 15, 2026
Merged

EHxuban11 merged 1 commit into
rf100vl-harnessfrom
rf100vl-ddp-lanes

Conversation

@EHxuban11

Copy link
Copy Markdown
Contributor

What

  • New --gpus-per-job N on rf100vl-train and rf100vl-campaign. Groups --gpus into lanes of N cards; each dataset trains on a whole lane via LibreYOLO's ddp_aware model.train(device="0,1,...") auto-spawn.
  • Global batch untouched. LibreYOLO splits batch // world_size per rank and derives accumulation from nbs against the global batch. Effective batch 16 holds.
  • Lane width in the run signature ONLY when > 1. Single-GPU signatures hash exactly as before; banked checkpoints and in-flight campaigns unaffected.
  • Batch plan records ddp_world_size and per_gpu_physical_batch. Plan errors loudly if the batch does not split evenly.
  • Orchestrator refuses --gpus-per-job > 1 on a LibreYOLO without libreyolo.training.ddp_spawn. Capability probe reports ddp.
  • Solo OOM drain retries at full lane width, never a single GPU (a narrower retry would change the signature and refuse the resume).
  • Timeout/interrupt kill the child's process tree. mp.spawn ranks are non-daemonic grandchildren and used to survive a killed worker holding VRAM.
  • gpu_trace maps a DDP dataset to every card in its lane instead of losing it on int("4,5").
  • Bugfix: rf100vl-train read args.keep_cache but never defined --keep-cache (AttributeError). rf100vl-campaign defined it but never forwarded it (flag was a no-op).
  • Runbook: new "Multi-GPU lanes" section with the rfdetr-m/l numbers and caveats.

Why

rfdetr-m and rfdetr-l need ~31 GB at protocol batch 16; batch 8 already OOMs a 16 GB card (measured 2026-08). The harness was one-dataset-one-GPU everywhere, so the only options were a 40+ GB rental or the grad-accum deviation the recipe just moved away from. 4x16 GB at batch 4/rank (~7 GB each) runs the protocol batch on hardware already rented. RF-DETR multi-scale draws its per-step scale from random.Random(step), so all ranks resize identically; DDP keeps one scale per optimizer step over 16 images, closer to the reference semantics than grad-accum's four scales per step.

Not bit-identical to single GPU (sampler sharding, per-rank aug seeds), hence the signature entry.

rfdetr-l still carries the physical_batch: 2 size override from the 16 GB era. Running l at batch 16 over a lane needs a recipe change and a fresh campaign; separate decision.

Validation

  • Full suite: 175 passed (was 167). New tests: lane parsing, batch split validation, signature back-compat (legacy plan without the new keys hashes identically), device string, lane grouping, full-width OOM drain, mutual exclusion, DDP-less refusal.
  • Not yet run on a real multi-GPU box. Needs a 2-epoch smoke of rfdetr-m with --gpus-per-job before a paid campaign.

Code provenance

All code written from scratch for this PR against the existing harness. No external code copied.

The orchestrator was one-dataset-one-GPU everywhere, so a model whose
protocol batch does not fit a single card (rfdetr-m/l need ~31 GB at
batch 16) could only run on bigger rented hardware. A lane can now span
N GPUs: --gpus is grouped into lanes, the worker gets the whole group in
CUDA_VISIBLE_DEVICES, and LibreYOLO's ddp_aware train() splits the
recipe's GLOBAL batch across ranks, so effective batch 16 is unchanged.

- lane width validated against the batch (must split evenly) and against
  the installed LibreYOLO (refuses a build without ddp_spawn)
- width recorded in the batch plan and stats; part of the run signature
  ONLY when >1, so every existing single-GPU signature and banked
  checkpoint hashes exactly as before
- solo OOM drain retries at full lane width, never one GPU
- timeout/interrupt now reap the child's process tree: mp.spawn ranks
  are non-daemonic grandchildren and used to outlive a killed worker
- gpu_trace attributes a DDP dataset to every card in its lane instead
  of dropping it on int("4,5")
- drive-by: rf100vl-train was reading args.keep_cache without defining
  --keep-cache, and rf100vl-campaign defined it but never forwarded it
@EHxuban11
EHxuban11 merged commit fd6c8d8 into rf100vl-harness Aug 15, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant