Repository navigation
DDP lanes: train one dataset across several GPUs (--gpus-per-job) - #15
Merged
Merged
Conversation
The orchestrator was one-dataset-one-GPU everywhere, so a model whose
protocol batch does not fit a single card (rfdetr-m/l need ~31 GB at
batch 16) could only run on bigger rented hardware. A lane can now span
N GPUs: --gpus is grouped into lanes, the worker gets the whole group in
CUDA_VISIBLE_DEVICES, and LibreYOLO's ddp_aware train() splits the
recipe's GLOBAL batch across ranks, so effective batch 16 is unchanged.
- lane width validated against the batch (must split evenly) and against
the installed LibreYOLO (refuses a build without ddp_spawn)
- width recorded in the batch plan and stats; part of the run signature
ONLY when >1, so every existing single-GPU signature and banked
checkpoint hashes exactly as before
- solo OOM drain retries at full lane width, never one GPU
- timeout/interrupt now reap the child's process tree: mp.spawn ranks
are non-daemonic grandchildren and used to outlive a killed worker
- gpu_trace attributes a DDP dataset to every card in its lane instead
of dropping it on int("4,5")
- drive-by: rf100vl-train was reading args.keep_cache without defining
--keep-cache, and rf100vl-campaign defined it but never forwarded it
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
--gpus-per-job Nonrf100vl-trainandrf100vl-campaign. Groups--gpusinto lanes of N cards; each dataset trains on a whole lane via LibreYOLO'sddp_awaremodel.train(device="0,1,...")auto-spawn.batch // world_sizeper rank and derives accumulation fromnbsagainst the global batch. Effective batch 16 holds.ddp_world_sizeandper_gpu_physical_batch. Plan errors loudly if the batch does not split evenly.--gpus-per-job > 1on a LibreYOLO withoutlibreyolo.training.ddp_spawn. Capability probe reportsddp.gpu_tracemaps a DDP dataset to every card in its lane instead of losing it onint("4,5").rf100vl-trainreadargs.keep_cachebut never defined--keep-cache(AttributeError).rf100vl-campaigndefined it but never forwarded it (flag was a no-op).Why
rfdetr-m and rfdetr-l need ~31 GB at protocol batch 16; batch 8 already OOMs a 16 GB card (measured 2026-08). The harness was one-dataset-one-GPU everywhere, so the only options were a 40+ GB rental or the grad-accum deviation the recipe just moved away from. 4x16 GB at batch 4/rank (~7 GB each) runs the protocol batch on hardware already rented. RF-DETR multi-scale draws its per-step scale from
random.Random(step), so all ranks resize identically; DDP keeps one scale per optimizer step over 16 images, closer to the reference semantics than grad-accum's four scales per step.Not bit-identical to single GPU (sampler sharding, per-rank aug seeds), hence the signature entry.
rfdetr-lstill carries thephysical_batch: 2size override from the 16 GB era. Running l at batch 16 over a lane needs a recipe change and a fresh campaign; separate decision.Validation
--gpus-per-jobbefore a paid campaign.Code provenance
All code written from scratch for this PR against the existing harness. No external code copied.