Conversation
0b6f4d8 to
e124158
Compare
Section 6A: a pure-Python planner that consumes a calibrated per-block/per-phase profile plus a block inventory and returns an auditable resident-prefix / offloaded-suffix plan. No torch, no CUDA, no allocation, no checkpoint read. Reports predicted per-rank peak memory, per-node pinned host demand, exposed transfer stall and total phase cost, rejects every candidate that violates a budget or capability, refuses to treat full residency as a fallback, and refuses to plan without a measured phase. An executor with no measured residency bound is charged conservatively rather than with a two-slot model it has not verified. Refs Tencent-Hunyuan#483 (section 6A) Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
e124158 to
57cec42
Compare
H100 supplementary validation recordThe archived 2026-09-28 results have been checked against the raw files and published as test artifacts and repository-external harnesses on 2026-10-02. Pinned PR head:
Each rank compares 214 tensors after two identical AdamW steps: logits, loss, parameter gradients, clipping norm, updated parameters and both Adam moments. FP32 tolerance is atol=1e-6 / rtol=1e-5; BF16 is atol=5e-3 / rtol=1e-2. The evidence includes all six successful per-rank JSON files, process exit records, the NCCL-only failure logs, the standalone FSDP harness, the planner CPU log and SHA-256 checksums. Empty success logs are recorded separately in the manifest. These are standalone FSDP2 capability/correctness checks. P08, H100 planner calibration and steady-state performance remain NOT_RUN. The raw timings retain their diagnostic-only label (cold iterations and existing serving occupancy), so speedup and latency reduction are N/A. No new GPU run was performed for this publication. |
Incremental review: Files changed against
main.Summary
Section 6A of the trainside optimization roadmap: decide how much to offload from
measurements instead of a boolean. This PR adds the measurement contract, a pure planner and
the first real calibration numbers; it adds no new executor.
unirl/train/backend/fsdp/offload_plan.pyconsumes a measuredPhaseProfileplus aBlockInventoryand returns an auditablePlacementPlan:can be re-derived and diffed from a JSON profile on any host.
candidate must satisfy the tightest rank budget and the node-summed pinned-host budget.
There is no per-rank plan, because ranks sharing an FSDP shard group must agree on ownership.
PhaseCost.resident_gpu_bytesmeasures the phase's per-rankhigh-water excluding all block parameter storage; the planner then adds
root_bytes,resident block staging, streamed staging and headroom itself. That contract is what makes
the ledger check below exact instead of double-counted.
live_streamed_bound=None— thehonest default for native FSDP2
CPUOffloadPolicy— every streamed block is charged atpeak. The two-slot model applies only when the executor declares that bound.
stream, a group split across the boundary is rejected, an unmeasured phase rejects every
candidate instead of extrapolating, and full residency is never an automatic fallback: no
feasible candidate yields
INFEASIBLEwith per-candidate reasons and nothing enabled.to_json()carries the selected plan, the predicted per-rank peaks,per-node host demand,
transfer_busy_ms,no_compute_transfer_msand every candidate'spredictions and rejection reasons.
unirl/train/backend/fsdp/README.mdrecords the profile contract that a calibration run musthonour, including the two footguns a reviewer should check first:
compute_msmust excludetransfer waits, and only
bandwidth_record_kind="shard"may be summed withshard_bytes.Related Issue
Refs #483 (section 6A). This is the planner and measurement contract only; §6B's executor and
§6C's phase-aware streaming are explicitly not part of it.
Test Plan
Archived supplementary evidence (published 2026-10-02): Raw results, harnesses and SHA-256 checksums. Test execution date: 2026-09-28; no new GPU run was performed for this publication.
H100 follow-up — 2026-09-28
Rechecked PR head
57cec429b5978d7da4bfd5f751764c0fd5da4c06. New hardware: 2× NVIDIA H100 80GB HBM3, NVLink, driver 595.71.05, Python 3.12, PyTorch 2.13.0+cu130. The original environment was retained. H100 validation is accepted for this review, as agreed with the maintainer.The recovered planner CPU harness passes P01–P07 on this host; P08 still needs the §6B executor. A separate two-rank FSDP2 capability harness compares identical initialization, input and two AdamW updates against full residency:
cpu:gloo,cuda:ncclEach comparison includes logits, loss, every parameter gradient, clipping norm, updated parameters and both Adam moments. FP32 tolerances: atol=1e-6 / rtol=1e-5; BF16: atol=5e-3 / rtol=1e-2. These cold-start correctness runs do not establish a performance speedup. The capability result resolves the CPU-collective failure on this H100 / Torch 2.13 stack by registering Gloo for CPU and NCCL for CUDA; it does not retroactively change the historical L20 measurements below or implement §6B.
Repository-external evidence:
validation/fsdp_tensor_equivalence.py,artifacts/remote/fsdp-{dual-fp32,dual-bf16,alloffload-fp32}-rank{0,1}.jsonand corresponding logs/exit statuses. The NCCL-only failure is retained. The harness contains a small model; this is not a real UniRL RL-training or planner-error benchmark. No H100 planner calibration or steady-state performance result is claimed.Original L20 validation
Environment:
jk01, 8× L20 (driver 570.86.10), torch 2.11.0+cu128, base SHAf8d95462837df00188dab1376e19347b43996307.Lint and guards:
CPU contract harness (repo-external, not committed), 8/8 not-FAIL:
The partial-suffix case is reachable through an exposure limit, which is the realistic reason
to keep blocks resident; a memory budget alone cannot select it for the native executor
(because its predicted peak is flat in the resident-prefix length under the conservative
bound). Both properties are asserted, not assumed.
Real L20 calibration and a real plan:
The transfer number is the planning-relevant one: at this block size a per-block copy is
almost entirely launch overhead.
Executor capability spike for the next milestone (§6B gate, 2 ranks, real FSDP2
fully_shard, per-blockCPUOffloadPolicymix, AdamW +clip_grad_norm_, 2 steps):bf16 mixed precision reproduces the same picture. So selective offload is supported and
numerically equivalent on this stack, while an all-offloaded plan is not usable on the current
gradient path — which is why the planner's partial suffix matters, and why §6B must not be
merged before the clip/Optimizer question is settled.
Compatibility / Risk
behaviour change, no dependency added (the module is stdlib-only).
nothing in the repository consumes it yet.
SCHEMA_VERSIONis stamped into the fingerprint and every plan.Reviewer Notes
stepon CPU acrossoffload/onload), perf(veomni-ep): reduce expert gather peak memory #422 (expert gather peak), feat(flowgrpo): add bounded-memory chunked replay backward #465 (chunked replay backward) and feat(fsdp): support copy-engine all-gather #467 (merged
copy-engine all-gather) are inputs to the profile key and the memory ledger, not
re-implementations; refactor(diffusion): make role residency one choice per role, not per phase #428's role residency is untouched.
_live_staging_bytes(the conservative-vs-declared bound split),_simulate_phase(ready time is transfer completion, not transfer start) and theresident_gpu_bytesexclusion contract in the README.offloaded_validated=falseis carried through the JSON and P08 isNOT_RUN.