Skip to content

feat(fsdp): add a measured resident/streamed block placement planner - #486

Open
0z5a wants to merge 1 commit into
Tencent-Hunyuan:mainfrom
0z5a:feat/483-6a-offload-planner
Open

0z5a wants to merge 1 commit into
Tencent-Hunyuan:mainfrom
0z5a:feat/483-6a-offload-planner

Conversation

@0z5a

@0z5a 0z5a commented Sep 18, 2026 •

Copy link
Copy Markdown

Incremental review: Files changed against main.

Summary

Section 6A of the trainside optimization roadmap: decide how much to offload from
measurements instead of a boolean. This PR adds the measurement contract, a pure planner and
the first real calibration numbers; it adds no new executor.

unirl/train/backend/fsdp/offload_plan.py consumes a measured PhaseProfile plus a
BlockInventory and returns an auditable PlacementPlan:

  • No torch, no CUDA, no allocation, no checkpoint read. Pure data in, plan out, so a plan
    can be re-derived and diffed from a JSON profile on any host.
  • Group-common plan. Every rank is evaluated for the same resident prefix; the chosen
    candidate must satisfy the tightest rank budget and the node-summed pinned-host budget.
    There is no per-rank plan, because ranks sharing an FSDP shard group must agree on ownership.
  • Exclusive memory ledger. PhaseCost.resident_gpu_bytes measures the phase's per-rank
    high-water excluding all block parameter storage; the planner then adds root_bytes,
    resident block staging, streamed staging and headroom itself. That contract is what makes
    the ledger check below exact instead of double-counted.
  • Executor capabilities are not interchangeable. With live_streamed_bound=None — the
    honest default for native FSDP2 CPUOffloadPolicy — every streamed block is charged at
    peak. The two-slot model applies only when the executor declares that bound.
  • Capacity and capability refusals. Alias groups move together, fixed/root blocks never
    stream, a group split across the boundary is rejected, an unmeasured phase rejects every
    candidate instead of extrapolating, and full residency is never an automatic fallback: no
    feasible candidate yields INFEASIBLE with per-candidate reasons and nothing enabled.
  • Auditable output. to_json() carries the selected plan, the predicted per-rank peaks,
    per-node host demand, transfer_busy_ms, no_compute_transfer_ms and every candidate's
    predictions and rejection reasons.

unirl/train/backend/fsdp/README.md records the profile contract that a calibration run must
honour, including the two footguns a reviewer should check first: compute_ms must exclude
transfer waits, and only bandwidth_record_kind="shard" may be summed with shard_bytes.

Related Issue

Refs #483 (section 6A). This is the planner and measurement contract only; §6B's executor and
§6C's phase-aware streaming are explicitly not part of it.

Test Plan

Archived supplementary evidence (published 2026-10-02): Raw results, harnesses and SHA-256 checksums. Test execution date: 2026-09-28; no new GPU run was performed for this publication.

H100 follow-up — 2026-09-28

Rechecked PR head 57cec429b5978d7da4bfd5f751764c0fd5da4c06. New hardware: 2× NVIDIA H100 80GB HBM3, NVLink, driver 595.71.05, Python 3.12, PyTorch 2.13.0+cu130. The original environment was retained. H100 validation is accepted for this review, as agreed with the maintainer.

The recovered planner CPU harness passes P01–P07 on this host; P08 still needs the §6B executor. A separate two-rank FSDP2 capability harness compares identical initialization, input and two AdamW updates against full residency:

Validation Baseline Candidate / result Speedup Latency reduction
Original CPU planner contracts P01–P07 PASS P01–P07 PASS; P08 NOT_RUN N/A N/A
Selective offload, NCCL-only process group Resident passes Fails at CPU gradient-norm collective N/A N/A
Selective FP32, cpu:gloo,cuda:nccl Resident PASS; 428 tensor comparisons, maximum absolute difference 3.949e-7 N/A N/A
Selective BF16, same dual backend Resident PASS; 428 comparisons, maximum absolute difference 2.384e-7 N/A N/A
All-offloaded FP32, same dual backend Resident PASS; 428 comparisons, maximum absolute difference 2.384e-7 N/A N/A
Planner versus real §6B executor No executor NOT_RUN N/A N/A

Each comparison includes logits, loss, every parameter gradient, clipping norm, updated parameters and both Adam moments. FP32 tolerances: atol=1e-6 / rtol=1e-5; BF16: atol=5e-3 / rtol=1e-2. These cold-start correctness runs do not establish a performance speedup. The capability result resolves the CPU-collective failure on this H100 / Torch 2.13 stack by registering Gloo for CPU and NCCL for CUDA; it does not retroactively change the historical L20 measurements below or implement §6B.

Repository-external evidence: validation/fsdp_tensor_equivalence.py, artifacts/remote/fsdp-{dual-fp32,dual-bf16,alloffload-fp32}-rank{0,1}.json and corresponding logs/exit statuses. The NCCL-only failure is retained. The harness contains a small model; this is not a real UniRL RL-training or planner-error benchmark. No H100 planner calibration or steady-state performance result is claimed.

Original L20 validation

Environment: jk01, 8× L20 (driver 570.86.10), torch 2.11.0+cu128, base SHA
f8d95462837df00188dab1376e19347b43996307.

Lint and guards:

ruff check unirl/train/backend/fsdp/                  -> All checks passed!
ruff format --check unirl/train/backend/fsdp/         -> 5 files already formatted
python lint/check_docstring_lines.py                  -> all one line
python lint/check_core_dependencies.py                -> ok

CPU contract harness (repo-external, not committed), 8/8 not-FAIL:

python evidence/harness/test_6a_planner.py --out evidence/validation_results_6a_cpu.json
  P01 PASS  N=0/1, bool/float/negative/NaN/inf, zero bandwidth, unknown block, missing phase
  P02 PASS  all-offloaded / infeasible / exposure-forced PARTIAL suffix; uneven bytes audited
  P03 PASS  fixed blocks resident, alias split rejected, duplicate id/order/phase rejected
  P04 PASS  two-slot bound only when declared; double buffering forces more resident blocks
  P05 PASS  backward exposes a stall the forward order hides; unmeasured phase rejects
  P06 PASS  max-rank timing, tightest-rank budget and node-summed host gate the shared plan
  P07 PASS  deterministic plans, JSON round-trip, missing budget row, bad units
  P08 NOT_RUN  planner-vs-observed error needs the §6B executor

The partial-suffix case is reachable through an exposure limit, which is the realistic reason
to keep blocks resident; a memory budget alone cannot select it for the native executor
(because its predicted peak is flat in the resident-prefix length under the conservative
bound). Both properties are asserted, not assumed.

Real L20 calibration and a real plan:

python evidence/harness/gpu_6a_profile.py --out evidence/validation_results_6a_gpu.json
  inventory     4 JointTransformerBlock, root 72960 B, blocks 151552/151552/151552/97024 B
  forward       1.154 / 1.247 / 1.143 / 0.969 ms per block (total 5.32 ms)
  backward      8.788 / 2.276 / 2.389 / 2.525 ms per block (total 16.31 ms)
  AdamW step    1.603 ms
  pinned H2D    32.6 us for a 151 KB block -> fitted launch overhead 31.8 us, 189 GB/s
  ledger check  all-resident predicted == observed exactly (delta 0 B) for both phases
  pass=True

The transfer number is the planning-relevant one: at this block size a per-block copy is
almost entirely launch overhead.

Executor capability spike for the next milestone (§6B gate, 2 ranks, real FSDP2
fully_shard, per-block CPUOffloadPolicy mix, AdamW + clip_grad_norm_, 2 steps):

torchrun --nproc_per_node=2 evidence/harness/gpu_6b_spike.py [--mixed-precision]
  all_resident       trains  placements 26 cuda   grad_norm 2.6192903519
  selective_suffix_2 trains  placements 14 cuda / 12 cpu
                             grad_norm 2.6192903519 (identical)  params 2e-06 abs / 6.3e-08 rel
  all_offloaded       trains  clip_grad_norm_ raises:
                             RuntimeError: No backend type associated with device type cpu
                             params 0.130 abs / 4.1e-03 rel

bf16 mixed precision reproduces the same picture. So selective offload is supported and
numerically equivalent on this stack, while an all-offloaded plan is not usable on the current
gradient path — which is why the planner's partial suffix matters, and why §6B must not be
merged before the clip/Optimizer question is settled.

Compatibility / Risk

  • Additive file plus a README. No existing module imports it, no config key, no runtime
    behaviour change, no dependency added (the module is stdlib-only).
  • The profile schema is a new contract: any field rename breaks the evidence directory, but
    nothing in the repository consumes it yet.
  • SCHEMA_VERSION is stamped into the fingerprint and every plan.

Reviewer Notes

@github-actions github-actions Bot added the need review Ready and waiting for review label Sep 18, 2026
@0z5a
0z5a force-pushed the feat/483-6a-offload-planner branch from 0b6f4d8 to e124158 Compare September 24, 2026 01:15
Section 6A: a pure-Python planner that consumes a calibrated per-block/per-phase
profile plus a block inventory and returns an auditable resident-prefix /
offloaded-suffix plan. No torch, no CUDA, no allocation, no checkpoint read.

Reports predicted per-rank peak memory, per-node pinned host demand, exposed
transfer stall and total phase cost, rejects every candidate that violates a
budget or capability, refuses to treat full residency as a fallback, and refuses
to plan without a measured phase. An executor with no measured residency bound is
charged conservatively rather than with a two-slot model it has not verified.

Refs Tencent-Hunyuan#483 (section 6A)

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a

0z5a commented Oct 2, 2026

Copy link
Copy Markdown
Author

H100 supplementary validation record

The archived 2026-09-28 results have been checked against the raw files and published as test artifacts and repository-external harnesses on 2026-10-02.

Pinned PR head: 57cec429b5978d7da4bfd5f751764c0fd5da4c06. The recorded stack is 2× NVIDIA H100 80GB HBM3, PyTorch 2.13.0+cu130; the host topology reports NV18 between the GPUs.

Check Result Baseline / candidate details Speedup
Planner CPU contracts P01–P07 PASS; P08 NOT_RUN Recovered original harness on the H100 host N/A
Selective FP32, NCCL-only FAIL, retained Resident arm completes; offload arm fails at the CPU gradient-norm collective; both ranks exit 1 N/A
Selective FP32, cpu:gloo,cuda:nccl PASS Full residency versus selective offload; 428 tensor comparisons; max abs 3.948807716369629e-7; both ranks exit 0 N/A
Selective BF16, same dual backend PASS 428 comparisons; max abs 2.384185791015625e-7; both ranks exit 0 N/A
All-offloaded FP32, same dual backend PASS 428 comparisons; max abs 2.384185791015625e-7; both ranks exit 0 N/A

Each rank compares 214 tensors after two identical AdamW steps: logits, loss, parameter gradients, clipping norm, updated parameters and both Adam moments. FP32 tolerance is atol=1e-6 / rtol=1e-5; BF16 is atol=5e-3 / rtol=1e-2.

The evidence includes all six successful per-rank JSON files, process exit records, the NCCL-only failure logs, the standalone FSDP harness, the planner CPU log and SHA-256 checksums. Empty success logs are recorded separately in the manifest.

These are standalone FSDP2 capability/correctness checks. P08, H100 planner calibration and steady-state performance remain NOT_RUN. The raw timings retain their diagnostic-only label (cold iterations and existing serving occupancy), so speedup and latency reduction are N/A. No new GPU run was performed for this publication.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

need review Ready and waiting for review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant