Skip to content

feat(inf-train parity): add Qwen3-MoE train-inference parity plugin, part 2/2 — training integration - #505

Open
cboss6 wants to merge 1 commit into
Tencent-Hunyuan:mainfrom
cboss6:feat/qwen3-moe-parity-training
Open

cboss6 wants to merge 1 commit into
Tencent-Hunyuan:mainfrom
cboss6:feat/qwen3-moe-parity-training

Conversation

@cboss6

@cboss6 cboss6 commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

Summary

Part 2 of 2: training integration and end-to-end verification for the Qwen3-MoE train–inference parity plugin.

This two-PR series splits #430 into reviewable components that together implement one end-to-end train–inference parity plugin feature, following #431. Neither part alone provides the complete verified training-and-rollout workflow.

Part Scope Branch
Part 1/2 — prerequisite PR vLLM plugin, shared numerical providers, model-specific inference patches, and reload lifecycle feat/qwen3-moe-parity-vllm-plugin
Part 2/2 — this PR Training-side model adapter, core integration, exact parity gate, verification pipeline, and example feat/qwen3-moe-parity-training

Split the training/core half out of #430. Add an opt-in Qwen3-30B-A3B
train/inference parity experiment with FSDP4 actor replay, direct vLLM TP4
rollouts, exact selected-token FP32 comparison before backward, and a two-phase
verification artifact covering initial weights and post-update IPC reload.

  • Keep experimental model adapters outside the UniRL wheel; baseline recipes do
    not import or enable them.
  • Add narrow core seams for exact replay providers, rollout capabilities,
    publication evidence, and rollout-complete callbacks.
  • Reuse baseline GRPO loss construction; honor model revisions in meta-init
    weight downloads.
  • Fail closed on incomplete micro-batch evidence, invalid frozen configuration,
    dirty source, or runtime failure.

Merge order: Part 1/2 → Part 2/2. This PR deliberately does not contain
vllm_plugin/; it consumes the numerical providers supplied by Part 1.
Install/merge Part 1 before running the experimental recipe. Together, both
parts provide the complete train–inference parity plugin feature.

Related Issue

Refs #431. Part 2/2 supersedes the training/core portion of #430; Part 1/2
supersedes its vLLM plugin portion. Do not merge the original large PR alongside
these replacements.

Test Plan

  • Local-only /tmp/parity_review_checks.py: eight regression groups PASS,
    including 18 exact baseline/refactored GRPO loss-and-gradient comparisons and
    bounded two-rank Gloo failure-injection checks.
  • Reran those checks from the split training worktree with the companion plugin:
    PASS.
  • pre-commit run --hook-stage manual --files $(git diff --name-only c8af5d6 HEAD):
    all applicable hooks PASS.
  • Full frozen-config preflight: PASS.
  • Full GPU run with Part 1/2 + Part 2/2 together: PASS, exit code 0,
    clean integration commit b9e8ede. This validates the complete feature,
    not either split branch in isolation.
    Command:
    CUDA_VISIBLE_DEVICES=0,1,2,3 python -m experimental.train_inference_parity.run --config-name=qwen3_moe_30b_a3b_fsdp_tp4
    on 4 × H20, Python 3.13, Torch 2.13.0+cu130, Transformers 5.6.0, vLLM 0.27.0;
    full GSM8K, Qwen3-30B-A3B revision
    ad44e777bcd18fa416d9da3bd8f70d33ebb85d39, 256 prompt / 1024 response limit,
    4 prompts × 8 samples, two rollouts.
  • Both initial_checkpoint and post_update_reload passed with 32768 tokens,
    torch_equal_fp32=true, zero mismatches, zero maximum FP32 difference, and
    zero K3 mean/max. Each performed one optimizer update; gradient norms were
    0.65234375 and 0.7421875. Parameter-change evidence and all four TP receipts
    were present after reload (18867 tensors, model version 1).
  • Artifact:
    outputs/train_inference_parity/review_20260922_214630/qwen3-moe-tp4.json,
    SHA-256 a4230d10baf5c19d5cf12dc3ca21867b7732256a9be465520a8d93a43f6b7eba.
  • No tests/harness files committed, per repository policy; commands and local
    results are recorded in the handoff.

Compatibility / Risk

Default core behavior is unchanged unless the exact provider/configuration is
selected. The experimental recipe requires the companion plugin and dedicated
processes. No new transport or checkpoint format is introduced.

The guarantee is selected-token FP32 forward log-probability equality on this
matrix, not identical gradients, optimizer state, or portability. Python expert
loops and surrogate backward prioritize correctness over throughput.

Reviewer Notes

  • AI-assisted review/refactor; the human submitter must review and own the diff.
  • Baseline: c8af5d6; branch: feat/qwen3-moe-parity-training.
  • Its diff and the companion diff are disjoint; the merged tree is byte-identical
    to integration commit b9e8ede407d3bb0c0dd47a773c1074527c988e04.

Checklist

  • Human submitter reviewed the full diff and accepts ownership.
  • No unrelated notes, generated artifacts, or one-off tests are committed.
  • Documentation and recipe accompany the implementation.
  • Full two-phase GPU validation completed; Test Plan updated.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

need review Ready and waiting for review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant