Repository navigation
feat(inf-train parity): add Qwen3-MoE train-inference parity plugin, part 2/2 — training integration - #505
Open
cboss6 wants to merge 1 commit into
Open
feat(inf-train parity): add Qwen3-MoE train-inference parity plugin, part 2/2 — training integration#505cboss6 wants to merge 1 commit into
cboss6 wants to merge 1 commit into
Conversation
cboss6
requested review from
CjhHa1,
Ideny42,
Jayce-Ping,
Zcchill,
celve,
haonan3,
leviking98z-rgb and
zzhuoxin1508
as code owners
September 23, 2026 06:41
2 tasks
This was referenced Oct 5, 2026
2 tasks done
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Part 2 of 2: training integration and end-to-end verification for the Qwen3-MoE train–inference parity plugin.
This two-PR series splits #430 into reviewable components that together implement one end-to-end train–inference parity plugin feature, following #431. Neither part alone provides the complete verified training-and-rollout workflow.
feat/qwen3-moe-parity-vllm-pluginfeat/qwen3-moe-parity-trainingSplit the training/core half out of #430. Add an opt-in Qwen3-30B-A3B
train/inference parity experiment with FSDP4 actor replay, direct vLLM TP4
rollouts, exact selected-token FP32 comparison before backward, and a two-phase
verification artifact covering initial weights and post-update IPC reload.
not import or enable them.
publication evidence, and rollout-complete callbacks.
weight downloads.
dirty source, or runtime failure.
Merge order: Part 1/2 → Part 2/2. This PR deliberately does not contain
vllm_plugin/; it consumes the numerical providers supplied by Part 1.Install/merge Part 1 before running the experimental recipe. Together, both
parts provide the complete train–inference parity plugin feature.
Related Issue
Refs #431. Part 2/2 supersedes the training/core portion of #430; Part 1/2
supersedes its vLLM plugin portion. Do not merge the original large PR alongside
these replacements.
Test Plan
/tmp/parity_review_checks.py: eight regression groups PASS,including 18 exact baseline/refactored GRPO loss-and-gradient comparisons and
bounded two-rank Gloo failure-injection checks.
PASS.
pre-commit run --hook-stage manual --files $(git diff --name-only c8af5d6 HEAD):all applicable hooks PASS.
clean integration commit
b9e8ede. This validates the complete feature,not either split branch in isolation.
Command:
CUDA_VISIBLE_DEVICES=0,1,2,3 python -m experimental.train_inference_parity.run --config-name=qwen3_moe_30b_a3b_fsdp_tp4on 4 × H20, Python 3.13, Torch 2.13.0+cu130, Transformers 5.6.0, vLLM 0.27.0;
full GSM8K, Qwen3-30B-A3B revision
ad44e777bcd18fa416d9da3bd8f70d33ebb85d39, 256 prompt / 1024 response limit,4 prompts × 8 samples, two rollouts.
initial_checkpointandpost_update_reloadpassed with 32768 tokens,torch_equal_fp32=true, zero mismatches, zero maximum FP32 difference, andzero K3 mean/max. Each performed one optimizer update; gradient norms were
0.65234375 and 0.7421875. Parameter-change evidence and all four TP receipts
were present after reload (18867 tensors, model version 1).
outputs/train_inference_parity/review_20260922_214630/qwen3-moe-tp4.json,SHA-256
a4230d10baf5c19d5cf12dc3ca21867b7732256a9be465520a8d93a43f6b7eba.results are recorded in the handoff.
Compatibility / Risk
Default core behavior is unchanged unless the exact provider/configuration is
selected. The experimental recipe requires the companion plugin and dedicated
processes. No new transport or checkpoint format is introduced.
The guarantee is selected-token FP32 forward log-probability equality on this
matrix, not identical gradients, optimizer state, or portability. Python expert
loops and surrogate backward prioritize correctness over throughput.
Reviewer Notes
c8af5d6; branch:feat/qwen3-moe-parity-training.to integration commit
b9e8ede407d3bb0c0dd47a773c1074527c988e04.Checklist