Conversation
cboss6
requested review from
CjhHa1,
Ideny42,
Jayce-Ping,
Zcchill,
celve,
haonan3,
leviking98z-rgb and
zzhuoxin1508
as code owners
September 9, 2026 17:15
2 tasks done
2 tasks
Make the exact gate collective-safe, strengthen reload evidence, and keep custom Qwen3-MoE expert and o-projection layouts valid across vLLM sleep and native layerwise weight reload.
cboss6
force-pushed
the
qwen3-moe-parity-tp4
branch
from
September 22, 2026 11:33
57d8289 to
b7c06fe
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR has been split into #504 and #505.
This PR introduces an experimental train–inference parity path for Qwen3-MoE.
In on-policy RL, numerical drift between vLLM rollout log-probabilities and differentiable actor replay can introduce artificial KL, move the importance ratio away from 1, and trigger clipping before any intended policy difference exists.
This experiment turns parity into a fail-closed runtime contract:
torch.equalbefore backward;token_count,torch_equal_fp32,mismatch_count,max_absdiff_fp32,k3_mean, andk3_maxmetrics;Related PR
This PR can be merged before below PR merged:
#480
Supported case
The currently validated configuration is:
public_referenceThe normal Qwen3-MoE training path remains unchanged and does not load the experimental plugin.

Post-update weight reload
The validation covers both the initial checkpoint and a non-zero optimizer update followed by FSDP → vLLM TP weight reload.
The dedicated synchronization path provides:
Validation
A clean two-rollout TP4 run passed both phases:
initial_checkpointpost_update_reloadFor both phases:
torch_equal_fp32 = truemismatch_count = 0max_absdiff_fp32 = 0.0k3_mean = 0.0k3_max = 0.0The post-update reload synchronized 18,867 tensors across 114 buckets, committed the same model version on TP0–TP3, and verified that actor parameters changed.
Scope
The bitwise guarantee applies to selected-token FP32 forward log-probabilities on the validated hardware/software matrix.
It does not claim bitwise-identical gradients, optimizer states, or portability across arbitrary GPU architectures, CUDA versions, kernels, or vLLM releases.
Related Issue
Test Plan
Compatibility / Risk
Reviewer Notes
Checklist