fix(ppo): protect against NaN/Inf gradient corruption in before_step - #2230
fix(ppo): protect against NaN/Inf gradient corruption in before_step#2230zhoufengen wants to merge 1 commit into
Conversation
NaN or Inf gradients can silently corrupt model weights when they reach optimizer.step() during PPO training. A single bad batch propagates NaN into all downstream weights, ruining multi-day training runs without any error signal. Add torch.nan_to_num_ after gradient clipping in before_step() to replace NaN with zero and clamp Inf to max_grad_norm bounds, preventing silent weight corruption. This is a backward-compatible safety net: valid gradients are unchanged, and the additional iteration over parameters is negligible compared to the backward pass. Fixes facebookresearch#2226
|
Hi @zhoufengen! Thank you for your pull request and welcome to our community. Action RequiredIn order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you. ProcessIn order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA. Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks! |
|
Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks! |
Summary
NaN or Inf gradients can silently corrupt model weights when they reach
optimizer.step()during PPO training. A single bad batch propagates NaN into all downstream weights, ruining multi-day training runs without any error signal.Changes
torch.nan_to_num_inPPO.before_step()after gradient clipping to replace NaN with zero and clamp Inf tomax_grad_normboundstest/test_ppo_nan_gradients.pycovering NaN replacement, Inf clamping, valid gradient preservation, all-parameter coverage, and end-to-end optimizer step safetyTesting
Fixes
Fixes #2226