Skip to content

fix(ppo): protect against NaN/Inf gradient corruption in before_step - #2230

Open
zhoufengen wants to merge 1 commit into
facebookresearch:mainfrom
zhoufengen:fix/nan-gradient-protection
Open

fix(ppo): protect against NaN/Inf gradient corruption in before_step#2230
zhoufengen wants to merge 1 commit into
facebookresearch:mainfrom
zhoufengen:fix/nan-gradient-protection

Conversation

@zhoufengen

Copy link
Copy Markdown

Summary

NaN or Inf gradients can silently corrupt model weights when they reach optimizer.step() during PPO training. A single bad batch propagates NaN into all downstream weights, ruining multi-day training runs without any error signal.

Changes

  • Add torch.nan_to_num_ in PPO.before_step() after gradient clipping to replace NaN with zero and clamp Inf to max_grad_norm bounds
  • Add 5 unit tests in test/test_ppo_nan_gradients.py covering NaN replacement, Inf clamping, valid gradient preservation, all-parameter coverage, and end-to-end optimizer step safety

Testing

test/test_ppo_nan_gradients.py::test_nan_gradients_replaced_with_zero PASSED
test/test_ppo_nan_gradients.py::test_inf_gradients_clamped_to_max_norm PASSED
test/test_ppo_nan_gradients.py::test_valid_gradients_unchanged PASSED
test/test_ppo_nan_gradients.py::test_all_parameters_cleaned PASSED
test/test_ppo_nan_gradients.py::test_optimizer_step_with_nan_gradients PASSED

Fixes

Fixes #2226

NaN or Inf gradients can silently corrupt model weights when they reach
optimizer.step() during PPO training. A single bad batch propagates NaN
into all downstream weights, ruining multi-day training runs without any
error signal.

Add torch.nan_to_num_ after gradient clipping in before_step() to replace
NaN with zero and clamp Inf to max_grad_norm bounds, preventing silent
weight corruption.

This is a backward-compatible safety net: valid gradients are unchanged,
and the additional iteration over parameters is negligible compared to
the backward pass.

Fixes facebookresearch#2226
@meta-cla

meta-cla Bot commented May 20, 2026

Copy link
Copy Markdown

Hi @zhoufengen!

Thank you for your pull request and welcome to our community.

Action Required

In order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you.

Process

In order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA.

Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with CLA signed. The tagging process may take up to 1 hour after signing. Please give it that time before contacting us about it.

If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks!

@meta-cla

meta-cla Bot commented May 20, 2026

Copy link
Copy Markdown

Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks!

@meta-cla meta-cla Bot added the CLA Signed Do not delete this pull request or issue due to inactivity. label May 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed Do not delete this pull request or issue due to inactivity.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Silent weight corruption: PPO.before_step lets NaN gradients reach optimizer.step()

1 participant