Repository navigation
fix: resize model vocab to match checkpoint directory before FSDP wrap - #115
Open
Papaercold wants to merge 1 commit into
Open
Papaercold wants to merge 1 commit into
Papaercold wants to merge 1 commit into
Conversation
When resuming FSDP2 training from a checkpoint directory (not a single .safetensors file), the model is created with the processor's vocabulary size (151667) but the checkpoint weights carry the action-expanded vocabulary (155765). The single-file resume path already handles this via _maybe_resize_token_embeddings_for_load, but the directory resume path skips that step. FSDP shards are then created at the wrong vocab size, and set_model_state_dict fails with a tensor-size mismatch. Fix: inspect the checkpoint directory's model.safetensors embedding shape and call resize_token_embeddings BEFORE _wrap_model(), so sharded DTensor shards match the checkpoint when resume_from_checkpoint() loads it.
Papaercold
force-pushed
the
fix/vocab-resize-on-checkpoint-dir-resume
branch
from
July 26, 2026 03:46
d5f7610 to
a0a4c9f
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Resuming FSDP2 training from a checkpoint directory fails with a vocabulary-size mismatch, even when using the same config, processor, dataset, and world size.
When
checkpoint.resume_frompoints to a directory (not a single.safetensorsfile), the model is created with the processor's vocabulary (151667), but checkpoint weights carry the action-expanded vocabulary (155765). FSDP shards are created at the wrong vocab size, andset_model_state_dictfails:The single-file resume path already handles this via
_maybe_resize_token_embeddings_for_load(incheckpoint_io.py), but the directory resume path (_load_fsdp2_or_dmuon_full) skips that step becauseload_weights()resizes to the processor vocab and the directory path skips the subsequent single-fileload_state_dict.See related issue: #114
Fix
Inspect the checkpoint directory's
model.safetensorsembedding shape and callresize_token_embeddingsbefore_wrap_model(), so FSDP shards match the checkpoint whenresume_from_checkpoint()loads it after wrapping.This mirrors the vocabulary adaptation already implemented for single-file loading (
_maybe_resize_token_embeddings_for_load), extended to the directory resume path.Reproduction
resume_frompointing to the pretrained.safetensorsfileresume_fromto the checkpoint directory and restart trainingset_model_state_dicttensor-size mismatchWith this fix, step 3 succeeds and training resumes from the saved epoch with full optimizer/scheduler/RNG state restored.
Notes
fsdp_trainer.pyis modified;checkpoint_io.pyis untouched