Skip to content

fix: resize model vocab to match checkpoint directory before FSDP wrap - #115

Open
Papaercold wants to merge 1 commit into
X-Square-Robot:mainfrom
Papaercold:fix/vocab-resize-on-checkpoint-dir-resume
Open

Papaercold wants to merge 1 commit into
X-Square-Robot:mainfrom
Papaercold:fix/vocab-resize-on-checkpoint-dir-resume

Conversation

@Papaercold

@Papaercold Papaercold commented Jul 26, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Resuming FSDP2 training from a checkpoint directory fails with a vocabulary-size mismatch, even when using the same config, processor, dataset, and world size.

When checkpoint.resume_from points to a directory (not a single .safetensors file), the model is created with the processor's vocabulary (151667), but checkpoint weights carry the action-expanded vocabulary (155765). FSDP shards are created at the wrong vocab size, and set_model_state_dict fails:

RuntimeError: The size of tensor a (18959) must match the size of tensor b (19471) at non-singleton dimension 0

The single-file resume path already handles this via _maybe_resize_token_embeddings_for_load (in checkpoint_io.py), but the directory resume path (_load_fsdp2_or_dmuon_full) skips that step because load_weights() resizes to the processor vocab and the directory path skips the subsequent single-file load_state_dict.

See related issue: #114

Fix

Inspect the checkpoint directory's model.safetensors embedding shape and call resize_token_embeddings before _wrap_model(), so FSDP shards match the checkpoint when resume_from_checkpoint() loads it after wrapping.

This mirrors the vocabulary adaptation already implemented for single-file loading (_maybe_resize_token_embeddings_for_load), extended to the directory resume path.

Reproduction

  1. Train Wall-OSS-0.5 with resume_from pointing to the pretrained .safetensors file
  2. Save a checkpoint (epoch complete, full state: model + optimizer + scheduler + RNG + epoch)
  3. Set resume_from to the checkpoint directory and restart training
  4. Observe the set_model_state_dict tensor-size mismatch

With this fix, step 3 succeeds and training resumes from the saved epoch with full optimizer/scheduler/RNG state restored.

Notes

  • Only fsdp_trainer.py is modified; checkpoint_io.py is untouched
  • The fix checks multiple common embedding key names for robustness across model variants
  • If the checkpoint vocab already matches the model, the resize is a no-op

When resuming FSDP2 training from a checkpoint directory (not a single
.safetensors file), the model is created with the processor's vocabulary
size (151667) but the checkpoint weights carry the action-expanded
vocabulary (155765). The single-file resume path already handles this via
_maybe_resize_token_embeddings_for_load, but the directory resume path
skips that step. FSDP shards are then created at the wrong vocab size,
and set_model_state_dict fails with a tensor-size mismatch.

Fix: inspect the checkpoint directory's model.safetensors embedding shape
and call resize_token_embeddings BEFORE _wrap_model(), so sharded DTensor
shards match the checkpoint when resume_from_checkpoint() loads it.
@Papaercold
Papaercold force-pushed the fix/vocab-resize-on-checkpoint-dir-resume branch from d5f7610 to a0a4c9f Compare July 26, 2026 03:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant