This guide covers the minimum path for adding an external model to Relax with SGLang rollout and Megatron training. Use the current dots.mcore / dots.mocr support as the reference implementation.
An external model must work through two compatible paths:
- Rollout: SGLang can load the model and process text or multimodal requests.
- Training: Megatron can build the same architecture and load / export weights through Megatron Bridge.
Use one HuggingFace checkpoint as the source of architecture, tokenizer, processor, and initial weights. Prefer Bridge mode:
--megatron-to-hf-mode bridgeKeep model-specific logic under relax/models/<model_name>/; keep launch scripts as configuration only.
For models not natively supported by SGLang, register an external model package:
--sglang-external-model-package relax.models.dots_ocr.sglangRelax sets the SGLang external model and multimodal processor environment variables before spawning SGLang. The package should provide:
- A SGLang-loadable model class exposed through
EntryClass. - A multimodal processor when the model has image, video, or audio inputs.
- A
load_weights()path that accepts HF-format parameter names.
For dots.mocr, relax.models.dots_ocr.sglang provides DotsOCRForCausalLM and DotsOCRImageProcessor. Its image token rule is:
<|img|><|imgpad|><|endofimg|>
Do not reuse a similar VLM processor unless the special tokens and feature layout match exactly.
Megatron integration normally consists of:
- A training model, such as
relax/models/dots_ocr/megatron/model.py. - A provider, such as
relax/models/dots_ocr/megatron/provider.py. - A Bridge adapter, such as
relax/models/dots_ocr/megatron/bridge.py.
The bridge registers the HF architecture name:
@MegatronModelBridge.register_bridge(
source="DotsOCRForCausalLM",
target=DotsOCRModel,
)
class DotsOCRBridge(MegatronModelBridge):
...The source value must match the HF checkpoint architecture. Import the Megatron package from relax/models/__init__.py so the registration runs before AutoBridge.from_hf_pretrained(...).
In mapping_registry(), map HF names to Megatron names. Common mappings include:
model.embed_tokens.weighttolanguage_model.embedding.word_embeddings.weightlm_head.weighttolanguage_model.output_layer.weight- HF
q_proj/k_proj/v_projto Megatronlinear_qkv - HF
gate_proj/up_projto Megatronlinear_fc1 - Multimodal towers, such as
vision_tower.**tovision_model.**
Use raw conversion only when Bridge cannot cover the model. Raw mode requires custom converters under relax/backends/megatron/weight_conversion/.
Split launch scripts into clear argument blocks:
source "${MODEL_CONFIG_DIR}/<model>.sh"
CKPT_ARGS=(...)
ROLLOUT_ARGS=(...)
PERF_ARGS=(...)
SGLANG_ARGS=(...)
MISC_ARGS=(...)Key checkpoint options:
--hf-checkpoint ${MODEL_DIR}/rednote-hilab/dots.mocr
--ref-load ${MODEL_DIR}/rednote-hilab/dots.mocr
--megatron-to-hf-mode bridgeKey rollout options for multimodal data:
--multimodal-keys '{"image":"image"}'Key SGLang option for external models:
--sglang-external-model-package relax.models.dots_ocr.sglangChoose exactly one execution mode:
--colocate--fully-async--hybrid
Pure fully async mode does not support --balance-data; use colocate or hybrid when data balancing is required.
Do not start full training before alignment. Check in this order:
- HF single-sample forward works.
- SGLang single-engine text and multimodal generation works.
- Megatron can build the model and load HF weights through Bridge.
- HF / SGLang / Megatron token logprobs match on a fixed sample.
- Packed sequence, context parallel, and dynamic batch paths are verified.
- One small rollout → train → update weights loop succeeds.
For dots.mocr, use:
scripts/debug/run-compare-dotsocr.sh
scripts/debug/run-compare-dotsocr-packed.shCommon alignment failures:
- Wrong multimodal special tokens or processor behavior.
- RoPE / position id mismatch.
- Missing Bridge mappings for fused qkv, fused MLP, or vision tower weights.
- SGLang
load_weights()not accepting the exported HF names.
Model checkpoint:
Relevant files:
relax/models/dots_ocr/
├── configuration.py
├── vision.py
├── sglang/
│ ├── model.py
│ └── processor.py
└── megatron/
├── bridge.py
├── model.py
└── provider.py
Launch files:
scripts/models/dotsocr2.sh
scripts/training/multimodal/run-dotsocr2-8xgpu.sh
scripts/training/multimodal/run-dotsocr2-8xgpu-hybrid.shBefore treating a model as integrated:
- The SGLang external package imports and exposes
EntryClass. - The multimodal processor uses the model's native token rules.
-
AutoBridge.from_hf_pretrained(...)finds the custom bridge. -
bridge.load_hf_weights(...)loads the checkpoint. -
mapping_registry()covers language, output, and multimodal tower weights. - The launch script uses
--megatron-to-hf-mode bridge. -
--multimodal-keysmatches the dataset fields. - Fixed-sample HF / SGLang / Megatron logprobs are aligned.
- A small rollout → train → update weights loop succeeds.