Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
32 commits
Select commit Hold shift + click to select a range
da4b2f9
feat(rollout): move the vllm-omni engine to 0.28.0rc1 on transformers…
celve Sep 1, 2026
db9c987
fix(rollout): let vLLM report an unconfigured MoE workspace lane itself
celve Sep 2, 2026
ca2ac49
Merge branch 'main' into LIN-981
CjhHa1 Sep 14, 2026
45112cb
Merge remote-tracking branch 'upstream/main' into fix/pr413-stable
CjhHa1 Sep 14, 2026
e076e85
feat(rollout): adapt vllm-omni to stable 0.28
CjhHa1 Sep 14, 2026
5378e0f
test(rollout): cover vllm-omni 0.28 contracts
CjhHa1 Sep 14, 2026
1e0ff2f
test(rollout): add real CuMem workspace smoke
CjhHa1 Sep 14, 2026
c08fc96
ci: install vllm-omni contract dependencies
CjhHa1 Sep 15, 2026
caf6195
ci: install diffusion contract dependency
CjhHa1 Sep 15, 2026
144fdbb
fix(rollout): keep Qwen3-Omni thinker LoRA-capable
CjhHa1 Sep 15, 2026
bb029ec
fix(rollout): retain HI3 expert mapping compatibility
CjhHa1 Sep 15, 2026
e73fea5
fix(rollout): defer HI3 diffusion handoff until decode
CjhHa1 Sep 15, 2026
d2e8076
fix(rollout): probe wrapped diffusion parameters
CjhHa1 Sep 15, 2026
4668869
fix(rollout): traverse stable diffusion worker wrappers
CjhHa1 Sep 15, 2026
279c37e
fix(rollout): skip empty parameter wrappers
CjhHa1 Sep 15, 2026
b39050f
Merge remote-tracking branch 'upstream/main' into fix/pr413-stable
CjhHa1 Sep 15, 2026
73e563b
fix(rollout): require stable stage metadata
CjhHa1 Sep 15, 2026
d03c1b7
fix(rollout): preserve HI3 AR generation mode
CjhHa1 Sep 15, 2026
acad292
ci: exercise installed vllm-omni contracts
CjhHa1 Sep 15, 2026
44441e0
fix(rollout): pin HI3 AR stages to mp
CjhHa1 Sep 15, 2026
43e5ca9
Merge branch 'main' into LIN-981
Jayce-Ping Sep 18, 2026
aa11992
test(rollout): remove vllm-omni PR tests
CjhHa1 Sep 18, 2026
1023b89
Merge remote-tracking branch 'celve/LIN-981' into fix/pr413-stable
CjhHa1 Sep 18, 2026
d1065ab
Merge remote-tracking branch 'upstream/main' into fix/pr413-stable
CjhHa1 Sep 21, 2026
15b801f
fix(rollout): support direct vllm on 0.28
CjhHa1 Sep 21, 2026
78014c6
refactor(rollout): simplify vllm-omni backend helpers
CjhHa1 Sep 21, 2026
a4e46ff
refactor(rollout): tighten vllm-omni backend boundaries
CjhHa1 Sep 21, 2026
775de27
docs(install): align extras notes with the 0.28 pins
CjhHa1 Sep 23, 2026
fc9ef9a
Merge remote-tracking branch 'upstream/main' into fix/pr413-stable
CjhHa1 Sep 23, 2026
3578e08
Merge branch 'main' into LIN-981
CjhHa1 Sep 23, 2026
c23fd66
refactor(rollout): tidy vllm-omni 0.28 docs and boundaries
CjhHa1 Sep 24, 2026
a02b484
Merge branch 'main' into LIN-981
CjhHa1 Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 11 additions & 16 deletions INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ virtualenv, never `--all-extras`. `train` and `infer` do not pull a rollout engi
| Engine extra | PyTorch | CUDA |
|---|---|---|
| `vllm` (vLLM + vLLM-Omni) | `2.13.0+cu130` | 13.0 |
| `sglang` | `2.11.0+cu130` | 13.0 |
| `sglang` | `2.13.0+cu130` | 13.0 |

SGLang's wheel needs glibc >= 2.34. Put NVIDIA's CUDA 13 forward-compat
libraries on `LD_LIBRARY_PATH` before launch; the launchers do not do this.
Expand All @@ -20,7 +20,7 @@ libraries on `LD_LIBRARY_PATH` before launch; the launchers do not do this.

```bash
uv venv --python 3.12 --seed .venv && source .venv/bin/activate
uv pip install -e ".[vllm,train,infer]" --prerelease=allow
uv pip install -e ".[vllm,train,infer]"
```

## sglang
Expand All @@ -47,7 +47,7 @@ Limit build parallelism with `MAX_JOBS` if host RAM is constrained, then install
the SGLang extra:

```bash
uv pip install -e ".[sglang,train,infer]" --prerelease=allow
uv pip install -e ".[sglang,train,infer]"
```

## Extras
Expand All @@ -58,7 +58,7 @@ uv pip install -e ".[sglang,train,infer]" --prerelease=allow
| `sglang` | `sglang[diffusion]`, `checkpoint-engine`, `flash-attn-4`, `flash-linear-attention[conv1d]`, torch +cu130 stack, PyAV | SGLang-based AR/VLM and diffusion recipes |
| `fastvideo` | FastVideo pinned to an upstream Git commit | WAN 2.1 / 2.2 rollout; [the extra does not currently resolve](#fastvideo-installation-blocker) |
| `train` | `wandb`, `aiohttp`, `math-verify` | Training runs and local math-answer scoring |
| `cosmos3` | `diffusers>=0.39` | [Cosmos3 SFT](unirl/models/cosmos3/README.md); apply the [version constraint](#cosmos3-version-prerequisite) |
| `cosmos3` | `diffusers>=0.39` | [Cosmos3 SFT](unirl/models/cosmos3/README.md); uv's `diffusers==0.40.0` override already satisfies this extra |
| `infer` | `accelerate`, `timm` | HunyuanImage3, Janus-Pro, and similar models |
| `eval` | `torchvision`, `paddlepaddle`, `paddleocr`, `python-Levenshtein` | OCR-based reward components |
| `veomni` | `veomni` | Recipes using the [VeOmni training backend](unirl/train/backend/veomni/) |
Expand All @@ -76,9 +76,9 @@ engine extras; install it when you need OCR rewards.
For development tools (lint and tests):

```bash
uv pip install -e ".[vllm,train,infer,eval,dev]" --prerelease=allow
uv pip install -e ".[vllm,train,infer,eval,dev]"
# or, for the sglang engine:
uv pip install -e ".[sglang,train,infer,eval,dev]" --prerelease=allow
uv pip install -e ".[sglang,train,infer,eval,dev]"
```

Prefer these extras over the legacy [`requirements.txt`](requirements.txt) and
Expand All @@ -89,24 +89,19 @@ Prefer these extras over the legacy [`requirements.txt`](requirements.txt) and
The `fastvideo` extra pins
[hao-ai-lab/FastVideo@2095477](https://github.com/hao-ai-lab/FastVideo/blob/2095477eac7e289c7a7ab13acb367ca60687c304/pyproject.toml),
which requires `transformers==4.57.3` and `wandb>=0.21.0`. The transformers pin
conflicts with UniRL's `transformers>=5.6,<5.7`, so `.[fastvideo]` does not
conflicts with UniRL's `transformers>=5.12,<5.13`, so `.[fastvideo]` does not
resolve — a separate venv does not help, because UniRL's base deps still apply.
Adding `train` also conflicts on `wandb`. Use `$FASTVIDEO_PATH` as in the
[FastVideo engine README](unirl/rollout/engine/fastvideo/README.md) until the extra
is solvable.

### Cosmos3 version prerequisite
### Cosmos3

`cosmos3` asks for `diffusers>=0.39`, but uv's override `diffusers>=0.38.0`
[replaces](https://docs.astral.sh/uv/concepts/resolution/#dependency-overrides)
that floor instead of intersecting with it. Include `cosmos3` in the extras and
pass `--constraint` when installing:
`cosmos3` asks for `diffusers>=0.39`. The uv override pins `diffusers==0.40.0`,
which already satisfies that floor, so install it as a normal extra:

```bash
COSMOS3_CONSTRAINTS="$(mktemp)"
printf '%s\n' 'diffusers>=0.39' > "$COSMOS3_CONSTRAINTS"
uv pip install -e ".[vllm,train,infer,cosmos3]" --prerelease=allow \
--constraint "$COSMOS3_CONSTRAINTS"
uv pip install -e ".[vllm,train,infer,cosmos3]"
```

## Environment
Expand Down
2 changes: 1 addition & 1 deletion datasets/ucf101/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ training environment can be installed with:
```bash
uv venv --python 3.12 --seed .venv
source .venv/bin/activate
uv pip install -e ".[vllm,train,infer]" --prerelease=allow
uv pip install -e ".[vllm,train,infer]"
```

Both supported engine extras (`vllm` and `sglang`) install PyAV for raw video
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -114,7 +114,7 @@ rollout:
_target_: unirl.rollout.engine.vllm_omni.config.VLLMOmniEngineConfig
model_path: ${oc.env:QWEN3_OMNI_PATH,/path/to/Qwen3-Omni-30B-A3B-Instruct}
modality: qwen3_omni_thinker
stage_yaml_override: qwen3_omni_thinker_only_rl_1x4.yaml
deploy_config_override: qwen3_omni_thinker_only_rl_1x4.yaml
enable_sleep_mode: true
max_prompt_length: 16384
video_fps: 2.0
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -124,7 +124,7 @@ rollout:
model_path: ${oc.env:QWEN3_OMNI_PATH,/path/to/Qwen3-Omni-30B-A3B-Instruct}
modality: qwen3_omni_thinker
# Audio-in-video must serialize packed forwards to avoid elevated rollout/replay K3.
stage_yaml_override: qwen3_omni_thinker_only_rl_audio_video_1x4.yaml
deploy_config_override: qwen3_omni_thinker_only_rl_audio_video_1x4.yaml
enable_sleep_mode: true
max_prompt_length: 16384
video_fps: 2.0
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -114,7 +114,7 @@ rollout:
model_path: ${oc.env:QWEN3_OMNI_PATH,/path/to/Qwen3-Omni-30B-A3B-Instruct}
modality: qwen3_omni_thinker
# Audio-in-video must serialize packed forwards to avoid elevated rollout/replay K3.
stage_yaml_override: qwen3_omni_thinker_only_rl_audio_video_1x4.yaml
deploy_config_override: qwen3_omni_thinker_only_rl_audio_video_1x4.yaml
enable_sleep_mode: true
max_prompt_length: 16384
video_fps: 2.0
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -109,7 +109,7 @@ rollout:
_target_: unirl.rollout.engine.vllm_omni.config.VLLMOmniEngineConfig
model_path: ${oc.env:QWEN3_OMNI_PATH,/path/to/Qwen3-Omni-30B-A3B-Instruct}
modality: qwen3_omni_thinker
stage_yaml_override: qwen3_omni_thinker_only_rl_1x4.yaml
deploy_config_override: qwen3_omni_thinker_only_rl_1x4.yaml
enable_sleep_mode: true
max_prompt_length: 16384
image_max_pixels: ${bundle.config.image_max_pixels}
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -122,7 +122,7 @@ rollout:
_target_: unirl.rollout.engine.vllm_omni.config.VLLMOmniEngineConfig
model_path: ${oc.env:QWEN3_OMNI_PATH,/path/to/Qwen3-Omni-30B-A3B-Instruct}
modality: qwen3_omni_thinker
stage_yaml_override: qwen3_omni_thinker_only_rl_1x4.yaml # TP=4 (audio tower 20 heads)
deploy_config_override: qwen3_omni_thinker_only_rl_1x4.yaml # TP=4 (audio tower 20 heads)
enable_sleep_mode: true
omni_extra:
stage_init_timeout: 1200
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -109,7 +109,7 @@ rollout:
_target_: unirl.rollout.engine.vllm_omni.config.VLLMOmniEngineConfig
model_path: ${oc.env:QWEN3_OMNI_PATH,/path/to/Qwen3-Omni-30B-A3B-Instruct}
modality: qwen3_omni_thinker
stage_yaml_override: qwen3_omni_thinker_only_rl_1x4.yaml # TP=4 (audio tower 20 heads)
deploy_config_override: qwen3_omni_thinker_only_rl_1x4.yaml # TP=4 (audio tower 20 heads)
enable_sleep_mode: true
omni_extra:
stage_init_timeout: 1200
Expand Down
4 changes: 2 additions & 2 deletions examples/diffusion/bagel/bagel_it2i_managed_editscore.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -137,7 +137,7 @@ backend:
total_steps: 100000
lora_cfg:
_target_: unirl.train.configs.LoraConfig
rank: 64 # must be <= the stage YAML's max_lora_rank (64)
rank: 64 # must be <= the deploy config's max_lora_rank (64)
alpha: 128
dropout: 0.0
bias: none
Expand All @@ -162,7 +162,7 @@ rollout:
# Required; same checkpoint the bundle loads.
model_path: ${oc.env:BAGEL_PATH,ByteDance-Seed/BAGEL-7B-MoT}
# BAGEL single-stage editing modality (registers BagelIt2iAdapter + boots
# stage_configs/bagel_t2i_rl.yaml — one YAML serves both image-out modalities —
# deploy_configs/bagel_t2i_rl.yaml — one YAML serves both image-out modalities —
# with RLBagelPipeline).
modality: bagel_it2i
# Required for colocate: the trainer calls sleep()/wake_up() around the train
Expand Down
4 changes: 2 additions & 2 deletions examples/diffusion/bagel/bagel_it2i_vllmomni.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -134,7 +134,7 @@ backend:
total_steps: 100000
lora_cfg:
_target_: unirl.train.configs.LoraConfig
rank: 64 # must be <= the stage YAML's max_lora_rank (64)
rank: 64 # must be <= the deploy config's max_lora_rank (64)
alpha: 128
dropout: 0.0
bias: none
Expand All @@ -159,7 +159,7 @@ rollout:
# Required; same checkpoint the bundle loads.
model_path: ${oc.env:BAGEL_PATH,ByteDance-Seed/BAGEL-7B-MoT}
# BAGEL single-stage editing modality (registers BagelIt2iAdapter + boots
# stage_configs/bagel_t2i_rl.yaml — one YAML serves both image-out modalities —
# deploy_configs/bagel_t2i_rl.yaml — one YAML serves both image-out modalities —
# with RLBagelPipeline).
modality: bagel_it2i
# Required for colocate: the trainer calls sleep()/wake_up() around the train
Expand Down
2 changes: 1 addition & 1 deletion examples/diffusion/bagel/bagel_vllmomni.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -119,7 +119,7 @@ rollout:
# Required; same checkpoint the bundle loads.
model_path: ${oc.env:BAGEL_PATH,/root/hf_model/BAGEL-7B-MoT}
# BAGEL single-stage T2I diffusion modality (registers BagelT2iAdapter +
# boots stage_configs/bagel_t2i_rl.yaml with RLBagelPipeline).
# boots deploy_configs/bagel_t2i_rl.yaml with RLBagelPipeline).
modality: bagel_t2i
# Required for colocate: trainer calls sleep()/wake_up() around the train
# phase, a no-op unless CuMemAllocator is enabled at vLLM-Omni init time.
Expand Down
2 changes: 1 addition & 1 deletion examples/diffusion/bagel/bagel_vllmomni_async.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -108,7 +108,7 @@ rollout:
# Required; same checkpoint the bundle loads.
model_path: ${oc.env:BAGEL_PATH,/root/hf_model/BAGEL-7B-MoT}
# BAGEL single-stage T2I diffusion modality (registers BagelT2iAdapter +
# boots stage_configs/bagel_t2i_rl.yaml with RLBagelPipeline).
# boots deploy_configs/bagel_t2i_rl.yaml with RLBagelPipeline).
modality: bagel_t2i
# Separate slabs don't time-share GPUs, so sleep/wake is unnecessary.
enable_sleep_mode: false
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@
# sync (sd3/qwen v2 use it). Pushes the trained LoRA adapter into the
# co-located sibling engine in-process; the engine runs base + adapter,
# which is mathematically the merged model the separate recipe pushed.
# hv15's stage config sets enable_lora/max_lora_rank=64 so the adapter
# hv15's deploy config sets enable_lora/max_lora_rank=64 so the adapter
# actually applies. This avoids the CUDA-IPC path's SGLang dependency,
# which the vllm-omni-only venv (two-venv image) does not provide.
#
Expand Down
4 changes: 2 additions & 2 deletions examples/diffusion/sd3/sd3_vllmomni_lora_separate.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
# - LoRA training (lora_cfg kept), pushed as the bare ADAPTER each sync
# (RemoteLoraWeightSync ships lora_A/lora_B to the engine's set_lora_from_tensors).
# - No NCCL group / merged-model broadcast: rank 0 ships the few-MB adapter over a
# plain Ray RPC to each rollout Worker. The vLLM-Omni sd35_t2i stage config
# plain Ray RPC to each rollout Worker. The vLLM-Omni sd35_t2i deploy config
# enables LoRA (enable_lora: true, max_lora_rank: 32), so the engine accepts it.

num_devices: 8
Expand Down Expand Up @@ -87,7 +87,7 @@ rollout:
config:
_target_: unirl.rollout.engine.vllm_omni.config.VLLMOmniEngineConfig
model_path: ${oc.env:PRETRAINED_MODEL,stabilityai/stable-diffusion-3.5-medium}
# sd35_t2i stage config enables LoRA (enable_lora: true, max_lora_rank: 32),
# sd35_t2i deploy config enables LoRA (enable_lora: true, max_lora_rank: 32),
# so the engine accepts the adapter pushed by RemoteLoraWeightSync.
modality: sd3_t2i
# Separate slabs do not time-share GPUs, so sleep/wake is unnecessary.
Expand Down
2 changes: 1 addition & 1 deletion examples/unified_model/hi3_vllmomni.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -118,7 +118,7 @@ backend:
# texts per prompt; the DiT engine (modality dit_recaption, single diffusion
# stage, GPUs 4-7) renders M distinct-noise images per recaption. UnifiedModelTrainer
# wires both with remote() in the shared placement; each engine clears
# CUDA_VISIBLE_DEVICES for its multi-GPU HI3 modality and its stage YAML's
# CUDA_VISIBLE_DEVICES for its multi-GPU HI3 modality and its deploy config's
# runtime.devices pins the physical cards (a real partition, NOT anchor+pop).
# Both run enable_sleep_mode so the trainer can sleep/wake them around the
# colocate train phase. The two engines share ONE backbone/LoRA via sync.
Expand Down
2 changes: 1 addition & 1 deletion examples/unified_model/hi3_vllmomni_veomni_ep.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -124,7 +124,7 @@ backend:
# texts per prompt; the DiT engine (modality dit_recaption, single diffusion
# stage, GPUs 4-7) renders M distinct-noise images per recaption. UnifiedModelTrainer
# wires both with remote() in the shared placement; each engine clears
# CUDA_VISIBLE_DEVICES for its multi-GPU HI3 modality and its stage YAML's
# CUDA_VISIBLE_DEVICES for its multi-GPU HI3 modality and its deploy config's
# runtime.devices pins the physical cards (a real partition, NOT anchor+pop).
# Both run enable_sleep_mode so the trainer can sleep/wake them around the
# colocate train phase. The two engines share ONE backbone/LoRA via sync.
Expand Down
26 changes: 12 additions & 14 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -41,10 +41,9 @@ dependencies = [
"tensordict>=0.5",
]

# Parent/inline engines load this. Multi-stage StageDiffusionProc does not —
# VLLMOmniHijack.hijack() reinstalls the same hook in spawn children.
# vllm-omni 0.28 loads this in parent and spawned stage processes.
[project.entry-points."vllm_omni.general_plugins"]
unirl_capture_flush = "unirl.rollout.engine.vllm_omni.plugin:register_capture_flush"
unirl_runtime = "unirl.rollout.engine.vllm_omni.plugin:register_unirl_runtime"

[project.optional-dependencies]
# sglang pinned to the upstream release the UniRL sglang patch package
Expand Down Expand Up @@ -83,10 +82,12 @@ sglang = [
]
# CUDA 13: vllm >=0.26 wants a bare `torch==2.13.0` whose wheel is cu130, and a
# +cu129 pin would resolve anyway, linking a CUDA-12 torch under a CUDA-13 vllm.
# Prerelease, so installs need --prerelease=allow, plus sglang's compat layer.
vllm = [
"vllm==0.27.0 ; sys_platform == 'linux'",
"vllm-omni==0.27.0rc1 ; sys_platform == 'linux'",
"vllm==0.28.0 ; sys_platform == 'linux'",
"vllm-omni==0.28.0 ; sys_platform == 'linux'",
# vllm-omni 0.28 needs >=5.10.1,<5.15 (5.10.0 is yanked, 5.15 has a known
# construction regression); exact because there is no uv.lock to hold it.
"transformers==5.12.1 ; sys_platform == 'linux'",
"torch==2.13.0+cu130 ; sys_platform == 'linux'",
"torchvision==0.28.0+cu130 ; sys_platform == 'linux'",
# torchaudio is not required directly by the vLLM backend, so this extra
Expand Down Expand Up @@ -169,17 +170,14 @@ index-strategy = "unsafe-best-match"
# transformers 5.12.1 accepts tokenizers<=0.23.0, but tokenizers 0.23.0(rc)
# dropped the `cls` kwarg from RobertaProcessing.__new__, so deserializing a
# saved CLIP tokenizer.json raises TypeError at rollout init (the SD3 text
# encoders, e2e-confirmed on H20). The sglang install's --prerelease=allow flag
# (see INSTALL.md) otherwise selects 0.23.0rc*; this override outranks it AND
# guards against 0.23.0 shipping stable. tokenizers 0.22.2 verified on both
# engine venvs.
# encoders, e2e-confirmed on H20). Keep both engine stacks on the verified
# tokenizers 0.22 family.
override-dependencies = [
"kernels>=0.14.1,<0.15",
"tokenizers>=0.22,<0.23",
# sglang[diffusion] hard-pins diffusers==0.37.0 (still pinned on sglang main),
# colliding with the base diffusers>=0.38.0 floor (#91); override to the floor
# so ".[sglang,...]" stays solvable (same tactic as kernels/tokenizers above).
"diffusers>=0.38.0",
# Override sglang[diffusion]'s stale 0.37 pin with the exact version required
# by vllm-omni 0.28.0. Both mutually exclusive engine stacks use this version.
"diffusers==0.40.0",
]
environments = [
"sys_platform == 'linux' and platform_machine == 'x86_64'",
Expand Down
4 changes: 2 additions & 2 deletions unirl/rollout/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,7 +90,7 @@ Engine dirs use two layouts. The compact engines (`trainside`, `fastvideo`,
management), `utils/`, `weight_sync.py`, and a
runtime-patch dir for the pinned upstream (`sglang_diffusion/_patches/`,
`vllm_omni/patches/`). `vllm_omni` additionally carries worker-subprocess code
(`pipelines/`, `worker/`) and stage boot configs (`stage_configs/`).
(`pipelines/`, `worker/`) and deployment configs (`deploy_configs/`).

Model onboarding is per-engine, and the adapter file is usually **not** the whole
change surface:
Expand All @@ -107,7 +107,7 @@ change surface:
model needs a new upstream patch.
- **`vllm_omni`:** add an `adapters/<family>.py` binder (keyed by modality),
register it, import it in `adapters/__init__.py`, and add the appropriate boot
YAML under `stage_configs/`. DiT families additionally need a worker-side
YAML under `deploy_configs/`. DiT families additionally need a worker-side
`pipelines/<model>/pipeline.py`; if the AR/DiT worker needs new behavior, add a
`worker/` extension or `patches/compat_<model>.py`.

Expand Down
2 changes: 1 addition & 1 deletion unirl/rollout/engine/vllm/runtime.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@
from typing import Any, Dict, List

_PROTOCOL_VERSION = 1
_SUPPORTED_VLLM_VERSION = "0.27.0"
_SUPPORTED_VLLM_VERSION = "0.28.0"


class _ProtocolError(RuntimeError):
Expand Down
2 changes: 1 addition & 1 deletion unirl/rollout/engine/vllm_omni/adapters/bagel.py
Original file line number Diff line number Diff line change
Expand Up @@ -245,7 +245,7 @@ def build_conditions(self, sample: Sample, per_request: List[List[OmniRawResult]
class BagelAdapter(ModelAdapter):
"""Bind BAGEL t2i and it2i to one single-stage DiT worker."""

stage_yaml = "bagel_t2i_rl.yaml"
deploy_config = "bagel_t2i_rl.yaml"
omni_mode = "text-to-image"
needs_driver_tokenizer = False
image_input: bool = False # Whether the modality requires an edit-source image.
Expand Down
6 changes: 3 additions & 3 deletions unirl/rollout/engine/vllm_omni/adapters/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ class ModelAdapter(ABC):

modality: str = ""

stage_yaml: str = ""
deploy_config: str = ""
omni_mode: Optional[str] = None
needs_sigmas: bool = True
needs_driver_tokenizer: bool = True
Expand Down Expand Up @@ -75,9 +75,9 @@ def resolve_sde_label(strategy: Any) -> Optional[str]:

def boot_kwargs(self) -> Dict[str, Any]:
"""Model-specific boot intent beyond the generic config spelling."""
require(bool(self.stage_yaml), f"{type(self).__name__} must set stage_yaml")
require(bool(self.deploy_config), f"{type(self).__name__} must set deploy_config")
kwargs: Dict[str, Any] = {
"stage_yaml": self.stage_yaml,
"deploy_config": self.deploy_config,
"needs_driver_tokenizer": bool(self.needs_driver_tokenizer),
"clear_cuda_visible": bool(self.clear_cuda_visible),
}
Expand Down
Loading
Loading