[NPU] Sync Ascend with main and consolidate NPU adaptations - #413
Meihan-chen wants to merge 70 commits into
Conversation
…llm-project#238) * refactor(test): align CPU unit test layout with slime and add _unit_stubs Move tests/unit/* into tests/ and tests/utils/, replace nested conftest stubs with shared helpers for optional deps, and register the moved files in cpu-unittest CI. Write generated workflow YAML as UTF-8 on Windows. Signed-off-by: kaiyuan <kyxiezju@163.com> * fix(vllm_engine): guard HTTPError.add_note for Python 3.10 CI Exception.add_note() requires Python 3.11+; cpu-unittest runs on 3.10. Signed-off-by: kaiyuan <kyxiezju@163.com> * refactor(test): drop pytest pythonpath, bootstrap tests/ in test files Remove redundant pythonpath from pyproject.toml; each test that imports _unit_stubs prepends tests/ to sys.path (including test_vllm_rollout.py). Signed-off-by: kaiyuan <kyxiezju@163.com> * ci: bump cpu-unittest to Python 3.11 and revert add_note guard Run cpu-unittest on 3.11 so HTTPError.add_note is available in CI without Py3.10 guards in vllm_engine or conditional __notes__ assertions in tests. Signed-off-by: kaiyuan <kyxiezju@163.com> --------- Signed-off-by: kaiyuan <kyxiezju@163.com>
* docs: align vllm docs with current defaults Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: fix router policy examples Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: simplify ray multi-node wording Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: restore quick start hardware notes Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: clarify cache-aware routing support Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: remove extra router policy example Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: restore cache aware policy example Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: restore hardware support wording Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: remove conda fallback note Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: update docs logo Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: replace legacy logo assets Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: add README docs badges Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: replace logo image Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: remove redundant logo jpg Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: adjust hardware support wording Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: refine hardware support wording Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: restore hardware support notes Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: add GB support guard notes Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: clarify profiling flow Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: shorten profiling note Signed-off-by: aoshen02 <aoshen@inferact.ai> * tools: simplify rollout profiler helper Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: restore profiling request range Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: remove profiling note Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: switch logo asset to jpg Signed-off-by: aoshen02 <aoshen@inferact.ai> * update docs logo Signed-off-by: aoshen02 <aoshen@inferact.ai> * update docs logo asset Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: move eplb note to faq Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: rephrase profiling wording Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: remove eplb note Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: restore eplb example Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: restore eplb example in english Signed-off-by: aoshen02 <aoshen@inferact.ai> --------- Signed-off-by: aoshen02 <aoshen@inferact.ai> docs: align vLLM docs with current defaults CI: merged with buildkite/vime-ci failure (infra); GA checks passed. See PR vllm-project#240 comment.
… extension (vllm-project#246) * refactor(weight-sync): use vLLM native /update_weights instead of worker extension Remove the custom `vLLMColocateWorkerExtension` and `_VLLMHijack` monkey-patch introduced in PR vllm-project#104. These reimplemented vLLM's built-in `IPCWeightTransferEngine.receive_weights()` (UUID routing, device_index remapping, layerwise reload) inside a worker extension, bypassing the native `/update_weights` endpoint in favor of `/collective_rpc`. vLLM's native IPC engine already handles everything the extension did. This change routes `update_weights_from_tensor` back through `POST /update_weights` (the path that was working before PR vllm-project#104) and deletes ~160 lines of redundant code. Changes: - `update_weight_from_tensor.py`: delete `_VLLMHijack`, `vLLMColocateWorkerExtension`, and fix `ipc_handles` format to store `ipc_args` (not `(rebuild_func, ipc_args)`) for native API compat - `vllm_engine.py`: revert `update_weights_from_tensor` from `/collective_rpc` to `/update_weights`, delete `update_weights_chunk`, remove `--worker-extension-cls` injection Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(reloadable-pg): add missing gather_object patch; switch IPC gather to match slime ReloadableProcessGroup monkey-patched all_gather_object but not gather_object, causing "Group is not registered" after Megatron offload/reload. This forced _send_to_colocated_engine to use the less efficient all_gather_object (all ranks receive, only leader uses). Fix: patch gather_object in the same list, then switch to dist.gather_object (only leader allocates the receive buffer) to align with slime's implementation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…move PR test on GHA (vllm-project#239) * ci: add Buildkite pipeline for always-on CPU jobs Port the always-on jobs from .github/workflows/pr-test.yml.j2 (pre-commit gate, plugin contracts, agent adapter, in-image unit tests) to a single dynamically generated Buildkite pipeline targeting the vLLM elastic-stack CPU queues. GitHub Actions keeps running in parallel and stays authoritative; GPU suites are not migrated yet. https://claude.ai/code/session_01BSqHKH1FafdRobA7ZeRzBQ Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci: replace Buildkite generator with static pipeline.yml Drop generate_pipeline.py in favor of a plain static .buildkite/pipeline.yml defining the four always-on CPU steps directly (pre-commit gate, plugin contracts, agent adapter, in-image unit tests). Simpler to read and review for a first cut; the GHA workflow stays authoritative and GPU suites are still out of scope. Pass GIT_CONFIG_PARAMETERS into every container so git (in pre-commit) doesn't abort with "dubious ownership" on the host-owned checkout, and fix the depends_on typo in the README. https://claude.ai/code/session_01BSqHKH1FafdRobA7ZeRzBQ Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci(buildkite): fix CPU test deps; add manual gate for GPU suites plugin-contracts failed in build vllm-project#3 on tests/utils/test_hf_checkpoint_saver.py (ModuleNotFoundError: safetensors): the dep list predated the slime sync in vllm-project#232 which added requests/ray/safetensors to the GHA template. Mirror it. GitHub PR labels can't trigger Buildkite jobs, so expose the run-ci-* GPU suites behind a block step instead: unblocking offers a multi-select of suites (short / vllm-config / megatron / precision / ckpt) and gpu_suites.py uploads one step per test with the same gpu_lock_exec + docker invocations and per-test DEEPEP/FP8/EVAL env combos as the GHA jobs. blocked_state: passed keeps the commit status green when the gate is left untouched. GPU steps target a new vime-gpu agent queue (self-hosted hosts; see README). https://claude.ai/code/session_01BSqHKH1FafdRobA7ZeRzBQ Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci(buildkite): run GPU suites on mithril-h100-pool; pin gloo to loopback Build vllm-project#4's unblock test showed the CI cluster rejects uploads targeting a nonexistent queue, and rather than minting a new queue, follow the pattern vllm-omni already uses for mithril-h100-pool: each GPU job is a Kubernetes pod (agent-stack-k8s kubernetes plugin) on an H100 SXM node with nvidia.com/gpu limits (4 or 8), memory-backed /dev/shm, and /mnt/hf-cache mounted as HF_HOME. vime tests hf-download their models, so the warm HF cache replaces the GHA runners' /mnt/nvme0n1/vime_ci mounts; the docker-run wrapper goes away since the pod runs the vime CI image directly. Also pin GLOO/TP_SOCKET_IFNAME=lo in the plugin-contracts container: test_metric_report_dist hung intermittently (build vllm-project#4 timed out at 30 min) because gloo can pick a non-loopback interface inside a bridge-network container. https://claude.ai/code/session_01BSqHKH1FafdRobA7ZeRzBQ Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci(buildkite): expandable_segments for the borderline OOM short test test_qwen3.5_0.8B_gsm8k_async_short OOMed in compute_log_probs on the mithril pool's 80 GB H100s (build vllm-project#6) with 7 GiB reserved-but-unallocated — the allocator-fragmentation case PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True exists for. Scope it to this test's pod only (vLLM sleep-mode CuMemAllocator can conflict with expandable segments) via verbatim pass-through of non-VIME env overrides. The other short tests passed on H100 pods unchanged. https://claude.ai/code/session_01BSqHKH1FafdRobA7ZeRzBQ Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci(buildkite): soft-fail the two known-H100-incompatible GPU tests Builds vllm-project#6/vllm-project#7 isolated two test-level failures on the mithril 80 GB H100s, neither a pipeline issue: - gsm8k_async_short OOMs as tuned (67 GiB live on the actor GPU after expandable_segments eliminated fragmentation; its sync twin passes). - parallel_check's CP=2 grad norm diverges ~4% from the same-node baseline recording, a topology-sensitive numerical invariance question. Mark exactly these two soft_fail so they keep running and stay visible on Buildkite without failing the build; their authoritative gate remains the GHA label jobs. https://claude.ai/code/session_01BSqHKH1FafdRobA7ZeRzBQ Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci(buildkite): keep the two H100-incompatible GPU tests failing loudly Revert the soft_fail: per review, the gsm8k_async_short OOM and the parallel_check CP-invariance divergence should stay visible as hard failures on Buildkite until the underlying issues are fixed. Keep the diagnostic comments and the test-scoped expandable_segments setting. https://claude.ai/code/session_01BSqHKH1FafdRobA7ZeRzBQ Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci(buildkite): soft-fail the two H100-incompatible GPU tests after all Re-apply b334784 (reverted in 0a98010): per the follow-up decision, mark gsm8k_async_short and parallel_check soft_fail so they keep running visibly on mithril without failing the build, with the GHA label jobs as their authoritative gate until the OOM tuning and CP-invariance questions are resolved. https://claude.ai/code/session_01BSqHKH1FafdRobA7ZeRzBQ Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(ci): resolve 0.8B async OOM on H100 by reducing max-tokens-per-gpu Root cause: Qwen3.5's 248K vocab produces [T, 248320] fp32 logits tensors. calculate_log_probs_and_entropy holds 5 copies simultaneously (2 clones + 2 intermediates + original). At max-tokens-per-gpu=9216, each copy is ~8.5 GB → 42.6 GB from logits alone, exceeding H100 80 GB with activations and reserved pool fragmentation. Fix: reduce max-tokens-per-gpu from 9216 to 2048. Peak drops from 117.6 GB to 39.6 GB (measured on H200), well within H100's 80 GB. GSM8K's longest sequence is ~1200 tokens, so 2048 still fits all samples. Also removes gsm8k_async_short from SOFT_FAIL_ON_H100 (no longer needed) and the expandable_segments workaround. parallel_check remains soft-fail: ~11% flake rate on TP4+per-token-loss, confirmed same behavior in slime (Megatron FP reduction-order issue). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * style: format update_weight_from_tensor Signed-off-by: aoshen02 <aoshen@inferact.ai> * remove github workflows Signed-off-by: khluu <khluu000@gmail.com> --------- Signed-off-by: aoshen02 <aoshen@inferact.ai> Signed-off-by: khluu <khluu000@gmail.com> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: aoshen02 <aoshen@inferact.ai>
…t#249) * docs: update docs favicon to Vime branding Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci: fix unit test path in buildkite Signed-off-by: aoshen02 <aoshen@inferact.ai> * test: force triton stub in CPU unit tests Signed-off-by: aoshen02 <aoshen@inferact.ai> * test: align update_weights_from_tensor with native vllm Signed-off-by: aoshen02 <aoshen@inferact.ai> * test: isolate deep_gemm and triton in cpu utils Signed-off-by: aoshen02 <aoshen@inferact.ai> --------- Signed-off-by: aoshen02 <aoshen@inferact.ai>
* fix(arguments): include DP in TP auto-compute default The `vllm_parse_args` default for `vllm_tensor_parallel_size` was computed as `rollout_num_gpus_per_engine // pp_size`, missing the `dp_size` divisor. This made the default TP too large when `--vllm-data-parallel-size > 1` (required for expert parallelism). Align the formula with `_resolve_vllm_parallel_sizes` in `vllm_engine.py`, which correctly uses `gpus // (pp * dp)`. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * test(arguments): add DP>1 cases for TP auto-compute Cover the fix: TP = gpus_per_engine // (PP * DP). - test_parse_args_tp_default_with_dp: 8 GPU, DP=4 → TP=2 - test_parse_args_tp_default_with_pp_and_dp: 8 GPU, PP=2, DP=2 → TP=2 - test_parse_args_tp_default_dp1_unchanged: DP=1 regression guard Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * Update arguments.py Signed-off-by: aoshen02 <aoshen@inferact.ai> * chore: rerun Buildkite Signed-off-by: aoshen02 <aoshen@inferact.ai> --------- Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(tau-bench): add vLLM multi-turn GRPO example with trainable agent Introduce generate_with_tau custom rollout (render + /inference/v1/generate), trainable_agents for tau-bench env interaction, vllm_tool_parser, and TAU_CONFIGS-based user simulator configuration in generate_with_tau.py. Signed-off-by: xky <kyxiezju@163.com> * style(tau-bench): apply pre-commit formatting Signed-off-by: xky <kyxiezju@163.com> * ci: retrigger buildkite Signed-off-by: xky <kyxiezju@163.com> * ci: retrigger buildkite Signed-off-by: xky <kyxiezju@163.com> --------- Signed-off-by: xky <kyxiezju@163.com>
* fix buildkite cpu utils ci Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs: remove gha references from buildkite ci Signed-off-by: aoshen02 <aoshen@inferact.ai> * deps: add transformers and cloudpickle Signed-off-by: aoshen02 <aoshen@inferact.ai> * test: stub optional utils deps in cpu ci Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix: keep vllm http error notes py310 compatible Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci: run cpu jobs on python 3.11 Signed-off-by: aoshen02 <aoshen@inferact.ai> * test: inline utils stubs in affected files Signed-off-by: aoshen02 <aoshen@inferact.ai> --------- Signed-off-by: aoshen02 <aoshen@inferact.ai>
…project#253) * docker: upgrade base to vLLM 0.23.0, remove CUDA 13 build path - Base image: v0.22.0-cu129-ubuntu2404 → v0.23.0-cu129-ubuntu2404 - Remove `ENABLE_CUDA_13` ARG and all conditional cu13 blocks: - cu13 apt dev headers (libcublas-dev-13-0, cuda-nvrtc-dev-13-0, etc.) - TE source build (cu13 wheel didn't exist; cu129 wheel works on arm64) - fzyzcjy triton source build (cu13 specific) - TMS_CUDA_MAJOR export (no longer needed) - Simplify cublas-dev to unconditional libcublas-dev-12-9 - Simplify TE install to wheel-only - justfile: remove `build-cu13` target and cu13 tag scheme - vllm.patch: adapt line numbers for 0.23.0 (776/1896 vs 750/1844), preserve `with self.log_iteration_details(None):` wrapper cu129 nvcc already supports sm100/sm120 (Blackwell), so cu13 build path was unnecessary — it caused build failures on gb300 (cu13 apt packages hijacked /etc/alternatives/cuda, breaking TE CMake). Tested: built successfully on gb200 (arm64), h200 (x86), gb300 (arm64). All three confirmed vLLM 0.23.0. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * docker/patch: replace core.py sleep fix (already in v023) with all2all_utils weight-reload fix (#45989) v0.23.0 already includes the sleep/scheduler guard from #44483, so the core.py patch is no longer needed. Replace it with the FP8+DeepEP weight-reload fix (vllm-project/vllm#45989): snapshot max_num_batched_tokens from FusedMoEConfig instead of calling get_current_vllm_config() during layerwise reload. Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Signed-off-by: Ao Shen <aoshen@inferact.ai> Signed-off-by: aoshen02 <aoshen@inferact.ai> --------- Signed-off-by: aoshen02 <aoshen@inferact.ai> Signed-off-by: Ao Shen <aoshen@inferact.ai> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…ion (vllm-project#264) * fix(vllm_engine): drop router_ip/port fallback to args in init() PD mode passes router_ip=None to engine.init() intentionally — the router launches AFTER engines are ready and collects their URLs. The fallback `self.args.vllm_router_ip` reads a pre-allocated-but-not-yet- listening address, causing ConnectionRefused on POST /workers. Drop the fallback; assign directly like main does. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(arguments): unblock 6 skipped params for --vllm-* forwarding skipped_args in arguments.py blocked 6 params that should be user-configurable via --vllm-* CLI flags: - served_model_name, tokenizer, tokenizer_mode, tokenizer_revision - dtype (only hardcoded when --fp16, otherwise user should set) - tool_call_parser (was manually registered then skipped; let monkey-patch auto-register like slime does) Also removes the redundant manual --vllm-tool-call-parser registration since the monkey-patch now handles it automatically. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-authored-by: Liron Kesem <52330564+LironKesem@users.noreply.github.com>
…llm-project#284) PR vllm-project#264 moved top-level and lazy vllm imports that break the bare python:3.11 CPU CI image (no vllm installed). Three changes — no production code touched: 1. `tests/_unit_stubs.py`: extend `install_vllm_cli_stubs()` with stubs for `vllm.utils.system_utils` (top-level import in vllm_engine.py) and the `vllm.entrypoints.*` hierarchy (lazy imports in arguments.py and vllm_engine._vllm_server_field_names). 2. `tests/utils/test_vllm_engine.py`: add an autouse fixture that patches `_VLLM_SERVER_FIELDS` to an empty frozenset, so tests calling `_compute_server_args` without an explicit monkeypatch don't trigger the real `_vllm_server_field_names()`. 3. `tests/utils/test_vllm_arguments.py`: evaluate `real_module_available` before calling `install_vllm_cli_stubs()` (which registers a fake vllm in sys.modules), then mark two tests that require real FrontendArgs field names with `@requires_vllm` so they skip on the CPU image. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
…ect#257) * fix(data): reuse stored multimodal_inputs in length filter filter_long_prompt re-extracted vision info from sample.prompt via process_vision_info in the multimodal branch. When apply_chat_template is set, sample.prompt is the rendered *string* (not a conversation list), so process_vision_info -> qwen_vl_utils crashed with "TypeError: string indices must be integers, not 'str'". This made prompt-length filtering unusable for any VLM dataset: setting --rollout-max-context-len (which derives rollout_max_prompt_len) or --rollout-max-prompt-len / --eval-max-prompt-len activates the filter and hits the crash. Reuse the multimodal inputs already computed during dataset construction (sample.multimodal_inputs) instead of recomputing them from the string prompt. Add CPU unit tests covering the multimodal branch and a mixed text-only + multimodal dataset. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: Meihan-chen <zr010426ztt@outlook.com> * Delete tests/test_filter_long_prompt_multimodal.py Signed-off-by: aoshen02 <aoshen@inferact.ai> * Update data.py Signed-off-by: aoshen02 <aoshen@inferact.ai> --------- Signed-off-by: Meihan-chen <zr010426ztt@outlook.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: aoshen02 <aoshen@inferact.ai>
…llm-project#283) Replace all references to `inferactinc/public:vime-*` with `vllm/vime:*` across CI, docs, and the release justfile. The image is now published under the official vllm DockerHub namespace (`vllm/vime:latest`, `vllm/vime:test-latest`) as a multi-arch manifest (amd64 + arm64). Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
…llm-project#280) tau-bench RunConfig defines agent_strategy, not agent. The old key was silently ignored by Pydantic, and resolve_tau_config accessed tau_config.agent which raises AttributeError at rollout time. Signed-off-by: kaiyuan <kyxiezju@163.com>
… 20B support (vllm-project#260) * restore: bring back deleted examples, scripts, and agent doc Restore files that were either deleted by vllm-project#126 ("trim examples to qwen3 only") or never synced from slime: **Reverted from pre-vllm-project#126 (translated):** - scripts/low_precision/run-qwen3-4b-fp8.sh - scripts/low_precision/run-qwen3-30b-a3b-fp8.sh - scripts/run-glm4-9B.sh - scripts/run-moonlight-16B-A3B.sh - scripts/run-qwen3-4B-base-sft.sh - scripts/run-qwen3-32B.sh - scripts/run-qwen3.5-35B-A3B-sft.sh **New from slime@44d29ee (translated):** - docs/en/get_started/agent.md - examples/fully_async/run-qwen2.5-0.5B-fully_async.sh All sglang engine flags translated to vllm equivalents (§2.4). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: unify pkill pattern to '[v]llm serve|VLL[M]::' Standardize all scripts to use the bracket-escaped pkill pattern that avoids matching pkill itself and also catches vLLM's renamed subprocesses (VLLM::EngineCore, VLLM::Worker_TP*). Matches the canonical pattern in command_utils.py. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * scripts: complete slime-exact translation of all 29 run scripts Translate all slime scripts to vime following SGLANG_TO_VLLM_TRANSLATION.md: - sglang→vllm prefix swap for CLI flags and variables - _slime→_vime for checkpoint paths - EP: --sglang-ep-size N → --vllm-enable-expert-parallel (boolean) - Speculative: multi-param → --vllm-speculative-config JSON (§5.2) - Delete genuinely sglang-coupled params (DP-attention, DeepEP, NSA, etc.) - flashinfer → FLASHINFER case fix (§2.4) 23 new scripts + 6 existing updated to match slime@cutoff. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(scripts): correct model config source path in FP8 low_precision scripts The FP8 scripts used `${SCRIPT_DIR}/../scripts/models/` which resolves to `scripts/scripts/models/` (non-existent). Changed to `../models/` to match the INT4 scripts. Same fix as slime PR #2094. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gpt-oss): fused BF16 format, bridge API patch, bshd qkv format Three fixes needed to run GPT-OSS 20B RLHF on vLLM backend: 1. hf_weight_iterator_bridge: match Megatron-Bridge 0.5.0 API _patch_bridge_expert_cache_to_cpu monkey-patches GPTOSSBridge. maybe_modify_converted_hf_weight gained a 4th `hf_state_dict` parameter; the patched wrapper only accepted 3, causing TypeError during weight sync. 2. run-gpt-oss-20B: point --hf-checkpoint at fused BF16 format vLLM's _load_weights_other expects gate_up_proj [E, hidden, 2*ffn] (fused). The old per-expert split format (experts.{e}.gate_proj.weight) causes KeyError on bias loading. Use tools/convert_gpt_oss_to_fused.py to convert an existing per-expert checkpoint, or re-run preprocess_gpt_oss.py to produce fused format directly. 3. run-gpt-oss-20B: add --qkv-format bshd + fix seq-length GPT-OSS uses learnable softmax (sink attention). TransformerEngine disables all attention backends when softmax_type=learnable and qkv_format=thd (packed sequences). --qkv-format bshd avoids this. --use-dynamic-batch-size is incompatible with bshd; replaced with fixed --seq-length 10240 (covers 8192 max response + prompt headroom). tools/convert_gpt_oss_to_fused.py: new tool to convert per-expert BF16 checkpoint (output of old preprocess_gpt_oss.py) to the fused HF format expected by vLLM without re-running the slow MXFP4 dequantization. Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com> * fix(scripts): replace pkill -9 vllm with precise -f pattern (21 files) pkill -9 vllm matches any process named "vllm" and can inadvertently kill unrelated vllm processes (e.g. background services). Use the same pattern as PR vllm-project#220 which targets only vllm serve and Ray VLL[M]:: actors: pkill -9 -f '[v]llm serve|VLL[M]::' Also updates the inline form used in multi-node SSH worker restart commands (run-qwen3-235B-A22B*.sh, run-qwen3.5-27B.sh, etc.). Skipped: scripts/run-gpt-oss-20B.sh (uses pkill -9 -f "vllm serve" already), scripts/run-minimax-m2.sh and run-glm4.7-*.sh (already used -f "vllm serve"). Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com> * chore(scripts): remove run-qwen3-4B-amd.sh from this PR AMD-specific script is out of scope for the gb300-complete-port PR. Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com> * sync(docs+scripts): port docs/examples from slime-44d29ee, fix script translations - Add missing EN/ZH docs: low-precision, on-policy-distillation, get_started/agent, pd-disaggregation (heterogeneous server groups fix), examples zh docs - Add missing examples: on_policy_distillation, eval_multi_task, delta_weight_sync, geo3k images - Fix vLLM flag translations across all example docs: - --vllm-mem-fraction-static → --vllm-gpu-memory-utilization - Remove non-existent dp-attention flags (--vllm-enable-dp-attention, --vllm-dp-size, --vllm-moe-dense-tp-size, --vllm-enable-dp-lm-head, --vllm-ep-size) - --vllm-ep-num-redundant-experts → --vllm-eplb-config - --vllm-cuda-graph-bs → --vllm-max-cudagraph-capture-size - sglang speculative flags → --vllm-speculative-config JSON - GLM-4.7 MTP: method=eagle → method=mtp, num_speculative_tokens=4 → 3 - sgl-router → vllm-router; THUDM/vime → vllm-project/vime - Fix scripts: restore run-kimi-k2-Instruct/Thinking/qwen3-4B/qwen3-235B-A22B to slime-44d29ee-as-vime + pkill precision fix only; restore int4 python3 path Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com> * revert(scripts): pkill -9 -f pattern back to pkill -9 vllm, align with slime Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com> * revert(pkill): align all remaining kill patterns with slime (pkill -9 vllm) Covers examples/, docs/, tests/, and vime/utils -- previously missed in the scripts/ revert. Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com> * chore: remove gpt-oss-20B script and convert tool (moved to separate PR) Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…-project#295) The rollout_buffer README (English and Chinese) linked to docs/{en,zh}/models/qwen3-4B.md, but there is no models/ directory — the referenced doc lives under examples/. Point both links to docs/{en,zh}/examples/qwen3-4B.md so the setup instructions resolve. Signed-off-by: ajinkya-metica <ajinkya@metica.com> Co-authored-by: ajinkya-metica <ajinkya@metica.com>
…eview] (vllm-project#286) * sync(slime #2014..#2125): diff3 3-way merge, conflicts preserved Mechanical commit 1 of 2 (per knowledge/rl/slime-to-vime-sync-sop.md §2). diff3 translated 3-way merge on upstream/main (incl vllm-project#260/vllm-project#280/vllm-project#283/vllm-project#257): ours = vime@main, base = translate(slime@#2013), theirs = translate(slime@#2125) Translation fixes vs prior attempt: - casing: SGLang->vLLM (prose) / SGLang<X>->VLLM<X> (identifiers); killed VLlm artifact (was 35 files) - dotted module refs slime.X->vime.X now translated (was leaking 'from slime.backends') These resolved 9 spurious conflicts (46->37 files). Results: 61 clean / 37 conflict (diff3 markers preserved) / 50 new-to-vime / 5 del. Conflict markers use readable -L labels (ours/base/theirs). Resolve in commit 2. Non-conflict provenance fix: vimerl/vime -> vllm/vime in 2 example docs. Engine patch handling (docker/patch/) deferred to commit 2 per SOP §4.5. Signed-off-by: aoshen02 <aoshen@inferact.ai> * sync(slime #2014..#2125): resolve all conflicts (commit 2) Resolved all 37 conflict files / 84 diff3 blocks per agent_run RESOLUTION_POLICY. Principle: keep vime vLLM impl (ours) + incorporate slime's new features (theirs). Highlights: - vLLM API form kept everywhere: /inference/v1/generate, choices parsing, AsyncEngineArgs, vLLM flag names (--vllm-gpu-memory-utilization etc). - Dropped all vllm.srt.* imports (non-existent in real vLLM). - Accepted new slime features: delta-weight-sync CLI args, append_response_tokens (Sample), get_server_info/start_external_rollout_servers/get_rollout_num_engines, old-router(<=0.2.1) compat, TrajectoryManager adapter design (vime already adopted it). - Kept vime-only: --rollout-external, add_router_arguments, _get_metrics_router_addr, reinit_wandb_primary_with_open_metrics, update_tracking_open_metrics, modal sandbox, VIME_AGENT_* env names, local-vLLM tau-bench user sim. - Engine patches (docker/patch/): kept ours vllm.patch (22-line MoE fix), dropped theirs sglang 2674-line content; deleted sglang-only vllm-top_p.patch (per SOP 4.5). - Dockerfile kept ours (vllm/vllm-openai base); version.txt accepted theirs nightly. - run-deepseek-r1.sh: dropped /sgl-workspace dead-path env. Deviations from policy (documented): vllm_rollout.py abort path kept ours pause/drain (abort_servers_until_idle would break partial-rollout drain + leave paused_workers unbound). README ecosystem section left empty (ours) pending de-translation of provenance. All changed .py py_compile clean; zero conflict markers; no sglang/slime leakage. Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): pre-commit lint/format on resolved files - re-add dropped `from vllm_router.launch_router import RouterArgs` import in vime/utils/arguments.py (add_router_arguments uses it; F821 from conflict resolution) - drop unused base_top_p_token_ids/offsets in vllm_streaming_rollout.py (F841; came from theirs but ours's choices-parsing path doesn't use them) - black/isort autoformat (anthropic.py, test_agent/*, arguments.py) - pipeline.yml: agent tests moved to tests/test_agent/; wire new CPU tests pre-commit: all hooks pass (ruff/autoflake/isort/black/yaml). Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): make CPU CI green (engine import, docker build, agent/utils tests) Local CPU CI (pre-commit + plugin + agent + utils) all green on 8xH200 host in python:3.11 containers. Fixes found while running it: - vllm_engine.py: make 'import vllm_router'/'packaging.parse' lazy (inside _register_to_router); top-level import broke CPU import (vllm_router absent in CPU CI). The old-router(<=0.2.1) compat branch is theirs-accepted. - docker/Dockerfile: TMS_CUDA_MAJOR=12 for torch_memory_saver pin (its build backend now requires it for CUDA wheels; base is cu129). Unblocks image build. - agent/adapters/common.py: _run_turn called parse_model_output() without the required tokenizer= kwarg -> 500s in adapter tests. Pass tokenizer=tok. - tests/test_agent/_fakes.py: FakeVLLMServer served sglang /generate + meta_info; retarget to vime /inference/v1/generate + choices shape + x-session-id header. - tests/test_agent/test_adapters.py: parse_model_output(tokenizer=...) + assert vime body keys (token_ids/max_tokens). - tests/utils/test_vllm_config.py: vLLMConfig->VllmConfig (4 sites); fake router returns 3-tuple (ip,port,prom) matching _start_router; drop spurious resolve(). - tests/test_megatron_argument_validation.py: add num_gpus_per_node=8 to the vime_validate_args fixture (vime colocate override needs it). - .buildkite/pipeline.yml: agent tests -> tests/test_agent/*; +cispo_loss, +logprob_response_spans (CPU-safe); test_rollout_metrics stays GPU-only (imports vllm). Engine patch verdict (PATCH_ASSESSMENT.md): P1-P7 vLLM doesn't need (NIXL/Mooncake native); kept ours vllm.patch, dropped sglang content + top_p.patch. Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): unblock GPU run (scipy pin, router arg dup, top-p-replay gate) Found running GPU CI on 8xH200 with vllm/vime:latest: - docker/Dockerfile: pin scipy<1.14 next to numpy<2 (scipy drifted to 1.18 which needs numpy>=2 and uses removed np.long -> 'import vllm' crash). - vllm_utils/arguments.py: drop sync-added RouterArgs.add_cli_args + its import in add_vllm_router_arguments. main exposes the full router surface only in utils.add_router_arguments; the duplicate re-registered --router-request-timeout-secs -> argparse conflict at train startup. - megatron_utils/loss.py: get_rollout_top_p_logprob_kwargs falls back to full-vocab logprob when top-p nucleus token ids are absent instead of raising. slime's top-p-replay needs engine-returned top-p tokens; vime's vLLM /inference/v1/generate does not expose them (sglang-only). Matches vime pre-sync behavior; flagged in OVERNIGHT_REPORT for review. Image import smoke + Megatron ckpt load + 4x VLLMEngine bringup + NCCL weight transfer all confirmed working in-image before this. Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci(gpu): drop deleted test_qwen2.5_0.5B_ppo_critic_only_short from short suite slime deleted tests/test_qwen2.5_0.5B_ppo_critic_only_short.py this window (#2014..#2125); gpu_suites.py still listed it -> 'no such file' exit 2. Critic-only path is still covered by test_qwen3_4B_ppo_train_critic_only (megatron suite). Other 3 short tests (gsm8k_async, gsm8k, fully_async) pass on 8xH200. Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): restore base_log_probs init in streaming rollout (None+list crash) GPU test_qwen3_4B_streaming_partial_rollout hit 'TypeError: unsupported operand +: NoneType and list' at vllm_streaming_rollout.py:234. The conflict resolution changed base_log_probs from main's `list(sample.rollout_log_probs or [])` to a None-able form; a fresh sample (rollout_log_probs=None) then did None + call_log_probs. Restored main's form. Real sync-resolution regression caught by GPU CI. Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): streaming rollout uses append_response_tokens (was renamed update_from_meta_info) slime #2110 renamed Sample.update_from_meta_info -> append_response_tokens; vllm_rollout was updated but vllm_streaming_rollout still called the old name (AttributeError at generate_streaming). Streaming already accumulates tokens incrementally for partial-rollout, so call append_response_tokens(meta_info=meta) with tokens omitted -> metadata-only finalize (no double-append). Caught by GPU test_qwen3_4B_streaming_partial_rollout. Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): restore vime unconditional colocate rollout_num_gpus re-derive 3-way compare (slime / vime-main / PR) showed the merge made a broken hybrid: it KEPT vime's num_gpus_per_node colocate override (which assumes rollout_num_gpus is forced to actor_num_gpus_per_node*actor_num_nodes) but REPLACED vime's unconditional re-derive (`!= -> re-derive`) with slime's `is None`-only form. When a colocate test's rollout_num_gpus is non-None but mismatches, it was left mis-sized -> engine/GPU misplacement -> mixed_offload IPC-UUID mismatch + ckpt 'Free memory < util'. slime passes (no override, self-consistent is-None); vime main passes (override + unconditional re-derive, coupled). Restore vime's re-derive; keep slime's new rollout_num_gpus==0 branch (checked first). Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): slime-consistency review pass (docs, rollout routed_experts, comments) Docs (EN+zh parity): - fault-tolerance: revert junk 'server'->'engine' mistranslation (keep correct /health endpoint, verified vs vllm_engine.py) - vllm-config: 'ServerArgs'->'EngineArgs' (sglang class -> vLLM AsyncEngineArgs u FrontendArgs); fix duplicated 'vllm-router (vllm-router)' alias - customization: restore over-deleted 'custom_generate -> list[Sample]' section + signature (dropped only the vime-absent search-r1 example link) Rollout: - routed_experts now flows through the slime-identical Sample._apply_meta_info (single assignment site, torch.int32 tensor matching downstream) instead of an inline numpy assign; vLLM .npy-on-choice decode stays (engine wire-format delta). Both vllm_rollout and vllm_streaming_rollout. Comments for future syncers: - --opd-teacher-model + on_policy_distillation: engine-driven divergence (vLLM model field; sglang /generate has none) - overrides / _vllm_server_field_names: AsyncEngineArgs u FrontendArgs == slime's sglang ServerArgs Examples/docker/etc: drop vime-absent npu/retool/search-r1/tau-bench files; restore eval_multi_task; docker alignment with slime. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(ci): pre-commit green — define base in streaming MM render (F821) + isort/black on test_agent Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(ci): correct two mistranslated CPU tests (colocate rollout-gpu re-derive; parse_model_output tokenizer) - test_..._preserves_larger_rollout_gpus_under_colocate asserted slime behavior (==12); vime re-derives to actor*nodes=8 under colocate (commit 9701304). Renamed + assert ==8 + divergence note. vime-main never had the test; slime does. - test_parse_model_output_plain_text_no_parsers called parse_model_output without the required tokenizer kwarg (vllm-project#198 made it required for vLLM parsers). Pass tokenizer=None (unused on the no-parser path). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): streaming rollout posts to /inference/v1/generate, not sglang /generate The slime diff3 merge took slime's sglang endpoint (/generate) for the streaming rollout POST instead of keeping vime's vLLM endpoint (/inference/v1/generate). main (7198547) had the correct URL; the sync regressed it (and dropped the base var). Result: 404 Not Found at vllm_streaming_rollout.py:182 -> test_qwen3_4B_streaming_partial_rollout fails. Caught on a clean h200 node. The file's own docstrings already say /inference/v1/generate throughout. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): thread rollout port cursor globally across multi-model engines The diff3 merge adopted slime's deferred rollout-engine init (start_rollout_servers now returns pending_init_handles awaited by the caller instead of ray.get-ing each model's engines before the next). That broke an implicit invariant the per-model port_cursors reset relied on: in main, model 0's engines were fully bound before model 1 allocated ports, so the free-port bind-test in _allocate_rollout_engine_addr_and_ports_normal skipped model 0's ports. With deferred init, model 1 allocates while model 0 is unbound, the bind-test sees the base ports free, and a second model (e.g. mixed_offload's frozen "ref") lands on the same 15000-15003 as the actor. The actor's POST /update_weights to :15002 then hits the never-started ref engine -> vLLM 500 "start_weight_update must be called before update_weights" (test_vllm_config_mixed_offload[_ft]). Fix: initialize port_cursors once before the model loop so the per-node next-free cursor is monotonic across all models, keeping every engine's ports disjoint regardless of bind timing. Single-model behaviour is unchanged; the per-model reset only existed to scope cursors that are already node-keyed. Caught on h200 GPU CI (new nightly-dev-20260618a image, which added the start_weight_update-before-update_weights enforcement that exposed the collision). Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): restore robust pkill pattern so ckpt cleanup kills vLLM ray actors PR vllm-project#260 ("complete slime-exact port") changed execute_train's pre-launch cleanup from pkill -9 -f '[v]llm serve|VLL[M]::' to pkill -9 vllm as an over-literal sglang->vllm translation. But the vLLM rollout engine runs as Ray actor processes whose process *name* is python/ray, with "VLLM::" only in the command line — so `pkill -9 vllm` (name match, no -f) does not kill them. Leftover engine processes from the ckpt test's save phase survive into the load phase, holding ~115 GiB, so the load-phase engine starts with ~24/139 GiB free and dies with "Free memory ... less than desired GPU memory utilization (0.8, 111.84 GiB)" (test_qwen3_4B_ckpt.py, both --async-save and not). Restore the cmdline-match pattern `-f '[v]llm serve|VLL[M]::'`. Bisected on h200: ckpt PASSES at 289ee6d / e62d44f (old pattern, 4/4 runs) and FAILS at 7198547/main + PR (new pattern, 0/3), same old image -> code regression in vllm-project#260. Verified: pkill-fixed PR code + new pr286 image -> ckpt PASS (579s). Signed-off-by: aoshen02 <aoshen@inferact.ai> * chore(sync): replace all `pkill -9 vllm` with cmdline-match pattern Same root cause as 764e1e1 (command_utils.py): `pkill -9 vllm` matches by process *name*, but vLLM rollout engines run as Ray actor processes (python/ray named, "VLLM::" only in the command line), so the name match never kills them. Apply the robust cmdline pattern `pkill -9 -f '[v]llm serve|VLL[M]::'` everywhere the bare `pkill -9 vllm` cleanup was used across run/example scripts, so leftover engines don't squat GPUs across runs. No logic change beyond the kill pattern. Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(docker): restore vLLM core.py partial-wake sleep-guard (#44483) Commit d41f0aa removed the core.py sleep-guard hunk from docker/patch/latest/vllm.patch on the assumption it was "already in v0.23.0". It is not: stock v0.23.0 `vllm/v1/engine/core.py` calls `resume_scheduler()` and `execute_dummy_batch()` even during a partial (weights-only) wake. So a colocate DP+EP pd_mooncake rollout, right after `POST /wake_up?tags=weights` (KV cache still released under level-2 sleep), has its DP busy-loop fire a decode-shaped dummy batch that touches freed KV -> the scheduler_metadata write in flashattn_mla.py:234 (MLA, glm4.7) and flash_attn.py:547 (FA3, qwen3.6) raises `CUDA error: invalid argument`. Restore the guard (`if not self.model_executor.is_sleeping` around resume_scheduler; `if not self.is_sleeping()` around execute_dummy_batch), keeping the all2all_utils weight-reload fix. This is the vllm-project#173 sleep-guard patch re-expressed against v0.23.0 line numbers. Verified: git-apply --check clean against stock v0.23.0; both guards land; glm4.7/qwen3.6 pd_mooncake reproduced the crash without it (the 8-day-old image that still carried the guard passes both). Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(docker): drop scipy<1.14 pin, mirror slime numpy<2 only The scipy<1.14 pin (added in 0cda11c while chasing the pd_mooncake crash) was a red herring: the real cause was the dropped core.py partial-wake sleep-guard, now restored. slime-2125-as-vime pins only `numpy<2`; this restores that exact line. numpy 1.26.4 + scipy resolved naturally matches the working baseline. Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(docker): keep FlashQLA install gated behind INSTALL_FLASHQLA=0 slime installs FlashQLA unconditionally, but it is sm90/Hopper-only. Restore vime's original gated form (default off; --qwen-gdn-backend fla elsewhere). CI build passes --build-arg INSTALL_FLASHQLA=1 to include it. Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs(vllm-config): fix inference-only FAQ — vime launches engines in-process The mechanical mirror of slime's answer steered users to vLLM's standalone `vllm serve` (slime's `launch_server` analog) for inference-only. That is misleading for vime: like slime, vime launches the vLLM engines in-process from `--vllm-config` (same in-process path as training), so a rollout-only run serves directly with no separate server process. Point standalone users to `--rollout-external-engine-addrs` instead. EN + ZH. Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs(debug): restore INT4 / Compressed-Tensors checkpoint section vime's debug.md was missing slime's "INT4 / Compressed-Tensors Quantization Checkpoint Issues" section (slime #1642) — dropped in an earlier sync, not present on main. Restore it (EN + ZH), translated sglang→vLLM / Megatron→vLLM. Covers the quantization_config.ignore list, all-zero MoE router weights (mlp.gate.weight) when mis-quantized, missing safetensors shards, and diagnosis via --check-weight-update-equal / --debug-rollout-only. Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs(vllm-config): use _run_vllm_server for inference-only FAQ Keep slime's wording; the only engine-coupled fix is the standalone launcher name. slime's `launch_server` is its in-process engine entry; vime's analog is `_run_vllm_server` (vllm_engine.py, launched via multiprocessing.Process), not the standalone `vllm serve` CLI. EN + ZH. Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(docker): restore scipy pin (scipy<1.18) — vime base needs it Reverts the scipy-pin removal in b359b24, which was wrong: vime's vllm/vllm-openai base ships no scipy, so unpinned the build pulls scipy>=1.18, which hard-requires numpy>=2 and uses np.long (removed numpy>=1.24) -> crashes against the numpy<2 reinstall (Megatron needs numpy 1.x). slime's sglang base resolves scipy 1.17.1 natively (numpy-1.x compatible), so slime needs no pin; this is a base-image divergence, not a red herring. Pin boundary is 1.18 (slime runs 1.17.1), not the earlier 1.14 over-estimate. Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(scripts): properly translate sglang args in glm5.2-744B + glm4.7-355B-delta These two were prefix-swapped (--sglang-X -> --vllm-X) without semantic mapping, leaving ~20 args that aren't vLLM AsyncEngineArgs (argparse would reject). Apply the knowledge/rl/sglang-to-vllm-translation.md §5.5 mappings: - dp-size->data-parallel-size, ep-size->enable-expert-parallel, max-running-requests-> max-num-seqs, cuda-graph-max-bs->max-cudagraph-capture-size - 5x/4x --speculative-* -> one --vllm-speculative-config JSON (§5.2) - DeepEP: per-group deepep_mode auto/low_latency -> all2all_backend deepep_high_throughput/ low_latency in the --vllm-config overrides (vLLM has no 'auto'; PD encodes it per-role) - watchdog-timeout -> env VLLM_ENGINE_ITERATION_TIMEOUT_S - drop sglang-only: dp-attention / dp-lm-head / moe-dense-tp / disable-overlap-schedule / NSA backends (vLLM selects DeepSeek sparse attn per model) / engine delta-receiver knobs - flag PD mooncake transport (-> --vllm-kv-transfer-config) as fabric-specific TODO These are 744B/355B scripts not runnable in CI — translations are SOP-mapped but hardware-unvalidated (flagged inline). Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(args): hard-guard unverified delta weight-sync mode --update-weight-mode=delta (PR vllm-project#278 lineage) is not yet validated on vime+vLLM (vLLM exposes dense/sparse_flat only, not slime's gap-delta/ zstd encoding). Raise NotImplementedError at arg-validation so it fails fast at startup instead of crashing mid weight-sync. Downstream delta code is kept untouched; remove this raise once a real delta-load run passes. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs,test: correct rollout engine endpoint to /inference/v1/generate The agent rollout engine is reached at vime's vLLM ``/inference/v1/generate`` (see vllm_rollout.get_model_url default + common.call_vllm_generate), not the bare ``/generate`` of sglang. Fix the imprecise path in adapter/test docstrings and comments, and rewrite the vllm-config.md custom-rollout examples (en+zh): they were still sglang-shaped (``/generate`` path + ``{"text":..., "return_logprob": True}`` body). Use vime's real request schema instead -- ``{"model","token_ids", "sampling_params"}`` with ``max_tokens``/``logprobs``, ``prompt_logprobs`` for fixed-sequence scoring, and the ``choices[0]`` response shape. No code/logic change: comments, docstrings, and doc examples only. The Megatron training server's own ``/generate`` endpoint and the sglang citation in arguments.py are correct and left untouched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * style(args): collapse delta-guard message to one line (black) The delta-guard NotImplementedError message was split across two adjacent string literals; black on the CI (line-length 119) collapses/normalizes it. Make it a single clean literal so pre-commit is green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * test(args): assert delta weight-sync is guarded off (not per-condition) The hard delta guard (7bb19e6) raises NotImplementedError at the top of the delta branch, making the downstream colocate / unknown-transport rejections unreachable. Replace test_update_weight_delta_rejects_colocate and test_update_weight_delta_rejects_unknown_transport (whose ValueError paths no longer fire) with a single test_update_weight_delta_disabled that asserts the guard raises for any delta config. Breadcrumb left to restore the per-condition tests when delta is verified and the guard is lifted. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(sync): restore dropped weight-sync metrics chain + delta dispatch (slime parity) A mirror cross-check of the delta-weight-sync surface (slime #1806/#1991) found the #2014..#2125 sync had silently dropped several slime-faithful pieces: * extra_metrics logging chain — slime threads weight-update metrics from the actor through log_perf_data -> log_perf_data_raw. vime dropped the param at all three layers, so weight-update metrics were never logged. Restored: train_metric_utils.log_perf_data_raw(extra_metrics=...), data.log_perf_data passthrough, and actor passing self.weight_updater.pop_metrics(). * pop_metrics on UpdateWeightFromDistributed — the default (non-colocate, nccl) weight_updater. slime gives all three updaters a pop_metrics() stub so the actor can call it uniformly; vime kept it on tensor/disk but dropped it on distributed, which would AttributeError once the actor calls it. Restored the ~5-line stub (delta-specific plumbing stays dropped — vime+vLLM has no DeltaSpec). * actor delta-mode dispatch branch — restores the elif selecting UpdateWeightFromDistributedDelta. Dead code behind the validation guard that rejects --update-weight-mode=delta, so vime mirrors slime with the guard as the single divergence. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * chore(sync): align comments to the mechanical mirror Comment-only pass; no behavior change. Aligns vime comments to what a faithful slime->vime translation would carry: * Strip vime-divergence rationale markers (delta guard, opd-teacher-model, top-p fallback, colocate/delta tests). The rationale belongs in the divergence manifest, not inline; the guarded code + self-explanatory NotImplementedError messages stand on their own. * De-verbose vLLM-specific comments to mirror scale: the router-args block, the AsyncEngineArgs u FrontendArgs docstring, and the MoE-replay / streaming-rollout blocks that slime does not carry at that length. * Restore slime-original comments the sync had dropped or naively translated, with judgment translation of sglang-specific terms: session_id routing ("vLLM router", not the mechanical "Model Gateway"), "Prepare payload for vLLM server", the unique-session_id loop, and the pending-tasks wait. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * docs(delta): note delta weight-sync not yet verified on vime+vLLM (PR vllm-project#286 review) Per review on PR vllm-project#286: delta weight sync is documented here but the arg guard disables --update-weight-mode=delta. Add a top-of-page note (en + zh) so users see it before hitting NotImplementedError at argparse. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> --------- Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ect#291) * Add examples/mem_agent for long-context MemAgent RL Co-authored-by: Cursor Signed-off-by: kaiyuan <kyxiezju@163.com> * Address review: harden mem_agent convert and eval scripts Co-authored-by: Cursor Signed-off-by: kaiyuan <kyxiezju@163.com> --------- Signed-off-by: kaiyuan <kyxiezju@163.com>
…m-project#296) Under --partial-rollout, abort() (pause -> drain -> resume) deadlocks: /pause?mode=abort puts the scheduler in PAUSED_NEW, and a /generate that races in after the pause parks in the waiting queue and never returns until /resume, which runs after the drain. Reordering to pause -> resume -> drain avoids the hang, but resume reopens the whole queue so the long tail runs to COMPLETION -- breaking partial rollout's "truncate the tail, resume it next step" semantics. Switch to a delete-type abort instead: - vLLM: add POST /abort_requests to the RLHF api_router -> EngineClient.abort() (removes queued requests from the waiting queue and finish-aborts running ones, whose partial output returns on the original /generate stream). It does not pause the scheduler, so there is no /resume and no deadlock. Shipped as a build-time patch in docker/patch/latest/vllm.patch. - vime: server_control.abort_inflight_requests() replaces the unused, slime-mirrored abort_servers_until_idle / _v1_loads helper (vLLM has neither /abort_request nor /v1/loads). abort() re-issues the sweep across drain waves and converges on state.pendings, with a timeout bounding how long a late multi-turn straggler can run before being truncated to partial. - vllm_engine: drop the legacy version gate in _register_to_router. vime ships its own vllm-router, so only the /workers payload path is needed. Adds delete-type abort unit tests. AI-assisted change; reviewed by a human before submission. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-authored-by: Josephasafg <ajgard7@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…-project#303) rollout/prefix_cache_hit_rate (and avg_cached_tokens_per_sample) were structurally 0 due to three coupled gaps: 1. vLLM's non-streaming /inference/v1/generate builds a `usage` block but silently drops it: GenerateResponse (v0.23.0 serve/disagg/protocol.py) never declared a `usage` field and has no extra="allow", so pydantic discards it. Patched in docker/patch/latest/vllm.patch, mirroring GenerateStreamResponse. (Upstream fix filed against vllm-project/vllm.) 2. The rollout parser read usage.prompt_tokens/completion_tokens but never usage.prompt_tokens_details.cached_tokens -> PrefixCacheInfo numerator pinned to 0. Now read in both vllm_rollout and vllm_streaming_rollout. 3. Streaming additionally needs stream_options.include_usage=True for vLLM to emit the terminal usage SSE chunk. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…-project#317) * fix(docker): abort-all in the /abort_requests vLLM patch must abort by internal ids The bundled /abort_requests endpoint (merged in vllm-project#296) populated request_ids from output_processor.request_states (internal ids) but called engine.abort() with the default internal=False, so they were treated as external, matched nothing, and POST /abort_requests {} silently aborted no requests under default request-id randomization. Abort the all-in-flight list as internal. Mirrors vllm-project/vllm#47173. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * style: reject malformed JSON in /abort_requests patch with 400 Match the sibling dev endpoints and the Rust frontend (400 on malformed JSON) instead of silently treating it as empty. Mirrors vllm-project/vllm#47173. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(docker): abort-all patch must also abort parallel-sampling parents The /abort_requests patch enumerated request_states (child internal ids only), so with n>1 the ParentRequest entry leaked. Include parent_requests keys in the abort-all set. Mirrors vllm PR #47173. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> * feat(docker): cu13 image variant + TMS cu13 preload Port the cu13 build support from vllm-project#307: ENABLE_CUDA_13 branches the apt dev headers, cublas header, TransformerEngine (source-built for cu13), TMS_CUDA_MAJOR auto-detect, and the cudnn pin. justfile gains a build-cu13 target and a VARIANT-prefixed manifest. actor_group preloads the cu13 TMS .so. Also switch the vLLM patch apply to --allow-empty so the build survives once the patch is emptied upstream. Excludes vllm-project#307's NCCL_CUMEM_ENABLE default flip (0->1) and the glm5.2 scripts by request. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> --------- Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* Added ROCm quick start guide * changed wording * changed docker image * updated instructions * Add initial working ROCm 7.0.2 build (Dockerfile.rocm + patch) Initial working setup for building/running vime on ROCm 7.0.2 (gfx950). * ROCm: re-base Dockerfile on vime mainline Megatron + shared patch * ROCm: apply checkpoint-writer patch in HF->torch_dist converter * ROCm: align async run script cleanup with run-qwen3-4B.sh format * ROCm: move megatron.patch to docker/amd_patch/ (slime layout) * ROCm: re-add fused-kernels init patch and AMD doc/template references --------- Co-authored-by: pancake0003 <146360951+pancake0003@users.noreply.github.com> Co-authored-by: indianspeedster <cp3793@nyu.edu>
…+ rewrite multi-node (vllm-project#322) * [Doc] Fix stale Qwen3-30B-A3B example: dataset download + align with actual script The Qwen3-30B-A3B example documented an elaborate env-var-driven multi-node script (ACTOR_NUM_NODES / MEGATRON_TP / ENABLE_R3 auto-behavior, "script skips Ray", default batch 4/2/8, --vllm-moe-backend triton). None of that exists in the current slime-exact scripts/run-qwen3-30B-A3B.sh — it was orphaned when the script was reverted to the slime-exact single-node port. The real defaults are rollout-batch-size 32 / n-samples 8 / global-batch 256. Changes (zh + en, kept concise): - Environment Preparation: add the concrete `hf download` commands for the model, dapo-math-17k train data, and aime-2024 eval data (previously the section only pointed at the Qwen3-4B doc). - Restore the "BF16 training + FP8 inference" section (present upstream in slime, dropped here); matches the commented --hf-checkpoint line in the script. - Replace the orphaned multi-node section with the concise, accurate slime-style guidance (manual modifications + vLLM EPLB example), fixing the broken zh anchor to the Chinese quick_start heading. Doc now matches scripts/run-qwen3-30B-A3B.sh; no script changes. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * [Doc] Expand Qwen3-30B-A3B multi-node section (code-grounded) Replace the concise multi-node bullets with an actionable but verified guide: Ray cross-node startup (from quick_start), the exact script edits (--actor-num-nodes, colocate auto rollout-num-gpus), parallelism scaling as constraints (defer concrete large-scale ratios to the GLM/DeepSeek examples rather than fabricating a 2-node table), and the real pitfalls we hit (keep each vLLM engine within a node to avoid cross-node TP=16 slowdown/ numerics, MASTER_ADDR / NCCL_SOCKET_IFNAME, global-batch identity, --num-gpus-per-node for <8 GPUs/node). Every claim traces to vime/utils/ arguments.py, quick_start, or observed multi-node runs; no "tested" claims. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * [Doc] Drop model-download line from Qwen3-30B-A3B example Model download is already covered by the referenced Qwen3-4B doc; keep only the train/eval dataset download (dapo-math-17k, aime-2024) surfaced inline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * [Doc] Drop redundant dataset download; keep env-prep slime-exact The intro sentence already covers env/model/data/ckpt via the Qwen3-4B reference, and slime's 30B doc carries no dataset block. Remove it; keep the 30B-specific torchrun checkpoint conversion (differs from 4B's single-process python). Env-prep is now byte-identical to upstream. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * [Doc] Multi-node: also skip the script's Ray-cleanup preamble Codex review (P2): for the manual multi-node path, removing only the `ray start --head` line is not enough — the script's initial cleanup block (`ray stop --force` + `pkill -9 ray/python/redis`) would tear down the manually-started head before `ray job submit`, failing submission and orphaning the workers. Tell users to remove/comment both the cleanup block and the `ray start --head` line. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The page title says "8xH100" but the CPU-Adam bullet said "8xH800 environment", an inconsistency inherited from slime. Both are 80GB cards so the point is unchanged; align the text with the title. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ction (vllm-project#334) Signed-off-by: kaiyuan <kyxiezju@163.com> Co-authored-by: XiaoxinWang <wangxiaoxin7@huawei.com> Co-authored-by: FulinGao_HW <gaofulin1@huawei.com> Co-authored-by: meihanc <zr010426ztt@outlook.com> Co-authored-by: wuxiang <498160096@qq.com> Co-authored-by: yuxinshan <syx_ctyg@126.com> Co-authored-by: flb_dayo <floatlibai@gmail.com> Co-authored-by: Windfeng8 <523758380@qq.com> Co-authored-by: YZY <532183776@qq.com>
Sync slime a897e1f4..680824dd using the mechanical merge plus vLLM adaptation workflow.
Revert the prematurely merged slime sync and restore the pre-vllm-project#338 tree.
…roject#329) * Fix pp for distributed update weights Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com> * Clean Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com> * Add test Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com> * Fix pre-commit Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com> * Add test to CI Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com> * Resolve comment Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com> * Recover pipeline code Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com> * Fix PP sync state handling and test module name Avoid mutating source and group state while staging pipeline weight sync, and rename the E2E test to a valid Python module name. Generated with Codex. Signed-off-by: aoshen02 <aoshen@inferact.ai> * Fix Bridge export with pipeline parallelism Signed-off-by: aoshen02 <aoshen@inferact.ai> * refactor: name persistent weight sync condition Signed-off-by: aoshen02 <aoshen@inferact.ai> * refactor: centralize persistent weight sync policy Signed-off-by: aoshen02 <aoshen@inferact.ai> * refactor: share raw weight sync barriers Signed-off-by: aoshen02 <aoshen@inferact.ai> * test: cover raw and bridge PP weight sync Signed-off-by: aoshen02 <aoshen@inferact.ai> * test: clean up vllm after stopping ray Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix: shut down rollout engines on dispose Signed-off-by: aoshen02 <aoshen@inferact.ai> --------- Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com> Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-authored-by: aoshen02 <aoshen@inferact.ai>
* sync: mechanically merge slime through 680824dd Signed-off-by: aoshen02 <aoshen@inferact.ai> * sync: adapt slime changes for vLLM Signed-off-by: aoshen02 <aoshen@inferact.ai> * ci: register synced CPU and GPU tests Signed-off-by: aoshen02 <aoshen@inferact.ai> * update Signed-off-by: aoshen02 <aoshen@inferact.ai> * update Signed-off-by: aoshen02 <aoshen@inferact.ai> * update Signed-off-by: aoshen02 <aoshen@inferact.ai> * sync: tighten vLLM adaptations after review Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix: preserve vLLM top-p fallback Signed-off-by: aoshen02 <aoshen@inferact.ai> * sync: complete slime adaptation audit Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(ci): update external PD test for native vLLM * fix(ci): use NIXL default UCX device selection * fix(ci): configure Qwen3.6 Mamba state layout Signed-off-by: aoshen02 <aoshen@inferact.ai> * fix(ci): honor explicit NCCL NVLS override * fix(ci): preserve explicit UCX device override --------- Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Use native stateful HCCL/NPU IPC backends and rank-local payloads. Add native Qwen3-VL loading and fix IPC/optimizer memory lifecycles. Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
There was a problem hiding this comment.
Code Review
This pull request introduces significant updates to the vime framework, including the addition of a delta weight synchronization mechanism for training/inference disaggregation, support for agentic RL workflows with new sandbox backends, and improvements to the CI/CD pipeline using Buildkite. My review identified two critical issues: a logic error in the pull_checkpoint version rollback condition and a token duplication/length mismatch error in the MemAgent rollout logic. Both issues require immediate attention to ensure correct checkpoint management and training stability.
| + start -= 1 | ||
| + | ||
| + if applied is None or start > applied: | ||
| + seed_dir = base_dir if start == 0 else _version_dir(source_dir, start) |
There was a problem hiding this comment.
In pull_checkpoint, if target_version < applied (e.g., during a rollback or reset to an older checkpoint), start (which is target_version) will be less than applied. Since start > applied is False, the checkpoint reset logic will be skipped, and the local checkpoint will remain at the newer applied version instead of rolling back to the requested target_version. Changing the condition to start != applied ensures that both forward-resets and backward-rollbacks are correctly handled.
if applied is None or start != applied:
|
|
||
| sample.prompt = final_messages[0]["content"] | ||
| sample.response = _strip_stop_tokens(final_out.response) | ||
| sample.tokens = prompt_token_ids + cat_token_ids |
There was a problem hiding this comment.
In generate, cat_token_ids already starts with the prompt of the first turn (first_out.prompt_token_ids) because turns[0]['tokens'] is constructed as out.prompt_token_ids + out.token_ids. Prepending prompt_token_ids to sample.tokens here duplicates the first prompt. Furthermore, since sample.loss_mask and sample.rollout_log_probs are set directly to cat_loss_mask and cat_log_probs (which do not have prompt_token_ids prepended), this causes a length mismatch between tokens, loss_mask, and log_probs, which will crash the training process. Setting sample.tokens = cat_token_ids resolves both the duplication and the length mismatch.
| sample.tokens = prompt_token_ids + cat_token_ids | |
| sample.tokens = cat_token_ids |
Register the NPU accelerator before MindSpeed bootstrap, preserve engine-group resource hooks, and handle empty optimizer/payload boundaries. Rebase vendor patches onto vLLM e6bfe03a and Ascend fd815467, isolate MRV2 KV allocations across sleep/wake, and disable NZ in the NPU E2E cases. Validation: Qwen3-4B, Qwen3-30B-A3B and Qwen3-VL-8B E2E passed on 2026-09-08. Existing runtime dependencies retained; NPU MTP, Omni and pull weights remain deferred. Original vendor IPC test fixture incompatibilities remain documented. Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Record the resolved upstream/ascend merge from the existing index. Keep working-tree adaptation in a separate follow-up commit. Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Record current MegatronAdaptor, patch-ordering, native VL and Qwen3.5 GDN adaptation with their tests. Keep torch_dist/ref-load opt-in and remove its temporary output-directory override as requested. Qwen3.5 serving/FLA environment isolation remains a follow-up working-tree change; this snapshot does not claim a passing Qwen3.5 E2E. Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Preserve the current S7 Qwen3.5 serving diagnostics: isolate training FLA libraries before serving startup and select spawn for vLLM TP workers. CPU contracts pass (132 tests); the latest E2E passes startup and reaches rollout but is blocked by stale vLLM-Ascend conv1d binary schemas. Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Revert f5b8491 (vllm-project#409) and the associated NPU adaptations in 7a27808 and 81ae79d. Defer Qwen3.5 validation to the next stage with a fresh, matched serving and training environment. Remove the NPU-only GDN interfaces, FLA bootstrap and serving isolation, forced spawn, dedicated scripts/tests/CI entry, and FLA build recipe. Restore the pre-409 Bridge pin and patch while preserving the common TransformerLayer implementation. Retain S7 vllm-project#385 training-stack migration, vllm-project#396 torch_dist/ref-load, Qwen3-VL fixes, and main Qwen3.5 model code. Serving patches and latest patches are unchanged. No installed environment is rolled back. Validation: 183 grouped CPU tests passed; Ruff and runner shell syntax passed; common-to-NPU Megatron and reverted Bridge patches pass clean-base apply checks. Existing basic E2E and torch_dist PASS logs are retained; no fresh post-revert E2E was run. Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Default both prepare() and the script entry point to torch_dist conversion and reference checkpoint loading. Retain VIME_TEST_TORCH_DIST_REF_LOAD=0 for the native HF path, and cover the default and explicit overrides in CPU contracts. Validation: 9 targeted CPU tests passed; Ruff and diff whitespace checks passed. No E2E was run for this default switch. Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
2ddcff4 to
713b48b
Compare
Apply the pinned isort and Black formatting to the ten files reported by pre-commit-npu. Keep runtime logic, CI configuration, and vendor patches unchanged. Validation: all nine pre-commit hooks passed across all files; 118 targeted regression tests passed. Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
713b48b to
95f6309
Compare
Merge ascend 083aea6 with main-friendly native loading and non-colocate training. Reuse the shared GLM model configuration, disable MTP until G2, and align NPU NZ and RoPE settings. Add GLM CI coverage and native MLA/MoE mapping contracts; keep obsolete Bridge files and deferred MTP vendor hunks out of the merge result. Validation: user-reported GLM G1 E2E passed; 291 CPU tests passed; Ruff, Black, isort, shell syntax and diff checks passed. Local TEST ONLY model/dataset overrides, companion tests and the existing Qwen3-30B checkpoint override remain uncommitted. Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
Add native Ascend worker draft target lifecycle and port the GLM MTP graph-friendly position mask. Preserve main's shared training and weight-transfer orchestration. Keep GLM G1 as default and expose independent MTP/eager test switches, with mapping and launch contracts. Validation: 319 CPU regression tests and all nine pre-commit hooks passed on the exact commit candidate. GLM G2 eager (20260910T122325Z) and graph (20260910T123259Z) passed with local 128-token responses, including two training steps and main/draft updates; graph capture completed. Production test retains 2048 tokens. Explicit only-MTP CI gating remains deferred; nonzero-GRPO G2 is not claimed validated. Exclude local model/dataset/checkpoint overrides, companion tests and the temporary response-length change. Signed-off-by: Meihan-chen <zr010426ztt@outlook.com>
| diff --git a/megatron/legacy/fused_kernels/__init__.py b/megatron/legacy/fused_kernels/__init__.py | ||
| index 87cceac3..ac686d74 100644 | ||
| --- a/megatron/legacy/fused_kernels/__init__.py | ||
| +++ b/megatron/legacy/fused_kernels/__init__.py |
| new file mode 100644 | ||
| index 000000000..dc388fd7f | ||
| --- /dev/null | ||
| +++ b/megatron/core/ssm/chunk_gated_delta_rule.py |
There was a problem hiding this comment.
The GDN Triton operator has been removed. Is that because Qwen3.5’s training side is already supported in another third-party repo?

Summary
Bring Ascend up to date with main through
ce92eff1, including the new nativemodel-loading, weight-transfer and accelerator interfaces, while incorporating
Ascend's training-stack updates and GLM-4.7-Flash support through
083aea63.The integration follows main's shared training and rollout flow. Necessary NPU
differences are kept in small platform/provider interfaces and
docker/npu_patch, rather than maintaining a separate copy of main'sorchestration. This keeps subsequent main merges easier to maintain.
Main changes and corresponding NPU adaptations
hccl/npu_ipcand their initialization types through the NPU provider; implement the matching native worker contracts in vendor patches and remove obsolete Vime worker extensions.engine_group.pysplit centralize device operations and engine management.Native loading here refers to the model construction/loading/export path.
The NPU image still contains Bridge-related dependencies; removing all of those
packages is not part of this PR.
Additional Ascend integration
e6bfe03aand vLLM-Ascendfd815467.VIME_TEST_TORCH_DIST_REF_LOAD=0.VIME_TEST_GLM_MTP=1;VIME_TEST_GLM_EAGER=1independently enables eager diagnosis.Validation
Local E2E results collected during integration:
Not yet adapted / follow-ups