Qwen3.8-Flash-Next (125B MoE, NVFP4) on 4x V100-SXM2-32GB and DeepSeek-V4.1-flash on 8x V100 — a Volta port of SGLang for agentic coding.
moe multi-gpu volta dgx-1 v100 nvlink long-context llm-inference qwen speculative-decoding sglang tesla-v100 agentic-coding nvfp4 sm70 deepseek-v41-flash
-
Updated
Sep 28, 2026 - Python