Popular repositories Loading
-
deepseek-v4-flash-gb10
deepseek-v4-flash-gb10 PublicServe DeepSeek-V4-Flash-0731 on 2x NVIDIA GB10 / DGX Spark (sm_121) with vLLM — tuned dual-Spark config (TP=2+EP, DSpark spec decode, NCCL 2.30.4/RDMA, 384K), runbook + verify harness. ~1.7k tok/s …
-
vllm-gfx1201-qwen3.6-256k
vllm-gfx1201-qwen3.6-256k PublicQwen3.6-27B-FP8 at 256k context on dual AMD RDNA4 (gfx1201) — working vLLM recipe, audited measurement harness, and the full debugging record
Python 4
-
vllm-metrics-dashboard
vllm-metrics-dashboard PublicReal-time single-screen dashboard for vLLM servers — throughput, latency, KV/prefix cache, and a cost/ROI panel vs hosted APIs. Standalone or fleet.
HTML 1
-
flashnext-rocm10-r9700
flashnext-rocm10-r9700 PublicExperimental ROCm 10 compatibility build for Qwen3.8-Flash-Next on 4x AMD R9700
Dockerfile 1
-
-
If the problem persists, check the GitHub status page or contact support.