面向 8×RTX 4090(SM89)的 DeepSeek-V4-Flash-0731 MXFP4 SGLang 实验分支
-
Updated
Sep 28, 2026 - Python
面向 8×RTX 4090(SM89)的 DeepSeek-V4-Flash-0731 MXFP4 SGLang 实验分支
4K video super-resolution on a 24 GB RTX 4090 — a free, local ComfyUI FlashVSR node (VIDEO in, VIDEO out) plus a Triton/FP8 operator pack at ~1.3×. No CUDA build, every number backed by benchmark JSONs.
QSA HiSparse for SGLang: 256K KV offload, CUDA Graph benchmarks, and patches tested on dual RTX 4090 48GB
Feeding the Tensor Cores: a dense FP16 GEMM for NVIDIA Ada (sm_89) at 96.5% of cuBLAS, and what it teaches about how each GPU generation handles async copy and Tensor Core issue.
Run DeepSeek-V4-Flash with 1M-token context on pre-Hopper NVIDIA GPUs (sm_89) — a vLLM 0.26.0 overlay replacing all SM90+ DeepGEMM/FlashMLA operators with Triton/TileLang/Marlin kernels + CPU-KV-pool (pinned host memory via CUDA UVA).
Port NInfer, a single-GPU CUDA inference engine, to the NVIDIA L20 (Ada sm_89, 92 SMs, 48 GB): patch set, build tooling, and measured results
To associate your repository with the sm89 topic, visit your repo's landing page and select "manage topics."