Popular repositories Loading
-
GLM-5.3-Flash-2x-DGX-Spark
GLM-5.3-Flash-2x-DGX-Spark PublicFirst GLM-5.3-Flash on DGX Spark (2x GB10): 24.7/30.3/19.6 tok/s with MTP-5. NoPE-MLA zero-pad mod + marlin MoE on the stock vLLM image. Tested, measured, honest numbers.
Python 5
-
strix-halo-quant-lab
strix-halo-quant-lab PublicRun large LLMs locally on AMD Ryzen AI Max+ 395 (Strix Halo, gfx1151) with ROCmFP4 4-bit quantization. Measured benchmarks, build + serving recipes, and 55 ready-to-run GGUF models.
Python 1
-
-
ROCmFPX
ROCmFPX PublicForked from charlie12345/ROCmFPX
ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants
C++
If the problem persists, check the GitHub status page or contact support.
