Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
-
Updated
Aug 25, 2026 - Python
Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
Measured LLM benchmarks for AMD Strix Halo / Ryzen AI Max+ 395 (Radeon 8060S, 128 GB unified): llama.cpp Vulkan & ROCm — decode pace, TTFA, prompt cache, quants, sustained load. Every number links to raw runs.
vLLM serving poolside Laguna S 2.1 (118B MoE, 8B active, INT4) on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151, 128GB unified) via TheRock ROCm nightly. OpenAI-compatible, 256K context, DFlash speculative decoding.
Run large LLMs locally on AMD Ryzen AI Max+ 395 (Strix Halo, gfx1151) with ROCmFP4 4-bit quantization. Measured benchmarks, build + serving recipes, and 55 ready-to-run GGUF models.
Add a description, image, and links to the ryzen-ai-max-395 topic page so that developers can more easily learn about it.
To associate your repository with the ryzen-ai-max-395 topic, visit your repo's landing page and select "manage topics."