Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
-
Updated
Aug 25, 2026 - Python
Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
A turnkey, fully-local AI workstation engineered for the AMD Ryzen AI Max+ 395. LLM inference, voice, document parsing, browser automation, agents — all on-device.
Measured LLM benchmarks for AMD Strix Halo / Ryzen AI Max+ 395 (Radeon 8060S, 128 GB unified): llama.cpp Vulkan & ROCm — decode pace, TTFA, prompt cache, quants, sustained load. Every number links to raw runs.
Run large LLMs locally on AMD Ryzen AI Max+ 395 (Strix Halo, gfx1151) with ROCmFP4 4-bit quantization. Measured benchmarks, build + serving recipes, and 55 ready-to-run GGUF models.
Docker stack: Ollama v0.21.0 built from source against ROCm 7.2.2 with native gfx1151 (Strix Halo) — serves Gemma 4 up to 256K context on AMD Ryzen AI MAX+ 395 / Radeon 8060S. Includes a 9-layer make validate ladder for the host firmware, ROCm runtime, container, and long-context inference.
High-performance Vulkan runner & deployment package for Ornith-1.0-35B on AMD Strix Halo (Radeon 8060S). Delivering 115+ t/s aggregate decode, 256k context, and high-accuracy 16-agent burst tool calling.
🚀 Automated nightly builds & multi-backend releases (ROCm 7.x TheRock, Vulkan RADV, CUDA, Metal) for CachyLLama & llama-ai on AMD Strix Halo (8060S), Strix Point (890M), Phoenix (780M), and Steam Deck.
Add a description, image, and links to the radeon-8060s topic page so that developers can more easily learn about it.
To associate your repository with the radeon-8060s topic, visit your repo's landing page and select "manage topics."