gfx1201
Here are 29 public repositories matching this topic...
Run Gemma-4-31B at full 256K context on a $1,400 AMD RDNA4 GPU (gfx1201): TurboQuant KV cache + HIP-graph-safe Flash-Attention for llama.cpp, fully measured on real hardware.
-
Updated
Jul 23, 2026 - Python
llama.cpp with native AMD RDNA4 (gfx1201) ROCm 7.11 support - 98.97 tok/s AI inference, competitive with RTX 4070 Ti, 32GB VRAM
-
Updated
Jan 3, 2026 - C++
Fine-tune your own LLM on an AMD Radeon GPU — the easy, tested way. QLoRA via ROCm on Windows/WSL2 & Linux, a worked Gemma-4 example, a reusable live training dashboard, and a smoke test that proves the loss falls.
-
Updated
Jun 20, 2026 - HTML
Hardware-focused llama.cpp fork for Windows and dual AMD RDNA4 RX 9070 XT GPUs: ROCm/HIP + Vulkan backends, PyQt6 GUI (RDNA LLM Studio), MTP speculative decoding, FP8 attention, long-context Qwen3.8 benchmarks
-
Updated
Sep 7, 2026 - C++
Qwen3.6-27B-FP8 at 256k context on dual AMD RDNA4 (gfx1201) — working vLLM recipe, audited measurement harness, and the full debugging record
-
Updated
Aug 7, 2026 - Python
A from-scratch LLM inference engine in Zig, specialized for AMD RDNA 4 GPUs.
-
Updated
Aug 26, 2026 - Zig
Validated Graphdeco 3D Gaussian Splatting port for ROCm 7.2 and RDNA 4 gfx1201, with HIP training extensions tested on Radeon AI PRO R9700.
-
Updated
Aug 2, 2026 - Python
-
Updated
Aug 2, 2026 - Cuda
Community ROCm 7.2 / RDNA4 gfx1201 patch stack for nvdiffrast. Non-commercial use restrictions may apply; see the NVIDIA Source Code License.
-
Updated
Aug 2, 2026 - Shell
Qualified AMD ROCm 7.2 / RDNA4 gfx1201 integration for Nerfstudio's Nerfacto training chain on Radeon AI PRO R9700.
-
Updated
Aug 5, 2026 - Python
-
Updated
Aug 2, 2026 - Cuda
-
Updated
Aug 2, 2026 - Cuda
Add this topic to your repo
To associate your repository with the gfx1201 topic, visit your repo's landing page and select "manage topics."