Multi-GPU LLM inference runtime - combines NVIDIA and AMD GPUs with speculative decoding for 2x speedup
-
Updated
Jan 25, 2026 - C++
Multi-GPU LLM inference runtime - combines NVIDIA and AMD GPUs with speculative decoding for 2x speedup
Add a description, image, and links to the super-llama-cpp topic page so that developers can more easily learn about it.
To associate your repository with the super-llama-cpp topic, visit your repo's landing page and select "manage topics."