Custom LM Studio CUDA backend adding sm_60 (P100) and sm_70 (V100) support.
Stock LM Studio 2.25.2 CUDA builds omit these architectures; Volta GPUs crash in
libggml-cuda.so (mul_mat_vec_q — no sm_70 kernel).
No llama.cpp source patches required — only rebuild at the pinned commit with
expanded CMAKE_CUDA_ARCHITECTURES.
Status: build scripts + integration tooling for LM Studio 2.25.2 (llama.cpp cb295bf59 / b9888).
Inspired by Delitants/lmstudio-cuda-kepler-patch.
| GPU | Architecture | Compute capability | Notes |
|---|---|---|---|
| Tesla P100 | Pascal | 6.0 | sm_60 |
| Tesla V100 | Volta | 7.0 | sm_70 — primary target |
| GTX 1080 / 1070 | Pascal | 6.1 | covered by stock sm_61; included in multi-arch build |
| RTX 20xx+ | Turing+ | 7.5+ | still included (75;80) |
- LM Studio 0.4.x with CUDA backend 2.25.2 installed
- Engine protocol enabled (
useLlamaCppEngineProtocolRuntime3: true) - CUDA Toolkit 11.x — stock backend links
libcudart.so.11.0 - cmake, ninja, gcc-11 / g++-11
- Linux x86_64 with AVX2
./scripts/install-deps.sh # optional apt helper# 1. Fetch pinned llama.cpp source
./scripts/fetch-llama-cpp.sh
# 2. Build CUDA libraries + llama-server
./build-scripts/build_cuda_v100_avx2.sh
# 3. Install custom backend alongside stock (does not overwrite stock)
./integrate.sh
# 4. Restart LM Studio → Runtime → "CUDA llama.cpp (Linux, V100 sm_70)"strings install-cuda-v100-avx2/bin/libggml-cuda.so | rg '\.target sm_' | sort -u
# expect sm_60 sm_70 sm_75 sm_80 (and possibly sm_50)From your CUDA build (open-source llama.cpp artifacts):
libggml-base.so,libggml-cpu.so,libggml-cuda.solibllama.so,libllama-common.so*libllama-server-impl.so,libmtmd.so,llama-server
Kept from stock (LM Studio proprietary — do not rebuild):
llm_engine_cuda.node,liblmstudio_bindings_cuda.nodelibllm_engine.so,liblmstudiocore.so,libggml_llamacpp.so
├── build-scripts/build_cuda_v100_avx2.sh
├── scripts/fetch-llama-cpp.sh
├── scripts/prepare-cuda-shadow.sh
├── scripts/install-deps.sh
├── integrate.sh
├── lm-studio-manifest/backend-manifest-cuda.json
├── docs/cuda-volta-lmstudio.md
└── screenshots/runtime-v100-sm70.png
llama.cpp/ is fetched locally and gitignored.
Before restart, in Developer settings:
autoDeleteExtensionPacks→ off (or custom backend may be purged on update)useLlamaCppEngineProtocolRuntime3→ on
rm -rf ~/.lmstudio/extensions/backends/llama.cpp-linux-x86_64-nvidia-cuda-avx2-v100-2.25.2Stock ...-nvidia-cuda-avx2-2.25.2 is never modified by integrate.sh.
LMS_BACKENDS_DIR=/path/to/lmstudio/extensions/backends ./integrate.sh| Symptom | Likely cause |
|---|---|
| Backend not in UI | Invalid JSON manifest or bad folder name |
| Same CUDA crash | libggml-cuda.so still lacks sm_70 — recheck strings |
llama-server won't start |
Wrong llama.cpp commit / ABI mismatch |
| cuBLAS errors | Built against CUDA 12+, stock expects 11.x |
| Backend vanishes after update | autoDeleteExtensionPacks enabled |
nvcc cospi/sinpi errors on Ubuntu 24.04+ |
Auto-handled via repo-local CUDA shadow toolkit (scripts/prepare-cuda-shadow.sh) |
See docs/cuda-volta-lmstudio.md for full details.
- Build scripts and docs: MIT
- llama.cpp: MIT (upstream)
- LM Studio
.node/ wrapper binaries: property of Element Labs — not redistributed here
