Skip to content

About

LM Studio CUDA backend patch adding sm_60/sm_70 (P100/V100) support for backend 2.25.2

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

4 Commits

Folders and files

Repository files navigation

Patched CUDA Runtime for NVIDIA Pascal / Volta — LM Studio

Custom LM Studio CUDA backend adding sm_60 (P100) and sm_70 (V100) support. Stock LM Studio 2.25.2 CUDA builds omit these architectures; Volta GPUs crash in libggml-cuda.so (mul_mat_vec_q — no sm_70 kernel).

No llama.cpp source patches required — only rebuild at the pinned commit with expanded CMAKE_CUDA_ARCHITECTURES.

Status: build scripts + integration tooling for LM Studio 2.25.2 (llama.cpp cb295bf59 / b9888).

Inspired by Delitants/lmstudio-cuda-kepler-patch.

LM Studio Runtime — V100 sm_70 backend selected


Supported hardware

GPU Architecture Compute capability Notes
Tesla P100 Pascal 6.0 sm_60
Tesla V100 Volta 7.0 sm_70 — primary target
GTX 1080 / 1070 Pascal 6.1 covered by stock sm_61; included in multi-arch build
RTX 20xx+ Turing+ 7.5+ still included (75;80)

Prerequisites

  • LM Studio 0.4.x with CUDA backend 2.25.2 installed
  • Engine protocol enabled (useLlamaCppEngineProtocolRuntime3: true)
  • CUDA Toolkit 11.x — stock backend links libcudart.so.11.0
  • cmake, ninja, gcc-11 / g++-11
  • Linux x86_64 with AVX2
./scripts/install-deps.sh   # optional apt helper

Quick start

# 1. Fetch pinned llama.cpp source
./scripts/fetch-llama-cpp.sh

# 2. Build CUDA libraries + llama-server
./build-scripts/build_cuda_v100_avx2.sh

# 3. Install custom backend alongside stock (does not overwrite stock)
./integrate.sh

# 4. Restart LM Studio → Runtime → "CUDA llama.cpp (Linux, V100 sm_70)"

Verify build before integrating

strings install-cuda-v100-avx2/bin/libggml-cuda.so | rg '\.target sm_' | sort -u
# expect sm_60 sm_70 sm_75 sm_80 (and possibly sm_50)

What gets replaced

From your CUDA build (open-source llama.cpp artifacts):

  • libggml-base.so, libggml-cpu.so, libggml-cuda.so
  • libllama.so, libllama-common.so*
  • libllama-server-impl.so, libmtmd.so, llama-server

Kept from stock (LM Studio proprietary — do not rebuild):

  • llm_engine_cuda.node, liblmstudio_bindings_cuda.node
  • libllm_engine.so, liblmstudiocore.so, libggml_llamacpp.so

Repository layout

├── build-scripts/build_cuda_v100_avx2.sh
├── scripts/fetch-llama-cpp.sh
├── scripts/prepare-cuda-shadow.sh
├── scripts/install-deps.sh
├── integrate.sh
├── lm-studio-manifest/backend-manifest-cuda.json
├── docs/cuda-volta-lmstudio.md
└── screenshots/runtime-v100-sm70.png

llama.cpp/ is fetched locally and gitignored.


LM Studio settings

Before restart, in Developer settings:

  • autoDeleteExtensionPacks → off (or custom backend may be purged on update)
  • useLlamaCppEngineProtocolRuntime3 → on

Rollback

rm -rf ~/.lmstudio/extensions/backends/llama.cpp-linux-x86_64-nvidia-cuda-avx2-v100-2.25.2

Stock ...-nvidia-cuda-avx2-2.25.2 is never modified by integrate.sh.

Test against a local LM Studio backends copy

LMS_BACKENDS_DIR=/path/to/lmstudio/extensions/backends ./integrate.sh

Troubleshooting

Symptom Likely cause
Backend not in UI Invalid JSON manifest or bad folder name
Same CUDA crash libggml-cuda.so still lacks sm_70 — recheck strings
llama-server won't start Wrong llama.cpp commit / ABI mismatch
cuBLAS errors Built against CUDA 12+, stock expects 11.x
Backend vanishes after update autoDeleteExtensionPacks enabled
nvcc cospi/sinpi errors on Ubuntu 24.04+ Auto-handled via repo-local CUDA shadow toolkit (scripts/prepare-cuda-shadow.sh)

See docs/cuda-volta-lmstudio.md for full details.


License

  • Build scripts and docs: MIT
  • llama.cpp: MIT (upstream)
  • LM Studio .node / wrapper binaries: property of Element Labs — not redistributed here

About

LM Studio CUDA backend patch adding sm_60/sm_70 (P100/V100) support for backend 2.25.2

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages