This guide brings up the traditional vLLM-backed stack on a single-GPU workstation where GPU 0 may be reserved for the desktop and GPU 1 is free for inference. It starts with the tiniest vLLM model to prove the plumbing works, then switches to progressively larger vLLM profiles.
For the simpler Pascal/1080 Ti path that runs Open WebUI -> Ollama with no
LiteLLM and no predeclared model registry, use
ollama_direct_quickstart.md.
Rendered shape for this guide:
Open WebUI -> LiteLLM -> vLLM
Default ports for this profile family:
- LiteLLM: http://127.0.0.1:14042/v1
- Open WebUI: http://127.0.0.1:13000
After setup, all state and generated artifacts live under
/data/service/docker/infer-stack/:
/data/service/docker/infer-stack/
generated/ <- docker-compose.yml, .env, plan.yaml
hf-cache/ <- downloaded Hugging Face weights
vllm-cache/ <- compiled vLLM artifacts
open-webui/ <- Open WebUI state
postgres-open-webui/ <- Open WebUI database
postgres-litellm/ <- LiteLLM database, only when LiteLLM is enabled
runtime/ <- runtime bind-mount configs such as litellm_config.yaml
~/.config/infer_stack/
config.yaml <- active profile, backend, paths, ports
models.yaml <- optional custom vllm_models, ollama_models, profiles
- Docker with the NVIDIA container runtime (
nvidia-smivisible inside containers). infer-stackinstalled in your Python environment.- A Hugging Face token for gated models, stored in the managed
.env. Compose never readsHF_TOKENfrom your shell:
infer-stack env HF_TOKEN=hf_...The first profile in this guide, gpt2-single, is public and does not need
HF_TOKEN.
Start with gpt2-single: GPT-2 124M, ~250 MB download, no auth,
completions-only. The goal is to validate Docker, GPU placement, LiteLLM, and
Open WebUI before committing to a large model.
infer-stack setup \
--backend compose \
--profile gpt2-single \
--data-dir /data/service/docker/infer-stackRender and start:
infer-stack render --yes
infer-stack up -dvLLM downloads GPT-2 and compiles a small graph. Expect roughly 30-60 seconds on first start and a few seconds on warm restart.
Watch it come up:
infer-stack ps
infer-stack logs -fOnce all containers are healthy, smoke-test the API:
infer-stack smoke-testManual completion check:
curl -s http://127.0.0.1:14042/v1/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $(infer-stack env --key LITELLM_MASTER_KEY)" \
-d '{
"model": "gpt2",
"prompt": "Once upon a time, ",
"max_tokens": 32
}' | python3 -m json.toolOpen WebUI is up at http://127.0.0.1:13000, but GPT-2 is a base model, so switch to a chat-capable profile before using it interactively.
infer-stack switch smollm2-135m-single --applyswitch --apply updates config.yaml, re-renders the stack, removes orphaned
vLLM containers, and refreshes LiteLLM/Open WebUI in place. Cached weights stay
on disk.
Smoke test again:
infer-stack smoke-testThen open http://127.0.0.1:13000. Open WebUI now has a chat-capable route.
After the plumbing is proven, switch to a real workstation profile. The
workstation-safe profile uses first-fit placement so it avoids display GPUs
when the policy reserves them:
infer-stack switch workstation-safe --applyFirst start downloads the model and warms the vLLM cache; subsequent restarts
reuse hf-cache/ and vllm-cache/.
infer-stack smoke-testIf you want a profile that explicitly pins a runtime to a particular GPU, copy
one of the built-in profile definitions into ~/.config/infer_stack/models.yaml
and change providers.vllm.runtimes.<name>.placement.gpu_indices.
Stop without deleting state:
infer-stack downStart after a reboot:
infer-stack up -dTail logs and inspect status:
infer-stack logs
infer-stack psRun the smoke test:
infer-stack smoke-testRead the LiteLLM key from .env:
infer-stack env --key LITELLM_MASTER_KEYinfer-stack up does a pre-flight check on enabled component ports. For this
profile family, the usual ports are 14042 for LiteLLM and 13000 for Open WebUI.
Direct Ollama profiles may also publish 11434.
Find what is holding a port:
ss -tlnp 'sport = :14042'
sudo lsof -nP -iTCP:14042 -sTCP:LISTEN
docker ps --filter publish=14042Common fixes:
- Stop the old stack with
infer-stack down. - Stop a stale container, for example
docker stop litellm && docker rm litellm. - Change ports with setup flags such as
infer-stack setup --litellm-port 14001 --open-webui-port 13001, then render and start again.
- Could not connect: give the stack more time; check
infer-stack ps. - Connection closed before a response: the router is up but the upstream model
may still be loading; inspect
infer-stack logs vllm-*. - 401/403: the key in
.envdoes not match the running LiteLLM container; restart withinfer-stack down && infer-stack up -d. - 503: vLLM is still loading; inspect
infer-stack logs vllm-*.
Delete databases, Open WebUI state, and runtime configs while keeping model caches:
infer-stack purge --yesDelete everything, including model caches:
infer-stack purge --yes --delete-cache