A desktop app (PyQt6) for KDE that started as a simple Ollama service manager and grew into a control panel for a whole local AI stack: multi-host model management, a model aggregator (LiteLLM) that exposes one OpenAI-compatible endpoint for tools like Continue/VS Code, and a browser chat UI (Open WebUI). No terminal, no Docker, no cloud — everything runs on your LAN.
(Polska wersja tego pliku: README_PL.md)
Status: stable (v0.5.0). Daily driver on my own setup, working as expected. Still actively developed — new options and features keep coming; bug reports and feedback are welcome.
- Python 3 + PyQt6, requests
- systemd + polkit (
pkexec) — standard on KDE/Debian;install.shoffers to installpkexecviaaptif it's missing (needed for every action requiring admin rights) - curl — used by the app's own install buttons (Ollama, and uv for Open WebUI/LiteLLM);
install.shoffers to install it viaaptif it's missing - Ollama (if you don't have it, the app installs it with one click)
- Optional: uv (for installing Open WebUI and LiteLLM — the app installs it itself if needed)
- Optional:
ffmpeg,pandoc,zstd(full Open WebUI functionality — voice, document RAG; the app offers to install these viaaptright after a successful Open WebUI install)
First, clone the repository and enter its directory — the install scripts below
expect to be run from there (they look for ollama_manager.py next to themselves):
git clone https://github.com/cyryllo/Ollama-manager.git
cd Ollama-manager
As a menu app, no root (recommended) — installs dependencies via pip, copies the app to
~/.local/share/ollama-manager and adds a menu entry (Utilities section). Detects an existing
install and offers to update/reinstall accordingly:
./install.sh
Uninstall with ./install.sh --uninstall.
As a .deb package (Debian/Ubuntu) — dependencies come from apt, easy to uninstall:
./build-deb.sh
sudo apt install ./ollama-manager_*_all.deb
Manually, for development — no copying, no menu entry:
pip install PyQt6 requests
python3 ollama_manager.py
Ollama service
- Start / stop the systemd service, live status detection
- Autostart on system boot
- Detects a missing install + a one-click install button
- On-demand update check against the latest GitHub release, with a prompt to reinstall (the official install script updates an existing install in place) if you're behind — useful since newer models sometimes need a newer Ollama to even load
Models
- One search for two sources — type a model name and the app searches the official Ollama library (ollama.com) and Hugging Face (GGUF models only) at once. Every result has a Source column, so it's always clear where a model comes from. Pick a result to see all its variants (tags/quantizations) with their real file sizes, pick a variant, click Download — with a progress bar. You never need to know the exact tag up front. MLX variants (Apple-only) and "cloud" models (not downloadable) are skipped
- Pasting a Hugging Face link (or
user/repo) jumps straight to that repo's variants; typing an exact name (e.g.llama3.2:3borhf.co/...:Q4_K_M) lets you download without searching - Hugging Face is an optional, community source — anyone can publish there and the app doesn't vet repo contents, so a note in the card recommends sticking to well-known quantization publishers (e.g. bartowski, mradermacher, unsloth)
- Approximate memory-usage estimate for the chosen variant (real file size + KV cache + buffers, then a recommended-RAM figure; details in a tooltip)
- Table of installed models — name, source (Ollama / Hugging Face), disk size and an estimated/live VRAM column — + deletion. Hover a row for its source (with the quantization author for HF), model family, parameter size, quantization and download date
- Preview of models currently loaded into memory (VRAM)
- Note: ollama.com has no public search API, so the app reads its search/tags web pages — if the site's layout changes, Ollama search may need a fix (Hugging Face uses a real API)
Open WebUI
- One-click install of the browser chat panel (no Docker)
- Start / stop, autostart on login
- The "Open WebUI" button opens the panel in the browser (no automatic opening)
- On-demand update check against the latest PyPI release, with a prompt to upgrade via
uv tool upgrade(no admin rights needed) if you're behind - After a fresh install, offers to install the optional
ffmpeg/pandoc/zstdpackages needed for voice transcription and document RAG (the only step that needs admin rights)
Server switcher
- Choose the Ollama host (localhost or any host on the LAN, e.g. BC-250) for model operations
- Add/remove servers from the window, remembered between runs
Model aggregator (LiteLLM)
- One-click install, start/stop and autostart for LiteLLM (no Docker)
- Exposes a single endpoint (OpenAI-compatible) combining models from ALL servers on the switcher list — VS Code/Continue only needs to point at this one address
- Preview of which models and hosts will end up in the config, before starting
- Generates a ready-to-paste Continue.dev
config.yamlfor the exposed models — copy/paste only, the app never writes to your files - Connection details for Unsloth Studio — Studio has no config file (connections are added in its Settings → Connections), so the app shows exactly what to enter (connection type, Base URL, API key) plus the list of model IDs to copy
- On-demand update check against the latest PyPI release, with a prompt to upgrade via
uv tool upgrade(no admin rights needed) if you're behind
Advanced (Ollama environment variables)
OLLAMA_KEEP_ALIVE,OLLAMA_CONTEXT_LENGTH(default 4096 is too small for agentic work),OLLAMA_MAX_LOADED_MODELS,OLLAMA_NUM_PARALLEL,OLLAMA_FLASH_ATTENTION,OLLAMA_KV_CACHE_TYPE,OLLAMA_VULKAN(Vulkan backend instead of ROCm, useful on AMD cards without full ROCm support e.g. BC-250),OLLAMA_IGPU_ENABLE,OLLAMA_HOST(listen on the LAN instead of localhost-only),GGML_VK_VISIBLE_DEVICES(Vulkan device index — needed on some APUs),OLLAMA_GPU_OVERHEAD(VRAM reserved for the rest of the system — desktop/server preset picker)- A single "Save" applies every field at once and restarts the service once, plus an "Apply recommended values" button that pre-fills a sane starting profile for review before saving
- A free-form field for any other environment variable not covered by the form above
- Full descriptions of each variable live in a separate "Help" tab, keeping this tab compact
Stats bar
- Ollama and Open WebUI status
- VRAM usage on the currently selected server
- Number of installed models
Event log — a log of all operations, always visible at the bottom of the window.
Interface language — Polish, English, German, Spanish, French, Portuguese and Italian, switchable from the window, remembered between runs.
- Ollama service control always targets the local machine — even if a remote server (e.g. BC-250) is selected in the window, start/stop/autostart act locally.
