barry-server setup is a one-command bootstrap for a fresh
WSL2 / Linux box with an NVIDIA GPU. It handles host detection,
container image pull, model download, config + token issuance,
optional systemd unit, and a daemon-spawn probe.
barry-server doctor re-runs the same checklist read-only — useful
for diagnosing "why won't barry start".
barry-server upgrade --llama-tag bXXXX [--model q4_k_m|q5_k_m|q6_k]
records a new pin in state.json; re-run setup to fetch the new
artefacts.
15 idempotent steps, in order (16 with --with-search). Each prints
[ ok | skip | warn | fail ]. A real failure aborts; warnings continue.
| # | Step | Description |
|---|---|---|
| 1 | platform | Detect WSL2 / native Linux. Refuse macOS, refuse non-WSL Windows. |
| 2 | gpu | Parse nvidia-smi for name + VRAM. Refuse < 16 GiB. |
| 3 | podman | Require ≥ 4.0 (CDI device passthrough). |
| 4 | cdi | Detect /etc/cdi/nvidia.yaml. Detect-only — print the exact sudo nvidia-ctk cdi generate … if missing. |
| 5 | disk | Need ≥ 30 GiB free at the model dir, on ext4/btrfs (refuses 9P / drvfs / cifs / ntfs). |
| 6 | image | podman pull the pinned llama.cpp tag. Captures the digest. |
| 7 | quant | VRAM-driven quant pick: <28 GiB → q4_k_m, <40 → q5_k_m, ≥40 → q6_k. |
| 8 | model | HTTP download the GGUF from HuggingFace with Range: resume + streaming SHA-256. |
| 9 | token | ~/.local/state/barry/token (mode 0600), 32 random bytes hex. |
| 10 | server-config | Copy bundled defaults/server.toml to ~/.config/barry/server.toml if absent. |
| 11 | client-config | Copy bundled defaults/config.toml to ~/.config/barry/config.toml if absent. |
| 12 | systemd | Install ~/.config/systemd/user/barry.service if systemd --user is up. WSL2 falls back to a warning + nohup hint. |
| 13 | searxng-pull | (only with --with-search) podman pull the pinned SearXNG tag. |
| 14 | searxng-run | (only with --with-search) Mint settings.yml (with persisted secret_key), podman run barry-searxng bound to 127.0.0.1:8888, splice [search].url into the client config.toml, health-probe. Doctor mode probes /healthz if the container is running. |
| 15 | probe | Spawn barry-server on a free port, poll /v1/server/health, kill. Skip with --no-probe or in doctor mode. |
| 16 | state | Persist ~/.local/state/barry/state.json with image digest, model SHA, last-setup-at, etc. |
The server-config / client-config steps copy the bundled defaults
straight from the repo (defaults/server.toml, defaults/config.toml)
so the file you edit on disk is the same one shipped in source. The
bundled file is also include_str!-ed into the server binary and merged
under your edits at runtime, so new profiles or knobs added in a future
release apply automatically — re-running setup is only needed if you
want the new bundled file alongside your edits as a reference.
The web_search tool talks to a SearXNG instance — a self-hosted
meta-search engine that aggregates Google / Bing / DuckDuckGo / etc.
without API keys. Install it alongside the daemon:
barry-server setup --with-searchThis pulls the pinned docker.io/searxng/searxng:<tag> image, mints
a settings.yml (with a secret_key persisted to state.json),
and starts the container as barry-searxng bound to 127.0.0.1:8888
(localhost only — never exposed to the LAN). Setup then writes
[search] url = "http://127.0.0.1:8888" into the client config.toml
so the web_search tool just works.
Bump the pinned image with barry-server upgrade --searxng-tag <YYYY.M.D-sha>; re-run setup --with-search to start the new
container. To point the client at a self-hosted SearXNG instead, set
[search].url manually before running setup — the splice step leaves
existing [search] blocks alone.
| Path | Purpose |
|---|---|
~/.cache/barry/models/ |
model GGUFs (override via BARRY_MODEL_DIR or [model] dir = "…" in server.toml) |
~/.local/share/barry/conversations.db |
SQLite (conversations, turns, memories, audit, artefacts dir) |
~/.local/share/barry/artefacts/<conv>/ |
spilled tool results > 6 KiB |
~/.local/share/barry/searxng/settings.yml |
SearXNG config (only with --with-search) |
~/.local/state/barry/state.json |
what setup did + when (incl. SearXNG secret_key) |
~/.local/state/barry/token |
bearer token (mode 0600) |
~/.local/state/barry/server.log |
daemon logs (when launched via make server-up) |
~/.config/barry/server.toml |
server config — [agent], [model], [profile.*], default_profile (see below) |
~/.config/barry/config.toml |
client config — [search], [shell], [[mcp.servers]] (see below) |
<project>/.barry/server.toml |
optional project-local server config (ancestor walk; first hit wins over user-global) |
<project>/.barry/config.toml |
optional project-local client config (same walk) |
~/.config/systemd/user/barry.service |
optional auto-start unit |
barry-searxng (podman container) |
localhost-bound SearXNG (only with --with-search) |
These are documented in ThreatModel.md too; calling them out here so
the runbook reads honest:
- CDI is detect-only. Setup refuses to silently sudo. If
/etc/cdi/nvidia.yamlis missing, the step prints the exact command to run yourself. - Model SHA-256 is computed not pinned. HuggingFace doesn't publish authoritative manifests, so we record the SHA on first download and warn on drift on later runs. No fail-fast against a hardcoded list.
barry-server upgradeis explicit-flag only — no auto-latest GHCR discovery in v1.0.- The probe doesn't bring up llama.cpp. It only verifies the
daemon's REST surface. End-to-end
barry once "hello"is the manual chat smoke test.
Barry splits its config along the client/server boundary. Both files
are optional — bundled defaults from defaults/server.toml and
defaults/config.toml (in the repo root) apply at runtime when
fields aren't set locally.
| File | Loaded by | Owns sections |
|---|---|---|
server.toml |
barry-server |
default_profile, [agent], [model], [profile.<name>] |
config.toml |
barry |
[search], [shell], [[mcp.servers]] |
Both files use the same discovery chain (first match wins):
- CLI flag (
--configfor server) or env (BARRY_CONFIGfor server,BARRY_CLIENT_CONFIGfor client). <ancestor>/.barry/<filename>— walked up from cwd toward/. Lets a project ship its own MCP fleet, sandbox tweaks, profile tweaks alongside its.barry/skills/.~/.config/barry/<filename>— user-global.
For full schema details (every field, every section, complete
worked examples) see docs/Configuration.md.
The three shipped model profiles (general-qwen-3.6-35,
coder-qwen-3-30, coder-glm-4.7-30) live in
defaults/server.toml as [profile."<name>"] blocks (note the
quotes — TOML treats . as a key separator unless quoted). To
tweak one:
# ~/.config/barry/server.toml
[profile."coder-glm-4.7-30"]
ctx = 65536 # tighter than the bundled 200000
temperature = 0.5To add a new one, supply every required field (see
defaults/server.toml for the field list and inline docs).
Pre-split, both halves of Barry shared ~/.config/barry/config.toml.
If you're upgrading from that version: server-side sections ([agent],
[model], default_profile, [profile.*]) need to move into
~/.config/barry/server.toml. The server logs a one-time warning at
startup if it spots them in the legacy file. Client sections
([search], [shell], [[mcp.servers]]) keep working in
config.toml unchanged — they're now the only thing that file owns.
For a clean setup re-run without re-downloading the 25 GiB model:
make nuke-state # remove state.json, token, server.toml, config.toml, systemd unit
make setup # `barry-server setup --quick` under the hood
make doctor # read-only re-checkThe model + image are preserved.
- "
df: options -T and --output are mutually exclusive" — already fixed; if you see this, you're on an old binary, runmake server-restart. - "Model file not found" in model-up — the path resolution walks
symlinks, but the parent directory of the resolved file must
exist. Make sure
BARRY_MODEL_DIR(or[model].dir) points at the real directory, not a symlink that resolves elsewhere. /systemshows "(server didn't expose its identity)" — old server binary, runmake server-restart.- Status line stays empty after a turn — old server (no
stream_options.include_usage). Runmake server-restart. - Setup says "container exited unexpectedly" during model-up —
tail the log:
make model-logs. Most common cause: not enough VRAM for the chosen quant. Either pick a smaller quant viabarry-server upgrade --model q4_k_mor close other GPU users.