Skip to content

Latest commit

 

History

History
186 lines (151 loc) · 9.72 KB

File metadata and controls

186 lines (151 loc) · 9.72 KB

Setup runbook

barry-server setup is a one-command bootstrap for a fresh WSL2 / Linux box with an NVIDIA GPU. It handles host detection, container image pull, model download, config + token issuance, optional systemd unit, and a daemon-spawn probe.

barry-server doctor re-runs the same checklist read-only — useful for diagnosing "why won't barry start".

barry-server upgrade --llama-tag bXXXX [--model q4_k_m|q5_k_m|q6_k] records a new pin in state.json; re-run setup to fetch the new artefacts.

What setup does

15 idempotent steps, in order (16 with --with-search). Each prints [ ok | skip | warn | fail ]. A real failure aborts; warnings continue.

# Step Description
1 platform Detect WSL2 / native Linux. Refuse macOS, refuse non-WSL Windows.
2 gpu Parse nvidia-smi for name + VRAM. Refuse < 16 GiB.
3 podman Require ≥ 4.0 (CDI device passthrough).
4 cdi Detect /etc/cdi/nvidia.yaml. Detect-only — print the exact sudo nvidia-ctk cdi generate … if missing.
5 disk Need ≥ 30 GiB free at the model dir, on ext4/btrfs (refuses 9P / drvfs / cifs / ntfs).
6 image podman pull the pinned llama.cpp tag. Captures the digest.
7 quant VRAM-driven quant pick: <28 GiB → q4_k_m, <40 → q5_k_m, ≥40 → q6_k.
8 model HTTP download the GGUF from HuggingFace with Range: resume + streaming SHA-256.
9 token ~/.local/state/barry/token (mode 0600), 32 random bytes hex.
10 server-config Copy bundled defaults/server.toml to ~/.config/barry/server.toml if absent.
11 client-config Copy bundled defaults/config.toml to ~/.config/barry/config.toml if absent.
12 systemd Install ~/.config/systemd/user/barry.service if systemd --user is up. WSL2 falls back to a warning + nohup hint.
13 searxng-pull (only with --with-search) podman pull the pinned SearXNG tag.
14 searxng-run (only with --with-search) Mint settings.yml (with persisted secret_key), podman run barry-searxng bound to 127.0.0.1:8888, splice [search].url into the client config.toml, health-probe. Doctor mode probes /healthz if the container is running.
15 probe Spawn barry-server on a free port, poll /v1/server/health, kill. Skip with --no-probe or in doctor mode.
16 state Persist ~/.local/state/barry/state.json with image digest, model SHA, last-setup-at, etc.

The server-config / client-config steps copy the bundled defaults straight from the repo (defaults/server.toml, defaults/config.toml) so the file you edit on disk is the same one shipped in source. The bundled file is also include_str!-ed into the server binary and merged under your edits at runtime, so new profiles or knobs added in a future release apply automatically — re-running setup is only needed if you want the new bundled file alongside your edits as a reference.

Optional: bundled SearXNG (--with-search)

The web_search tool talks to a SearXNG instance — a self-hosted meta-search engine that aggregates Google / Bing / DuckDuckGo / etc. without API keys. Install it alongside the daemon:

barry-server setup --with-search

This pulls the pinned docker.io/searxng/searxng:<tag> image, mints a settings.yml (with a secret_key persisted to state.json), and starts the container as barry-searxng bound to 127.0.0.1:8888 (localhost only — never exposed to the LAN). Setup then writes [search] url = "http://127.0.0.1:8888" into the client config.toml so the web_search tool just works.

Bump the pinned image with barry-server upgrade --searxng-tag <YYYY.M.D-sha>; re-run setup --with-search to start the new container. To point the client at a self-hosted SearXNG instead, set [search].url manually before running setup — the splice step leaves existing [search] blocks alone.

Where things live

Path Purpose
~/.cache/barry/models/ model GGUFs (override via BARRY_MODEL_DIR or [model] dir = "…" in server.toml)
~/.local/share/barry/conversations.db SQLite (conversations, turns, memories, audit, artefacts dir)
~/.local/share/barry/artefacts/<conv>/ spilled tool results > 6 KiB
~/.local/share/barry/searxng/settings.yml SearXNG config (only with --with-search)
~/.local/state/barry/state.json what setup did + when (incl. SearXNG secret_key)
~/.local/state/barry/token bearer token (mode 0600)
~/.local/state/barry/server.log daemon logs (when launched via make server-up)
~/.config/barry/server.toml server config — [agent], [model], [profile.*], default_profile (see below)
~/.config/barry/config.toml client config — [search], [shell], [[mcp.servers]] (see below)
<project>/.barry/server.toml optional project-local server config (ancestor walk; first hit wins over user-global)
<project>/.barry/config.toml optional project-local client config (same walk)
~/.config/systemd/user/barry.service optional auto-start unit
barry-searxng (podman container) localhost-bound SearXNG (only with --with-search)

Trade-offs baked in

These are documented in ThreatModel.md too; calling them out here so the runbook reads honest:

  • CDI is detect-only. Setup refuses to silently sudo. If /etc/cdi/nvidia.yaml is missing, the step prints the exact command to run yourself.
  • Model SHA-256 is computed not pinned. HuggingFace doesn't publish authoritative manifests, so we record the SHA on first download and warn on drift on later runs. No fail-fast against a hardcoded list.
  • barry-server upgrade is explicit-flag only — no auto-latest GHCR discovery in v1.0.
  • The probe doesn't bring up llama.cpp. It only verifies the daemon's REST surface. End-to-end barry once "hello" is the manual chat smoke test.

Config files

Barry splits its config along the client/server boundary. Both files are optional — bundled defaults from defaults/server.toml and defaults/config.toml (in the repo root) apply at runtime when fields aren't set locally.

File Loaded by Owns sections
server.toml barry-server default_profile, [agent], [model], [profile.<name>]
config.toml barry [search], [shell], [[mcp.servers]]

Both files use the same discovery chain (first match wins):

  1. CLI flag (--config for server) or env (BARRY_CONFIG for server, BARRY_CLIENT_CONFIG for client).
  2. <ancestor>/.barry/<filename> — walked up from cwd toward /. Lets a project ship its own MCP fleet, sandbox tweaks, profile tweaks alongside its .barry/skills/.
  3. ~/.config/barry/<filename> — user-global.

For full schema details (every field, every section, complete worked examples) see docs/Configuration.md.

Profiles are config-driven

The three shipped model profiles (general-qwen-3.6-35, coder-qwen-3-30, coder-glm-4.7-30) live in defaults/server.toml as [profile."<name>"] blocks (note the quotes — TOML treats . as a key separator unless quoted). To tweak one:

# ~/.config/barry/server.toml
[profile."coder-glm-4.7-30"]
ctx = 65536        # tighter than the bundled 200000
temperature = 0.5

To add a new one, supply every required field (see defaults/server.toml for the field list and inline docs).

Migration from the old single config.toml

Pre-split, both halves of Barry shared ~/.config/barry/config.toml. If you're upgrading from that version: server-side sections ([agent], [model], default_profile, [profile.*]) need to move into ~/.config/barry/server.toml. The server logs a one-time warning at startup if it spots them in the legacy file. Client sections ([search], [shell], [[mcp.servers]]) keep working in config.toml unchanged — they're now the only thing that file owns.

Re-test loop

For a clean setup re-run without re-downloading the 25 GiB model:

make nuke-state          # remove state.json, token, server.toml, config.toml, systemd unit
make setup               # `barry-server setup --quick` under the hood
make doctor              # read-only re-check

The model + image are preserved.

Troubleshooting

  • "df: options -T and --output are mutually exclusive" — already fixed; if you see this, you're on an old binary, run make server-restart.
  • "Model file not found" in model-up — the path resolution walks symlinks, but the parent directory of the resolved file must exist. Make sure BARRY_MODEL_DIR (or [model].dir) points at the real directory, not a symlink that resolves elsewhere.
  • /system shows "(server didn't expose its identity)" — old server binary, run make server-restart.
  • Status line stays empty after a turn — old server (no stream_options.include_usage). Run make server-restart.
  • Setup says "container exited unexpectedly" during model-up — tail the log: make model-logs. Most common cause: not enough VRAM for the chosen quant. Either pick a smaller quant via barry-server upgrade --model q4_k_m or close other GPU users.