A low-latency voice companion for a self-hosted AI stack.
You talk to your phone. A fast interactive model answers immediately and, when the request warrants it, hands work to a background agent that keeps working while the conversation continues. The background agent reports back, and the interactive model surfaces those results in conversation at a natural moment.
The experience target is Gemini Live / GPT-Live: under 1.2s from end of speech to first audio, interruptible mid-sentence, and never silent while thinking.
The deployment target is your own hardware. Talaria owns no weights — every model and voice service is an external HTTP endpoint you point it at.
Full requirements and design: talaria-requirements-and-design.md.
Section references throughout the code and this README point at it.
The repo ships a dummy LLM backend (apps/mock-services) that answers every
endpoint Talaria talks to with canned, deterministic responses: an
OpenAI-compatible chat endpoint that streams tokens and emits grammar-shaped
tool calls, an ASR endpoint, a TTS endpoint that returns real decodable audio at
22.05kHz (so the server's normalizer is genuinely exercised), and a mock Hermes
Gateway with timed task progress.
pnpm install
pnpm build
# Terminal 1 — the dummy LLM backend on :8899
pnpm dev:mock
# Terminal 2 — Talaria on :8790, DemoBackend for background work
TALARIA_TOKEN=demo-token-please-change TALARIA_CONFIG=docs/demo.yaml pnpm dev:server
# Terminal 3 — drive one full turn through the real socket
node --experimental-strip-types scripts/drive-session.mjs "Research the latency budget and write it up."That last command streams synthetic speech at the server and prints what comes back:
you Research the latency budget and write it up.
talaria On it — I'll dig into research the latency budget and write it up.
task task-demo1 [queued] Research the latency budget and write it up.
task task-demo1 [running] Breaking that down into steps.
alert Breaking that down into steps.
The acknowledgement is spoken while the background agent starts (§6.2), the task list updates live, and the alert is held until the room is quiet (§6.3).
The admin panel is at http://127.0.0.1:8790/admin — it will ask you to set a password before it does anything (§12).
The shutter button attaches a photo to your next question, and the server routes it to whichever model you have configured for the foreground role. That is the feature no generic OpenAI client gives you.
The browser client is at http://127.0.0.1:8790/app/ once you have run
pnpm --filter talaria-client build:spa. The browser build uses a
getUserMedia fallback; shipping audio goes through the native plugins for
the reasons in §10.
talaria/
├── apps/
│ ├── client/ Quasar + Capacitor (Android + iOS), native audio plugin
│ ├── admin/ Quasar SPA — server admin panel (§9)
│ ├── server/ Node + TS — the whole pipeline (§5)
│ └── mock-services/ Dummy LLM backend: chat, ASR, TTS, Hermes gateway
├── packages/
│ ├── protocol/ Zod schemas + inferred types — the single source of truth for the wire (§4)
│ ├── config/ Config loading, validation, defaults (§7)
│ └── i18n/ Shared locale catalogs, client + admin (§11)
├── scripts/
│ ├── install-voice/ Voice service installer (§8)
│ ├── drive-session.mjs One-turn smoke test against a running server
│ └── lint-i18n.mjs Fails the build on hardcoded user-facing strings (§11)
├── docs/ Config profiles and design notes
└── fastlane/ Signing and store submission for both platforms (§13)
| Command | What it does |
|---|---|
pnpm build |
Builds the workspace packages and the server |
pnpm test |
The whole suite: protocol, chunker, arbiter, VAD fixtures, integration |
pnpm coverage |
Same, with the §14 coverage floors enforced |
pnpm typecheck |
strict: true across every workspace |
pnpm lint:i18n |
Fails on any hardcoded user-facing string (§11) |
pnpm dev:mock |
The dummy LLM backend |
pnpm dev:server |
The Talaria server |
pnpm dev:admin |
The admin panel with hot reload, proxying /api to :8790 |
pnpm --filter talaria-client dev |
The client in a browser |
pnpm --filter talaria-client dev:android |
The client on a device or emulator |
pnpm --filter talaria-client dev:ios |
The client on a simulator or device |
docs/home-lab.yaml is the reference deployment in §2, filled in and measured:
| Role | Where | Notes |
|---|---|---|
| Interactive | WARLOCK 192.168.1.151:8081, RTX 3060 |
Gemma 4 12B, --reasoning off |
| ASR | WARLOCK :8082, RTX 3050 |
faster-whisper small.en |
| TTS | WARLOCK :8083, RTX 3050 |
Kokoro af_heart, already 24kHz |
| Background | hermes.local 192.168.1.192:8642 |
Hermes async runs |
| Deep thought | INFINITY 192.168.1.99:8080, RTX 4090 |
Qwen 3.8 27B, reasoning_effort: low |
export TALARIA_TOKEN="$(openssl rand -base64 24)"
export HERMES_API_KEY="$(ssh hermes 'grep ^API_SERVER_KEY= ~/.hermes/.env | cut -d= -f2-')"
scripts/make-certs.sh # phones need wss://, see docs/device-clients.md
TALARIA_CONFIG=docs/home-lab.yaml pnpm dev:serverCheck every endpoint before blaming the app:
$ node apps/server/dist/cli/verify-voice.js --config docs/home-lab.yaml --full
Voice round-trip
ok tts 410ms 3250ms of audio, normalized to 24kHz
ok asr 364ms transcribed 44 characters
heard back: "The quick brown fox jumps over the lazy dog."
word recall: 100%
ok foreground 542ms model gemma-4-12b-it answered
ok background 15ms hermes gateway reachable, model hermes-agent
ok deepThought 1133ms model Qwen3.8-27B answered
Measured end to end on that hardware, from the workstation across a subnet boundary, warm:
| Stage | Measured p50 | Budget |
|---|---|---|
| endpointing | 320ms | 300ms |
| asr | 294ms | 250ms |
| foreground TTFT | 629ms | 400ms |
| tts first chunk | 209ms | 250ms |
| end to end | ~1420ms | 1200ms |
About 200ms over, and the gap is foreground TTFT. Running the server on the
same subnet does not close it — measured from HULK, TTFT was 459ms median
versus 397ms from the workstation, so the variance is model-side (prefill plus
grammar compilation), not network. See docs/latency.md.
Copy a profile and edit it:
cp docs/single-endpoint.yaml talaria.config.yaml # everything on one box
cp docs/split-gpu.yaml talaria.config.yaml # the §2 reference deployment
export TALARIA_TOKEN="$(openssl rand -base64 24)"
pnpm dev:serverEvery field is validated, and every invalid field produces a specific message rather than a stack trace:
Invalid Talaria config in talaria.config.yaml
- roles.foreground.endpoint: must be an absolute http(s) URL, e.g. http://warlock.local:8081/v1
- voice.vad.speculativeSilenceMs: must be less than voice.vad.commitSilenceMs (500)
Nobody should have to hand-edit YAML over SSH to change a model endpoint, so the
admin panel writes the same file: validated first, written atomically, previous
version kept as .bak.
sudo scripts/install-voice/install-voice.sh --asr-gpu 0 --tts-gpu 0Detects CUDA and falls back to CPU cleanly, writes systemd units, prints the config block to paste in, and self-verifies by round-tripping a test phrase through both services and reporting measured latency per stage. The verification calls the same code path as the admin panel's Test button, so the installer's "it works" and the panel can never disagree (§8).
If every endpoint points at one machine, the §6.1 latency budget does not
apply unless roles.background.preemptAbort is on — and it is, by default,
because auto resolves to true when foreground and background share a host. A
background generation already in flight otherwise occupies the model server and
the foreground request queues behind it at exactly the wrong moment.
The consequence is a contract on whoever writes background tasks: steps must be idempotent, because an aborted step gets re-run. Anything with side effects must be safe to repeat or checkpointed before the model call, not after.
- On-device inference. Talaria is a client for remote endpoints.
- Your networking. WAN/LAN reachability is your problem;
0.0.0.0plus a forwarded port with no TLS is a bad idea, and the server says so at startup. - Backends other than Hermes Agent. The interface exists, and
DemoBackendproves it, but Hermes is the only real implementation shipped. - Multi-user. One server, one person. Auth exists to stop strangers, not to partition tenants.
- Using the deep-thought model directly. It is health-checked and configured here because the topology in §2 puts it behind the background agent, which is what actually calls it. Talaria never sends a turn to it.
- Locales other than English. The architecture is not English-only (§11) and
there is a test that exercises the chunker interface with a non-Latin
implementation, but only
en-USships.
This is a microphone reachable from the internet, so:
- Auth arrives in the first control frame, not an upgrade header and not a query string. Sockets that have not authenticated within 2 seconds are closed, and connection attempts are rate-limited per source IP.
- The admin panel binds to localhost only by default, has its own password-based login (scrypt, cookie session, CSRF on every mutation, lockout on repeated failures), and refuses to open remotely without a password set.
- No conversation audio or transcripts hit disk unless
debug.persistis on — one class in the server can write conversational content, so the guarantee is a single check rather than a property of every code path. Audio stays opt-in even within debug mode. - Tokens, credentials and transcript content are never logged at default level — there is a test that asserts it.
pnpm test # 385 tests
pnpm coverage # with the §14 floors: 80% overall, 95% on arbiter and protocolThe suite covers what §14 asks for: protocol round-trips and malformed-input rejection for every message type; the sentence chunker against abbreviations, decimals, ellipses, code blocks and missing punctuation; the arbiter's full barge-in state machine including the false-barge-in path; VAD against fixtures containing coughs, door slams, background speech and TTS echo bleed; session resume inside and outside the TTL; the backend adapter against a mock gateway including abort-then-reissue; admin API auth, CSRF and atomic config writes; and a full turn end to end asserting event ordering and that first audio is emitted before generation completes. There is also a parity test that fails the build if a method is added to the audio plugin interface and only implemented on one platform — which is exactly how a platform starts falling behind.
Five things have to be true, and each fails misleadingly:
TLS is mandatory (a webview refuses ws://), the CSP must allow the user's
server, the phone must trust the certificate, the native plugin must be
registered and not just compiled, and the build needs to know the address.
All five, with the symptom each produces, are in
docs/device-clients.md.
See docs/self-build.md. If you cannot get from a clean
clone to an APK by following it, that is a bug — please open an issue.