Always-on LAN voice service on a single AMD Ryzen AI ("Krackan") mini-PC: speech-to-text and text-to-speech served as two OpenAI-compatible HTTP endpoints, deliberately split across two different accelerators so the two workloads never contend for the same silicon.
- STT → AMD XDNA2 NPU via FastFlowLM (Whisper large-v3-turbo)
- TTS → Radeon 840M iGPU via qwentts.cpp (Qwen3-TTS 1.7B, Vulkan)
The host is the audio edge only — ears and mouth. The reasoning happens elsewhere (an agent on the LAN consumes both endpoints); no chat LLM is required on this box.
Interactive version:
docs/architecture.html
LAN clients (192.168.8.0/24)
│
├── POST :8081/v1/audio/transcriptions ──► whisper-proxy ──► FLM ASR :8090 ──► NPU (XDNA2, /dev/accel0)
│ (LAN face) (localhost)
│
└── POST :8092/v1/audio/speech ──────────► tts-server ────────────────────► iGPU (Radeon 840M, /dev/dri)
(LAN face)
| Port | Bind | Service | Device | Purpose |
|---|---|---|---|---|
| 8081 | 0.0.0.0 |
whisper-proxy |
— | LAN-facing STT, OpenAI shape |
| 8090 | 127.0.0.1 |
flm-asr |
NPU | ASR backend (localhost only) |
| 8092 | 0.0.0.0 |
qwen-tts |
iGPU | LAN-facing TTS |
| Role | Model | Quant | Size |
|---|---|---|---|
| STT | whisper-v3:turbo (FLM) |
NPU2 | 623 MB |
| STT co-load | llama3.2:1b (forced by --asr) |
NPU2 | 1.3 GB |
| TTS talker | qwen-talker-1.7b-base |
Q8_0 | 2.0 GB |
| TTS codec | qwen-tokenizer-12hz |
Q8_0 | 278 MB |
RAM footprint ≈ 10 GB of 30 GB.
All figures measured on-host, 2026-08-06. Raw log: benchmarks/stage3-latency.log.
The complete sanitized technical record of the FLM ASR latency investigation — including source review, audio_chunk analysis, instrumented-build stage timings, orderly test procedure, restoration checks, and update review — is available as docs/flm-latency-investigation-session.pdf and its browsable source docs/flm-latency-investigation-session.html.
Key result: the 3,000-frame mel calculation took approximately 6 ms; the fixed ~2.2 s floor was dominated by the NPU Whisper encode_audio() stage on FLM's fixed 30-second acoustic window.
| Path | Metric | Value |
|---|---|---|
| STT (NPU) | 2.5 s utterance | 2.40 s |
| STT (NPU) | 11 s clip | 2.75 s |
| TTS (iGPU) | full WAV | 1.23 s |
| TTS (iGPU) | time to first audio (streaming PCM) | 0.20 s |
| End-to-end | STT complete → TTS first audio | 3.73 s |
One STT and one TTS request running strictly simultaneously, versus each running alone:
| Workload | Idle | Concurrent | Delta |
|---|---|---|---|
| STT | 2.379 s | 2.422 s | +1.8 % |
| TTS (total) | 1.404 s | 1.450 s | +3.3 % |
| TTS (TTFA) | 0.202 s | 0.203 s | +0.8 % |
Confirmed structurally by open file descriptors — zero overlap:
flm (STT): 1 × /dev/accel0 0 × /dev/dri
tts-server (TTS): 0 × /dev/accel0 2 × /dev/dri
Flood tests agree: 5 concurrent TTS requests serialise (--max-batch 1) while STT is
unaffected (+0.9 %); 5 concurrent STT requests queue on the NPU while TTS is unaffected (+5.8 %).
End-to-end is 3.73 s against a <1 s working target. The cause is not the device:
| Audio length | STT wall time |
|---|---|
| 0.5 s | 2.16 s |
| 2.5 s | 2.53 s |
| 11 s | 2.75 s |
Latency is nearly flat against input length — roughly 2.2 s of fixed per-request overhead
inside the ASR runtime, not compute. The proxy contributes nothing (2.23 s direct vs 2.23 s
proxied). A Vulkan whisper.cpp backend on the same host shows the same pathology (~1.68 s
fixed), so this is a large-v3-turbo request-overhead characteristic rather than an
NPU-vs-iGPU question. A smaller ASR model is the lever, not a different accelerator.
TTS is already conversational at 0.20 s TTFA when streaming.
On this hardware the NPU is slower than the iGPU for the same ASR model:
| Audio | NPU (FLM) | iGPU (whisper.cpp/Vulkan) |
|---|---|---|
| 2.5 s | 2.32 s | 1.62 s |
| 11 s | 2.70 s | 1.74 s |
| 55 s | 8.16 s | 5.62 s |
| CPU cost/req | 1.03 cpu-s | 0.41 cpu-s |
The NPU was chosen anyway, deliberately: it keeps STT off the iGPU so TTS owns that device outright. The isolation numbers above are what that ~1.5× latency buys.
Requires: FastFlowLM (flm) on PATH, a built qwentts.cpp, and the GGUF models.
# 1. memlock — FLM cannot initialise the NPU under the default 8 MB limit
sudo install -m 644 config/99-flm-memlock.conf /etc/security/limits.d/
# (re-login; verify with: ulimit -l → unlimited)
flm validate # must report "Memlock Limit: infinity"
# 2. services
sudo install -m 644 systemd/*.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now flm-asr qwen-tts whisper-proxy
# 3. open the TTS port to the LAN (STT 8081 assumed already allowed)
sudo ufw allow from 192.168.8.0/24 to any port 8092 proto tcpEvery unit sets LimitMEMLOCK=infinity; the whole stack is verified to come back after a
cold reboot.
# STT
curl -X POST http://<host>:8081/v1/audio/transcriptions \
-F file=@clip.wav -F model=whisper-1
# TTS → file
curl -X POST http://<host>:8092/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3-tts-1.7b","input":"Hello.","response_format":"wav"}' \
-o out.wav
# TTS → streaming PCM (0.20 s to first audio); omit response_formatThere are no built-in voices — a voice is a reference clip you register, and the
model clones it. The bundled scripts/slug-voice-push does the durable version:
it converts the sample (mono 16 kHz), copies it into the voice store on the TTS
host, registers the voice via the API, and verifies it actually speaks.
# ICL clone mode (supply the exact transcript of the clip — better fidelity)
slug-voice-push myvoice ~/Downloads/reference.wav \
--ref-text "the exact words spoken in the recording"
# base clone mode (no transcript)
slug-voice-push myvoice ~/Downloads/reference.wav --no-ref-text
# inspect what is live vs what is persisted
slug-voice-push --list
# make it the default Hermes TTS voice
slug-voice-push --set-default myvoiceReboot persistence (important). Registered voices live in the server's memory
only. A restart or reboot silently drops them, and TTS then returns HTTP 200 with
zero bytes — no error. qwen-tts.service runs scripts/slug-voice-restore as
ExecStartPost, which re-registers every voice from ~/.config/slug-voices/ on the
TTS host. slug-voice-push writes the sample + metadata there, so a voice added
once survives every reboot unattended. The leading - on that line means a restore
failure can never take the TTS service down.
ref_text must match. A transcript from a different clip makes ICL run away —
measured 163.8 s of noise for a 10-word line versus 4.0 s correct. slug-voice-push
detects the runaway (audio far longer than the probe implies) and refuses to leave a
broken voice configured.
Then speak it by passing "voice": "myvoice" in the speech request, or set it as the
Hermes default above.
flm serve --asr 1is not ASR-only. It requires an LLM tag and silently pullsllama3.2:1b(1.2 GB) onto the NPU alongside Whisper. Budget the RAM.- Memlock is a hard blocker. Under the default 8 MB,
flm validatefails outright and the NPU never initialises. - TTS streams by default. No
response_formatyields chunkedaudio/pcm, so naive clients writing to a file get 0 bytes with HTTP 200. Pass"response_format":"wav". - Base TTS models have no named voices. Sending
"voice":"default"returnsunknown voice 'default'. Omit the field, or register a clone first. - The proxy needed a route fix. whisper.cpp exposes
/inference; FLM exposes/v1/audio/transcriptions. The script now readsBACKEND_PATH(default/inference, so existing whisper.cpp deployments are unaffected). - systemd
Wants=cannot be reliably cleared from a drop-in when the same file also adds newWants=entries — replace the unit file instead. - Connecting to Hermes TTS.
hermes config set tts.provider openai,tts.openai.base_url http://<host>:8092/v1,tts.openai.voice <name>(or empty for the base voice),tts.openai.api_key not-needed. Thetext_to_speechtool derives its output format from the file extension and defaults tomp3, which this backend rejects (400 response_format must be 'pcm' or 'wav') — pass anoutput_pathending in.wav. Streaming voice replies are unaffected (they hardcodepcm). The model name Hermes sends (gpt-4o-mini-tts) is ignored by the server. - Voices vanish on restart. Covered above — use
slug-voice-push(writes to the reboot-proof store) and rely onslug-voice-restoreviaExecStartPost. A wrong voice name returns HTTP 200 with 0 bytes, so verify by RMS, not by status code.
systemd/ flm-asr, qwen-tts, whisper-proxy unit files
config/ memlock limits drop-in
scripts/ whisper_openai_proxy.py (OpenAI shape → backend, route-configurable)
slug-voice-push (install a WAV as a reboot-proof cloned voice)
slug-voice-restore (ExecStartPost hook: re-register voices at boot)
benchmarks/ measurement harnesses + raw latency log
docs/ architecture diagram (HTML + PNG)
No credentials, tokens, model weights, or private hostnames are stored in this repository. Both endpoints are unauthenticated plain HTTP and are intended for a trusted LAN only — do not expose them to the internet.
