Skip to content

Latest commit

 

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

slug-tts-stt

Always-on LAN voice service on a single AMD Ryzen AI ("Krackan") mini-PC: speech-to-text and text-to-speech served as two OpenAI-compatible HTTP endpoints, deliberately split across two different accelerators so the two workloads never contend for the same silicon.

  • STT → AMD XDNA2 NPU via FastFlowLM (Whisper large-v3-turbo)
  • TTS → Radeon 840M iGPU via qwentts.cpp (Qwen3-TTS 1.7B, Vulkan)

The host is the audio edge only — ears and mouth. The reasoning happens elsewhere (an agent on the LAN consumes both endpoints); no chat LLM is required on this box.

architecture

Interactive version: docs/architecture.html


Topology

LAN clients (192.168.8.0/24)
        │
        ├── POST :8081/v1/audio/transcriptions ──► whisper-proxy ──► FLM ASR :8090 ──► NPU  (XDNA2, /dev/accel0)
        │                                          (LAN face)        (localhost)
        │
        └── POST :8092/v1/audio/speech ──────────► tts-server ────────────────────► iGPU (Radeon 840M, /dev/dri)
                                                   (LAN face)
Port Bind Service Device Purpose
8081 0.0.0.0 whisper-proxy — LAN-facing STT, OpenAI shape
8090 127.0.0.1 flm-asr NPU ASR backend (localhost only)
8092 0.0.0.0 qwen-tts iGPU LAN-facing TTS

Models

Role Model Quant Size
STT whisper-v3:turbo (FLM) NPU2 623 MB
STT co-load llama3.2:1b (forced by --asr) NPU2 1.3 GB
TTS talker qwen-talker-1.7b-base Q8_0 2.0 GB
TTS codec qwen-tokenizer-12hz Q8_0 278 MB

RAM footprint ≈ 10 GB of 30 GB.


Measured results

All figures measured on-host, 2026-08-06. Raw log: benchmarks/stage3-latency.log.

Detailed FLM latency investigation

The complete sanitized technical record of the FLM ASR latency investigation — including source review, audio_chunk analysis, instrumented-build stage timings, orderly test procedure, restoration checks, and update review — is available as docs/flm-latency-investigation-session.pdf and its browsable source docs/flm-latency-investigation-session.html.

Key result: the 3,000-frame mel calculation took approximately 6 ms; the fixed ~2.2 s floor was dominated by the NPU Whisper encode_audio() stage on FLM's fixed 30-second acoustic window.

Latency

Path Metric Value
STT (NPU) 2.5 s utterance 2.40 s
STT (NPU) 11 s clip 2.75 s
TTS (iGPU) full WAV 1.23 s
TTS (iGPU) time to first audio (streaming PCM) 0.20 s
End-to-end STT complete → TTS first audio 3.73 s

Device isolation — the point of the design

One STT and one TTS request running strictly simultaneously, versus each running alone:

Workload Idle Concurrent Delta
STT 2.379 s 2.422 s +1.8 %
TTS (total) 1.404 s 1.450 s +3.3 %
TTS (TTFA) 0.202 s 0.203 s +0.8 %

Confirmed structurally by open file descriptors — zero overlap:

flm        (STT): 1 × /dev/accel0   0 × /dev/dri
tts-server (TTS): 0 × /dev/accel0   2 × /dev/dri

Flood tests agree: 5 concurrent TTS requests serialise (--max-batch 1) while STT is unaffected (+0.9 %); 5 concurrent STT requests queue on the NPU while TTS is unaffected (+5.8 %).

Known gap — the <1 s conversational target is not met

End-to-end is 3.73 s against a <1 s working target. The cause is not the device:

Audio length STT wall time
0.5 s 2.16 s
2.5 s 2.53 s
11 s 2.75 s

Latency is nearly flat against input length — roughly 2.2 s of fixed per-request overhead inside the ASR runtime, not compute. The proxy contributes nothing (2.23 s direct vs 2.23 s proxied). A Vulkan whisper.cpp backend on the same host shows the same pathology (~1.68 s fixed), so this is a large-v3-turbo request-overhead characteristic rather than an NPU-vs-iGPU question. A smaller ASR model is the lever, not a different accelerator.

TTS is already conversational at 0.20 s TTFA when streaming.

Device choice, honestly

On this hardware the NPU is slower than the iGPU for the same ASR model:

Audio NPU (FLM) iGPU (whisper.cpp/Vulkan)
2.5 s 2.32 s 1.62 s
11 s 2.70 s 1.74 s
55 s 8.16 s 5.62 s
CPU cost/req 1.03 cpu-s 0.41 cpu-s

The NPU was chosen anyway, deliberately: it keeps STT off the iGPU so TTS owns that device outright. The isolation numbers above are what that ~1.5× latency buys.


Install

Requires: FastFlowLM (flm) on PATH, a built qwentts.cpp, and the GGUF models.

# 1. memlock — FLM cannot initialise the NPU under the default 8 MB limit
sudo install -m 644 config/99-flm-memlock.conf /etc/security/limits.d/
#    (re-login; verify with: ulimit -l  →  unlimited)
flm validate            # must report "Memlock Limit: infinity"

# 2. services
sudo install -m 644 systemd/*.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now flm-asr qwen-tts whisper-proxy

# 3. open the TTS port to the LAN (STT 8081 assumed already allowed)
sudo ufw allow from 192.168.8.0/24 to any port 8092 proto tcp

Every unit sets LimitMEMLOCK=infinity; the whole stack is verified to come back after a cold reboot.

Usage

# STT
curl -X POST http://<host>:8081/v1/audio/transcriptions \
     -F file=@clip.wav -F model=whisper-1

# TTS  → file
curl -X POST http://<host>:8092/v1/audio/speech \
     -H 'Content-Type: application/json' \
     -d '{"model":"qwen3-tts-1.7b","input":"Hello.","response_format":"wav"}' \
     -o out.wav

# TTS  → streaming PCM (0.20 s to first audio); omit response_format

Voice cloning (zero-shot)

There are no built-in voices — a voice is a reference clip you register, and the model clones it. The bundled scripts/slug-voice-push does the durable version: it converts the sample (mono 16 kHz), copies it into the voice store on the TTS host, registers the voice via the API, and verifies it actually speaks.

# ICL clone mode (supply the exact transcript of the clip — better fidelity)
slug-voice-push myvoice ~/Downloads/reference.wav \
    --ref-text "the exact words spoken in the recording"

# base clone mode (no transcript)
slug-voice-push myvoice ~/Downloads/reference.wav --no-ref-text

# inspect what is live vs what is persisted
slug-voice-push --list

# make it the default Hermes TTS voice
slug-voice-push --set-default myvoice

Reboot persistence (important). Registered voices live in the server's memory only. A restart or reboot silently drops them, and TTS then returns HTTP 200 with zero bytes — no error. qwen-tts.service runs scripts/slug-voice-restore as ExecStartPost, which re-registers every voice from ~/.config/slug-voices/ on the TTS host. slug-voice-push writes the sample + metadata there, so a voice added once survives every reboot unattended. The leading - on that line means a restore failure can never take the TTS service down.

ref_text must match. A transcript from a different clip makes ICL run away — measured 163.8 s of noise for a 10-word line versus 4.0 s correct. slug-voice-push detects the runaway (audio far longer than the probe implies) and refuses to leave a broken voice configured.

Then speak it by passing "voice": "myvoice" in the speech request, or set it as the Hermes default above.


Gotchas worth knowing

  1. flm serve --asr 1 is not ASR-only. It requires an LLM tag and silently pulls llama3.2:1b (1.2 GB) onto the NPU alongside Whisper. Budget the RAM.
  2. Memlock is a hard blocker. Under the default 8 MB, flm validate fails outright and the NPU never initialises.
  3. TTS streams by default. No response_format yields chunked audio/pcm, so naive clients writing to a file get 0 bytes with HTTP 200. Pass "response_format":"wav".
  4. Base TTS models have no named voices. Sending "voice":"default" returns unknown voice 'default'. Omit the field, or register a clone first.
  5. The proxy needed a route fix. whisper.cpp exposes /inference; FLM exposes /v1/audio/transcriptions. The script now reads BACKEND_PATH (default /inference, so existing whisper.cpp deployments are unaffected).
  6. systemd Wants= cannot be reliably cleared from a drop-in when the same file also adds new Wants= entries — replace the unit file instead.
  7. Connecting to Hermes TTS. hermes config set tts.provider openai, tts.openai.base_url http://<host>:8092/v1, tts.openai.voice <name> (or empty for the base voice), tts.openai.api_key not-needed. The text_to_speech tool derives its output format from the file extension and defaults to mp3, which this backend rejects (400 response_format must be 'pcm' or 'wav') — pass an output_path ending in .wav. Streaming voice replies are unaffected (they hardcode pcm). The model name Hermes sends (gpt-4o-mini-tts) is ignored by the server.
  8. Voices vanish on restart. Covered above — use slug-voice-push (writes to the reboot-proof store) and rely on slug-voice-restore via ExecStartPost. A wrong voice name returns HTTP 200 with 0 bytes, so verify by RMS, not by status code.

Repository layout

systemd/     flm-asr, qwen-tts, whisper-proxy unit files
config/      memlock limits drop-in
scripts/     whisper_openai_proxy.py (OpenAI shape → backend, route-configurable)
             slug-voice-push      (install a WAV as a reboot-proof cloned voice)
             slug-voice-restore  (ExecStartPost hook: re-register voices at boot)
benchmarks/  measurement harnesses + raw latency log
docs/        architecture diagram (HTML + PNG)

Security

No credentials, tokens, model weights, or private hostnames are stored in this repository. Both endpoints are unauthenticated plain HTTP and are intended for a trusted LAN only — do not expose them to the internet.

About

LAN voice service on AMD Ryzen AI: STT on the XDNA2 NPU (FastFlowLM/Whisper) + TTS on the Radeon iGPU (qwentts.cpp/Qwen3-TTS), split across accelerators so they never contend. Measured, systemd-persisted, reboot-verified.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages