Skip to content

Start TTS model loads earlier without loading at every startup - #150

Merged
JRufer merged 1 commit into
developmentfrom
feat/tts-preload
Sep 24, 2026
Merged

JRufer merged 1 commit into
developmentfrom
feat/tts-preload

Conversation

@JRufer

@JRufer JRufer commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Part of re-landing 1f13d62. This one is independent and based directly on development.

What changes

  • preload_tts() now also works in always-loaded mode. Without prewarm, that mode loads the model on first use. Recording against a speech target, or a command spotted mid-speech (Detect voice commands while the user is still speaking #148), now starts that load while the user is still talking. The worker ignores the request if the model is already loaded.
  • Switching from on-demand to always-loaded (for example from the tray) loads the model right away when the engine's prewarm is on, and logs a failure instead of hiding it.
  • The TTS worker starts before AppState is built, so tts_handle is set from the start. The speak callback uses it directly when the lock is free, skipping a runtime hop.
  • The "load the selected engine" match is now shared by Preload and the mode switch.

Fixes over the original commit

  • prewarm controls startup loading again. The original loaded the model at every launch in always-loaded mode (the default), including Breeze-TTS-2 at about 5 GB. That made the four prewarm toggles in the TTS settings do nothing.
  • Voice Command Router hotkeys no longer count as speech targets. The first-run wizard binds the first hotkey to the router, so with that change on-demand mode loaded the model on every dictation and effectively never unloaded it.

docs/tts.md is updated.

Test plan

  • cargo test -p voxctrl-tts -p voxctrl-app: 172 + 92 pass
  • Manual: always-loaded with prewarm off, so the model isn't loaded at launch. Dictate to a Say target: loading starts on key-down

🤖 Generated with Claude Code

…rtup

- preload_tts() now also works in always-loaded mode, where a model without
  prewarm is loaded lazily; recording against a speech target (or an early
  spoken command) starts that first load while the user is still talking.
  The worker ignores the request when the model is already resident.
- Switching from on-demand to always-loaded loads the model immediately
  when the engine's prewarm setting is on, and logs a load failure.
- The per-engine prewarm settings still decide whether the model is loaded
  at startup (an earlier draft loaded it unconditionally, including
  Breeze-TTS-2's ~5 GB, and made the prewarm toggles do nothing).
- Voice Command Router hotkeys no longer count as speech targets, so
  on-demand mode doesn't load the model for every ordinary dictation.
- Start the TTS worker before AppState is built so tts_handle is set from
  the start, and let the speak callback use it without a runtime hop when
  it's uncontended.
- Share the "load the selected engine" match between Preload and the mode
  switch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@JRufer JRufer mentioned this pull request Sep 24, 2026
6 tasks
@JRufer
JRufer merged commit 12ea455 into development Sep 24, 2026
5 checks passed
@JRufer
JRufer deleted the feat/tts-preload branch September 28, 2026 11:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant