Skip to content

[🐞 BUG] "Faster first result" doesn't keep Fluid-1 ready β€” first dictation after launch stalls ~8.5 s (model loads only on hotkey press)Β #705

Description

@matejcermak

Description

With "Faster first result" enabled, the first dictation after every app launch still stalls for ~8.5 s before the enhanced text appears. The toggle's tooltip says "Keeps Fluid-1 ready so your first dictation finishes sooner", but Fluid-1 is not actually kept ready: the MLX helper (fluid-intelligence-mlx serve-json) sits at ~0 MB RSS after app launch, and the model only starts loading when the dictation hotkey is pressed.

Environment

  • FluidVoice 1.6.5 (16), MLX backend, model fluid-1-nvfp4-mlx + MTP drafter gemma-4-E2B-it-qat-assistant-bf16-mlx-mtp (draft-block-size 6)
  • macOS 26.5.2, MacBook Pro M1 Pro, 16 GB
  • "Faster first result" (prefix KV cache) ON

Measurements

Cold pipeline in a fresh helper process (fluid-intelligence-mlx run, warm OS caches):

latency_ms=9307  ttft_ms=8550  generated_tokens=8  tokens_per_second=10.57
  • warmup subcommand alone (model + drafter load): ~4.7 s β€” the remaining ~4 s is the first prefill of the ~1300-token dictation prompt + kernel warm-up.
  • Warm requests afterwards are great: totalMs=1336–3914, ttftMs=836–922, tps=25–33 (prefix cache clearly working).

So the ~8.5 s is paid exactly once per app launch β€” but it lands on the user's first dictation. Short first utterances (a few words) feel the entire stall; only utterances longer than ~9 s of speech hide it, because prewarm overlaps with recording.

Cause (from source)

prewarmDictation() is only triggered from prewarmPrivateAIDictationIfNeeded, which is called inside onCaptureStarted β€” i.e. at hotkey press, never at app launch:

Since there is no idle unload (unloadCachedRuntime only fires on termination/config change), preloading at launch would have no recurring memory downside beyond holding the weights β€” which is what the toggle already implies.

Expected behavior / proposal

When "Faster first result" is ON and the backend is the local FluidIntelligence bridge, call prewarmDictation() once at app launch (and after model/backend switches), not just at capture start. That would make the toggle's behavior match its tooltip and remove the once-per-launch 8.5 s stall.

If holding ~3.3 GB from launch is a concern on 8 GB machines, gating the launch-time prewarm on available memory (or making it a separate "Preload at launch" option) would still solve it for machines that can afford it.

Workarounds found

  • Make the first dictation after launch a long one (>9 s of speech) so prewarm overlaps with recording.
  • External fluid-intelligence-mlx run at login primes the OS page cache / Metal shader cache, but can't warm the app's own stdio-connected server process, so it only trims the post-reboot worst case.

Happy to test a build β€” and thanks for open-sourcing this; the codebase was a pleasure to read through.

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions