Skip to content

Detect voice commands while the user is still speaking - #148

Merged
JRufer merged 2 commits into
developmentfrom
feat/interim-command-detection
Sep 24, 2026
Merged

JRufer merged 2 commits into
developmentfrom
feat/interim-command-detection

Conversation

@JRufer

@JRufer JRufer commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Part 2 of re-landing 1f13d62. Stacked on #147: the base is feat/hey-vox-trigger, and GitHub retargets it to development once #147 merges.

While you're still holding the hotkey, short interim transcription passes run over the start of the recording. A spoken command ("Hey Vox, say …") then shows its overlay and starts loading the TTS model before you let go.

Fixes over the original commit

  • No billed OpenAI calls from interim passes. Interim requests carry the hotkey's binding, so each pass (about every 400ms) ran that hotkey's OpenAI rewrite and threw the result away.
  • Bounded cost. Only the first 6s of audio are transcribed early. The final transcription never waits long behind a pass, and long dictations don't keep the model busy for their whole length. The original re-transcribed all audio so far on every pass.
  • Only runs when useful. Passes are skipped when no command target exists, for the remote backend, and for CPU-only medium/large Whisper models.
  • The early match can't route text. The final transcript alone decides routing and must contain the trigger itself. The original fell back to the early match whenever the final text contained the target name anywhere, so "Hey folks, say how are you" could go to a Say target.
  • Unconfirmed commands are withdrawn. If the final transcript doesn't confirm an early command (no command, empty result or an error), a new command-withdrawn event takes the overlay pill down right away.
  • New setting features.early_command_detection: default on, and older configs read as on. It's under Settings → Features → Voice Commands and documented in docs/configuration.md and docs/api.md.

Test plan

  • cargo test for voxctrl-app (92), voxctrl-inference (67), voxctrl-config (26, including a default-on test for older configs), voxctrl-routing (69)
  • npx vitest run: 240 pass; svelte-check: 0 errors
  • Manual: say "Hey Vox, say hello" while holding the hotkey. The overlay should appear before release, and dictation without a trigger should behave as before

🤖 Generated with Claude Code

@JRufer
JRufer added this pull request to stack #154 September 24, 2026 01:24
Base automatically changed from feat/hey-vox-trigger to development September 24, 2026 01:28
JRufer and others added 2 commits September 23, 2026 20:28
Run short interim transcription passes over the opening of a recording so a
spoken command ("Hey Vox, say ...") is recognised mid-speech: the command
overlay appears early and the TTS model starts loading before release.

Compared with the first cut of this feature:
- Interim passes skip the hotkey's OpenAI post-processing, which otherwise
  made a billed request every ~400ms while recording.
- Only the first 6s of audio is transcribed early, so passes stay short,
  the final transcription never waits long behind one, and long dictations
  don't keep the model busy for their whole length.
- Interim passes run only when a command could match (a non-router target
  exists) and never for remote or CPU medium/large Whisper backends.
- The early match is a head start only. Routing is decided by the final
  transcript alone, which must carry the trigger itself, so a misheard
  interim ("hey folks" -> "hey box") can no longer send ordinary dictation
  to a command target.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…detection off

- When an interim pass announced a command but the final transcript
  doesn't confirm it (no command, empty result or a transcription error),
  the command overlay is taken down right away through a new
  "command-withdrawn" event instead of lingering for its full duration.
- New features.early_command_detection setting (default on; older configs
  read as on), with a Settings → Features → Voice Commands toggle and docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@JRufer
JRufer force-pushed the feat/interim-command-detection branch from f1f2fdf to 62f8997 Compare September 24, 2026 01:28
@JRufer
JRufer merged commit 362b37a into development Sep 24, 2026
5 checks passed
@JRufer
JRufer deleted the feat/interim-command-detection branch September 28, 2026 11:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant