Detect voice commands while the user is still speaking - #148
Merged
Merged
Conversation
This was referenced Sep 24, 2026
Merged
JRufer
added this pull request to stack #154
September 24, 2026 01:24
Run short interim transcription passes over the opening of a recording so a
spoken command ("Hey Vox, say ...") is recognised mid-speech: the command
overlay appears early and the TTS model starts loading before release.
Compared with the first cut of this feature:
- Interim passes skip the hotkey's OpenAI post-processing, which otherwise
made a billed request every ~400ms while recording.
- Only the first 6s of audio is transcribed early, so passes stay short,
the final transcription never waits long behind one, and long dictations
don't keep the model busy for their whole length.
- Interim passes run only when a command could match (a non-router target
exists) and never for remote or CPU medium/large Whisper backends.
- The early match is a head start only. Routing is decided by the final
transcript alone, which must carry the trigger itself, so a misheard
interim ("hey folks" -> "hey box") can no longer send ordinary dictation
to a command target.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…detection off - When an interim pass announced a command but the final transcript doesn't confirm it (no command, empty result or a transcription error), the command overlay is taken down right away through a new "command-withdrawn" event instead of lingering for its full duration. - New features.early_command_detection setting (default on; older configs read as on), with a Settings → Features → Voice Commands toggle and docs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
JRufer
force-pushed
the
feat/interim-command-detection
branch
from
September 24, 2026 01:28
f1f2fdf to
62f8997
Compare
6 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part 2 of re-landing 1f13d62. Stacked on #147: the base is
feat/hey-vox-trigger, and GitHub retargets it todevelopmentonce #147 merges.While you're still holding the hotkey, short interim transcription passes run over the start of the recording. A spoken command ("Hey Vox, say …") then shows its overlay and starts loading the TTS model before you let go.
Fixes over the original commit
command-withdrawnevent takes the overlay pill down right away.features.early_command_detection: default on, and older configs read as on. It's under Settings → Features → Voice Commands and documented indocs/configuration.mdanddocs/api.md.Test plan
cargo testforvoxctrl-app(92),voxctrl-inference(67),voxctrl-config(26, including a default-on test for older configs),voxctrl-routing(69)npx vitest run: 240 pass;svelte-check: 0 errors🤖 Generated with Claude Code