Repository navigation
windows_build
| Tool | Version | Download |
|---|---|---|
| Rust (via rustup) | 1.75+ | https://rustup.rs/ |
| Node.js | 18+ | https://nodejs.org/ |
| Visual Studio Build Tools | 2019+ | https://visualstudio.microsoft.com/visual-cpp-build-tools/ |
| WebView2 Runtime | Any | Pre-installed on Windows 10 21H2+ and Windows 11 |
During installation, select the "Desktop development with C++" workload. This provides MSVC, the Windows SDK, and the linker required by Rust.
After installation, run builds from a Visual Studio Developer Command Prompt or ensure cl.exe is on your PATH. The easiest way to ensure this is to install and use rustup with the default stable-x86_64-pc-windows-msvc toolchain.
cargo install tauri-clinpm installnpm run tauri devThis starts the Vite dev server and compiles the Rust backend in debug mode. Hot-reload is active for Svelte changes; Rust changes require a recompile (~5–30s).
npm run tauri buildOutput artifacts land in src-tauri\target\release\bundle\:
-
nsis\VoxCtrl_<version>_x64-setup.exe— NSIS installer -
msi\VoxCtrl_<version>_x64.msi— MSI package
This is what the published Windows release actually ships — the CPU-only build above and the WebGPU build measure within 0.1 MB of each other, so there is no size cost to always enabling it, and it falls back to the CPU cleanly when no usable Direct3D 12 GPU is found.
To build it, accelerating Moonshine speech recognition via ONNX Runtime's WebGPU execution provider over Direct3D 12:
npx tauri build --bundles nsis --features moonshine-webgpu,parakeet-webgpuThis accelerates Moonshine and Parakeet on any modern Direct3D 12 capable GPU (NVIDIA, AMD, Intel) with no vendor SDK needed at build time.
If you have an NVIDIA GPU and the CUDA Toolkit installed (11.x or 12.x), you can enable GPU-accelerated inference:
# Set CUDA path if needed (adjust to your installed version)
$env:CUDA_PATH = "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.3"
$env:CUDA_COMPUTE_CAP = "86" # Set to your GPU's compute capability
npm run tauri build -- --features cudaWithout the cuda feature flag, Whisper inference runs on the CPU. The cuda
feature is opt-in and never required.
whisper.cpp Vulkan note: whisper.cpp's Vulkan
backend fails to register on Windows MSVC static builds — which is what
whisper-rs-sys produces for Rust — and falls back to the CPU while reporting
that it found a GPU (whisper.cpp#3750).
Because of this, the published Windows installer runs Whisper/Parakeet on the CPU while accelerating Moonshine via Direct3D 12, and the S1-mini sidecar utilizes Vulkan/CPU.
A PowerShell helper script automates prerequisite checks and the build:
# Standard build
.\scripts\build_windows.ps1
# With CUDA
.\scripts\build_windows.ps1 -Cuda
# Debug build
.\scripts\build_windows.ps1 -DebugPlace .bin model files in %LOCALAPPDATA%\voxctrl\models\ (created on first run).
Download a model manually:
$model = "ggml-large-v3.bin"
$url = "https://huggingface.co/ggerganov/whisper.cpp/resolve/main/$model"
$dest = "$env:LOCALAPPDATA\voxctrl\models\$model"
New-Item -ItemType Directory -Force (Split-Path $dest) | Out-Null
Invoke-WebRequest $url -OutFile $destSupported sizes: tiny, base, small, medium, large-v3, large-v3-turbo.
Piper is the default TTS engine, and on Windows it is the one thing VoxCtrl cannot install for you yet. Download the Windows release and place the binary at:
%LOCALAPPDATA%\voxctrl\piper\piper.exe
Download from: https://github.com/rhasspy/piper/releases
Voice models go in %LOCALAPPDATA%\voxctrl\piper-voices\. The Settings UI has a
download button for each supported voice, and those do work on Windows — it is
only the engine binary that has to be placed by hand.
Or pick an engine that needs nothing. Settings → Text-to-Speech offers
Pocket-TTS and Breeze-TTS-2, which are pure Rust plus ONNX and work out of the box.
eSpeak-NG and Inflect-Micro both shell out to espeak-ng, which is not on a stock
Windows machine, so they need it installed and on PATH first.
VoxCtrl types dictated text with the Win32 SendInput API, one KEYEVENTF_UNICODE
event per UTF-16 code unit (see crates/voxctrl-winput). The character travels as
itself, so no keyboard layout is consulted and no escaping layer can misread it.
Transcriptions longer than 2000 characters go via the clipboard and a synthesised Ctrl+V instead, because the receiving application processes one message per character and a long paragraph visibly crawls into editors that do syntax work per keystroke. The previous clipboard contents are restored afterwards.
Every synthesised event carries a marker in dwExtraInfo so VoxCtrl's own
keyboard hook ignores it. Without that, dictating text that completes a shortcut
would re-trigger that shortcut from VoxCtrl's own output.
Previously: this path shelled out to PowerShell and called
SendKeys::SendWait. The payload was base64-encoded so that no shell metacharacter could escape the string — a real defence, and it worked — butSendKeysthen applied its own escaping to the decoded text, in which+ ^ % ~ ( ) { } [ ]are syntax. "50% (a+b)" arrived as "50" plus two stray chords and "array[0]" as "array0", so any dictation containing ordinary punctuation came out wrong.
SendInput cannot deliver into a window owned by a more-privileged process
(Windows calls this UIPI). If dictation works everywhere except one application,
that application is almost certainly running elevated; VoxCtrl has to be elevated
too to type into it. VoxCtrl reports this rather than failing silently.
Shortcuts arrive through a Win32 low-level keyboard hook
(crates/voxctrl-hotkeys/src/windows/). Keys are identified by scan code,
not virtual-key code, because scan codes are physical positions: a shortcut
recorded on the key left of S fires from that key whatever the layout calls it,
which is what KeyboardEvent.code in the settings UI records. The mapping from
scan code to VoxCtrl's key names lives in crates/voxctrl-hotkeys/src/win_keys.rs
and is deliberately outside the platform gate so its tests run on every platform.
All four gesture styles work — hold, toggle, double-tap, double-tap-hold — the same as on Linux.
Two limits are inherent to the mechanism and cannot be worked around:
- Windows does not deliver keys to the hook while an elevated application has focus, unless VoxCtrl is elevated too.
- The hook never sees the secure desktop — the UAC prompt, the lock screen, Ctrl+Alt+Del — so shortcuts do not fire there.
The non-modifier key that completes a shortcut is swallowed, so binding
Super+Space does not also open Windows Search. Modifiers are never swallowed:
eating a bare Ctrl or Super would break every other shortcut on the machine.
A low-level keyboard hook is called for every keystroke on the machine — the same exposure as the evdev and X11 backends on Linux, and unlike the XDG portal, where the compositor owns the grab and VoxCtrl is told only that its own shortcut fired. The Hotkeys tab says so plainly on Windows and does not show the padlock it shows for the portal. Keys are matched against your shortcuts and discarded; nothing is stored, logged, or sent anywhere.
Without a code signing certificate, Windows SmartScreen will display an "Unknown Publisher" warning when users run the installer.
To sign your build:
- Obtain an Authenticode certificate (EV certificate eliminates SmartScreen entirely; standard OV certificates reduce warnings after enough users install the app).
- Set the following environment variables before building:
$env:TAURI_SIGNING_PRIVATE_KEY = "path\to\key.pem"
$env:TAURI_SIGNING_PRIVATE_KEY_PASSWORD = "your-passphrase"See the Tauri signing docs for full details.
MSVC linker is missing. Install Visual Studio Build Tools with the C++ workload, or run rustup target add x86_64-pc-windows-msvc.
Wrong target selected. Ensure rustup default stable-x86_64-pc-windows-msvc.
Download the WebView2 Evergreen Bootstrapper from Microsoft and run it before launching the app. On Windows 11 and Windows 10 21H2+, WebView2 ships with the OS.
- Confirm
CUDA_PATHpoints to an installed CUDA Toolkit. - Ensure Visual Studio Build Tools are installed (whisper-rs compiles CUDA kernels with MSVC).
- Try without
--features cudato rule out a non-CUDA issue first.
VoxCtrl uses WASAPI via cpal. If no microphone is listed, check Windows privacy settings: Settings → Privacy & security → Microphone → Allow apps to access your microphone.