A lightweight, cross-platform realtime dictation app that types what you speak into any focused text field.
- 🎤 Real-time speech-to-text dictation
- 📝 Works in any app (browser, editor, chat, etc.)
- 🔒 Local-only processing — your voice never leaves your machine
- ⚡ Lightweight — runs on CPU, no GPU required
- 🫧 Tiny floating bubble UI — always accessible, never in the way
- 🔌 Pluggable ASR backends (faster-whisper, more coming)
(Coming soon)
Mic Input → VAD → Audio Buffer → ASR (faster-whisper) → Text Injection → Active App
↓ ↑
[phrase detection] [partial/final display]
- Python 3.11 or 3.12 (NOT 3.14)
- Windows 10/11 or Linux (X11)
- Microphone
# 1. Clone
git clone https://github.com/t957095/meow.git
cd meow
# 2. Create virtual environment
python -m venv .venv
.venv\Scripts\Activate.ps1
# 3. Install
pip install -e ".[dev]"
# 4. Run
python -m meow# 1. Clone
git clone https://github.com/t957095/meow.git
cd meow
# 2. Create virtual environment
python3 -m venv .venv
source .venv/bin/activate
# 3. Install system dependencies (Ubuntu/Debian)
sudo apt-get install portaudio19-dev xdotool
# 4. Install
pip install -e ".[dev]"
# 5. Run
python -m meow- Launch Meow — a small bubble appears
- Click the 🎤 button to start dictation
- Speak naturally
- Gray text = processing, White text = finalized
- Finalized text is automatically typed into the focused app
- Click ⏹ to stop
Create a .env file in the project root:
MEOW_ASR_BACKEND=faster_whisper
MEOW_MODEL_SIZE=base
MEOW_DEVICE=auto
MEOW_LANGUAGE=en
MEOW_VAD_AGGRESSIVENESS=2
MEOW_PHRASE_TIMEOUT=1.5| Model | Speed | Accuracy | VRAM | Best For |
|---|---|---|---|---|
| tiny | Fastest | Basic | ~1GB | Quick testing |
| base | Fast | Good | ~1GB | Default |
| small | Moderate | Better | ~2GB | Accuracy priority |
- No telemetry — zero data collection
- No cloud — all processing is local
- No audio storage — audio is processed in memory and discarded
- No logs of transcript content — only timing and status
- Open source — audit the code yourself
# Run tests
pytest tests/ -v
# Lint
ruff check src/
ruff format src/
# Type check
mypy src/meow/- Phase 1: Speech-to-text dictation
- Phase 2: TTS (text-to-speech) with VibeVoice-Realtime
- Phase 3: Custom vocabulary / hotwords
- Phase 4: Global hotkeys
- Phase 5: Linux Wayland support
Uses faster-whisper with CTranslate2 for efficient CPU/GPU inference.
Microsoft's VibeVoice-ASR (9B params) is designed for batch transcription, not real-time dictation. May be added as an optional backend.
Important: microsoft/VibeVoice-Realtime-0.5B is a text-to-speech model (speaks text aloud), NOT speech-to-text. It will be used for Phase 2 TTS features, not dictation.
MIT License — see LICENSE
Contributions welcome! Please read our Contributing Guide (coming soon).