Skip to content

Repository files navigation

Meow 🎤

A lightweight, cross-platform realtime dictation app that types what you speak into any focused text field.

Features

  • 🎤 Real-time speech-to-text dictation
  • 📝 Works in any app (browser, editor, chat, etc.)
  • 🔒 Local-only processing — your voice never leaves your machine
  • ⚡ Lightweight — runs on CPU, no GPU required
  • 🫧 Tiny floating bubble UI — always accessible, never in the way
  • 🔌 Pluggable ASR backends (faster-whisper, more coming)

Screenshots

(Coming soon)

Architecture

Mic Input → VAD → Audio Buffer → ASR (faster-whisper) → Text Injection → Active App
   ↓              ↑
[phrase detection] [partial/final display]

Installation

Prerequisites

  • Python 3.11 or 3.12 (NOT 3.14)
  • Windows 10/11 or Linux (X11)
  • Microphone

Windows

# 1. Clone
git clone https://github.com/t957095/meow.git
cd meow

# 2. Create virtual environment
python -m venv .venv
.venv\Scripts\Activate.ps1

# 3. Install
pip install -e ".[dev]"

# 4. Run
python -m meow

Linux

# 1. Clone
git clone https://github.com/t957095/meow.git
cd meow

# 2. Create virtual environment
python3 -m venv .venv
source .venv/bin/activate

# 3. Install system dependencies (Ubuntu/Debian)
sudo apt-get install portaudio19-dev xdotool

# 4. Install
pip install -e ".[dev]"

# 5. Run
python -m meow

Usage

  1. Launch Meow — a small bubble appears
  2. Click the 🎤 button to start dictation
  3. Speak naturally
  4. Gray text = processing, White text = finalized
  5. Finalized text is automatically typed into the focused app
  6. Click ⏹ to stop

Configuration

Create a .env file in the project root:

MEOW_ASR_BACKEND=faster_whisper
MEOW_MODEL_SIZE=base
MEOW_DEVICE=auto
MEOW_LANGUAGE=en
MEOW_VAD_AGGRESSIVENESS=2
MEOW_PHRASE_TIMEOUT=1.5

Model Sizes

Model Speed Accuracy VRAM Best For
tiny Fastest Basic ~1GB Quick testing
base Fast Good ~1GB Default
small Moderate Better ~2GB Accuracy priority

Privacy

  • No telemetry — zero data collection
  • No cloud — all processing is local
  • No audio storage — audio is processed in memory and discarded
  • No logs of transcript content — only timing and status
  • Open source — audit the code yourself

Development

# Run tests
pytest tests/ -v

# Lint
ruff check src/
ruff format src/

# Type check
mypy src/meow/

Roadmap

  • Phase 1: Speech-to-text dictation
  • Phase 2: TTS (text-to-speech) with VibeVoice-Realtime
  • Phase 3: Custom vocabulary / hotwords
  • Phase 4: Global hotkeys
  • Phase 5: Linux Wayland support

Model Backends

Current: faster-whisper

Uses faster-whisper with CTranslate2 for efficient CPU/GPU inference.

Future: VibeVoice-ASR

Microsoft's VibeVoice-ASR (9B params) is designed for batch transcription, not real-time dictation. May be added as an optional backend.

VibeVoice Clarification

Important: microsoft/VibeVoice-Realtime-0.5B is a text-to-speech model (speaks text aloud), NOT speech-to-text. It will be used for Phase 2 TTS features, not dictation.

License

MIT License — see LICENSE

Contributing

Contributions welcome! Please read our Contributing Guide (coming soon).

About

No description, website, or topics provided.

Resources

Contributing

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages