Ultra-low-latency, real-time live captions and floating desktop subtitles for Windows.
Empowering deaf and hard-of-hearing individuals with instant speech-to-text across meetings, video calls, streaming, and daily conversations.
For deaf and hard-of-hearing people, standard computer audio is a daily accessibility hurdle:
- Video meetings on Zoom, Teams, or Google Meet often have delayed, unreliable, or unavailable auto-captions.
- YouTube videos, online courses, tutorials, and social media frequently lack subtitles in specific languages or local dialects.
- Desktop voice/video calls on WhatsApp, Discord, or Telegram provide no native real-time captioning.
OmniCaption AI solves this directly at the operating system level. It captures any sound playing on your Windows PCโplus your microphoneโand displays clean, instantaneous, rolling subtitles on top of your screen in under 80 milliseconds.
- Floats over any application: Stays visible over Zoom, Microsoft Teams, Google Meet, YouTube, VLC player, Netflix, or web browsers.
- Zero clutter design: Compact 130px banner specifically styled to show only active speech without vertical cutoffs or historical screen pollution.
- Fully customizable: Drag anywhere on the screen, toggle transparency (65% to 98%), or expand to full Studio mode with one click.
- Direct WebSocket streaming powered by Deepgramโs state-of-the-art Nova-3 speech model.
- TCP Nagle algorithm bypass (
TCP_NODELAY = 1): Audio chunks are pushed directly to the network interface without Windows packet-bundling delay. - Instant queue draining: Prevents speech packet backlog during momentary network jitter.
- ๐น๐ณ Tunisian Darija (
ar-tn/ ุงูุฏุงุฑุฌุฉ ุงูุชููุณูุฉ): Specialized vocabulary prompting tuned for Tunisian idioms, daily expressions, and television culture (e.g. Choufli Hal characters and colloquial phrases). - ๐ซ๐ท French (
fr): Formatted for technical, business, and everyday French speech. - ๐ฌ๐ง English (
en): Cloud, computing, data engineering, and conversational vocabulary. - Smart Numerals & Formatting: Converts spoken numbers into digits and symbols automatically (e.g.,
300,90%,$10,1er).
- Captures raw digital audio straight from your default speaker output device using Windows WASAPI loopback.
- Zero virtual cables needed: No complex third-party virtual audio cables (VB-Cable/Voicemeeter) required.
- Three audio modes:
- System Audio: Transcribes YouTube, meetings, movies, and PC sound.
- Microphone: Transcribes your room / in-person speech.
- Both (Calls Mode): Transcribes both sides of a phone/video conversation simultaneously.
- Detects and separates different speakers in real time.
- Assigns distinct, high-contrast colors (
[Person 1],[Person 2]) to make group conversations and conference calls intuitive to follow.
- Need to take notes or review a lecture? Toggle Keep History to record the conversation with millisecond timestamps.
- Export full transcripts to
.srt(SubRip Subtitles) or.txtwith one click. - Instant copy to clipboard.
flowchart LR
A["Windows Audio Source\n(WASAPI Loopback / Mic)"] --> B["Acoustic Pre-Processing\n- 80Hz Low-Cut Rumble Filter\n- Vectorized 16kHz Resampling"]
B --> C["WebSocket Client\n- TCP_NODELAY Enabled\n- Zero-Drop Bounded FIFO"]
C <--> D["Deepgram Nova-3 API\n- endpointing=300ms\n- utterance_end_ms=1000ms\n- Keyterm Boosting"]
D --> E["Anti-Flicker Layer\n- Cyan: Real-Time Interim Draft\n- White: Locked Committed Text"]
E --> F["Always-On-Top Floating Bar\n(CustomTkinter GUI)"]
๐ก For Deaf & Non-Technical Users (No Python Needed):
You can download the ready-to-use standalone package directly:
๐ Download OmniCaption AI v1.0.0 (Windows 64-bit)
Unzip the folder, double-clickOmniCaption_AI.exe, and start transcribing immediately!
- Windows 10 or Windows 11 (64-bit)
- Python 3.10, 3.11, or 3.12
- A Deepgram API Key (Free tier provides $200 in free credits)
git clone https://github.com/msouid/omnicaption-ai.git
cd omnicaption-aipython -m venv venv
venv\Scripts\activate
pip install -r requirements.txtCreate a .env file from the provided example:
copy .env.example .envOpen .env and add your Deepgram API Key:
DEEPGRAM_API_KEY=your_deepgram_api_key_herepython app.pyPrefer running live captions directly in your Command Prompt / terminal without a GUI? Use the built-in interactive CLI runner:
# French with PC system audio
python cli_test.py --lang fr --source system
# Tunisian Darija with PC audio
python cli_test.py --lang ar-tn --source system
# English with microphone & speaker diarization
python cli_test.py --lang en --source mic --diarizeOr double-click test_in_cmd.bat on Windows!
You can easily customize the words, acronyms, colleague names, or local slang that the AI prioritizes by editing predefined_languages_keyterms.json:
{
"ar-tn": {
"keyterms": [
"ุดูููู ุญู", "ุณุจูุนู", "ุณููู
ุงู", "ุจุฑุดุง", "ูุนูุดู", "ูุง ุฎููุง"
]
},
"fr": {
"keyterms": [
"Delta Lake", "Parquet", "Kubernetes", "Docker", "API REST"
]
},
"en": {
"keyterms": [
"PySpark", "Machine Learning", "Microservices", "CI/CD"
]
}
}The app automatically reads this file on startup and injects up to 90 unweighted keyterm prompts into Deepgram Nova-3, boosting keyword recall by up to 90%!
To generate a portable .exe that deaf family members or friends can run without needing Python installed:
pip install pyinstaller
pyinstaller --noconfirm OmniCaption_AI.specThe compiled, zero-dependency standalone app will be created inside the dist/OmniCaption_AI/ folder.
We warmly welcome contributions from the deaf, hard-of-hearing, and accessibility engineering communities!
- Suggest UI/UX accessibility enhancements (high contrast themes, dyslexic fonts, haptic feedback).
- Expand dialect keyterm dictionaries for other Arabic varieties, French regional expressions, or specialized industry terms.
- Report bugs or submit feature requests via GitHub Issues.
This project is open-source and released under the MIT License.