A lightweight, on-demand (and optional auto-narration) text-to-speech speaker plugin. High-quality neural voice via edge-tts + playsound (no media-player popup), with native/pyttsx3 fallbacks. Zero extra LLM tokens — it only speaks text that was already generated.
Rebrand note: This project was previously known as
claude-code-voiceand the GitHub repo is nowrhishi99/OutLoud(the old URL auto-redirects).
Config/state directories remainclaude-code-voiceinternally (%APPDATA%\claude-code-voice,~/.config/claude-code-voice) so existing setups keep working — only the repo name changed.
The problem — long AI replies are walls of text that tire eyes and kill flow.
OutLoud's fix — instant voice narration, with zero extra tokens via hotkey or auto-read.
flowchart LR
A[LLM reply] --> B[Stop hook saves text]
B --> C[Hotkey / /speak / autoSpeak]
C --> D["Neural voice (edge-tts default)"]
Placeholder ready — real 20-40s recording of
/speak last+/speak on(with terminal + actual audio) coming soon.
Drop-in instructions for the real demo asset:
# Recommended: use OBS Studio / Windows Game Bar / ffmpeg
# Example (macOS):
# ffmpeg -f avfoundation -i "1:0" -t 35 assets/demo.mp4
# Then embed below and commit:
# <video src="assets/demo.mp4" controls width="100%"></video>[ assets/demo.mp4 — terminal capture + voice of /speak last and /speak on ]
See it in action in the interactive visualizer too.
This project was designed and built end-to-end with Grok Build — full credit to Grok for the architecture, the iterative build loop, the multi-agent docs sweep, and the live tooling. The entire build journey (every milestone, every fix) is captured in an interactive timeline:
- 📈 build-journey.html — the Grok Build story, milestone by milestone (live, rendered)
- 🎛️ voice-plugin-visualizer.html — live architecture + interactive config console
Open either file in a browser to see how it came together.
Credits & roles: OutLoud was fully designed and built by Grok Build — all architecture, code, engines, hooks, and the build journey. Claude Code only performed cosmetic doc edits, repo validation, and the final publish/push (repo creation, link rename to
OutLoud, WSL quickstart wording). No core functionality was authored by Claude.
- 🎙️ Natural neural voice —
edge-tts(Microsoften-US-AriaNeuralby default) for genuinely human-sounding output. - 🔇 No popup —
playsoundplays the MP3 directly. No media player hijacking your screen. - 💸 Zero extra tokens with the hotkey or auto-read — the speaker reads text Claude/Grok already produced and never calls an LLM itself. The
/speakslash command is the exception: Claude Code sends it to the model like any prompt, so it costs one normal Claude turn (bigger in long sessions). Use the hotkey or! python scripts/speaker.py --lastto stay token-free. - ⌨️ Multiple triggers — hotkey,
/speakslash command, CLI, or status-line badge. - 🪝 Automatic capture — a Stop hook quietly saves the last response so it's ready to speak on demand.
- 🔁 Optional autoSpeak — opt-in automatic narration after Stop (with limits, code skipping, and modes).
- 🧩 Real plugin — proper hooks, commands, and skills.
- 🤖 Multi-agent — one shared backend powers both Claude Code (hook-based) and Grok Build (direct invoke).
- 🔁 Pluggable engines —
edge-tts,native(OS built-in, offline, zero deps),pyttsx3.kokoro(offline neural) is paused/experimental. - ♿ Accessibility-first — listen while you work; great for eyes-off and screen-reader-adjacent workflows.
- 🔇 Global mute —
OUTLOUD_MUTE=1kill switch (respected everywhere).
/plugin marketplace add rhishi99/OutLoud
/plugin install speaker@outloudAfter install, use /speak (it may appear as /speaker:speak in the menu).
Recommended first step:
/speaker:setup
This dedicated command forces validation + prints the exact minimal setup instructions for your platform (or "✅ setup completed" if everything is ready).
WSL/Linux/mac users: run /speaker:setup — it will give you the precise sudo apt + pip lines to copy-paste.
git clone https://github.com/rhishi99/OutLoud.git
cd OutLoud
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
python scripts/speaker.py --set engine edge-tts --set voice "en-US-AriaNeural"See the dedicated Status Line section below.
After /plugin install speaker@outloud, just type:
/speak last
or
/speaker:config
Run /speaker:setup — it will detect issues and print the exact minimal commands (or confirm ✅ setup completed).
WSL/Linux/mac (one time in your terminal):
sudo apt update && sudo apt install -y python3-pip python3-venv mpg123
pip install --user edge-tts playsound==1.2.2Then use /speak again. The plugin now drives most of this for you on first use.
Audio in WSL: WSLg (Windows 11+) gives audio for free. Without it, run the speaker from native Windows instead.
| Command | Effect |
|---|---|
/speak last |
speak the last response on demand |
/speak "text" |
speak arbitrary text immediately |
/speak on |
enable autoSpeak (auto-narrate after Stop hook) |
/speak off |
disable auto-narration |
/speak stop |
best-effort stop current playback |
Tip:
/speakwith no argument is ambiguous (the agent will ask what you want). Use/speak lastor/speak onto act directly.
The default engine is edge-tts (natural neural voice). OUTLOUD_MUTE=1 is a global kill switch honored everywhere. If a response wasn't captured yet, capture happens on the next Stop — ask something first, then /speak last.
Preferred flow:
/plugin marketplace add rhishi99/OutLoud/plugin install speaker@outloud- Inside Claude Code type
/speak last
The plugin will print the exact 2 lines you need to run in your WSL terminal (apt + pip).
Only if you want to clone manually for development, see the older detailed steps in previous versions of this README or the repo history.
Audio note: WSLg (Windows 11) is the easiest. Otherwise run the speaker commands from a Windows PowerShell instead of inside WSL.
OutLoud can show a compact badge in Claude Code's footer / status bar:
OutLoud(on-demand ready)auto(autoSpeak enabled)OutLoud [muted](whenOUTLOUD_MUTE=1)/speak(fallback label)
Honest note: The status line is text-only and NOT clickable. It is purely informational.
There is no button or tap target — use your global hotkey or type /speak (or /speak last) in the prompt.
- Add (or merge) the following into
~/.claude/settings.json(create the file if it does not exist):
{
"statusLine": {
"command": "node \"/absolute/path/to/claude-code-voice/scripts/status.js\""
}
}Use an absolute path to scripts/status.js (example for Windows):
"command": "node \"C:\\Users\\YourName\\path\\to\\claude-code-voice\\scripts\\status.js\""- Fully restart Claude Code.
The script reads your live config.json + the OUTLOUD_MUTE env var on every render.
- Ask Claude (or Grok) something.
- Hear the last response any of these ways:
- Press your hotkey (e.g.
Ctrl+Alt+S) - Type
/speak(or/speak last) - Run it directly:
python scripts/speaker.py --last
- Press your hotkey (e.g.
- Speak arbitrary text:
python scripts/speaker.py "This is a very natural voice" /speak Some custom text here - Toggle auto-narration:
/speak on # enable autoSpeak /speak off # disable autoSpeak /speak stop # best-effort stop current playback
Large code blocks are spoken as
[code block](or omitted entirely under autoSpeak) so you're not read a wall of syntax.
autoSpeak lets OutLoud automatically read responses aloud after the Stop hook (opt-in only). It is disabled by default.
It always runs non-blocking (detached process) and never holds the hook.
| Key | Type | Default | Description |
|---|---|---|---|
autoSpeak |
boolean | false |
Master switch. When true, the Stop hook will speak a processed slice of the response. |
autoSpeakMaxChars |
number | 1200 |
Hard cap on characters spoken automatically (shorter & friendlier than the on-demand max_chars). |
autoSpeakSkipCodeBlocks |
boolean | true |
When true, fenced + inline code is stripped entirely (clean narration). When false, code becomes [code block]. |
autoSpeakMode |
string | "full" |
How to truncate when over the char limit: • "full" — take the first N chars (word-aware)• "summary" — first sentence + last sentence• "first-paragraph" — up to first blank line (or cutoff) |
Enable a nice summary mode (recommended for long answers):
{
"autoSpeak": true,
"autoSpeakMaxChars": 900,
"autoSpeakSkipCodeBlocks": true,
"autoSpeakMode": "summary"
}Minimal "first paragraph only":
python scripts/speaker.py --set autoSpeak true
python scripts/speaker.py --set autoSpeakMaxChars 600
python scripts/speaker.py --set autoSpeakMode first-paragraphVia slash command (Claude Code):
/speak on
/speak off
Via the CLI (works everywhere):
python scripts/speaker.py --autospeak on
python scripts/speaker.py --autospeak off
python scripts/speaker.py --configSafety: autoSpeak is completely ignored when OUTLOUD_MUTE=1 is set.
On-demand hotkey + /speak (and /speak last) always use the full saved last response (subject only to the regular max_chars + strip_code settings).
All settings live in a single JSON file:
- Windows:
%APPDATA%\claude-code-voice\config.json - macOS / Linux:
~/.config/claude-code-voice/config.json
See config.example.json for a complete starting point.
| Key | Type | Default | Notes |
|---|---|---|---|
engine |
string | "edge-tts" |
edge-tts (recommended), native, pyttsx3, kokoro (experimental) |
voice |
string | "en-US-AriaNeural" |
Engine-specific voice name/ID |
rate |
number | 1.0 |
0.5–2.0 speed multiplier |
volume |
number | 1.0 |
0.0–1.0 |
language |
string | "en" |
Used by some engines (e.g. kokoro) |
strip_code |
boolean | true |
On-demand: replace code blocks with [code block] |
max_chars |
number | 6500 |
Hard truncation for on-demand speech |
autoSpeak |
boolean | false |
Enable automatic narration after Stop |
autoSpeakMaxChars |
number | 1200 |
Char limit for auto narration |
autoSpeakSkipCodeBlocks |
boolean | true |
Strip code when auto-speaking |
autoSpeakMode |
string | "full" |
"full" | "summary" | "first-paragraph" |
Global kill switch (environment variable, works everywhere):
OUTLOUD_MUTE=1When set, all speech (manual + autoSpeak) is suppressed. Perfect for CI, pair programming, or recording.
Change settings live:
python scripts/speaker.py --set engine edge-tts
python scripts/speaker.py --set voice en-GB-SoniaNeural
python scripts/speaker.py --set rate 1.05
python scripts/speaker.py --config
python scripts/speaker.py --list-voices| Engine | Quality | Install | Offline | Best for |
|---|---|---|---|---|
| edge-tts (default) | Excellent | pip install edge-tts playsound==1.2.2 |
No | Natural listening |
| native | Basic (fast) | None (built-in) | Yes | Zero dependencies |
| pyttsx3 | Good | pip install pyttsx3 |
Yes | Better native control |
| kokoro (paused/experimental) | Very good (local neural) | contributions welcome | Yes | Fully offline neural |
Change engine/voice anytime (see Configuration section above).
| Platform | edge-tts | native | pyttsx3 | kokoro (exp.) |
|---|---|---|---|---|
| Windows | ✅ full (needs net) | ✅ full, offline | ✅ full, offline | |
| macOS | ✅ full (needs net) | ✅ full, offline (say) |
✅ full, offline | |
| Linux | ✅ full (needs net) | ✅ full, offline | ||
| WSL | ✅ (WSLg/PulseAudio + net) |
✅ = recommended / works with audio device & deps
⚠️ = works with extra setup or limited quality
All playback is local after synthesis (edge-tts only step needing internet).
Minimal:
edge-tts
playsound==1.2.2
- All playback is local after install.
edge-ttsneeds internet to synthesize. - The
nativeengine works fully offline with zero extra packages. - Config + last-response live under the
claude-code-voicedirectory (see above).
Bind a global hotkey so you can hear the last response from anywhere without touching the mouse or terminal.
The exact command to run against the saved last response:
python /full/path/to/claude-code-voice/scripts/speaker.py --lastEquivalent convenient wrappers (recommended on Windows):
- Windows:
powershell -ExecutionPolicy Bypass -File "C:\full\path\claude-code-voice\speak.ps1" -Last - Linux/macOS helper:
./speak.sh --last(orbash speak.sh --last)
Always prefer absolute paths in hotkey bindings.
One command, nothing to install (recommended):
powershell -ExecutionPolicy Bypass -File scripts\install-hotkey.ps1Adds Ctrl + Alt + S (speak last response) and Ctrl + Alt + X (stop) as native Windows shortcut keys. Pick other keys with -SpeakKey 'CTRL+ALT+R' -StopKey 'CTRL+ALT+Q'; undo with -Remove.
PowerToys Keyboard Manager (zero code):
- Open PowerToys → Keyboard Manager → Remap a shortcut.
- New shortcut:
Ctrl + Alt + S(or your preference). - Action: "Launch program".
- Program: full path to
python.exe - Arguments:
"C:\path\to\claude-code-voice\scripts\speaker.py" --last - (Optional) Start in: the project folder.
AutoHotkey v2 example (outloud.ahk):
#Requires AutoHotkey v2.0
; Ctrl+Alt+S
^!s::{
Run 'python "E:\path\to\claude-code-voice\scripts\speaker.py" --last', , "Hide"
}Run the script (or compile to .exe and put in Startup folder).
Alternative: use the root dispatcher for native routing:
^!s::{
Run 'powershell -ExecutionPolicy Bypass -File "E:\path\to\claude-code-voice\speak.ps1" -Last', , "Hide"
}Shortcuts app (built-in):
- Open Shortcuts → File → New Shortcut.
- Add action "Run Shell Script".
- Paste:
python3 /Users/you/path/to/claude-code-voice/scripts/speaker.py --last - Give it a name (e.g. "OutLoud Speak Last").
- System Settings → Keyboard → Keyboard Shortcuts → App Shortcuts (or Services) → assign
Cmd + Shift + S.
Karabiner-Elements (advanced):
Create a complex modification that runs the shell command above.
sxhkd (popular with bspwm, i3, etc.) in ~/.config/sxhkd/sxhkdrc:
ctrl + alt + s
python3 /home/you/claude-code-voice/scripts/speaker.py --last
Then pkill -USR1 sxhkd (or restart).
xbindkeys in ~/.xbindkeysrc:
"python3 /home/you/claude-code-voice/scripts/speaker.py --last"
control + alt + s
Run xbindkeys -p to reload.
Use your window manager's native keybinding tool if preferred.
- Claude Code: a Stop hook (
hooks/save-last.js) writes cleaned text tolast-response.txtafter every final response./speakor your hotkey plays it. - Grok Build: no Stop hook, so Grok invokes the speaker scripts directly (see
skills/grok-voice/SKILL.md). - autoSpeak (opt-in): when enabled, the same hook also fires a detached speaker call on a processed slice of the text.
- Both agents share the same
config.json, engines, and capture file — one backend, two (soon more) agents. - Speaking is always explicit (hotkey, CLI,
/speak) or opt-in viaautoSpeak. Nothing speaks by default.
See the interactive visualizer for the full flow diagram.
No audio in WSL
WSL2 has no audio device by default. Use WSLg (built-in PulseAudio on Win11+). Test: python scripts/speaker.py "hello".
Fallback: run from native Windows PowerShell instead, or install mpg123 + use edge-tts fallback players.
edge-tts needs internet / want offline
edge-tts requires net for synthesis only. Switch instantly:
python scripts/speaker.py --engine native "offline test"
(native is zero-dep, built-in on Win/mac, needs espeak-ng on Linux).
playsound / package issues
Run /speaker:setup — it will print the exact pip (and apt for WSL) commands needed.
pip install playsound==1.2.2 (exact pin recommended). Speaker falls back automatically if needed.
Still stuck? Ask Claude Code to "just set up OutLoud for me" — /speaker:setup will delegate the install to a cheap sub-agent (haiku) that runs the platform-specific commands and re-validates automatically.
Stop hook not firing / no last response
- Make sure the plugin is installed in Claude Code:
/plugin install speaker@outloud(after marketplace add). - Check hooks: the
hooks/hooks.jsonregisters on Stop. Restart Claude Code after install. - Use
/speak lastor hotkey; first message must complete fully.
Status line not showing
Add absolute path to ~/.claude/settings.json (see Status Line section).
Restart Claude Code completely (not just reload). The badge reads live config + OUTLOUD_MUTE.
OUTLOUD_MUTE stuck / everything silent
OUTLOUD_MUTE=1 is a hard global kill switch (respected by CLI, hooks, and autoSpeak).
Unset it (close terminal or set OUTLOUD_MUTE=) or explicitly use a shell without the var. Check with python scripts/speaker.py --config.
General debug
python scripts/speaker.py --config
python scripts/speaker.py --engine native "quick test"
OUTLOUD_MUTE=1 python scripts/speaker.py --last # safePRs and issues welcome — especially:
- Reviving kokoro for fully-offline neural voice.
- A native speaker button inside the Grok Build / Claude Code UI.
- More voices, engines, and platform testing (macOS / Linux / WSL).
- Polish on the upcoming folder rename.
Fork, branch, and open a PR. Keep it lightweight.
MIT © OutLoud contributors. Built with Grok Build.
claude-code · outloud · tts · edge-tts · voice · accessibility · grok · developer-tools · cli · neural-voice