PolyVoice transcribes meeting audio locally and runs a small crew of local-LLM agents (via Ollama) to produce summaries, action items, and decisions — keeping sensitive audio on-device.
Built on the "Antigravity" model.
# 1. Clone
git clone https://github.com/Kimosabey/poly-voice.git
cd poly-voice
# 2. Install
# (see docs/GETTING_STARTED.md for the full setup)
# 3. Run
docker compose up- Fully local transcription + summarization
- Multi-agent orchestration over one meeting
- Action-item and decision extraction
- Privacy-preserving (no audio leaves the machine)
%%{init: {'theme':'base','themeVariables':{'primaryColor':'#ffffff','lineColor':'#2563eb','mainBkg':'#ffffff'}}}%%
graph LR
A([Meeting Audio])
B([Transcription])
C([Agent Summarizer])
D([Structured Notes])
A --> B
B --> C
C --> D
style A fill:#eff6ff,stroke:#2563eb,stroke-width:2px,color:#1e40af
style B fill:#eff6ff,stroke:#2563eb,stroke-width:2px,color:#1e40af
style C fill:#eff6ff,stroke:#2563eb,stroke-width:2px,color:#1e40af
style D fill:#eff6ff,stroke:#2563eb,stroke-width:2px,color:#1e40af
Orchestrating local models well enough to rival cloud summaries under a tight hardware budget.
See docs/ARCHITECTURE.md for the full HLD/LLD and design decisions.
| Layer | Technology | Role |
|---|---|---|
| Ollama | Ollama |
Local LLM runtime |
| LangGraph | LangGraph |
Stateful multi-agent orchestration |
| Whisper | Whisper |
Speech-to-text model |
- Architecture — high- and low-level design, decision log
- Getting Started — prerequisites, setup, environment
- Failure Scenarios — fault analysis and recovery
- Interview Q&A — deep-dive walkthrough
- Speaker-attributed summaries
- Topic segmentation
- Export to task trackers
Released under the MIT License.
Harshan Aiyappa Senior Full-Stack Hybrid AI Engineer Voice AI • Distributed Systems • Infrastructure