Tired of hitting Claude Code token limits? claude-polyrouter silently routes every query to the right model — stop paying Opus prices for simple questions. 82% less token waste, 10 languages at full parity, zero setup.
- Live HUD —
ctx:%, real-time5h/wk/sntrate-limit bars with threshold-based colors, animated Poly mascot, idle fallback- Dynamic effort per query —
medium→high→xhighchosen from a multi-signal score (architecture, file count, tool intensity)- Opus 4.8 sub-effort routing — deep-tier queries get the right Opus effort level automatically; architectural calls escalate to
xhigh + adv- 10 languages, full parity — EN · ES · FR · DE · PT · JA · KO · ZH · AR · RU all share the same deep/arch pattern coverage
| Query | Tier | Model | Effort | Reason |
|---|---|---|---|---|
| "hola" / "ok" / "what is X?" | Fast | Haiku 4.5 | low | Short or simple question |
| "create a function" / "fix this bug" | Standard | Sonnet 5 | medium | Coding task |
| "debug this stack trace across 3 files" | Deep | Opus 4.8 | high | Multi-file, tool-heavy |
| "redesign the auth architecture" | Deep | Opus 4.8 | xhigh + adv | Architectural scope → Advisor engaged |
Routing happens automatically on every query via a UserPromptSubmit hook. No manual intervention needed.
- Context-aware prompt enhancer: prompts de implementación largos se elevan automáticamente a Sonnet/Opus — ya no van a Haiku
- Slash commands inteligentes: /spec siempre → Opus, /work analiza el spec activo para decidir el modelo
- Soul Map: aprende tus preferencias explícitas de modelo (cuando mencionas "usa opus", "con sonnet", etc.)
- exec_model real: el HUD ahora muestra el modelo que CC realmente usó para ejecutar, no solo el que poly predijo
- Prompt quality scorer: q:N% en el HUD cuando el prompt puede mejorar (oculto cuando es >= 80%)
- ⏱ Tiempo de sesión: muestra cuánto lleva la sesión activa (verde <30m, amarillo 30-60m, rojo >60m → señal para /compact)
- 🤖99+: soporte para Dynamic Workflows con cientos de subagentes
- Session state por proyecto: cada proyecto tiene su propio estado de routing — las sesiones ya no se interfieren entre sí
- Fallback model support: si Opus no está disponible, poly degrada a Sonnet automáticamente
- Effort skew detection: detecta cuando CC usa un effort diferente al que poly esperaba
- claude-opus-4-8: modelo actualizado al último Opus
- CC max → xhigh: el nivel max de CC mapea al techo de poly
- 983 tests passing
- Verifiability routing (Karpathy) — Tasks with verifiable outcomes route to Opus; heuristic-only tasks stay on Sonnet.
✓/~indicator in the HUD model segment - Native OAuth usage polling —
5h/wk/snt/extrabars without ccusage dependency, stale-while-revalidate cache - 🧠 thinking indicator — appears on the exec model when the active subagent runs opus / deep / xhigh
- 📁 CWD in HUD — current directory basename visible on every prompt
- HUD fully independent — runs standalone with no OMC or external plugin dependency
- HUD always complete — no width-based truncation; all segments render
- OAuth field mapping bugfixes — v1.9.2 (nested response), v1.9.3 (
used_credits/100+is_enabledgate) - 909 tests passing (14 new verifiability unit tests)
- exec segment in HUD —
⚙ exec:opus·xhigh·advnow appears as a separate segment afterprompt:, making the executor model visible at a glance - 🤖N subagent counter — Active subagent count shown in HUD (
🤖1,🤖2, …) whenever a subagent is running - Swap detection expanded —
⚠swapglyph now covers both fast and standard tiers viaPreToolUse:Taskhook (not just deep) - SessionStart reset —
subagent_countandactivestate reset on every new session to prevent stale carry-over between sessions - 895 tests passing · 98.9% routing accuracy
- HUD v1.6 — New format
[poly v1.6.2], prompt/exec split,ctx:%, rate-limit bars (5h/wk/snt),⚠compactat ctx≥70%, new mascot states[>.^](ctx high) and[x.x](critical) - Idle fallback — Stale sessions emit
[poly v1.6.2] [^.^]~ idlefor non-OMC users instead of a blank statusline - Accurate savings calc — Per-token formula (1k input + 500 output) with corrected Opus 4.7 pricing ($15/$75 per 1M tokens)
- DE/FR/PT deep patterns — Multi-file refactor queries in German, French, and Portuguese now promote to
deep + xhigh - 613 tests passing
- Pinned model IDs —
claude-haiku-4-5,claude-sonnet-4-6,claude-opus-4-7— display stays compact, versions stay explicit - Dynamic deep-tier effort —
medium→high→xhighchosen per-query from multi-signal score (architecture, file count, code blocks, tool intensity) - xhigh tier — polyrouter-only display label for architectural/critical work; normalizes to
highfor theCLAUDE_CODE_EFFORT_LEVELenv var - Advisor Strategy —
requires_advisor=trueflag set automatically onxhigh; surfaces asadvin the HUD so the executor can engage the Advisor (Opus on-demand) for architectural calls - Subagent lifecycle tracking — HUD shows
(subagente)while the spawned executor is running;SubagentStophook clears it - Architectural promotion — non-deep queries with architecture keywords + standard signals auto-promote to
deep + xhigh(e.g. "plan a major refactor across services")
| Query | Tier | Effort | Advisor | Why |
|---|---|---|---|---|
| "fix typo in README" | fast | low | — | Short, no signals |
| "add pagination to /api/users" | standard | medium | — | Single-file standard task |
| "debug this 200-line stack trace" | deep | high | — | Deep + tool-intensive |
| "redesign the billing subsystem" | deep | xhigh | adv | Architectural keyword match |
| "refactor auth.py login.py session.py as a strategic migration" | deep | xhigh | adv | Arch keyword + 3 files + orchestration |
- 82% token reduction — additionalContext reduced from ~150 tokens (v1.3) to ~27 tokens
- 100% classification accuracy — 30/30 on multilingual test suite across all 10 languages
- Multi-signal scoring — 9-signal weighted engine replaces simple pattern counting
- Poly mascot HUD — Animated ASCII mascot with cache freshness bar, zero token cost
- Auto-hook injection — Zero-config installation, hooks configured automatically
- Multi-signal scoring — 9-signal weighted scoring engine (patterns, code blocks, error traces, file paths, prompt length, tool results, conversation depth, effort level, universal tech symbols)
- 10 languages — English, Spanish, Portuguese, French, German, Russian, Chinese, Japanese, Korean, Arabic, plus Spanglish detection
- Zero API keys — Pure rule-based classification with pre-compiled regex patterns (~3ms latency)
- Cost savings — 50-80% reduction by routing simple queries to cheaper models
- Two-level cache — In-memory LRU + file-based persistent cache for repeated queries
- Multi-turn awareness — Detects follow-up queries and maintains conversation context
- Dynamic effort — Automatic effort level mapping (low/medium/high) based on routing tier
- Cache keep-alive — PostToolUse hook detects prompt cache expiration risk
- Compact advisory — Recommends context compaction when stale tool results accumulate
- Poly mascot HUD — Animated ASCII mascot in statusLine with 6 states, cache freshness bar, zero token cost
- Project learning — Optional knowledge base that fine-tunes routing per project
- Analytics — Terminal stats and HTML dashboard with Charts.js visualizations
- Claude Code v2.1.70+
- Python 3.8+
- Node.js 18+
# Step 1: Add marketplace (one-time)
claude plugin marketplace add SonyHarv/claude-polyrouter
# Step 2: Install plugin
claude plugin install claude-polyrouter@claude-polyrouter
# Step 3: Restart Claude CodeThat's it. On first run, the plugin auto-configures:
UserPromptSubmithook insettings.jsonfor automatic routing- HUD symlink for the Poly mascot statusLine
currentsymlink for version-agnostic paths
| Integration | Purpose | Install |
|---|---|---|
ccusage |
Shows 5h/weekly/sonnet rate-limit bars in the HUD | npm install -g ccusage |
Claude Code supports one statusLine command at a time. If you use poly
alongside other plugins that also provide a statusLine, you can combine them
with a simple wrapper script.
Example wrapper (~/.claude/statusline-wrapper.sh):
#!/bin/bash
# Combine multiple statusLine outputs into one line.
# Each plugin emits its own section — the wrapper concatenates them.
PLUGIN_A=$(node ~/.claude/plugins/plugin-a/hud.mjs 2>/dev/null)
POLY=$(node ~/.claude/plugins/cache/claude-polyrouter/claude-polyrouter/current/hud/polyrouter-hud.mjs 2>/dev/null)
parts=()
[[ -n "$PLUGIN_A" ]] && parts+=("$PLUGIN_A")
[[ -n "$POLY" ]] && parts+=("$POLY")
printf '%s' "$(IFS=' '; echo "${parts[*]}")"Wire it up in ~/.claude/settings.json:
{
"statusLine": {
"type": "command",
"command": "bash ~/.claude/statusline-wrapper.sh"
}
}Note: poly is designed to be the primary statusLine. If you use a wrapper, test that the combined output fits your terminal width or enable multi-line mode in your terminal.
Poly lives in your statusLine and shows routing state at zero token cost.
No subagent:
[poly v1.9.5] [^.^]~ haiku·fast✓ │ cache:████░ ctx:8% │ 5h:45%(1h2m) wk:9%(6d19h) snt:3%(6d19h) extra:1%($2.09/$150.00) │ 📁mi-proyecto $0.03↓ es
With subagent (thinking model):
[poly v1.9.5] [^.^]~ prompt:haiku·fast ⚙ exec:opus·xhigh·adv🧠 │ 🤖1 cache:████░ ctx:15% │ 5h:45%(1h2m) wk:9%(6d19h) snt:3%(6d19h) extra:1%($2.09/$150.00) │ 📁mi-proyecto $9.50↓ es
High context (compact advisory):
[poly v1.9.5] [^.^]~ haiku·fast✓ ⚠compact │ cache:████░ ctx:78% │ 5h:45%(1h2m) wk:9%(6d19h) │ 📁mi-proyecto $0.03↓ es
Stale session (>30 min, no OMC):
[poly v1.9.5] [^.^]~ idle
| Element | When shown | Meaning |
|---|---|---|
[poly v1.9.5] |
Always | Plugin prefix + version |
[^.^]~ / [^-^] / [>.^] / [x.x] |
Always | Mascot state (see below) |
haiku·fast / sonnet·std / opus·deep |
After a route | Model + tier, dot-separated |
✓ / ~ |
After a route | Verifiability — ✓ task is verifiable, ~ heuristic only |
·high / ·xhigh |
Deep tier only | Sub-effort — medium is elided |
·adv |
requires_advisor=true |
Advisor (Opus on-demand) engaged |
prompt:… / ⚙ exec:…[🧠] |
Subagent active | Prompt model vs executor model split |
🧠 |
Subagent uses opus/deep/xhigh | "Thinking" indicator on exec model |
🤖N |
Subagent active | Count of active subagents |
cache:… |
Session active | Freshness bar (5-block Unicode) |
ctx:N% |
transcript_path available |
Context window used (from Claude Code) |
5h:N%(T) |
OAuth usage available | 5-hour limit usage + time to reset |
wk:N%(T) |
OAuth usage available | Weekly limit usage + time to reset |
snt:N%(T) |
OAuth usage available | Sonnet weekly limit usage + time to reset |
extra:N%($U/$L) |
Max plan, extra credits enabled | Extra credit usage (dollars used / monthly limit) |
📁<name> |
Always | Current working directory (basename) |
⚠compact |
ctx ≥ 70% | Context approaching limit — run /compact |
$x.xx↓ |
Savings > $0 | Cumulative estimated cost saved vs. always-Opus |
es / en / pt / … |
Language detected | ISO language code |
| State | Display | Trigger |
|---|---|---|
| Idle | [^.^]~ / [^-^] |
Default — session active, no pressure |
| Routing | [^o^]» |
Within 3 s of a new query |
| Thinking | [^.^]... |
3–10 s after query |
| Keepalive | [~_~]zzz |
Cache elapsed > 40 min |
| Danger | [°O°]!!! |
Cache elapsed > 50 min |
| Compact | [^.^]~~~ |
compact.advisory_active flag set |
| Ctx High | [>.^] |
ctx ≥ 70% |
| Critical | [x.x] |
ctx ≥ 90% or any rate-limit ≥ 90% |
| Time | Display | Meaning |
|---|---|---|
| 0–10 min | cache:█████ |
Fresh — cache is warm |
| 10–30 min | cache:████░ |
Warm — still healthy |
| 30–50 min | cache:███░░ ! |
Warning — consider a keep-alive query |
| 50+ min | cache:░░░░░ exp |
Expired — triggers Danger state |
The polyrouter HUD is self-contained. OMC (oh-my-claudecode) is optional — if present, its output is prepended to the poly line; if absent, polyrouter renders on its own.
| Command | Description |
|---|---|
/polyrouter:route <tier> [query] |
Manual routing override |
/polyrouter:stats |
View routing statistics |
/polyrouter:dashboard |
Open HTML analytics dashboard |
/polyrouter:config |
Show active configuration |
/polyrouter:learn |
Extract routing insights from conversation |
/polyrouter:learn-on |
Enable continuous learning mode |
/polyrouter:learn-off |
Disable continuous learning mode |
/polyrouter:knowledge |
View knowledge base status |
/polyrouter:learn-reset |
Clear knowledge base |
/polyrouter:retry |
Retry with escalated tier |
Global config at ~/.claude/polyrouter/config.json:
{
"default_level": "fast",
"confidence_threshold": 0.7,
"levels": {
"fast": { "model": "haiku", "model_id": "claude-haiku-4-5", "agent": "fast-executor" },
"standard": { "model": "sonnet", "model_id": "claude-sonnet-5", "agent": "standard-executor" },
"deep": { "model": "opus", "model_id": "claude-opus-4-8", "agent": "deep-executor" }
},
"scoring": {
"thresholds": { "fast_max": 0.35, "standard_max": 0.65 }
},
"effort": { "auto": true },
"keepalive": { "enabled": true, "threshold_minutes": 50 },
"compact": { "enabled": true, "keep_last_n": 5 },
"hud": { "mascot_enabled": true, "statusline_native": true }
}Project override at <project>/.claude-polyrouter/config.json:
{
"default_level": "standard",
"confidence_threshold": 0.8
}When new models release, update config only — no code changes needed:
{
"levels": {
"fast": { "model": "haiku-next", "agent": "fast-executor" }
}
}- Exception check — Slash commands, meta-queries, and continuations bypass routing
- Intent override — Natural language model forcing ("use opus") takes max priority
- Cache lookup — Two-level cache (memory + file) for repeated queries
- Language detection — Stopword-based scoring identifies the query language
- Pattern extraction — Raw signal counting from language-specific regex patterns
- Multi-signal scoring — Weighted 9-signal composite score maps to tier (fast <0.35, standard <0.65, deep >=0.65)
- Context boost — Multi-turn awareness adjusts confidence for follow-up queries
- Architectural promotion — non-deep tiers with arch keywords + signals auto-promote to deep + xhigh
- Dynamic effort — deep tier receives a sub-effort (medium/high/xhigh) from score + signal mix
- Advisor flag —
xhighsetsrequires_advisor=true, surfaced asadvin the HUD - Learned adjustments — Optional project knowledge base fine-tunes routing
| Language | Code | Notes |
|---|---|---|
| English | en | Native patterns |
| Spanish | es | Native patterns (accent-tolerant) |
| Portuguese | pt | Native patterns (accent-tolerant) |
| French | fr | Native patterns |
| German | de | Native patterns |
| Russian | ru | Native patterns (declension-aware) |
| Chinese | zh | Native patterns + CJK word counting |
| Japanese | ja | Native patterns (SOV word order) |
| Korean | ko | Native patterns + CJK word counting |
| Arabic | ar | Native patterns |
| Spanglish | en+es | Auto-detected |
To add a language: create languages/<code>.json with stopwords and patterns. Auto-discovered — no code changes needed.
- Pinned model IDs (haiku-4-5 / sonnet-4-6 / opus-4-7) with compact display names
- Dynamic deep-tier effort classification (medium / high / xhigh)
- Architectural promotion: non-deep → deep + xhigh on arch keywords + signals
- Advisor Strategy:
requires_advisorflag +advHUD tag - Subagent lifecycle tracking:
(subagente)HUD tag +SubagentStophook
- Multi-signal 9-weighted scoring engine
- 10-language support with accent-tolerant patterns
- Poly mascot HUD with animated states
- Cache freshness bar with prefix and warning indicators
- Auto-hook injection for zero-config setup
- 82% token reduction in additionalContext
- HUD v1.6 redesign:
[poly v1.6.2]prefix, prompt/exec split,ctx:%, rate-limit bars - Idle fallback for non-OMC users (stale session emits
[poly v1.6.2] [^.^]~ idle) - Accurate per-token savings calc with corrected Opus 4.7 pricing
- DE/FR/PT multi-file refactor patterns → deep + xhigh
- Spanish
rediseñonoun form fix
- Native
xhigheffort (no longer downgrades tohighin env) - Silent model swap detection (
⚠swapglyph when CC overrides poly's choice) - Retry-escalation arrow in HUD (
fast → deep,⚠maxat ceiling) - Advisor hand-off protocol: structured
[POLY:ADVISOR]block + manual/polyrouter:advisor - Effort override via
/polyrouter:effort <low|medium|high|xhigh>slash command - Per-session routing breakdown via
/polyrouter:stats(tier·effort, methods, languages, savings, retries) - Corpus expansion to 336 prompts with EN/ES balance (98.8% benchmark accuracy)
- Tokenizer calibration (
tokenizer_factor×1.35 for Claude 4.x family — recalibratesSavingsfigure) - Deep-pattern parity sweep — 26 patterns per language across all 10 supported langs (advanced architectural clusters)
- Stats export via
/polyrouter:export csv|jsonfor pandas / Excel / jq pipelines - Coverage report + badge — 87.6% on
hooks/lib/(details, regenerate viascripts/poly-coverage.sh) - Tier infrastructure is data-driven — adding a new tier (e.g.
ultra) is a config-only change; see docs/ADDING-A-TIER.md (CALIDAD #16) -
session_namepassthrough — per-session-name stats bucket survives/clear(requires Claude Code v2.1.120+); display truncates to 20 chars +…; verification runbook in docs/CLEAR-VERIFICATION.md (CALIDAD #17)
Items evaluated and deliberately discarded to keep polyrouter deterministic, debuggable, and cheap. Listed here so the rationale survives in repo memory and we don't re-litigate the same designs.
maxtier abovexhigh. Multi-pass Opus, Opus 1M pinning, and opus-vs-opus consensus were evaluated. Discarded —xhighremains the recommended ceiling for Opus 4.8; amaxtier would risk overthinking without measurable quality gains, while doubling per-prompt cost. The escape hatch for genuinely architectural prompts is/polyrouter:advisor, which already locks todeep/xhigh + opus-orchestrator.- Adaptive scoring thresholds (
/polyrouter:learn-on) (CALIDAD #14). A learner that would tunefast_max/standard_maxfrom routing history was evaluated. Discarded — runs against poly's deterministic philosophy ("same prompt → same tier, today and tomorrow") and carries a real risk of silent routing degradation: a single noisy retry burst could nudge thresholds into a regime where common prompts start mis-routing, and the symptom ("why did this go to deep today?") is exactly the kind of bug that's painful to triage. If routing quality drifts in practice, the right fix is to add a deterministic pattern or update the gold corpus, not to drift the boundaries.
-
execsegment in HUD (⚙ exec:opus·xhigh·adv) -
🤖Nsubagent counter - Swap detection expanded to fast/standard tiers via
PreToolUse:Task - SessionStart reset for
subagent_countandactivestate - 895 tests passing
- Verifiability routing (Karpathy) —
✓/~indicator on model segment - Native OAuth usage polling (5h/wk/snt/extra) — no ccusage dependency
- 🧠 thinking indicator when exec uses opus/deep/xhigh
- 📁 CWD basename in HUD tail
- HUD fully independent (no OMC or external plugin dep)
- HUD always complete (no width-based truncation)
- OAuth field mapping bugfixes (v1.9.2 nested response, v1.9.3 cents division +
is_enabledgate) - 909 tests passing (14 new verifiability tests)
- Context-aware prompt enhancer — long implementation prompts promote off Haiku
- Intelligent slash command routing (
/spec→ opus,/workreads the active spec) - Soul Map — learns explicit model preferences ("usa opus", "con sonnet")
- exec_model real — HUD shows the model CC actually executed (read from transcript)
- Prompt quality scorer (
q:N%) — hidden when ≥ 80% - ⏱ session elapsed time (green <30m, yellow 30-60m, red >60m →
/compactsignal) - 🤖99+ subagent counter cap for Dynamic Workflows
- Per-project session state — concurrent sessions no longer interfere
- Fallback model support — Opus → Sonnet degrade when unavailable
- Effort skew detection
- claude-opus-4-8 + CC
max→xhighmapping - 983 tests passing
- Multi-agent support: Codex CLI, Gemini CLI
- Ultra tier for next-gen models
- Weighted ensemble classification (rules + embeddings)
- Auto-escalation on repeated low-confidence routes
- Fork the repository
- Create a branch:
git checkout -b feat/my-feature - Add tests in
tests/ - Run the test suite:
python -m pytest tests/ -v - Ensure all 983+ tests pass before submitting
- Commit using conventional commits:
feat:,fix:,refactor:,test:,docs: - Open a pull request with a clear description of the change
Keep classification latency under 5ms. Maintain test coverage for all routing paths.
MIT — by SonyHarv