Skip to content

About

Intelligent model routing for Claude Code — routes haiku/sonnet/opus by complexity. Live HUD with ctx%, rate limits & savings. 10 languages, zero config.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

claude-polyrouter

claude-polyrouter

Version License Tests Languages Token Reduction Routing Accuracy Effort Accuracy Coverage

Tired of hitting Claude Code token limits? claude-polyrouter silently routes every query to the right model — stop paying Opus prices for simple questions. 82% less token waste, 10 languages at full parity, zero setup.

  • Live HUD — ctx:%, real-time 5h/wk/snt rate-limit bars with threshold-based colors, animated Poly mascot, idle fallback
  • Dynamic effort per query — medium → high → xhigh chosen from a multi-signal score (architecture, file count, tool intensity)
  • Opus 4.8 sub-effort routing — deep-tier queries get the right Opus effort level automatically; architectural calls escalate to xhigh + adv
  • 10 languages, full parity — EN · ES · FR · DE · PT · JA · KO · ZH · AR · RU all share the same deep/arch pattern coverage

How It Routes

Query Tier Model Effort Reason
"hola" / "ok" / "what is X?" Fast Haiku 4.5 low Short or simple question
"create a function" / "fix this bug" Standard Sonnet 5 medium Coding task
"debug this stack trace across 3 files" Deep Opus 4.8 high Multi-file, tool-heavy
"redesign the auth architecture" Deep Opus 4.8 xhigh + adv Architectural scope → Advisor engaged

Routing happens automatically on every query via a UserPromptSubmit hook. No manual intervention needed.


v1.9.5 Highlights

Routing más inteligente

  • Context-aware prompt enhancer: prompts de implementación largos se elevan automáticamente a Sonnet/Opus — ya no van a Haiku
  • Slash commands inteligentes: /spec siempre → Opus, /work analiza el spec activo para decidir el modelo
  • Soul Map: aprende tus preferencias explícitas de modelo (cuando mencionas "usa opus", "con sonnet", etc.)

HUD honesto

  • exec_model real: el HUD ahora muestra el modelo que CC realmente usó para ejecutar, no solo el que poly predijo
  • Prompt quality scorer: q:N% en el HUD cuando el prompt puede mejorar (oculto cuando es >= 80%)
  • ⏱ Tiempo de sesión: muestra cuánto lleva la sesión activa (verde <30m, amarillo 30-60m, rojo >60m → señal para /compact)
  • 🤖99+: soporte para Dynamic Workflows con cientos de subagentes

Robustez

  • Session state por proyecto: cada proyecto tiene su propio estado de routing — las sesiones ya no se interfieren entre sí
  • Fallback model support: si Opus no está disponible, poly degrada a Sonnet automáticamente
  • Effort skew detection: detecta cuando CC usa un effort diferente al que poly esperaba

Core

  • claude-opus-4-8: modelo actualizado al último Opus
  • CC max → xhigh: el nivel max de CC mapea al techo de poly
  • 983 tests passing

v1.9 Highlights

  • Verifiability routing (Karpathy) — Tasks with verifiable outcomes route to Opus; heuristic-only tasks stay on Sonnet. ✓/~ indicator in the HUD model segment
  • Native OAuth usage polling — 5h / wk / snt / extra bars without ccusage dependency, stale-while-revalidate cache
  • 🧠 thinking indicator — appears on the exec model when the active subagent runs opus / deep / xhigh
  • 📁 CWD in HUD — current directory basename visible on every prompt
  • HUD fully independent — runs standalone with no OMC or external plugin dependency
  • HUD always complete — no width-based truncation; all segments render
  • OAuth field mapping bugfixes — v1.9.2 (nested response), v1.9.3 (used_credits/100 + is_enabled gate)
  • 909 tests passing (14 new verifiability unit tests)

v1.8 Highlights

  • exec segment in HUD — ⚙ exec:opus·xhigh·adv now appears as a separate segment after prompt:, making the executor model visible at a glance
  • 🤖N subagent counter — Active subagent count shown in HUD (🤖1, 🤖2, …) whenever a subagent is running
  • Swap detection expanded — ⚠swap glyph now covers both fast and standard tiers via PreToolUse:Task hook (not just deep)
  • SessionStart reset — subagent_count and active state reset on every new session to prevent stale carry-over between sessions
  • 895 tests passing · 98.9% routing accuracy

v1.6 Highlights

  • HUD v1.6 — New format [poly v1.6.2], prompt/exec split, ctx:%, rate-limit bars (5h/wk/snt), ⚠compact at ctx≥70%, new mascot states [>.^] (ctx high) and [x.x] (critical)
  • Idle fallback — Stale sessions emit [poly v1.6.2] [^.^]~ idle for non-OMC users instead of a blank statusline
  • Accurate savings calc — Per-token formula (1k input + 500 output) with corrected Opus 4.7 pricing ($15/$75 per 1M tokens)
  • DE/FR/PT deep patterns — Multi-file refactor queries in German, French, and Portuguese now promote to deep + xhigh
  • 613 tests passing

v1.5 Highlights

  • Pinned model IDs — claude-haiku-4-5, claude-sonnet-4-6, claude-opus-4-7 — display stays compact, versions stay explicit
  • Dynamic deep-tier effort — medium → high → xhigh chosen per-query from multi-signal score (architecture, file count, code blocks, tool intensity)
  • xhigh tier — polyrouter-only display label for architectural/critical work; normalizes to high for the CLAUDE_CODE_EFFORT_LEVEL env var
  • Advisor Strategy — requires_advisor=true flag set automatically on xhigh; surfaces as adv in the HUD so the executor can engage the Advisor (Opus on-demand) for architectural calls
  • Subagent lifecycle tracking — HUD shows (subagente) while the spawned executor is running; SubagentStop hook clears it
  • Architectural promotion — non-deep queries with architecture keywords + standard signals auto-promote to deep + xhigh (e.g. "plan a major refactor across services")

Dynamic Escalation Examples

Query Tier Effort Advisor Why
"fix typo in README" fast low — Short, no signals
"add pagination to /api/users" standard medium — Single-file standard task
"debug this 200-line stack trace" deep high — Deep + tool-intensive
"redesign the billing subsystem" deep xhigh adv Architectural keyword match
"refactor auth.py login.py session.py as a strategic migration" deep xhigh adv Arch keyword + 3 files + orchestration

v1.4.0 Highlights

  • 82% token reduction — additionalContext reduced from ~150 tokens (v1.3) to ~27 tokens
  • 100% classification accuracy — 30/30 on multilingual test suite across all 10 languages
  • Multi-signal scoring — 9-signal weighted engine replaces simple pattern counting
  • Poly mascot HUD — Animated ASCII mascot with cache freshness bar, zero token cost
  • Auto-hook injection — Zero-config installation, hooks configured automatically

Features

  • Multi-signal scoring — 9-signal weighted scoring engine (patterns, code blocks, error traces, file paths, prompt length, tool results, conversation depth, effort level, universal tech symbols)
  • 10 languages — English, Spanish, Portuguese, French, German, Russian, Chinese, Japanese, Korean, Arabic, plus Spanglish detection
  • Zero API keys — Pure rule-based classification with pre-compiled regex patterns (~3ms latency)
  • Cost savings — 50-80% reduction by routing simple queries to cheaper models
  • Two-level cache — In-memory LRU + file-based persistent cache for repeated queries
  • Multi-turn awareness — Detects follow-up queries and maintains conversation context
  • Dynamic effort — Automatic effort level mapping (low/medium/high) based on routing tier
  • Cache keep-alive — PostToolUse hook detects prompt cache expiration risk
  • Compact advisory — Recommends context compaction when stale tool results accumulate
  • Poly mascot HUD — Animated ASCII mascot in statusLine with 6 states, cache freshness bar, zero token cost
  • Project learning — Optional knowledge base that fine-tunes routing per project
  • Analytics — Terminal stats and HTML dashboard with Charts.js visualizations

Requirements

  • Claude Code v2.1.70+
  • Python 3.8+
  • Node.js 18+

Installation

# Step 1: Add marketplace (one-time)
claude plugin marketplace add SonyHarv/claude-polyrouter

# Step 2: Install plugin
claude plugin install claude-polyrouter@claude-polyrouter

# Step 3: Restart Claude Code

That's it. On first run, the plugin auto-configures:

  • UserPromptSubmit hook in settings.json for automatic routing
  • HUD symlink for the Poly mascot statusLine
  • current symlink for version-agnostic paths

Optional integrations

Integration Purpose Install
ccusage Shows 5h/weekly/sonnet rate-limit bars in the HUD npm install -g ccusage

Combining with other plugins

Claude Code supports one statusLine command at a time. If you use poly alongside other plugins that also provide a statusLine, you can combine them with a simple wrapper script.

Example wrapper (~/.claude/statusline-wrapper.sh):

#!/bin/bash
# Combine multiple statusLine outputs into one line.
# Each plugin emits its own section — the wrapper concatenates them.

PLUGIN_A=$(node ~/.claude/plugins/plugin-a/hud.mjs 2>/dev/null)
POLY=$(node ~/.claude/plugins/cache/claude-polyrouter/claude-polyrouter/current/hud/polyrouter-hud.mjs 2>/dev/null)

parts=()
[[ -n "$PLUGIN_A" ]] && parts+=("$PLUGIN_A")
[[ -n "$POLY" ]]     && parts+=("$POLY")

printf '%s' "$(IFS='  '; echo "${parts[*]}")"

Wire it up in ~/.claude/settings.json:

{
  "statusLine": {
    "type": "command",
    "command": "bash ~/.claude/statusline-wrapper.sh"
  }
}

Note: poly is designed to be the primary statusLine. If you use a wrapper, test that the combined output fits your terminal width or enable multi-line mode in your terminal.


HUD — Poly Mascot

Poly lives in your statusLine and shows routing state at zero token cost.

No subagent:

[poly v1.9.5] [^.^]~ haiku·fast✓ │ cache:████░ ctx:8% │ 5h:45%(1h2m) wk:9%(6d19h) snt:3%(6d19h) extra:1%($2.09/$150.00) │ 📁mi-proyecto $0.03↓ es

With subagent (thinking model):

[poly v1.9.5] [^.^]~ prompt:haiku·fast ⚙ exec:opus·xhigh·adv🧠 │ 🤖1 cache:████░ ctx:15% │ 5h:45%(1h2m) wk:9%(6d19h) snt:3%(6d19h) extra:1%($2.09/$150.00) │ 📁mi-proyecto $9.50↓ es

High context (compact advisory):

[poly v1.9.5] [^.^]~ haiku·fast✓ ⚠compact │ cache:████░ ctx:78% │ 5h:45%(1h2m) wk:9%(6d19h) │ 📁mi-proyecto $0.03↓ es

Stale session (>30 min, no OMC):

[poly v1.9.5] [^.^]~ idle

HUD Element Reference

Element When shown Meaning
[poly v1.9.5] Always Plugin prefix + version
[^.^]~ / [^-^] / [>.^] / [x.x] Always Mascot state (see below)
haiku·fast / sonnet·std / opus·deep After a route Model + tier, dot-separated
✓ / ~ After a route Verifiability — ✓ task is verifiable, ~ heuristic only
·high / ·xhigh Deep tier only Sub-effort — medium is elided
·adv requires_advisor=true Advisor (Opus on-demand) engaged
prompt:… / ⚙ exec:…[🧠] Subagent active Prompt model vs executor model split
🧠 Subagent uses opus/deep/xhigh "Thinking" indicator on exec model
🤖N Subagent active Count of active subagents
cache:… Session active Freshness bar (5-block Unicode)
ctx:N% transcript_path available Context window used (from Claude Code)
5h:N%(T) OAuth usage available 5-hour limit usage + time to reset
wk:N%(T) OAuth usage available Weekly limit usage + time to reset
snt:N%(T) OAuth usage available Sonnet weekly limit usage + time to reset
extra:N%($U/$L) Max plan, extra credits enabled Extra credit usage (dollars used / monthly limit)
📁<name> Always Current working directory (basename)
⚠compact ctx ≥ 70% Context approaching limit — run /compact
$x.xx↓ Savings > $0 Cumulative estimated cost saved vs. always-Opus
es / en / pt / … Language detected ISO language code

Mascot States

State Display Trigger
Idle [^.^]~ / [^-^] Default — session active, no pressure
Routing [^o^]» Within 3 s of a new query
Thinking [^.^]... 3–10 s after query
Keepalive [~_~]zzz Cache elapsed > 40 min
Danger [°O°]!!! Cache elapsed > 50 min
Compact [^.^]~~~ compact.advisory_active flag set
Ctx High [>.^] ctx ≥ 70%
Critical [x.x] ctx ≥ 90% or any rate-limit ≥ 90%

Cache Freshness Bar

Time Display Meaning
0–10 min cache:█████ Fresh — cache is warm
10–30 min cache:████░ Warm — still healthy
30–50 min cache:███░░ ! Warning — consider a keep-alive query
50+ min cache:░░░░░ exp Expired — triggers Danger state

Works without OMC

The polyrouter HUD is self-contained. OMC (oh-my-claudecode) is optional — if present, its output is prepended to the poly line; if absent, polyrouter renders on its own.


Commands

Command Description
/polyrouter:route <tier> [query] Manual routing override
/polyrouter:stats View routing statistics
/polyrouter:dashboard Open HTML analytics dashboard
/polyrouter:config Show active configuration
/polyrouter:learn Extract routing insights from conversation
/polyrouter:learn-on Enable continuous learning mode
/polyrouter:learn-off Disable continuous learning mode
/polyrouter:knowledge View knowledge base status
/polyrouter:learn-reset Clear knowledge base
/polyrouter:retry Retry with escalated tier

Configuration

Global config at ~/.claude/polyrouter/config.json:

{
  "default_level": "fast",
  "confidence_threshold": 0.7,
  "levels": {
    "fast":     { "model": "haiku",  "model_id": "claude-haiku-4-5",  "agent": "fast-executor" },
    "standard": { "model": "sonnet", "model_id": "claude-sonnet-5", "agent": "standard-executor" },
    "deep":     { "model": "opus",   "model_id": "claude-opus-4-8",   "agent": "deep-executor" }
  },
  "scoring": {
    "thresholds": { "fast_max": 0.35, "standard_max": 0.65 }
  },
  "effort": { "auto": true },
  "keepalive": { "enabled": true, "threshold_minutes": 50 },
  "compact": { "enabled": true, "keep_last_n": 5 },
  "hud": { "mascot_enabled": true, "statusline_native": true }
}

Project override at <project>/.claude-polyrouter/config.json:

{
  "default_level": "standard",
  "confidence_threshold": 0.8
}

When new models release, update config only — no code changes needed:

{
  "levels": {
    "fast": { "model": "haiku-next", "agent": "fast-executor" }
  }
}

How It Works

  1. Exception check — Slash commands, meta-queries, and continuations bypass routing
  2. Intent override — Natural language model forcing ("use opus") takes max priority
  3. Cache lookup — Two-level cache (memory + file) for repeated queries
  4. Language detection — Stopword-based scoring identifies the query language
  5. Pattern extraction — Raw signal counting from language-specific regex patterns
  6. Multi-signal scoring — Weighted 9-signal composite score maps to tier (fast <0.35, standard <0.65, deep >=0.65)
  7. Context boost — Multi-turn awareness adjusts confidence for follow-up queries
  8. Architectural promotion — non-deep tiers with arch keywords + signals auto-promote to deep + xhigh
  9. Dynamic effort — deep tier receives a sub-effort (medium/high/xhigh) from score + signal mix
  10. Advisor flag — xhigh sets requires_advisor=true, surfaced as adv in the HUD
  11. Learned adjustments — Optional project knowledge base fine-tunes routing

Supported Languages

Language Code Notes
English en Native patterns
Spanish es Native patterns (accent-tolerant)
Portuguese pt Native patterns (accent-tolerant)
French fr Native patterns
German de Native patterns
Russian ru Native patterns (declension-aware)
Chinese zh Native patterns + CJK word counting
Japanese ja Native patterns (SOV word order)
Korean ko Native patterns + CJK word counting
Arabic ar Native patterns
Spanglish en+es Auto-detected

To add a language: create languages/<code>.json with stopwords and patterns. Auto-discovered — no code changes needed.


Roadmap

v1.5 (completed)

  • Pinned model IDs (haiku-4-5 / sonnet-4-6 / opus-4-7) with compact display names
  • Dynamic deep-tier effort classification (medium / high / xhigh)
  • Architectural promotion: non-deep → deep + xhigh on arch keywords + signals
  • Advisor Strategy: requires_advisor flag + adv HUD tag
  • Subagent lifecycle tracking: (subagente) HUD tag + SubagentStop hook

v1.4.0 (completed)

  • Multi-signal 9-weighted scoring engine
  • 10-language support with accent-tolerant patterns
  • Poly mascot HUD with animated states
  • Cache freshness bar with prefix and warning indicators
  • Auto-hook injection for zero-config setup
  • 82% token reduction in additionalContext

v1.6 (completed)

  • HUD v1.6 redesign: [poly v1.6.2] prefix, prompt/exec split, ctx:%, rate-limit bars
  • Idle fallback for non-OMC users (stale session emits [poly v1.6.2] [^.^]~ idle)
  • Accurate per-token savings calc with corrected Opus 4.7 pricing
  • DE/FR/PT multi-file refactor patterns → deep + xhigh
  • Spanish rediseño noun form fix

v1.7 (in progress)

  • Native xhigh effort (no longer downgrades to high in env)
  • Silent model swap detection (⚠swap glyph when CC overrides poly's choice)
  • Retry-escalation arrow in HUD (fast → deep, ⚠max at ceiling)
  • Advisor hand-off protocol: structured [POLY:ADVISOR] block + manual /polyrouter:advisor
  • Effort override via /polyrouter:effort <low|medium|high|xhigh> slash command
  • Per-session routing breakdown via /polyrouter:stats (tier·effort, methods, languages, savings, retries)
  • Corpus expansion to 336 prompts with EN/ES balance (98.8% benchmark accuracy)
  • Tokenizer calibration (tokenizer_factor ×1.35 for Claude 4.x family — recalibrates Savings figure)
  • Deep-pattern parity sweep — 26 patterns per language across all 10 supported langs (advanced architectural clusters)
  • Stats export via /polyrouter:export csv|json for pandas / Excel / jq pipelines
  • Coverage report + badge — 87.6% on hooks/lib/ (details, regenerate via scripts/poly-coverage.sh)
  • Tier infrastructure is data-driven — adding a new tier (e.g. ultra) is a config-only change; see docs/ADDING-A-TIER.md (CALIDAD #16)
  • session_name passthrough — per-session-name stats bucket survives /clear (requires Claude Code v2.1.120+); display truncates to 20 chars + …; verification runbook in docs/CLEAR-VERIFICATION.md (CALIDAD #17)

Cancelled / Future Research

Items evaluated and deliberately discarded to keep polyrouter deterministic, debuggable, and cheap. Listed here so the rationale survives in repo memory and we don't re-litigate the same designs.

  • max tier above xhigh. Multi-pass Opus, Opus 1M pinning, and opus-vs-opus consensus were evaluated. Discarded — xhigh remains the recommended ceiling for Opus 4.8; a max tier would risk overthinking without measurable quality gains, while doubling per-prompt cost. The escape hatch for genuinely architectural prompts is /polyrouter:advisor, which already locks to deep/xhigh + opus-orchestrator.
  • Adaptive scoring thresholds (/polyrouter:learn-on) (CALIDAD #14). A learner that would tune fast_max / standard_max from routing history was evaluated. Discarded — runs against poly's deterministic philosophy ("same prompt → same tier, today and tomorrow") and carries a real risk of silent routing degradation: a single noisy retry burst could nudge thresholds into a regime where common prompts start mis-routing, and the symptom ("why did this go to deep today?") is exactly the kind of bug that's painful to triage. If routing quality drifts in practice, the right fix is to add a deterministic pattern or update the gold corpus, not to drift the boundaries.

v1.8 (completed)

  • exec segment in HUD (⚙ exec:opus·xhigh·adv)
  • 🤖N subagent counter
  • Swap detection expanded to fast/standard tiers via PreToolUse:Task
  • SessionStart reset for subagent_count and active state
  • 895 tests passing

v1.9 (completed)

  • Verifiability routing (Karpathy) — ✓/~ indicator on model segment
  • Native OAuth usage polling (5h/wk/snt/extra) — no ccusage dependency
  • 🧠 thinking indicator when exec uses opus/deep/xhigh
  • 📁 CWD basename in HUD tail
  • HUD fully independent (no OMC or external plugin dep)
  • HUD always complete (no width-based truncation)
  • OAuth field mapping bugfixes (v1.9.2 nested response, v1.9.3 cents division + is_enabled gate)
  • 909 tests passing (14 new verifiability tests)

v1.9.5 (completed)

  • Context-aware prompt enhancer — long implementation prompts promote off Haiku
  • Intelligent slash command routing (/spec → opus, /work reads the active spec)
  • Soul Map — learns explicit model preferences ("usa opus", "con sonnet")
  • exec_model real — HUD shows the model CC actually executed (read from transcript)
  • Prompt quality scorer (q:N%) — hidden when ≥ 80%
  • ⏱ session elapsed time (green <30m, yellow 30-60m, red >60m → /compact signal)
  • 🤖99+ subagent counter cap for Dynamic Workflows
  • Per-project session state — concurrent sessions no longer interfere
  • Fallback model support — Opus → Sonnet degrade when unavailable
  • Effort skew detection
  • claude-opus-4-8 + CC max → xhigh mapping
  • 983 tests passing

v2 (planned)

  • Multi-agent support: Codex CLI, Gemini CLI
  • Ultra tier for next-gen models
  • Weighted ensemble classification (rules + embeddings)
  • Auto-escalation on repeated low-confidence routes

Contributing

  1. Fork the repository
  2. Create a branch: git checkout -b feat/my-feature
  3. Add tests in tests/
  4. Run the test suite: python -m pytest tests/ -v
  5. Ensure all 983+ tests pass before submitting
  6. Commit using conventional commits: feat:, fix:, refactor:, test:, docs:
  7. Open a pull request with a clear description of the change

Keep classification latency under 5ms. Maintain test coverage for all routing paths.


License

MIT — by SonyHarv

About

Intelligent model routing for Claude Code — routes haiku/sonnet/opus by complexity. Live HUD with ctx%, rate limits & savings. 10 languages, zero config.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages