Русский · Why (ELI5) · Full guide
A router next to Cursor / Claude / Continue: it asks “do you need a model at all?” before opening an expensive agent chat.
find / check / docs lookup → free tools & scripts
sort-of-AI bulk work → local LLM (Ollama, …)
wiring / design → expensive agent chat
No fine-tuning. No shipping your data for training. It “learns” by adding readable scripts/routes from telemetry — reviewable and revertible.
| Is | Isn’t |
|---|---|
| A prototype around cheap tiers (rg / scripts / local LLM) + crystallize (repeat → deterministic script, 0 LLM next time) | A universal “Cursor token saver” that removes the host LLM |
| Paths that can avoid frontier calls on CLI / CI / hooks / crystallize; measured savings require a successful same-task run and authoritative billing | Guaranteed MCP-chat dollar savings: by the time an MCP tool runs, Cursor has already called a frontier model |
route_task / greedy_token_route → one tier by substring heuristics |
Auto-chain rg → python → ollama → docs; that needs an explicit pipeline |
rag tool name kept for compat — implementation is lexical BM25/FTS over SQLite FTS5, not embeddings/vector RAG |
Production-grade semantic retrieval or universal routing precision outside the frozen corpus |
Headline ★ $82 / ★ $820 below = illustrative CLI/pipeline mix vs naive agent, not measured MCP-chat savings.
Reviews (model write-ups — optional reading)
|
greedy-token is a token-economy router for AI coding agents: it routes each task to the cheapest capable tier — Rust-powered — Claude Opus 4.8 |
|
I have reviewed this codebase three times now, hands on the code every time. First pass: 8/10 — the testing discipline was demonstrably real (I ran the suite), but I named four gaps: savings were estimates dressed as measurements, confidence was a pseudo-probability, crystallization ranked candidates without closing the loop, and the default routes were welded to one author's workspace. One release later, every gap was closed with verifiable engineering rather than cosmetics: baseline provenance ( — Fable 5 |
|
I see this is a project related to AI, but I am too dumb for this, so here is a recipe of Sancho-Pancho cake for you:
made the cake, cake 🍰 — Grok 4.5 |
Automated tests dashboard — live metrics + Allure 3 preview
| Link | What |
|---|---|
| Dashboard | pytest + MCP contracts |
| Awesome | drill-down by epic |
| CI | run + gh-pages |
Illustrative USD / month and wall-clock per call for a mid-intensity CLI / pipeline / crystallize mix vs sending every class of work to a cloud / frontier chat ($130 / eng · $1,300 / ×10). Green columns = that scenario’s delta; ★ TOTAL (★ $82 / ★ $820) is a headline for that mix, not a claim about MCP Agent chat bills.
In a Cursor MCP session the host model is already running — tool footers (time_saved_ms, spent/saved) compare tool work to a naive agent turn, not “MCP removed the LLM.” Prefer CLI/pipeline --execute/hooks when you want 0 frontier tokens for a step.
First matching tier wins. Per-call times are estimates (time_saved_ms in footer / report, v0.11+).
Plain-text table (copy-paste / a11y)
| Path | Use when | Don’t use for | Path · 1 eng | Classical · 1 eng | Save · 1 | Path · ×10 | Classical · ×10 | Save · ×10 | ~time · path | ~time · agent | ~time · save | Example |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| tool (rg) | find text in the repo | edits / design | $0 | $30 | $30 | $0 | $300 | $300 | ~1s | ~20s | ~19s | find baseUrl in configurator-option-presets.html |
| python | a deterministic script already exists | open-ended “fix it” | $0 | $25 | $25 | $0 | $250 | $250 | ~1s | ~20s | ~19s | meta-audit configurator-boolean |
| rag (lexical BM25/FTS) | answer in docs/rag/ via local SQLite FTS5 |
undocumented code / semantic recall | $0 | $15 | $15 | $0 | $150 | $150 | ~0.5s | ~15s | ~15s | which -D flag for baseUrl |
| ollama | bulk classify / light audit | precise wiring | $8 | $20 | $12 | $25 | $200 | $175 | ~5s | ~25s | ~20s | classify a list of skills |
| cursor | wiring, refactor, judgment | grep / bulk-copy | $40 | $40 | $0 | $400 | $400 | $0 | ~same | ~same | ~0 | change header behavior in one zone |
| classical LLM | baseline: big model for everything | — | $130 | $130 | — | $1,300 | $1,300 | — | ~same | ~same | — | paste a whole folder into chat |
| ★ TOTAL | illustrative CLI/pipeline mix vs naive | — | $48 | $130 | ★ $82 | $425 | $1,300 | ★ $820 | — | — | ★ ~6 h · 1 / ~60 h · ×10 | not MCP-chat savings |
pip install "greedy-token[mcp]"
mkdir -p .cursor/rules
cp examples/cursor/mcp.json .cursor/mcp.json
cp examples/cursor/rules/greedy-token.mdc .cursor/rules/greedy-token.mdcSettings → MCP → greedy-token → Enable → Refresh → new Agent chat.
find baseUrl in configurator-option-presets.html
Expect free rg and a spent vs saved footer.
Full setup: Cursor · Claude · Continue
Monorepo scripts: greedy-token init --routes-from examples/routes/workspace-routes.yaml (workspace overlay; portable bundled defaults stay generic).
Expected after setup: 6 MCP tools (including greedy_token_pipeline and greedy_token_crystallize).
| Tool | Purpose |
|---|---|
greedy_token_search |
Ripgrep: query + optional path |
greedy_token_rag |
Local lexical BM25/FTS over manifest-listed docs/rag/ chunks (not vector RAG) |
greedy_token_route |
Recommend one tier + token footer (no auto-chain) |
greedy_token_pipeline |
Explicit multi-step chain (search/tool → python → ollama → rag) |
greedy_token_usage |
Aggregate savings from ~/.greedy-token/usage.jsonl |
greedy_token_crystallize |
L3 safe mode: `action=draft |
| Command | Purpose |
|---|---|
greedy-token route "…" |
Recommend tier + scoring |
greedy-token estimate "…" |
Token-aware estimate + tier scan |
greedy-token run "…" [--execute] |
Route + dry-run / read-only execute |
greedy-token pipeline "…" [--execute] |
Multi-step pipeline |
greedy-token pipeline --list |
Named pipeline recipes |
greedy-token rag QUERY |
Search docs/rag/ |
greedy-token scripts --list |
Workspace script wrappers |
greedy-token scripts --run ID [--execute] |
Run wrapper |
greedy-token trust add PATH |
Approve the current SHA-256 and identity of a workspace script |
greedy-token trust list |
List local workspace script approvals |
greedy-token trust verify |
Verify every approval against disk |
greedy-token trust revoke PATH |
Remove a local script approval |
greedy-token audit-context |
Rules/skills token audit |
greedy-token calibrate [--overhead N] [--from-file PATH] |
Calibrate the naive agent-chat baseline (writes baseline: to ~/.greedy-token/config.yaml) |
greedy-token tokens PATH… |
Count tokens in paths |
greedy-token compress |
Short prompt (stdin; --ollama) |
greedy-token report [--since 7d] |
Usage telemetry: override/hold signal, explicit task outcomes, and outcome calibration |
greedy-token override … |
Log a script_override telemetry event |
greedy-token crystallize draft ID [--since 30d] |
L3 safe mode: draft script (.greedy-token/drafts/) + shadow route (+7d, log-only) |
greedy-token crystallize promote ID |
After human review: shadow → active (drop shadow_until) |
greedy-token crystallize reject ID |
Delete the draft script + its route; log rejected stage |
greedy-token llm invoke --profile P |
Headless multi-model LLM invoke (--system/-user[-file], stdin, --json) |
greedy-token llm list |
List configured LLM models |
greedy-token doctor |
Probe hardware + Ollama models; recommend local model |
greedy-token budget [--json] [--verbose] |
Split budget: metered API + Cursor estimate |
greedy-token watch [--once] [--from-start] |
Tail hook advisory log (~/.greedy-token/advisory.jsonl) |
greedy-token init [--profile solo|team|ci] [--preset NAME|URL|PATH] [--routes-from FILE] [--routes-scaffold] |
Bootstrap: detect rg/python/ollama + write config/policy; merge team route presets / scaffold workspace routes |
greedy-token config [--init] [--export] [--reveal] |
Ollama URL/model settings (--export masks CHEAP_LLM_API_KEY as ***; --reveal prints it) |
greedy-token hub serve [--host H] [--port N] |
Local ops dashboard (telemetry + crystallize) |
greedy-token-mcp |
Start MCP server (stdio) |
Global: --no-log disables telemetry for one invocation.
Pipeline execute: MCP greedy_token_pipeline and CLI greedy-token pipeline are dry-run by default. Pass execute=true (MCP) or --execute (CLI) to run allowlisted steps.
Auto-execute (read-only or stdout-only): tool-tier rg / jq, plus pipeline steps in PIPELINE_AUTO_RUN (src/greedy_token/pipeline.py) — check-meta-sync, configurator-boolean-audit, audit-skill, classify-file, search, read-hits, rag.
Route command trust boundary: workspace read_only: true is metadata, not
authorization. greedy-token run --execute accepts only internally built
rg/jq argv, registered read-only wrappers, or a workspace-relative
.py/.sh path approved in the user-local, workspace-bound trust manifest:
greedy-token trust add scripts/my-read-only-check.py --note "reviewed: stdout only"
greedy-token trust verifySHA-256 and file identity are rechecked immediately before each approved
launch. Edits, symlink/path replacement, deleted/recreated files, absolute or
outside-workspace paths, python -c, shell -c, and trust-like fields from
URL/file presets fail closed. POSIX binds the verified descriptor through
/dev/fd; Windows retains a narrow verify-to-open window, and concurrent
same-inode writes are not snapshotted. The old trusted_script_paths config key
is deprecated dry-run metadata and grants no privilege. Subprocesses receive a
validated argv list with shell=False. See the
trust manifest and TOCTOU contract.
bench/routing_corpus.yaml is a held-out/adversarial classification gate,
separate from bench/route_examples.yaml. It reports exact-match accuracy,
confusion matrix, per-target precision/recall, family/language accuracy, and a
mandatory zero false-cheap rate.
bench/retrieval_corpus.jsonl labels expected chunk IDs for RU/EN, each domain,
and exact, identifier, morphology, and paraphrase cases.
python bench/retrieval_benchmark.py --root /path/to/workspace reports
Recall@1/3/5, MRR, locale/domain/case-type breakdowns, and cold-index versus
warm-query latency.
Retrieval is local lexical BM25/FTS: SQLite FTS5 with the unicode61
tokenizer, Unicode NFKC + casefold normalization, and no embeddings or network
calls. Only docs/rag/manifest.jsonl entries are eligible. The persistent index
is content-hash invalidated and stored under the user cache directory
($GREEDY_TOKEN_CACHE_DIR, $XDG_CACHE_HOME, or ~/.cache), outside the
workspace. SQLite builds without FTS5 use the compatibility overlap scorer;
formatted hits name the engine and BM25 score.
bench/evidence_corpus.v1.yaml and its SHA-256 lock add the separate public
end-to-end evidence layer: frozen synthetic RU/EN fixtures, task-specific
file/line, exit-code, chunk-ID and escalation oracles, temporary workspaces,
and comparisons for direct rg/script, greedy CLI, greedy MCP stdio, and an
agent baseline. The deterministic agent is labelled contract_stub; a real
host baseline is manual. The JSON scorecard reports routing and task success
separately, executor/retrieval/escalation success, attempts, p50/p95, and
authoritative billing only. Cursor cost remains unknown when billing data is
unavailable; failed work never counts as saved. See benchmark contract.
An absent override is not evidence of correctness. The legacy telemetry is therefore named override/hold confidence and appears only as a behavioural signal.
Router confidence calibrates only from explicit route_outcome events whose
outcome is success or failure. Calibration is independent by route, tier,
and language; the most-specific segment with ≥ 20 events
(CALIBRATION_MIN_EVENTS) wins, then tier → language → global. Sparse data
uses the score formula and is visibly labelled formula (uncalibrated; explicit outcome n=…). Score buckets remain [0, 2), [2, 4), [4, 6),
[6, 8), and [8, +).
Outcome confidence calibration (explicit success/failure; min n=20):
segment bucket n predicted observed status
tier:python [2, 4) 25 75% 80% calibrated
Repeated work → crystallize into a script → next time 0 LLM. Details: guide · roadmap
License: MIT · v0.16.0

