Make a cheap model perform like an expensive one on your MCP server — by changing the server, not the prompt.
An Agent Skill that measures the gap between a weak model and a strong model on one MCP server, then closes it with small, additive, evidence-backed changes to the server's tool surface. Every change is proposed from a real failure, applied to a branch, and kept only if a re-run and a held-out set both agree it helped.
Status: experimental, v0.1.0. The harness is tested; the optimization loop costs real API tokens. Read What this does not do before you point it at production.
In Claude Code:
/plugin marketplace add akontadakis/whetstone-mcp
/plugin install whetstone-mcp@whetstone-mcp
Restart Claude Code and the skill is live. Other hosts and the manual route are in Install below.
A weak model does not fail on your server because it is weak. It fails because
the server asks it to think. Raw numbers instead of a verdict. Four primitives
instead of one workflow. A free-text argument instead of an enum. A
ValueError instead of an instruction.
Move that work into the server and the gap shrinks. The trouble is knowing
which change to make. Guessing is cheap and wrong. whetstone-mcp replaces
the guess with a measurement.
Seven steps, four gates you hold, and a loop that repeats until the gap closes or stops moving.
flowchart TD
S1["Step 1 · Setup<br/>load config, require a clean git tree"]
S2["Step 2 · Introspect<br/>list every tool, schema and annotation"]
G0{"Gate 0 — YOU<br/>label each tool safe or destructive"}
S3["Step 3 · Draft the task set<br/>60 tasks: 40 scored, 20 held out"]
G1{"Gate 1 — YOU<br/>approve the task set — it is the ruler"}
S4["Step 4 · Freeze references<br/>strong model repeats each task until the answer is stable"]
GC{"Cost gate — YOU<br/>see the exact run count and price first"}
S5["Step 5 · Baseline<br/>both models run all 60 tasks<br/>weak = target, strong = ceiling"]
subgraph SETUP["Steps 1-4 · Build the ruler — no paid runs yet"]
direction LR
S1 --> S2 --> G0 --> S3 --> G1 --> S4
end
SETUP --> GC
GC -->|approved| S5
GC -->|declined| STOP["Stop — nothing spent"]
S5 --> R1
subgraph LOOP["Step 6 · Round loop — max 6 rounds"]
direction TB
R1["A · Read the target's scored failures<br/>group them by shared server-side cause"]
R2["A · Proposal tournament<br/>3 proposers + judge, ranked cheapest class first<br/>predictions discounted by the calibration ledger<br/>proposer hidden while ranking · swap check for order bias"]
R3["A · Render round-N.md"]
G2{"Gate 2 — YOU<br/>approve all, some, none, or stop"}
R4["C · Apply the exact diffs on a round branch"]
R5["D · Build, then guardrails:<br/>corpus replay · baseline args · output compat · your test suite"]
GD{"All guardrails pass?"}
RV["Revert the whole bundle<br/>restore tasks, rebuild, record why"]
R6["E · Re-arm<br/>re-introspect, you classify any new tool, record its corpus entry"]
GE{"Bumped routes verified?"}
R7["F · Re-run target and ceiling<br/>held-out every round, scored every 2nd<br/>record the held-out delta"]
R1 --> R2 --> R3 --> G2
G2 -->|approved| R4 --> R5 --> GD
GD -->|no| RV
GD -->|yes| R6 --> GE
GE -->|no| RV
GE -->|yes| R7
end
G2 -->|none or stop| S7
RV --> GSTOP
R7 --> GSTOP{"Stop rule fires?<br/>gap closed · 6 rounds · 2 flat rounds<br/>no server-side finding · you say stop"}
GSTOP -->|no| R1
GSTOP -->|yes| S7["Step 7 · Final report<br/>rebuild, fresh baseline of every model on all 60 tasks<br/>write final.md: before, after, every kept bundle and its delta"]
classDef gate fill:#ffe9b8,stroke:#b8860b,color:#111
classDef step fill:#e8f0fe,stroke:#3b6ea5,color:#111
classDef bad fill:#fbe0e0,stroke:#b03030,color:#111
classDef done fill:#e2f3e4,stroke:#3a7d44,color:#111
class G0,G1,GC,G2 gate
class S1,S2,S3,S4,S5,R1,R2,R3,R4,R5,R6,R7 step
class GD,GE,GSTOP step
class RV,STOP bad
class S7 done
Two models run the same task set. The weak model's transcripts are the evidence. Each round proposes a bundle of changes, and a bundle survives only if the re-run improves the scored set without regressing the held-out set. Rejected bundles are reverted, not argued with.
Every stage above is checkpointed to state.json, so a crash or a rate-limit
pause resumes at the exact stage it stopped at — it never repeats an apply, a
build, or a paid run.
Changes are proposed cheapest first. You do not add a workflow tool before you have tried a better description.
| # | Class | What it changes |
|---|---|---|
| 1 | Descriptions as instructions | Tool docstrings that answer four questions: what it does, when to use it, what each argument accepts, what it returns — with one example call |
| 2 | Server instructions block |
Server-level orientation the model reads before any call |
| 3 | Constrained inputs | Enums, tight ranges, sensible defaults instead of free strings |
| 4 | Teaching errors | Errors that name the fix, not just the fault — with one valid call the model can copy |
| 5 | Verdict sentences | Pre-interpreted output fields, so the model relays instead of infers |
| 6 | Workflow tools | One composed tool replacing a sequence the model has to get right |
Everything is additive. Old tool names stay. Old output fields keep their
values. An optional argument stays optional. Three mechanical checks
(check_corpus_acceptance, check_baseline_args, check_output_compat)
enforce this after every apply, so an existing client cannot be broken by a
round it did not ask for.
Additive-only has a ceiling: the catalog only grows, and a bloated catalog is itself a routing hazard. When routing failures keep coming back after a class-6 consolidation, the skill reports that the server needs tools removed instead of proposing a seventh tool.
The judge is checked too. The three proposers and the judge are the same
model, so the ranking follows bias rules: proposals are ranked on evidence
fields with the proposer's name hidden, ties break mechanically toward the
cheaper class and then the lower risk, and a reversed-order swap check catches
a pick that was chosen by position. A proposer that keeps winning round after
round is reported under ## Judge calibration, not quietly accepted.
The skill never changes your server on its own judgement.
| Gate | You decide |
|---|---|
| 0 | Which tools are safe to call, and which are destructive and off-limits |
| 1 | The task set — the ruler everything is measured with |
| cost | The exact run count and its cost, before any model is spawned |
| 2 | Each proposed diff, in full, before it is applied |
Requires uv and Python 3.12+ for the harness.
/plugin marketplace add akontadakis/whetstone-mcp
/plugin install whetstone-mcp@whetstone-mcp
Updates then arrive with /plugin marketplace update whetstone-mcp.
Any host that reads a skills directory:
git clone https://github.com/akontadakis/whetstone-mcp.git
cp -r whetstone-mcp/skills/whetstone-mcp ~/.claude/skills/whetstone-mcp| Host | Skills directory |
|---|---|
| Claude Code, user-wide | ~/.claude/skills/ |
| Claude Code, one project | <project>/.claude/skills/ |
cd ~/.claude/skills/whetstone-mcp # or the plugin's cached path
PYTHONPATH=$PWD/harness uv run --project harness python -m pytest harness/tests -q384 tests, about two minutes. The harness builds its own virtualenv on first run, so nothing is installed into your global Python.
Point the skill at a server you own:
Optimize the MCP server at
~/code/my-mcpso Haiku performs like Opus.
It will introspect the surface, ask you to triage each tool for safety, draft a task set for your approval, show you the cost before spending anything, and then run rounds until the gap closes or stops moving.
To see a full round with no real API spend, use the sandbox below.
skills/whetstone-mcp/sandbox/ is a deliberately bad MCP server carrying one
planted defect per polish class, plus a twelve-task ruler (8 scored, 4 held
out) and an answer key in DEFECTS.md.
Tier A — free and deterministic. A weak/strong simulator whose competence is bounded by exactly those six defects speaks real MCP to the real server. Eleven tests, about a minute:
cd skills/whetstone-mcp
PYTHONPATH=$PWD/harness uv run --project harness \
python -m pytest harness/tests/test_sandbox_e2e.py -qTier B — real models, by hand. Runs the whole skill unchanged against the
sandbox. Roughly one to two hours and the cost of about 400 short sessions.
See sandbox/README.md.
Read the answer key after a run, never before.
The honest part.
- It cannot fix open-ended reasoning. Tasks that need novel synthesis across tool outputs still favour the bigger model. This closes the gap on bounded, objectively checkable work.
- It costs real tokens. A default round is 60 tasks × 2 models × 3 repeats. The cost gate shows you the number first; nothing is hidden.
- Results are observed, not settled. Every number the skill reports comes with its spread. Three repeats is not a significance test.
- It edits your server. Always on a branch, always additive, always after you approved the exact diff — but it does write code.
- It has no correctness guarantee. A round that raises the score has raised the score on your task set. The held-out set is the only defence against overfitting to it, and it is a partial one.
.claude-plugin/ plugin + marketplace manifests
skills/whetstone-mcp/
├── SKILL.md the skill — seven steps, four gates
├── references/
│ ├── polish-playbook.md the six change classes and the proposal format
│ ├── task-authoring.md task schema, anti-leakage rules, the 40/20 split
│ ├── model-traps.md how to run a phase and how to describe a number
│ └── report-template.md round and final report structure
├── harness/ the Python side — uv project, 16 modules, 16 test files
│ └── whetstone_mcp/ introspect, tasks, runner, scoring, guardrails,
│ changes, freeze, report, state, calibration
└── sandbox/ the bad server, its ruler, and the answer key
The model does the judgement — writing tasks, reading failures, proposing diffs. The harness does every mechanical part — spawning sessions, scoring, regression checks, atomic state, rendering reports. Nothing that must be reproducible is left to a model.
Issues and pull requests welcome. See CONTRIBUTING.md. The sandbox is the regression suite: a change to the skill that the sandbox cannot detect is a change that needs a new planted defect.
The idea that a weak model plus a well-engineered tool surface approaches a strong model on bounded tasks is not new — it is the practical core of Anthropic's writing tools for agents guidance. This repository is an attempt to make that loop measurable and repeatable rather than intuitive.
The judge bias rules, the four-question description contract, and the
copyable-call error rule were adapted from the tool-design and
advanced-evaluation skills in Murat Can Koylan's
Agent Skills for Context Engineering.
MIT — see LICENSE.