Skip to content

About

Make a cheap model perform like an expensive one on your MCP server — by changing the server, not the prompt. An Agent Skill that measures the weak/strong model gap and closes it with additive, evidence-backed tool-surface changes.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

Repository files navigation

whetstone-mcp

Make a cheap model perform like an expensive one on your MCP server — by changing the server, not the prompt.

An Agent Skill that measures the gap between a weak model and a strong model on one MCP server, then closes it with small, additive, evidence-backed changes to the server's tool surface. Every change is proposed from a real failure, applied to a branch, and kept only if a re-run and a held-out set both agree it helped.

Status: experimental, v0.1.0. The harness is tested; the optimization loop costs real API tokens. Read What this does not do before you point it at production.


Install in two lines

In Claude Code:

/plugin marketplace add akontadakis/whetstone-mcp
/plugin install whetstone-mcp@whetstone-mcp

Restart Claude Code and the skill is live. Other hosts and the manual route are in Install below.


Why

A weak model does not fail on your server because it is weak. It fails because the server asks it to think. Raw numbers instead of a verdict. Four primitives instead of one workflow. A free-text argument instead of an enum. A ValueError instead of an instruction.

Move that work into the server and the gap shrinks. The trouble is knowing which change to make. Guessing is cheap and wrong. whetstone-mcp replaces the guess with a measurement.

The whole workflow

Seven steps, four gates you hold, and a loop that repeats until the gap closes or stops moving.

flowchart TD
    S1["Step 1 · Setup<br/>load config, require a clean git tree"]
    S2["Step 2 · Introspect<br/>list every tool, schema and annotation"]
    G0{"Gate 0 — YOU<br/>label each tool safe or destructive"}
    S3["Step 3 · Draft the task set<br/>60 tasks: 40 scored, 20 held out"]
    G1{"Gate 1 — YOU<br/>approve the task set — it is the ruler"}
    S4["Step 4 · Freeze references<br/>strong model repeats each task until the answer is stable"]
    GC{"Cost gate — YOU<br/>see the exact run count and price first"}
    S5["Step 5 · Baseline<br/>both models run all 60 tasks<br/>weak = target, strong = ceiling"]

    subgraph SETUP["Steps 1-4 · Build the ruler — no paid runs yet"]
      direction LR
      S1 --> S2 --> G0 --> S3 --> G1 --> S4
    end

    SETUP --> GC
    GC -->|approved| S5
    GC -->|declined| STOP["Stop — nothing spent"]
    S5 --> R1

    subgraph LOOP["Step 6 · Round loop — max 6 rounds"]
      direction TB
      R1["A · Read the target's scored failures<br/>group them by shared server-side cause"]
      R2["A · Proposal tournament<br/>3 proposers + judge, ranked cheapest class first<br/>predictions discounted by the calibration ledger<br/>proposer hidden while ranking · swap check for order bias"]
      R3["A · Render round-N.md"]
      G2{"Gate 2 — YOU<br/>approve all, some, none, or stop"}
      R4["C · Apply the exact diffs on a round branch"]
      R5["D · Build, then guardrails:<br/>corpus replay · baseline args · output compat · your test suite"]
      GD{"All guardrails pass?"}
      RV["Revert the whole bundle<br/>restore tasks, rebuild, record why"]
      R6["E · Re-arm<br/>re-introspect, you classify any new tool, record its corpus entry"]
      GE{"Bumped routes verified?"}
      R7["F · Re-run target and ceiling<br/>held-out every round, scored every 2nd<br/>record the held-out delta"]

      R1 --> R2 --> R3 --> G2
      G2 -->|approved| R4 --> R5 --> GD
      GD -->|no| RV
      GD -->|yes| R6 --> GE
      GE -->|no| RV
      GE -->|yes| R7
    end

    G2 -->|none or stop| S7
    RV --> GSTOP
    R7 --> GSTOP{"Stop rule fires?<br/>gap closed · 6 rounds · 2 flat rounds<br/>no server-side finding · you say stop"}
    GSTOP -->|no| R1
    GSTOP -->|yes| S7["Step 7 · Final report<br/>rebuild, fresh baseline of every model on all 60 tasks<br/>write final.md: before, after, every kept bundle and its delta"]

    classDef gate fill:#ffe9b8,stroke:#b8860b,color:#111
    classDef step fill:#e8f0fe,stroke:#3b6ea5,color:#111
    classDef bad fill:#fbe0e0,stroke:#b03030,color:#111
    classDef done fill:#e2f3e4,stroke:#3a7d44,color:#111
    class G0,G1,GC,G2 gate
    class S1,S2,S3,S4,S5,R1,R2,R3,R4,R5,R6,R7 step
    class GD,GE,GSTOP step
    class RV,STOP bad
    class S7 done
Loading

Two models run the same task set. The weak model's transcripts are the evidence. Each round proposes a bundle of changes, and a bundle survives only if the re-run improves the scored set without regressing the held-out set. Rejected bundles are reverted, not argued with.

Every stage above is checkpointed to state.json, so a crash or a rate-limit pause resumes at the exact stage it stopped at — it never repeats an apply, a build, or a paid run.

The six polish classes

Changes are proposed cheapest first. You do not add a workflow tool before you have tried a better description.

# Class What it changes
1 Descriptions as instructions Tool docstrings that answer four questions: what it does, when to use it, what each argument accepts, what it returns — with one example call
2 Server instructions block Server-level orientation the model reads before any call
3 Constrained inputs Enums, tight ranges, sensible defaults instead of free strings
4 Teaching errors Errors that name the fix, not just the fault — with one valid call the model can copy
5 Verdict sentences Pre-interpreted output fields, so the model relays instead of infers
6 Workflow tools One composed tool replacing a sequence the model has to get right

Everything is additive. Old tool names stay. Old output fields keep their values. An optional argument stays optional. Three mechanical checks (check_corpus_acceptance, check_baseline_args, check_output_compat) enforce this after every apply, so an existing client cannot be broken by a round it did not ask for.

Additive-only has a ceiling: the catalog only grows, and a bloated catalog is itself a routing hazard. When routing failures keep coming back after a class-6 consolidation, the skill reports that the server needs tools removed instead of proposing a seventh tool.

The judge is checked too. The three proposers and the judge are the same model, so the ranking follows bias rules: proposals are ranked on evidence fields with the proposer's name hidden, ties break mechanically toward the cheaper class and then the lower risk, and a reversed-order swap check catches a pick that was chosen by position. A proposer that keeps winning round after round is reported under ## Judge calibration, not quietly accepted.

You hold four gates

The skill never changes your server on its own judgement.

Gate You decide
0 Which tools are safe to call, and which are destructive and off-limits
1 The task set — the ruler everything is measured with
cost The exact run count and its cost, before any model is spawned
2 Each proposed diff, in full, before it is applied

Install

Requires uv and Python 3.12+ for the harness.

Claude Code plugin (recommended)

/plugin marketplace add akontadakis/whetstone-mcp
/plugin install whetstone-mcp@whetstone-mcp

Updates then arrive with /plugin marketplace update whetstone-mcp.

Manual

Any host that reads a skills directory:

git clone https://github.com/akontadakis/whetstone-mcp.git
cp -r whetstone-mcp/skills/whetstone-mcp ~/.claude/skills/whetstone-mcp
Host Skills directory
Claude Code, user-wide ~/.claude/skills/
Claude Code, one project <project>/.claude/skills/

Check it works

cd ~/.claude/skills/whetstone-mcp     # or the plugin's cached path
PYTHONPATH=$PWD/harness uv run --project harness python -m pytest harness/tests -q

384 tests, about two minutes. The harness builds its own virtualenv on first run, so nothing is installed into your global Python.

Quick start

Point the skill at a server you own:

Optimize the MCP server at ~/code/my-mcp so Haiku performs like Opus.

It will introspect the surface, ask you to triage each tool for safety, draft a task set for your approval, show you the cost before spending anything, and then run rounds until the gap closes or stops moving.

To see a full round with no real API spend, use the sandbox below.

The sandbox

skills/whetstone-mcp/sandbox/ is a deliberately bad MCP server carrying one planted defect per polish class, plus a twelve-task ruler (8 scored, 4 held out) and an answer key in DEFECTS.md.

Tier A — free and deterministic. A weak/strong simulator whose competence is bounded by exactly those six defects speaks real MCP to the real server. Eleven tests, about a minute:

cd skills/whetstone-mcp
PYTHONPATH=$PWD/harness uv run --project harness \
  python -m pytest harness/tests/test_sandbox_e2e.py -q

Tier B — real models, by hand. Runs the whole skill unchanged against the sandbox. Roughly one to two hours and the cost of about 400 short sessions. See sandbox/README.md.

Read the answer key after a run, never before.

What this does not do

The honest part.

  • It cannot fix open-ended reasoning. Tasks that need novel synthesis across tool outputs still favour the bigger model. This closes the gap on bounded, objectively checkable work.
  • It costs real tokens. A default round is 60 tasks × 2 models × 3 repeats. The cost gate shows you the number first; nothing is hidden.
  • Results are observed, not settled. Every number the skill reports comes with its spread. Three repeats is not a significance test.
  • It edits your server. Always on a branch, always additive, always after you approved the exact diff — but it does write code.
  • It has no correctness guarantee. A round that raises the score has raised the score on your task set. The held-out set is the only defence against overfitting to it, and it is a partial one.

Repo layout

.claude-plugin/                 plugin + marketplace manifests
skills/whetstone-mcp/
├── SKILL.md                  the skill — seven steps, four gates
├── references/
│   ├── polish-playbook.md    the six change classes and the proposal format
│   ├── task-authoring.md     task schema, anti-leakage rules, the 40/20 split
│   ├── model-traps.md        how to run a phase and how to describe a number
│   └── report-template.md    round and final report structure
├── harness/                  the Python side — uv project, 16 modules, 16 test files
│   └── whetstone_mcp/        introspect, tasks, runner, scoring, guardrails,
│                             changes, freeze, report, state, calibration
└── sandbox/                  the bad server, its ruler, and the answer key

The model does the judgement — writing tasks, reading failures, proposing diffs. The harness does every mechanical part — spawning sessions, scoring, regression checks, atomic state, rendering reports. Nothing that must be reproducible is left to a model.

Contributing

Issues and pull requests welcome. See CONTRIBUTING.md. The sandbox is the regression suite: a change to the skill that the sandbox cannot detect is a change that needs a new planted defect.

Credits

The idea that a weak model plus a well-engineered tool surface approaches a strong model on bounded tasks is not new — it is the practical core of Anthropic's writing tools for agents guidance. This repository is an attempt to make that loop measurable and repeatable rather than intuitive.

The judge bias rules, the four-question description contract, and the copyable-call error rule were adapted from the tool-design and advanced-evaluation skills in Murat Can Koylan's Agent Skills for Context Engineering.

License

MIT — see LICENSE.

About

Make a cheap model perform like an expensive one on your MCP server — by changing the server, not the prompt. An Agent Skill that measures the weak/strong model gap and closes it with additive, evidence-backed tool-surface changes.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages