A skill for systematically improving other skills. Treats SKILL.md as a model trained over rounds, not prose iterated by feel.
Most skill optimization is ad-hoc: read it, feel it's weak, edit it, hope it's better. This skill enforces the discipline that catches four classes of failure:
- Self-evaluation bias — the agent that edits a skill cannot reliably score it
- Target confusion — similarly named installed, managed, and source copies can send a valid patch to the wrong artifact
- Imagination drift — optimizing without a real failure anchor produces churn, not progress
- Citation hallucination — LLM judges fabricate quotes; verify them against source
git clone <this-repo-url> ~/.claude/skills/optimizing-skillsgit clone <this-repo-url> ~/.agents/skills/optimizing-skills
# Or symlink if you already installed for Claude Code:
mkdir -p ~/.agents/skills
ln -s ~/.claude/skills/optimizing-skills ~/.agents/skills/optimizing-skillsAfter install, ask Claude/Codex something like "optimize the karpathy-guidelines skill" — it should auto-load this skill.
| File | When loaded |
|---|---|
SKILL.md |
Always (entry point: Iron Law, Core Loop, Red Flags) |
rubric.md |
When scoring (9-dim weighted rubric + dim 8 strength tiers + gap classification) |
failure-modes.md |
When blocked (if-then-fallback table for env / judge / citation issues) |
output-format.md |
When writing the round report (A-G section template) |
You suspect my-skill.md doesn't catch a class of bugs you've been seeing:
- Phase 0 — map the requested skill to one authoritative real path, ownership type, source repository, and publication target; stop on any identity conflict
- Phase 0.5 — write 2-3 test prompts based on actual past failures (not hypothetical edge cases)
- Phase 1 — score
my-skill.mdagainst the 9-dim rubric; scan recent chat history for the failure pattern; cluster by root cause - Phase 2 — propose ONE change targeting the weakest dim; commit on a branch; spawn ≥2 independent judges (ideally cross-model); keep iff score ↑ AND citations verbatim; otherwise
git revert - Phase 3 — write a report using
output-format.md(sections A-G)
Stop after two consecutive rounds with <2-point gain.
Synthesizes:
- writing-skills (Jesse Vincent / obra) — RED-GREEN-REFACTOR for skills via pressure scenarios
- darwin-skill (花叔 / alchaincyf) — 9-dim rubric + git ratchet
- SkillOpt (Microsoft Research) — validation-gated edits in text-space
Where each was the source for which mechanism:
| Mechanism | Source |
|---|---|
| RED-GREEN-REFACTOR cycle | writing-skills |
| 9-dim rubric + dim 8 weight 23 | darwin-skill |
| Phase structure + git ratchet | darwin-skill |
| Validation gate (keep iff score↑) | SkillOpt |
| Iron Law (editor ≠ judge) | this skill (original) |
| Anti-imagination rule | this skill (original) |
MIT — see LICENSE.