A Claude Code skill for multi-agent coding pipelines, built on research from Google DeepMind and Anthropic.
"The space of interesting harness combinations doesn't shrink as models improve. Instead, it moves." — Prithvi Rajasekaran, Anthropic Labs
Abdul Bari — CS (AI/ML), San Francisco State University.
In early 2026, two research groups published work on the same problem within weeks of each other. DeepMind's Aletheia team was building agents for autonomous math research. Anthropic's Labs team was building agents for long-running software engineering. Neither group was aware of the other's work, but they landed on the same core finding: if you separate the agent that writes code from the agent that judges it, the results improve more than any other single intervention.
What caught my attention was that each team also discovered techniques the other missed. Aletheia's verifier is decoupled from the generator's chain-of-thought, so it can't be led down the same reasoning path that produced a flawed solution. Anthropic's evaluator grades against concrete criteria and actually runs the code, clicking through live applications. But Aletheia had no planner (math problems arrive pre-specified), and Anthropic's evaluator shares the same agent context as the generator (no chain-of-thought decoupling).
This project combines both, and adds one thing of its own: blind pre-analysis. Before the evaluator sees any generated code, it reasons about what the correct approach should be. Aletheia decoupled the verifier from the generator's reasoning; this goes a step further and decouples it from the solution itself. The evaluator forms expectations first, then compares. That's not from either paper. That's mine.
The result is a Planner → Generator → Evaluator → Reviser pipeline that neither team built individually.
Models are bad at judging their own work. This isn't a skill issue that scales away with better models; it's structural. When you ask an agent to evaluate code it just wrote, it tends to praise it. Rajasekaran put it plainly: agents "tend to respond by confidently praising the work, even when, to a human observer, the quality is obviously mediocre."
The other failure modes are familiar to anyone who's used agentic coding tools. Models lose coherence on long tasks. They stub features and move on. They solve the easy version of the problem instead of the one you asked for. These patterns persist across model generations.
The question worth asking isn't whether the next model will be better at single-pass generation (it will), but whether you can get more out of it by structuring the workflow around its blind spots.
Paper: Towards Autonomous Mathematics Research Authors: Tony Feng, Trieu H. Trinh, et al. Code: github.com/google-deepmind/superhuman/tree/main/aletheia
Aletheia is a math research agent built on Gemini Deep Think. It produced publication-grade mathematics, including a full research paper with zero human intervention. The architecture is a Generator → Verifier → Reviser loop in natural language.
The important thing about Aletheia's verifier is what it doesn't see. It receives the problem and the candidate solution, but never the generator's chain-of-thought. The reasoning tokens that led to the solution are stripped away. This matters because those tokens can be persuasive even when they're wrong. From the paper: "Decoupling a reasoning model's final output from its intermediate thinking tokens, and adding well-chosen prompt scaffolding, enables the model to recognize flaws it initially overlooked during generation."
The verification prompt (Appendix A of the FirstProof companion paper) instructs the evaluator to "rigorously evaluate a problem and candidate solution," looking for "logical fallacies, unstated assumptions, calculation errors, or lack of rigor." The verifier returns one of three verdicts: CORRECT, FIXABLE, or WRONG.
Key results:
- The Generate-Verify-Revise loop beat raw Deep Think at equivalent compute budgets
- On PhD-level problems, the agent only answered ~60% of questions but was right >82% of the time on those it attempted. It knows when it doesn't know.
- Structured verdicts forced clear decisions. No hedging.
Post: anthropic.com/engineering/harness-design-long-running-apps Author: Prithvi Rajasekaran, Anthropic Labs
Anthropic's Labs team built a multi-agent architecture for full-stack application development, inspired by GANs. The architecture evolved through three versions:
- V1 (frontend): Generator + Evaluator. The evaluator used Playwright MCP to navigate live pages, take screenshots, and interact with the application before scoring. 5-15 iterations, about 4 hours per run.
- V1.5 (full-stack): Added a Planner. Introduced sprint contracts where the Generator and Evaluator negotiate what "done" means before each sprint.
- V2 (Opus 4.6): Dropped sprints entirely. The model sustains coherence well enough that sprint decomposition became overhead. Simplified to Planner + Generator + a single evaluation pass at the end.
The two techniques that matter most:
Criteria-based evaluation. Instead of asking "is this good?", the evaluator grades against specific dimensions: design quality, originality, craft, functionality. Each has a weight and a threshold. This turns a subjective judgment into something a model can do reliably. The criteria also shaped the generator's behavior before any evaluator feedback, just by being part of the system prompt.
Adversarial separation. "Tuning a standalone evaluator to be skeptical turns out to be far more tractable than making a generator critical of its own work."
Key results:
- Without the Planner, the generator under-scopes. Given a raw prompt, it builds less than it would if you spec'd the work first.
- The evaluator catches real bugs, but needs calibration. "Out of the box, Claude is a poor QA agent. In early runs, I watched it identify legitimate issues, then talk itself into deciding they weren't a big deal and approve the work anyway."
- The full harness costs about 10-20x a solo agent run (solo: ~20 min/$9 vs. full harness: ~6 hr/$200 on a retro game maker). Worth it for complex tasks. Not for simple ones.
Both teams found, independently, that separating generation from evaluation is the highest-leverage move. Beyond that, each discovered things the other didn't:
| Technique | Source | How it's used here |
|---|---|---|
| CoT-decoupled verification | DeepMind (Aletheia) | Evaluator is a separate agent, decoupled from the generator's reasoning |
| Blind pre-analysis | This harness (extending Aletheia) | Evaluator reasons about the correct approach before seeing any code |
| Structured verdicts (CORRECT / FIXABLE / WRONG) | DeepMind (Aletheia) | Forces clear decisions instead of vague feedback |
| Planner for spec expansion | Anthropic (V1.5+) | Short prompt becomes a full product spec before any code is written |
| Criteria-based grading | Anthropic | Concrete, weighted dimensions replace subjective quality judgments |
| Live application testing | Anthropic | Evaluator runs the code, not just reads it. Uses claude-in-chrome (teammate on macOS/Linux) or browser evidence files (Windows fallback) |
| Teammate-based evaluator | This harness (Claude Code adaptation) | Evaluator runs as a Claude Code teammate with its own session, MCP access, and full context isolation |
| Sprint contracts | Anthropic (V1.5 only; dropped in V2) | Historical. Replaced by single-pass evaluation. |
| WRONG is terminal | DeepMind (Aletheia) | WRONG stops the pipeline; no revision attempted on fundamentally flawed approaches |
| Abstention | DeepMind (Aletheia) | Pipeline can output "no correct solution found" rather than delivering broken code |
| Revision with specific feedback | DeepMind (Aletheia) | FIXABLE verdict triggers targeted fixes, not restarts |
| Few-shot evaluator calibration | Anthropic | Concrete scoring examples anchor the evaluator's judgment thresholds |
| Logging | Both | Full pipeline trace preserved for tuning |
The novel piece is blind pre-analysis. Aletheia hides the generator's reasoning from the verifier. Anthropic doesn't decouple them at all. This harness hides the entire solution from the evaluator until it's formed its own expectations about what correct looks like. It's a stronger form of the same principle Aletheia identified: if you want honest evaluation, limit what the evaluator can anchor on.
The entire evaluation runs as a teammate — an independent Claude Code session that has never seen the generator's conversation context. Teammates (unlike regular subagents) load the full project context including MCP servers, so the evaluator can use claude-in-chrome for live browser testing on macOS/Linux. On Windows, a known bug (anthropics/claude-code#30499) prevents teammates from accessing chrome MCP tools; the main agent collects browser evidence and the teammate evaluates from files. The pre-analysis focuses on testable behavioral expectations — what the code should do, not how it should be structured — to avoid penalizing valid alternative approaches.
The full evaluator sequence: blind pre-analysis → run and test the code → grade against criteria → structured verdict.
┌─────────────┐
│ PLANNER │ Expands task into full spec
└──────┬──────┘
│
▼
┌─────────────┐
┌────▶│ GENERATOR │ Builds against the spec
│ └──────┬──────┘
│ │
│ ▼
│ ┌─────────────┐
│ │ EVALUATOR │ 1. Blind pre-analysis
│ │ │ 2. Run and test the code
│ │ │ 3. Grade against criteria
│ │ │ 4. CORRECT / FIXABLE / WRONG
│ └──────┬──────┘
│ │
│ ┌──────┴──────┐
│ │ │
│ FIXABLE CORRECT ───▶ DONE
│ │
│ │ WRONG ────▶ STOP (explain to user)
└─────┘
Planner (from Anthropic)
- Takes a 1-4 sentence task description and expands it into a full specification
- Focuses on what the system should do, not how to implement every piece
- Rajasekaran found that over-specified plans cascade errors downstream: "If the planner tried to specify granular technical details upfront and got something wrong, the errors in the spec would cascade into the downstream implementation"
Generator (shared)
- Implements the spec feature by feature
- Does a quick self-check before handing off, but the real assessment comes from the Evaluator
- On revision: fixes specific issues from evaluator feedback without touching what already works
Evaluator (synthesis of both teams + this harness's blind pre-analysis)
The entire evaluation runs as a teammate — an independent Claude Code session, completely decoupled from the generator's context. Teammates have access to MCP tools (including claude-in-chrome for live browser testing on macOS/Linux).
- Step 1, Blind Pre-Analysis: Before looking at any code, reason about what correct behavior looks like. What should the code do? What edge cases must it handle? What would be fundamentally wrong? Focuses on testable expectations, not implementation preferences.
- Step 2, Concrete Testing: Run it. Execute tests, check types, lint. If it's a server, hit the endpoints. If it's a frontend, use chrome to interact with it (or read browser evidence from the main agent on Windows).
- Step 3, Criteria Grading: Score against weighted dimensions, calibrated with few-shot examples (see below).
- Step 4, Verdict: CORRECT, FIXABLE, or WRONG. FIXABLE lists specific issues with severity, location, and fix direction. WRONG is terminal — the pipeline stops and explains the architectural failure rather than attempting revision.
Adapted from Anthropic's framework, extended for general software:
| Criterion | Weight | What it checks |
|---|---|---|
| Correctness | Critical | Does it solve the actual problem? Are all code paths valid? |
| Completeness | High | Every requirement addressed? No stubs or placeholders? |
| Security | High | Input validation, auth, injection prevention, data protection |
| Resilience | Medium | Error handling, edge cases, graceful degradation |
| Code Quality | Medium | Readable, maintainable, follows existing conventions |
| Design Quality | Medium (frontend only) | Coherent visual identity. No purple-gradient-over-white-card AI slop. |
FIXABLE means all Critical criteria pass but there are identifiable issues. WRONG means the approach is fundamentally flawed — the pipeline stops and presents the evaluator's architectural critique rather than attempting revision (matching Aletheia's design where WRONG is terminal). The evaluator is calibrated with few-shot examples showing what each verdict level looks like in practice (see evaluator.md).
git clone https://github.com/zhadyz/aletheia-harness.git
mkdir -p ~/.claude/skills/aletheia
cp aletheia-harness/SKILL.md ~/.claude/skills/aletheia/SKILL.md
cp aletheia-harness/evaluator.md ~/.claude/skills/aletheia/evaluator.md
cp aletheia-harness/planner.md ~/.claude/skills/aletheia/planner.mdOr: cd aletheia-harness && bash install.sh
mkdir -p .claude/skills/aletheia
cp aletheia-harness/SKILL.md .claude/skills/aletheia/SKILL.md
cp aletheia-harness/evaluator.md .claude/skills/aletheia/evaluator.md
cp aletheia-harness/planner.md .claude/skills/aletheia/planner.md/aletheia Build a rate limiter middleware for Fastify using Redis with sliding window algorithm
/aletheia review src/routes/auth.ts
/aletheia architect Design a webhook processing system with retry logic and dead letter queues
/aletheia quick Fix the N+1 query in the user dashboard endpoint
/aletheia runs the full pipeline. quick skips the planner and does a single evaluation pass. review evaluates existing code. architect produces a design proposal with adversarial evaluation.
1. The generator will produce flawed output. Plan for it. Both papers found this. The harness exists to catch errors, not prevent them.
2. Evaluate blind, then evaluate concrete. The evaluator forms expectations before it sees the code, then tests the code against reality. This two-phase approach catches both conceptual flaws (wrong approach) and implementation flaws (right approach, buggy execution).
3. Grade against criteria, not feelings. "Is this code good?" is not a question a model can answer reliably. "Does this handle auth token expiry?" is.
4. Run the code. Anthropic's evaluator caught critical bugs by clicking through the live application. Static review misses integration issues, broken wiring, and behavior that only shows up at runtime.
5. Fix, don't restart. And know when to stop. FIXABLE means the approach is sound but there are specific problems. Fix those problems. WRONG means the approach is fundamentally flawed — stop, explain, and let the user decide. Published Aletheia artifacts show at most one revision pass; the 3-iteration cap on FIXABLE is my own design choice. If the pipeline can't produce a correct solution, it says so rather than delivering broken code with confidence.
6. Every harness component is a bet against the model. "Those assumptions are worth stress testing, both because they may be incorrect, and because they can quickly go stale as models improve." If the model gets good enough to not need a component, drop it.
7. Log everything. Both teams emphasized this. Aletheia published all prompts and outputs. Anthropic tuned their evaluator by reading its logs. The skill writes all phases to .aletheia/ for the same reason.
This is meant to be modified. What works depends on the model, the task, and the domain.
Use the full pipeline for:
- Complex, multi-feature builds
- Ambiguous scope that benefits from spec expansion
- Production-critical code
Skip the Planner for:
- Tasks with clear, well-specified requirements
- Bug fixes where you already know the scope
Skip the Evaluator for:
- Simple, isolated functions
- Tasks well within the model's comfort zone
- When speed matters more than thoroughness
The full pipeline costs roughly 10-20x what a solo agent costs. Rajasekaran: "The evaluator is not a fixed yes-or-no decision. It is worth the cost when the task sits beyond what the current model does reliably solo."
The evaluator prompt ships with few-shot calibration examples showing what CORRECT, FIXABLE, and WRONG verdicts look like in practice. Rajasekaran was direct about why this matters: "Out of the box, Claude is a poor QA agent." His team went through multiple cycles of reading evaluator logs and adjusting the prompt before the evaluator produced reliable judgments. The few-shot examples provide scoring anchors that reduce drift.
Watch for these patterns in your .aletheia/evaluation.md outputs:
- Rubber-stamping: finds real issues, then approves anyway
- Style over substance: obsesses over naming conventions while missing logic bugs
- False strictness: fails code for criteria that don't apply to the task
- Stub blindness: gives CORRECT to code with TODO comments and empty function bodies
- Structure bias: penalizes valid alternative approaches because they don't match the evaluator's preferred implementation style
Read the evaluator's output after your first few runs. Adjust where it's wrong. Replace the shipped few-shot examples with ones from your own domain. This is iterative, not one-shot.
Anthropic published numbers from their harness runs:
| Setup | Time | Cost |
|---|---|---|
| Solo agent | ~20 min | ~$9 |
| Full harness (retro game maker) | ~6 hr | ~$200 |
| Full harness (browser DAW) | ~3 hr 50 min | ~$125 |
That's a 10-20x multiplier. The tradeoff: the full pipeline is worth it when a solo agent would produce output that's subtly broken in ways that cost you more time to debug than the pipeline costs to run. For straightforward tasks, just use the model directly.
Anthropic identified "context anxiety" — models wrapping up prematurely as the context window fills. Their V1 harness used context resets between phases. However, their V2 harness (Opus 4.6) dropped context resets entirely, finding the model sustains coherence natively.
The harness defaults to trusting the model's native coherence. Running the evaluator in a separate subagent already provides a natural context boundary at the most quality-sensitive phase. For very long runs where coherence visibly degrades, the harness can write state to .aletheia/handoff.md and resume from a fresh context — but this is a fallback, not the default.
The first real use of this harness was evaluating the harness itself. /aletheia review was run against the project's own SKILL.md, evaluator.md, and planner.md, with all five source articles fetched and analyzed as part of the review.
What happened:
-
The pipeline read all project files and fetched the source articles (Aletheia paper, FirstProof companion, Rajasekaran's blog post, "Building Effective Agents," "Effective harnesses for long-running agents").
-
A blind pre-analysis subagent — given only the research summaries, without seeing the harness code — independently reasoned about what an optimal implementation should look like. It identified 9 requirements, 8 common implementation mistakes, and 6 conflicts between the research programs.
-
The evaluator compared the actual implementation against the source articles and the blind analysis. Verdict: FIXABLE — 2 critical, 4 major, 2 minor issues.
Issues found:
| # | Severity | Issue | Source |
|---|---|---|---|
| 1 | Critical | No few-shot calibration examples in the evaluator prompt | Rajasekaran: "I calibrated the evaluator using few-shot examples with detailed score breakdowns" |
| 2 | Critical | WRONG verdict triggered revision — Aletheia says WRONG is terminal | Aletheia paper: WRONG means "cannot be salvaged without complete rewrite" |
| 3 | Major | Only blind pre-analysis ran in the subagent; Steps 2-4 ran in the main agent (breaking CoT decoupling) | Aletheia: verifier must be decoupled from generator's context |
| 4 | Major | Context management implemented V1's resets, but V2 dropped them | Rajasekaran: "I was able to drop context resets from this harness entirely" |
| 5 | Major | No abstention path — pipeline always delivers something | Aletheia: system outputs "No solution found" when it can't solve a problem |
| 6 | Major | Pre-analysis prompt asked about code structure, risking false negatives | Blind pre-analysis subagent identified this: "judge by behavior, not structure" |
| 7 | Minor | All artifacts in markdown; research found JSON more robust | "Effective harnesses": "model is less likely to inappropriately change JSON" |
| 8 | Minor | Evaluation criteria not shared with the generator phase | Rajasekaran: criteria alone improved first-pass quality before any evaluator feedback |
All 8 issues were fixed in a single revision pass. Additional improvements emerged from the review: the evaluator was migrated from a subagent to a teammate (for MCP access and true context isolation), review and architect modes were fleshed out, and the harness was optimized for Claude Code on subscription rather than API.
What this demonstrates:
- The blind pre-analysis caught issues (#2, #6) that might have been rationalized away if the evaluator had seen the code first. The subagent formed expectations from the research alone, then found the implementation didn't match.
- The evaluator correctly distinguished between fixable issues (wrong algorithm, missing feature) and things that were fine (overall architecture, synthesis approach, documentation quality). It didn't try to rewrite the whole thing.
- The pipeline's own principle — "separate generation from evaluation" — was validated by the fact that it found flaws in its own design that the original author missed.
The full evaluation output is preserved in .aletheia/evaluation.md in the repository.
Feng, T., Trinh, T.H., Bingham, G., et al. (2026). "Towards Autonomous Mathematics Research." arXiv:2602.10177. arxiv.org/abs/2602.10177
Feng, T., et al. (2026). "Aletheia tackles FirstProof autonomously." arXiv:2602.21201. arxiv.org/abs/2602.21201
Rajasekaran, P. (2026). "Harness design for long-running application development." Anthropic Engineering Blog. anthropic.com/engineering/harness-design-long-running-apps
Anthropic. (2025). "Building Effective Agents." anthropic.com/research/building-effective-agents. Describes the Evaluator-Optimizer workflow pattern, the direct ancestor of the evaluator loop used here.
Anthropic. (2025). "Effective harnesses for long-running agents." anthropic.com/engineering/effective-harnesses-for-long-running-agents
Google DeepMind. (2026). Aletheia prompts and outputs. github.com/google-deepmind/superhuman/tree/main/aletheia
MIT
This is an open synthesis of published research. If you find techniques from other papers that fit the framework, open a PR with the research reference and proposed integration.