A research subagent for Claude Code. Drop the three files in .claude/agents/ and it works.
It investigates a question deeply across your local notes, the web, GitHub, Hacker News, X, and podcasts — and returns a structured, sourced brief. No raw search dumps. No marketing copy. Citations at the point of claim, with a verification pass that catches the citations that don't actually back what they're cited for.
I built this for my own Claude Code workflow and have used it on a few hundred research dispatches. It's adaptive — if you have a markdown knowledge base, it'll search that first; if you don't, it falls back to reading whatever local files you point it at. The agent itself is one file; the orchestration playbook for the main thread is another; the tool reference loaded on demand is a third.
git clone https://github.com/TheBitcoinBreakdown-95/claude-research-agent
cp claude-research-agent/.claude/agents/* <your-project>/.claude/agents/That's it. Open Claude Code in your project and ask it a research question. The agent fires automatically per its description ("use PROACTIVELY for any non-trivial research task").
If you want the agent to search your notes, point it at them by editing the tools: line in researcher.md (add your MCP search tool if you have one) and the User Context section (so the agent calibrates to your domain expertise — see Customize below).
The main thread picks the mode based on the shape of the question. You don't have to specify — the orchestrator playbook in RESEARCHER-ORCHESTRATION.md decides.
| Mode | When | What happens |
|---|---|---|
| Single dispatch (default) | Depth-shaped query — one topic, one rabbit hole | One researcher does the work end-to-end |
| Verify | Citation-dense or contested brief; or you explicitly ask | Re-dispatch the agent in verify-only mode against an existing brief; every citation gets walked and marked [confirmed] / [partially confirmed] / [unverified] / [contradicted] |
| Parallel fan-out | Breadth-shaped query — comparing K things, surveying a landscape, multiple independent sub-topics | K=3-5 researchers run in parallel, main thread synthesizes their briefs into one, citation verifier runs automatically on the unified brief |
The verifier is the load-bearing piece. Parallel mode auto-fires it because parallel errors compound — see The honest case study below.
Most research agents either assume zero local notes or hardcode a specific knowledge base tool. This one adapts to what you have:
- MCP search tool — if you have a markdown-search MCP (e.g. kb-mcp), the agent uses it first. Hybrid BM25 + vector search on your structured topic files.
- CLI search tool — if you have something like QMD on PATH, it'll use that. Hybrid search via shell.
- Filesystem fallback — always available.
Glob+Grep+Readon whatever local notes directory you've pointed the agent at. - Web sweep — only after the local pass.
If you have nothing local, the agent skips Tier 1 and starts at the web — but it'll surface a Setup Recommendation in the brief noting what you could install. The recommendation is one-line, not a sales pitch.
This means the agent works the same way whether you have a 1,500-source knowledge base or a notes/ folder with twelve files. The richer your second brain, the more of the agent's research happens locally.
The agent has one section that needs editing: User Context in researcher.md. It's a template explaining what to write — your domain expertise, topics you'll spot vague claims about, things you don't want pitched. Without this calibration, the agent will hedge on topics you know cold and over-explain things you understand fluently.
Example calibration:
The user is a senior backend engineer with deep Postgres experience.
Vague or hedged claims about Postgres, query planners, or replication will be called out — write with the precision the topic deserves. Cite the actual source code, the actual mailing list thread, the named committer.
The user is NOT looking for NoSQL alternatives.
The richer the User Context, the better the brief.
Full design journey, architectural decisions (D1-D4), comparative positioning vs. other agent setups, and named patterns are in DECISIONS.md. Headlines:
- D1 — Orchestration playbook lives in a sibling file, not in the agent prompt. The agent stays scannable; the orchestrator can evolve without redeploying.
- D2 — Synthesis happens in the main thread, not a synthesizer subagent. Keeps the user's original question grounded in synthesis.
- D3 — Parallel mode auto-fires citation verification. This is load-bearing — see the case study below.
- D4 — No cost guardrail for parallel mode. On a flat plan, K× dispatches cost wall-clock + rate-limit room, not dollars. Change this if you bill on tokens.
- Budget gating — Paid tools require an explicit
PAID TOOLS APPROVED:line in the dispatch prompt. Default is free-only research. - System prompts under ~200 lines — Tool reference docs live in
RESEARCHER-TOOLS.md, loaded on demand. The agent prompt itself stays scoped to behavior.
If you fork this agent, don't strip D3. The next section is why.
While building this I ran the Phase 2 smoke test on a breadth query about Bitcoin self-custody. The unified brief from K=4 parallel researchers looked clean — coherent narrative, sourced claims, no obvious gaps. The citation verifier (per D3) caught 5 material errors in it:
- A wrong Cash App Lightning support claim
- A wrong exit date for a major service (which nullified a downstream causation chain in the brief)
- A 5× overstatement of an insurance coverage figure
- A wrong product-rollout status
- A version-number error
I named the pattern parallel-error-non-cancellation: factual errors in independent sub-briefs do NOT average out in synthesis. They stack, because the synthesis step has no fact-checking machinery and the errors live in non-overlapping factual claims. Errors don't cancel because they aren't on the same claim.
The verifier is what makes parallel mode safe. Without it, K parallel briefs ship with compounding errors that read fluently. Most parallel-orchestration tutorials don't include verification; this one does because of this exact failure.
claude-researcher/
├── README.md ← you are here
├── DECISIONS.md ← design journey, D1-D4, comparative positioning, named patterns
├── LICENSE ← MIT
└── .claude/agents/
├── researcher.md ← the agent system prompt
├── RESEARCHER-ORCHESTRATION.md ← main-thread playbook (verify mode + parallel mode)
└── RESEARCHER-TOOLS.md ← tool invocation reference, loaded on demand
- Smoke-test transcripts. Verification artifacts, not demonstrations. The
parallel-error-non-cancellationpattern preserves the lesson. - The raw meta-audit brief. Summarized in DECISIONS.md. The raw brief had methodology-internal
[unverified]markers that would need extensive explanation. - Personal context. Stripped. User Context section is a customizable template.
- API keys, tokens, secrets. None in this repo. RESEARCHER-TOOLS.md uses generic env-var names and template paths.
Component patterns here (orchestrator-worker, citation verifier, parallel fan-out, tiered source hierarchy) are well documented in the public ecosystem. The contribution of this repo is the integration plus the operational specifics. See DECISIONS.md > Comparative Positioning for a comparison against:
- Anthropic's multi-agent research system (the canonical reference)
- wshobson/agents (185+ agents marketplace)
- VoltAgent/awesome-claude-code-subagents (100+ community agents)
- obra/superpowers (workflow library)
- karpathy/llm-council (peer-review patterns)
- Simon Willison's async research methodology
MIT. Use it, fork it, modify it. If you ship a meaningfully different orchestration design, I'd like to read it.