Build your agent team in pi with one command.
You don't need to hire a dev team. You need to define one. Tamandua gives you a team of specialized AI agents — planner, developer, verifier, tester, reviewer — that work together in reliable, repeatable workflows. One install. Zero infrastructure.
- Install from GitHub
- Install from local checkout
- Quickstart
- What You Get: Bundled Workflows
- Why It Works
- How It Works
- Build Your Own
- Native AutoResearch
- Security
- Commands
- Requirements
- License · Origins
curl -fsSL https://raw.githubusercontent.com/igorhvr/tamandua/main/scripts/install.sh | bashOr just tell your agent: "Clone github.com/igorhvr/tamandua to my home dir, install it and learn the skill included inside it."
git clone https://github.com/igorhvr/tamandua.git
cd tamandua
./build-and-installOr step by step:
./build # npm install + tsc
./install # symlink into ~/.local/binThe build script handles everything: checks Node.js >= 22, runs npm install, compiles TypeScript. The install script creates a symlink at ~/.local/bin/tamandua pointed at your checkout — so you can keep the source wherever you like and tamandua stays in sync. Both call into scripts/install.sh internally.
That's it. Run tamandua workflow list to see available workflows.
Not on npm. Tamandua is installed from source (or GitHub), not the npm registry.
Requires Node.js >= 22. If
tamanduafails with anode:sqliteerror, make sure you're running real Node.js 22+, not Bun's node wrapper.
Sixty seconds from install to a running agent team:
$ tamandua workflow install feature-dev
# Or install all bundled workflows at once
$ tamandua workflow install --all
✓ Installed workflow: feature-dev
$ tamandua workflow run feature-dev "Add user authentication with OAuth"
Run: a1fdf573
Workflow: feature-dev
Status: running
$ tamandua workflow status "OAuth"
Run: a1fdf573
Workflow: feature-dev
Steps:
[done ] plan (planner)
[done ] setup (setup)
[running] implement (developer) Stories: 3/7 done
[pending] verify (verifier)
[pending] test (tester)Then watch your team work in real time:
$ tamandua dashboard # web UI at http://localhost:3334Tamandua ships with 23 bundled workflows organized into six families. Use tamandua workflow list to see available workflows, and tamandua workflow install <id> to install one.
Worktree variants (*-worktree, *-merge-worktree) run in a detached git worktree
created from your origin repository. Your main working copy stays untouched until the
workflow completes. This gives you full isolation — continue working while agents
iterate — and a clean abort path: delete the worktree and nothing in your origin repo
has changed. The origin repository only sees changes when a -merge variant squashes
the result back into the original branch.
When a merge workflow (-merge, -merge-worktree) fails at the finalize_merge
step and the base branch tip has moved since the run started, Tamandua automatically
launches a fresh replacement run with the same parameters. This "rugpull" detection
runs after the final merge failure — if the base branch stayed put, no replacement is
triggered. Pass --no-relaunch-upon-rugpull to workflow run to suppress the
automatic replacement.
Rugpull replacement runs are narrowly scoped: they only apply to
finalize_merge step failures in merge workflows (*-merge,
*-merge-worktree) where the base branch tip moved during the run.
All other failures — mid-pipeline step retry exhaustion, expects
validation exhaustion, worker death — permanently fail the run
UNLESS the workflow declares on_fail.retry_step, in which case
the run reroutes to the named upstream producer (bounded by
max_reroutes, default 2 before falling through to permanent
failure). No automatic replacement is triggered for these failures.
Use tamandua workflow resume <run-id> to reattempt a permanently
failed run; fix the underlying issue before resuming.
Story-based feature development. The planner decomposes your task into ordered user stories. Each story goes through implement → verify → test before the next one starts.
Show the 5 Feature Development variants
| Variant | Workflow ID | Agents | Pipeline |
|---|---|---|---|
| Local-only | feature-dev |
5 | plan → setup → implement → verify → test |
| + Merge | feature-dev-merge |
6 | plan → setup → implement → verify → test → finalize_merge |
| Worktree | feature-dev-worktree |
5 | plan → setup → implement → verify → test |
| Worktree + Merge | feature-dev-merge-worktree |
6 | plan → setup → implement → verify → test → finalize_merge |
| GitHub PR | feature-dev-github-pr |
6 | plan → setup → implement → verify → test → pr → review |
Local-only stops after testing — commits stay on the feature branch, no merge or
PR. + Merge variants add a finalize_merge step that squash-merges all commits
back into the original branch. Worktree variants run isolated in a detached worktree.
GitHub PR variants create a pull request and run a code review step.
Bug triage and fix. The triager reproduces the bug, the investigator finds the root cause, the fixer patches it, and the verifier confirms the fix against acceptance criteria.
Show the 5 Bug Fix variants
| Variant | Workflow ID | Agents | Pipeline |
|---|---|---|---|
| Local-only | bug-fix |
5 | triage → investigate → setup → fix → verify |
| + Merge | bug-fix-merge |
6 | triage → investigate → setup → fix → verify → finalize_merge |
| Worktree | bug-fix-worktree |
5 | triage → investigate → setup → fix → verify |
| Worktree + Merge | bug-fix-merge-worktree |
6 | triage → investigate → setup → fix → verify → finalize_merge |
| GitHub PR | bug-fix-github-pr |
6 | triage → investigate → setup → fix → verify → pr |
Vulnerability scanning and patching. Scans for vulnerabilities, ranks by severity, patches each one, re-audits after all fixes are applied, and runs regression tests.
Show the 5 Security Audit variants
| Variant | Workflow ID | Agents | Pipeline |
|---|---|---|---|
| Local-only | security-audit |
6 | scan → prioritize → setup → fix → verify → test |
| + Merge | security-audit-merge |
7 | scan → prioritize → setup → fix → verify → test → finalize_merge |
| Worktree | security-audit-worktree |
6 | scan → prioritize → setup → fix → verify → test |
| Worktree + Merge | security-audit-merge-worktree |
7 | scan → prioritize → setup → fix → verify → test → finalize_merge |
| GitHub PR | security-audit-github-pr |
7 | scan → prioritize → setup → fix → verify → test → pr |
Detect failing tests, disable them minimally, and iterate until the full test suite passes. Useful for establishing a clean baseline on a branch with known test failures.
Show the 3 Quarantine Broken Tests variants
| Variant | Workflow ID | Agents | Pipeline |
|---|---|---|---|
| Local-only | quarantine-broken-tests |
3 | setup → quarantine → verify |
| + Merge | quarantine-broken-tests-merge |
4 | setup → quarantine → verify → finalize_merge |
| Worktree + Merge | quarantine-broken-tests-merge-worktree |
4 | setup → quarantine → verify → finalize_merge |
Single-agent workflows for quick one-off tasks and workflow auto-selection.
| Workflow ID | Agents | Pipeline | Description |
|---|---|---|---|
do-now |
1 | execute | Submit any task. Get back a success/failure report. No planning, no stories. |
just-do-it |
1 | dispatch | Describe what you want. Dispatches to the most appropriate workflow automatically. For coding tasks (feature-dev*, bug-fix*, security-audit*) it defaults to merge-worktree variants unless the prompt gives a specific reason otherwise. |
do-review-do-verify |
3 | do → review → do-again → verify | Two-pass execution: do the work, review it, revise, then verify the result. |
Workflows for auditing and validating the project itself.
| Workflow ID | Agents | Pipeline | Description |
|---|---|---|---|
frontend-test |
1 | test | Builds the project and validates the dashboard frontend: HTML structure, route definitions, and test coverage. Does not start a second dashboard server. |
skills-normalize-audit |
3 | scan → audit → report | Scans a skills directory, analyzes the skills for overlaps and redundancies, and produces consolidation recommendations in a structured report. |
Install all bundled workflows at once with:
$ tamandua workflow install --all- Deterministic workflows — Same workflow, same steps, same order. Not "hopefully the agent remembers to test."
- Agents verify each other — The developer doesn't mark their own homework. A separate verifier checks every story against acceptance criteria.
- Fresh context, every step — Each agent gets a clean session. No context window bloat. No hallucinated state from 50 messages ago.
- Retry and reroute — Failed steps retry automatically, and can be rerouted to upstream producers for fresh context. When budgets exhaust, the run fails — terminally and automatically. Nothing fails silently.
- Zero tokens when idle — Checking for work is a database peek, not a model call; agents spawn only when a step is ready, and completion nudges make step-to-step latency near zero. The old polling motor measured roughly 30% token overhead; the new motor: zero.
- Define — Agents and steps in YAML. Each agent gets a persona, workspace, and strict acceptance criteria. No ambiguity about who does what.
- Install — One command provisions everything: agent workspaces, scheduling, subagent permissions. No Docker, no queues, no external services.
- Run — The scheduler checks for work deterministically (a DB peek — no model, no tokens) and spawns an agent only when a step is ready. Claim a step, do the work, pass context to the next agent. SQLite tracks state.
flowchart LR
CLI["tamandua CLI<br/>workflow run"] -->|create run| DB[("SQLite<br/>~/.tamandua/tamandua.db")]
CLI -->|register run| Daemon["Background daemon<br/>control plane"]
Daemon -->|dispatches work| Agents["Agent team<br/>planner · developer · verifier · tester"]
Agents -->|"pi --print"| Harness["pi harness<br/>(or Hermes, alpha)"]
Agents -->|claim step / write results| DB
DB --> Dashboard["Dashboard :3334<br/>Kanban + AutoResearch panels"]
DB --> MCP["Remote MCP :3338<br/>14 tools"]
The motor's invariants are pinned by an engineering contract with acceptance tests and real-model baselines: tests/MOTOR-CONTRACT.md.
YAML + SQLite + deterministic dispatch. That's it. No Redis, no Kafka, no container orchestrator. Tamandua is a TypeScript CLI with zero external dependencies. It runs wherever pi runs. Checking for work never invokes a model — idle runs cost zero tokens.
Tamandua ships with a content-addressed test-suite ledger that skips redundant test re-execution across workflow runs. When the same working tree with the same test command has already passed, TSTX replays the recorded result instead of re-running the suite.
How it works:
tamandua-testwraps every test command (via{{test_cmd}}) and computes a content-hash of the working tree usinggit write-treeon a temporary index — the repository's real index is never touched- On a cache hit (same tree + same command, green within 24h), the result is
replayed with exit 0 — no re-execution needed; a
TAMANDUA-TEST CACHEDbanner identifies the replay - On a cache miss or any doubt, the command runs normally — the shim passes through stdout/stderr and exit code unchanged
- TSTX hashes the tree again after the command exits and records a result only when the pre/post hashes match. If tracked or untracked-not-ignored content changes (or the post-run hash is unavailable), the result is not cached and a stable-tree rerun is required. A passing command fails closed with shim exit code 86; an already-failing command keeps its original nonzero code. The abandoned single-flight claim is released so a waiter can rerun promptly.
Safety: TSTX is strictly monotone — it may only skip work that is provably redundant (a green result for the byte-identical tree and command). On any doubt, error, or unexpected condition, it degrades to running the real command unchanged (passthrough). It is impossible for TSTX to make a task slower, wrong, or uncompletable compared to not having TSTX at all. Passthrough is byte-identical to the raw command except for a single stderr notice.
Kill switch: Set TAMANDUA_TSTX=0 to disable TSTX entirely — all test
commands bypass the ledger and execute directly.
Submodule caveat: git write-tree records submodule pointers (commits),
not their dirty working tree contents. If your repository uses submodules,
changes inside a submodule won't be reflected in the tree hash until they're
committed and the pointer is updated in the parent repository.
The bundled workflows are starting points. Define your own agents, steps, retry logic, and verification gates in plain YAML and Markdown. If you can write a prompt, you can build a workflow.
id: my-workflow
name: My Custom Workflow
agents:
- id: researcher
name: Researcher
workspace:
files:
AGENTS.md: agents/researcher/AGENTS.md
steps:
- id: research
agent: researcher
input: |
Research {{task}} and report findings.
Reply with STATUS: done and FINDINGS: ...
expects: "STATUS: done"Full guide: docs/creating-workflows.md
Tamandua includes native AutoResearch primitives for measurable optimization loops. Unlike a normal workflow, AutoResearch stores durable project-local state so an agent can resume after restarts, learn from each measured run, and choose the next experiment from evidence.
Use AutoResearch when the task has a reliable numeric metric and the agent should run a sequence of experiments instead of one batch of edits. Typical examples are raising test coverage, reducing validation loss, improving latency, or lowering cost while preserving correctness.
tamandua autoresearch init \
--goal "reduce validation loss" \
--metric val_bpb \
--direction lower \
--command "uv run train.py"
tamandua autoresearch run-experiment
tamandua autoresearch log-experiment --status auto \
--description "try lower learning rate" \
--hypothesis "smaller LR improves stability" \
--learned "validation improved but training slowed" \
--next-focus "test warmup schedule"
tamandua autoresearch next
# Inspect the loop for a Tamandua workflow run
tamandua workflow autoresearch <run-id>AutoResearch can be driven manually from any project directory, or delegated to a Tamandua workflow agent. In both cases the project needs a metric command that prints one parseable number. The command should be deterministic enough to compare experiments and should exclude generated or third-party code when measuring a project-owned objective.
Manual loop:
cd /path/to/project
tamandua autoresearch init \
--goal "Increase unit test coverage to 1.000 without changing application code" \
--metric coverage \
--unit ratio \
--direction higher \
--command "./measure-test-coverage.sh" \
--metric-regex "^([0-9]\\.[0-9]{3})$" \
--checks-command "./measure-test-coverage.sh"
tamandua autoresearch run-experiment
tamandua autoresearch log-experiment --status auto \
--description "baseline coverage" \
--hypothesis "establish current coverage" \
--learned "baseline recorded" \
--next-focus "cover the lowest-risk uncovered module"
tamandua autoresearch nextWorkflow-driven loop:
tamandua workflow install do-now
tamandua dashboard start
tamandua workflow run do-now \
"In the target repo, create or verify ./measure-test-coverage.sh, initialize tamandua autoresearch, then run 10 bounded experiments. Before each edit run tamandua autoresearch next. Only add or change tests/fixtures/test config. After each experiment run tamandua autoresearch run-experiment and tamandua autoresearch log-experiment --status auto with description, hypothesis, learned, and next-focus. Stop and report best metric, commits, and remaining gaps." \
--working-directory-for-harness /path/that/contains/or/is/the/project \
--pi-as-harnessMonitor it while the workflow runs:
tamandua workflow status <run-id>
tamandua workflow autoresearch <run-id>
open http://localhost:3334The dashboard's AutoResearch panel reads the run's harness working directory,
discovers the nearest autoresearch.config.json / autoresearch.jsonl, and
renders the experiment trace. Gray points are attempted experiments; green points
and the green line are the kept best-so-far frontier.
Tamandua maintains a SQLite registry of AutoResearch sessions so the dashboard
can discover them directly without scanning workflow runs. The registry lives in
a table called autoresearch_sessions inside the main Tamandua database
(~/.tamandua/tamandua.db).
- Project-local files are the source of truth.
autoresearch.config.json,autoresearch.jsonl,autoresearch.md, andautoresearch.shremain on disk in your project. The DB registry is an index/cache for discovery and dashboard UX — it never modifies your project files. - Sessions are registered automatically. Every
tamandua autoresearchcommand (init, run-experiment, log-experiment, status, next, loop) updates or creates the registry entry for that project directory. - Backfill on dashboard start. When the dashboard starts, it scans recent workflow runs for harness directories that contain AutoResearch files and backfills any missing registry entries.
Use tamandua autoresearch prune to clean up stale registry rows without
removing any project-local files.
# Prune sessions not updated in 30 days
tamandua autoresearch prune --older-than 30d
# Prune only sessions whose project files no longer exist
tamandua autoresearch prune --older-than 7d --missing
# Preview what would be pruned without deleting
tamandua autoresearch prune --older-than 30d --dry-runThe prune command only touches the SQLite registry — your autoresearch.jsonl,
config files, and experiment history remain untouched on disk.
For a test-coverage loop, a single experiment should be narrow enough to explain before editing and measurable enough to keep or discard after the run.
# 1. Ask the ratchet what evidence should drive the next edit.
tamandua autoresearch next
# Example returned focus:
# Best run 1: 0.336 ratio
# Next focus: cover pure helpers in batch_processor without touching application code
# 2. Make one focused test-only change.
# Example hypothesis:
# "Adding unit tests for batch_processor pure helper functions will increase
# coverage without requiring Spark or changing runtime code."
# 3. Measure and log the result.
tamandua autoresearch run-experiment
tamandua autoresearch log-experiment --status auto \
--description "cover batch_processor pure helpers" \
--hypothesis "pure-helper tests increase coverage without Spark" \
--learned "coverage increased from 0.336 to 0.477; helper paths are now covered" \
--next-focus "cover utils.py pure helpers and runtime stubs"If the metric improves in the configured direction and checks pass, the logged run
is kept. If it regresses, crashes, or fails checks, it is logged as discarded,
crash, or checks_failed; with --revert-discard, Tamandua can revert non-state
experiment files while preserving autoresearch.jsonl.
Project files:
| File | Purpose |
|---|---|
autoresearch.config.json |
Session config: goal, metric, direction, command, parser, checks. |
autoresearch.md |
Agent-facing objective and operating loop. |
autoresearch.jsonl |
Append-only run history: measured results, decisions, learning, next focus. |
autoresearch.sh |
Benchmark command. |
autoresearch.checks.sh |
Optional correctness checks run after successful measurements. |
When a workflow run was started with --working-directory-for-harness, the
dashboard includes an AutoResearch panel that reads that directory's
autoresearch.jsonl and shows best/baseline metrics, kept/discarded counts,
failures, and the recent learning timeline.
The core loop is init -> run-experiment -> log-experiment -> next. log --status auto classifies a
run as baseline, keep, discard, crash, metric_not_found, or checks_failed by comparing the
latest metric with prior accepted results (metric_not_found when the command exits 0 but the metric cannot be parsed from its output — such runs do not update best/baseline). The next prompt carries the ratchet:
it restates the goal, best result, last learning, and next focus before the agent
starts another experiment.
You're installing agent teams that run code on your machine. We take that seriously.
- Curated repo only — Tamandua only installs workflows from the official repository. No arbitrary remote sources.
- Reviewed for prompt injection — Every workflow is reviewed for prompt injection attacks and malicious agent files before merging.
- Community contributions welcome — Want to add a workflow? Submit a PR. All submissions go through careful security review before they ship.
- Transparent by default — Every workflow is plain YAML and Markdown. You can read exactly what each agent will do before you install it.
If something isn't working as expected, start with the built-in diagnostic:
- Run
tamandua doctor— One-shot diagnostic that checks environment (Node.js >= 22, pi on PATH, gh on PATH), services (dashboard, daemon, MCP), daemon staleness (running daemon matches installed build), database state (run-level anomalies), and LLM prompt adherence (per-step key-emission rates from workflow runs, measuring how often agents deliver expected output keys). Each check prints pass/fail status and on failure prints the exact remedy command to run. - Check service status — Run
tamandua statusto verify dashboard, daemon, and MCP are running on their expected ports. - Check logs — Run
tamandua logsto see recent daemon events. For live tailing:tamandua logs-tail. - Restart services — If the daemon (control plane + motor) is unresponsive, run
tamandua daemon restart. The dashboard UI can be restarted independently withtamandua dashboard restart(safe — never touches the motor). To pick up a locally rebuilt tree, run./build-and-installfollowed bytamandua restart— this restarts all services with a real stop→ready barrier instead of blind sleeps.
| Command | Description |
|---|---|
tamandua get-ready |
Install bundled workflows and start dashboard/control plane |
tamandua source-path |
Print the Tamandua source checkout path |
tamandua skill-path |
Print the path to the bundled tamandua-agents agent skill |
tamandua update [--force] |
Pull the source checkout, rebuild, reinstall workflows (refreshes all installed bundled workflow files — local edits are overwritten), and restart previously running services. For local development, use ./build-and-install && tamandua restart instead. |
tamandua uninstall [--force] |
Full teardown (agents, crons, DB) |
| Command | Description |
|---|---|
tamandua workflow run <id> <task> [--working-directory-for-harness <dir>] [--wait [--timeout <dur>] [--json]] [--pi-as-harness | --hermes-as-harness] |
Start a run (defaults harness CWD to your current directory). With --wait, block until the run finishes |
tamandua workflow status <query> |
Check run status |
tamandua workflow runs |
List all runs |
tamandua workflow wait <selector...> [--all] [--timeout <dur>] [--json] [--quiet] |
Block until selected runs reach terminal status |
tamandua workflow resume <run-id> |
Resume a failed or paused run |
tamandua workflow stop <run-id> |
Stop/cancel a running workflow |
tamandua workflow cancel <run-id> |
Alias for stop — cancels a running workflow |
tamandua workflow delete <run-id> [--force] |
Permanently delete a workflow run and associated data |
tamandua workflow list |
List available workflows |
tamandua workflow install <id> [--all] |
Install one or all workflows. Installed bundled definitions are refreshed on every install/update — local edits are overwritten. To customize a workflow, copy it under a new workflow id. |
tamandua workflow uninstall <id> |
Remove a single workflow |
| Command | Description |
|---|---|
tamandua merge-branch --origin <repo> --branch <branch> --into <target> --expect-tip <sha> --message <message> |
Atomically land a plumbing-based squash merge with managed checkout parking. A clean attached target is refreshed in place (refreshed); a dirty attached target remains safely on a backup branch (parked:<branch>); a coherent owned no-op reports already-coherent; and a bare or unowned target reports not-applicable. Multiple owners, invalid or ambiguous ownership metadata, and an owner operation in progress remain bounded refusals. See Atomic merge-branch landing. |
| Command | Description |
|---|---|
tamandua restart [--force] |
Restart all services (daemon, dashboard, MCP) with stop→ready barrier — no sleep guessing. The sanctioned way to pick up a locally rebuilt tree (./build-and-install first, then tamandua restart). |
tamandua dashboard start|stop|restart|status [--port N] |
Manage the standalone dashboard UI server (safe anytime) |
tamandua daemon start|stop|restart|status |
Manage the daemon (control plane + scheduling motor) |
tamandua mcp start|stop|restart|status [--port N] |
Manage the standalone MCP server |
tamandua control-plane start|stop|restart|status [--port N] |
Alias for daemon commands (control plane is hosted by daemon) |
| `tamandua logs [ | |
| `tamandua logs-tail [ | |
tamandua nudge |
Trigger an immediate dispatch round for all running runs |
When you start the management dashboard (tamandua dashboard), Tamandua automatically starts the remote MCP server too.
- Dashboard:
http://localhost:3334(or your custom--port) - MCP endpoint:
http://localhost:3338/mcp(fixed port)
Use tamandua dashboard status to verify both endpoints are up.
By default, the dashboard and MCP servers bind to 127.0.0.1 (localhost only), so they are not reachable from other machines on the network. If you need remote access, set TAMANDUA_BIND_HOST=0.0.0.0 (or a specific IP) before starting the dashboard:
TAMANDUA_BIND_HOST=0.0.0.0 tamandua dashboard --port 3334This environment variable applies to both the dashboard HTTP server and the MCP HTTP server. The control plane already binds independently to 127.0.0.1 and is not affected by this setting.
Each run also has a swim-lane view at http://localhost:3334/runs/<run-id>/kanban
(linked from the run-ID in the dashboard's runs table). Lanes are derived
dynamically from the workflow's steps: single steps render one card per lane,
loop steps (e.g. the developer agent iterating over user stories) render one
card per story. Cards are colour-coded by status (todo / running / done /
failed) and the page polls /api/runs/<run-id>/kanban every 3 seconds. The
JSON endpoint is also useful for external integrations — see
src/server/kanban-data.ts for the response shape.
By default, Tamandua uses pi (pi --print) as its agent harness. You can
override this with the harness selection flags on tamandua workflow run:
| Flag | Description |
|---|---|
--pi-as-harness |
Use pi as the agent harness. This is the default. |
--hermes-as-harness |
Use Hermes as the agent harness instead of pi. |
These flags are mutually exclusive — specifying both is an error.
⚠️ Alpha quality. Hermes harness support is in alpha and has known limitations: it is very slow compared to pi. Token usage is read from hermes' state.db after each round (best-effort: falls back to 0 tokens with a warning if the hermes schema is unavailable or changed). Use pi (--pi-as-harness) for production workflows.
Tamandua resolves the Hermes binary through a three-tier chain. The resolver never creates, deletes, replaces, chmods, or otherwise mutates any user executable or symlink — discovery is entirely side-effect-free.
Tier 1 — Explicit environment variable (always wins):
export TAMANDUA_HERMES_BINARY=/path/to/hermesSet TAMANDUA_HERMES_BINARY to an absolute or relative path. Relative paths
are resolved against the daemon's working directory at scheduling time. If the
path is not executable, the run fails immediately with a clear actionable
error.
Tier 2 — Current process PATH:
If TAMANDUA_HERMES_BINARY is not set, Tamandua searches the daemon's own
PATH for hermes. When noHurrySaveTokensMode is enabled, Tier 2 first
searches for hermes-token-saver (a token-saving wrapper) before falling back
to a bare hermes binary.
Tier 3 — Login-shell fallback (bounded):
If neither the env var nor the process PATH yields a working Hermes,
Tamandua spawns zsh -lic 'command -v hermes' so Hermes installed via
nix/homebrew/npm in shell-specific paths is discoverable even when not on the
daemon's PATH. The returned path is realpath-resolved and validated with
X_OK. This fallback is bounded and only runs when the first two tiers fail.
Every resolved binary path is guaranteed to be absolute. Relative
TAMANDUA_HERMES_BINARY values and relative/empty PATH entries are resolved
against the daemon process's current working directory at validation time. This
prevents ./hermes: not found errors when the dispatcher invokes the binary
from a different working directory.
When dispatching a Hermes agent session, the resolved binary's directory is
prepended to the child's PATH so nested Hermes invocations within the agent
session find the same binary, even when the daemon's own PATH lacked it
(e.g. login-shell-discovered Hermes). The original PATH is preserved as a
suffix so standard system tools remain reachable.
Tamandua's Hermes discovery is entirely side-effect-free: it never
creates, deletes, replaces, chmods, or otherwise mutates ~/.local/bin/hermes
or any other user executable or symlink. The old behavior of automatically
managing a ~/.local/bin/hermes symlink has been removed.
The harness validation runs at scheduling time — if no Hermes binary is found through any tier, the run fails immediately with a clear error.
./run-hermes-e2e-canary is an opt-in end-to-end canary that validates
the full Hermes pipeline against the real Hermes binary. It launches a single
trivial workflow run (--hermes-as-harness) through the daemon, scheduler, and
Hermes harness, then audits the token-attribution chain:
session_id trailer → state.db lookup → runs.tokens_spent > 0.
⚠️ Spends real tokens and is very slow (30+ minutes). The canary is never part of./run-all-e2e-testsornpm test. Run it manually after Hermes upgrades or when changing the harness adapter.
./run-hermes-e2e-canaryThe test silently skips with a clear message when no Hermes binary is
found on PATH or via TAMANDUA_HERMES_BINARY. A temporary isolated Tamandua
home is created for each run, but ~/.hermes is symlinked in so the real
Hermes binary can find its credentials and config.
tamandua doctor includes a Hermes state.db contract check in its
ENVIRONMENT group. When a Hermes binary is found, the doctor probes
$HERMES_HOME/state.db (read-only, no Hermes invocation, no tokens) and
verifies the sessions table contains all columns required for token
accounting: input_tokens, output_tokens, cache_read_tokens,
cache_write_tokens.
- Contract OK →
info: "hermes state.db contract OK — token accounting available" - Contract broken →
warn: "hermes state.db contract broken: . Hermes runs will report 0 tokens." - No Hermes binary → the check is omitted entirely.
This is a cheap schema probe that catches Hermes-side breakage (new
state.db format, renamed columns) before a production run silently reports
zero tokens.
The remote MCP endpoint exposes 14 tools:
| Tool | Description |
|---|---|
tamandua.runs.list |
List recent Tamandua workflow runs. Accepts optional limit (integer, 1–200, default 50). |
tamandua.run.status |
Fetch detailed status for a run. Requires query (run id, prefix, or task substring). |
tamandua.run.start |
Start a workflow run. Requires workflowId and taskTitle. |
tamandua.run.pause |
Pause a running workflow run. Requires runId. Optional drain (boolean) to wait for in-flight work before pausing. |
tamandua.run.resume |
Resume a paused workflow run. Requires runId. |
tamandua.run.delete |
Permanently delete a workflow run and associated steps, stories, and worktree metadata. Requires runId. Optional force (boolean) cancels and deletes running or paused runs. |
| Tool | Description |
|---|---|
tamandua.events.recent |
List recent global Tamandua events. Accepts optional limit (integer, 1–500, default 50). |
tamandua.source.path |
Return the local Tamandua source checkout path. No parameters. |
tamandua.skill.path |
Return the path to the bundled tamandua-agents agent skill. No parameters. |
tamandua.update.command |
Return local CLI guidance for updating Tamandua safely. No parameters. |
| Tool | Description |
|---|---|
tamandua.autoresearch.init |
Create project-local AutoResearch state. Requires cwd, goal, metricName, direction, and command. Optional metricUnit, metricRegex, checksCommand, and overwrite. |
tamandua.autoresearch.run_experiment |
Run the configured experiment command in cwd, parse the metric, run optional checks, and append a run_result. Optional command, metricRegex, checksCommand, and timeoutMs. |
tamandua.autoresearch.log_experiment |
Append the decision and learning for the latest run. Requires cwd and description; optional status, metric, hypothesis, learned, nextFocus, commit, and revertDiscard. |
tamandua.autoresearch.status |
Summarize baseline, best result, failures, and the next ratchet prompt for cwd. |
| Parameter | Required | Description |
|---|---|---|
workflowId |
Yes | Workflow id to run. |
taskTitle |
Yes | Task description for the workflow run. |
workingDirectoryForHarness |
For direct workflows | Harness working directory for remote MCP runs. Required for direct workflows, invalid for worktree workflows. |
worktreeOriginRepository |
For worktree workflows | Repository path to create the worktree from. Required for worktree workflows, invalid for direct workflows. |
worktreeOriginRef |
No | Git ref (branch, tag, SHA) for the worktree. Optional. Only valid for worktree workflows. |
noHurrySaveTokensMode |
No | When true, work spawns prefer a <harness>-token-saver wrapper from PATH over the plain harness binary (pi-token-saver for pi runs, hermes-token-saver for hermes runs; per invocation; falls back to the plain binary when absent). Idle dispatch is free either way. Optional, defaults to false. |
workingDirectoryForHarness and worktreeOriginRepository are mutually exclusive: direct workflows require the former, worktree workflows require the latter. Supplying the wrong one or both results in an invalid-params error.
- Node.js >= 22
- pi installed on the host
- Tamandua uses pi for AI agent execution. Agents run via
pi --printin non-interactive mode.
- Tamandua uses pi for AI agent execution. Agents run via
ghCLI for PR creation steps
Tamandua began as a fork of antfarm and pursues the same goal — orchestrating teams of AI agents through deterministic, repeatable workflows — but is built on top of pi instead of OpenClaw. Credit to the original authors for the design and inspiration.
Built with Tamanduás in mind.


