One CLI for all your robots. Connect them, command them, and let them work together, each with an LLM for a brain.
quackd, pronounced “quacked”, began as the brain daemon the Microduck was missing, named like that robot's own robotd, mediad, padd and tofd. That is where the ducks come from: a task is a .duck file and a group of robots is a flock. Seven robots today, every one of them in a simulator or a mock so far, and one of them a duck you can print and build yourself.
Register each robot once, by name. Then state a goal from a terminal or from a chat with Claude, to one robot or to a flock of them, and the same contract decides which of each robot's skills its model may use, how many steps it gets, and when it has to ask you first.
Hey, my name is Rok and this is why I built quackd 👋
I see quackd as a ChatGPT like moment for robotics. Let me explain what I mean.
LLMs existed long before ChatGPT. What ChatGPT actually did was take LLMs and hand them to ordinary people in a chat interface everyone already knew, like Facebook Messenger or Instagram. That was the real unlock.
Right now, in 2026, most people still think robots belong in science fiction movies or in a lab at Tesla. That is not true anymore. There are already open source robots you can build yourself for under $1000. And they actually work. They can go to your fridge, open it, grab a can of Coke, close the fridge and bring it to you.
The problem is they have a huge limitation. You can teach them dozens of moves, like "get a coke". But the robot itself is still dumb. It knows the moves, it just cannot connect them on its own. For robots to become truly useful, they need to become AI first and agentic. You give them a goal and they figure out the steps themselves.
To get there, robots need a brain. And here is the catch. Today's robots simply do not have enough hardware on board to think, reason and plan. Their skull is too small for the brain this kind of intelligence needs. So the brain has to live outside the robot, in the cloud or on your own computer, where it can grow as big as you need. The robot itself stays small and light while all the heavy thinking happens somewhere else. That is what quackd started as. A brain for one robot.
But here is what I think comes next. In the future everyone will have a flock of robots. At home, in the office, wherever. And they will not all be the same robot. Different types, different capabilities, even different companies. The first challenge is having all of them in ONE place to command. That is what quackd is now. You connect every robot you own to one CLI, and you command all of them from there, with an LLM as the brain of each one.
But even when you have them all in one place, that is still not enough. What you ask for will be complex, and robots will probably be very specialized. One can walk, one can grab, one can carry. For a bunch of different robots to be useful with as little of your involvement as possible, they need to start working together towards the goals you give them. Which means they need to communicate between each other. So that is the other half of quackd. You give the flock a goal, each robot's brain reads what the others can do, and they divide the work between themselves.
Imagine telling your robots "I want to eat and drink something". The one with wheels goes to the fridge, checks what is inside and tells the others what it found. The one with arms grabs a plate and some cutlery. One of them brings it all to you and asks what you would like, you choose, and they go back for the food and wish you a good meal. Sounds like science fiction, right? We are closer than you think!
That is basically what ChatGPT did for LLMs. It took something powerful and put it in one place everyone could reach. And that is why I see quackd as a ChatGPT like moment for robotics.
— Rok Benko, September 2026
The same sentence, the same duck, with and without quackd. Left: you type walk in a square and a pilot picks the robot's own verbs one at a time, correcting as it reads the pose it actually reached. Nobody wrote a square. Right: the identical world, robot and walking policy, minus quackd. A Microduck takes a twist, which is three numbers, so an English sentence has nowhere to go and it stands there. The pilot here is scripted, so this needs no API key. How it was made.
quackd is a command line for the robots you own. Each one joins through an adapter that declares, as a manifest, what the body is and what it can do, and you register it once by name with how to reach it. Give a robot a goal like "find the ball and kick it" and a large language model picks one skill at a time from the list that manifest declares, quackd runs it, looks at the camera, and asks again until the job is done or clearly impossible. Give the same goal to a flock and every robot in it gets a model of its own, and they divide the work by telling each other what they are going to do. A goal can arrive from a chat, a command line or a .duck task file, and whichever way it comes, quackd enforces a contract the model cannot talk its way out of: which skills are allowed, how many steps, when a human must say yes, when to abort. Claude, OpenAI, Gemini, Grok, Mistral, DeepSeek, Cohere, Qwen, Kimi, GLM and Meta work over their APIs. Open source models work on your own machine through Ollama, vLLM, llama.cpp or LM Studio, with no key.
The first robot is the Microduck from Pollen Robotics: a 25 cm, 800 g biped with fifteen small servos, a camera in its head, a depth sensor, a speaker and an onboard computer, open source and about $399, which already knows how to walk, turn, kick, scoop something off the floor, look around and quack at 50 Hz on its own hardware. It is the robot quackd started on, independently and unofficially. Six more bodies follow it through adapters that declare what each can do: an Open Duck Mini v2 you can print and build yourself, an SO-101 class arm through LeRobot, any wheeled base over rosbridge, an XLeRobot dual-arm cart, an AlohaMini with two arms on a lift, and a ToddlerBot humanoid.
You do not need a robot to try it. Two simulators ship with quackd. The physics one puts the real Microduck in MuJoCo and runs the walking policy Pollen trained for it, so the duck walks instead of sliding and a command below its gait floor produces nothing at all. The cartoon starts in a second, downloads nothing, and is what the three other bodies that have a simulator use, along with every seeded sweep in CI. These goals succeed on 10 of 10 seeds with the scripted pilot and a ground truth check:
"Find the ball and kick it." · "Find the ball, walk up to it and say where it is." (an Open Duck Mini v2, which cannot kick) · "Split the search, the closest duck kicks." (a flock)
The first of those also passes 10 of 10 on the physics simulator, with the duck on its own gait rather than a sprite on rails: that is test_find_and_kick_on_the_real_duck, which needs upstream's model in the cache, so a nightly job fetches it the way your first run would and CI's own gating job runs the stand-in. The rest are cartoon only, because the other six bodies have no physics model here.
Nothing here has run on a real robot yet, on any of the seven adapters, and no flock has yet crossed from one machine to a second. Every hardware backend speaks names read from upstream source at a pinned commit and has only ever talked to fakes. For the Open Duck Mini and the ToddlerBot those fakes are the daemons quackd itself ships for the robot, exercised over loopback, so there only the body is untested. Goals like "find my keys", handed to a flock that sorts out who does what, are where this is going, not what it does yet. The honest label for today is LLM driven, goal directed control of simulated robots, alone or in flocks, and Which robots work says exactly how far each one has got.
- Try it in 60 seconds
- Why?
- How it works (the simple version)
- Example
- Status
- Which robots work
- Architecture
- Installation
- Usage
- Connect any robot
- Flock mode
- The browser demo
- Configuration
- Performance
- Limitations
- Roadmap
- Contributing
- Safety
- Acknowledgements
- Star history
- License
uvx --from "quackd[mujoco]" quackd run --goal "walk in a square" --robot microduck:mujoco --provider fake # the duck above: real physics, its own trained gait (first run fetches about 10 MB)
uvx quackd run find-and-kick --provider fake # the cartoon: no download, done in a second
claude mcp add quackd -- uvx quackd serve-mcp --robot microduck:sim2d # or just chat with it: "find the ball and kick it"
uvx quackd run open-duck-scout --provider fake # a duck you can build: it finds the ball and walks up, no kick
uvx --from "quackd[mujoco,anthropic]" quackd run find-and-kick --provider anthropic --robot microduck:mujoco # a real model on the real gait, needs ANTHROPIC_API_KEY
uvx --from "quackd[openai]" quackd run find-and-kick --provider ollama --model qwen3:8b # local model, no key
open runs/*/run.gif # a GIF in either simulator, a transcript every timeOr open the browser demo and install no quackd at all. It is live at https://www.quackd.org/simulator, with the same physics, the same two upstream policies, seven of the same verbs and a contract of its own, in a page. Type a sentence, paste your own API key or point it at Ollama, and watch what the model chose. The keyboard beside the box is live at the same time, so a key can take the duck off the model mid-run. The browser demo says what it does and how to run it from a checkout.
Put keys in the environment or in a .env file (copy .env.example). quackd doctor tells you what is missing. Needs Python 3.11 or newer and uv, nothing else.
Every line above runs on the released package, which is 0.8. Naming robots and giving a flock one pilot each are the next release, so quackd robot, quackd flock and flock-hello want a checkout (uv sync) until then, and everything below that shows them is written for one.
A modern robot is not short of skills. The Microduck's onboard controllers already balance it, walk, kick, sit, stand up after a fall and scoop with its beak. An arm picks with its own learned policy, a wheeled base drives. Each is the robot's own skill, trained, written or recorded, and each works without any help from an AI model. What the robot lacks is any idea of what those skills are for.
Traditional control: walk forward, turn left, walk, look down, scoop, ... (you plan every step)
One robot: "Pick up the ball." (you state the goal)
A flock of them: "Pick up the ball." (they also settle who does it)
Low level skills and high level goals are different layers. The robot knows the words, but it cannot hold a conversation. quackd is an open source attempt to connect the two layers, with an LLM doing the planning and the robot's own controllers doing the moving.
The second gap opens with the second robot. Every body speaks its own protocol, the arm's SDK, the cart's ZeroMQ host, the base's ROS topics, and every body has a different set of skills, so the robots you own end up commanded from as many terminals as there are robots and none of them knows the others exist. quackd puts them behind one command line, under names you choose, and lets one goal go to several of them at once. Each robot's pilot is told what the others are and what they can do, so the work gets divided on data rather than on guesses.
flowchart TD
YOU["You<br/>“find the ball and kick it”"]
LLM["LLM<br/>looks at the camera, the robot's state and the last result<br/>picks ONE of the robot's own skills (a verb) and its parameters"]
Q["quackd<br/>checks the rules: allowed? task judged possible? budget left? needs confirmation?<br/>then runs the verb"]
R["Robot or simulator<br/>executes the skill with its own controllers<br/>(walking, grasping and looking are not the LLM's job)"]
O["quackd observes the result<br/>new camera frame, new state"]
YOU --> LLM --> Q --> R --> O
O -- "next verb, until done or impossible" --> LLM
The verbs the model can pick from are the robot's real, existing capabilities and nothing more. They come from its manifest, and a verb that is not in the manifest does not exist:
| Kind | Verbs | What they are |
|---|---|---|
| Core | observe report_state stop say move go_to search_scan approach_and |
on any robot whose manifest satisfies their requirements (a camera, a twist intent, a sound intent) |
| Microduck | sit stand stand_up kick grab gaze quack |
one each per behaviour the robot ships with, each an intent the robot's own controllers execute |
| LeRobot arm | move_joints gripper place pick |
an SO-101 class arm. pick is one skill intent the arm's own learned policy executes, confirm gated and present only when a policy is loaded |
| rosbridge base | introspect |
a wheeled base over ROS 2. It gets move, stop and report_state, plus observe, go_to, search_scan and approach_and once an image topic is configured. introspect asks the bridge what the body is: the topic list and the robot's own description, which is where its weight and its joint count come from |
| Open Duck Mini v2 | gaze express quack |
a 42 cm biped. No sit, no kick, no stand_up: its runtime has no such skill, so the verb does not exist rather than being refused |
| XLeRobot | move_joints gripper |
a dual-arm cart. The arm joints are a normalised -100..100 range, not degrees, and gripper takes a side because there are two of them |
| AlohaMini | lift move_joints gripper home_arms |
two arms on a motorised lift. The arm verbs refuse until quackd's own host wrapper is running on the robot, because upstream's leaves the arms limp |
| ToddlerBot | look stand perform grip |
a humanoid. look turns a two joint neck, stand slews to the safe pose and is not a way up from a fall, and perform plays only the keyframe motions the daemon actually loaded. grip appears on the gripper builds, which carry two more motors |
| Aliases | get_frame walk_to walk |
the 0.3 names of observe, go_to and move. They keep working in every .duck file |
| Learned | (none yet) | v2: policies trained from LLM written rewards, registered like any other verb |
go_to (still spelled walk_to in the older starter files) is a small closed loop in plain Python that steers toward whatever the camera sees, ten times a second, without asking the model. The LLM says "go to the ball" and never "turn 4° left". The same code steers a duck, a cart and a wheeled base, clamped to each manifest's speed limits. On a body with no locomotion, such as the arm, go_to does not exist at all. On the ToddlerBot, which looks with a two joint neck, search_scan sweeps the head instead of turning the body.
A flock is that same loop once per robot, all at once, with a bus between the pilots so one can tell another what it is about to do (Flock mode).
A find-and-kick run in the cartoon simulator, from its transcript (runs/<timestamp>-find-and-kick/transcript.jsonl). Every run opens the same way, with the pilot judging whether this body can do this task at all, because nothing that moves the duck runs until it has. This one is the scripted pilot, so model says so, the verdict says a rule has no judgement of a body, and usage is an estimate from character counts (no tokenizer). A real provider records the API's own counts and weighs the task against the datasheet in its prompt.
The same thing as a conversation, through MCP in Claude Code or Claude Desktop:
You: List the duck's verbs, then find the ball and kick it. Claude: (calls
robot_list_verbs,robot_observe,robot_assess_task("feasible"),robot_run_verb("search_scan"),robot_run_verb("go_to"),robot_run_verb("kick"),robot_say) Done. The ball moved about half a metre.
And the same shape with two robots instead of one, from the flock.jsonl of a flock-hello run: a simulated duck and a mock arm, one pilot each, on wall clock. Nothing here moves a joint on purpose, because this is the smallest honest test that two pilots in two different bodies can find each other and exchange a fact.
{"t": 0.102, "kind": "bus", "msg": {"src": "duck", "kind": "TALK", "to": null, "text": "duck here and ready; say hello back"}}
{"t": 0.106, "kind": "bus", "msg": {"src": "arm", "kind": "TALK", "to": null, "text": "arm here and ready; say hello back"}}
{"t": 0.123, "kind": "member_end", "duck": "duck", "outcome": "success", "reason": "said hello and heard back from arm", "llm_calls": 3}
{"t": 0.128, "kind": "member_end", "duck": "arm", "outcome": "success", "reason": "said hello and heard back from duck", "llm_calls": 3}
{"t": 0.128, "kind": "flock_end", "outcome": "success", "reason": "every member declared success: duck, arm", "messages": 2, "llm_calls": 6}tell is what put those two lines on the bus. It moves nothing, costs no step, and arrives in the addressee's next observation. Each member then declared for itself, and the flock succeeds only when all of them did. This one is the scripted pilot again, so the sentences are a rule's and not a model's.
Version 0.8, simulator and mocks. What has been built, and how far each piece has actually been exercised:
| Piece | Status |
|---|---|
sim2d bundled simulator (default) |
✅ 10 of 10 seeds on find-and-kick, GIF and transcript per run |
mujoco physics simulator (quackd[mujoco]) |
✅ 10 of 10 seeds on find-and-kick twice over: once on the kinematic stand-in and once with the duck walking on upstream's own trained policy, both ground truth checked, and both named tests rather than remembered numbers. The trained-gait sweep needs upstream's model in the cache, so a nightly job runs it and the gating job on every push runs the stand-in. The model and the policy are fetched from upstream at a pinned commit and hash checked, never shipped |
Browser demo (web/) |
🧪 the same physics, the same two upstream policies, seven of the same verbs and the same allowlist-and-budget machinery in a static page, with the sentence box and the keyboard live on one duck at the same time. Bring your own key, or point it at Ollama. CI checks everything that can be checked without a browser, which tests/test_web.py lists. The page has been booted in a browser twice and a held W walks the duck, but a full model-driven run, a barge-in out of one and the recording have never been watched. Live at https://www.quackd.org/simulator |
Manifests and core verbs (quackd list-adapters, quackd list-verbs --robot) |
✅ seven adapters, eight core verbs that appear only where the manifest meets their requirements, speed limits from the manifest, manifest.schema.json generated and drift tested |
MCP server (quackd serve-mcp) |
✅ Claude Code and Claude Desktop, one robot or a flock with --robots or --flock NAME (nine robot_* tools, tested in process against the simulator and the mocks), no Claude Desktop session on record |
Memory between runs (quackd memory, remember) |
✅ one JSONL file per adapter:backend, or per registered robot name, notes and run outcomes into the next prompt, tested end to end offline, 🧪 the remember tool itself exercised by one local model on one machine and by no cloud model (docs/memory.md) |
| Providers: eleven cloud vendors, fake | ✅ implemented, tested offline against stubbed SDK clients, with one hand curated catalogue of 115 model ids that --model is checked against before any call (quackd list-models), real model hero recording pending an API key |
| Local models (Ollama, vLLM, llama.cpp, LM Studio, any OpenAI compatible server) | ✅ implemented and tested against the OpenAI wire format, 🧪 two live runs by a contributor (Qwen 2.5 Coder 14B on LM Studio, seeds 5 and 6), never on this machine, transcripts in docs/assets/transcripts/, more welcome |
Registered robots and flocks (quackd robot, quackd flock) |
✅ both command groups over ~/.quackd/robots.json and ~/.quackd/flocks.json, so --robot NAME means the same thing in every command that takes a robot and --flock NAME in run and serve-mcp, tested offline, robot list --probe answered by the mocks, 🧪 never pointed at hardware (docs/registry.md) |
Pilot flocks, several robots on one task (--flock NAME) |
✅ one whole pilot per body, any backend, same or different bodies, 2 to 8, each with its own executor, allowlist, budget, heartbeat, memory and verdict, dividing the work with tell over the bus, 🧪 experimental, exercised on mock and sim2d bodies with the scripted pilot, by no real model and on no hardware |
Coordinator flock, one referee instead (--flock N) |
✅ deterministic auction and bus, one planner LLM call at most, ground truth checked in tests, 🧪 experimental and sim2d Microducks only. Its capability aware role auction (spotter/kicker) is unit tested but has no bundled two-robot demo today |
LAN discovery (quackd discover, quackd announce, quackd[lan]) |
✅ record format and both commands on fakes in the suite, 🧪 real zeroconf exercised once on one machine, never between two (docs/lan.md) |
MQTT flock bus (MqttBus, library only) |
✅ every message kind and a full flock run on a fake broker, 🧪 exercised once between two nodes through a local broker on one machine, never a flock across machines (no distributed clock yet) (docs/lan.md) |
| Learned verbs | 🗺️ v2, interface and docs only (docs/learned-verbs.md) |
Everything quackd assumes about each robot's API, and how sure we are: docs/adapter-status.md. quackd doctor prints the unverified ones for your machine.
Seven robots, and one table for how far each one has actually got. Each name links to that robot's own page. The distinction that matters is between code we have run and hardware we have not: no robot of any kind has run quackd, so the honest question is how much of the path to one is tested.
| How far it has got | What that means |
|---|---|
| ✅ simulator | Runs a whole task in the bundled 2D simulator, with a seeded acceptance sweep in CI that checks the simulator's ground truth, not the model's claim |
| ✅ physics | Runs a whole task in MuJoCo on the robot's own trained gait, checked against the physics world's ground truth rather than the model's claim. Needs quackd[mujoco]: CI installs it for the stand-in body on every push, and fetches upstream's model nightly for the trained gait |
| ✅ mock | Every verb runs offline against a scripted double, in the test suite |
| 🧪 daemon | The wire protocol runs end to end against the real on-robot daemon over loopback in CI. Everything except the robot is exercised |
| 🧪 names | Every upstream name read from upstream source at a pinned commit, exercised against fakes. Never connected to anything real |
| ⏳ stub | Refuses with a link, waiting for upstream to ship the thing it would talk to |
| Robot | --robot |
The body | How far it has got |
|---|---|---|---|
| Microduck | microduck:sim2d, mock |
a 25 cm biped from Pollen Robotics | ✅ simulator, ✅ mock |
microduck:mujoco |
the same robot in MuJoCo, on its own walking policy | ✅ physics. find-and-kick 10 of 10 seeds while it really walks (test_find_and_kick_on_the_real_duck), nightly, because the model is fetched rather than shipped (ADR-0030) |
|
microduck:jsonrpc |
the real one, over robotd |
🧪 names. Early pre-orders arrive around Christmas 2026, later orders in four to six months (checklist) | |
microduck:websocket |
upstream's planned agent gateway | ⏳ stub | |
| Open Duck Mini v2 | open_duck:sim2d, mock |
a 42 cm 3D printed biped you can build yourself | ✅ simulator, ✅ mock |
open_duck:bridge |
the real one, through a daemon quackd ships for its Raspberry Pi | 🧪 daemon. The nearest of these to a first real run, because the hardware is buildable today (checklist) | |
| LeRobot arm | lerobot:mock |
an SO-101 class desktop arm | ✅ mock |
lerobot:real |
the real one, through LeRobot | 🧪 names, behind quackd[lerobot], Python 3.12 or newer (checklist) |
|
| Any ROS base | rosbridge:mock |
any wheeled base that takes a Twist | ✅ mock |
rosbridge:ws |
the real one, over rosbridge_server |
🧪 names, behind quackd[rosbridge] |
|
| XLeRobot | xlerobot:mock |
a dual-arm mobile manipulator on an IKEA cart, about $660 to build | ✅ mock |
xlerobot:zmq |
the real one, over the ZeroMQ host it already ships | 🧪 names, behind quackd[xlerobot]. The whole wire format is exercised against a fake host over loopback (checklist) |
|
| AlohaMini | alohamini:mock, sim2d |
two arms on a lift, on a wheeled base | ✅ mock, ✅ simulator, alohamini-lookout 10 of 10 seeds |
alohamini:zmq |
the real one, over the ZeroMQ host it already ships | 🧪 names, behind quackd[alohamini]. The wire is exercised against a fake host over loopback. Its arms need quackd's own host on the robot, because upstream's leaves them limp (checklist) |
|
| ToddlerBot | toddlerbot:mock, sim2d |
a small open source humanoid you can build | ✅ mock, ✅ simulator, toddlerbot-lookout 10 of 10 seeds |
toddlerbot:bridge |
the real one, through a daemon quackd ships for it | 🧪 daemon: the protocol and the daemon's own safety machinery exercised against a fake body over loopback. It has no walk policy unless you stage one, and it cannot get up if it falls (checklist) |
open-duck-scout on open_duck:sim2d, seed 3, driven by the scripted pilot. It finds the ball and walks up to it, because this duck has no kick.
If you own one of these, the Open Duck Mini is where help is worth the most. It is a body a stranger can build from scratch, the daemon and the protocol are already exercised against each other, and the only untested part left is the duck. docs/open-duck-hardware-checklist.md is the order to try it in, feet off the ground until step 10.
Three loops, three rates, three owners. The LLM decides what, at 0.2 to 1 Hz. The steering loop decides how to get there, at 10 Hz, and never waits for the model. The robot's own controllers do the moving, at their own rate: balance on a biped, a pick policy on an arm, a gait policy on a humanoid, the driver on a wheeled base. The table with the rates and the owners is in docs/architecture.md.
The robot is an adapter that declares a manifest: what body it has, which intents and sensors, which verbs, what stops it. The registry, the tool list, the verbs a .duck may allow and the system prompt are built from that manifest when the robot connects. A verb that is not in it does not exist. All seven bodies go through the same loop, executor and contract (ADR-0017). A flock of pilots is that whole stack once per robot, side by side, with one bus between the pilots and one kill switch that reaches every executor.
flowchart LR
HUMAN["Human<br/>goal in human language"]
LLM["LLM<br/>Claude · OpenAI · Gemini · Grok · seven more cloud vendors · local (Ollama, vLLM, llama.cpp) · fake"]
subgraph quackd
LOOP["agent loop<br/>observe → think → enforce → act"]
EXEC["safety executor<br/>allowlist · confirm gates · budgets · abort rules · heartbeat"]
VERBS["verb registry<br/>built from the robot's manifest: core · the robot's own · aliases · learned (v2)"]
PERC["perception<br/>frame → detections → “ball at bearing 18° left, ~0.6 m”"]
ADAPTER["robot adapter<br/>microduck · lerobot · rosbridge · open_duck · xlerobot · alohamini · toddlerbot<br/>returns a manifest (embodiment, intents, sensors, verbs, limits, safety authority)<br/>sends intents, never motor writes<br/>backends: sim2d ✅ · mujoco ✅ · mock ✅ · jsonrpc, real, ws, zmq, bridge 🧪 never run on a robot · websocket ⏳"]
end
ROBOT["Robot<br/>its own controllers: robotd at 50 Hz on a Microduck, the position controller and pick policy on an arm, the driver on a base"]
SIM["simulators and mocks<br/>the cartoon world and the MuJoCo one, duck cam and head cam, offline doubles for every adapter"]
HUMAN --> LLM
LLM -- "exactly one tool call per turn" --> LOOP
LOOP --> EXEC --> VERBS --> ADAPTER
ADAPTER -- "intents: twist, skill, gaze, sound, joint, pose, gripper" --> ROBOT
ADAPTER --> SIM
ADAPTER -- "frame and state" --> PERC --> LOOP
LOOP -- "observation: text and image" --> LLM
Why predefined skills matter. The LLM never generates motor commands. Every verb is an intent the robot already understands: a velocity, a named skill (kick_left or ground_pick on the Microduck, pick as a LeRobot policy on the arm), a gaze target, a sound, a joint goal, a gripper command. The robot's own controllers do the physical part, on the Microduck policies trained in microduck_rl and exported to ONNX at 50 Hz, so a slow or confused model degrades the task, never the balance. Where a body has a deadman it stops itself when commands stall. The Microduck's robotd has one, on the Open Duck and the ToddlerBot the daemon quackd ships is the deadman, the XLeRobot's and the AlohaMini's hosts stop the wheels but not the arms, and on the arm and a rosbridge base quackd's heartbeat and stop are the only stop authority. The LLM names the skill, the body performs it.
Enforcement order. Every verb call passes Executor.run_verb, which applies the contract in a fixed order: abort flag, allowlist, the pilot's feasibility verdict, parameter validation, confirm gate, budgets, abort_when, preconditions, dry run, then execution with a timeout that races the abort, so a kill switch cancels the verb that is running. The preconditions are named by the manifest and supplied by the adapter, so a body's own rules are its own: not fallen on a duck, torque on and a cool servo for an arm, calibrated and not fallen on the humanoid. The full order and what each step means: docs/safety.md.
Prompts. The system prompt opens with the robot's own one line introduction from its manifest, then the contract in prose, what the robot remembers from earlier runs, and the .duck body verbatim. Tools are JSON schemas generated from each verb's parameter model, plus assess_task, declare_success and declare_failure, plus remember when memory is on and tell when the run is a flock of pilots, and the model must return exactly one tool call. Only the last two observations keep their images. Local models get one extra line describing the JSON shape to answer with when native tool calling is unavailable. The system prompt and the tools that are not verbs are in quackd/agent/prompts.py.
Perception: features, not frames. The default detector is an HSV colour threshold, about 1 ms per frame, no model download. Bearing comes from horizontal position through the camera's focal length. Distance comes from apparent size, so --fov-deg matters on a real camera. The simulator draws the ball in a known orange, so it works out of the box. For a real ball you tune one HSV range (FAQ). A YOLO detector is an optional extra.
Talking to the robots. Each adapter speaks its body's own protocol and spells every upstream name in one upstream_api.py, tagged VERIFIED (read from upstream source at a pinned commit) or UNVERIFIED, and a test proves the unverified ones are only reachable from the experimental backends. Two bodies are reached through an installed SDK (the arm through LeRobot, the base through roslibpy), two by speaking the ZeroMQ host they already ship because neither is an installable package (the XLeRobot and the AlohaMini), and two through a daemon quackd ships for the robot because neither runtime has a network API at all (the Open Duck Mini and the ToddlerBot). The Microduck's robotd speaks JSON RPC 2.0 over a unix socket, and quackd re-sends robot.move every 100 ms while walking on purpose, because the robot zeroes its velocity when those stop. Every name is tabulated in docs/adapter-status.md, each of the other six bodies has a page under docs/adapters/, and the traps that recur when you read a robot you cannot run are collected in docs/reading-robots.md.
Safety layer. Heartbeat failure, Ctrl+C and q all mean the same thing: stop, then abort. A verb that times out or raises stops the robot and comes back as a failed result, not an abort. --dry-run sends nothing. And stop always means stop, never collapse, on every body: quackd sends no robot's go limp call, ever. Session end is different on the arm, where LeRobot's own disconnect() releases torque by its default, so the arm can sag when a run ends. What actually stops each body when quackd goes quiet differs enough to be worth a table of its own, and each manifest declares its own answer in safety_authority: docs/safety.md.
The full map, with a "why it exists" line per module: docs/architecture.md. Decisions and their reasons: docs/adr/.
Requirements: Python 3.11 or newer and uv. Windows, macOS and Linux. No GPU. The default install is about 250 MB (OpenCV is most of it). Provider SDKs, robot SDKs and the LAN libraries are optional extras, so uvx stays fast and the default install never imports a robot SDK.
uvx quackd --version # nothing to install, uvx fetches it
uv pip install "quackd[anthropic]" # or: openai, gemini, grok, mistral, deepseek, cohere, qwen, kimi, glm, meta, all, yolo, live
uv pip install "quackd[lerobot]" # Python 3.12+; or: rosbridge, xlerobot, alohamini, lan. Never imported by default
git clone https://github.com/rokbenko/quackd && cd quackd && uv sync --extra dev # contributors# a goal in human language (bundled simulator, scripted pilot, no key needed)
uvx quackd run --goal "find the ball and kick it" --provider fake
# the same goal with Claude
uvx --from "quackd[anthropic]" quackd run --goal "find the ball and kick it" --provider anthropic
# a task file (thirteen ship with the package, the starter table below lists them)
uvx quackd run find-and-kick --provider fake --seed 3Every run writes runs/<timestamp>-<name>/ (--runs-dir replaces runs/) with transcript.jsonl (every prompt, tool call, gate, intent, result and token count, plus the robot's manifest in run_start), every frame quackd captured, summary.json, and run.gif on the simulator. quackd trace replays any of it afterwards.
Cloud or local, same command.
| Provider | Extra | Key | Run |
|---|---|---|---|
| Claude | quackd[anthropic] |
ANTHROPIC_API_KEY |
uvx --from "quackd[anthropic]" quackd run find-and-kick --provider anthropic |
| OpenAI | quackd[openai] |
OPENAI_API_KEY |
uvx --from "quackd[openai]" quackd run find-and-kick --provider openai |
| Gemini | quackd[gemini] |
GEMINI_API_KEY |
uvx --from "quackd[gemini]" quackd run find-and-kick --provider gemini |
| Grok | quackd[grok] |
XAI_API_KEY |
uvx --from "quackd[grok]" quackd run find-and-kick --provider grok |
| Mistral | quackd[mistral] |
MISTRAL_API_KEY |
uvx --from "quackd[mistral]" quackd run find-and-kick --provider mistral |
| DeepSeek | quackd[deepseek] |
DEEPSEEK_API_KEY |
uvx --from "quackd[deepseek]" quackd run find-and-kick --provider deepseek |
| Cohere | quackd[cohere] |
COHERE_API_KEY |
uvx --from "quackd[cohere]" quackd run find-and-kick --provider cohere |
| Qwen | quackd[qwen] |
DASHSCOPE_API_KEY |
uvx --from "quackd[qwen]" quackd run find-and-kick --provider qwen |
| Kimi | quackd[kimi] |
MOONSHOT_API_KEY |
uvx --from "quackd[kimi]" quackd run find-and-kick --provider kimi |
| GLM | quackd[glm] |
ZAI_API_KEY |
uvx --from "quackd[glm]" quackd run find-and-kick --provider glm |
| Meta | quackd[meta] |
META_API_KEY |
uvx --from "quackd[meta]" quackd run find-and-kick --provider meta |
| fake (scripted) | none | none | uvx quackd run find-and-kick --provider fake |
| Ollama (local) | quackd[openai] |
none | uvx --from "quackd[openai]" quackd run find-and-kick --provider ollama --model qwen3:8b |
| vLLM (local) | quackd[openai] |
none | uvx --from "quackd[openai]" quackd run find-and-kick --provider vllm --model Qwen/Qwen3-8B |
| llama.cpp (local) | quackd[openai] |
none | uvx --from "quackd[openai]" quackd run find-and-kick --provider llamacpp |
| LM Studio (local) | quackd[openai] |
none | uvx --from "quackd[openai]" quackd run find-and-kick --provider lmstudio |
| any OpenAI compatible server | quackd[openai] |
optional | uvx --from "quackd[openai]" quackd run find-and-kick --provider local --base-url http://host:8000/v1 |
Every row above runs the cartoon, which is the default robot. To put the same model on the physics simulator instead, ask for both extras and name the backend: uvx --from "quackd[mujoco,anthropic]" quackd run find-and-kick --provider anthropic --robot microduck:mujoco. The extras are independent, so quackd[anthropic] alone gives you the model and no physics. Nobody stands in the physics arena, so follow-me, whose whole task is to follow somebody, cannot succeed there and nothing stops you pointing it at that backend anyway.
A cloud model that takes an image sees the camera frame. Where a vendor does not document image input, quackd list-models marks that model no frames and quackd sends it the detections as text instead. Local models get the text detections by default and the frame too with --vision, which also overrides a no frames mark. --no-vision is the other direction, for a vision model you would rather send text to. The scripted pilot only reads the detection summary. Local setup, tool calling flags per server and what to expect from small models: docs/local-llms.md.
| Command | What it does |
|---|---|
quackd run <duck> or quackd run --goal "..." |
Run a task. --provider picks the model, --robot the body (a spec or a registered name), --flock runs several at once, --dry-run sends nothing, --live opens a window, --no-trace stops it narrating. quackd run --help has the rest, grouped. It exits 1 when a run does not succeed, and 3 when the pilot judged the task beyond this body and nothing moved |
quackd validate ducks/*.duck |
Check task files against the spec and a robot's manifest (--robot, a registered name or a spec, repeatable, --robots for a flock, or the file's own robots: if it has one). Exits 1 with field level errors such as requires kick, but arm-01 (lerobot-so101) does not provide it. --json prints one object per file and keeps the same exit code |
quackd serve-mcp |
Expose a robot (--robot <adapter>:<backend> or a registered name), or a flock of them with --robots name=<adapter>:<backend>,... or --flock NAME for a stored one, as MCP tools over stdio. --duckfile starts with a contract loaded on the default robot, --yes allows confirm gated verbs, --seed, --address, --dry-run, --no-memory, --memory-dir and --no-trace |
quackd doctor |
Keys, extras, adapters, local LLM servers, and every upstream assumption on this machine, ending in one line saying whether anything can run here (--robot for one robot's manifest, --address to ask a real robot what it is running, --json for a script). It exits 1 when nothing here can run, so a setup script can branch on it |
quackd list-verbs |
The vocabulary with parameters and safety classes (--robot for another robot, --json for a script) |
quackd list-adapters |
The robot adapters this build knows, their backends and status (--json for a script) |
quackd list-models |
Every model this build knows for every cloud vendor: the id --model takes, a label, one of five statuses (current, legacy, preview, specialised, open) and notes, which mark each vendor's default, the OpenAI models that need the Responses API, and the models quackd sends text detections to rather than a frame. --provider NAME prints one vendor. Local presets have no rows, because they take any id their server serves. --json for a script |
quackd discover |
The quackd robots answering on the LAN (zeroconf, needs quackd[lan]). --timeout seconds to listen, --json one object per robot. See docs/lan.md |
quackd announce --robot <adapter>:<backend> |
Advertise a robot's identity on the LAN (a static manifest, no robot connection). --name sets the manifest id, --for seconds to stay announced, default until Ctrl+C |
quackd memory show|add|clear |
What one robot remembers between runs: the notes a pilot saved and how recent runs ended. --robot picks the body by spec or by registered name, --raw prints the file, --memory-dir points elsewhere, clear --yes skips the prompt. See docs/memory.md |
quackd robot add|list|show|edit|remove |
The robots you have named: which body, where it is, its token and camera, and optionally the provider and model that pilot it. Kept in ~/.quackd/robots.json, so --robot NAME means the same thing in every command. list --probe connects to each and says whether it answered. --registry-dir points elsewhere. See docs/registry.md |
quackd flock create|list|show|edit|delete |
Named groups of registered robots, for --flock NAME on run and serve-mcp. create with no --robot prints what you have registered, numbered, and asks which to include. Kept in ~/.quackd/flocks.json. Not the flock: block of a task file, which says how the work is shared out. See docs/registry.md |
quackd record <duck> |
run pinned to microduck:sim2d (no --robot) that always writes a GIF. --seed defaults to 0 and gated verbs are auto accepted, as with --yes. --no-trace and --no-trace-prompt work here too |
quackd trace [run] |
Replay a finished run from its transcript, on stdout, as the same lines it printed while it ran. No argument means the newest run under --runs-dir, and a name, a timestamp prefix or a transcript file all work. --no-prompt, `--thinking all |
Reaching a real robot takes a spec, an address, a token and a camera URL, and retyping them on every command puts that token in your shell history. quackd robot add keeps them instead: which body, where it is, its token and camera, and optionally the provider and model that pilot it. A flock is a named list of those entries. Register once, and --robot NAME means the same thing in every command that takes a robot, --flock NAME in run and serve-mcp.
quackd robot add duck microduck:mock
quackd robot add arm lerobot:mock
quackd robot add scout open_duck:bridge --address tcp://10.0.0.5:9871 --token 8f2c... # a real one, when you have it
quackd robot list --probe # who is actually answering
quackd flock create pair --robot duck --robot arm
quackd run flock-hello --flock pair --provider fake # one pilot per member, talking
quackd serve-mcp --flock pair # the same flock behind one MCP server--probe connects to every registered robot at once, asks how it is, and closes again, which is the nearest thing here to looking across your robots:
robots (--robot NAME)
+---------------------------------------------------------------------------------+
| name | robot | address | flocks | reachable |
|-------+------------------+---------------------+--------+-----------------------|
| arm | lerobot:mock | | pair | + ok |
| duck | microduck:mock | | pair | + ok, battery 88% |
| scout | open_duck:bridge | tcp://10.0.0.5:9871 | | x timed out after 5 s |
+---------------------------------------------------------------------------------+
The two mocks answer because a mock always answers. scout is a real robot's address with nothing at it, which is what a robot that is switched off looks like, and it makes the command exit 1 so a script can branch on it. Both files live under ~/.quackd/, --registry-dir or QUACKD_REGISTRY_DIR moves them, and tokens are stored there in plain text. What a name changes, and what happens when a flock's member goes missing: docs/registry.md.
A task file is a contract plus instructions, deliberately shaped like a SKILL.md. The YAML frontmatter is enforced by quackd. The Markdown body is read by the model.
---
duck: 0
name: find-and-kick
description: Search the area for a ball, walk to it, kick it.
verbs:
allow: [search_scan, walk_to, kick, quack, get_frame, stop]
confirm: [] # verbs that ask a human y/N first
budgets: {max_steps: 40, max_minutes: 5, max_llm_calls: 40}
success:
- Ball displaced more than 0.3 m in sim, or human confirms the kick landed.
abort_when: [Battery below 15%, Same verb fails 3 times in a row]
persona: Determined and cheerful. Quack once when you succeed.
---
# Task
Find the ball and kick it.
## Strategy
1. `search_scan`. 2. `walk_to` the ball, stop ~0.25 m away. 3. `kick`. 4. Verify, and retry if it did not move.That is a duck: 0 file, the contract since 0.1, and every bundled v0 file still parses. A duck: 1 file can also say which body it is for and what it truly needs:
duck: 1
robots: microduck:sim2d # the default body, so `quackd run` needs no --robot (or one robot per flock member)
requires: [search_scan, walk_to, kick] # the honest minimum a body must providequackd validate --robot checks requires against a robot's manifest before anything moves: quackd validate find-and-kick --robot lerobot:mock exits 1 with requires kick, but arm-01 (lerobot-so101) does not provide it. For a duck: 0 file the whole allowlist counts as required. Of the thirteen bundled starters, the six written before 0.4 keep their 0.3 spellings at duck: 0 and the seven written since are duck: 1. A duck: 2 file can also correct the robot's datasheet for the build in front of you with a datasheet: block, and the prompt labels those numbers as coming from the task file.
| Starter | Goal | Notes |
|---|---|---|
hello-world |
quack, one step forward, quack | the smoke test |
find-and-kick |
find the ball and kick it | the flagship, ground truth checked in tests |
patrol-and-quack |
wander, quack twice on a person or pet | the scripted pilot quacks at the sighting but hits its budget on seeds 0 to 9, no pilot has completed it yet. Nobody is in the physics arena, so on microduck:mujoco it is a patrol with nobody to announce |
follow-me |
keep a person in view and follow at 0.5 m | cartoon only, nobody stands in the physics arena and this task is to follow somebody, so it cannot succeed on microduck:mujoco. The scripted pilot has no strategy for it and declares success after two steps without a single walk_to, no real model run yet |
fetch |
scoop the ball up and bring it back | experimental, the scoop is open loop and fails about 40 % of the time in sim, by design, and the scripted pilot has no strategy for it either, no real model run yet |
flock-kick |
multiple ducks split the search, the closest one kicks | flock mode, cooperation over a bus and an auction |
flock-hello |
a duck and an arm introduce themselves to each other | pilot flock, one LLM per body, talking over the bus |
open-duck-scout |
find the ball, walk up to it, say where it is | Open Duck Mini v2 (--robot open_duck:sim2d is its default), the kick free shape of find-and-kick, ground truth checked on 10 of 10 seeds |
open-duck-lookout |
stand still, look around, say what you can see | Open Duck Mini v2, and the task to point at a real duck first: nothing in its allowlist moves a leg, and it works on a duck with no head at all |
microduck-lookout |
stand still, look around, say what you can see | the same idea for a Microduck: nothing in its allowlist moves a leg, it copes with having no camera, and it stops and says so if posture reads unknown, which is the one thing worth knowing before letting the duck walk |
xlerobot-lookout |
stand still and report what is in front of you | an XLeRobot, and the task to point at a real cart first: nothing in its allowlist moves a wheel or an arm. This robot has no head control and no voice, so a human aims it and it reports in text |
alohamini-lookout |
stand still and report what is in front of you | an AlohaMini, and the task to point at a real robot first: nothing in its allowlist moves a wheel, an arm or the lift. Like the XLeRobot it has no head and no voice, so a human aims it and it reports in text |
toddlerbot-lookout |
stand still, look around with the head, and report what you can see | a ToddlerBot, and the task to point at a real humanoid first: nothing in its allowlist moves a leg, an arm or the waist. Put it on its safety stand before you try it |
lerobot-lookout |
move nothing, read the arm back, and report what it says about itself | a LeRobot SO-101 arm, and the task to point at a real arm first: nothing in its allowlist moves a joint, and it asks for report_state rather than observe because the real backend has no camera |
Full spec: docs/duck-spec.md. Add yours to ducks/.
claude mcp add quackd -- uvx quackd serve-mcp --robot microduck:sim2dThen, in Claude Code or Claude Desktop: "List the duck's verbs, then find the ball and kick it." Before the first verb that moves the body, the model has to say whether this body can do the task at all, judged against the datasheet in its robot_list row, and robot_run_verb refuses anything that moves until it has. Without a .duck loaded the session runs on a default budget of 40 verb steps and five minutes. Load one with robot_load_duckfile, or start with --duckfile, and its allowlist and budgets apply instead. Pass --robots duck=microduck:sim2d,arm=lerobot:mock to front several robots at once, or --flock NAME to front a stored flock, with one executor, budget and heartbeat per robot. Simulated robots here each get their own world (a shared arena over MCP is future work), and this session is one model driving each of them in turn, so for robots that pilot themselves and divide the work use flock mode instead (quackd run <duck> --flock NAME). Config for both clients, the nine robot_* tools, and a two minute script: docs/mcp.md.
A run does not start from nothing. Each robot has a small memory under ~/.quackd/memory/, keyed adapter:backend so the simulator and a real duck keep separate files, or by its name once you have registered it with quackd robot add, so two ducks of one kind do too. It holds: the notes the pilot saved with the remember tool, and one line per earlier run that quackd writes itself (outcome, reason, the last few verb results). The newest of both go into the system prompt at the next run, and remember costs no step. quackd memory show, add and clear manage it, and --no-memory runs fresh. Over MCP the same file sits behind robot_recall and robot_remember. In a flock the same rule applies, so two members of one kind share a file unless you have registered them, and what one needs another to know during a run it says with tell rather than writing it down. Details and what it is not (a learning loop, a search index): docs/memory.md.
Any robot with a way in can join: through its own SDK, through the protocol its host already speaks, or through a small daemon quackd puts on the robot when its runtime has no network API at all. Two of the seven are in that last category, and a third gets a host wrapper from quackd because the one upstream ships leaves its arms limp. A robot joins as an adapter that answers one question, what is this body and what can it do, as a manifest: its embodiment, the intents its controllers accept (a velocity, a named skill, a gaze, a sound, a joint goal, a pose, a gripper), its sensors, the limits its verbs clamp to, who stops it when quackd goes quiet, its verbs, and a datasheet of what it weighs, can carry and can reach, with a confidence and a source on every number. Everything else (the loop, the executor, the contract, the MCP server) is shared.
uvx quackd list-adapters # the seven adapters, for your build
uvx quackd list-verbs --robot open_duck:sim2d # a buildable duck's vocabulary
uvx quackd run open-duck-scout --provider fake # it finds the ball and walks up, 10 of 10 seeds
uvx quackd validate ducks/find-and-kick.duck --robot open_duck:mock # exit 1: requires kick, but open-duck-01 (open-duck-mini-v2) does not provide it
uvx quackd serve-mcp --robots duck=microduck:sim2d,arm=lerobot:mock # a duck and an arm behind one MCP serverThe backends that need a library sit behind extras (quackd[lerobot], quackd[rosbridge], quackd[xlerobot], quackd[alohamini]), import it only on connect, spell every upstream name in one pinned upstream_api.py, and never use a body's go limp call as stop. open_duck:bridge and toddlerbot:bridge need no extra at all, because the part that touches the robot runs on the robot: quackd ships a daemon for the duck's Raspberry Pi and another for the ToddlerBot, since neither runtime has a network control API to talk to. Adding a body of your own takes a manifest and a mock, about a day: docs/adapters.md, the fields in docs/manifest-spec.md, what has and has not run in docs/adapter-status.md.
Several robots can work one task together, and every member gets its own LLM pilot: its own executor, allowlist, budgets, heartbeat, memory and feasibility verdict, all at once, on wall clock, on any backend. Same bodies or different ones, 2 to 8 of them.
uvx quackd run flock-hello --provider fake # the bundled demo: a duck and an arm, no key and no registry
uvx quackd run <duck> --flock pair --provider anthropic # your own robots, from the registryThey divide the work with a tell tool that reaches the addressee in its next observation, and each pilot's prompt carries every peer's datasheet, so the work is divided on data rather than on guesses. Each one declares for itself and the flock succeeds only when all of them did, and one Ctrl+C reaches every executor.
Nothing about a member is a special case, so what a flock can do is what one pilot can do times the number of bodies. Each is handed the part of the contract its own body can answer for, so an arm in a walking flock is not turned away at the door for having no legs, while what the task requires is still checked against every body together before anything connects. Say which bodies with a stored flock, or with flock.members and robots: in a duck: 1 file and no registry at all, which is what flock-hello does.
The honest part: N simulated members are N separate worlds with no shared arena and no ground truth to check a claim against, it costs N budgets and N times the tokens, a seed does not make it reproducible, nothing here has run on hardware, and tell has been exercised by the scripted pilot and by no real model. Details: docs/flock.md.
The other kind of flock has one referee instead of N pilots, and it is the one that ships a choreography: the flock splits the search for a ball, holds a quick auction, and the closest duck takes the shot.
uvx quackd run flock-kick --provider fake --seed 3
The first choreography: one flock, one auction, one kicker. Scripted planner, deterministic coordinator. Every message is in the transcript.
The interesting part is not the kick, it is the talking. The ducks coordinate over an in process bus with nine message kinds (TASK, BID, CLAIM, ROLE, HINT, VERDICT, HB, RESULT and TALK), the same bus tell uses, every one logged in flock.jsonl, and a deterministic Contract Net auction decides which duck acts, from each duck's own camera distance estimate. Every action goes through verbs the duck already has, so the machinery is task agnostic and what a flock can do is bounded by its skills, not by the ball. The LLM contributes at most one planning call per run, and each duck still enforces the .duck contract on itself. The outcome is judged from sim ground truth, not from a model's claim. Add a flock: block to any .duck or pass --flock N (2 to 4 ducks). This kind is simulator only (every member must be a sim2d Microduck) and its per duck pilots are deterministic rules, on purpose. Details: docs/flock.md.
A coordinator flock can also assign roles by capability, not just split the same search. A duck: 1 file may declare flock.roles (spotter, kicker), and bids carry a capability term, so each robot bids only for a role its manifest can fill. A duck: 2 role can ask for a body rather than a vocabulary, enough payload or a gripper rather than a beak, matched against each robot's datasheet. This machinery is unit tested but has no bundled multi-role starter today: see docs/flock.md.
web/ is the same duck with no quackd to install: a static page that loads MuJoCo compiled
to WebAssembly, the Microduck's own model and two of the policies Pollen trained, and re-implements
this repository's loop, verbs and executor in plain JavaScript modules with no build step.
You bring the key, or point it at Ollama and bring none. About 45 MB arrives the first time, the
libraries from jsDelivr and the model and the policies from upstream's own repositories at the same
pins Python uses, and the browser caches it afterwards. Nothing upstream is vendored here.
It is live at https://www.quackd.org/simulator, served there by quackd-web, a separate
repository, whose build fetches this directory at a pinned commit, so a change here reaches the
page on that project's next build. /simulator/source.json records which commit the deployed copy
came from. To run the same page from a checkout instead:
python web/serve.py # stdlib only, no dependencies, no build step
# then open http://localhost:8000/simulator/A plain python -m http.server --directory web will not do, because the page is mounted at
/simulator rather than at a root: web/README.md says why that mount is not
a style choice.
Both ways of driving are live at once, and that is the argument the page is making. The
sentence box and the keyboard hold the same duck at the same time, with no mode to flip. A key
that would move the robot takes it mid-run: the run aborts, the request in flight to the model
aborts with it, and the transcript names the key that did it. A key that only reads never barges
in. There is deliberately no key for say, because a key carries a command and a sentence
needs something to read it, and the switch marked quackd is on removes that reading layer and
only that layer: the physics and the policies are identical either way.
The key map, the API key handling, and every way this differs from the Python backend (the arena,
perception, the missing hash check, what a seed means, the absent scripted pilot) are one list in
web/README.md. That list is the canonical one, and this section deliberately
does not keep a second copy of it.
It has been booted, and that is all. The page opened clean in a browser twice while it was
being built, and a held W walked the duck. Nobody has yet watched a full model-driven run, a
barge-in out of one, or the recording, on that machine or any other. tests/test_web.py holds
what can be checked without a browser and runs in the ordinary suite, which is a floor and not a
browser test.
| What | How |
|---|---|
| API keys | ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, XAI_API_KEY, MISTRAL_API_KEY, DEEPSEEK_API_KEY, COHERE_API_KEY (or CO_API_KEY), DASHSCOPE_API_KEY, MOONSHOT_API_KEY, ZAI_API_KEY, META_API_KEY (or MODEL_API_KEY) in the environment or a .env file (see .env.example) |
| Model | --model or QUACKD_MODEL, an id from the catalogue. An id a cloud vendor does not list is refused before any call, and the refusal prints the ids that vendor does take. quackd list-models prints them all. The defaults are claude-opus-5, gpt-5.6-sol, gemini-3.8-flash, grok-4.6, mistral-medium-3-5, deepseek-flash, command-a-plus-05-2026, qwen3.8-max, kimi-k3, glm-5.3 and muse-spark-1.3 |
| Claude reasoning effort | QUACKD_EFFORT (low to max, default medium). QUACKD_ANTHROPIC_FALLBACKS=0 disables server side refusal fallbacks. QUACKD_THINKING_DISPLAY=omitted stops Claude returning a summary of its reasoning, and QUACKD_GEMINI_THOUGHTS=0 does the same for Gemini |
| OpenAI API and effort | QUACKD_OPENAI_API=responses opens on the Responses API instead of Chat Completions, and QUACKD_OPENAI_REASONING_EFFORT sets the effort on either. Neither is usually needed: quackd already knows which models want Responses, and moves a run there by itself when one says so (FAQ) |
| Local models | --provider ollama, vllm, llamacpp, lmstudio or local --base-url http://host:port/v1. No key. --model takes any id the server serves, and without it quackd uses the first model the server lists. The catalogue is for cloud vendors only, so nothing here is refused for being unlisted. --vision sends frames. QUACKD_TOOL_CHOICE=auto, required or none for picky servers. --extra-body or QUACKD_EXTRA_BODY merges a JSON object into every request body, which is how Qwen3 is told not to think on vLLM, and it works on every vendor that speaks OpenAI's API. See docs/local-llms.md |
| Robot | --robot <adapter>:<backend> or a name from quackd robot add, or a robots: line in the .duck, the flag wins. Default microduck:sim2d. quackd list-adapters lists the seven that ship, quackd list-verbs --robot X what each can do |
| Physics simulator | --robot microduck:mujoco, with quackd[mujoco]. The model and the policies are fetched once into ~/.quackd/cache, where QUACKD_CACHE_DIR moves them and QUACKD_MICRODUCK_ASSETS points at your own microduck_rl checkout instead. QUACKD_MUJOCO_BODY=puppet runs the kinematic stand-in, which downloads nothing and is the body the tests build. --live opens MuJoCo's own viewer |
| Determinism | --seed N makes a simulator run repeatable |
| Budgets | in the .duck. --max-steps overrides for one run |
| Human in the loop | verbs.confirm in the .duck prompts y/N, and so does a pilot that answers uncertain when it judges the task, where a no ends the run. --yes auto accepts both, and quackd record always passes it. MCP refuses gated verbs unless started with --yes |
| Dry run | --dry-run sends nothing, and the trace shows every verb it would have run, with its parameters |
| Trace | on by default, on stderr: the prompt, what the model thought and chose, every executor decision, every intent sent to the robot, every result, tokens and timings, a rule per step and a glyph per outcome. --no-trace or QUACKD_TRACE=0 turns it off and leaves the one line status that says what the run is waiting on, --no-trace-prompt or QUACKD_TRACE_PROMPT=0 drops just the system prompt, QUACKD_TRACE_THINKING caps the reasoning shown per turn (default 2000 characters, all for everything). The transcript keeps all of it either way. See docs/architecture.md |
| Colour | on when the output is a terminal. quackd --no-color <command> or NO_COLOR=1 turns it off, FORCE_COLOR=1 keeps it in a pipe. --no-color belongs to quackd itself, so it goes before the command rather than after it. Glyphs fall back to ASCII on a codepage that cannot carry them |
| Machine readable | --json on validate, list-verbs, list-adapters, list-models, doctor and discover. One object per line on stdout, nothing else, and the exit code is unchanged |
| Memory | on by default, under ~/.quackd/memory/. --no-memory runs fresh, --memory-dir or QUACKD_MEMORY_DIR moves it |
| Registered robots | ~/.quackd/robots.json and ~/.quackd/flocks.json, written by quackd robot and quackd flock. --registry-dir or QUACKD_REGISTRY_DIR moves both. Tokens are stored in plain text there. See docs/registry.md |
Real robots. Each needs --robot and --address, and the extra named. None has been run against its target by us, so all seven are 🧪 (docs/adapter-status.md).
| Body | --robot ... --address ... |
Needs |
|---|---|---|
| Microduck | microduck:jsonrpc --address unix:///run/robotd.sock on the robot, or tcp://127.0.0.1:9870 after ssh -L 9870:/run/robotd.sock <robot>. For a picture, --camera-url webrtc://<robot>:8443, because robotd serves no frames |
quackd[microduck-camera] for the camera |
| Open Duck Mini v2 | open_duck:bridge --address tcp://open-duck.local:9871 --camera-url http://open-duck.local:9872/snapshot.jpg --token <the bridge token> |
nothing: the daemon runs on the robot |
| LeRobot arm | lerobot:real --address /dev/ttyACM0 (the arm's serial port, COM5 on Windows) |
quackd[lerobot], Python 3.12 or newer |
| Any ROS base | rosbridge:ws --address "ws://robot.local:9090?cmd_vel=/cmd_vel&odom=/odom&image=/camera/image/compressed" |
quackd[rosbridge] |
| XLeRobot | xlerobot:zmq --address tcp://xlerobot.local:5555 (add ?variant=diff2 or ?variant=mecanum for a base other than the default three-omniwheel one, or ?swap_colour=0, if you need them) |
quackd[xlerobot] |
| AlohaMini | alohamini:zmq --address tcp://alohamini.local:5555 |
quackd[alohamini] |
| ToddlerBot | toddlerbot:bridge --address tcp://toddlerbot.local:9873 --token <the daemon token> |
nothing: the daemon runs on the robot |
On the simulator with the scripted pilot, find-and-kick takes 3 to 8 verb steps, one model call each plus one to declare success, and under a second of loop wall clock per run on a laptop. Interpreter start and GIF rendering add a few seconds to the whole command, and simulated time runs as fast as the CPU allows. With a real model each decision is one API call: the system prompt and the tool schemas are about 7 k characters (roughly 2 k tokens) with memory on, each observation a few hundred characters plus a 256 px PNG for vision models, and the transcript records each provider's own usage per turn. Model latency never affects control, because the steering loop runs at 10 Hz and the robot's own controllers run regardless of how long the model thinks. That holds for local models too. The default install is about 250 MB, needs no GPU, and the simulator renders at 256 px (--gif-size for prettier GIFs).
The physics simulator costs what physics costs. Measured here on one Windows laptop with an integrated GPU, walk in a circle on microduck:mujoco took about 8 seconds of wall clock without a GIF and 15 with one, against under a second of loop time in the cartoon, and the first run downloads about 10 MB of model and policy into ~/.quackd/cache and leaves 23 MB on disk. Rendering is the cost rather than physics, which steps at roughly 24 times real time, so shadows are off unless QUACKD_MUJOCO_SHADOWS=1 asks for them and the recorder samples half as often as the cartoon's. The arena is upstream's own scene: the blue checker floor, the gradient sky and the lighting come from the scene*.xml wrappers in microduck_rl, so a duck here stands where a duck there stands. The head camera is the exception, and ADR-0030 says why.
- The default simulator is a cartoon on purpose. It tests the agent loop, not physics, and will not tell you whether a gait works.
microduck:mujocois the one that can, and only for the Microduck. - The physics simulator runs upstream's walking and standing policies and nothing else of theirs.
kickandgrabuse the cartoon's contact rules,sitis refused, and a fall is recovered by standing the model up, because upstream's episodic policies did nothing from a standing pose when they were tried. The gait floor, no step below about 0.22 m/s or 1.0 rad/s and roughly 0.42 of what is asked above it, was measured here on one machine with the model's own actuators and is tagged UNVERIFIED, because upstream deploys a different actuator model. All six are listed instate.extras.assumptions, so a transcript never implies more than happened. - Nothing has run on a real robot of any kind. What each body cannot report or detect on hardware (posture inferred from a policy name on the Microduck,
holdinginferred from the gripper stopping short on the arm, no verified deadman on a rosbridge base, no fall detection and no battery on an Open Duck) is spelled out in docs/adapter-status.md and the adapter pages. - The datasheets were read from the makers' pages, repositories and one paper on 2026-09-13. Nothing was measured here, which is what the confidence label on every number is for, and a body whose maker never published a figure says so rather than having one invented for it.
- Whether a task fits a body is the model's own judgement, recorded before anything moves and weighed against numbers that carry their own confidence. Every gate below it still applies: the allowlist, the budgets, the confirm gates and the robot's own safety authority.
- The hero GIF is the scripted pilot, not an LLM, because this repository was built without an API key. The real model code paths are tested against stubbed SDK clients.
- Success is the model's own claim (
declare_success) on a solo run. In the simulator, tests also check ground truth, and a coordinator flock's success needs a member's kick report (or the spotter's verdict) and sim ground truth to agree. A pilot flock has no shared world to ask, so its success is every member's own claim and nothing vetoes it. On hardware, the.duckbodies insist on verifying with a fresh frame. - Memory between runs is a file, not a memory system: no embedding, no search, no sharing between bodies, and nothing the executor ever trusts. The scripted pilot never writes a note, so with
--provider fakeonly run outcomes accumulate. Notes have been exercised by one local model on one machine, and by no cloud model at all (docs/memory.md). - No robot here has text to speech. The Microduck has seven duck sounds, so
quack("hello")andsaypick a tone. The arm, the base, the XLeRobot, the AlohaMini and the ToddlerBot do not getsayat all. grabis open loop upstream and unreliable here on purpose.fetchsays so in its file.- A manifest can be smaller than the robot. The LeRobot arm's
realbackend claims no camera and nopickuntil it connects, and even thenpickappears only when a policy object was injected in code. A rosbridge base overwshas no camera verbs unless the address names an image topic, and a ToddlerBot has nomoveunless a walk checkpoint is staged. - The model catalogue is hand curated. It was read off the eleven vendors' own documentation on 2026-09-12 and it is a snapshot of that day, not a live list. A vendor can retire, rename or add an id between quackd releases, and this build would then refuse an id that is real and offer one that is gone.
quackd list-modelsprints exactly what this build knows, which is the only thing--modelaccepts for a cloud vendor. - Local model quality is unmeasured. The JSON text fallback and the one retry exist because small models often miss native tool calls. One contributor ran
find-and-kickagainst Qwen 2.5 Coder 14B through LM Studio on two seeds, both successes, one of them reading an earlier run's memory. The two transcripts are indocs/assets/transcripts/and read in docs/local-llms.md. - The pilot flock takes any body and has run on
mockandsim2dones only: N simulated members are N separate worlds with no shared arena, it costs one budget and one model call per member per turn, andtellhas been exercised by the scripted pilot and by no real model. The coordinator flock is simulator only, ships one choreography and exactly two roles (spotter and kicker, unit tested but with no bundled multi-role starter), and knows only the Microduck. Separation uses sim ground truth, and two robots share no frame of reference on hardware. - LAN discovery and the MQTT bus have each been exercised once, on one machine. Nothing has crossed to a second machine, the MQTT bus is a library with no
--busflag, and a coordinator flock across machines also needs a clock across machines, which does not exist yet. A pilot flock needs no such clock and has simply never been tried across two.
Why a task can refuse a body, whether two robots can share a task, and more: docs/faq.md.
Non goals for now, on purpose: no RL training or reward generation (that is v2, and only the registry hook exists), no features that require hardware, and no vendoring of Pollen Robotics assets. No logo, mesh, policy or sound of theirs is committed here. The physics simulator and the browser demo fetch the model and the policies from upstream at run time, and the one exception in this repository is the hero recording, which renders that model and carries its CC BY-NC-SA terms (docs/licenses.md).
- Hardware: the Open Duck Mini v2 is the nearest first real run (its checklist). An SO-101 arm, a rosbridge base, an XLeRobot, an AlohaMini and a ToddlerBot also exist today, so their backends can flip from 🧪 to ✅ with one real run each.
microduck:jsonrpcwaits for a Microduck to arrive and thewebsocketstub waits for upstream to ship its WebSocket surface. Open an issue withquackd doctoroutput and the first lines oftranscript.jsonl. - Flocks next: a pilot flock against a real model rather than the scripted one, then a pilot flock across two machines over the MQTT bus (docs/lan.md), which needs a
--busflag and a run that proves it rather than a clock, and then hardware flocks. For the coordinator: more choreographies from the verbs the robots already have (a patrol that splits the area, a follow chain), a clock that crosses machines so that bus can carry one across a room instead of a process, and a second body so Open Ducks can join one. - More bodies: whichever robots people own. An adapter is a manifest and a mock, about a day (docs/adapters.md).
- Talk to it from anywhere: the MCP server speaks
stdiotoday, so it is a local subprocess of Claude Code or Claude Desktop. An HTTP or SSE transport would make it a remote connector, which is what a phone talks to. That needs a long lived process, a reachable address and auth the server does not have yet (docs/mcp.md). - v1: a starter task on a real duck, on video. An Open Duck Mini can get there first, and a Microduck once it ships.
- v2, learned verbs. LLM written rewards (Eureka and DrEureka style) train new policies in
microduck_rlthat register as one more verb. The registry hook exists today. The training loop does not.
Help wanted: a recorded browser session with web/, because the page boots and a held W walks the duck but nobody has watched a model drive a whole run, a key barge in out of one, or the Record button work, on any machine but the one that wrote it, a real model recording in either simulator (see docs/assets), a transcript from a local model run on any server, a run against any real hardware (an Open Duck Mini is the most reachable, see its checklist), a pilot flock driven by a real model rather than the scripted one, with its transcripts, and new .duck files.
Add your .duck to ducks/. PRs welcome. That is the community funnel and the number we actually care about. Adding a verb to a robot is one function plus one manifest entry. Both are described in CONTRIBUTING.md, and design decisions live in docs/adr/. Tests run with no network and no keys: uv sync --extra dev && uv run pytest.
Thank you to everyone who has sent quackd code. 0.6 was the first release built on other people's pull requests, and both of them changed the project: one gave every robot a memory between runs, the other closed a budget a slow model could walk straight through. A bug report or a .duck that mostly fails counts too, because that is data.
Run on the floor, not a table. Keep pets and kids clear of kick. quackd adds a heartbeat, a kill switch (Ctrl+C or q stops the robot, a second Ctrl+C quits), allowlists, confirmation gates and budgets, and stop always means stop rather than collapse, see docs/safety.md. Before a robot that cannot detect its own fall walks, quackd run asks once whether you are watching it, and no is the default (--yes skips the question). In a flock that one kill switch reaches every member's executor, so it stops every body rather than the one in front of you. Who stops the body when quackd goes quiet differs per robot, and each manifest says so honestly.
On a Microduck the gamepad preempts remote control and robotd is the safety authority. On an Open Duck Mini it is quackd's own daemon, running on the robot and zeroing the velocity after 300 ms of silence, inside the loop rather than on a timer, so a dead laptop still stops the duck. That duck cannot get up if it falls, so work with it on a stand until you trust the link, and keep a hand near the power switch, which is its only e-stop. A ToddlerBot cannot get up either, and on that body torque off is a fall, so the daemon quackd ships for it answers silence by slewing to a safe pose and holding, never by letting go. You are responsible for your robot.
They built the duck, and quackd began as its brain. Thanks to Pollen Robotics for microduck (the onboard daemon stack and its JSON RPC contract) and microduck_rl (the training stack behind the policies the robot runs), to the MCP Python SDK, and to the authors of DrEureka for the idea behind learned verbs. Thanks to Antoine Pirrone and the Open Duck Mini project for designing a biped anyone can print and build, and for publishing the runtime that makes it walk. Community: the Pollen Robotics Discord linked from the upstream README.
quackd is an independent community project, not affiliated with or endorsed by Pollen Robotics, Hugging Face or the Open Duck Mini project. "Microduck" is used nominatively to describe compatibility. No Pollen Robotics or Open Duck Mini logo, mesh, ONNX policy or sound is distributed here. The physics simulator fetches the Microduck's model and its policies from upstream at run time and checks every file against a recorded hash. The browser demo fetches the same files at the same pin straight into the visitor's browser and hashes nothing. The README hero renders that model, so it carries the model's own CC BY-NC-SA terms (docs/licenses.md).
Apache 2.0. Third party and asset licenses (including why the robot's CC BY NC SA meshes are never vendored) are in docs/licenses.md and NOTICE.

{"kind": "llm", "step": 0, "tool_calls": [{"name": "assess_task", "arguments": {"verdict": "feasible", "reason": "scripted pilot: a rule has no judgement of the body, so it assumes the task fits"}}], "usage": {"input_tokens": 1602, "output_tokens": 16}} {"kind": "assess", "step": 0, "verdict": "feasible", "reason": "scripted pilot: a rule has no judgement of the body, so it assumes the task fits"} {"kind": "llm", "step": 0, "tool_calls": [{"name": "search_scan", "arguments": {"target": "ball"}}]} {"kind": "verb", "step": 1, "name": "search_scan", "ok": true, "summary": "ball found: ball at bearing 18° left ~0.58 m (after 4 turn steps)"} {"kind": "llm", "step": 1, "tool_calls": [{"name": "walk_to", "arguments": {"target": "ball", "stop_distance": 0.22}}]} {"kind": "verb", "step": 2, "name": "walk_to", "canonical": "go_to", "ok": true, "summary": "reached the ball: ~0.22 m away, bearing +0°", "data": {"distance_m": 0.217, "ticks": 27}} {"kind": "llm", "step": 2, "tool_calls": [{"name": "kick", "arguments": {"leg": "right"}}]} {"kind": "verb", "step": 3, "name": "kick", "ok": true, "summary": "kicked with right leg, ball moved 0.53 m"} {"kind": "llm", "step": 3, "tool_calls": [{"name": "quack", "arguments": {"text": "yay, got it!"}}]} {"kind": "llm", "step": 4, "tool_calls": [{"name": "declare_success", "arguments": {"reason": "ball displaced by the kick"}}]}