Two interchangeable brains · pure simulation · full-body 3D humanoid.
Talk to it and a 3D humanoid walks, gestures and reacts, with an LED face that picks its own expression from what it's saying.
┌──────────────────────────────────────────┐
you speak ──────▶│ BRAIN (one of two, same tool dispatch) │
│ nvidia Web Speech ↔ NVIDIA (default) │
│ ultravox realtime WebRTC, needs a plan │
└────────────────────┬─────────────────────┘
│
┌────────────────────┴────────────────────┐
│ display device (browser IS the head) │
│ camera = eyes · in-browser face recog │
│ owns the session · dispatches tools │
└───────┬──────────────────────┬──────────┘
set_expression │ perform_motion │
+ auto sentiment ▼ ▼
LED face 3D humanoid
▲
│ relay (ws://:8080/relay)
▼
control dashboard /control
The FastAPI backend is deliberately thin: it mints Ultravox calls and proxies NVIDIA (so neither key reaches the browser), relays state between windows, and writes logs. It holds no conversation state.
copy .env.example .envFill in NVIDIA_API_KEY (https://build.nvidia.com). ULTRAVOX_API_KEY
(https://app.ultravox.ai) is optional — it's the nicer brain but needs a paid
plan. Then:
.\click_to_install.bat.\start.batstart.bat builds the frontend, starts the server on :8080, and opens the
windows tiled. Press Power on, allow the microphone, and talk.
| URL | Window | What it is |
|---|---|---|
/robot (also /) |
Display + Simulation | Both in one page: LED face beside the 3D humanoid, one brain driving both |
/control |
Dashboard | Power, mute, brain switch, conversation, activity, motion grid |
/face |
Display only | The robot's own screen, for the physical 8" panel |
/sim |
Simulation only | 3D humanoid on its own |
Try saying:
- "Hey, walk over here" → starts
walk_forward, keeps talking, stops when told - "Nice to meet you" →
handshake+set_expression("happy") - "Show me some push-ups" →
push_uploops until you say stop - "Do a squat", "Dance for me", "Spin around", "Sit down", "Take a bow"
- "My name is Shyam" →
remember_personlearns your face - Interrupt mid-sentence — Ultravox barge-in cuts the speaker instantly
Requires Python 3.10+ and Node 20+.
nvidia (default) |
ultravox |
|
|---|---|---|
| Ears | browser SpeechRecognition | WebRTC audio |
| Thinking | NVIDIA chat completions, real tool calling | Ultravox realtime |
| Voice | browser speechSynthesis | Ultravox TTS |
| Feel | turn-based, ~1 s behind | realtime, instant barge-in |
| Needs | NVIDIA key only | a paid Ultravox plan |
ROBOT_MODE in .env picks one; auto tries Ultravox and falls back to
nvidia when the call can't be created. Both brains share the same tool
dispatch, so expressions, motions and memory behave identically either way.
The dashboard has an Auto / Ultravox / Web Speech switch for changing brain
mid-demo without editing .env.
Default is nvidia because the Ultravox subscription kept returning
"Set up your subscription to make more calls" — a brain that always works beats
a better brain that intermittently doesn't. A fallback is always announced on
screen; it never happens silently.
You never ask for an expression. sentiment.js scores every line the robot says and sets the face itself:
| It says | Face |
|---|---|
| "Nice to meet you" | happy |
| "That's amazing!!" | excited |
| "Whoa, no way!" | surprised |
| "I'm so sorry to hear that" | sad |
| "Careful, watch out!" | scared |
| "Let me think…" | thinking |
| "What's your name?" | curious |
It also reacts while you are still talking, so the face isn't frozen
mid-turn. set_expression still works and outranks inference for 7 seconds
when the model chooses an emotion deliberately.
It's a keyword/punctuation heuristic on purpose: it has to resolve in the same frame the words start playing, and a network round trip for "is this happy" would land after the sentence had ended.
Display device owns everything — the session, the camera, in-browser face recognition, every tool call — and republishes what it decided on the relay.
Control dashboard (/control) is a thin remote. Power, mute and text all
round-trip through the relay rather than acting locally, so the dashboard can
never disagree with the robot about its own state.
Simulation is a pure consumer: it receives motion, expression and
speaking events and animates them. Its head renders the same LedFace
class as the display, so the two faces cannot drift apart.
/robot runs the brain, the face and the humanoid in one page. That needs
relay publishes to loop back locally — the server only fans messages out to
other clients, so a single page hosting both halves would otherwise never see
its own events.
Every window is optional. Close the simulator and the robot keeps talking.
All three consoles are built on css/theme.css — one set of tokens, buttons, pills, segmented controls and scroll areas. Light, warm and editorial: paper background, layered warm-tinted elevation, a single indigo accent, semantic colour only where it means something. Page stylesheets contain layout and nothing else.
The 3D stage is lit to match — white studio floor, soft-box rig, off-white composite shell with a clearcoat and machined-aluminium joints. A black viewport inside a paper-white UI reads as a broken embed, not a design choice.
The LED face stays black, because that is the hardware. On the combined page it's framed like a device screen rather than bled to the panel edge.
52 motions in frontend/src/motions.js — the single
source of truth. Adding one entry there makes it callable by voice, clickable in
the dashboard and animated in the simulator.
| Category | Motions |
|---|---|
| Locomotion | walk forward/backward, run, march, turn L/R, side-step L/R, spin |
| Posture | stand, squat, crouch, sit, stand up, lie down, balance on one leg, lean, T-pose |
| Social | wave hello/goodbye, handshake, high five, bow, salute, clap, thumbs up, celebrate, shrug, cross arms, hands on hips, thinking pose, hug, dance |
| Manipulation | point fwd/L/R, reach, pick up, place down, offer hand |
| Head | nod yes, shake no, tilt, look around |
| Exercise | push-ups, jumping jacks, jump, kick L/R, stretch, squat reps, warm-up twist |
Motions are procedural functions of time, not baked clips, so they compose and
can be swapped mid-flight. Looping motions (walk, run, dance, push-ups…) run
until stop_motion or the next motion.
Feet are grounded by solving actual contact. Each frame the simulator transforms the eight corners of both sole boxes into world space and slides the body so the lowest corner sits exactly on the floor. That one solver is why toe-off during a walk, tiptoes in a stretch, deep squats, kicks and the push-up plank all land correctly instead of needing per-motion fudge factors.
Camera presets (¾ / front / side / top / head), follow toggle, grid, recenter, searchable motion palette, expression pills, and keyboard shortcuts:
W/S walk · A/D turn · R run · Q squat · E wave · F handshake
C clap · J jump · B bow · X dance · Z sit · T t-pose · SPACE stop
| Tool | Goes to |
|---|---|
set_expression(emotion) |
LED face — 15 emotions, both screens (an override; the face already moves on its own) |
perform_motion(motion) |
3D humanoid — 52 whole-body motions |
stop_motion() |
back to a relaxed idle stand |
remember_person(name) / remember_fact(fact) / forget_person(name) |
local face memory (IndexedDB, never leaves the machine) |
execute_task(...) |
dashboard task queue |
execute_robot_action(...) |
look_at / grasp / release in camera frame |
Eight tools, ~3.1k tokens of context per turn. It was 10 tools and ~4.7k until the 52-motion catalogue turned out to be spelled out in both the system prompt and the tool description — enough redundancy that a small model lost the thread of a one-word greeting and started inventing topics. The enum in the schema is now the only copy.
| Path | What it is |
|---|---|
frontend/main.py |
FastAPI — Ultravox minting, /api/chat, /api/think, /api/mode, relay, logs |
frontend/robot.html |
Display + simulation in one page |
frontend/control.html |
Operator dashboard |
frontend/face.html |
Display device on its own |
frontend/simulation.html |
Simulation on its own |
frontend/css/theme.css |
Design system — tokens + primitives, shared by every console |
frontend/src/main.js |
Both brains + tool dispatch |
frontend/src/humanoid.js |
Three.js rig + procedural motion engine |
frontend/src/motions.js |
Motion catalogue — the single source of truth |
frontend/src/sentiment.js |
Automatic expression inference |
frontend/src/face-renderer.js |
LED dot-matrix face, shared by display and sim head |
frontend/src/config.js |
Settings + the robot's persona/system prompt |
frontend/src/tools.js |
Tool declarations |
frontend/src/robot.js |
World model + task queue |
frontend/src/faces.js |
In-browser face recognition |
frontend/src/camera-source.js |
ESP32-CAM feed → MediaStream |
frontend/esp32_cam.py |
Camera ingest hub — one upstream, many subscribers |
ros2/humanoid_control/ |
ROS 2 Jazzy — limb nodes, voice bridge, camera driver |
tools/check_esp32cam.py |
Camera diagnostic |
firmware/ |
ESP32 sketch (hardware arm, no longer driven by the app) |
ros2/humanoid_control/ drives the hardware from the same voice stack
that drives the simulator. Four limb controllers, a voice bridge and the camera
driver:
ros2 launch humanoid_control humanoid.launch.py laptop_ip:=192.168.137.1voice_bridge joins the backend's relay socket as one more listener, so it
receives the identical perform_motion event the simulator does. That is the
whole integration: the robot and the on-screen twin cannot disagree, and no code
in frontend/ knows ROS 2 exists.
motion_map.py translates the 52-motion catalogue into limb commands and marks
each one full (30), approximate (13, e.g. a bow plays as a crouch — no
waist joint) or unsupported (9, e.g. push-ups). Unsupported motions hold
position and say so in the activity feed rather than substituting a different
gesture.
Bring it up without commanding a servo, to check speech is arriving first:
ros2 launch humanoid_control humanoid.launch.py dry_run:=trueSee ros2/README.md.
All keys live in the repo-root .env (gitignored):
NVIDIA_API_KEY— the default brain (/api/chat) and/api/thinkULTRAVOX_API_KEY— the realtime brain, optionalROBOT_MODE—nvidia|ultravox|autoESP32_CAM_HOST— the robot's camera (not a secret, but it lives here too)
No key is ever sent to the browser. The client asks /api/ultravox for a
call and receives only a single-use joinUrl; NVIDIA is called entirely
server-side. Do not reintroduce a VITE_-prefixed API key — Vite inlines those
straight into the public bundle, which is exactly how the Ultravox key was
leaking into shipped JS before.
The camera is on the robot, not the laptop. The backend proxies it at
/api/camera/stream, and camera-source.js
paints that MJPEG feed into a canvas and hands captureStream() back as a
MediaStream — so face recognition, the tracking overlay and the WebRTC feed to
the dashboard are unchanged from when this was a webcam.
The proxy is not decoration. The ESP32's web server holds one connection for the
whole duration of a stream, so whichever client connects first locks out every
other one — and the browser head and the esp32_cam ROS 2 node both need it.
The backend takes that single upstream slot and fans frames out locally. Serving
same-origin also keeps the canvas untainted, which is what lets the face engine
call getImageData at all.
VITE_CAMERA_SOURCE picks the source: auto falls back to the laptop webcam
when the ESP32 is unreachable (right for a dev machine), esp32 refuses to fall
back (right for the robot — a silent substitution there means it narrates the
room it is sitting in front of, not the one it is looking at).
Camera not coming up?
python tools/check_esp32cam.pyThe physical ESP32 servo arm is still gone — tools, client dispatch, backend
proxy and dashboard status. perform_motion covers every gesture it could do,
with a whole body behind it instead of one limb. That firmware sketch is in
firmware/ if you want to wire it back up.
Gemini Live and a local-LLM mode used to live here too. They are gone. Three half-wired brains meant a failure in one silently became a different one; the two that remain share one dispatch and announce every switch on screen.
chromium --kiosk --app=http://localhost:8080/faceThe relay auto-reconnects if either side restarts.