Skip to content

Repository files navigation

Robot — Humanoid Cognitive Core

Two interchangeable brains · pure simulation · full-body 3D humanoid.

Talk to it and a 3D humanoid walks, gestures and reacts, with an LED face that picks its own expression from what it's saying.

                    ┌──────────────────────────────────────────┐
   you speak ──────▶│  BRAIN  (one of two, same tool dispatch)  │
                    │  nvidia   Web Speech ↔ NVIDIA (default)   │
                    │  ultravox realtime WebRTC, needs a plan   │
                    └────────────────────┬─────────────────────┘
                                         │
                    ┌────────────────────┴────────────────────┐
                    │   display device (browser IS the head)  │
                    │   camera = eyes · in-browser face recog │
                    │   owns the session · dispatches tools   │
                    └───────┬──────────────────────┬──────────┘
          set_expression    │      perform_motion  │
          + auto sentiment  ▼                      ▼
                      LED face              3D humanoid
                                     ▲
                                     │  relay (ws://:8080/relay)
                                     ▼
                             control dashboard  /control

The FastAPI backend is deliberately thin: it mints Ultravox calls and proxies NVIDIA (so neither key reaches the browser), relays state between windows, and writes logs. It holds no conversation state.


Quick start

copy .env.example .env

Fill in NVIDIA_API_KEY (https://build.nvidia.com). ULTRAVOX_API_KEY (https://app.ultravox.ai) is optional — it's the nicer brain but needs a paid plan. Then:

.\click_to_install.bat
.\start.bat

start.bat builds the frontend, starts the server on :8080, and opens the windows tiled. Press Power on, allow the microphone, and talk.

URL Window What it is
/robot (also /) Display + Simulation Both in one page: LED face beside the 3D humanoid, one brain driving both
/control Dashboard Power, mute, brain switch, conversation, activity, motion grid
/face Display only The robot's own screen, for the physical 8" panel
/sim Simulation only 3D humanoid on its own

Try saying:

  • "Hey, walk over here" → starts walk_forward, keeps talking, stops when told
  • "Nice to meet you" → handshake + set_expression("happy")
  • "Show me some push-ups" → push_up loops until you say stop
  • "Do a squat", "Dance for me", "Spin around", "Sit down", "Take a bow"
  • "My name is Shyam" → remember_person learns your face
  • Interrupt mid-sentence — Ultravox barge-in cuts the speaker instantly

Requires Python 3.10+ and Node 20+.


Two brains

nvidia (default) ultravox
Ears browser SpeechRecognition WebRTC audio
Thinking NVIDIA chat completions, real tool calling Ultravox realtime
Voice browser speechSynthesis Ultravox TTS
Feel turn-based, ~1 s behind realtime, instant barge-in
Needs NVIDIA key only a paid Ultravox plan

ROBOT_MODE in .env picks one; auto tries Ultravox and falls back to nvidia when the call can't be created. Both brains share the same tool dispatch, so expressions, motions and memory behave identically either way. The dashboard has an Auto / Ultravox / Web Speech switch for changing brain mid-demo without editing .env.

Default is nvidia because the Ultravox subscription kept returning "Set up your subscription to make more calls" — a brain that always works beats a better brain that intermittently doesn't. A fallback is always announced on screen; it never happens silently.

Expressions are automatic

You never ask for an expression. sentiment.js scores every line the robot says and sets the face itself:

It says Face
"Nice to meet you" happy
"That's amazing!!" excited
"Whoa, no way!" surprised
"I'm so sorry to hear that" sad
"Careful, watch out!" scared
"Let me think…" thinking
"What's your name?" curious

It also reacts while you are still talking, so the face isn't frozen mid-turn. set_expression still works and outranks inference for 7 seconds when the model chooses an emotion deliberately.

It's a keyword/punctuation heuristic on purpose: it has to resolve in the same frame the words start playing, and a network round trip for "is this happy" would land after the sentence had ended.

The windows

Display device owns everything — the session, the camera, in-browser face recognition, every tool call — and republishes what it decided on the relay.

Control dashboard (/control) is a thin remote. Power, mute and text all round-trip through the relay rather than acting locally, so the dashboard can never disagree with the robot about its own state.

Simulation is a pure consumer: it receives motion, expression and speaking events and animates them. Its head renders the same LedFace class as the display, so the two faces cannot drift apart.

/robot runs the brain, the face and the humanoid in one page. That needs relay publishes to loop back locally — the server only fans messages out to other clients, so a single page hosting both halves would otherwise never see its own events.

Every window is optional. Close the simulator and the robot keeps talking.

Design system

All three consoles are built on css/theme.css — one set of tokens, buttons, pills, segmented controls and scroll areas. Light, warm and editorial: paper background, layered warm-tinted elevation, a single indigo accent, semantic colour only where it means something. Page stylesheets contain layout and nothing else.

The 3D stage is lit to match — white studio floor, soft-box rig, off-white composite shell with a clearcoat and machined-aluminium joints. A black viewport inside a paper-white UI reads as a broken embed, not a design choice.

The LED face stays black, because that is the hardware. On the combined page it's framed like a device screen rather than bled to the panel edge.


Body motions

52 motions in frontend/src/motions.js — the single source of truth. Adding one entry there makes it callable by voice, clickable in the dashboard and animated in the simulator.

Category Motions
Locomotion walk forward/backward, run, march, turn L/R, side-step L/R, spin
Posture stand, squat, crouch, sit, stand up, lie down, balance on one leg, lean, T-pose
Social wave hello/goodbye, handshake, high five, bow, salute, clap, thumbs up, celebrate, shrug, cross arms, hands on hips, thinking pose, hug, dance
Manipulation point fwd/L/R, reach, pick up, place down, offer hand
Head nod yes, shake no, tilt, look around
Exercise push-ups, jumping jacks, jump, kick L/R, stretch, squat reps, warm-up twist

Motions are procedural functions of time, not baked clips, so they compose and can be swapped mid-flight. Looping motions (walk, run, dance, push-ups…) run until stop_motion or the next motion.

Feet are grounded by solving actual contact. Each frame the simulator transforms the eight corners of both sole boxes into world space and slides the body so the lowest corner sits exactly on the floor. That one solver is why toe-off during a walk, tiptoes in a stretch, deep squats, kicks and the push-up plank all land correctly instead of needing per-motion fudge factors.

Simulator controls

Camera presets (¾ / front / side / top / head), follow toggle, grid, recenter, searchable motion palette, expression pills, and keyboard shortcuts:

W/S walk · A/D turn · R run · Q squat · E wave · F handshake
C clap · J jump · B bow · X dance · Z sit · T t-pose · SPACE stop

Tools the model can call

Tool Goes to
set_expression(emotion) LED face — 15 emotions, both screens (an override; the face already moves on its own)
perform_motion(motion) 3D humanoid — 52 whole-body motions
stop_motion() back to a relaxed idle stand
remember_person(name) / remember_fact(fact) / forget_person(name) local face memory (IndexedDB, never leaves the machine)
execute_task(...) dashboard task queue
execute_robot_action(...) look_at / grasp / release in camera frame

Eight tools, ~3.1k tokens of context per turn. It was 10 tools and ~4.7k until the 52-motion catalogue turned out to be spelled out in both the system prompt and the tool description — enough redundancy that a small model lost the thread of a one-word greeting and started inventing topics. The enum in the schema is now the only copy.


Layout

Path What it is
frontend/main.py FastAPI — Ultravox minting, /api/chat, /api/think, /api/mode, relay, logs
frontend/robot.html Display + simulation in one page
frontend/control.html Operator dashboard
frontend/face.html Display device on its own
frontend/simulation.html Simulation on its own
frontend/css/theme.css Design system — tokens + primitives, shared by every console
frontend/src/main.js Both brains + tool dispatch
frontend/src/humanoid.js Three.js rig + procedural motion engine
frontend/src/motions.js Motion catalogue — the single source of truth
frontend/src/sentiment.js Automatic expression inference
frontend/src/face-renderer.js LED dot-matrix face, shared by display and sim head
frontend/src/config.js Settings + the robot's persona/system prompt
frontend/src/tools.js Tool declarations
frontend/src/robot.js World model + task queue
frontend/src/faces.js In-browser face recognition
frontend/src/camera-source.js ESP32-CAM feed → MediaStream
frontend/esp32_cam.py Camera ingest hub — one upstream, many subscribers
ros2/humanoid_control/ ROS 2 Jazzy — limb nodes, voice bridge, camera driver
tools/check_esp32cam.py Camera diagnostic
firmware/ ESP32 sketch (hardware arm, no longer driven by the app)

ROS 2 Jazzy — the physical robot

ros2/humanoid_control/ drives the hardware from the same voice stack that drives the simulator. Four limb controllers, a voice bridge and the camera driver:

ros2 launch humanoid_control humanoid.launch.py laptop_ip:=192.168.137.1

voice_bridge joins the backend's relay socket as one more listener, so it receives the identical perform_motion event the simulator does. That is the whole integration: the robot and the on-screen twin cannot disagree, and no code in frontend/ knows ROS 2 exists.

motion_map.py translates the 52-motion catalogue into limb commands and marks each one full (30), approximate (13, e.g. a bow plays as a crouch — no waist joint) or unsupported (9, e.g. push-ups). Unsupported motions hold position and say so in the activity feed rather than substituting a different gesture.

Bring it up without commanding a servo, to check speech is arriving first:

ros2 launch humanoid_control humanoid.launch.py dry_run:=true

See ros2/README.md.

Secrets

All keys live in the repo-root .env (gitignored):

  • NVIDIA_API_KEY — the default brain (/api/chat) and /api/think
  • ULTRAVOX_API_KEY — the realtime brain, optional
  • ROBOT_MODE — nvidia | ultravox | auto
  • ESP32_CAM_HOST — the robot's camera (not a secret, but it lives here too)

No key is ever sent to the browser. The client asks /api/ultravox for a call and receives only a single-use joinUrl; NVIDIA is called entirely server-side. Do not reintroduce a VITE_-prefixed API key — Vite inlines those straight into the public bundle, which is exactly how the Ultravox key was leaking into shipped JS before.

The eyes are an ESP32-CAM

The camera is on the robot, not the laptop. The backend proxies it at /api/camera/stream, and camera-source.js paints that MJPEG feed into a canvas and hands captureStream() back as a MediaStream — so face recognition, the tracking overlay and the WebRTC feed to the dashboard are unchanged from when this was a webcam.

The proxy is not decoration. The ESP32's web server holds one connection for the whole duration of a stream, so whichever client connects first locks out every other one — and the browser head and the esp32_cam ROS 2 node both need it. The backend takes that single upstream slot and fans frames out locally. Serving same-origin also keeps the canvas untainted, which is what lets the face engine call getImageData at all.

VITE_CAMERA_SOURCE picks the source: auto falls back to the laptop webcam when the ESP32 is unreachable (right for a dev machine), esp32 refuses to fall back (right for the robot — a silent substitution there means it narrates the room it is sitting in front of, not the one it is looking at).

Camera not coming up?

python tools/check_esp32cam.py

The physical ESP32 servo arm is still gone — tools, client dispatch, backend proxy and dashboard status. perform_motion covers every gesture it could do, with a whole body behind it instead of one limb. That firmware sketch is in firmware/ if you want to wire it back up.

Gemini Live and a local-LLM mode used to live here too. They are gone. Three half-wired brains meant a failure in one silently became a different one; the two that remain share one dispatch and announce every switch on screen.

Deploying to the robot's screen

chromium --kiosk --app=http://localhost:8080/face

The relay auto-reconnects if either side restarts.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages