Skip to content

About

Real-time voice AI agent shipped in 2 hours. Cloudflare Workers + OpenAI Realtime API + WebRTC — sub-second-latency voice-to-agent pipeline. Demonstrates edge-runtime, streaming audio, and modern AI-agent architecture.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

A voice front-door for Hermes. Speak into a Cloudflare Worker; your VPS agent ships a live site.

▶ 2-minute demo · live app · built in 2h for the Nous Research internship challenge.

demo


TL;DR

Hermes (Nous Research Hermes Agent) is a real, autonomous VPS agent — shell, browser, wrangler, messaging gateway. People text it on Telegram and it deploys websites, edits repos, runs commands.

Donna is the voice interface Hermes was missing. Tap the mic on a Cloudflare Worker, describe what you want, and Donna — running on OpenAI Realtime 2.1 — turns your speech into a precise instruction, fires it at Hermes over the same channel a human would use, then speaks the reply back with the live URL Hermes just deployed.

Try it live: open the app, tap Talk, and say:

"Deploy a one-page landing site to Cloudflare Pages that says 'hello from Donna' and give me the URL."

Architecture

    ┌────────────┐   WebRTC audio    ┌───────────────────┐
    │  Browser   │ ─────────────────▶│  OpenAI Realtime  │
    │  mic + UI  │◀───────────────── │   gpt-realtime-2.1│
    └──────┬─────┘   TTS + events    └─────────┬─────────┘
           │                                    │ tool_call:
           │  POST /api/session                 │  send_to_hermes(msg)
           ▼                                    ▼
    ┌───────────────────────────────────────────────────┐
    │           Cloudflare Worker  (this repo)          │
    │  • mints ephemeral Realtime client_secret         │
    │  • forwards send_to_hermes → Hermes VPS           │
    └───────────────────────┬───────────────────────────┘
                            │  HTTPS (cloudflared tunnel)
                            ▼
                ┌───────────────────────┐
                │  Hermes Agent on VPS  │
                │  Nous Research OSS    │
                │  shell · wrangler ·   │
                │  browser · files      │
                └───────────┬───────────┘
                            │  `wrangler pages deploy`
                            ▼
                ┌───────────────────────┐
                │   Cloudflare Pages    │
                │   live user site      │
                └───────────────────────┘

How it works — 6 hops

  1. Browser requests an ephemeral Realtime token from /api/session on the Worker.
  2. Worker calls POST /v1/realtime/client_secrets on OpenAI, returns a short-lived client_secret.value — never exposing the real API key.
  3. Browser opens a WebRTC session directly to POST /v1/realtime/calls, streams mic audio, listens for TTS + data-channel events.
  4. OpenAI Realtime 2.1 transcribes, reasons, and — when the user asks for anything actionable — calls the send_to_hermes tool over the data channel.
  5. Browser relays the tool call to /api/hermes on the Worker, which POSTs it to Hermes on the VPS (over a cloudflared tunnel or straight to the VPS's HTTP gateway). Fallback path uses the Telegram gateway that ships with Hermes.
  6. Hermes executes (e.g. wrangler pages deploy), returns the reply with the live URL. The Worker returns it to the browser, the browser feeds it back to Realtime as function_call_output, and Donna speaks the URL aloud.

Quickstart

npm install && npx wrangler dev        # local dev on :8787
npx wrangler deploy                    # ship to workers.dev

Then set the three secrets:

wrangler secret put OPENAI_API_KEY     # from platform.openai.com/api-keys
wrangler secret put HERMES_HTTP_URL    # https://<cloudflared>.trycloudflare.com/message
wrangler secret put HERMES_HTTP_TOKEN  # bearer for that endpoint

Telegram-gateway fallback (no HTTP on Hermes): set TELEGRAM_BOT_TOKEN + TELEGRAM_CHAT_ID instead — Donna will drop the instruction into the same chat a human uses. See src/index.js for the polling logic.

Stack

Layer Choice Why
Voice OpenAI Realtime 2.1 (gpt-realtime-2.1-mini) Direct WebRTC, sub-500ms latency, server-side VAD, native tool-calling on the data channel
Edge Cloudflare Worker + [assets] binding Zero-cold-start static + API in one runtime; ephemeral tokens minted at the edge
Agent Nous Research Hermes Agent on a VPS The actual product this challenge is about — messaging gateway, tool gateway, shell + wrangler
Deploy target Cloudflare Pages Hermes calls wrangler pages deploy — under 10s from tool-call to live URL

What I'd build next

  • Persist conversations to Workers KV so Donna can pick up mid-thought across reloads.
  • Second tool screenshot(url) — after Hermes deploys, screenshot the live page and hand the image back into Realtime for a spoken "looks good, it renders like X".
  • Voice-only auth: enroll a passphrase on first use, embed it into the Realtime session, refuse send_to_hermes calls unless the speaker matches.
  • Swap the polling Telegram fallback for a Worker Durable Object that reads getUpdates in the background.

Repo tour

src/index.js           Worker: session mint + Hermes bridge (HTTP + Telegram)
public/index.html      Mic UI
public/app.js          WebRTC + tool-call dispatch
wrangler.toml          Worker config with static-asset binding
docs/                  Demo assets (gif, screenshots)

Built in 2h for the Nous Research Hermes/Donna internship challenge — Sayuj Pillai.

About

Real-time voice AI agent shipped in 2 hours. Cloudflare Workers + OpenAI Realtime API + WebRTC — sub-second-latency voice-to-agent pipeline. Demonstrates edge-runtime, streaming audio, and modern AI-agent architecture.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages