Skip to content

Repository files navigation

Resonance

Resonance is a private, browser-based 3D voice assistant. A living neon waveform listens through your microphone, transcribes locally with Whisper, thinks locally with WebLLM, and speaks with a local Transformers.js voice model. There is no backend, account, API key, or remote inference service.

Screenshots: add desktop and mobile captures here after deployment.

Architecture

flowchart LR
  Mic[Microphone] --> Audio[Web Audio analyser]
  Audio --> STT[Whisper worker]
  STT --> LLM[WebLLM WebGPU worker]
  Text[Text input] --> LLM
  LLM --> Stream[Streaming response]
  Stream --> TTS[Local TTS worker]
  TTS --> Playback[Web Audio playback]
  Playback --> Audio
  Audio --> Wave[Adaptive 2D waveform]
  Registry[Model registry] --> Cache[Browser model caches]
  STT --- Cache
  LLM --- Cache
  TTS --- Cache
Loading

One adaptive visual surface reads a cached mutable analyser snapshot independently from React rerenders. It uses a capped Canvas 2D loop on capable devices and a lightweight SVG fallback on constrained mobile devices. Zustand owns the explicit assistant state machine, preferences, conversation context, and separate download jobs. The three expensive AI runtimes are lazy-loaded inside dedicated module workers.

Privacy model

  • Prompts, responses, conversation context, and microphone audio are processed in this browser.
  • The app has no backend and makes no calls to cloud inference APIs.
  • Static model files download from MLC and Hugging Face distribution infrastructure and may be cached by the browser. Those hosts receive ordinary file requests, not your prompts or audio.
  • Microphone recordings are held only long enough to transcribe and are never persisted.
  • If local TTS fails, Resonance clearly falls back to browser/OS SpeechSynthesis. That facility is not guaranteed offline and follows the browser or operating system’s own privacy behavior.
  • Preferences and advisory model metadata use local storage. Conversation messages are session-memory only.

Browser requirements

  • A current Chromium-family browser with WebGPU enabled is recommended.
  • A compatible GPU and enough graphics memory for the selected LLM are required.
  • Microphone capture requires a secure context (https or localhost).
  • Safari and Firefox support for WebGPU, WebAssembly model operations, codecs, and large cache entries varies. Unsupported browsers receive a visual demo and never silently use cloud AI.

Local development

Install pnpm and use Node.js 22 or newer.

pnpm install
pnpm dev

Quality checks:

pnpm lint
pnpm typecheck
pnpm test
pnpm build

First run and model downloads

No model download begins on page load. Setup reads the current WebLLM prebuilt model catalog, selects a compact instruction-tuned quantized model with a transparent heuristic, and shows the best available size/storage estimates.

  • Prepare full voice assistant queues the LLM, multilingual Whisper Tiny, and, on desktops, the multi-voice Supertonic 2 speech model. Constrained mobile devices use the operating system voice so local TTS does not compete with WebGPU inference for the browser’s memory budget.
  • Start with text only initializes only the LLM. Speech models are downloaded later only after an explicit action.

Setup also offers environment presets while keeping the model catalog available for manual selection:

  • Mobile pairs a compact 0.5B-class model and local listening with the system voice.
  • Fast pairs a responsive 1B-class model with the listening and speech models.
  • Medium pairs a balanced 3B-class model with full voice.
  • Smartest pairs an 8B-class instruct model with full voice for capable desktops.

The exact quantized model is selected from the live WebLLM catalog with language-aware scoring. The browser-reported device memory determines which preset receives the Suggested badge. Any manual model or voice-mode change switches the selection to Custom.

Progress is determinate only when the model library exposes real totals or a real progress value. Shader compilation, verification, cache loading, and other unknown-total work uses an indeterminate treatment. Successful files are normally reused from browser caches after the first load.

Models and storage

Use Existing files during setup, or open Settings → Models & storage afterward, to inspect recognized model bundles, browser usage/quota estimates, and local status. Models can be removed individually when a clean download is needed. Bulk removal lists the models and requires explicit confirmation.

Resonance never calls localStorage.clear(), never deletes every cache, and never clears unrelated site data. For WebLLM it uses the library’s per-model deletion API when present. For Transformers.js it deletes only cache requests whose URL matches the selected model identifier. A browser can still report approximate sizes, omit quota data, evict cached models, or change internal cache behavior. Partial cleanup failures are shown and can be retried.

The small registry is advisory: it tracks selection, status, approximate size, known identifiers, and last use. It is reconciled against detectable cache state after refresh and deletion.

GitHub Pages deployment

  1. Push the project to a GitHub repository and enable Settings → Pages → GitHub Actions.
  2. Push to main. .github/workflows/deploy.yml installs with pnpm, runs lint, type checking and tests, builds, then deploys dist.
  3. The workflow sets VITE_BASE_PATH to /<repository-name>/, so worker and asset URLs work below a repository subpath.

For a custom Pages base path, set it manually during build:

VITE_BASE_PATH=/my-path/ pnpm build

For a user/organization root site, use VITE_BASE_PATH=/.

Model configuration

Speech model IDs and approximate sizes live in src/config/models.ts. Compatible LLMs come from prebuiltAppConfig.model_list; the UI does not rely on one permanent hardcoded model. The default scoring favors chat/instruct models around 0.5B–1.5B parameters and low-memory quantization, with Llama 3.2 and Qwen 2.5 preferred for Spanish conversations.

Conversation languages

Choose English or Español during setup or under Settings → Voice & interaction. The selection controls Whisper transcription, the local LLM system prompt, Supertonic 2 speech, and the browser/OS speech fallback. Prompts, microphone audio, transcripts, and generated speech remain local under the same privacy model. The application interface itself is currently English.

Changing a selector never silently downloads a new model. Models download only through a labeled initialization or download action.

Performance guidance

Use a small quantized LLM first. Lower Wave detail, enable Reduced motion, and close other GPU-heavy tabs on integrated graphics. Visual frame rate and device pixel ratio are capped by device tier and assistant state; the loop stops when hidden or covered by a dialog. Geometry, gradients, analyser snapshots, and typed arrays are reused. Speech and LLM work remain isolated from the main thread, while microphone PCM capture and 16 kHz downsampling use an AudioWorklet when available.

The setup model list is a small curated catalog tied to the pinned WebLLM package version. This keeps the multi-megabyte WebLLM runtime out of a clean setup visit until the user explicitly initializes a model. When upgrading WebLLM, update WEB_LLM_CATALOG_VERSION, verify every curated model ID against the new prebuilt catalog, and run pnpm build:budget.

Known limitations

  • WebGPU model compatibility and memory limits vary widely by driver.
  • First-run downloads can be hundreds of megabytes or more.
  • Storage quota and usage values are estimates; private-browsing modes may offer very little space.
  • Browsers may evict cached files at any time.
  • Per-model Transformers.js deletion relies on identifiable model URLs because the library does not expose a universal per-model removal API.
  • MediaRecorder codecs and browser speech fallback behavior vary by platform.
  • Conversation input and responses support English and Spanish; the application interface is currently English-only.

Troubleshooting

  • WebGPU unavailable: update the browser and graphics driver, then check the browser’s WebGPU diagnostics. Visual demo mode remains available.
  • Out of memory/device lost: select a smaller quantized LLM, lower wave detail, close GPU-heavy tabs, and reload.
  • Microphone denied: allow microphone access in site settings or continue with text input.
  • Download failure: verify network and free storage, then retry. Cached partial files are kept when the model library supports reuse.
  • Audio does not start: interact with the page once to resume the browser AudioContext.
  • Model repeatedly disappears: request persistent storage during setup; the browser may deny it.

License

MIT. Model files retain the licenses published by their respective model authors.

About

Private, browser-based local AI voice assistant powered by WebGPU

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages