Skip to content

Repository files navigation

Signbridge

A real-time, on-device sign-language communication aid. It watches signs through the camera and turns them into on-screen text, and turns a hearing person's speech into captions back — a two-way bridge for a quick conversation, with no interpreter and no special hardware.

🔗 Live demo: signbridgebysd.vercel.app — open in Chrome/Edge or Safari and allow the camera to start signing.

Built with the Morflax "porcelain gallery" visual system: a monochrome porcelain canvas under a near-black nav, with a single cobalt spark rationed to one action per view — a calm, legible, trustworthy look for an accessibility tool.

Honest scope (please read). Sign languages are full natural languages with their own grammar. "Point a camera and get perfect English" is not a solved problem, and claiming it would be dishonest. Signbridge is an assistive aid, not a universal translator: it recognizes a defined, extensible vocabulary — the ASL fingerspelling alphabet plus signs you teach it — and assembles them into text. It is not for medical, legal, or emergency use. See DESIGN.md §2.

What it does

  • Sign → text. MediaPipe tracks your hand; a geometric classifier recognizes the static ASL fingerspelling alphabet; letters assemble into words and sentences on a big, legible caption surface.
  • Phrase gestures (full-body). MediaPipe Holistic adds head + shoulder tracking, so a dynamic-gesture engine (temporal window + motion gating) recognizes a bundled set of moving phrase signs — YES · NO · WHY · WHO · WHAT — that fingerspelling can't express. Toggle in Settings; it's heavier than hands-only tracking.
  • Teach a sign. Record five takes of your own sign; it's stored on-device and recognized immediately (few-shot, nearest-exemplar). This is how the vocabulary grows to fit you and how it sidesteps the scarce-dataset problem.
  • Speech → captions. The hearing side replies by voice (Web Speech API) or by typing; everything lands in a shared, copyable transcript.
  • Accessible by construction. Keyboard operable, screen-reader labelled, live-region captions, user-scalable text, high-contrast and reduce-motion modes.

Privacy

  • Camera: on-device, always. Video is processed locally (WASM/WebGL) and never leaves your device.
  • Speech: honestly labelled. In Chrome/Edge the Web Speech API sends microphone audio to the browser's speech service (i.e. not on-device); Safari uses Apple's dictation; Firefox has none, so replies are typed. The app tells you which is active at mic-permission time and in Settings, and typing is always available and fully local. True on-device speech (Whisper / SFSpeechRecognizer) is on the roadmap. See DESIGN.md §7.
  • Your taught signs live in IndexedDB on this device; export a JSON backup from the Teach screen.

Pipeline

flowchart LR
  cam[Camera frame] --> mp[MediaPipe HandLandmarker]
  mp --> feat[Normalize landmarks\ntranslation/scale/rotation invariant]
  feat --> cls{Classify}
  cls -->|built-in| heur[Geometric alphabet classifier]
  cls -->|personal| knn[KNN over taught exemplars]
  heur --> gate[Confidence gate + debounce\nhold Tms, no double-emits]
  knn --> gate
  gate --> asm[Text assembler\nfingerspelling → words → sentences]
  asm --> clean[Rules grammar cleanup]
  clean --> ui[Caption surface + transcript]
  mic[Microphone] --> speech[Web Speech API] --> ui
Loading

Run it

npm install
npm run dev      # http://localhost:5188
npm test         # unit + pipeline-replay tests (no camera needed)
npm run build    # production build → dist/
npm run eval     # recognition accuracy on a dataset → confusion matrix + eval-report.md

Open in Chrome/Edge or Safari and allow the camera. Add ?hud=1 to the URL for a dev overlay (FPS, recognizer state, top prediction).

Recognition approach & honest limitations

Tier 0 uses a geometric/prototype classifier rather than a trained neural net: each static letter is a set of fuzzy geometric conditions over normalized hand landmarks, and taught signs are matched by nearest-exemplar KNN in the same feature space. This is a deliberate, honest choice — there's no ethically-sourced signing dataset bundled here, and KNN-over-features is the Teach-a-sign mechanism. A trained TF.js head is the documented Tier 1 upgrade.

Known limitations, surfaced rather than hidden:

  • J and Z require motion and M, N hide the thumb — these static frames can't separate them reliably. Teach them as personal signs instead.
  • Phrase gestures are hand-authored heuristics, not a trained model. YES/NO/WHAT are pure hand-motion and most reliable; WHY/WHO depend on face-relative position and need your head and shoulders in frame with good lighting. Expect misses. The dynamic engine is proven self-consistent and discriminative on synthetic trajectories (tests/dynamic.test.js); the ?hud=1 overlay shows live phrase scores for tuning against real motion.
  • Accuracy depends on lighting, framing, and hand orientation; the confidence cue and adjustable sensitivity/hold-time exist to manage that.
  • Continuous, grammatical translation of free-form signing is explicitly not attempted.

Recognition eval & data capture (dev)

Recognition changes are measured, not guessed. A hidden capture tool at #/capture records labeled samples — raw image + world landmarks + handedness — and exports a JSON dataset. npm run eval -- your-dataset.json runs it through the exact feature extractor + classifier the app uses and prints per-class precision/recall/F1, a confusion matrix, and (for Phase-2 calibration) the mean world thumb-depth per class. npm run eval alone runs a synthetic smoke set.

World-landmark depth (N vs S/M): the fist-family letters hinge on whether the thumb is in front of the curled fingers (S/E) or tucked behind (N/M). The classifier now uses MediaPipe's metric-3D worldLandmarks for that cue (worldThumbDepth), which is far more reliable than the monocular image-z. Because the sign of that depth depends on how signs present, it's calibrated from captured data: record N and S at #/capture, run npm run eval, and if the two mean depths come out with the opposite sign to what's assumed, flip one constant (WORLD_FRONT_SIGN in heuristics.js).

Testing

npm test runs Vitest over the pure pipeline modules with no camera required:

  • Feature invariance — normalized landmarks are stable under translation and scale. (This suite caught a real wrist-aliasing bug during development.)
  • Gate/debounce — a token emits only after a stable hold; no flicker double-emits; deliberate double letters still work.
  • Pipeline replay — a synthetic prediction stream flows gate → assembler → cleanup deterministically ("HI" + gap + "YOU" → Hi you.).
  • Assembler, cleanup, KNN — fingerspelling merge, grammar rules, nearest-exemplar match.

Tech

Vite · vanilla JS · @mediapipe/tasks-vision HandLandmarker · Web Speech API · IndexedDB · Inter. Static and client-only — no backend.

Ethics

Sign languages are real languages belonging to real communities. Any bundled vocabulary beyond fingerspelling should be sourced with Deaf-community input and credited. This project is an accessibility aid built with that respect in mind, not a replacement for human interpreters.

See DESIGN.md for the full product spec, architecture, and rationale.

About

Real-time, on-device sign-language communication aid — sign-to-text and speech-to-caption in the browser. Accessibility engineering with an honest ML scope.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages