A real-time, on-device sign-language communication aid. It watches signs through the camera and turns them into on-screen text, and turns a hearing person's speech into captions back — a two-way bridge for a quick conversation, with no interpreter and no special hardware.
🔗 Live demo: signbridgebysd.vercel.app — open in Chrome/Edge or Safari and allow the camera to start signing.
Built with the Morflax "porcelain gallery" visual system: a monochrome porcelain canvas under a near-black nav, with a single cobalt spark rationed to one action per view — a calm, legible, trustworthy look for an accessibility tool.
Honest scope (please read). Sign languages are full natural languages with their own grammar. "Point a camera and get perfect English" is not a solved problem, and claiming it would be dishonest. Signbridge is an assistive aid, not a universal translator: it recognizes a defined, extensible vocabulary — the ASL fingerspelling alphabet plus signs you teach it — and assembles them into text. It is not for medical, legal, or emergency use. See DESIGN.md §2.
- Sign → text. MediaPipe tracks your hand; a geometric classifier recognizes the static ASL fingerspelling alphabet; letters assemble into words and sentences on a big, legible caption surface.
- Phrase gestures (full-body). MediaPipe Holistic adds head + shoulder tracking, so a dynamic-gesture engine (temporal window + motion gating) recognizes a bundled set of moving phrase signs — YES · NO · WHY · WHO · WHAT — that fingerspelling can't express. Toggle in Settings; it's heavier than hands-only tracking.
- Teach a sign. Record five takes of your own sign; it's stored on-device and recognized immediately (few-shot, nearest-exemplar). This is how the vocabulary grows to fit you and how it sidesteps the scarce-dataset problem.
- Speech → captions. The hearing side replies by voice (Web Speech API) or by typing; everything lands in a shared, copyable transcript.
- Accessible by construction. Keyboard operable, screen-reader labelled, live-region captions, user-scalable text, high-contrast and reduce-motion modes.
- Camera: on-device, always. Video is processed locally (WASM/WebGL) and never leaves your device.
- Speech: honestly labelled. In Chrome/Edge the Web Speech API sends microphone audio to
the browser's speech service (i.e. not on-device); Safari uses Apple's dictation; Firefox
has none, so replies are typed. The app tells you which is active at mic-permission time and
in Settings, and typing is always available and fully local. True on-device speech (Whisper /
SFSpeechRecognizer) is on the roadmap. See DESIGN.md §7. - Your taught signs live in IndexedDB on this device; export a JSON backup from the Teach screen.
flowchart LR
cam[Camera frame] --> mp[MediaPipe HandLandmarker]
mp --> feat[Normalize landmarks\ntranslation/scale/rotation invariant]
feat --> cls{Classify}
cls -->|built-in| heur[Geometric alphabet classifier]
cls -->|personal| knn[KNN over taught exemplars]
heur --> gate[Confidence gate + debounce\nhold Tms, no double-emits]
knn --> gate
gate --> asm[Text assembler\nfingerspelling → words → sentences]
asm --> clean[Rules grammar cleanup]
clean --> ui[Caption surface + transcript]
mic[Microphone] --> speech[Web Speech API] --> ui
npm install
npm run dev # http://localhost:5188
npm test # unit + pipeline-replay tests (no camera needed)
npm run build # production build → dist/
npm run eval # recognition accuracy on a dataset → confusion matrix + eval-report.mdOpen in Chrome/Edge or Safari and allow the camera. Add ?hud=1 to the URL for a dev
overlay (FPS, recognizer state, top prediction).
Tier 0 uses a geometric/prototype classifier rather than a trained neural net: each static letter is a set of fuzzy geometric conditions over normalized hand landmarks, and taught signs are matched by nearest-exemplar KNN in the same feature space. This is a deliberate, honest choice — there's no ethically-sourced signing dataset bundled here, and KNN-over-features is the Teach-a-sign mechanism. A trained TF.js head is the documented Tier 1 upgrade.
Known limitations, surfaced rather than hidden:
- J and Z require motion and M, N hide the thumb — these static frames can't separate them reliably. Teach them as personal signs instead.
- Phrase gestures are hand-authored heuristics, not a trained model. YES/NO/WHAT are pure
hand-motion and most reliable; WHY/WHO depend on face-relative position and need your head and
shoulders in frame with good lighting. Expect misses. The dynamic engine is proven
self-consistent and discriminative on synthetic trajectories (
tests/dynamic.test.js); the?hud=1overlay shows live phrase scores for tuning against real motion. - Accuracy depends on lighting, framing, and hand orientation; the confidence cue and adjustable sensitivity/hold-time exist to manage that.
- Continuous, grammatical translation of free-form signing is explicitly not attempted.
Recognition changes are measured, not guessed. A hidden capture tool at #/capture records
labeled samples — raw image + world landmarks + handedness — and exports a JSON dataset.
npm run eval -- your-dataset.json runs it through the exact feature extractor + classifier the
app uses and prints per-class precision/recall/F1, a confusion matrix, and (for Phase-2
calibration) the mean world thumb-depth per class. npm run eval alone runs a synthetic smoke set.
World-landmark depth (N vs S/M): the fist-family letters hinge on whether the thumb is in
front of the curled fingers (S/E) or tucked behind (N/M). The classifier now uses MediaPipe's
metric-3D worldLandmarks for that cue (worldThumbDepth), which is far more reliable than the
monocular image-z. Because the sign of that depth depends on how signs present, it's
calibrated from captured data: record N and S at #/capture, run npm run eval, and if the two
mean depths come out with the opposite sign to what's assumed, flip one constant
(WORLD_FRONT_SIGN in heuristics.js).
npm test runs Vitest over the pure pipeline modules with no camera required:
- Feature invariance — normalized landmarks are stable under translation and scale. (This suite caught a real wrist-aliasing bug during development.)
- Gate/debounce — a token emits only after a stable hold; no flicker double-emits; deliberate double letters still work.
- Pipeline replay — a synthetic prediction stream flows gate → assembler → cleanup
deterministically ("HI" + gap + "YOU" →
Hi you.). - Assembler, cleanup, KNN — fingerspelling merge, grammar rules, nearest-exemplar match.
Vite · vanilla JS · @mediapipe/tasks-vision
HandLandmarker · Web Speech API · IndexedDB · Inter. Static and client-only — no backend.
Sign languages are real languages belonging to real communities. Any bundled vocabulary beyond fingerspelling should be sourced with Deaf-community input and credited. This project is an accessibility aid built with that respect in mind, not a replacement for human interpreters.
See DESIGN.md for the full product spec, architecture, and rationale.