Live lecture captions, translation into 22 Indian languages, and a jargon glossary —
running entirely on a Snapdragon PC, with the network switched off.
Try it in your browser · Demo video (2 min) · Watch a recorded session · Measured results · Submission kit
Snapdragon® AI Lab Build & Present Challenge 2026 · A Hackathon Project by AndrinGodson
Sahaay (सहाय, "help") listens to any lecture on a laptop — a Zoom call, a YouTube video, a professor in the room — and does three things live: it captions the speech, translates each line into the student's own language with the technical terms kept intact, and explains the jargon as it is spoken. Whisper transcribes, NLLB-200 translates and Llama 3.2 explains, all on the device, with the speech model on the Snapdragon Hexagon NPU. No audio, no text and no account ever leaves the machine.
sahaay-offline.vercel.app runs Whisper and the translation model in your own browser, so anyone can try it without installing anything, and carries every measurement with its source. The full desktop pipeline — the NPU, all-system audio capture, the language-model glossary — is not hosted and cannot be: there is no NPU in a datacenter, and the whole claim is that the lecture never leaves your machine. Why, in full.
An engineering lecture in India is rarely in one language. It sounds like this:
"Matrix A ka determinant zero hoga, tabhi non-trivial solution milega."
A student following that in their second language has two problems at once. They have to parse a Hindi sentence, and they have to recognise the English technical terms inside it — the exact terms that will appear in the textbook and the exam.
Cloud captioning handles this badly and costs money per minute. It also needs connectivity that a lecture hall, a hostel room, or a rural college often does not have. And a student who is deaf or hard of hearing cannot wait for a laggy cloud round-trip mid-sentence.
Three things, live, while the lecture is happening:
- Captions the lecturer — capturing both the microphone and whatever is playing through the speakers, so an in-person class and a Zoom lecture both work with no setup.
- Translates each line into one of 22 Indian languages, while protecting the technical terms so "eigenvalue" stays "eigenvalue" instead of becoming a transliterated guess.
- Explains the jargon as it is spoken. When the lecturer says "eigenvalue", a one-line explanation appears in the student's language, in the sidebar, without them leaving the lecture to search for it.
When the session ends it writes structured notes, a glossary and a five-question self-test to a Markdown file.
Nothing leaves the device. There is no account, no API key, and no network call in the entire audio path.
This is the part that makes it a Snapdragon application rather than a Python app that happens to run on one — and it is measured, not asserted.
Sahaay runs two models concurrently and continuously for the length of a lecture: Whisper transcribing every few seconds, and Llama 3.2 generating glossary entries alongside it. Here is what the second one costs the first, on a CPU — Whisper Small, real recorded speech:
| Whisper latency | Mean | p95 | Real-time factor |
|---|---|---|---|
| Glossary idle | 3275.9 ms | 3422.3 ms | 0.409 |
| Glossary running | 12400.2 ms | 13241.6 ms | 1.55 |
On a CPU, the pipeline stops keeping up. Real-time factor crosses 1.0, which is the line that matters: above it, transcription is slower than the speech arriving, so captions fall further behind the lecturer every minute until they are useless. Not a tuning problem — two compute-bound models on one set of cores.
On a Snapdragon PC the Whisper encoder runs on the Hexagon NPU instead: 13.5 ms on X2 Elite, 129 of 129 layers on the NPU. The two models sit on separate silicon and the contention above disappears. Method and full numbers in docs/CONCURRENCY.md.
This number was wrong until 22 Sep 2026, and the way it was wrong is worth knowing. The harness reported RTF 0.043 → 0.299 and concluded the pipeline still kept up. It had measured
whisper_tiny_enagainst a synthetic test tone, and recorded neither fact —model_idisauto, so what gets benchmarked depends on which weights happen to be on the machine. The product resolves towhisper_small_portable. Both harnesses now name the model and the signal in their output.
A one-hour lecture, every day, for every student, is also precisely the workload that is absurd to send to the cloud — and precisely what a 45 TOPS NPU sitting idle in a laptop is for.
Captions arriving live, the jargon sidebar filling as terms are spoken, and the execution-provider badge reporting what the models are actually running on. Translation shows "not installed" here because this capture ran without the NLLB weights — the UI never claims a capability it does not have.
No install at all: sahaay-offline.vercel.app/live downloads Whisper once and runs it in your own browser — speak, or share a tab playing a lecture. Your audio never leaves the tab. Watch the RTF badge: after tuning it holds about 0.5 on a laptop CPU (docs/TUNING.md), because the browser runs one model. The measurement this project is about is what happens when a second model shares those cores (docs/CONCURRENCY.md). Pick a language and it translates too, into any of the 22 Indian languages the desktop app offers, with the same model (NLLB-200, about 900 MB, downloaded only when you pick one). The NPU and system-audio capture need the desktop build below.
No models, no audio device, no Snapdragon hardware:
git clone https://github.com/andringodson/Hackathon-SnapdragonAILab
cd Hackathon-SnapdragonAILab
powershell -ExecutionPolicy Bypass -File install.ps1
.\run.bat --mockA browser opens with a scripted code-mixed lecture running through the real pipeline — the same event plumbing, UI, glossary and notes writer the production path uses.
For the real thing:
.\.venv\Scripts\python.exe scripts\download_models.py --auto
.\run.batThen play any lecture and press Start.
The most common way a live demo fails is silently: the app is running, the models are loaded, and nothing is reaching it. An empty caption list looks exactly the same whether the lecturer is quiet, the output device is muted, or Windows is playing to a different device than the one being captured.
.\.venv\Scripts\python.exe scripts\check_audio.py --playThat plays a tone through the default output device and measures what the
capture path actually receives. It prints OK only if real audio arrived —
--selftest deliberately cannot tell you this, because it lists devices
rather than reading from them.
The header also carries a small input meter. If it moves, sound is reaching the pipeline. If it sits still while something is playing, the problem is the audio device, not the models.
Expect the first caption after about 5–10 seconds: the models warm up, and a segment is only flushed once the speaker pauses.
.
un.bat --selftest + onnxruntime version 1.30.0
~ Hexagon NPU CPU (QNN present but bound to CPU...)
-> Expected on any non-Snapdragon machine
+ live caption transport websockets
x speech recognition no Whisper weights found
-> python scripts/download_models.py --asr
~ translation captions will not be translated
-> python scripts/download_models.py --translate
+ end-to-end pipeline 2 captions in mock mode
3 ok, 5 degraded, 1 failed
Three states, and the distinction is deliberate: degraded means missing with a documented fallback, failed means the app will not work. A missing NPU is degraded — the product is designed to run without one. Missing Whisper weights are a failure, because there is no captioning without them. Every non-ok line says what to do about it.
Ask it:
.\run.bat --device Sahaay device report
----------------------------------------------
Execution provider : Hexagon NPU (QNN)
ONNX Runtime : 1.30.0
Architecture : ARM64 (ARM64)
Available providers: QNNExecutionProvider, CPUExecutionProvider
Hexagon NPU active : yes
The same information is a live badge in the UI. It reports the provider that ONNX Runtime actually bound to — not what the config asked for.
It will tell you "no" when the answer is no. On the x86 machine this was developed on, it prints:
Execution provider : CPU
Hexagon NPU active : no
Fallback reason : QNN present but bound to CPU, so there is no Hexagon NPU here
That distinction is not cosmetic, and finding it was the single most valuable hour of this build — see the hardware notes.
microphone ─┐
├─► 16 kHz mono ─► Silero VAD ─► segment on pause
speakers ───┘ (WASAPI) (CPU) │
(loopback) ▼
┌──────────────────┐
│ Whisper Small │ NPU
│ w8a16 QNN │
└────────┬─────────┘
│ caption
┌────────────────┼────────────────┐
▼ ▼
┌──────────────────┐ ┌──────────────────┐
│ NLLB-200 600M │ NPU │ Llama 3.2 3B │ NPU
│ INT8 │ │ glossary, low │
└────────┬─────────┘ │ priority │
│ └────────┬─────────┘
▼ ▼
translation jargon sidebar
└───────────┬───────────────────┘
▼
session notes + self-test (Markdown)
Details in docs/ARCHITECTURE.md.
| Stage | Model | Precision | Source |
|---|---|---|---|
| VAD | Silero VAD | fp32 | onnx-community/silero-vad |
| ASR | Whisper Small | w8a16 | qualcomm/Whisper-Small-Quantized — Qualcomm AI Hub, validated on Snapdragon X Elite |
| ASR (portable) | Whisper Small / Tiny | int8 / fp32 | onnx-community — the x86 fallback these tests ran on |
| Translation | NLLB-200 distilled 600M | int8 | Xenova/nllb-200-distilled-600M |
| Glossary + notes | Llama 3.2 3B Instruct | Hexagon assets | onnx-community/Llama-3.2-3B-instruct-hexagon-npu-assets |
| Glossary (portable) | Llama 3.2 3B / 1B Instruct | int4 | GENAI-ONNX builds — ship genai_config.json, which plain ONNX exports do not |
Which glossary model gets used depends on the hardware. Measured on x86 CPU: the 3B runs at 7.4 tok/s against the 1B's 16.9, and needs ~3.5 GB resident — under memory pressure one 146-token call took 35 minutes. So with the NPU active the 3B is preferred for its better explanations; on CPU the 1B is, because a glossary entry that arrives after the lecture has ended is not a glossary entry.
No weights are vendored. scripts/download_models.py --auto picks the Snapdragon or portable tier by architecture.
Every stage has a fallback, so the app is reviewable on any machine:
| If this is missing | What happens |
|---|---|
| Hexagon NPU | DirectML, then CPU. Same code path. |
onnxruntime-genai or the LLM |
Glossary uses a seeded STEM vocabulary + morphology filter |
| NLLB weights | Captions still work; the UI says translation is not installed |
| Silero VAD | Adaptive energy gating with a tracked noise floor |
| Everything | --mock runs the full pipeline on a scripted transcript |
The UI never claims a capability it does not have. A passthrough translation says so rather than showing English text under a Hindi heading.
Claims in a hackathon README are cheap, so here is exactly what was run, and on what.
The pipeline was tested end to end against real model weights and real speech with known ground truth — audio generated via Windows SAPI so the expected transcript is known in advance, rather than judged by ear.
| Stage | Result |
|---|---|
| Silero VAD → segmentation | 3 spoken sentences → exactly 3 segments, boundaries correct |
| Whisper → transcript | word-for-word correct on all three sentences |
| NLLB → Hindi/Tamil/Telugu/Malayalam | fluent output in all four |
| Term protection | eigenvalues, eigenvectors, determinant, SVD, backpropagation survive in Latin script inside a Devanagari sentence |
| Llama → glossary | real Hindi explanations of code-mixed terms, 24 tok/s on CPU int4 |
| Llama → notes + quiz | structured Topics / Key points, and real Q&A pairs |
| Full pipeline | audio in → captions → translation → glossary → saved notes |
| Accuracy | 0.0% WER on English, 14/14 technical terms kept — docs/ACCURACY.md |
| Concurrency | on CPU the glossary pushes captions past real time, RTF 1.55 — docs/CONCURRENCY.md |
| Local web server | real models load, EP badge live, WebSocket feed correct |
Example, straight out of the run:
CAPTION: So today we will start with eigenvalues and eigenvectors.
HI: तो आज हम eigenvalues और eigenvectors के साथ शुरू करेंगे।
kept: ['eigenvalues', 'eigenvectors']
The Whisper encoder has additionally been run on real Snapdragon hardware via Qualcomm's device farm — see the table above and docs/AIHUB.md.
What has not been verified: the full three-model pipeline end to end on a physical Snapdragon PC. Individual graphs were profiled there, but wall-clock numbers for the whole pipeline come from the x86 machine, and docs/BENCHMARKS.md says so on every row. Word error rate is measured, but against synthesised speech (docs/ACCURACY.md) — real-speaker accuracy needs a labelled code-mixed corpus and is still unknown. Both the 1B and 3B glossary models have now been run; the 3B writes visibly better explanations, and the app picks between them by hardware (see below).
These are measurements on physical Snapdragon hardware from Qualcomm's own device farm — not datasheet figures, not estimates. Every number has its AI Hub job ID. The job pages open after signing in with a Qualcomm ID; the full profiles are in docs/AIHUB.md.
| Whisper encoder | On-device | Peak memory | Layers on NPU |
|---|---|---|---|
| Snapdragon X Elite | 27.48 ms | 16.9 MB | 129 / 129 |
| Snapdragon X2 Elite | 13.5 ms | 9.2 MB | 129 / 129 |
| Snapdragon X Plus 8-Core | 26.71 ms | 16.2 MB | 129 / 129 |
129 of 129 layers ran on the Hexagon NPU — Qualcomm's profiler reports zero CPU fallback for the entire graph. Full table and job links in docs/AIHUB.md.
docs/BENCHMARKS.md covers the x86 development machine and the CPU fallback path, generated by scripts/bench.py on the machine that ran it — nothing typed in by hand. It reports p95 alongside the mean, because a captioner with a 200 ms mean and a 3 s p95 feels broken.
python scripts\bench.py --compare --write
python scripts\aihub_profile.py --all --writeThe primary user may be reading the lecture rather than hearing it, so this is treated as a feature area, not a checkbox:
- Caption text size is adjustable and remembered per viewer
- High-contrast dark theme by default, light theme honoured from the OS
- Captions are an
aria-liveregion so a screen reader announces new lines prefers-reduced-motiondisables the caption animation and smooth scroll- Space toggles start/stop; Escape closes the notes sheet
- The whole UI is keyboard reachable with visible focus rings
- The server binds to
127.0.0.1only — never the LAN - No telemetry, no analytics, no accounts, no API keys
- Audio is never written to disk; only the transcript you chose to save
- Sessions go to
sessions/as plain Markdown you can read, move or delete
tests/test_offline.py enforces this. It monkeypatches socket so any
connection or DNS lookup to anything but loopback raises, then runs a full
session — segmentation, ASR, translation, glossary, notes — and asserts
nothing tried. The guard self-tests first, because a guard that never fires
proves nothing. Nothing under sahaay/ may import an HTTP client at all.
$ pytest tests/test_offline.py
12 passed
powershell -ExecutionPolicy Bypass -File install.ps1 -Dev
.\.venv\Scripts\python.exe -m pytest # 400+ tests, no weights requiredsahaay/
runtime.py execution-provider selection — the Snapdragon-specific part
features.py Whisper log-mel front end, pure NumPy
audio.py WASAPI loopback + microphone capture
vad.py Silero VAD and pause-based segmentation
asr.py Whisper inference (AI Hub and Optimum packagings)
translate.py NLLB with technical-term protection
llm.py Llama via ORT GenAI, with a heuristic fallback
glossary.py the jargon sidebar worker
notes.py session notes, glossary and quiz
pipeline.py the three-thread orchestrator
server.py FastAPI + WebSocket
ui/ no framework, no build step
scripts/ benchmarks, profiling, accuracy, soak, and the site build
web/ the public site — generated from ui/ by scripts/build_web.py
docs/ every measurement, with its method and its caveats
MIT — see LICENSE. Model weights carry their own licences.
Built on Qualcomm AI Hub models, ONNX Runtime with the QNN execution provider, OpenAI Whisper, Meta's NLLB-200 and Llama 3.2, and Silero VAD.
The presentation kit - slide outline, three-minute demo script and a pre-recording checklist - is in docs/PRESENTATION.md.
