Skip to content

Transcription accuracy improvements - #69

Merged
wassgha merged 9 commits into
mainfrom
transcription-improvements
Aug 5, 2026
Merged

wassgha merged 9 commits into
mainfrom
transcription-improvements

Conversation

@wassgha

@wassgha wassgha commented Aug 5, 2026 •

Copy link
Copy Markdown
Owner

Summary

Four independent transcription fixes, plus the scaffolding a verbatim (CrisperWhisper) model needs — currently landed but commented out, since the export tooling moved to its own repo (be27e36).

1. Removed decoder-prefix conditioning (lib/models.ts, workers/transcription.worker.ts)

Whisper's <|startofprev|> filler prompt is gone, along with whisperFillerPrompt, MAX_VERBATIM_PROMPT_LENGTH, WHISPER_FILLER_PROMPTS, keepFillers, and buildPromptedDecoderIds.

The prompt was meant to give "Remove fillers" tokens to act on. Measured on an 11.5 s clip, decoding per VAD segment the way the worker actually does, Whisper Small:

slice plain + prompt
full 11.5 s complete, includes "uh" "Nice. How does it, uh,"
2.5–5.0 s "Nice. How does it work?" "Nice. How does it"
5.0–11.5 s complete sentence "Um,"

The tail segment collapses into an echo of the prompt. Medium is worse; the length cap doesn't help (a 20-char prompt still triggers it). CrisperWhisper's trained-in [verbatim_N] prefix fails the same way for the same reason. Nothing is lost: every checkpoint emits fillers unprompted, and anything swallowed still surfaces as a timed ... placeholder.

tests/vad-regression-test.ts now asserts the worker contains no decoder_input_ids; the full measurements live in a note above MODELS.

2. Parakeet decoder quantisation (workers/transcription.worker.ts)

In a TDT model the decoder/joint network emits the tokens, so quantising it lands directly on word accuracy — and it's ~1% of the download (fp32 72 MB vs. 1239 MB for the fp16 encoder). parakeet.js defaults both nets to int8; the decoder is now set explicitly (fp32 on WebGPU, fp16 on WASM) while the encoder stays a pure size decision.

Also switched preprocessorBackend from "js" to "onnx" — nemo128.onnx is NeMo's own featurisation graph (0.1 MB); the hand-written JS mel drifts from it and degrades every prediction.

Sizes were understated: label corrected to "Parakeet v3", size ~700 MB → ~1.3 GB.

3. WASM-only on WebKit (workers/transcription.worker.ts)

Safari and every iOS browser kill the tab for memory far sooner than Chromium, and onnxruntime's WebGPU path is what pushes it over — the JSEP build is 26 MB of wasm vs. 13 MB, all compiled up front by JSC, then every weight is uploaded into Metal buffers during session creation. navigator.vendor sniffing is deliberate: WebGPU is present and functional here, it's the memory ceiling around it that differs, and no API reports that.

4. @huggingface/transformers patch via patch-package (patches/)

WhisperTokenizer.decodeWithTimestamps and WhisperTimeStampLogitsProcessor test only the lower bound of the timestamp block, so every token above it is mistaken for a timestamp. Harmless for stock Whisper (nothing sits above the block); breaks any vocabulary-extending derivative:

  • Decoding: splitting on [UM] leaves an empty bucket that reaches decode([]) → token_ids must be a non-empty array of integers, thrown mid-transcript.
  • Generation: the processor sees [UH] as a timestamp and fill(-Infinity)s every text token — transcription stops at the first hesitation with no error, the rest coming back as ... placeholders.

The patch range-checks both bounds (matching what _decode_asr already does), uses the timestamp_begin getter, and skips empty buckets — that last one is a latent crash on stock Whisper too. Worth upstreaming; patches/README.md has the full writeup.

@huggingface/transformers is pinned to exact 4.2.0, not a caret range: patch-package only warns on version mismatch, and the unpatched failure mode is a silent truncation.

5. Diagnostics

A segment decoding to zero words is indistinguishable from silence downstream, so the worker now warns with the offset, chunk count, and raw text — enough to tell "ASR returned nothing" from "words were produced then dropped in post-processing".

6. UI

  • New components/SignalBars.tsx: three-bar strength icons for model rows. lucide's SignalLow omits bars above the level, giving each row a differently-shaped glyph; these always draw all three and fade the inactive ones.
  • ModelSelector icons are now resolved per source (iconForSource), and OptionTrigger.icon widened from LucideIcon to a local IconComponent.

@vercel

vercel Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
app.rescript Ready Ready Preview Aug 5, 2026 9:54am

@wassgha wassgha changed the title Transcription improvements Transcription accuracy improvements Aug 5, 2026
@wassgha
wassgha merged commit 778467a into main Aug 5, 2026
3 checks passed

This branch was successfully deployed

1 active deployment
Preview — daac3888 Deployed Aug 5, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant