Transcription accuracy improvements - #69
Merged
Merged
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Four independent transcription fixes, plus the scaffolding a verbatim (CrisperWhisper) model needs — currently landed but commented out, since the export tooling moved to its own repo (
be27e36).1. Removed decoder-prefix conditioning (
lib/models.ts,workers/transcription.worker.ts)Whisper's
<|startofprev|>filler prompt is gone, along withwhisperFillerPrompt,MAX_VERBATIM_PROMPT_LENGTH,WHISPER_FILLER_PROMPTS,keepFillers, andbuildPromptedDecoderIds.The prompt was meant to give "Remove fillers" tokens to act on. Measured on an 11.5 s clip, decoding per VAD segment the way the worker actually does, Whisper Small:
The tail segment collapses into an echo of the prompt. Medium is worse; the length cap doesn't help (a 20-char prompt still triggers it). CrisperWhisper's trained-in
[verbatim_N]prefix fails the same way for the same reason. Nothing is lost: every checkpoint emits fillers unprompted, and anything swallowed still surfaces as a timed...placeholder.tests/vad-regression-test.tsnow asserts the worker contains nodecoder_input_ids; the full measurements live in a note aboveMODELS.2. Parakeet decoder quantisation (
workers/transcription.worker.ts)In a TDT model the decoder/joint network emits the tokens, so quantising it lands directly on word accuracy — and it's ~1% of the download (fp32 72 MB vs. 1239 MB for the fp16 encoder). parakeet.js defaults both nets to int8; the decoder is now set explicitly (fp32 on WebGPU, fp16 on WASM) while the encoder stays a pure size decision.
Also switched
preprocessorBackendfrom"js"to"onnx"—nemo128.onnxis NeMo's own featurisation graph (0.1 MB); the hand-written JS mel drifts from it and degrades every prediction.Sizes were understated: label corrected to "Parakeet v3", size
~700 MB→~1.3 GB.3. WASM-only on WebKit (
workers/transcription.worker.ts)Safari and every iOS browser kill the tab for memory far sooner than Chromium, and onnxruntime's WebGPU path is what pushes it over — the JSEP build is 26 MB of wasm vs. 13 MB, all compiled up front by JSC, then every weight is uploaded into Metal buffers during session creation.
navigator.vendorsniffing is deliberate: WebGPU is present and functional here, it's the memory ceiling around it that differs, and no API reports that.4.
@huggingface/transformerspatch via patch-package (patches/)WhisperTokenizer.decodeWithTimestampsandWhisperTimeStampLogitsProcessortest only the lower bound of the timestamp block, so every token above it is mistaken for a timestamp. Harmless for stock Whisper (nothing sits above the block); breaks any vocabulary-extending derivative:[UM]leaves an empty bucket that reachesdecode([])→token_ids must be a non-empty array of integers, thrown mid-transcript.[UH]as a timestamp andfill(-Infinity)s every text token — transcription stops at the first hesitation with no error, the rest coming back as...placeholders.The patch range-checks both bounds (matching what
_decode_asralready does), uses thetimestamp_begingetter, and skips empty buckets — that last one is a latent crash on stock Whisper too. Worth upstreaming;patches/README.mdhas the full writeup.@huggingface/transformersis pinned to exact4.2.0, not a caret range: patch-package only warns on version mismatch, and the unpatched failure mode is a silent truncation.5. Diagnostics
A segment decoding to zero words is indistinguishable from silence downstream, so the worker now warns with the offset, chunk count, and raw text — enough to tell "ASR returned nothing" from "words were produced then dropped in post-processing".
6. UI
components/SignalBars.tsx: three-bar strength icons for model rows. lucide'sSignalLowomits bars above the level, giving each row a differently-shaped glyph; these always draw all three and fade the inactive ones.ModelSelectoricons are now resolved per source (iconForSource), andOptionTrigger.iconwidened fromLucideIconto a localIconComponent.