CTC forced alignment for word timings (supersedes #61, #65) - #70
Merged
Merged
Conversation
# Conflicts: # workers/transcription.worker.ts
The CTC aligner loaded itself with bare from_pretrained calls, so its 86 MB download happened behind a static "Aligning words…" with no bar — indistinguishable from a hang on a cold cache. Register it in the existing ModelManager instead, so it shares the download progress, cache labeling and unload-on-GPU-loss path with the ASR models. It needs the processor, model and tokenizer separately rather than a pipeline(), so it is a plain ModelDefinition wired to transformersProgress rather than transformersModel(). The aligner warms in the background while Whisper runs, so its bytes must not fight the "Transcribing…" line: speech models always own the progress line, and the aligner only claims it once transcription is done and we are actually waiting on it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts: # workers/transcription.worker.ts
wav2vec2 is not in the same family as the ASR models — it is an acoustic aligner, not a transcriber, and not something the user picks. But the deciding factor is lifecycle, not taxonomy. The ASR registry exists under a WebGPU→WASM fallback policy whose recovery step is unloadAll(). The aligner never asks for a device, so transformers.js runs it on WASM (DEFAULT_DEVICE), where a lost GPU cannot touch it. Sharing the registry meant a GPU loss threw away 86 MB of perfectly good weights and forced a needless session rebuild. Splitting also removes the id special-casing the shared registry forced into the progress subscription. The ASR handler goes back to "whatever is loading", and the aligner's own subscription is scoped to the await in forceAlign — the subscription's lifetime is the gate, so there is no module-level flag left to fall out of step if that path is ever re-entered. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
subscribe() only fires on the next change, so a load that had started but not yet received its first byte left the UI on the previous message. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Share the VAD → CTC → expand post-pass between Whisper and Parakeet, and pick a language-specific aligner (English wav2vec2, MMS for es/fr/de, XLS-R for zh) so non-English runs no longer warm an English-only model. Co-authored-by: Wassim Gharbi <wassgha@gmail.com>
…rced-alignment Brings CTC forced alignment (PR #65) onto main. Supersedes PR #61, whose commits are all contained here — #61 aligns every language against an English-only wav2vec2 and skips Parakeet entirely. One conflict, in workers/transcription.worker.ts, and purely additive: main added the CrisperWhisper token block where the branch added the aligner registry. Both kept, in that order. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both derivations in alignerModel() took a shortcut that is correct on two of
the three shipped aligners and silently wrong on the third. Neither failure
raises; both just cost timing accuracy on es/fr/de.
blankId came from `pad_token_id`. For wav2vec2-base-960h and XLS-R-zh, <pad>
*is* the CTC blank and sits at 0, so that happened to be right. The MMS
forced-aligner declares them separately — <blank> at 0, <pad> at 1 — so MMS
scored the lattice's blank states against a column the model never activates.
Viterbi does not fail there, it avoids blank states, which smears every
boundary it returns.
delimiterId came from `encode("|")`, which cannot detect absence: a vocab
without "|" maps it to <unk>, one id, indistinguishable from a real hit. MMS
has no "|", so <unk> was threaded between every pair of words. The old comment
("MMS has none") described the intent correctly; the check did not implement
it.
Both now go through convert_tokens_to_ids, which yields undefined rather than
<unk> for a token the vocab lacks. Moved out of the worker into
lib/forcedAlign.ts as ctcVocabFromTokenizer() so it is unit-testable without
loading a model; the new tests use fakes built from the ids the three real
tokenizers actually report.
Verified against the live tokenizers — all five languages now resolve blank 0,
delimiter 4 for en/zh and none for es/fr/de, and no sample text falls through
to <unk>. Before: es/fr/de got blank 1 and delimiter 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Words read consistently a hair late even with CTC alignment placing the boundaries. Measurement fixed the part of the old ALIGN_LEAD_S that was compensating for a late reference — Silero raises its speech flag ~40 ms after speech starts — but not the reason the lead existed in the first place: a highlight landing slightly early reads as in sync, one landing slightly late reads as lagging, and the two are not equally forgiving. So it comes back, as applyAlignLead() in lib/align.ts, applied in refineWordTimestamps() where Whisper and Parakeet, CTC and heuristic fallback, all converge. Applied last, after insertDisfluencyPlaceholders, for two reasons. A uniform shift over the finished list preserves every relative gap, so the "..." placeholders travel with the words around them instead of drifting against them. And nothing downstream can snap the nudge back onto a VAD onset, which is what the old placement inside alignWordsToSpeech had to work around. Both edges move together; leading only the starts would stretch every word and eventually overlap its neighbour. Starts clamp at 0. The tests bound the constant to (0, 0.15] rather than asserting a value — it is a perceptual setting tuned by ear, and pinning it would make retuning look like a regression. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
CTC forced alignment for word timings
Replaces Whisper's DTW-derived word boundaries (and Parakeet's TDT ones) with boundaries measured against the audio: a language-specific CTC acoustic model, Viterbi over the standard CTC lattice, and an envelope pass to widen the peaky spans that CTC returns.
Supersedes #61 and #65. This is #65's branch merged onto current
main, plus a fix for two bugs in its multilingual layer. #61 should be closed unmerged — its commits are all contained here, and on its own it aligns every language against an English-only wav2vec2 (stripping accents to nothing and Han characters entirely) and skips Parakeet.What lands
lib/forcedAlign.ts— the algorithm, pure and model-free so it's testable without loading weights. Viterbi overblank t0 blank t1 … blank, word spans read off the path, unspellable words (digits, bare punctuation) interpolated from their neighbours, batching at the widest available pause so the frames × tokens lattice stays bounded.lib/alignModels.ts— language → aligner map. English keeps the small Librispeech wav2vec2 (~86 MB q4); es/fr/de share the MMS forced-aligner; zh uses an XLS-R CTC model with Han in its vocab (~240 MB q4 each). A language with no entry skips CTC and keeps the heuristic path.lib/align.ts— new work that needs no model at all and stands on its own:speechEnvelope— loudness sampled every 5 ms, as the max of a broadband and a >2 kHz envelope each normalised against its own 95th percentile. The high band is what lets an unvoiced fricative register: broadband alone puts the onset of "shot" on the vowel, ~0.2 s late.estimateLagFromEnvelope— estimates the global shift from loudness rises rather than the VAD mask, which goes flat (and silently returns ~0) on speech that never pauses.refineOnsets— pulls each VAD onset back off the vowel onto the sound that actually starts the word.repairCollapsedWords— re-splits words Whisper's DTW collapsed to a sliver and handed to a neighbour. On the test clip "evening" spanned 0.06 s while the preceding "and" took 0.40 s.ALIGN_LEAD_S, the fixed 80 ms perceptual nudge, is removed — see the open question below.workers/transcription.worker.ts— the aligner gets its ownModelManagerrather than joining the ASR registry. The deciding factor is lifecycle: the ASR registry runs under a WebGPU→WASM fallback policy whose recovery step isunloadAll(), and the aligner never asks for a device, so a GPU loss would otherwise throw away perfectly good WASM weights. Whisper and Parakeet share one post-ASR timing pass; disfluency placeholders run after alignment, since...has no CTC spelling.The two bugs fixed on top (
b207e41)Both derivations in
alignerModel()were correct on two of the three shipped aligners and silently wrong on the third. Neither raises; both just cost timing accuracy on es/fr/de.Blank.
blankIdcame frompad_token_id. For wav2vec2-base-960h and XLS-R-zh,<pad>is the CTC blank and sits at 0, so that was accidentally right. The MMS forced-aligner declares them separately —<blank>at 0,<pad>at 1 — so MMS scored the lattice's blank states against a column the model never activates. Viterbi doesn't fail there; it avoids blank states, which smears every boundary.Delimiter.
delimiterIdcame fromencode("|"), which cannot detect absence: a vocab without|maps it to<unk>, one id, indistinguishable from a real hit. MMS has no|, so<unk>was threaded between every pair of words. The old comment ("MMS has none") described the intent correctly; the check didn't implement it.Both now go through
convert_tokens_to_ids, which yieldsundefinedrather than<unk>for an absent token. Moved out of the worker intoctcVocabFromTokenizer()so it's unit-testable.Verification
tsc --noEmit,eslint, and all 17 test files pass.Checked against the live tokenizers, all five languages:
Before the fix, es/fr/de resolved
blank=1anddelim=3. Eight new assertions intests/forced-align-test.tspin this using fakes built from the ids the real tokenizers report.Not verified:
vad-regression-test.tsskips its audio-coverage check when ffmpeg isn't on PATH, which was the case in my environment. Worth a run somewhere with ffmpeg before merge.Open questions
ALIGN_LEAD_Sremoval. The 80 ms nudge existed because a highlight landing early reads as in-sync while one landing late reads as lagging. CTC measures properly so the nudge shouldn't be needed — butforceAlignfalls back to the heuristic path on any failure (cold-cache error, failed batch, unsupported language), and that fallback now has no lead. Worth deciding whether to keep the nudge on the fallback path only.ROADMAP.md:33already marks "Better word↔audio alignment (speech-onset anchoring / lag correction)" as done. This is the tier beyond that — the entry probably wants splitting or updating.🤖 Generated with Claude Code