Skip to content

CTC forced alignment for word timings (supersedes #61, #65) - #70

Merged
wassgha merged 10 commits into
mainfrom
forced-alignment
Aug 6, 2026
Merged

wassgha merged 10 commits into
mainfrom
forced-alignment

Conversation

@wassgha

@wassgha wassgha commented Aug 6, 2026

Copy link
Copy Markdown
Owner

CTC forced alignment for word timings

Replaces Whisper's DTW-derived word boundaries (and Parakeet's TDT ones) with boundaries measured against the audio: a language-specific CTC acoustic model, Viterbi over the standard CTC lattice, and an envelope pass to widen the peaky spans that CTC returns.

Supersedes #61 and #65. This is #65's branch merged onto current main, plus a fix for two bugs in its multilingual layer. #61 should be closed unmerged — its commits are all contained here, and on its own it aligns every language against an English-only wav2vec2 (stripping accents to nothing and Han characters entirely) and skips Parakeet.

What lands

lib/forcedAlign.ts — the algorithm, pure and model-free so it's testable without loading weights. Viterbi over blank t0 blank t1 … blank, word spans read off the path, unspellable words (digits, bare punctuation) interpolated from their neighbours, batching at the widest available pause so the frames × tokens lattice stays bounded.

lib/alignModels.ts — language → aligner map. English keeps the small Librispeech wav2vec2 (~86 MB q4); es/fr/de share the MMS forced-aligner; zh uses an XLS-R CTC model with Han in its vocab (~240 MB q4 each). A language with no entry skips CTC and keeps the heuristic path.

lib/align.ts — new work that needs no model at all and stands on its own:

  • speechEnvelope — loudness sampled every 5 ms, as the max of a broadband and a >2 kHz envelope each normalised against its own 95th percentile. The high band is what lets an unvoiced fricative register: broadband alone puts the onset of "shot" on the vowel, ~0.2 s late.
  • estimateLagFromEnvelope — estimates the global shift from loudness rises rather than the VAD mask, which goes flat (and silently returns ~0) on speech that never pauses.
  • refineOnsets — pulls each VAD onset back off the vowel onto the sound that actually starts the word.
  • repairCollapsedWords — re-splits words Whisper's DTW collapsed to a sliver and handed to a neighbour. On the test clip "evening" spanned 0.06 s while the preceding "and" took 0.40 s.

ALIGN_LEAD_S, the fixed 80 ms perceptual nudge, is removed — see the open question below.

workers/transcription.worker.ts — the aligner gets its own ModelManager rather than joining the ASR registry. The deciding factor is lifecycle: the ASR registry runs under a WebGPU→WASM fallback policy whose recovery step is unloadAll(), and the aligner never asks for a device, so a GPU loss would otherwise throw away perfectly good WASM weights. Whisper and Parakeet share one post-ASR timing pass; disfluency placeholders run after alignment, since ... has no CTC spelling.

The two bugs fixed on top (b207e41)

Both derivations in alignerModel() were correct on two of the three shipped aligners and silently wrong on the third. Neither raises; both just cost timing accuracy on es/fr/de.

Blank. blankId came from pad_token_id. For wav2vec2-base-960h and XLS-R-zh, <pad> is the CTC blank and sits at 0, so that was accidentally right. The MMS forced-aligner declares them separately — <blank> at 0, <pad> at 1 — so MMS scored the lattice's blank states against a column the model never activates. Viterbi doesn't fail there; it avoids blank states, which smears every boundary.

Delimiter. delimiterId came from encode("|"), which cannot detect absence: a vocab without | maps it to <unk>, one id, indistinguishable from a real hit. MMS has no |, so <unk> was threaded between every pair of words. The old comment ("MMS has none") described the intent correctly; the check didn't implement it.

Both now go through convert_tokens_to_ids, which yields undefined rather than <unk> for an absent token. Moved out of the worker into ctcVocabFromTokenizer() so it's unit-testable.

Verification

tsc --noEmit, eslint, and all 17 test files pass.

Checked against the live tokenizers, all five languages:

en  blank=0  delim=4          "HELLOTHERE"      unk-hits=0/10
es  blank=0  delim=undefined  "cancioncafe"     unk-hits=0/11
fr  blank=0  delim=undefined  "ouetesvous"      unk-hits=0/10
de  blank=0  delim=undefined  "gruesseschoen"   unk-hits=0/13
zh  blank=0  delim=4          "你好世界"           unk-hits=0/4

Before the fix, es/fr/de resolved blank=1 and delim=3. Eight new assertions in tests/forced-align-test.ts pin this using fakes built from the ids the real tokenizers report.

Not verified: vad-regression-test.ts skips its audio-coverage check when ffmpeg isn't on PATH, which was the case in my environment. Worth a run somewhere with ffmpeg before merge.

Open questions

  • ALIGN_LEAD_S removal. The 80 ms nudge existed because a highlight landing early reads as in-sync while one landing late reads as lagging. CTC measures properly so the nudge shouldn't be needed — but forceAlign falls back to the heuristic path on any failure (cold-cache error, failed batch, unsupported language), and that fallback now has no lead. Worth deciding whether to keep the nudge on the fallback path only.
  • Download and memory cost lands hardest on WebKit. +86 MB (en) or +240 MB (multilingual) on top of Parakeet's 1.3 GB, plus a full-audio CTC forward pass. Transcription accuracy improvements #69 just forced Safari and all iOS browsers onto WASM specifically because of memory-driven tab reloads; this pushes against that ceiling. The aligner is already WASM-only so there's no direct conflict, but it's more pressure on the platform we just hardened.
  • ROADMAP.md:33 already marks "Better word↔audio alignment (speech-onset anchoring / lag correction)" as done. This is the tier beyond that — the entry probably wants splitting or updating.

🤖 Generated with Claude Code

wassgha and others added 9 commits August 2, 2026 11:09
# Conflicts:
#	workers/transcription.worker.ts
The CTC aligner loaded itself with bare from_pretrained calls, so its 86 MB
download happened behind a static "Aligning words…" with no bar — indistinguishable
from a hang on a cold cache.

Register it in the existing ModelManager instead, so it shares the download
progress, cache labeling and unload-on-GPU-loss path with the ASR models. It
needs the processor, model and tokenizer separately rather than a pipeline(),
so it is a plain ModelDefinition wired to transformersProgress rather than
transformersModel().

The aligner warms in the background while Whisper runs, so its bytes must not
fight the "Transcribing…" line: speech models always own the progress line, and
the aligner only claims it once transcription is done and we are actually
waiting on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	workers/transcription.worker.ts
wav2vec2 is not in the same family as the ASR models — it is an acoustic
aligner, not a transcriber, and not something the user picks. But the deciding
factor is lifecycle, not taxonomy.

The ASR registry exists under a WebGPU→WASM fallback policy whose recovery step
is unloadAll(). The aligner never asks for a device, so transformers.js runs it
on WASM (DEFAULT_DEVICE), where a lost GPU cannot touch it. Sharing the registry
meant a GPU loss threw away 86 MB of perfectly good weights and forced a
needless session rebuild.

Splitting also removes the id special-casing the shared registry forced into the
progress subscription. The ASR handler goes back to "whatever is loading", and
the aligner's own subscription is scoped to the await in forceAlign — the
subscription's lifetime is the gate, so there is no module-level flag left to
fall out of step if that path is ever re-entered.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
subscribe() only fires on the next change, so a load that had started but not
yet received its first byte left the UI on the previous message.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Share the VAD → CTC → expand post-pass between Whisper and Parakeet, and
pick a language-specific aligner (English wav2vec2, MMS for es/fr/de, XLS-R
for zh) so non-English runs no longer warm an English-only model.

Co-authored-by: Wassim Gharbi <wassgha@gmail.com>
…rced-alignment

Brings CTC forced alignment (PR #65) onto main. Supersedes PR #61, whose
commits are all contained here — #61 aligns every language against an
English-only wav2vec2 and skips Parakeet entirely.

One conflict, in workers/transcription.worker.ts, and purely additive: main
added the CrisperWhisper token block where the branch added the aligner
registry. Both kept, in that order.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both derivations in alignerModel() took a shortcut that is correct on two of
the three shipped aligners and silently wrong on the third. Neither failure
raises; both just cost timing accuracy on es/fr/de.

blankId came from `pad_token_id`. For wav2vec2-base-960h and XLS-R-zh, <pad>
*is* the CTC blank and sits at 0, so that happened to be right. The MMS
forced-aligner declares them separately — <blank> at 0, <pad> at 1 — so MMS
scored the lattice's blank states against a column the model never activates.
Viterbi does not fail there, it avoids blank states, which smears every
boundary it returns.

delimiterId came from `encode("|")`, which cannot detect absence: a vocab
without "|" maps it to <unk>, one id, indistinguishable from a real hit. MMS
has no "|", so <unk> was threaded between every pair of words. The old comment
("MMS has none") described the intent correctly; the check did not implement
it.

Both now go through convert_tokens_to_ids, which yields undefined rather than
<unk> for a token the vocab lacks. Moved out of the worker into
lib/forcedAlign.ts as ctcVocabFromTokenizer() so it is unit-testable without
loading a model; the new tests use fakes built from the ids the three real
tokenizers actually report.

Verified against the live tokenizers — all five languages now resolve blank 0,
delimiter 4 for en/zh and none for es/fr/de, and no sample text falls through
to <unk>. Before: es/fr/de got blank 1 and delimiter 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
app.rescript Ready Ready Preview Aug 6, 2026 2:44am

Words read consistently a hair late even with CTC alignment placing the
boundaries. Measurement fixed the part of the old ALIGN_LEAD_S that was
compensating for a late reference — Silero raises its speech flag ~40 ms after
speech starts — but not the reason the lead existed in the first place: a
highlight landing slightly early reads as in sync, one landing slightly late
reads as lagging, and the two are not equally forgiving.

So it comes back, as applyAlignLead() in lib/align.ts, applied in
refineWordTimestamps() where Whisper and Parakeet, CTC and heuristic fallback,
all converge.

Applied last, after insertDisfluencyPlaceholders, for two reasons. A uniform
shift over the finished list preserves every relative gap, so the "..."
placeholders travel with the words around them instead of drifting against
them. And nothing downstream can snap the nudge back onto a VAD onset, which is
what the old placement inside alignWordsToSpeech had to work around.

Both edges move together; leading only the starts would stretch every word and
eventually overlap its neighbour. Starts clamp at 0.

The tests bound the constant to (0, 0.15] rather than asserting a value — it is
a perceptual setting tuned by ear, and pinning it would make retuning look like
a regression.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@wassgha
wassgha merged commit 9d0deb1 into main Aug 6, 2026
2 of 3 checks passed

This branch was successfully deployed

1 active deployment
Preview — 3899aea0 Deployed Aug 6, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants