fix(stt): start each transcript word at the end of the token before it - #950
Conversation
whisper.cpp's t_dtw marks where a token ends, so taking a word's first-token t_dtw as its start put every word one token late (+175 ms median). The helper now starts a word at the previous token's t_dtw, ends it at its own last token's, and sends the first-token time as `anchor` to tell which stretch of speech owns it. The RMS lookback snap goes; phrase edges are anchored on the VAD in both directions. Adds the word-timing harness used to measure it (issue #948).
|
Warning Review limit reachedYou've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Next included review available in 2 minutes. View limit detailsLimit details: You’ve used all 8 included reviews currently available. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (18)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
@coderabbitai review |
|
Summary
Transcript word boundaries move from ~105 ms late to ~30 ms, with no download. This is Phase 1 of #948.
Root cause. whisper.cpp's
t_dtwmarks where a token ends. The helper used a word's first-tokent_dtwas its start, so every word started one token late.Changes
whisper-stt/src/main.cpp: a word starts at thet_dtwof the token before it and ends at its own last token's. The previous token carries across segments. It also emitsanchor, the first-token time, which always lies inside the word.electron/stt/snapWordBoundaries.ts: the RMS lookback snap is gone; kept on top of the fix, it drags correct boundaries 120 ms early. Phrase edges are anchored on the VAD in both directions, and a stretch owns a word by itsanchor.MAX_ANCHOR_SEC, the tail rule and punctuation collapsing are unchanged.anchorstays off the IPC contract.tools/stt-eval/word-timing/: the harness from Transcript word timings: measure them, then make them precise enough for text-based cuts #948, dependency-free, with a README.+15 ms offset (config F): not kept. It gains 4 points of inner boundaries within 50 ms but drops noisy phrase deletes to 82%, under the 84% bar.
Results (harness, Vulkan,
ggml-small-q8_0, times in ms)real-check.mjs). The first-token time sits a median 200 ms after the new start. After the post-pass, all 9 phrase starts land on the VAD onset; the raw starts were at −8.8 s to +0.58 s.Related issue
Part of #948 (Phase 1 of 4).
Type of change
Release impact
Desktop impact
Screenshots / video
None: no UI change.
Testing
run-helper.mjsthenevaluate.mjson the 96-clip corpus, with a local Windows Vulkan build of the helper. The baseline is reproduced exactly byevaluate.mjs vk-stock --snap <main's snapWordBoundaries.ts>.vitest --run electron/stt(87 pass).tsc --noEmitand-p tsconfig.test.jsonreport no new errors. Biome is clean.build-whisper-stt.ymlbuilds the helper on macOS, Linux and Windows.🤖 Generated with Claude Code