Transcribe podcast episodes from RSS feeds using the whisper.cpp HTTP server and index them into a local SQLite FTS5 database for full-text search.
Walks one or more RSS feeds, downloads each episode's MP3, and POSTs it to a running whisper.cpp HTTP server. The server decodes and resamples the audio itself, so no local transcoding step is needed and the MP3 is what gets kept on disk. It returns verbose_json, from which the pipeline renders a WebVTT transcript into transcript.json and writes the per-word timings alongside as words.jsonl.gz. Designed to run on a schedule (e.g. cron) to keep transcripts up to date.
A Quarkus HTTP server that serves a full-text search interface over the SQLite database produced by index.py. Supports filtering by podcast, collection, episode type, and duration. Search results are ranked by BM25 relevance. The episode detail page shows a clickable transcript synced to the audio player. Opened from a search result it carries the query along (/episode/{id}?q=…), highlights every cue the query matches, scrolls to the first, and steps between them with Prev/Next or the n/p keys. Runs on port 8080 by default.
The FTS index is one row per episode, so SQLite reports which episodes match but not where. The episode page re-applies the query cue by cue, reading it the way FTS5 does: quoted phrases stay whole, AND/OR/NOT/NEAR are operators only in upper case, a trailing * is a prefix, and words are compared case-insensitively with diacritics removed, as the unicode61 tokenizer indexed them.
Note: whisper-server also defaults to 8080. They are rarely up at the same time, but if they are, move one —
--porton whisper-server,quarkus.http.porton the web module.
Reads the transcript.json files in the data directory and writes them into a SQLite FTS5 database. Run this after the pipeline to make new transcripts searchable. After the first run it reads only what has changed, so running it after every pipeline run costs seconds rather than minutes.
Requires Python 3 with an FTS5-capable SQLite. The script prefers pysqlite3 when installed and falls back to the standard library sqlite3 module otherwise, exiting with a clear error if neither has FTS5:
pip install pysqlite3Note: The default system
sqlite3on some platforms (e.g. Synology NAS) does not include FTS5.pysqlite3bundles a SQLite build that does.
- JDK 21+
- A running whisper.cpp HTTP server
- Python 3 with an FTS5-capable SQLite (for the indexing script;
pip install pysqlite3if the system build lacks FTS5)
No ffmpeg is required. MP3s are uploaded as-is and the whisper.cpp server decodes them with its
built-in miniaudio decoder, resampling to 16 kHz mono internally.
./gradlew buildThe pipeline POSTs audio to the whisper.cpp /inference endpoint. It does not
start or manage the server — start it yourself and leave it up for the run:
whisper-server \
-m /path/to/models/ggml-large-v3.bin \
--port 8080The whisper-server binary is built alongside whisper-cli when you compile
whisper.cpp. Set PIPELINE_WHISPER_SERVER_URL in pipeline/.env to its base
URL.
Only the model is chosen at launch. Everything else that matters is a
per-request form field, and server.cpp overrides launch defaults with
whatever a request actually sends. So the pipeline's own settings win, and the
only knob you pick when starting the server is -m.
Do not pass --convert (it shells out to ffmpeg on the server host — MP3s are
uploaded as-is and decoded internally), and do not pass -nt, which would
suppress the timestamps the pipeline exists to capture.
The pipeline is run against whisper.cpp v1.9.4 (927cfce3), serving ggml-large-v3 from two
builds: Metal with the CoreML encoder on Apple Silicon, and CUDA on Linux. This is not a requirement.
It is the version the repair's retries were measured on. The corpus was first transcribed, and
repaired once, on v1.9.2 (306c88f4).
v1.9.4 re-seeds the sampler on every request
(whisper.cpp#4025), so a server gives the same
output for the same request, even when it falls back to sampling. Older servers drift between calls,
which is what a repeated request used to get its second chance from. So --repair-windows tries each
window prompted and unprompted, then both again at temperature 0.2, then both again with a beam of 8:
each retry asks for something different.
whisper.cpp uses its own GGML model format (.bin files), not the OpenAI
Python .pt files. Download from HuggingFace:
curl -L -o ggml-large-v3.bin \
https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3.binPut the matching ggml-large-v3-encoder.mlmodelc next to the .bin, or the
encoder silently falls off the Neural Engine and everything gets much slower.
Turbo prunes the decoder and keeps the encoder, and under Core ML the encoder is where most of the wall clock goes. Measured over 13 episodes, large-v3 costs 1.36x turbo's decode time — not the 2-3x that "turbo for speed" implies — and produced a usable transcript on 10 of 13 episodes turbo could not decode at all.
It also fixes a failure nothing else touches. Turbo shreds some episodes into runs of one- and two-word cues: 1,256 episodes in one corpus. An A/B on 4 shredded episodes across 4 arms took fragmentation to zero in all sixteen cells, including the no-prompt no-VAD control. That is a model difference, and neither the prompt nor VAD substitutes for it.
These are the request fields in Transcriber.kt. Each one is load-bearing.
Whisper intermittently decodes an entire episode with no punctuation and no capitals. That is not cosmetic: whisper segments on sentence structure, so with no full stops the cue boundaries stop tracking speech and every timestamp derived from them is fiction. 654 episodes were hit, and 13 resisted every attempt to re-decode them.
An initial_prompt of ordinary punctuated prose fixed all 13:
| approach | fixed |
|---|---|
| initial_prompt | 13 / 13 |
| large-v3 | 10 / 13 |
| best VAD parameter | 9 / 13 |
| Silero VAD v6.2.0 | 6 / 13 |
| re-decode at another VAD threshold | 0 / 13 |
It works because this is a decoder mode, not a segmentation problem. VAD only changes what audio reaches the decoder; a prompt conditions the decoder itself, and punctuation is a style. It also prevents runaway repetition loops — without it, 4 of 4 test decodes collapsed, one returning 539 usable words out of 10,660, and ran 2.5x slower because a looping decode burns tokens.
carry_initial_prompt=true matters: without it only the first window is
conditioned, and an episode that degrades later still degrades (13/13 with,
12/13 without).
Keep the prompt generic. It biases vocabulary as well as style, so anything
domain-specific contaminates transcripts. The default in Transcriber.kt was
checked against a repaired episode — zero occurrences of any prompt fragment,
word count within 5% of the original.
The prompt only rides with the language it is written in. The default is
English prose, so a podcast set to another language decodes without it — an
English prompt on French audio is the same vocabulary contamination as a
domain-specific one, and carry_initial_prompt would apply it to every window.
Under language: auto it is worse: the prompt would skew whisper's own language
detection toward English before it decoded anything, breaking the very thing
auto is for — so auto never takes a prompt, however explicitly one is set.
A feed in another language keeps the lever by supplying its own prompt:
- name: Une Emission Francaise
url: https://example.com/fr.rss
language: fr
initial_prompt: Bonjour, et bienvenue dans cette emission. Aujourd'hui, nous allons parler de plusieurs choses.The prompt's language is the podcast's language, so there is no second key to
keep in sync — and a podcast that sets initial_prompt must set language too,
or the run refuses to start rather than guessing. A top-level initial_prompt
works the same way against the top-level language.
The built-in English prompt stays bound to English, whatever the top-level
language is. Only a prompt you actually set belongs to that language; an unset
one is not "the default, in French".
Keep any prompt generic, and check a new one the way the English default was checked: decode an episode with it and confirm no prompt fragment appears in the transcript and the word count has not moved.
Not merely omitted. A request that omits vad inherits whatever the server was
launched with, silently, which is how one corpus ended up with no record of its
own VAD state.
It has to be off because VAD breaks word timestamps. whisper.cpp keeps token timestamps in VAD-compressed time while remapping segment timestamps to real time, so the two drift apart by however much silence VAD removed:
| offset, first word vs its own segment | |
|---|---|
| VAD on | −1.79s at the start, −6.51s by the end |
| VAD off | +0.00s to +0.20s |
Same episode, same model, same fields — on a six-minute episode. Every word time would be early by a growing, episode-dependent, invisible amount.
Nothing is lost by turning it off. VAD's real value was suppressing the unpunctuated collapse, and the prompt does that better (13/13 against 9/13).
The server can return WebVTT directly. It is derived from the JSON here instead,
because per-word start/end come only from verbose_json and the decode
computes them either way — rendering to VTT throws them away.
Both artifacts come from one parse of one response. Asking for both formats would mean two decodes, and whisper is not deterministic across runs, so their cues and words could disagree in ways nothing downstream could detect.
Output per episode:
transcript.json metadata plus the WebVTT string
words.jsonl.gz {"w":" Doritos","s":1423.44,"e":1423.79,"p":0.94,"seg":118}
p is the decoder's own confidence and seg the cue the word came from, so the
sidecar joins back to the WebVTT without re-alignment. Roughly 274 KB per
episode before compression.
e is clamped to at least s on the way out, because whisper.cpp sometimes
stamps a word with its end before its start.
whisper_exp_compute_token_level_timestamps guards its monotonicity fix-up on
j > 0, so it never repairs a segment's first token, and it runs per segment,
so it cannot order across a segment boundary. Over a 17,553-episode corpus that
reached 2.1% of episodes, and about 4.8% of the words inside an affected one,
with 54% of the inversions sitting on the first token of their segment.
e is the half that moves, matching the t1 = max(t0, t1) whisper.cpp applies
to the tokens it does repair; lowering s instead would drag the word back
past whatever precedes it. The word is clamped rather than dropped, because it
still marks a real position in the stream and dropping it would shift every
index built on the sidecar. Sidecars written before this change keep their
inversions until the episode is decoded again, so a reader that must also cope
with the existing corpus still needs its own min/max.
whisper.cpp only applies max_len when token_timestamps is on — the wrap call
is nested inside if (params.token_timestamps) in whisper_full, so sending
max_len alone is silently ignored. Without them a whole episode can come back
as a single cue; 139 episodes in one corpus did, the worst covering over 7,200
seconds. A cue that long cannot carry a usable timestamp.
split_on_word cuts at word boundaries rather than mid-token.
whisper-server defaults to greedy and whisper-cli to beam_size=5; both run
strategy = beam_size > 1 ? BEAM_SEARCH : GREEDY. Adopting the server without
sending this field silently put the pipeline on greedy.
Greedy's characteristic failure is repetition: it locks onto a phrase and emits it for minutes. It hit 0.7%–5.0% of episodes per show across the first eleven regenerated shows, and a repair pass running beam has fixed 57 of 57 of them, most on the first attempt.
A paired trial on one show — same audio, same model, same fields, only
beam_size moved — found beam equal or better on healthy material too:
| greedy | beam | |
|---|---|---|
| median punctuation/word | 0.1546 | 0.1611 |
| sub-threshold loops cleared | — | 5 of 5 |
| clamped / unpunctuated episodes | 0 | 0 |
and it rescued a shredded episode outright, 0.74 s/cue to 2.42 with punctuation 0.109 → 0.208.
The costs are real but small: roughly 30% more decode time, and about 12% fewer
cues. Coarser cues used to matter because the cue was the floor on boundary
precision; with per-word times in words.jsonl.gz it no longer is.
Set beamSize = 1 for greedy.
Acceleration is determined by how the whisper.cpp server binary was compiled:
- macOS (Apple Silicon) — Metal acceleration, built in by default
- Linux — CUDA acceleration if built with
WHISPER_CUDA=1; CPU fallback otherwise
With pipeline/.env filled in (see Pipeline configuration below), no arguments are needed:
./transcribe./transcribe builds the distribution and runs it, so a source change can never leave you
running a stale binary. It passes its arguments straight through and exits with the
pipeline's own status, and it resolves .env and relative paths against your current
directory — it is the long form below with the path memorised for you:
./gradlew :pipeline:installDist
./pipeline/build/install/pipeline/bin/pipeline./gradlew :pipeline:run also works, with the caveats under
Arguments below.
Every .env setting also has a flag, which takes precedence over it — run
./transcribe --help for the list.
Give each instance its own pods.yaml and its own whisper.cpp server. In one terminal:
./transcribe --config ~/pods-a.yaml --whisper-url http://localhost:8081and in another:
./transcribe --config ~/pods-b.yaml --whisper-url http://localhost:8082Don't use ./gradlew :pipeline:run for this — two concurrent Gradle invocations in one
checkout serialise on the project lock. ./transcribe takes that lock only for its build
and has released it by the time the pipeline starts, so a second launch waits a second at
most.
Both can share one data directory, but nothing coordinates them: list a show in only one
of the two files. If both walk the same feed they will download the same episode to the
same audio.mp3.part staging file, and whichever finishes first deletes the other's.
python3 index.py /path/to/data_directory
# Specify a custom database path
python3 index.py /path/to/data_directory --db /path/to/podcasts.db
# Read every transcript again, instead of only what changed
python3 index.py /path/to/data_directory --fullThe database defaults to podcasts.db inside the data directory. Designed to run directly on the machine hosting the files to avoid network filesystem overhead.
The first run against a database reads everything. After that it stats each
transcript.json and opens only the ones whose modification time or size has changed,
along with any it has not seen before; rows whose file has gone are deleted. Reading the
files is the expensive part on a network share, so a run that finds nothing new finishes
in about the time it takes to walk the tree.
Each row records the file it came from, and a second table records the files that were
walked but deliberately not indexed — no _id, no transcript, or unreadable — so those
are not reopened on every run either. They are picked up as soon as they change.
episodes_fts is an external-content table, so it is maintained alongside: rows are
deleted from and inserted into the index by hand as their episodes change, or the whole
index is rebuilt in one go once more than a fifth of the corpus is affected.
Three things are worth knowing:
- A rewrite that lands with both the same modification time and the same size as the
version already indexed is not noticed. Filesystems with coarse timestamps make this
possible in principle;
--fullis the repair. - An empty data directory leaves the database untouched rather than emptying it, which is
what an unmounted share looks like. So does a run that would remove every indexed
episode — a restore that leaves the transcripts truncated but freshly stamped looks
exactly like a corpus that has legitimately gone.
--fullis how you say you meant it. - A file that cannot be read or stat'ed this time is left as it was, rather than treated as deleted, so a share that blinks costs nothing. The next run picks it up. A rebuild keeps only what it reads, so it refuses outright rather than dropping what it could not look at — unless nothing is indexed yet, where there is nothing to lose.
--full rebuilds both tables from every transcript, and is taken automatically when the
database predates incremental indexing.
Copy .env.example to .env and fill in your values (.env is gitignored):
cd web
cp .env.example .envAPP_DB_PATH=/path/to/podcasts.db
# Required. The pipeline's data directory, served over HTTP. Word timings are read from under it.
APP_DATA_URL=http://your-nas:9280
# Optional. The pipeline's audio directory, served over HTTP; defaults to APP_DATA_URL.
#APP_AUDIO_URL=http://your-nas:9281Quarkus picks up .env automatically. Alternatively, override properties inline:
Development (live reload on template/code changes):
./gradlew :web:quarkusDevProduction — build a runnable JAR then launch it:
./gradlew :web:build
java -jar web/build/quarkus-app/quarkus-run.jarConfiguration properties can also be overridden at launch without editing the file:
java -Dapp.db.path=/data/podcasts.db \
-Dapp.data.url=http://nas:9280 \
-jar web/build/quarkus-app/quarkus-run.jar| Path | What it is |
|---|---|
/search |
Full-text search with filters for duration, podcast, collection, tag, year and episode type, and a relevance/newest/oldest sort |
/episode/{id} |
One episode: metadata, audio player, and the transcript as clickable cues. ?q= highlights the query's matches and steps between them; #t=<seconds> opens on a cue and starts playback there |
/podcasts |
What the corpus holds: one card per podcast with artwork, episode count, total hours and date range, plus when index.py last wrote the database |
The pipeline writes words.jsonl.gz beside every episode's audio: one line per word
with its start, end and the decoder's own confidence. The web module fetches it from
APP_DATA_URL, so that must be an absolute http(s) URL; with a relative one the
feature stays hidden.
With it set, the transcript gains a Mark low confidence toggle, and playback
highlights the current word rather than only the current cue. Words whisper scored
below p = 0.4 are faded when the toggle is on, so a reader can see where the decoder
was guessing. Dimming is opt-in: a transcript permanently mottled with faded words is
harder to read than one that never shows its confidence at all.
The sidecar is roughly 60 KB per episode, so it is never fetched on page load — the first play pulls it, and so does ticking the toggle. Episodes transcribed before word timestamps were emitted have no sidecar; the route returns 404 and the toggle says No word timings rather than sitting there doing nothing. Three other cases say something more useful than that:
- a query that matched every line leaves nothing to split, so the toggle says Hidden on matched lines and stays live — clearing the query brings the words back;
- a sidecar written against a different decode of the episode says Word timings out
of date, which re-running
index.pyover the new transcript fixes. It is judged whole rather than line by line: one cue in common, by ordinal or by coincidence, is not agreement; - whisper does not always time every word, and the pipeline drops the ones it did not, so a line the sidecar cannot rebuild keeps its own text. If that is every line, the toggle says Word timings incomplete.
A line is only ever split when its words rebuild it exactly, so the transcript itself is never rewritten by the sidecar.
The file is fetched by the web module from APP_DATA_URL and served from
/episode/{id}/words, rather than fetched by the page from the data host: it needs no
CORS grant on a server that only has to serve files, and Content-Encoding: gzip lets
the browser inflate it instead of the page carrying a decompressor. It is sent gzipped
whatever the request's Accept-Encoding says, since the file is only gzip on the data
host — curl it with --compressed.
The URL is APP_DATA_URL, the episode's directory (the database's relative audio path
minus audio.mp3, each segment percent-encoded) and words.jsonl.gz. A path segment of
.. or ., or an empty one, is refused rather than sent, so a bad database value
cannot walk up out of the data URL. APP_DATA_URL needs a host and no query string or
fragment, or the feature stays hidden.
A 404 from the data host is a 404 here, and the toggle says No word timings. Any
other outcome — a timeout, a refused connection, a 5xx — is a 502, which the page
treats as transient: it says nothing and tries again on the next play.
The response revalidates rather than being held: re-transcribing an episode rewrites
the sidecar and the cue ordinals it is keyed to together, and an hour-old sidecar
against a fresh transcript mis-times every word. The browser's If-None-Match is passed
on to the data host, so an unchanged file is a 304 on both legs and no body crosses either.
The tag is the data host's own ETag when it sends one. A host that sends none (Python's
http.server) gets a SHA-256 of the bytes just fetched, which is always right but can
only save the browser leg. The trade-off: a host whose ETag is derived from a stat
can name a version the body is not, so a same-size file put back with its old mtime
would keep the old tag, and no-cache would pin the stale timings in the browser until
the file next changed.
On a line the search query matched, the <mark> highlighting wins and the line is
not split into words — word timing is the lesser feature there.
For reading the corpus from a shell or a notebook:
curl 'http://localhost:8080/api/search?q=climate+change&sort=newest&pageSize=5'
curl 'http://localhost:8080/api/episode/abcd1234'/api/search takes the same parameters as /search (q, duration, podcast,
collection, tag, year, episodeType, sort, page) plus pageSize, capped at 100.
It returns the result page with totalCount, totalPages, hasNext and hasPrevious, and
each episode carries a plain-text snippetText of the matched passage. The transcript
field is absent from that payload rather than null, so a consumer can tell "not included"
from "this episode has none"; /api/episode/{id} returns one episode with it, or a 404
with a JSON body.
episodeSummary comes from the feed. A summary containing a < is treated as HTML and
sanitised before it goes out, the same rule the episode page uses; anything else is passed
through untouched, since running prose through an HTML sanitiser turns its ampersands and
quotes into entities. Prose that happens to contain a < — an address in angle brackets,
say — is sanitised and loses it, on both the page and the API.
What that sanitising buys is narrow: it strips scripting from markup so the value can be
inserted as HTML. It does not make the value safe to drop into an HTML attribute —
sanitised markup still contains quotes, and a prose summary is returned with its quotes
intact, so either can break out of an unescaped attr="...".
Every other feed-supplied string in the payload is raw — episodeTitle,
podcastTitle, snippetText, the links and the tags are whatever the feed said, because
an API cannot know how a consumer will render them. Escape everything at render time for
the context you are rendering into; the HTML pages here do it with th:text.
A page number far enough past the end returns no rows rather than wrapping around to the first page.
Copy pipeline/.env.example to pipeline/.env and fill in your values (.env is gitignored):
PIPELINE_DATA_DIRECTORY=/path/to/download-directory
#PIPELINE_AUDIO_DIRECTORY=/path/to/audio-directory
PIPELINE_WHISPER_SERVER_URL=http://localhost:8080
PIPELINE_CONFIG_PATH=/path/to/pods.yaml
#PIPELINE_VERBOSE=truePIPELINE_AUDIO_DIRECTORY is optional. Unset, the mp3s sit beside their transcripts in the
data directory. Set, the mp3s go there instead, under the same <podcast>/<episode> layout,
and the data directory keeps everything else. Moving an existing tree from one layout to the
other is not supported.
PIPELINE_VERBOSE is optional; when set to a non-blank value it overrides the verbose value from
pods.yaml. Leave it commented out (or blank) to let pods.yaml decide — note that any value other
than true counts as false, so PIPELINE_VERBOSE=false will override verbose: true.
The .env file is resolved relative to the working directory: pipeline/.env is tried
first (running from the repo root), then ./.env (running from inside pipeline/, or
next to an installed distribution). A missing .env is fine as long as the arguments
below supply the three required values.
| Flag | Overrides |
|---|---|
--config <path> |
PIPELINE_CONFIG_PATH |
--data-dir <path> |
PIPELINE_DATA_DIRECTORY |
--audio-dir <path> |
PIPELINE_AUDIO_DIRECTORY |
--whisper-url <url> |
PIPELINE_WHISPER_SERVER_URL |
--whisper-model <name> |
PIPELINE_WHISPER_MODEL, or whisper_model in pods.yaml |
--verbose / --no-verbose |
PIPELINE_VERBOSE |
--recover-orphans / --no-recover-orphans |
recover_orphans in pods.yaml |
--orphan-limit <n> |
orphan_recovery_limit in pods.yaml |
--quality-retry / --no-quality-retry |
quality_retry in pods.yaml |
--dry-run |
No equivalent; see Dry run |
--dump-feed-markup <url>, --dump-limit <n> |
No equivalent; see What else a feed carries |
--dump-audio-chapters <dir> |
No equivalent; see Chapters inside the audio |
--retranscribe <dir>, --retranscribe-id <hex8>, --retranscribe-list <file>, --retranscribe-flagged, --retranscribe-limit <n>, --retranscribe-force |
No equivalent; see Re-transcribing an episode |
--verify-pairs |
No equivalent; see Checking pairs |
--list-defects |
No equivalent; see Repairing windows |
--repair-gaps |
No equivalent; see Filling gaps |
Precedence is argument, then .env, then pods.yaml. A flag that is not passed falls
through, so --whisper-url alone leaves everything else coming from .env.
Under Gradle the flags go through --args. Gradle splits that string itself rather than
handing it to a shell, so write $HOME or an absolute path — a ~ inside the quotes
arrives at the pipeline literally and the config file is not found:
./gradlew :pipeline:run --args="--config $HOME/pods-a.yaml --whisper-url http://localhost:8081"A bad flag exits 2 with its message on stderr. Unusable configuration exits 1 — that
covers a missing --config, a data directory that is missing or not writable, and a
whisper server that does not answer the preflight (see
When whisper is not there). A run that found nothing new to
transcribe exits 0, so a wrapper script can tell a failed launch from a quiet one.
verbose— enable debug logging (optional, defaultfalse; overridden byPIPELINE_VERBOSEwhen that is set)skip_after_consecutive— stop walking a feed once this many consecutive already-transcribed episodes are seen (optional, default20)exclude_title_keywords— titles matching any of these are skipped for every feed (optional, see Non-content exclusions; set to[]to disable)min_episode_duration_seconds— skip episodes shorter than this (optional, default150; set to0to disable)recover_orphans— transcribe episodes that aged out of their feed before they were processed (optional, defaulttrue; see Orphan recovery)orphan_recovery_limit— at most this many orphans per run, across all podcasts (optional, default0, meaning no limit)language— ISO 639-1 code whisper decodes in, orautoto detect from the audio (optional, defaulten). Case does not matter; it is lower-cased before being sent. A code whisper does not recognise fails the run at startup rather than being sent:whisper_lang_idreturns-1for an unmatched code and its caller adds that to the language token's base index, so the decode would proceed in the wrong language with nothing reporting it — and the quality gate cannot catch it, since a wrong-language decode can be fluent, punctuated and loop-free. A region tag (en-US), a 639-2 code (eng) and a stray space are all rejected on those grounds. A language name (english) is rejected too, but for a different reason: whisper does accept it, and would decode correctly, but the initial prompt is matched on the code — so a name would silently drop the one lever that fixes unpunctuated decodesinitial_prompt— the prompt sent with this feed's decodes, written in itslanguage(optional; the top-level default is English prose). A podcast that sets it must setlanguagetoo, and neither may beauto. Seepromptandcarry_initial_promptquality_retry— decode a flagged transcript a second time and keep the better one (optional, defaulttrue; see Transcript quality gate)notify_url— POST the run's summary line here when a run finishes (optional; see Run report)podcasts— list of RSS feeds to process, each withname,url, optionalcollections, optionalexcludes, an optionalmin_episode_duration_secondsthat overrides the global floor, and an optionallanguagethat overrides the global one
name becomes the show's directory name, so changing it moves every episode of
that feed. Re-capitalising it used to create a second directory for the same
show; the pipeline now reuses a directory that differs only by case, and logs
when it does. Renaming it any other way still starts a fresh directory and
re-transcribes the feed.
Two filters run before an episode is downloaded, so excluded episodes cost nothing.
exclude_title_keywords matches whole words, case-insensitively, against the episode
title. Whole-word matching is what makes the list safe to apply globally: a substring
match on repeat also swallows "Repeating FRB Mystery", and on archives it swallows
"Inside the Archives", an actual interview series. The default list is trailers,
cross-promos and repeat markers: trailer, introducing, encore, classic episode,
rewind, re-release, re-run, rerun, rebroadcast, best of, repeat, replay,
from the archives. Setting the key replaces the list rather than adding to it; the
per-podcast excludes list is separate, still a plain substring match, and still applies
on top.
coming soon is deliberately not in the default list. It reads as a safe global term
but matches Planetary Radio's "2012 DA14--Coming Soon to a Planet Near You!", a real
29-minute episode. Every genuine hit for it sits in one of two feeds, so it belongs in
their excludes rather than the global list.
min_episode_duration_seconds uses the feed's itunes:duration. An episode whose feed
omits the tag is never filtered on length. The 150 default sits at the point where
short-form content starts to outnumber promos — below it a feed is almost entirely
trailers, hiatus notices and "coming soon" stubs. Feeds that publish genuine short-form
episodes need the floor lifted per podcast:
- name: A Short-Form Show
url: https://example.com/feed.rss
min_episode_duration_seconds: 0_id is md5(guid)[:8], and the episode table is keyed on it. Publishers do
occasionally ship two entries under one GUID — HBR IdeaCast has two such pairs —
which under INSERT OR REPLACE silently collapsed them into a single row, losing
one episode from search while its transcript sat on disk.
index.py now suffixes the later members of a clashing group (<id>-2, -3, …),
ordered by audio path so the ids are stable across re-indexes, and warns on stderr
naming every episode involved. Both episodes stay searchable and keep a working
/episode/{id} permalink. Widening the hash would fix nothing here — the inputs
really are identical — and would rename every episode directory, forcing a full
re-transcription.
Feeds are typically ordered newest-first. Rather than stat'ing every episode directory (expensive for feeds with thousands of entries), the transcriber walks the feed and stops on a podcast once it sees skip_after_consecutive transcribed episodes in a row. The counter resets on any gap, so a cancelled run that left untranscribed holes will be picked up on the next invocation.
The pipeline only transcribes what the feed offers, and publishers age entries out on hard
caps or date cutoffs. An episode downloaded but not yet transcribed when that happens is
stranded: its directory holds an audio.mp3 no later run will ever look at. 659 of 17,517
episode directories were in that state when this was written.
After walking a feed, the pipeline lists the podcast's directory once and compares each
episode directory's YYYY-MM-DD-<hex8> prefix against the feed. Directories whose prefix
is absent are opened and, if they hold audio but no transcript, transcribed in place —
newest first, and never renamed, because indexed rows already point at the existing path.
The feed still describes the show, so every podcast_* field is accurate. The entry is
gone, so only what the directory name carries can be recovered:
| Recovered | Source |
|---|---|
_id |
the <hex8> in the directory name — the same id the episode always had |
episode_published_on |
the date in the directory name |
episode_title |
the rest of the directory name, dashes back to spaces (lossy: punctuation is gone) |
episode_duration |
where the decoded speech ends — approximate, and short of the file by any trailing silence |
all_tags |
feed-level categories only |
episode_metadata_recovered |
true, present only on recovered files |
episode_audio_link, episode_web_link, episode_image, episode_summary,
episode_subtitle, episode_authors, episode_number, episode_season and
episode_type are written as null. The web UI hides a null field and would render an
empty string as a dead link, so the distinction matters.
Audio that decodes to nothing leaves a recovery-failed marker in the directory. Without
it the file would be re-uploaded to whisper on every run forever — an orphan has no feed
entry whose download could fail and stop it. A transcriber error (server down, timeout)
writes no marker and is retried.
Two situations are reported as errors rather than recovered, because the feed entry has better metadata than the directory name ever could:
- a directory whose id is in the feed under a different date — a publisher re-issued or re-dated an old episode
- an episode still in the feed, still untranscribed, that
skip_after_consecutivestopped before reaching. Raise the threshold to pick it up.
Recovery is on by default and costs about 20 seconds across a full run, because only the
directories absent from the feed are opened. --no-recover-orphans turns it off;
--orphan-limit <n> bounds how many a single run will transcribe, which is worth setting
when the backlog is large enough to crowd out new episodes.
Three failure modes were previously only ever found by an external pass reading the
finished corpus: an episode decoded with no punctuation at all, a greedy-style
repetition loop, and cues shredded into one- and two-word fragments. The pipeline now
scores every transcript as it writes it, and records the score in transcript.json
under episode_quality:
"episode_quality": {
"punctuation_per_word": 0.1611,
"seconds_per_cue": 2.42,
"repeated_share": 0.0,
"longest_repeated_cue_run": 1,
"mean_word_probability": 0.87,
"low_confidence_share": 0.04,
"word_count": 8123,
"cue_count": 1044,
"flags": []
}flags is empty for a healthy decode, and otherwise names what tripped:
| Flag | Trips when | Measured |
|---|---|---|
no-speech |
the decode produced no words at all | every other check needs words to measure, so without this an empty decode scores clean |
unpunctuated |
punctuation per word below 0.03 |
healthy episodes sit near 0.15 |
shredded-cues |
under 1.0 seconds per cue, with at least 50 cues |
a shredded episode measured 0.74 against 2.42 re-decoded |
repetition-loop |
one 4-gram repeats over 5% of the words, or 4+ consecutive cues are identical | greedy decoding hit 0.7%–5.0% of episodes per show |
low-confidence |
over 20% of words scored under p = 0.3 |
skipped entirely for episodes decoded before word timestamps existed |
When a decode is flagged the pipeline decodes the episode once more and keeps the better of the two: a decode that produced speech always beats one that did not, then fewer flags wins, and on a tie the more punctuated one. Whisper is not deterministic, and the repair passes that inspired this cleared 57 of 57 repetition cases, most on the first re-decode. A retry doubles decode time for the 1–5% of episodes that trip a flag.
no-speech is why the retry is a comparison rather than a preference: an empty decode
trips none of the other checks, so without a flag of its own it would score clean, win on
flag count, and replace a real transcript. One flag is not enough on its own, though —
it still beat a decode bad enough to trip two — so an empty decode loses to any decode
with speech in it before flags are counted at all. A poor transcript is worth more than
none, and on the recovery path none is permanent: an episode that decodes to nothing
gets its one retry and then recovery-failed abandons it for good.
If the kept decode is still flagged it is written anyway, with its flags recorded, and
a warning goes to the error log — the transcript is still worth having, and
episode_quality.flags is what a later re-transcription pass selects on.
Turn the retry off with --no-quality-retry, or quality_retry: false in pods.yaml.
The thresholds above are starting points measured on the real corpus, not tuned
constants. They live together at the top of
pipeline/src/main/kotlin/com/rsstowhisper/pipeline/TranscriptQuality.kt, each with the
number it came from.
decode_without_history: true in pods.yaml makes the first decode of every episode
send max_context=0, so whisper conditions no window on the text before it. That text
is what feeds repetition loops and stretch-copies. whisper.cpp skips the initial prompt
along with the history, so no prompt is sent, and a decode that comes back flagged is
retried with the prompt and the history; the better of the two is kept.
Measured on the same audio, same model and server, 2026-09-25:
| sample | request | clean | loops | stretch-copy | unpunctuated | words |
|---|---|---|---|---|---|---|
| 24 random episodes | default | 18 | 5 | 1 | 0 | 1.00 |
max_context=0 |
23 | 0 | 0 | 1 | 1.01 | |
| 16 with stretch-copies or loops | default | 9 | 3 | 6 | 0 | 1.00 |
max_context=0 |
15 | 0 | 0 | 1 | 1.02 |
The cost is punctuation: median marks per word fell from 0.140 to 0.114 without the prompt. It is off by default.
./transcribe --repair-windows --retranscribe-list episodes.txtRe-decodes only the stretches of each target around its defects, instead of the whole episode, and splices them into the pair on disk. A defect is a stretch-copy, a loop (four or more identical cues, a few cues repeated in turn, or one cue repeating a phrase faster than anyone speaks), an echo or long copy of the cues before it, or a cue of 10 s or more that is nothing but a sentence of the initial prompt or one of whisper's stock phrases ("Thanks for watching."): over music or silence, a decode voices them.
- Windows are the defective cues plus two good cues either side. whisper-server
decodes just that stretch (
offset_t,duration); nothing is cut from the mp3. - Anchors. The outer good cue at each end is kept exactly as it was; the one against the defect is re-decoded, since it is the cue most often damaged. The new decode is cut after the last three words of the left anchor and before the first three of the right one, found within 2 s of their old time. Otherwise the decode's words are walked over the anchor's, past a word the anchor does not have, and only then cut by time, in the clock the other anchor's words set. The new window's clock is mapped onto the old one between the two anchors.
- Attempts. Prompted, then without history (
max_context=0), then both again at temperature 0.2, then both again with a beam of 8: whisper can skip speech on one decode and keep it on another, but since v1.9.4 only if the request differs. A window is applied only if it comes back with fewer defects than it had, counting any sentence of the prompt it now holds; an attempt holding a prompt sentence as long as a leak is refused outright, as is one that drops a cue the base had right where speech is heard under it. - Lost speech. An attempt that leaves more than 2 s of speech without a word near it is refused: whisper can skip a whole window and carry on, and the words either side still match. Without VAD, the speech is the base's own words in the window's good cues.
- Silence. With
vad_binaryandvad_modelset inpods.yaml, Silero VAD (whisper.cpp'swhisper-vad-speech-segments) says where speech is instead. Cues the new decode put over non-speech are dropped, keeping one♪where it marked music, and a window may come back empty only where VAD hears nothing. A VAD that cannot run stops the batch; one that hears nothing under most of an episode's cues is ignored for it. The server's own VAD cannot be used for this: withvadon it appliesoffset_tto VAD-compressed time. - Only a pair already from one decode is repaired. The result gets a new run id;
whisper_run.base_runis the pair it was spliced into andwhisper_run.repairsrecords each window, its anchors and gaps, and what was refused.
--list-defects prints every episode the repair would find something in, with counts by kind,
as a list that works with --retranscribe-list. It only reads.
./transcribe --repair-gaps --retranscribe-list episodes.txtwhisper decodes in 30 s windows, and when the model ends one early, whisper.cpp skips the rest of it ("single timestamp ending"). Over an ad jingle, music or another language the model says so wrongly, and the speech later in the window is lost: nothing is written, or a filler such as "Thank you." is held to the window's end. Re-decoding the same window from the same cue edges meets the same wall.
- Gaps are stretches VAD hears with no word starting within 1 s, 4 s or more of them, from the cue before to the cue after. Filler cues over a gap, and crammed cues at its edges (whisper often crams part of what it skipped into no time at the skip), are replaced with it. A gap at an episode's very start or end has no anchor on that side and is skipped.
- Chunks. The speech between the anchors is cut and merged from VAD's spans into pieces under 30 s, each decoded alone and starting on speech, the way WhisperX avoids the skip.
- Attempts. Unprompted first, then prompted, then warmer and with a wider beam. A prompt on every short chunk invites skips and echoes; unprompted chunks sometimes come back without punctuation, so a punctuated attempt within a second of the best coverage wins.
- Kept only if it covers at least 2 s more of the speech, keeps every heard cue around it,
holds no defect of its own, no prompt sentence, and is not mostly the text of the cues either
side (the next line written early over music). Recorded with
"kind": "gap". - Needs
vad_binaryandvad_model. It can run with--repair-windowsin the same pass; a gap inside a defect window is left to that window.
Redoing an episode used to mean deleting its transcript.json by hand. With the quality
gate recording flags, the loop closes:
./transcribe --retranscribe-flagged --retranscribe-limit 50
./transcribe --retranscribe "Ask-a-Spaceman/2024-01-02-abcd1234-some-episode"
./transcribe --retranscribe-id abcd1234
./transcribe --retranscribe-list episodes.txt--retranscribe-list takes one <podcast>/<episode> path or id per line. Blank lines and
# comments are skipped and anything after a tab is ignored, so the output of
--verify-pairs can be passed straight back in. It counts as naming its targets, so
--retranscribe-force applies to it.
A path is the episode's directory as it sits on disk, so it carries the escaped podcast
name (Ask-a-Spaceman, not Ask a Spaceman) — every character that is not a letter or
a digit became a dash when the directory was created.
--retranscribe and --retranscribe-id are repeatable, and any of the three skips the
feeds entirely — every target is already on disk, and an episode that aged out of its
feed has nothing left to fetch. --retranscribe-flagged reads every transcript.json
under the data directory, which is slow on a network volume; --retranscribe-limit
caps it, since a first pass over the corpus can select hundreds.
The limit is a rolling window, not a fixed prefix. Each attempt writes the time into
a retranscribe-attempted file in the episode directory, and the flagged scan takes the
least recently attempted first — never-attempted episodes ahead of everything, then
oldest attempt, with the directory name breaking ties so a run over the same corpus picks
the same episodes. Without that, an episode whose flags cannot clear — genuinely bad
audio, a music-heavy show, one that is mostly silence — would sit at the front of the
sorted list forever and starve everything behind it, costing decode time on every run and
changing nothing.
The time is recorded for every attempt, not every success, because the episodes that starve the selection are exactly the ones each re-decode refuses. The exception is a decode that could not reach the whisper server: nothing was learned about that episode, so it keeps its place rather than rotating away untried — otherwise a server that died partway through would send the rest of the window to the back of the queue unexamined.
Each target needs its audio.mp3; one without it is reported and skipped. The rewrite
keeps every existing field and replaces only episode_transcript and episode_quality
— all the rest came from a feed entry that may no longer exist, so what is on disk is
the best there is. episode_duration is replaced too, but only when
episode_metadata_recovered is true, because that number came from the previous decode
rather than from the feed.
--retranscribe-force keeps the new decode whatever it scores. Use it when you are
redoing an episode for a reason the score cannot see — a corrected language, a fixed
initial_prompt, a model change. The comparison counts flags and then compares
punctuation; it knows nothing about which language the old decode was in, so a
wrong-language transcript that happens to score a hair better would otherwise win.
It only applies to explicitly named targets, and is refused outright alongside
--retranscribe-flagged: that scan is the one that runs at scale, and it must never be
able to make the corpus worse. A flag that silently covered only the named half of a
mixed run would be worse than one that says so.
Force skips the quality comparison and nothing else. An episode with no audio, a decode
that found no speech, and one that came back without word timestamps are all still
refused — force means "I know better than the score", not "ignore a server that stopped
sending token_timestamps".
The no-speech refusal is a check in its own right rather than a consequence of the comparison, because a blank transcript is not caught by the emptiness test ahead of it: segments carrying timestamps and no text render a VTT that is not blank.
It needs a target. --retranscribe-force on its own is refused rather than ignored,
since the run would otherwise fall through to an ordinary feed run — which is what a
script whose target list came out empty would get, having asked for the opposite.
Flags are counted only where both decodes could have raised them. low-confidence
needs word probabilities, so a decode scored without them cannot trip it and would
otherwise win on raw count against one that did — keeping a transcript with no word
timings forever, since every re-decode that finally produced them would be discarded.
With the flags level, the decode that has word timings wins; it is worth having, but not
worth trading a better-scoring transcript for.
A re-decode is kept only if it scores at least as well as the transcript it would
replace, by the same measure the quality gate's retry uses — fewer flags, then better
punctuation, then word times. Whisper is not deterministic, so a redo can come back worse than what it
overwrites, and that write is the only copy: re-transcribing can improve an episode or
leave it alone, never cost it the better decode. A transcript written before the quality
gate has no score to compare against, so it is simply replaced. Nor is a transcript
whose words.jsonl.gz came from another decode: its score describes one half of a pair
neither half of which is usable, so it is replaced whatever it scores.
After the batch every target is checked with the same test as --verify-pairs, written
or not, and the run fails listing any whose pair still disagrees.
A re-decode that comes back with no word timestamps at all is refused outright when the
decode on disk had them — judged by its recorded score rather than by whether
words.jsonl.gz is there, since writing the sidecar is allowed to fail. The comparison
below only reaches word times once the flags and the punctuation are level, so a decode
that lost them can still win outright on a flag the other tripped — and would take the
words.jsonl.gz with it. It means the server ignored token_timestamps, and the warning
says so.
Otherwise both files are staged beside their originals and moved into place, with the
sidecar absent for the whole swap: cleared first, restored only once the transcript it
belongs to is there. words.jsonl.gz addresses cues by position, so either file left
beside the other's transcript mis-times every word — silently and for good, since an
episode that has a transcript.json is one nothing revisits. Interrupted anywhere in
between, the episode is left visibly missing a sidecar instead of quietly holding the
wrong one. If the new sidecar cannot be written at all — a full disk, an I/O error — the episode is
left exactly as it was rather than committed without one.
Staged files are named per process, because two instances sharing a data directory select
the same episodes: --retranscribe-flagged scans the whole tree, whatever podcasts its
config names. That stops one run moving another's half-written file, and nothing more —
the swap is two moves, not one, and nothing here locks, so do not point two runs at the
same episode: they can interleave into a transcript and a sidecar from different decodes.
An id is part of a path, not a unique key, so --retranscribe-id redoes every copy it
finds rather than guessing which was meant.
--dry-run walks every feed and applies every filter exactly as a real run does, then
reports what it would have done instead of doing it:
Would transcribe Ask a Spaceman/2024-03-14-a1b2c3d4-what-is-dark-matter
Ask a Spaceman: would transcribe 4 episodes
Would recover Ask a Spaceman/2019-01-01-deadbeef-aged-out
Dry run: would transcribe 12 and recover 3 episodes
Nothing is downloaded, nothing is decoded, and no podcast or episode directory is
created — useful because the global exclusions, per-podcast excludes, the duration floor,
skip_after_consecutive and orphan recovery interact, and reading the result is cheaper
than spending GPU hours on a misconfiguration.
The whisper server is never contacted, so a dry run works with it switched off. It exits
0. Combined with --retranscribe, it lists the episodes that would be redone without
touching their transcripts — worth doing before --retranscribe-flagged, which selects
across the whole corpus.
--dump-feed-markup <url> fetches one feed and prints, per entry, everything ROME parsed
that no registered module claimed — which is every namespace the pipeline does not map.
--dump-limit caps the entries sampled (default 10). It needs no pods.yaml, no data
directory and no whisper server, and exits without running anything:
./transcribe --dump-feed-markup https://example.com/feed.rss --dump-limit 10feed: Some Podcast
modules: http://www.itunes.com/dtds/podcast-1.0.dtd
channel foreign markup:
<podcast:medium>podcast</podcast:medium>
10 of 300 entries
[1] An Episode
published: Wed Sep 10 09:00:00 UTC 2026
modules: http://www.itunes.com/dtds/podcast-1.0.dtd
foreign markup:
<podcast:chapters url="https://example.com/ep1/chapters.json" type="application/json+chapters">
<podcast:soundbite startTime="1234.5" duration="60.0">A clip</podcast:soundbite>
distinct foreign elements across the sample:
10 x podcast:chapters
3 x podcast:soundbite
The closing tally is the point: a tag that only some episodes carry is easy to miss one
entry at a time. Elements a module did claim — anything itunes:, dc:, media: —
are absent from the markup and named in the modules line instead, so the report
separates what is already reachable through ROME from what would need new parsing.
--dump-audio-chapters <dir> reads the ID3v2 tag of every audio.mp3 under a data
directory and reports which episodes carry CHAP chapter frames. It never touches the
audio itself -- only the tag at the head of each file -- so a pass over a large library
costs little, and it needs no network:
./transcribe --dump-audio-chapters /path/to/data --dump-limit 10scanned audio.mp3 in 1284 episodes
with ID3 chapters: 37 (2.9%)
showing 10 of 37
Some Podcast/2026-01-02-a1b2c3d4-an-episode
0:00:00-0:01:32 Advertisement
0:01:32-0:24:10 Part One
chapter titles by episodes carrying them:
37 episodes, 94 uses Advertisement
37 episodes, 37 uses Intro
1 episodes, 1 uses The Chemical Barons
The ordering is the point. A title carried by many episodes is a structural segment -- an intro, a break, a sponsor read -- while a title carried once is that episode's own content, so ranking by episode count puts the recurring furniture at the top.
Chapters read here describe the file on disk, which is the same file the transcript was decoded from. That matters because a feed's own chapter timestamps describe the unstitched master: where a host inserts ads at download time, the two disagree.
The server is asked for its base URL once, before the first feed is fetched; nothing
listening, and the run exits 1 having downloaded nothing. Any answer counts, a 404
included — --request-path moves whisper.cpp's page off /, and a proxy may route only
/inference, neither of which is a reason to refuse to run. Only a gateway saying its
upstream is gone (502, 503, 504) or nothing answering at all fails it.
If whisper goes away during a run, the run carries on to the end and then exits
non-zero, naming how many episodes could not reach it. It does not stop early: the
downloads are not wasted, since a completed audio.mp3 is kept and the next run skips
straight to decoding it. What the exit code buys is that a run which downloaded the whole
backlog and transcribed none of it cannot look successful to whatever is driving it.
Only failures that say nothing is there count — a connection that goes nowhere or dies
mid-response, a gateway 502/503/504, or an empty body. A status the server chose is
the server talking, and what it is talking about is that request: whisper.cpp answers
400 for an mp3 it cannot decode and 500 for one it cannot process, and those episodes
fail on their own merits as they always did, without failing the run.
The error log says whether anything went wrong. It does not say what was done, so every
run also writes <data-dir>/logs/run-<yyyyMMdd-HHmmss>.json and copies it to
logs/latest-run.json — the stamped file is the history, the stable name is what
anything watching the directory can read without listing it first.
{
"started_at": "2026-09-14T02:00:03Z",
"finished_at": "2026-09-14T06:41:55Z",
"duration_seconds": 16912,
"totals": { "transcribed": 41, "recovered": 3, "failed": 1, "skipped": 112 },
"podcasts": {
"Ask a Spaceman": {
"transcribed": 4,
"recovered": 0,
"failed": 0,
"skipped": { "global_keyword": 2, "too_short": 1 }
}
}
}The skip counts are by the reason the pipeline actually applied, which is what answers
"why did this feed transcribe nothing this week". failed counts episodes this run tried
for and did not get, whether they fell over before the download or during the decode.
Reports are kept for 14 days, matching the error log. Two instances sharing a data
directory each write their own stamped file; latest-run.json is whichever finished last.
A report that cannot be written is a warning and never fails the run, and it is written
even when a run dies part-way, since that is when knowing how far it got matters most.
Set notify_url in pods.yaml to have the run POST its summary line — the same
Run finished: N warnings, M errors that goes to stdout — as text/plain when it
finishes. That is exactly what ntfy.sh expects, and other services can
adapt. There is no auth and no retry: the run has already finished by then, so a failed
notification is a warning and nothing more.
Warnings and errors are mirrored to <data-dir>/logs/pipeline-errors.log, rolling daily
and keeping two weeks, and the run prints a count of both when it finishes. A 13,000-episode
run otherwise buries its failures in scrollback.
For each episode, the following files are created:
{data_directory}/{podcast_name}/{YYYY-MM-DD-<hex8>-episode-title}/
audio.mp3 # Downloaded audio, kept as-is and served by the web module
transcript.json # Full metadata + WebVTT transcript
words.jsonl.gz # One line per word, with its own start/end
recovery-failed # Only when recovery found no speech in the audio
With PIPELINE_AUDIO_DIRECTORY set, audio.mp3 moves to
{audio_directory}/{podcast_name}/{same episode directory}/ and the rest stays put.
The episode_transcript field in transcript.json is a raw WebVTT string,
rendered from the same verbose_json response that produced words.jsonl.gz.
episode_quality alongside it records how that decode scored; see
Transcript quality gate.
words.jsonl.gz is written before transcript.json, because the latter
existing is what marks an episode done — so a crash between the two leaves the
episode to be redone rather than permanently without its sidecar. A decode that
could not write its sidecar leaves the episode undone for the same reason —
writing transcript.json anyway would mark done an episode that will never get
one. A decode that simply carries no word times, from a server that ignored
token_timestamps, still writes its transcript; otherwise no episode could ever
complete against such a server.
Both files carry the decode's run id: whisper_run.run_id in transcript.json, and a
run field on every line of words.jsonl.gz — per line rather than as a header row,
since every existing reader takes each line to be a word. whisper_run also records
when the decode ran, the server, the model (from --whisper-model, since whisper does
not report it), every form field sent, the audio's SHA-256 and size, the pipeline's
git describe, and how many words the sidecar holds. Both files are staged in full
before either is moved into place, and the pair is checked on disk after every write.
./transcribe --verify-pairs > diverged.txtReads every episode under the data directory and prints one <podcast>/<episode>, a
tab, and the reason for each whose transcript.json and words.jsonl.gz did not come
from one decode, exiting 1 if there are any. It never contacts whisper. Pairs with run
ids are compared on them; older ones by whether each cue's words rebuild that cue, the
test the web player applies before it trusts a sidecar. A transcript from before word
timings, with no sidecar and no run id, is counted but not listed. A pair that could not
be read just then is listed as "could not be checked" and also fails the check; nothing
replaces a pair on the strength of a failed read. Logs go to stderr, so the list can be
fed straight to --retranscribe-list. A data directory that is missing, or a directory
that cannot be listed, fails the check rather than passing with nothing checked.
The <hex8> in the directory name is md5(entry.uri) truncated to 8
characters. It is part of a path, not a unique key: date and title slug
disambiguate it.
./gradlew test./gradlew ktlintCheckTo auto-format:
./gradlew ktlintFormat