Skip to content

Repository files navigation

rss_to_whisper

Transcribe podcast episodes from RSS feeds using the whisper.cpp HTTP server and index them into a local SQLite FTS5 database for full-text search.

Modules

pipeline (Kotlin)

Walks one or more RSS feeds, downloads each episode's MP3, and POSTs it to a running whisper.cpp HTTP server. The server decodes and resamples the audio itself, so no local transcoding step is needed and the MP3 is what gets kept on disk. It returns verbose_json, from which the pipeline renders a WebVTT transcript into transcript.json and writes the per-word timings alongside as words.jsonl.gz. Designed to run on a schedule (e.g. cron) to keep transcripts up to date.

web (Kotlin)

A Quarkus HTTP server that serves a full-text search interface over the SQLite database produced by index.py. Supports filtering by podcast, collection, episode type, and duration. Search results are ranked by BM25 relevance. The episode detail page shows a clickable transcript synced to the audio player. Opened from a search result it carries the query along (/episode/{id}?q=…), highlights every cue the query matches, scrolls to the first, and steps between them with Prev/Next or the n/p keys. Runs on port 8080 by default.

The FTS index is one row per episode, so SQLite reports which episodes match but not where. The episode page re-applies the query cue by cue, reading it the way FTS5 does: quoted phrases stay whole, AND/OR/NOT/NEAR are operators only in upper case, a trailing * is a prefix, and words are compared case-insensitively with diacritics removed, as the unicode61 tokenizer indexed them.

Note: whisper-server also defaults to 8080. They are rarely up at the same time, but if they are, move one — --port on whisper-server, quarkus.http.port on the web module.

Python scripts

index.py

Reads the transcript.json files in the data directory and writes them into a SQLite FTS5 database. Run this after the pipeline to make new transcripts searchable. After the first run it reads only what has changed, so running it after every pipeline run costs seconds rather than minutes.

Requires Python 3 with an FTS5-capable SQLite. The script prefers pysqlite3 when installed and falls back to the standard library sqlite3 module otherwise, exiting with a clear error if neither has FTS5:

pip install pysqlite3

Note: The default system sqlite3 on some platforms (e.g. Synology NAS) does not include FTS5. pysqlite3 bundles a SQLite build that does.

Prerequisites

  • JDK 21+
  • A running whisper.cpp HTTP server
  • Python 3 with an FTS5-capable SQLite (for the indexing script; pip install pysqlite3 if the system build lacks FTS5)

No ffmpeg is required. MP3s are uploaded as-is and the whisper.cpp server decodes them with its built-in miniaudio decoder, resampling to 16 kHz mono internally.

Building

./gradlew build

Running the whisper.cpp server

The pipeline POSTs audio to the whisper.cpp /inference endpoint. It does not start or manage the server — start it yourself and leave it up for the run:

whisper-server \
  -m /path/to/models/ggml-large-v3.bin \
  --port 8080

The whisper-server binary is built alongside whisper-cli when you compile whisper.cpp. Set PIPELINE_WHISPER_SERVER_URL in pipeline/.env to its base URL.

Only the model is chosen at launch. Everything else that matters is a per-request form field, and server.cpp overrides launch defaults with whatever a request actually sends. So the pipeline's own settings win, and the only knob you pick when starting the server is -m.

Do not pass --convert (it shells out to ffmpeg on the server host — MP3s are uploaded as-is and decoded internally), and do not pass -nt, which would suppress the timestamps the pipeline exists to capture.

Known-good version

The pipeline is run against whisper.cpp v1.9.4 (927cfce3), serving ggml-large-v3 from two builds: Metal with the CoreML encoder on Apple Silicon, and CUDA on Linux. This is not a requirement. It is the version the repair's retries were measured on. The corpus was first transcribed, and repaired once, on v1.9.2 (306c88f4).

v1.9.4 re-seeds the sampler on every request (whisper.cpp#4025), so a server gives the same output for the same request, even when it falls back to sampling. Older servers drift between calls, which is what a repeated request used to get its second chance from. So --repair-windows tries each window prompted and unprompted, then both again at temperature 0.2, then both again with a beam of 8: each retry asks for something different.

Choosing a model

whisper.cpp uses its own GGML model format (.bin files), not the OpenAI Python .pt files. Download from HuggingFace:

curl -L -o ggml-large-v3.bin \
  https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3.bin

Put the matching ggml-large-v3-encoder.mlmodelc next to the .bin, or the encoder silently falls off the Neural Engine and everything gets much slower.

large-v3, not large-v3-turbo

Turbo prunes the decoder and keeps the encoder, and under Core ML the encoder is where most of the wall clock goes. Measured over 13 episodes, large-v3 costs 1.36x turbo's decode time — not the 2-3x that "turbo for speed" implies — and produced a usable transcript on 10 of 13 episodes turbo could not decode at all.

It also fixes a failure nothing else touches. Turbo shreds some episodes into runs of one- and two-word cues: 1,256 episodes in one corpus. An A/B on 4 shredded episodes across 4 arms took fragmentation to zero in all sixteen cells, including the no-prompt no-VAD control. That is a model difference, and neither the prompt nor VAD substitutes for it.

What the pipeline sends, and why

These are the request fields in Transcriber.kt. Each one is load-bearing.

prompt and carry_initial_prompt — the single biggest lever

Whisper intermittently decodes an entire episode with no punctuation and no capitals. That is not cosmetic: whisper segments on sentence structure, so with no full stops the cue boundaries stop tracking speech and every timestamp derived from them is fiction. 654 episodes were hit, and 13 resisted every attempt to re-decode them.

An initial_prompt of ordinary punctuated prose fixed all 13:

approach fixed
initial_prompt 13 / 13
large-v3 10 / 13
best VAD parameter 9 / 13
Silero VAD v6.2.0 6 / 13
re-decode at another VAD threshold 0 / 13

It works because this is a decoder mode, not a segmentation problem. VAD only changes what audio reaches the decoder; a prompt conditions the decoder itself, and punctuation is a style. It also prevents runaway repetition loops — without it, 4 of 4 test decodes collapsed, one returning 539 usable words out of 10,660, and ran 2.5x slower because a looping decode burns tokens.

carry_initial_prompt=true matters: without it only the first window is conditioned, and an episode that degrades later still degrades (13/13 with, 12/13 without).

Keep the prompt generic. It biases vocabulary as well as style, so anything domain-specific contaminates transcripts. The default in Transcriber.kt was checked against a repaired episode — zero occurrences of any prompt fragment, word count within 5% of the original.

The prompt only rides with the language it is written in. The default is English prose, so a podcast set to another language decodes without it — an English prompt on French audio is the same vocabulary contamination as a domain-specific one, and carry_initial_prompt would apply it to every window. Under language: auto it is worse: the prompt would skew whisper's own language detection toward English before it decoded anything, breaking the very thing auto is for — so auto never takes a prompt, however explicitly one is set.

A feed in another language keeps the lever by supplying its own prompt:

- name: Une Emission Francaise
  url: https://example.com/fr.rss
  language: fr
  initial_prompt: Bonjour, et bienvenue dans cette emission. Aujourd'hui, nous allons parler de plusieurs choses.

The prompt's language is the podcast's language, so there is no second key to keep in sync — and a podcast that sets initial_prompt must set language too, or the run refuses to start rather than guessing. A top-level initial_prompt works the same way against the top-level language.

The built-in English prompt stays bound to English, whatever the top-level language is. Only a prompt you actually set belongs to that language; an unset one is not "the default, in French".

Keep any prompt generic, and check a new one the way the English default was checked: decode an episode with it and confirm no prompt fragment appears in the transcript and the word count has not moved.

vad=false — sent explicitly, and off

Not merely omitted. A request that omits vad inherits whatever the server was launched with, silently, which is how one corpus ended up with no record of its own VAD state.

It has to be off because VAD breaks word timestamps. whisper.cpp keeps token timestamps in VAD-compressed time while remapping segment timestamps to real time, so the two drift apart by however much silence VAD removed:

offset, first word vs its own segment
VAD on −1.79s at the start, −6.51s by the end
VAD off +0.00s to +0.20s

Same episode, same model, same fields — on a six-minute episode. Every word time would be early by a growing, episode-dependent, invisible amount.

Nothing is lost by turning it off. VAD's real value was suppressing the unpunctuated collapse, and the prompt does that better (13/13 against 9/13).

response_format=verbose_json

The server can return WebVTT directly. It is derived from the JSON here instead, because per-word start/end come only from verbose_json and the decode computes them either way — rendering to VTT throws them away.

Both artifacts come from one parse of one response. Asking for both formats would mean two decodes, and whisper is not deterministic across runs, so their cues and words could disagree in ways nothing downstream could detect.

Output per episode:

transcript.json     metadata plus the WebVTT string
words.jsonl.gz      {"w":" Doritos","s":1423.44,"e":1423.79,"p":0.94,"seg":118}

p is the decoder's own confidence and seg the cue the word came from, so the sidecar joins back to the WebVTT without re-alignment. Roughly 274 KB per episode before compression.

e is clamped to at least s on the way out, because whisper.cpp sometimes stamps a word with its end before its start. whisper_exp_compute_token_level_timestamps guards its monotonicity fix-up on j > 0, so it never repairs a segment's first token, and it runs per segment, so it cannot order across a segment boundary. Over a 17,553-episode corpus that reached 2.1% of episodes, and about 4.8% of the words inside an affected one, with 54% of the inversions sitting on the first token of their segment.

e is the half that moves, matching the t1 = max(t0, t1) whisper.cpp applies to the tokens it does repair; lowering s instead would drag the word back past whatever precedes it. The word is clamped rather than dropped, because it still marks a real position in the stream and dropping it would shift every index built on the sidecar. Sidecars written before this change keep their inversions until the episode is decoded again, so a reader that must also cope with the existing corpus still needs its own min/max.

token_timestamps, max_len, split_on_word

whisper.cpp only applies max_len when token_timestamps is on — the wrap call is nested inside if (params.token_timestamps) in whisper_full, so sending max_len alone is silently ignored. Without them a whole episode can come back as a single cue; 139 episodes in one corpus did, the worst covering over 7,200 seconds. A cue that long cannot carry a usable timestamp.

split_on_word cuts at word boundaries rather than mid-token.

beam_size=5 — not the server's greedy default

whisper-server defaults to greedy and whisper-cli to beam_size=5; both run strategy = beam_size > 1 ? BEAM_SEARCH : GREEDY. Adopting the server without sending this field silently put the pipeline on greedy.

Greedy's characteristic failure is repetition: it locks onto a phrase and emits it for minutes. It hit 0.7%–5.0% of episodes per show across the first eleven regenerated shows, and a repair pass running beam has fixed 57 of 57 of them, most on the first attempt.

A paired trial on one show — same audio, same model, same fields, only beam_size moved — found beam equal or better on healthy material too:

greedy beam
median punctuation/word 0.1546 0.1611
sub-threshold loops cleared — 5 of 5
clamped / unpunctuated episodes 0 0

and it rescued a shredded episode outright, 0.74 s/cue to 2.42 with punctuation 0.109 → 0.208.

The costs are real but small: roughly 30% more decode time, and about 12% fewer cues. Coarser cues used to matter because the cue was the floor on boundary precision; with per-word times in words.jsonl.gz it no longer is.

Set beamSize = 1 for greedy.

Platform acceleration

Acceleration is determined by how the whisper.cpp server binary was compiled:

  • macOS (Apple Silicon) — Metal acceleration, built in by default
  • Linux — CUDA acceleration if built with WHISPER_CUDA=1; CPU fallback otherwise

Transcribing

With pipeline/.env filled in (see Pipeline configuration below), no arguments are needed:

./transcribe

./transcribe builds the distribution and runs it, so a source change can never leave you running a stale binary. It passes its arguments straight through and exits with the pipeline's own status, and it resolves .env and relative paths against your current directory — it is the long form below with the path memorised for you:

./gradlew :pipeline:installDist
./pipeline/build/install/pipeline/bin/pipeline

./gradlew :pipeline:run also works, with the caveats under Arguments below.

Every .env setting also has a flag, which takes precedence over it — run ./transcribe --help for the list.

Running two instances at once

Give each instance its own pods.yaml and its own whisper.cpp server. In one terminal:

./transcribe --config ~/pods-a.yaml --whisper-url http://localhost:8081

and in another:

./transcribe --config ~/pods-b.yaml --whisper-url http://localhost:8082

Don't use ./gradlew :pipeline:run for this — two concurrent Gradle invocations in one checkout serialise on the project lock. ./transcribe takes that lock only for its build and has released it by the time the pipeline starts, so a second launch waits a second at most.

Both can share one data directory, but nothing coordinates them: list a show in only one of the two files. If both walk the same feed they will download the same episode to the same audio.mp3.part staging file, and whichever finishes first deletes the other's.

Indexing

python3 index.py /path/to/data_directory

# Specify a custom database path
python3 index.py /path/to/data_directory --db /path/to/podcasts.db

# Read every transcript again, instead of only what changed
python3 index.py /path/to/data_directory --full

The database defaults to podcasts.db inside the data directory. Designed to run directly on the machine hosting the files to avoid network filesystem overhead.

What a re-index actually reads

The first run against a database reads everything. After that it stats each transcript.json and opens only the ones whose modification time or size has changed, along with any it has not seen before; rows whose file has gone are deleted. Reading the files is the expensive part on a network share, so a run that finds nothing new finishes in about the time it takes to walk the tree.

Each row records the file it came from, and a second table records the files that were walked but deliberately not indexed — no _id, no transcript, or unreadable — so those are not reopened on every run either. They are picked up as soon as they change.

episodes_fts is an external-content table, so it is maintained alongside: rows are deleted from and inserted into the index by hand as their episodes change, or the whole index is rebuilt in one go once more than a fifth of the corpus is affected.

Three things are worth knowing:

  • A rewrite that lands with both the same modification time and the same size as the version already indexed is not noticed. Filesystems with coarse timestamps make this possible in principle; --full is the repair.
  • An empty data directory leaves the database untouched rather than emptying it, which is what an unmounted share looks like. So does a run that would remove every indexed episode — a restore that leaves the transcripts truncated but freshly stamped looks exactly like a corpus that has legitimately gone. --full is how you say you meant it.
  • A file that cannot be read or stat'ed this time is left as it was, rather than treated as deleted, so a share that blinks costs nothing. The next run picks it up. A rebuild keeps only what it reads, so it refuses outright rather than dropping what it could not look at — unless nothing is indexed yet, where there is nothing to lose.

--full rebuilds both tables from every transcript, and is taken automatically when the database predates incremental indexing.

Serving the web UI

Copy .env.example to .env and fill in your values (.env is gitignored):

cd web
cp .env.example .env
APP_DB_PATH=/path/to/podcasts.db
# Required. The pipeline's data directory, served over HTTP. Word timings are read from under it.
APP_DATA_URL=http://your-nas:9280
# Optional. The pipeline's audio directory, served over HTTP; defaults to APP_DATA_URL.
#APP_AUDIO_URL=http://your-nas:9281

Quarkus picks up .env automatically. Alternatively, override properties inline:

Development (live reload on template/code changes):

./gradlew :web:quarkusDev

Production — build a runnable JAR then launch it:

./gradlew :web:build
java -jar web/build/quarkus-app/quarkus-run.jar

Configuration properties can also be overridden at launch without editing the file:

java -Dapp.db.path=/data/podcasts.db \
     -Dapp.data.url=http://nas:9280 \
     -jar web/build/quarkus-app/quarkus-run.jar

Pages

Path What it is
/search Full-text search with filters for duration, podcast, collection, tag, year and episode type, and a relevance/newest/oldest sort
/episode/{id} One episode: metadata, audio player, and the transcript as clickable cues. ?q= highlights the query's matches and steps between them; #t=<seconds> opens on a cue and starts playback there
/podcasts What the corpus holds: one card per podcast with artwork, episode count, total hours and date range, plus when index.py last wrote the database

Word timings

The pipeline writes words.jsonl.gz beside every episode's audio: one line per word with its start, end and the decoder's own confidence. The web module fetches it from APP_DATA_URL, so that must be an absolute http(s) URL; with a relative one the feature stays hidden.

With it set, the transcript gains a Mark low confidence toggle, and playback highlights the current word rather than only the current cue. Words whisper scored below p = 0.4 are faded when the toggle is on, so a reader can see where the decoder was guessing. Dimming is opt-in: a transcript permanently mottled with faded words is harder to read than one that never shows its confidence at all.

The sidecar is roughly 60 KB per episode, so it is never fetched on page load — the first play pulls it, and so does ticking the toggle. Episodes transcribed before word timestamps were emitted have no sidecar; the route returns 404 and the toggle says No word timings rather than sitting there doing nothing. Three other cases say something more useful than that:

  • a query that matched every line leaves nothing to split, so the toggle says Hidden on matched lines and stays live — clearing the query brings the words back;
  • a sidecar written against a different decode of the episode says Word timings out of date, which re-running index.py over the new transcript fixes. It is judged whole rather than line by line: one cue in common, by ordinal or by coincidence, is not agreement;
  • whisper does not always time every word, and the pipeline drops the ones it did not, so a line the sidecar cannot rebuild keeps its own text. If that is every line, the toggle says Word timings incomplete.

A line is only ever split when its words rebuild it exactly, so the transcript itself is never rewritten by the sidecar.

The file is fetched by the web module from APP_DATA_URL and served from /episode/{id}/words, rather than fetched by the page from the data host: it needs no CORS grant on a server that only has to serve files, and Content-Encoding: gzip lets the browser inflate it instead of the page carrying a decompressor. It is sent gzipped whatever the request's Accept-Encoding says, since the file is only gzip on the data host — curl it with --compressed.

The URL is APP_DATA_URL, the episode's directory (the database's relative audio path minus audio.mp3, each segment percent-encoded) and words.jsonl.gz. A path segment of .. or ., or an empty one, is refused rather than sent, so a bad database value cannot walk up out of the data URL. APP_DATA_URL needs a host and no query string or fragment, or the feature stays hidden.

A 404 from the data host is a 404 here, and the toggle says No word timings. Any other outcome — a timeout, a refused connection, a 5xx — is a 502, which the page treats as transient: it says nothing and tries again on the next play.

The response revalidates rather than being held: re-transcribing an episode rewrites the sidecar and the cue ordinals it is keyed to together, and an hour-old sidecar against a fresh transcript mis-times every word. The browser's If-None-Match is passed on to the data host, so an unchanged file is a 304 on both legs and no body crosses either.

The tag is the data host's own ETag when it sends one. A host that sends none (Python's http.server) gets a SHA-256 of the bytes just fetched, which is always right but can only save the browser leg. The trade-off: a host whose ETag is derived from a stat can name a version the body is not, so a same-size file put back with its old mtime would keep the old tag, and no-cache would pin the stale timings in the browser until the file next changed.

On a line the search query matched, the <mark> highlighting wins and the line is not split into words — word timing is the lesser feature there.

JSON API

For reading the corpus from a shell or a notebook:

curl 'http://localhost:8080/api/search?q=climate+change&sort=newest&pageSize=5'
curl 'http://localhost:8080/api/episode/abcd1234'

/api/search takes the same parameters as /search (q, duration, podcast, collection, tag, year, episodeType, sort, page) plus pageSize, capped at 100. It returns the result page with totalCount, totalPages, hasNext and hasPrevious, and each episode carries a plain-text snippetText of the matched passage. The transcript field is absent from that payload rather than null, so a consumer can tell "not included" from "this episode has none"; /api/episode/{id} returns one episode with it, or a 404 with a JSON body.

episodeSummary comes from the feed. A summary containing a < is treated as HTML and sanitised before it goes out, the same rule the episode page uses; anything else is passed through untouched, since running prose through an HTML sanitiser turns its ampersands and quotes into entities. Prose that happens to contain a < — an address in angle brackets, say — is sanitised and loses it, on both the page and the API.

What that sanitising buys is narrow: it strips scripting from markup so the value can be inserted as HTML. It does not make the value safe to drop into an HTML attribute — sanitised markup still contains quotes, and a prose summary is returned with its quotes intact, so either can break out of an unescaped attr="...".

Every other feed-supplied string in the payload is raw — episodeTitle, podcastTitle, snippetText, the links and the tags are whatever the feed said, because an API cannot know how a consumer will render them. Escape everything at render time for the context you are rendering into; the HTML pages here do it with th:text.

A page number far enough past the end returns no rows rather than wrapping around to the first page.

Pipeline configuration

Environment (.env)

Copy pipeline/.env.example to pipeline/.env and fill in your values (.env is gitignored):

PIPELINE_DATA_DIRECTORY=/path/to/download-directory
#PIPELINE_AUDIO_DIRECTORY=/path/to/audio-directory
PIPELINE_WHISPER_SERVER_URL=http://localhost:8080
PIPELINE_CONFIG_PATH=/path/to/pods.yaml
#PIPELINE_VERBOSE=true

PIPELINE_AUDIO_DIRECTORY is optional. Unset, the mp3s sit beside their transcripts in the data directory. Set, the mp3s go there instead, under the same <podcast>/<episode> layout, and the data directory keeps everything else. Moving an existing tree from one layout to the other is not supported.

PIPELINE_VERBOSE is optional; when set to a non-blank value it overrides the verbose value from pods.yaml. Leave it commented out (or blank) to let pods.yaml decide — note that any value other than true counts as false, so PIPELINE_VERBOSE=false will override verbose: true.

The .env file is resolved relative to the working directory: pipeline/.env is tried first (running from the repo root), then ./.env (running from inside pipeline/, or next to an installed distribution). A missing .env is fine as long as the arguments below supply the three required values.

Arguments

Flag Overrides
--config <path> PIPELINE_CONFIG_PATH
--data-dir <path> PIPELINE_DATA_DIRECTORY
--audio-dir <path> PIPELINE_AUDIO_DIRECTORY
--whisper-url <url> PIPELINE_WHISPER_SERVER_URL
--whisper-model <name> PIPELINE_WHISPER_MODEL, or whisper_model in pods.yaml
--verbose / --no-verbose PIPELINE_VERBOSE
--recover-orphans / --no-recover-orphans recover_orphans in pods.yaml
--orphan-limit <n> orphan_recovery_limit in pods.yaml
--quality-retry / --no-quality-retry quality_retry in pods.yaml
--dry-run No equivalent; see Dry run
--dump-feed-markup <url>, --dump-limit <n> No equivalent; see What else a feed carries
--dump-audio-chapters <dir> No equivalent; see Chapters inside the audio
--retranscribe <dir>, --retranscribe-id <hex8>, --retranscribe-list <file>, --retranscribe-flagged, --retranscribe-limit <n>, --retranscribe-force No equivalent; see Re-transcribing an episode
--verify-pairs No equivalent; see Checking pairs
--list-defects No equivalent; see Repairing windows
--repair-gaps No equivalent; see Filling gaps

Precedence is argument, then .env, then pods.yaml. A flag that is not passed falls through, so --whisper-url alone leaves everything else coming from .env.

Under Gradle the flags go through --args. Gradle splits that string itself rather than handing it to a shell, so write $HOME or an absolute path — a ~ inside the quotes arrives at the pipeline literally and the config file is not found:

./gradlew :pipeline:run --args="--config $HOME/pods-a.yaml --whisper-url http://localhost:8081"

A bad flag exits 2 with its message on stderr. Unusable configuration exits 1 — that covers a missing --config, a data directory that is missing or not writable, and a whisper server that does not answer the preflight (see When whisper is not there). A run that found nothing new to transcribe exits 0, so a wrapper script can tell a failed launch from a quiet one.

pods.yaml

  • verbose — enable debug logging (optional, default false; overridden by PIPELINE_VERBOSE when that is set)
  • skip_after_consecutive — stop walking a feed once this many consecutive already-transcribed episodes are seen (optional, default 20)
  • exclude_title_keywords — titles matching any of these are skipped for every feed (optional, see Non-content exclusions; set to [] to disable)
  • min_episode_duration_seconds — skip episodes shorter than this (optional, default 150; set to 0 to disable)
  • recover_orphans — transcribe episodes that aged out of their feed before they were processed (optional, default true; see Orphan recovery)
  • orphan_recovery_limit — at most this many orphans per run, across all podcasts (optional, default 0, meaning no limit)
  • language — ISO 639-1 code whisper decodes in, or auto to detect from the audio (optional, default en). Case does not matter; it is lower-cased before being sent. A code whisper does not recognise fails the run at startup rather than being sent: whisper_lang_id returns -1 for an unmatched code and its caller adds that to the language token's base index, so the decode would proceed in the wrong language with nothing reporting it — and the quality gate cannot catch it, since a wrong-language decode can be fluent, punctuated and loop-free. A region tag (en-US), a 639-2 code (eng) and a stray space are all rejected on those grounds. A language name (english) is rejected too, but for a different reason: whisper does accept it, and would decode correctly, but the initial prompt is matched on the code — so a name would silently drop the one lever that fixes unpunctuated decodes
  • initial_prompt — the prompt sent with this feed's decodes, written in its language (optional; the top-level default is English prose). A podcast that sets it must set language too, and neither may be auto. See prompt and carry_initial_prompt
  • quality_retry — decode a flagged transcript a second time and keep the better one (optional, default true; see Transcript quality gate)
  • notify_url — POST the run's summary line here when a run finishes (optional; see Run report)
  • podcasts — list of RSS feeds to process, each with name, url, optional collections, optional excludes, an optional min_episode_duration_seconds that overrides the global floor, and an optional language that overrides the global one

name becomes the show's directory name, so changing it moves every episode of that feed. Re-capitalising it used to create a second directory for the same show; the pipeline now reuses a directory that differs only by case, and logs when it does. Renaming it any other way still starts a fresh directory and re-transcribes the feed.

Non-content exclusions

Two filters run before an episode is downloaded, so excluded episodes cost nothing.

exclude_title_keywords matches whole words, case-insensitively, against the episode title. Whole-word matching is what makes the list safe to apply globally: a substring match on repeat also swallows "Repeating FRB Mystery", and on archives it swallows "Inside the Archives", an actual interview series. The default list is trailers, cross-promos and repeat markers: trailer, introducing, encore, classic episode, rewind, re-release, re-run, rerun, rebroadcast, best of, repeat, replay, from the archives. Setting the key replaces the list rather than adding to it; the per-podcast excludes list is separate, still a plain substring match, and still applies on top.

coming soon is deliberately not in the default list. It reads as a safe global term but matches Planetary Radio's "2012 DA14--Coming Soon to a Planet Near You!", a real 29-minute episode. Every genuine hit for it sits in one of two feeds, so it belongs in their excludes rather than the global list.

min_episode_duration_seconds uses the feed's itunes:duration. An episode whose feed omits the tag is never filtered on length. The 150 default sits at the point where short-form content starts to outnumber promos — below it a feed is almost entirely trailers, hiatus notices and "coming soon" stubs. Feeds that publish genuine short-form episodes need the floor lifted per podcast:

- name: A Short-Form Show
  url: https://example.com/feed.rss
  min_episode_duration_seconds: 0

Duplicate GUIDs

_id is md5(guid)[:8], and the episode table is keyed on it. Publishers do occasionally ship two entries under one GUID — HBR IdeaCast has two such pairs — which under INSERT OR REPLACE silently collapsed them into a single row, losing one episode from search while its transcript sat on disk.

index.py now suffixes the later members of a clashing group (<id>-2, -3, …), ordered by audio path so the ids are stable across re-indexes, and warns on stderr naming every episode involved. Both episodes stay searchable and keep a working /episode/{id} permalink. Widening the hash would fix nothing here — the inputs really are identical — and would rename every episode directory, forcing a full re-transcription.

Skip heuristic

Feeds are typically ordered newest-first. Rather than stat'ing every episode directory (expensive for feeds with thousands of entries), the transcriber walks the feed and stops on a podcast once it sees skip_after_consecutive transcribed episodes in a row. The counter resets on any gap, so a cancelled run that left untranscribed holes will be picked up on the next invocation.

Orphan recovery

The pipeline only transcribes what the feed offers, and publishers age entries out on hard caps or date cutoffs. An episode downloaded but not yet transcribed when that happens is stranded: its directory holds an audio.mp3 no later run will ever look at. 659 of 17,517 episode directories were in that state when this was written.

After walking a feed, the pipeline lists the podcast's directory once and compares each episode directory's YYYY-MM-DD-<hex8> prefix against the feed. Directories whose prefix is absent are opened and, if they hold audio but no transcript, transcribed in place — newest first, and never renamed, because indexed rows already point at the existing path.

The feed still describes the show, so every podcast_* field is accurate. The entry is gone, so only what the directory name carries can be recovered:

Recovered Source
_id the <hex8> in the directory name — the same id the episode always had
episode_published_on the date in the directory name
episode_title the rest of the directory name, dashes back to spaces (lossy: punctuation is gone)
episode_duration where the decoded speech ends — approximate, and short of the file by any trailing silence
all_tags feed-level categories only
episode_metadata_recovered true, present only on recovered files

episode_audio_link, episode_web_link, episode_image, episode_summary, episode_subtitle, episode_authors, episode_number, episode_season and episode_type are written as null. The web UI hides a null field and would render an empty string as a dead link, so the distinction matters.

Audio that decodes to nothing leaves a recovery-failed marker in the directory. Without it the file would be re-uploaded to whisper on every run forever — an orphan has no feed entry whose download could fail and stop it. A transcriber error (server down, timeout) writes no marker and is retried.

Two situations are reported as errors rather than recovered, because the feed entry has better metadata than the directory name ever could:

  • a directory whose id is in the feed under a different date — a publisher re-issued or re-dated an old episode
  • an episode still in the feed, still untranscribed, that skip_after_consecutive stopped before reaching. Raise the threshold to pick it up.

Recovery is on by default and costs about 20 seconds across a full run, because only the directories absent from the feed are opened. --no-recover-orphans turns it off; --orphan-limit <n> bounds how many a single run will transcribe, which is worth setting when the backlog is large enough to crowd out new episodes.

Transcript quality gate

Three failure modes were previously only ever found by an external pass reading the finished corpus: an episode decoded with no punctuation at all, a greedy-style repetition loop, and cues shredded into one- and two-word fragments. The pipeline now scores every transcript as it writes it, and records the score in transcript.json under episode_quality:

"episode_quality": {
  "punctuation_per_word": 0.1611,
  "seconds_per_cue": 2.42,
  "repeated_share": 0.0,
  "longest_repeated_cue_run": 1,
  "mean_word_probability": 0.87,
  "low_confidence_share": 0.04,
  "word_count": 8123,
  "cue_count": 1044,
  "flags": []
}

flags is empty for a healthy decode, and otherwise names what tripped:

Flag Trips when Measured
no-speech the decode produced no words at all every other check needs words to measure, so without this an empty decode scores clean
unpunctuated punctuation per word below 0.03 healthy episodes sit near 0.15
shredded-cues under 1.0 seconds per cue, with at least 50 cues a shredded episode measured 0.74 against 2.42 re-decoded
repetition-loop one 4-gram repeats over 5% of the words, or 4+ consecutive cues are identical greedy decoding hit 0.7%–5.0% of episodes per show
low-confidence over 20% of words scored under p = 0.3 skipped entirely for episodes decoded before word timestamps existed

When a decode is flagged the pipeline decodes the episode once more and keeps the better of the two: a decode that produced speech always beats one that did not, then fewer flags wins, and on a tie the more punctuated one. Whisper is not deterministic, and the repair passes that inspired this cleared 57 of 57 repetition cases, most on the first re-decode. A retry doubles decode time for the 1–5% of episodes that trip a flag.

no-speech is why the retry is a comparison rather than a preference: an empty decode trips none of the other checks, so without a flag of its own it would score clean, win on flag count, and replace a real transcript. One flag is not enough on its own, though — it still beat a decode bad enough to trip two — so an empty decode loses to any decode with speech in it before flags are counted at all. A poor transcript is worth more than none, and on the recovery path none is permanent: an episode that decodes to nothing gets its one retry and then recovery-failed abandons it for good.

If the kept decode is still flagged it is written anyway, with its flags recorded, and a warning goes to the error log — the transcript is still worth having, and episode_quality.flags is what a later re-transcription pass selects on.

Turn the retry off with --no-quality-retry, or quality_retry: false in pods.yaml.

The thresholds above are starting points measured on the real corpus, not tuned constants. They live together at the top of pipeline/src/main/kotlin/com/rsstowhisper/pipeline/TranscriptQuality.kt, each with the number it came from.

Decoding without history

decode_without_history: true in pods.yaml makes the first decode of every episode send max_context=0, so whisper conditions no window on the text before it. That text is what feeds repetition loops and stretch-copies. whisper.cpp skips the initial prompt along with the history, so no prompt is sent, and a decode that comes back flagged is retried with the prompt and the history; the better of the two is kept.

Measured on the same audio, same model and server, 2026-09-25:

sample request clean loops stretch-copy unpunctuated words
24 random episodes default 18 5 1 0 1.00
max_context=0 23 0 0 1 1.01
16 with stretch-copies or loops default 9 3 6 0 1.00
max_context=0 15 0 0 1 1.02

The cost is punctuation: median marks per word fell from 0.140 to 0.114 without the prompt. It is off by default.

Repairing windows

./transcribe --repair-windows --retranscribe-list episodes.txt

Re-decodes only the stretches of each target around its defects, instead of the whole episode, and splices them into the pair on disk. A defect is a stretch-copy, a loop (four or more identical cues, a few cues repeated in turn, or one cue repeating a phrase faster than anyone speaks), an echo or long copy of the cues before it, or a cue of 10 s or more that is nothing but a sentence of the initial prompt or one of whisper's stock phrases ("Thanks for watching."): over music or silence, a decode voices them.

  • Windows are the defective cues plus two good cues either side. whisper-server decodes just that stretch (offset_t, duration); nothing is cut from the mp3.
  • Anchors. The outer good cue at each end is kept exactly as it was; the one against the defect is re-decoded, since it is the cue most often damaged. The new decode is cut after the last three words of the left anchor and before the first three of the right one, found within 2 s of their old time. Otherwise the decode's words are walked over the anchor's, past a word the anchor does not have, and only then cut by time, in the clock the other anchor's words set. The new window's clock is mapped onto the old one between the two anchors.
  • Attempts. Prompted, then without history (max_context=0), then both again at temperature 0.2, then both again with a beam of 8: whisper can skip speech on one decode and keep it on another, but since v1.9.4 only if the request differs. A window is applied only if it comes back with fewer defects than it had, counting any sentence of the prompt it now holds; an attempt holding a prompt sentence as long as a leak is refused outright, as is one that drops a cue the base had right where speech is heard under it.
  • Lost speech. An attempt that leaves more than 2 s of speech without a word near it is refused: whisper can skip a whole window and carry on, and the words either side still match. Without VAD, the speech is the base's own words in the window's good cues.
  • Silence. With vad_binary and vad_model set in pods.yaml, Silero VAD (whisper.cpp's whisper-vad-speech-segments) says where speech is instead. Cues the new decode put over non-speech are dropped, keeping one ♪ where it marked music, and a window may come back empty only where VAD hears nothing. A VAD that cannot run stops the batch; one that hears nothing under most of an episode's cues is ignored for it. The server's own VAD cannot be used for this: with vad on it applies offset_t to VAD-compressed time.
  • Only a pair already from one decode is repaired. The result gets a new run id; whisper_run.base_run is the pair it was spliced into and whisper_run.repairs records each window, its anchors and gaps, and what was refused.

--list-defects prints every episode the repair would find something in, with counts by kind, as a list that works with --retranscribe-list. It only reads.

Filling gaps

./transcribe --repair-gaps --retranscribe-list episodes.txt

whisper decodes in 30 s windows, and when the model ends one early, whisper.cpp skips the rest of it ("single timestamp ending"). Over an ad jingle, music or another language the model says so wrongly, and the speech later in the window is lost: nothing is written, or a filler such as "Thank you." is held to the window's end. Re-decoding the same window from the same cue edges meets the same wall.

  • Gaps are stretches VAD hears with no word starting within 1 s, 4 s or more of them, from the cue before to the cue after. Filler cues over a gap, and crammed cues at its edges (whisper often crams part of what it skipped into no time at the skip), are replaced with it. A gap at an episode's very start or end has no anchor on that side and is skipped.
  • Chunks. The speech between the anchors is cut and merged from VAD's spans into pieces under 30 s, each decoded alone and starting on speech, the way WhisperX avoids the skip.
  • Attempts. Unprompted first, then prompted, then warmer and with a wider beam. A prompt on every short chunk invites skips and echoes; unprompted chunks sometimes come back without punctuation, so a punctuated attempt within a second of the best coverage wins.
  • Kept only if it covers at least 2 s more of the speech, keeps every heard cue around it, holds no defect of its own, no prompt sentence, and is not mostly the text of the cues either side (the next line written early over music). Recorded with "kind": "gap".
  • Needs vad_binary and vad_model. It can run with --repair-windows in the same pass; a gap inside a defect window is left to that window.

Re-transcribing an episode

Redoing an episode used to mean deleting its transcript.json by hand. With the quality gate recording flags, the loop closes:

./transcribe --retranscribe-flagged --retranscribe-limit 50
./transcribe --retranscribe "Ask-a-Spaceman/2024-01-02-abcd1234-some-episode"
./transcribe --retranscribe-id abcd1234
./transcribe --retranscribe-list episodes.txt

--retranscribe-list takes one <podcast>/<episode> path or id per line. Blank lines and # comments are skipped and anything after a tab is ignored, so the output of --verify-pairs can be passed straight back in. It counts as naming its targets, so --retranscribe-force applies to it.

A path is the episode's directory as it sits on disk, so it carries the escaped podcast name (Ask-a-Spaceman, not Ask a Spaceman) — every character that is not a letter or a digit became a dash when the directory was created.

--retranscribe and --retranscribe-id are repeatable, and any of the three skips the feeds entirely — every target is already on disk, and an episode that aged out of its feed has nothing left to fetch. --retranscribe-flagged reads every transcript.json under the data directory, which is slow on a network volume; --retranscribe-limit caps it, since a first pass over the corpus can select hundreds.

The limit is a rolling window, not a fixed prefix. Each attempt writes the time into a retranscribe-attempted file in the episode directory, and the flagged scan takes the least recently attempted first — never-attempted episodes ahead of everything, then oldest attempt, with the directory name breaking ties so a run over the same corpus picks the same episodes. Without that, an episode whose flags cannot clear — genuinely bad audio, a music-heavy show, one that is mostly silence — would sit at the front of the sorted list forever and starve everything behind it, costing decode time on every run and changing nothing.

The time is recorded for every attempt, not every success, because the episodes that starve the selection are exactly the ones each re-decode refuses. The exception is a decode that could not reach the whisper server: nothing was learned about that episode, so it keeps its place rather than rotating away untried — otherwise a server that died partway through would send the rest of the window to the back of the queue unexamined.

Each target needs its audio.mp3; one without it is reported and skipped. The rewrite keeps every existing field and replaces only episode_transcript and episode_quality — all the rest came from a feed entry that may no longer exist, so what is on disk is the best there is. episode_duration is replaced too, but only when episode_metadata_recovered is true, because that number came from the previous decode rather than from the feed.

--retranscribe-force keeps the new decode whatever it scores. Use it when you are redoing an episode for a reason the score cannot see — a corrected language, a fixed initial_prompt, a model change. The comparison counts flags and then compares punctuation; it knows nothing about which language the old decode was in, so a wrong-language transcript that happens to score a hair better would otherwise win.

It only applies to explicitly named targets, and is refused outright alongside --retranscribe-flagged: that scan is the one that runs at scale, and it must never be able to make the corpus worse. A flag that silently covered only the named half of a mixed run would be worse than one that says so.

Force skips the quality comparison and nothing else. An episode with no audio, a decode that found no speech, and one that came back without word timestamps are all still refused — force means "I know better than the score", not "ignore a server that stopped sending token_timestamps".

The no-speech refusal is a check in its own right rather than a consequence of the comparison, because a blank transcript is not caught by the emptiness test ahead of it: segments carrying timestamps and no text render a VTT that is not blank.

It needs a target. --retranscribe-force on its own is refused rather than ignored, since the run would otherwise fall through to an ordinary feed run — which is what a script whose target list came out empty would get, having asked for the opposite.

Flags are counted only where both decodes could have raised them. low-confidence needs word probabilities, so a decode scored without them cannot trip it and would otherwise win on raw count against one that did — keeping a transcript with no word timings forever, since every re-decode that finally produced them would be discarded. With the flags level, the decode that has word timings wins; it is worth having, but not worth trading a better-scoring transcript for.

A re-decode is kept only if it scores at least as well as the transcript it would replace, by the same measure the quality gate's retry uses — fewer flags, then better punctuation, then word times. Whisper is not deterministic, so a redo can come back worse than what it overwrites, and that write is the only copy: re-transcribing can improve an episode or leave it alone, never cost it the better decode. A transcript written before the quality gate has no score to compare against, so it is simply replaced. Nor is a transcript whose words.jsonl.gz came from another decode: its score describes one half of a pair neither half of which is usable, so it is replaced whatever it scores.

After the batch every target is checked with the same test as --verify-pairs, written or not, and the run fails listing any whose pair still disagrees.

A re-decode that comes back with no word timestamps at all is refused outright when the decode on disk had them — judged by its recorded score rather than by whether words.jsonl.gz is there, since writing the sidecar is allowed to fail. The comparison below only reaches word times once the flags and the punctuation are level, so a decode that lost them can still win outright on a flag the other tripped — and would take the words.jsonl.gz with it. It means the server ignored token_timestamps, and the warning says so.

Otherwise both files are staged beside their originals and moved into place, with the sidecar absent for the whole swap: cleared first, restored only once the transcript it belongs to is there. words.jsonl.gz addresses cues by position, so either file left beside the other's transcript mis-times every word — silently and for good, since an episode that has a transcript.json is one nothing revisits. Interrupted anywhere in between, the episode is left visibly missing a sidecar instead of quietly holding the wrong one. If the new sidecar cannot be written at all — a full disk, an I/O error — the episode is left exactly as it was rather than committed without one.

Staged files are named per process, because two instances sharing a data directory select the same episodes: --retranscribe-flagged scans the whole tree, whatever podcasts its config names. That stops one run moving another's half-written file, and nothing more — the swap is two moves, not one, and nothing here locks, so do not point two runs at the same episode: they can interleave into a transcript and a sidecar from different decodes.

An id is part of a path, not a unique key, so --retranscribe-id redoes every copy it finds rather than guessing which was meant.

Dry run

--dry-run walks every feed and applies every filter exactly as a real run does, then reports what it would have done instead of doing it:

Would transcribe Ask a Spaceman/2024-03-14-a1b2c3d4-what-is-dark-matter
Ask a Spaceman: would transcribe 4 episodes
Would recover Ask a Spaceman/2019-01-01-deadbeef-aged-out
Dry run: would transcribe 12 and recover 3 episodes

Nothing is downloaded, nothing is decoded, and no podcast or episode directory is created — useful because the global exclusions, per-podcast excludes, the duration floor, skip_after_consecutive and orphan recovery interact, and reading the result is cheaper than spending GPU hours on a misconfiguration.

The whisper server is never contacted, so a dry run works with it switched off. It exits 0. Combined with --retranscribe, it lists the episodes that would be redone without touching their transcripts — worth doing before --retranscribe-flagged, which selects across the whole corpus.

What else a feed carries

--dump-feed-markup <url> fetches one feed and prints, per entry, everything ROME parsed that no registered module claimed — which is every namespace the pipeline does not map. --dump-limit caps the entries sampled (default 10). It needs no pods.yaml, no data directory and no whisper server, and exits without running anything:

./transcribe --dump-feed-markup https://example.com/feed.rss --dump-limit 10
feed: Some Podcast
modules: http://www.itunes.com/dtds/podcast-1.0.dtd
channel foreign markup:
    <podcast:medium>podcast</podcast:medium>

10 of 300 entries

[1] An Episode
  published: Wed Sep 10 09:00:00 UTC 2026
  modules: http://www.itunes.com/dtds/podcast-1.0.dtd
  foreign markup:
      <podcast:chapters url="https://example.com/ep1/chapters.json" type="application/json+chapters">
      <podcast:soundbite startTime="1234.5" duration="60.0">A clip</podcast:soundbite>

distinct foreign elements across the sample:
  10 x podcast:chapters
  3 x podcast:soundbite

The closing tally is the point: a tag that only some episodes carry is easy to miss one entry at a time. Elements a module did claim — anything itunes:, dc:, media: — are absent from the markup and named in the modules line instead, so the report separates what is already reachable through ROME from what would need new parsing.

Chapters inside the audio

--dump-audio-chapters <dir> reads the ID3v2 tag of every audio.mp3 under a data directory and reports which episodes carry CHAP chapter frames. It never touches the audio itself -- only the tag at the head of each file -- so a pass over a large library costs little, and it needs no network:

./transcribe --dump-audio-chapters /path/to/data --dump-limit 10
scanned audio.mp3 in 1284 episodes
with ID3 chapters: 37 (2.9%)

showing 10 of 37

Some Podcast/2026-01-02-a1b2c3d4-an-episode
  0:00:00-0:01:32  Advertisement
  0:01:32-0:24:10  Part One

chapter titles by episodes carrying them:
  37 episodes, 94 uses  Advertisement
  37 episodes, 37 uses  Intro
  1 episodes, 1 uses  The Chemical Barons

The ordering is the point. A title carried by many episodes is a structural segment -- an intro, a break, a sponsor read -- while a title carried once is that episode's own content, so ranking by episode count puts the recurring furniture at the top.

Chapters read here describe the file on disk, which is the same file the transcript was decoded from. That matters because a feed's own chapter timestamps describe the unstitched master: where a host inserts ads at download time, the two disagree.

When whisper is not there

The server is asked for its base URL once, before the first feed is fetched; nothing listening, and the run exits 1 having downloaded nothing. Any answer counts, a 404 included — --request-path moves whisper.cpp's page off /, and a proxy may route only /inference, neither of which is a reason to refuse to run. Only a gateway saying its upstream is gone (502, 503, 504) or nothing answering at all fails it.

If whisper goes away during a run, the run carries on to the end and then exits non-zero, naming how many episodes could not reach it. It does not stop early: the downloads are not wasted, since a completed audio.mp3 is kept and the next run skips straight to decoding it. What the exit code buys is that a run which downloaded the whole backlog and transcribed none of it cannot look successful to whatever is driving it.

Only failures that say nothing is there count — a connection that goes nowhere or dies mid-response, a gateway 502/503/504, or an empty body. A status the server chose is the server talking, and what it is talking about is that request: whisper.cpp answers 400 for an mp3 it cannot decode and 500 for one it cannot process, and those episodes fail on their own merits as they always did, without failing the run.

Run report

The error log says whether anything went wrong. It does not say what was done, so every run also writes <data-dir>/logs/run-<yyyyMMdd-HHmmss>.json and copies it to logs/latest-run.json — the stamped file is the history, the stable name is what anything watching the directory can read without listing it first.

{
  "started_at": "2026-09-14T02:00:03Z",
  "finished_at": "2026-09-14T06:41:55Z",
  "duration_seconds": 16912,
  "totals": { "transcribed": 41, "recovered": 3, "failed": 1, "skipped": 112 },
  "podcasts": {
    "Ask a Spaceman": {
      "transcribed": 4,
      "recovered": 0,
      "failed": 0,
      "skipped": { "global_keyword": 2, "too_short": 1 }
    }
  }
}

The skip counts are by the reason the pipeline actually applied, which is what answers "why did this feed transcribe nothing this week". failed counts episodes this run tried for and did not get, whether they fell over before the download or during the decode.

Reports are kept for 14 days, matching the error log. Two instances sharing a data directory each write their own stamped file; latest-run.json is whichever finished last. A report that cannot be written is a warning and never fails the run, and it is written even when a run dies part-way, since that is when knowing how far it got matters most.

Set notify_url in pods.yaml to have the run POST its summary line — the same Run finished: N warnings, M errors that goes to stdout — as text/plain when it finishes. That is exactly what ntfy.sh expects, and other services can adapt. There is no auth and no retry: the run has already finished by then, so a failed notification is a warning and nothing more.

Error log

Warnings and errors are mirrored to <data-dir>/logs/pipeline-errors.log, rolling daily and keeping two weeks, and the run prints a count of both when it finishes. A 13,000-episode run otherwise buries its failures in scrollback.

Output structure

For each episode, the following files are created:

{data_directory}/{podcast_name}/{YYYY-MM-DD-<hex8>-episode-title}/
    audio.mp3                      # Downloaded audio, kept as-is and served by the web module
    transcript.json                # Full metadata + WebVTT transcript
    words.jsonl.gz                 # One line per word, with its own start/end
    recovery-failed                # Only when recovery found no speech in the audio

With PIPELINE_AUDIO_DIRECTORY set, audio.mp3 moves to {audio_directory}/{podcast_name}/{same episode directory}/ and the rest stays put.

The episode_transcript field in transcript.json is a raw WebVTT string, rendered from the same verbose_json response that produced words.jsonl.gz. episode_quality alongside it records how that decode scored; see Transcript quality gate.

words.jsonl.gz is written before transcript.json, because the latter existing is what marks an episode done — so a crash between the two leaves the episode to be redone rather than permanently without its sidecar. A decode that could not write its sidecar leaves the episode undone for the same reason — writing transcript.json anyway would mark done an episode that will never get one. A decode that simply carries no word times, from a server that ignored token_timestamps, still writes its transcript; otherwise no episode could ever complete against such a server.

Both files carry the decode's run id: whisper_run.run_id in transcript.json, and a run field on every line of words.jsonl.gz — per line rather than as a header row, since every existing reader takes each line to be a word. whisper_run also records when the decode ran, the server, the model (from --whisper-model, since whisper does not report it), every form field sent, the audio's SHA-256 and size, the pipeline's git describe, and how many words the sidecar holds. Both files are staged in full before either is moved into place, and the pair is checked on disk after every write.

Checking pairs

./transcribe --verify-pairs > diverged.txt

Reads every episode under the data directory and prints one <podcast>/<episode>, a tab, and the reason for each whose transcript.json and words.jsonl.gz did not come from one decode, exiting 1 if there are any. It never contacts whisper. Pairs with run ids are compared on them; older ones by whether each cue's words rebuild that cue, the test the web player applies before it trusts a sidecar. A transcript from before word timings, with no sidecar and no run id, is counted but not listed. A pair that could not be read just then is listed as "could not be checked" and also fails the check; nothing replaces a pair on the strength of a failed read. Logs go to stderr, so the list can be fed straight to --retranscribe-list. A data directory that is missing, or a directory that cannot be listed, fails the check rather than passing with nothing checked.

The <hex8> in the directory name is md5(entry.uri) truncated to 8 characters. It is part of a path, not a unique key: date and title slug disambiguate it.

Testing

./gradlew test

Linting

./gradlew ktlintCheck

To auto-format:

./gradlew ktlintFormat

About

Transcribe podcasts with whisper.cpp

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages