Repository navigation
feat: put transcription on one contract and route every Whisper call through it - #357
Merged
Merged
Conversation
…through it Twelve places in the CLI built their own Whisper client (create, initialize, read the file, wrap a Blob, transcribe), and jump-cut posted to the OpenAI API itself with a hardcoded "whisper-1". - `Transcriber` (ai-providers/src/transcription): `transcribeAudio(request)` returns a timed transcript or throws `ProviderError`; `createTranscriber(provider, key)` opens one. WhisperProvider implements it over `providerRequest`, takes its model from the catalog, refuses audio over Whisper's 25 MB limit before sending, and names the upload from the audio's magic bytes (it sent every file as "audio.webm"). - `transcribeAudioFile(path, ...)` in the CLI replaces all twelve call sites (audio, captions, highlights, edit, scene narration) and jump-cut's raw request. - `vibe audio transcribe` reports a missing file as NOT_FOUND instead of an API error. Claude-Session: https://claude.ai/code/session_011S1BLTaFctZNQzGERpkCuA
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Transcription counterpart of the video, image (#355), and speech (#356) provider work.
Why
Twelve places in the CLI built their own Whisper client: create, initialize, read the file, wrap a
Blob, transcribe.jump-cutposted to the OpenAI API itself with a hardcoded"whisper-1". Whisper also sent every upload namedaudio.webm, whatever the real format was, and OpenAI reads the format from that extension.What changes
Transcribercontract (ai-providers/src/transcription):transcribeAudio(request)returns a timed transcript or throwsProviderError, andcreateTranscriber(provider, key)opens one.WhisperProviderimplements it overproviderRequest.transcribeAudioFile(path, ...)replaces all twelve call sites (audio, captions, highlights, edit, scene narration) and jump-cut's raw request.vibe audio transcribeon a missing file now returnsNOT_FOUNDinstead ofAPI_ERROR.Verification
auth. The Whisper network-error test now reflects the shared retries.pnpm build,lint,typecheck, tests (providers 244, CLI 1343 including integration, MCP 73),providers:checkwith 0 warnings, pre-push gate.vibe audio transcribe openai.mp3gave the exact sentence.NOT_FOUND.https://claude.ai/code/session_011S1BLTaFctZNQzGERpkCuA