Skip to content

feat: put transcription on one contract and route every Whisper call through it - #357

Merged
kiyeonjeon21 merged 1 commit into
mainfrom
feat/transcription-contract
Oct 5, 2026
Merged

kiyeonjeon21 merged 1 commit into
mainfrom
feat/transcription-contract

Conversation

@kiyeonjeon21

Copy link
Copy Markdown
Contributor

Transcription counterpart of the video, image (#355), and speech (#356) provider work.

Why

Twelve places in the CLI built their own Whisper client: create, initialize, read the file, wrap a Blob, transcribe. jump-cut posted to the OpenAI API itself with a hardcoded "whisper-1". Whisper also sent every upload named audio.webm, whatever the real format was, and OpenAI reads the format from that extension.

What changes

  • Transcriber contract (ai-providers/src/transcription): transcribeAudio(request) returns a timed transcript or throws ProviderError, and createTranscriber(provider, key) opens one.
  • Whisper on the contract: WhisperProvider implements it over providerRequest.
    • It takes its model from the catalog.
    • It refuses audio over Whisper's 25 MB limit before sending.
    • It names the upload from the audio's magic bytes: WAV, MP3, M4A/MP4, OGG, FLAC, WebM.
  • One CLI helper: transcribeAudioFile(path, ...) replaces all twelve call sites (audio, captions, highlights, edit, scene narration) and jump-cut's raw request.
  • Missing file: vibe audio transcribe on a missing file now returns NOT_FOUND instead of API_ERROR.

Verification

  • Tests: a new transcription contract suite covers filename detection for each container, word and segment timings on the catalog default, the 25 MB refusal before sending, and a bad key as auth. The Whisper network-error test now reflects the shared retries.
  • Checks: pnpm build, lint, typecheck, tests (providers 244, CLI 1343 including integration, MCP 73), providers:check with 0 warnings, pre-push gate.
  • Live:
    • vibe audio transcribe openai.mp3 gave the exact sentence.
    • The same for an extensionless WAV, whose format was detected from its bytes.
    • A missing file gives NOT_FOUND.

https://claude.ai/code/session_011S1BLTaFctZNQzGERpkCuA

…through it

Twelve places in the CLI built their own Whisper client (create, initialize,
read the file, wrap a Blob, transcribe), and jump-cut posted to the OpenAI
API itself with a hardcoded "whisper-1".

- `Transcriber` (ai-providers/src/transcription): `transcribeAudio(request)`
  returns a timed transcript or throws `ProviderError`;
  `createTranscriber(provider, key)` opens one. WhisperProvider implements
  it over `providerRequest`, takes its model from the catalog, refuses
  audio over Whisper's 25 MB limit before sending, and names the upload
  from the audio's magic bytes (it sent every file as "audio.webm").
- `transcribeAudioFile(path, ...)` in the CLI replaces all twelve call
  sites (audio, captions, highlights, edit, scene narration) and jump-cut's
  raw request.
- `vibe audio transcribe` reports a missing file as NOT_FOUND instead of
  an API error.

Claude-Session: https://claude.ai/code/session_011S1BLTaFctZNQzGERpkCuA
@vercel

vercel Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
vibeframe Ready Ready Preview Oct 5, 2026 7:43am UTC

Request Review

@kiyeonjeon21
kiyeonjeon21 merged commit 4ed3e3c into main Oct 5, 2026
6 checks passed
@kiyeonjeon21
kiyeonjeon21 deleted the feat/transcription-contract branch October 5, 2026 07:48

This branch was successfully deployed

1 active deployment
Preview — 30dfd942 Deployed Oct 5, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant