Transcribe any video, audio or podcast URL to text: full transcript with timestamped segments plus SRT/VTT subtitles, 99+ languages auto-detected. Works on YouTube and TikTok links, podcast RSS feeds and direct media files with no API key needed.
Video & Audio Transcriber on Apify →
No credit card required. No commitment. Cancel anytime.
- Click sign up — pick GitHub, Google, or email; takes ~30 seconds
- Open this actor — input is pre-filled with a working example
- Click Start — export results as JSON, CSV, or Excel
Your $5 monthly platform credit is enough to run this actor right away — and again every month — scraping typically several hundred to several thousand results per run, depending on your input.
Proxy support — Route traffic through Apify Proxy or your own proxy group to avoid regional blocks and rate limits.
Export anywhere — Download as JSON, CSV, or Excel. Stream via Apify API, webhooks, or integrations with Make, Zapier, Airbyte, Keboola.
Structured data — Every transcript returns the same schema with consistent field naming. All fields always present — null when unavailable, never omitted.
Data pipeline automation Integrate with your ETL pipeline to collect structured transcripts from the source on a schedule. Export to CSV, JSON, or directly to your database.
Market research Monitor transcripts, track trends, and analyze market dynamics with structured, deduplicated data from the source.
AI / LLM training data Structured JSON per transcript is ready for RAG pipelines, embeddings, and agent workflows.
{
"mediaUrls": [
"https://www.youtube.com/watch?v=jNQXAC9IVRw"
],
"maxMinutesPerItem": 5
}| Parameter | Type | Default | Description |
|---|---|---|---|
mediaUrls |
array | — | Video, audio, podcast-feed or direct-file URLs to transcribe (up to 50 per run). Works with YouTube, TikTok, Instagram, Facebook, X, Rumble, SoundCloud, Dailymotion, Twitch and 1,800+ other sites, podcast RSS/Atom feeds (newest episodes are expanded automatically) and direct media files (mp3, mp4, wav, m4a, flac, ogg, webm, mov). |
language |
string | "" |
ISO 639-1 code of the spoken language, e.g. en, es, de, pt. Leave empty to auto-detect — 99+ languages are recognised. |
translateToEnglish |
boolean | false |
Output an English translation of the speech instead of a transcript in the original language. Subtitles are translated too. |
wordTimestamps |
boolean | false |
Add a words array with start/end times for every word to each segment. Useful for karaoke-style captions and precise clipping. |
maxMinutesPerItem |
integer | 120 |
Per-URL cap on how many minutes of media are transcribed and billed (max 300). Longer media is transcribed up to the cap; you pay only for transcribed minutes. Free-plan runs are additionally limited to 30 minutes of media per run in total. Runs with a cap of 30 minutes or less use 1024 MB of memory, longer caps use 2048 MB. |
maxEpisodesPerFeed |
integer | 1 |
When a URL is a podcast RSS/Atom feed, transcribe this many of the newest episodes. Each episode becomes one output item. |
cookies |
string | — | Contents of a cookies.txt exported from a logged-in browser session. Needed for YouTube's "Sign in to confirm you're not a bot" gate and for age- or region-restricted content. Stored as a secret; leave empty for public media. |
proxyConfiguration |
object | {"useApifyProxy":true,"apifyProxyGroups":["RESIDENTIAL"]} |
Proxy settings for media platforms that limit automated access. The default works for most sources; change it only if a platform keeps refusing downloads. |
excludeEmptyFields |
boolean | false |
Drop null, empty-string and empty-array fields from every output record to keep exports compact. |
Every transcript returns the same 23-field schema. Missing values are null — never omitted.
urlinputUrlsourceTypeplatformtitleuploaderpublishedAtthumbnailUrldurationSecondstranscribedSecondsbilledMinuteslanguagelanguageProbabilitytasktextwordCountsegmentssrtvttsrtFileUrlvttFileUrlerrorscrapedAt
One object per transcript. Here is a real example from a production run:
{
"url": "https://traffic.megaphone.fm/FSI5284996555.mp3",
"inputUrl": "https://feed.syntax.fm/rss",
"sourceType": "podcast-episode",
"platform": "podcast",
"title": "1035: Why everyone is moving to Stylex?",
"uploader": null,
"publishedAt": "2026-09-02T11:00:00.000Z",
"thumbnailUrl": null,
"durationSeconds": 1596.656325,
"transcribedSeconds": 119.1,
"billedMinutes": 2,
"language": "en"
}Truncated — full records contain 23 fields. See Output fields for the complete schema.
Try Video & Audio Transcriber now — $5 free credit, no credit card →
Pay only for what you extract. No subscription required — Apify's free $5 credit covers thousands of results.
| Event | Price (USD) |
|---|---|
| Actor Start | $0.01 |
| Minute of media transcribed | $0.03 |
See the actor on Apify for current pricing.
How much does it cost? Pay-per-event pricing — you only pay for ${NOUNS} extracted. Apify's free $5 credit is enough to run thousands of results before you pay anything.
Do I need an API key or credentials? No. Just sign up for Apify, paste your input, and click Start. No credit card required.
Browse all Black Falcon Data actors →
New to Apify? Create a free account with $5 credit — no credit card required.
- Sign up — $5 platform credit included
- Open Video & Audio Transcriber and configure your input
- Click Start — export results as JSON, CSV, or Excel
Need more later? See Apify pricing.
Black Falcon Data builds production-grade web scrapers for job boards and marketplace data. Browse our full actor catalog at www.blackfalcondata.com.
Last updated: 2026 09