Audio & Video Transcriber — MP3, Podcast & Speech to Text
Pricing
$5.00 / 1,000 audio minute transcribeds
Audio & Video Transcriber — MP3, Podcast & Speech to Text
Transcribe an audio or video file to text straight from its URL — MP3, MP4, WAV, OGG, M4A and podcast episodes. Transcript, timed segments, ready-made SRT and WebVTT subtitles, optional translation, 99+ languages. Silence and music cost nothing. $0.005 per audio minute.
Pricing
$5.00 / 1,000 audio minute transcribeds
Rating
0.0
(0)
Developer
⚡ Blitzdata
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
12 hours ago
Last modified
Categories
Share
Audio & Video Transcriber
Give it a file URL, get the words. This Actor transcribes an MP3, MP4, WAV, OGG, M4A or podcast episode to text: the full transcript, timed segments, ready-made SRT and WebVTT subtitles, and optionally a translation and a one-line hook and summary. Links to Instagram Reels, TikToks and YouTube videos work just as well.
It bills by the audio minute, and only for audio that actually contains speech. A voice-activity gate runs before the model, so a silent stretch, a music bed or a failed URL costs nothing — and you never get a sentence the model invented over a soundtrack.
Audio to text · MP3 to text · speech to text · podcast transcription · video to text · SRT and VTT subtitles · 99+ languages.
What this Actor does
- Transcript of the speech in an Instagram Reel, TikTok, YouTube video or Short
- Timed segments (
start,end,text) and the same blocks as SRT and WebVTT files - Translation into any language you name, from the same model, at no extra charge
- Hook and summary: a one-line hook and a one-sentence summary, when you ask for them
- Who said what: turn on
diarizeand every segment and subtitle line is labelled Speaker 1, Speaker 2 — for interviews, podcasts and panels, at no extra charge - Whole channels: paste a YouTube channel, playlist or TikTok profile and set how many latest posts to transcribe
- No charge for music or silence: a voice-activity gate runs first, so a clip without speech costs nothing and returns no invented text
- Direct media too: an
.mp3or.mp4URL is transcribed like any platform link
All three platforms are verified from a datacenter IP, the same kind of address this Actor runs on. Vimeo is not supported: it serves no media to logged-out clients (0 of 4 public videos). X and Facebook are untested and therefore not claimed.
Music and silence cost nothing
Reels are full of clips with no speech at all: a music-only edit, a montage, a silent product shot. Most transcribers run the model anyway, bill you for it, and hand back a sentence the model invented over the soundtrack.
A voice-activity gate runs first here. No speech means no text and no charge.
Measured against a leading competitor on six real clips, five of them without speech:
| This Actor | Competitor | |
|---|---|---|
| Billed | 1 of 6 | 6 of 6 |
| On a music-only clip | (empty) | шум — Russian for "noise" |
How accurate is the transcript?
It runs on Voxtral Small 24B, not Whisper. Over a 120-clip corpus in 12 languages it scores a word error rate of 6.9 against 9.7 for whisper-large-v3 (lower is better: 6.9 means about seven wrong words in a hundred). With music under the speech, the normal case for a Reel, the gap widens: 9.3 against 43.9 on the six speech-with-music clips, and Whisper dropped one of them entirely.
Per language, on the same corpus (word error rate; character error rate for zh, ja, ko):
| en | nl | de | fr | es | it | pt | hi | ar | zh | ja | ko |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 7.7 | 6.8 | 2.9 | 9.4 | 1.7 | 2.2 | 5.1 | 10.1 | 16.4 | 7.4 | 5.4 | 8.1 |
These are read-speech benchmark clips. A Reel with a beat under the voice, a phone microphone and street noise will score worse than this; the speech-with-music number above is the better guide for that case. Languages outside the twelve are handled by the same model but are not measured here.
On clean speech the transcript is level with the best competitors: two NASA clips came back word-for-word equal. The audio is cut at silences and the pieces are transcribed in parallel, at roughly 5 to 9 times realtime: transcribing a 24-minute video takes about three minutes and a 69-minute podcast about nine, plus the download.
Pass the language if you know it
Speech models decode better when they are told the language, and the difference is not small. On the same 120-clip benchmark:
| word error rate | |
|---|---|
language set to the right code | 6.8 |
language: auto | 7.6 |
auto gets part of the way there by transcribing one chunk first, reading the
language off it, and running the rest with that. Each row reports what
happened in language_source: given, probe, or none.
If you are scraping one creator, one market or one campaign, you already know the language. Passing it is the cheapest accuracy you will get anywhere.
SRT and WebVTT subtitles from any Reel, TikTok or YouTube video
segments gives {start, end, text}; srt and vtt are the same blocks in
the two subtitle formats, ready to upload.
100:00:00,390 --> 00:00:04,270NASA is building a moon base, a place where astronauts will live, work, and conduct200:00:04,270 --> 00:00:08,530science on and around the moon. Now building a moon base won't happen all at once.
Cue boundaries come from the voice-activity gate, so a cue starts and ends where someone is actually speaking. Inside a cue the words are spread by length, because the model gives no per-word timestamps and inventing them would look more precise while being less true. Blocks stay within two lines of 42 characters and six seconds, and break on sentence ends where they can. Japanese and Chinese get shorter lines, as subtitles in those scripts should.
How to get a transcript from an Instagram Reel, TikTok or YouTube video
- Open Audio & Video Transcriber on Apify and click Try for free.
- Paste one or more video links in Video URLs. Reels, TikToks, YouTube videos and Shorts can be mixed in one run.
- Set Language to the ISO code if you know it (
en,nl,es, ...). Leaveautoif you do not. - Tick Add hook + summary or fill Translate to if you want them. They cost nothing extra.
- Click Start. One dataset row per video appears with the transcript, segments, SRT and VTT.
- Export the dataset as JSON, CSV or Excel, or read it through the API, the Python or JavaScript client, or the MCP server (see the API tab).
The same input works as a JSON body through the API:
{"urls": ["https://www.instagram.com/nasa/reel/CvNoLm8Ouaa/","https://www.tiktok.com/@nasa/video/7685115221458930957","https://www.youtube.com/watch?v=aircAruvnKk"],"language": "en","include_hook": true}
Input
| Field | Type | Default | Meaning |
|---|---|---|---|
urls | array | — | One or more video URLs, or channel/profile URLs (required) |
posts_per_profile | int | 0 | Transcribe this many latest posts per channel or profile URL. 0 disables it |
language | string | auto | ISO code, or auto. Pass the code when you know it — see above |
include_hook | bool | false | Also return a one-line hook and a one-sentence summary |
diarize | bool | false | Work out who speaks when; labels every segment and subtitle line |
num_speakers | int | — | Exact number of speakers, when you know it. An interview is 2 |
translate_to | string | — | Translate the transcript into this language |
use_proxy | string | auto | auto routes YouTube through a residential pool (it refuses datacenter IPs) and retries anything else through it if the direct attempt is refused; always / never override |
proxy_region | string | — | ISO-2 country for the residential exit (NL, US, …) |
max_duration_seconds | int | 7200 | Skip longer audio so one upload cannot run up the bill |
Duplicate URLs in one run, including a video that arrives both on its own and through a profile, are transcribed and billed once.
Output
One row per video:
| Field | What |
|---|---|
text | the full transcript |
language | ISO code, read off the transcript |
language_source | given, probe or none — how the language was decided |
segments | {start, end, text} blocks |
srt, vtt | the same blocks as subtitle files |
translation, translation_language | when translate_to was set |
hook, summary | when include_hook was set |
speakers, speaker_turns | who spoke and when, with speaker on every segment — when diarize was set |
duration_seconds, speech_seconds | audio length, and how much of it was speech |
speech_spans | the raw voice-activity windows |
platform, video_id, title, uploader, upload_date, view_count, like_count, webpage_url | source metadata, as far as the platform exposes it |
billed_audio_minutes | what this row cost, in minutes |
from_profile | the channel or profile a video was expanded from |
proxy_used | whether the download went through the residential pool |
error | set when one URL fails — a broken URL never stops the run |
Output example
A public-domain speech recording, shortened:
{"url": "https://upload.wikimedia.org/wikipedia/commons/1/1f/George_W_Bush_Columbia_FINAL.ogg","platform": "Generic","video_id": "George_W_Bush_Columbia_FINAL","title": "George_W_Bush_Columbia_FINAL","text": "My fellow Americans, this day has brought terrible news and great sadness to our country. At 9 o'clock this morning, Mission Control in Houston lost contact with ...","language": "en","language_source": "given","duration_seconds": 198.73,"speech_seconds": 115.46,"segments": [{ "start": 0.99, "end": 6.62, "text": "My fellow Americans, this day has brought terrible news and great sadness to our" },{ "start": 6.62, "end": 8.0, "text": "country. At 9 o'clock this morning," }],"srt": "1\n00:00:00,990 --> 00:00:06,620\nMy fellow Americans, this day has brought terrible news and gre...","vtt": "WEBVTT\n\n00:00:00.990 --> 00:00:06.620\nMy fellow Americans, ...","translation": null,"hook": null,"summary": null,"billed_audio_minutes": 4,"error": null}
Note speech_seconds (115.46) against duration_seconds (198.73): the
recording has long pauses, and the gate measured them rather than guessing.
Transcribe a whole YouTube channel or TikTok profile
Paste https://www.youtube.com/@NASA/videos with posts_per_profile: 3 and the
three latest videos come back as three rows, each carrying from_profile.
Channel URLs, playlists and ytsearch5:... all work.
TikTok profiles work when TikTok feels like it. Instagram profiles cannot be listed at all, so give Reel URLs there — you get a row saying so rather than an empty result.
How much does it cost to transcribe a Reel, TikTok or YouTube video?
$0.005 per audio minute, rounded up per video, and only when speech was found. Everything is included: the download, the residential proxy YouTube needs, the transcription, the subtitles, the translation and the hook. There is no start fee and no per-result fee.
| Video | Billed | Cost |
|---|---|---|
| A 30-second Reel | 1 minute | $0.005 |
| An 82-second TikTok | 2 minutes | $0.01 |
| A 19-minute YouTube video | 19 minutes | $0.095 |
| A 70-minute podcast | 70 minutes | $0.35 |
| 1,000 Reels under a minute each | 1,000 minutes | $5 |
A clip with no speech costs nothing. A URL that fails costs nothing. Apify's free plan includes $5 of usage a month, which covers up to 1,000 audio minutes here without a card. Set a maximum charge on the run if you want a hard cap; the Actor stops at the limit and tells you which URLs it did not reach.
Use cases for a video transcript API
- Content research: pull the scripts of a creator's last fifty Reels and see what they actually say
- Ad and competitor monitoring: turn TikTok and Reel campaigns into searchable text, with the hook already extracted
- Subtitles and repurposing: SRT and VTT for re-uploads, translated in the same call
- LLM and RAG pipelines: clean, timed text from short video, one JSON row per clip
- Social listening and brand safety: what is being said about a brand in video, not just in captions
FAQ
Does it work without an Instagram or TikTok login? Yes. No cookies, no session ID, no account. Public content only.
Does it transcribe YouTube Shorts?
Yes. Shorts, regular videos, playlists, channels and ytsearch queries all work. YouTube refuses datacenter IPs, so those downloads go through a residential pool; that is included in the price.
Which languages are supported?
The model is multilingual; twelve languages are measured above (English, Dutch, German, French, Spanish, Italian, Portuguese, Hindi, Arabic, Chinese, Japanese, Korean). Pass language when you know it.
What happens on a video with no speech?
The row comes back with speech: false, an empty text and billed_audio_minutes: 0. No charge, and no invented sentence.
How long can a video be?
Up to two hours per video by default (max_duration_seconds). Long audio is cut at silences and transcribed in parallel; a 69-minute podcast returned 13,757 words, 1,111 timed segments and a 111 KB SRT.
How is the translation billed? It is not. Translation and the hook are text passes on the same model and are included in the per-minute price.
Are the segment timestamps word-accurate? Cue starts and ends sit on measured speech. Within a cue, words are spread by length, so expect them to be close rather than frame-exact. See the subtitles section above.
What comes back for a private, deleted or mistyped URL? An error row that says so, not an empty transcript, so the two are easy to tell apart. The rest of the run continues and the failed URL is not billed.
Can I call it from Python, JavaScript, curl or an AI agent? Yes. The API tab shows ready-made calls for the Apify API, the Python and JavaScript clients, the CLI and the MCP server. Runs can be scheduled and results delivered through webhooks and integrations like any Apify Actor.
Does it support Vimeo, X or Facebook?
Vimeo no: it serves no media to logged-out clients. X and Facebook are untested and are not claimed. Direct .mp3 and .mp4 URLs work.
Limits
- Instagram profiles cannot be expanded to their latest posts; give Reel URLs.
- TikTok profile expansion depends on TikTok and is not guaranteed.
- One video is billed on its full audio length, rounded up to the minute, once speech is found anywhere in it.
- Korean is the weakest of the measured languages (8.1 character error rate); Arabic the weakest in word error rate (16.4).
Is it legal to transcribe Instagram Reels, TikToks and YouTube videos?
This Actor reads publicly available video and turns speech into text. It does not bypass logins and does not access private content. The audio is deleted after transcription and is not used for training. You are responsible for using the output in line with the platforms' terms, copyright and the privacy law that applies to you, in particular when the transcripts contain personal data.
Support
Something wrong, a platform that stopped working, a field you need? Open an issue on the Issues tab of this Actor and it will be looked at.