Video & Audio Transcriber - Word-Level SRT & VTT
Pricing
from $20.00 / 1,000 transcribed minutes
Video & Audio Transcriber - Word-Level SRT & VTT
Transcribe any video or audio URL into text. You get the full text, segment timestamps and word-level timestamps. Download the result as SRT, VTT or TXT. Works with mp4, mov, webm, mp3, wav and m4a. It finds the language on its own and handles many files at once. $0.02 per transcribed minute.
Pricing
from $20.00 / 1,000 transcribed minutes
Rating
5.0
(1)
Developer
Dami's Studio
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
0
Monthly active users
4 days ago
Last modified
Categories
Share
Video & Audio Transcriber: transcripts with word-level timing, from any media URL
Give it a public link to a video or an audio file and get back the full text, sentence-level
segments, and a start and end time for every single word. It also writes ready-made .srt, .vtt
and .txt files into the run's storage.
You bring your own OpenAI key, so the transcription itself is billed to you by OpenAI on top of what you pay here.
| Input | One public media URL, or a list of them. mp4, mov, webm, mp3, wav, m4a |
| Output | One row per file, with the transcript, segments, words and file links |
| Ceiling | About 5 GB per file at the default memory. Long recordings are the real limit, see below |
| Account needed | Your own OpenAI API key |
| Price | $0.02 per transcribed minute, flat on every plan |
🎙️ What Video & Audio Transcriber does
It downloads the file you point it at, pulls the speech out as compressed mono audio, and sends that for transcription. What comes back is the plain text, the segments with their timings, and the word list with a start and end time on each word. That word list is what karaoke-style captions need, and most transcript tools do not hand it over.
Give it mediaUrls instead of mediaUrl and it walks the list, writing one row per file. A file
that fails does not stop the ones after it.
The language is detected for you unless you name it. Naming it is usually a little more accurate on short or noisy clips.
📥 What you give it
{"mediaUrl": "https://example.com/podcast.mp3","language": "auto","wordTimestamps": true,"outputFormats": ["srt", "vtt", "txt"],"openaiApiKey": "sk-..."}
| Field | Default | What it is |
|---|---|---|
mediaUrl | none | A public, direct link to one video or audio file. |
mediaUrls | none | A list, for a batch. One dataset row per URL. You can use this instead of mediaUrl or alongside it. |
language | auto | ISO code of the spoken language, or auto to detect it. |
wordTimestamps | true | Adds the per-word words array to the row. Turn it off for a much smaller row on long files. |
outputFormats | ["srt", "vtt"] when the field is absent, box starts at srt, vtt, txt | Which files to write into the run's storage. |
openaiApiKey | none | Your own key. Marked secret, so it is not stored with the run input. |
model | whisper-1 | Any transcription model your key can reach. |
baseUrl | https://api.openai.com/v1 | Point it at any OpenAI-compatible endpoint. |
The URL has to be the file itself, not a page that plays it. A link ending in .mp4 or .mp3
works; a video page does not.
📤 What you get back
A real row, with the long arrays and URLs cut short:
{"ok": true,"sourceUrl": "https://example.com/podcast.mp3","language": "en","text": "Welcome back to the show. Today we cover the deep ocean.","wordCount": 11,"segmentCount": 2,"durationSeconds": 8,"segments": [{ "start": 0, "end": 4, "text": "Welcome back to the show." },{ "start": 4, "end": 8, "text": "Today we cover the deep ocean." }],"words": [{ "word": "Welcome", "start": 0, "end": 0.4 }, "..."],"srtKey": "transcript-1757178142775-0.srt","srtUrl": "https://api.apify.com/v2/key-value-stores/.../records/transcript-...","vttUrl": "https://api.apify.com/v2/key-value-stores/.../records/transcript-..."}
| Field | What it is |
|---|---|
words | One entry per word with its own start and end, in seconds. Present only when wordTimestamps is on. |
segments | Sentence-sized cues. This is what the .srt and .vtt files are built from. |
durationSeconds | Where the last speech segment ends, rounded. This is the length that gets billed, not the file's own runtime. |
text | The whole transcript as one string. |
srtUrl, vttUrl, txtUrl | Download links, present only for the formats you asked for. |
sourceUrl | Which input URL this row belongs to. On a batch this is the only way to tell rows apart. |
🧾 Reading the output
Three kinds of row can land in your dataset.
| Row | How to spot it | Charged a minute |
|---|---|---|
| A transcript | ok: true and a sourceUrl | yes |
| A failed file | ok: false and an error string | no |
| The sample row | _demo: true | no |
Check ok before you count rows. A failed file still writes a row, so a batch of ten can show
ten rows with only six transcripts among them. The error field on those rows says what went wrong
in plain words: a download that failed, a file over the size limit, or no speech in the audio.
The default table view hides sourceUrl and error. Switch the dataset to All fields, or
export as JSON, when you are working through a batch.
You get the _demo sample row instead of real work when no media URL was given, or no key was.
▶️ How to run it
- Open Video & Audio Transcriber and click Try for free.
- Paste a direct file link into Media URL, or a list into Media URLs (batch).
- Put your key into OpenAI API key (BYO).
- Leave Include word timestamps on if you want karaoke captions, then click Start.
- Read the rows in the dataset, or open the run's Key-value store tab for the subtitle files.
💰 How much does it cost?
$0.02 per transcribed minute. Flat on every Apify plan, no volume tiers. A 12 minute clip is 12 minutes, rounded up to the next whole minute, with one minute as the smallest charge per file.
The length billed is the speech in the file, measured to where the last segment ends. A file that fails and a sample row are not billed any minutes. Your OpenAI key is charged separately by OpenAI for the transcription itself.
💡 What people use it for
- Word-timed captions for Shorts and Reels, where each word pops as it is spoken.
- Turning a podcast back catalogue into searchable text, one run per batch of episodes.
- Pulling quotes out of recorded interviews with a timestamp you can jump to.
- Getting a
.vtttrack ready to upload alongside a video for accessibility.
🚧 What it does not do
- It does not pay for the model. Transcription runs on your own key and shows up on your OpenAI bill.
- It does not split long recordings. The audio goes up as one upload, and transcription endpoints commonly stop at 25 MB. At the bitrate used here that lands somewhere around 50 minutes of speech, so a full hour is likely to be refused. Cut long files before sending them.
- No speaker labels. You get the words and the timings, not who said them.
- It does not resolve a video page. Give it the file URL, not the page the player sits on.
- The run finishes even when every file failed. Failures are rows, not a failed run, so always
read
ok. - A run stops at one hour and a single download stops after 15 minutes, whichever comes first.
- Rows get large. An hour of speech with word timestamps is a heavy JSON row. Turn
wordTimestampsoff when you only need the text. - Accuracy is the model's. Accents, crosstalk and background noise affect it, and naming the language usually helps more than changing anything else here.
🧭 Which audio tool do you need?
| If you want | Use |
|---|---|
| A transcript with word-level timings | This one |
| That transcript translated into other languages | Subtitle Translator |
| Captions burned into the picture | Auto Caption Burner |
| A video re-voiced in another language | AI Video Dubber |
| Text turned into a spoken audio file | AI Text-to-Speech Voiceover |
❓ Questions people ask
What counts as a minute? The speech in the file, measured to the end of the last segment and rounded up. Silence at the end of a recording does not add to it.
Can I transcribe several files at once? Yes. Put them in Media URLs (batch) and you get one row per file.
Why is durationSeconds shorter than my file? Because it measures speech, not runtime. A clip
with a long musical outro ends its last segment well before the file does.
Can I use a different provider? Yes. Set baseUrl to any OpenAI-compatible endpoint and model
to whatever it serves.
Why did one file in my batch come back empty? Read its error field. Most often the link was
not a direct file link, or the audio had no speech in it.
Is it legal to transcribe this? Transcribing media you own or have the right to use is normally fine. Recordings of people carry personal data, which GDPR and similar laws cover, so have a reason for holding it. Apify's write-up on scraping and the law is a reasonable starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the media URL you used. The error
field on the failing row usually names the problem on its own.