Video & Audio Transcriber - Word-Level SRT & VTT avatar

Video & Audio Transcriber - Word-Level SRT & VTT

Pricing

from $20.00 / 1,000 transcribed minutes

Go to Apify Store
Video & Audio Transcriber - Word-Level SRT & VTT

Video & Audio Transcriber - Word-Level SRT & VTT

Transcribe any video or audio URL into text. You get the full text, segment timestamps and word-level timestamps. Download the result as SRT, VTT or TXT. Works with mp4, mov, webm, mp3, wav and m4a. It finds the language on its own and handles many files at once. $0.02 per transcribed minute.

Pricing

from $20.00 / 1,000 transcribed minutes

Rating

5.0

(1)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

0

Monthly active users

4 days ago

Last modified

Share

Video & Audio Transcriber: transcripts with word-level timing, from any media URL

Give it a public link to a video or an audio file and get back the full text, sentence-level segments, and a start and end time for every single word. It also writes ready-made .srt, .vtt and .txt files into the run's storage.

You bring your own OpenAI key, so the transcription itself is billed to you by OpenAI on top of what you pay here.

InputOne public media URL, or a list of them. mp4, mov, webm, mp3, wav, m4a
OutputOne row per file, with the transcript, segments, words and file links
CeilingAbout 5 GB per file at the default memory. Long recordings are the real limit, see below
Account neededYour own OpenAI API key
Price$0.02 per transcribed minute, flat on every plan

🎙️ What Video & Audio Transcriber does

It downloads the file you point it at, pulls the speech out as compressed mono audio, and sends that for transcription. What comes back is the plain text, the segments with their timings, and the word list with a start and end time on each word. That word list is what karaoke-style captions need, and most transcript tools do not hand it over.

Give it mediaUrls instead of mediaUrl and it walks the list, writing one row per file. A file that fails does not stop the ones after it.

The language is detected for you unless you name it. Naming it is usually a little more accurate on short or noisy clips.

📥 What you give it

{
"mediaUrl": "https://example.com/podcast.mp3",
"language": "auto",
"wordTimestamps": true,
"outputFormats": ["srt", "vtt", "txt"],
"openaiApiKey": "sk-..."
}
FieldDefaultWhat it is
mediaUrlnoneA public, direct link to one video or audio file.
mediaUrlsnoneA list, for a batch. One dataset row per URL. You can use this instead of mediaUrl or alongside it.
languageautoISO code of the spoken language, or auto to detect it.
wordTimestampstrueAdds the per-word words array to the row. Turn it off for a much smaller row on long files.
outputFormats["srt", "vtt"] when the field is absent, box starts at srt, vtt, txtWhich files to write into the run's storage.
openaiApiKeynoneYour own key. Marked secret, so it is not stored with the run input.
modelwhisper-1Any transcription model your key can reach.
baseUrlhttps://api.openai.com/v1Point it at any OpenAI-compatible endpoint.

The URL has to be the file itself, not a page that plays it. A link ending in .mp4 or .mp3 works; a video page does not.

📤 What you get back

A real row, with the long arrays and URLs cut short:

{
"ok": true,
"sourceUrl": "https://example.com/podcast.mp3",
"language": "en",
"text": "Welcome back to the show. Today we cover the deep ocean.",
"wordCount": 11,
"segmentCount": 2,
"durationSeconds": 8,
"segments": [
{ "start": 0, "end": 4, "text": "Welcome back to the show." },
{ "start": 4, "end": 8, "text": "Today we cover the deep ocean." }
],
"words": [{ "word": "Welcome", "start": 0, "end": 0.4 }, "..."],
"srtKey": "transcript-1757178142775-0.srt",
"srtUrl": "https://api.apify.com/v2/key-value-stores/.../records/transcript-...",
"vttUrl": "https://api.apify.com/v2/key-value-stores/.../records/transcript-..."
}
FieldWhat it is
wordsOne entry per word with its own start and end, in seconds. Present only when wordTimestamps is on.
segmentsSentence-sized cues. This is what the .srt and .vtt files are built from.
durationSecondsWhere the last speech segment ends, rounded. This is the length that gets billed, not the file's own runtime.
textThe whole transcript as one string.
srtUrl, vttUrl, txtUrlDownload links, present only for the formats you asked for.
sourceUrlWhich input URL this row belongs to. On a batch this is the only way to tell rows apart.

🧾 Reading the output

Three kinds of row can land in your dataset.

RowHow to spot itCharged a minute
A transcriptok: true and a sourceUrlyes
A failed fileok: false and an error stringno
The sample row_demo: trueno

Check ok before you count rows. A failed file still writes a row, so a batch of ten can show ten rows with only six transcripts among them. The error field on those rows says what went wrong in plain words: a download that failed, a file over the size limit, or no speech in the audio.

The default table view hides sourceUrl and error. Switch the dataset to All fields, or export as JSON, when you are working through a batch.

You get the _demo sample row instead of real work when no media URL was given, or no key was.

▶️ How to run it

  1. Open Video & Audio Transcriber and click Try for free.
  2. Paste a direct file link into Media URL, or a list into Media URLs (batch).
  3. Put your key into OpenAI API key (BYO).
  4. Leave Include word timestamps on if you want karaoke captions, then click Start.
  5. Read the rows in the dataset, or open the run's Key-value store tab for the subtitle files.

💰 How much does it cost?

$0.02 per transcribed minute. Flat on every Apify plan, no volume tiers. A 12 minute clip is 12 minutes, rounded up to the next whole minute, with one minute as the smallest charge per file.

The length billed is the speech in the file, measured to where the last segment ends. A file that fails and a sample row are not billed any minutes. Your OpenAI key is charged separately by OpenAI for the transcription itself.

💡 What people use it for

  • Word-timed captions for Shorts and Reels, where each word pops as it is spoken.
  • Turning a podcast back catalogue into searchable text, one run per batch of episodes.
  • Pulling quotes out of recorded interviews with a timestamp you can jump to.
  • Getting a .vtt track ready to upload alongside a video for accessibility.

🚧 What it does not do

  • It does not pay for the model. Transcription runs on your own key and shows up on your OpenAI bill.
  • It does not split long recordings. The audio goes up as one upload, and transcription endpoints commonly stop at 25 MB. At the bitrate used here that lands somewhere around 50 minutes of speech, so a full hour is likely to be refused. Cut long files before sending them.
  • No speaker labels. You get the words and the timings, not who said them.
  • It does not resolve a video page. Give it the file URL, not the page the player sits on.
  • The run finishes even when every file failed. Failures are rows, not a failed run, so always read ok.
  • A run stops at one hour and a single download stops after 15 minutes, whichever comes first.
  • Rows get large. An hour of speech with word timestamps is a heavy JSON row. Turn wordTimestamps off when you only need the text.
  • Accuracy is the model's. Accents, crosstalk and background noise affect it, and naming the language usually helps more than changing anything else here.

🧭 Which audio tool do you need?

If you wantUse
A transcript with word-level timingsThis one
That transcript translated into other languagesSubtitle Translator
Captions burned into the pictureAuto Caption Burner
A video re-voiced in another languageAI Video Dubber
Text turned into a spoken audio fileAI Text-to-Speech Voiceover

❓ Questions people ask

What counts as a minute? The speech in the file, measured to the end of the last segment and rounded up. Silence at the end of a recording does not add to it.

Can I transcribe several files at once? Yes. Put them in Media URLs (batch) and you get one row per file.

Why is durationSeconds shorter than my file? Because it measures speech, not runtime. A clip with a long musical outro ends its last segment well before the file does.

Can I use a different provider? Yes. Set baseUrl to any OpenAI-compatible endpoint and model to whatever it serves.

Why did one file in my batch come back empty? Read its error field. Most often the link was not a direct file link, or the audio had no speech in it.

Is it legal to transcribe this? Transcribing media you own or have the right to use is normally fine. Recordings of people carry personal data, which GDPR and similar laws cover, so have a reason for holding it. Apify's write-up on scraping and the law is a reasonable starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID and the media URL you used. The error field on the failing row usually names the problem on its own.