Whisper Audio & Video URL Transcriber — Speech to Text avatar

Whisper Audio & Video URL Transcriber — Speech to Text

Pricing

from $14.00 / 1,000 audio minute (small model)s

Go to Apify Store
Whisper Audio & Video URL Transcriber — Speech to Text

Whisper Audio & Video URL Transcriber — Speech to Text

Speech-to-text transcription of audio and video files from direct URLs (mp3, mp4, m4a, wav, webm, ogg, flac) into text, timed segments, SRT and VTT. Open-source Whisper, 99 languages. From $0.36 per audio hour.

Pricing

from $14.00 / 1,000 audio minute (small model)s

Rating

0.0

(0)

Developer

drop-in apis

drop-in apis

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 hours ago

Last modified

Categories

Share

Transcribe audio and video files from direct URLs into text, timed segments, SRT and VTT subtitles: mp3, mp4, m4a, wav, webm, ogg, flac, mkv and more, in 99 languages with automatic language detection. It runs open-source Whisper (faster-whisper) on Apify, and you pay per minute of audio.

At a glance: $0.014 per minute of audio with the default small model ($0.84 per hour), or $0.006 per minute with base · failed files are not charged · measured: a 10.3-minute speech took about 7 minutes with small on the default 2 GB memory (under 4 minutes with 4 GB). Try it: click Start with the prefilled 24-second public sample file.

Last updated: 2026-10-02

  • ✅ One dataset row per file: text, segments (start, end, text), srt, vtt, detected language, durationSec, billedMinutes
  • ✅ Audio and video files: the audio track is extracted automatically
  • ✅ 99 languages, auto-detected per file, or set language (for example en, ar, es)
  • ✅ Two models: small (default, more accurate, $0.84 per audio hour) and base (2–3× faster, $0.36 per audio hour)
  • ✅ Silence is skipped (voice activity detection), so long recordings with pauses process faster
  • ✅ maxMinutesPerFile caps cost: long files are cut at that point and you pay only for what was transcribed
  • 🔒 Files are downloaded to a temporary file, transcribed, and deleted. Nothing is sent to a third-party API.
  • ❌ No YouTube, TikTok, Instagram or other platform page URLs. The Actor does not download from video or music platforms. Pass a direct link to a media file you are allowed to use (your own storage, S3, a CDN, a podcast enclosure URL).

Input

{
"urls": ["https://upload.wikimedia.org/wikipedia/commons/d/dd/Armstrong_Small_Step.ogg"],
"language": "",
"model": "small",
"outputFormats": ["text", "segments", "srt", "vtt"],
"maxMinutesPerFile": 60
}
FieldDefaultMeaning
urlsrequiredDirect media file URLs, up to 500 per run, up to 1 GB each
languageautoISO-639-1 code. Setting it helps on short or noisy files
modelsmallsmall ($0.014/min) or base ($0.006/min, 2–3× faster, less accurate)
outputFormatsall fourAny of text, segments, srt, vtt
maxMinutesPerFile601–180. Transcribe at most this many minutes from the start of each file

Output (real run on Apify, 2026-10-02, default small model)

{
"url": "https://upload.wikimedia.org/wikipedia/commons/d/dd/Armstrong_Small_Step.ogg",
"language": "en",
"durationSec": 24.11,
"transcribedSec": 24.11,
"truncated": false,
"model": "small",
"billedMinutes": 1,
"text": "I'm going to step off the land now. That's one small step for man. One giant leap for mankind.",
"segments": [
{ "start": 3.38, "end": 15.18, "text": "I'm going to step off the land now." },
{ "start": 15.18, "end": 20.81, "text": "That's one small step for man." },
{ "start": 20.81, "end": 23.81, "text": "One giant leap for mankind." }
],
"srt": "1\n00:00:03,380 --> 00:00:15,180\nI'm going to step off the land now.\n\n...",
"vtt": "WEBVTT\n\n00:00:03.380 --> 00:00:15.180\nI'm going to step off the land now.\n\n...",
"error": null
}

A file that cannot be processed (404, a web page instead of a media file, no audio track, an unsupported format) produces a row with error set and billedMinutes: 0. The run continues with the next file.

Run it from code

Python (apify-client)

import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("dropin-apis/media-url-transcriber").call(run_input={
"urls": ["https://example.com/podcast/episode-12.mp3"],
"outputFormats": ["text", "srt"],
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
print(row["billedMinutes"], row.get("text", row.get("error"))[:200])

Node.js (apify-client)

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('dropin-apis/media-url-transcriber').call({ urls: ['https://example.com/talk.mp4'], language: 'en' });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].vtt);

AI agents (Apify MCP server)

Agents connected to the Apify MCP server can find this Actor with search-actors ("transcribe audio") and run it with call-actor. The input is the JSON above.

Limits

  • Up to 500 URLs per run and 1 GB per file; maxMinutesPerFile is 1–180 (default 60) and only the first N minutes of each file are transcribed.
  • Direct URLs to public media files only. A web page, a 404, a file with no audio track or an unsupported format produces a row with error set and no charge; use media you are entitled to process.
  • Timestamps are per segment (sentence or phrase), not per word, and accuracy is that of the open Whisper small/base models.

Pricing

Pay per started minute of transcribed audio, plus Apify's standard run-start charge ($0.00005 per GB of run memory). There is no subscription.

ModelEventPer minutePer hour
small (default)audio-minute$0.014$0.84
baseaudio-minute-base$0.006$0.36
Audiosmallbase
30-second voice note$0.014$0.006
10-minute speech$0.14$0.06
1-hour podcast$0.84$0.36

Failed files are not charged. If you set a maximum total charge on the run, the Actor checks it before transcribing each file and stops cleanly when the next file would exceed it.

Speed and accuracy

  • Measured on Apify with the default 2 GB memory (half a CPU core): a 10.3-minute speech took 7 minutes with small. With 4 GB (1 core) the same file took under 4 minutes, and base took 1.5 minutes. More memory gives more CPU cores and a faster run; the per-minute price does not change.
  • Accuracy is that of the open Whisper small model: good on clear speech, weaker than large commercial models on noisy audio, heavy accents and rare languages. base is noticeably less accurate.
  • Timestamps are per segment (a sentence or phrase), not per word.

FAQ

How do I transcribe an MP4 video to text?

Pass the direct URL of the .mp4 file in urls. The Actor extracts the audio track and returns the transcript, segments, SRT and VTT.

Can I make SRT or VTT subtitles from an audio file?

Yes. Include srt and/or vtt in outputFormats (both are included by default). Each row then has ready-to-save subtitle text.

No. The Actor does not download from video platforms. Use a direct link to a media file you have the right to process.

Which languages are supported?

All 99 Whisper languages, including English, Arabic, Spanish, French, German, Hindi, Chinese and Japanese. The language is detected per file unless you set language.

How long can a file be?

Up to 1 GB, and up to 180 minutes per file are transcribed (maxMinutesPerFile, default 60). Longer files are cut at the limit, marked truncated: true, and billed only for the transcribed minutes.

What does it cost?

$0.014 per started minute with the default small model ($0.84 per hour), or $0.006 per minute with base ($0.36 per hour). Errors are not charged.

Is my audio stored?

No. Each file is downloaded to a temporary file inside the run, transcribed, and deleted. Only the transcript is saved to your dataset.

I need an OpenAI whisper-1 compatible API instead

Use the sister Actor whisper-1 alternative: the same engine behind OpenAI's /v1/audio/transcriptions request and response format, for apps built on the OpenAI SDK.