Audio & Video to Text - Whisper Transcription API avatar

Audio & Video to Text - Whisper Transcription API

Pricing

from $3.20 / 1,000 audio minute transcribeds

Go to Apify Store
Audio & Video to Text - Whisper Transcription API

Audio & Video to Text - Whisper Transcription API

Transcribe audio and video files from direct URLs (mp3, wav, m4a, ogg, mp4, mov, webm) with Whisper large-v3-turbo. Get text, SRT, VTT, timestamped segments and RAG chunks. $0.004 per audio minute, platform usage included.

Pricing

from $3.20 / 1,000 audio minute transcribeds

Rating

0.0

(0)

Developer

Rosario Vitale

Rosario Vitale

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

What does Audio & Video to Text do?

Audio & Video to Text turns recordings into text. Give it direct links to audio or video files, and it returns one row per file with the transcript, timestamped segments, SRT and WebVTT subtitles and RAG-ready chunks. It runs the Whisper large-v3-turbo speech recognition model, detects the spoken language automatically and can also translate speech to English.

No API key and no model hosting are needed. The Actor downloads each file, extracts the audio track with ffmpeg, converts it to compact 16 kHz mono mp3, cuts long recordings at natural pauses into parts of about 10 minutes, transcribes the parts in parallel and merges everything back into one transcript with correct timestamps. Supported formats include mp3, wav, m4a, ogg, opus, flac, webm, mp4, mov and mkv, up to 2 GB and 180 minutes per file.

Who is it for?

  • Podcasters and media teams who need transcripts, show notes and subtitle files for episodes and interviews.
  • Developers and data teams who feed call recordings, meetings or lectures into search, analytics or LLM pipelines.
  • AI and RAG builders who want timestamped chunks ready for embeddings and a vector database.
  • Researchers, journalists and educators who transcribe interviews, archives and recorded courses in bulk.
  • Video editors and creators who need SRT or WebVTT subtitle files from their own media files.

Output fields

FieldDescription
urlThe media link that was processed.
statusCOMPLETE, PARTIAL, VALID_EMPTY, INVALID_INPUT, UPSTREAM_FAILED or TOO_LONG (see below).
durationSecondsLength of the audio in seconds.
billedMinutesStarted audio minutes that were charged for this file.
languageDetected (or selected) spoken language code, for example en.
textFull transcript as plain text.
segmentsList of {start, end, text} with times in seconds from the start of the file.
srt, vttSubtitles as text, when the format is selected.
srtUrl, vttUrlDirect links to the .srt and .vtt files saved in the key-value store of the run.
chunksRAG chunks: {index, text, startSeconds, endSeconds, charCount} limited to the chosen chunk size.
wordCount, task, model, partsTotal, partsTranscribed, error, processedAtDetails and diagnostics.

Only the formats you select in Output formats are filled, the other fields stay null.

Row status values:

  • COMPLETE: the whole file was transcribed. Charged per started audio minute.
  • PARTIAL: some parts of the file were transcribed, others failed or the run's maximum charge was reached. Charged only for the minutes that were transcribed; error lists what is missing.
  • VALID_EMPTY: the file was processed but no speech was found. Free.
  • INVALID_INPUT: not a public direct media link, a web page, a file without audio, a YouTube or social-media link, or a private network address. Free.
  • TOO_LONG: the file is longer than Maximum duration per file (hard cap 180 minutes) or larger than 2 GB. Free.
  • UPSTREAM_FAILED: the file could not be downloaded (for example HTTP 404) or the transcription service was unavailable. Free.

Pricing at a glance

Price per 1,000 audio minutes
This Actor$4.00
Median of 12 audio and video transcription Actors in Apify Store with a per-minute price (October 2026)$17.50
Cheapest comparable Actor found (October 2026)$3.00

That is about $0.24 per hour of audio, roughly 75% below the median per-minute price for this task in Apify Store.

Cost example: a 45-minute podcast episode is 45 × $0.004 = $0.18. A 90-second clip counts as 2 started minutes ($0.008). Files with no speech, invalid or too long files and failed downloads cost nothing, and you pay only for the parts that were actually transcribed.

Subscriber discounts: on a paid Apify plan you pay less per minute: Bronze −10%, Silver −15%, Gold and higher −20% ($3.20 per 1,000 minutes on Gold).

Apify platform usage (compute, storage, data transfer) is included. You pay only the audio-minute price plus a small run start fee of $0.0001 per GB of run memory.

Use the run's maximum charge setting to cap spending. The Actor stops cleanly when the limit is reached and returns what was transcribed so far.

How to use

  1. Open the Actor and paste direct links to your audio or video files into Audio or video file URLs, one per line. The prefilled example is a short public-domain speech recording.
  2. Optionally choose the Spoken language, switch Task to translate to English, and select the Output formats you need.
  3. Click Start. A short clip finishes in a few seconds; a one-hour file typically takes a few minutes.
  4. Open the Output tab to read the transcripts, or export as JSON, CSV, Excel or HTML. Download the .srt and .vtt files from the links in each row.

Good to know

  • Direct media links only. The link must return the audio or video file itself (for example https://example.com/episode.mp3). Web pages, YouTube, TikTok, Instagram, Facebook, X, Vimeo and similar links are rejected as INVALID_INPUT. For YouTube videos use the YouTube Transcript Actor.
  • Public files only. Private, local and reserved network addresses are blocked, and links that need a login or cookies are not supported.
  • Shared daily capacity. Transcription runs on a hosted model with a daily capacity limit. If it is used up during your run, the remaining files come back as UPSTREAM_FAILED and are not charged; run them again later.
  • Privacy. Audio is processed in temporary storage, deleted after each file, and sent in short parts to the hosted Whisper model for transcription.
  • Accuracy. Quality depends on the recording. Choosing the spoken language, adding a context prompt with names and terms, and enabling Skip silence helps with noisy or music-heavy audio.

Input example

{
"mediaUrls": ["https://upload.wikimedia.org/wikipedia/commons/4/46/1941_Roosevelt_speech_pearlharbor_p1.ogg"],
"language": "auto",
"task": "transcribe",
"outputFormats": ["text", "srt"],
"maxDurationMinutes": 60
}

Output example

{
"recordType": "transcript",
"url": "https://upload.wikimedia.org/wikipedia/commons/4/46/1941_Roosevelt_speech_pearlharbor_p1.ogg",
"status": "COMPLETE",
"durationSeconds": 26.1,
"billedMinutes": 1,
"language": "en",
"task": "transcribe",
"model": "whisper-large-v3-turbo",
"text": "Yesterday, December 7th, 1941, a state which will live in infamy. The States of America was suddenly and deliberately attacked by naval and air forces of the Empire of Japan.",
"segments": [
{ "start": 0.24, "end": 6.56, "text": "Yesterday, December 7th, 1941," },
{ "start": 7.38, "end": 12.81, "text": "a state which will live in infamy." }
],
"srt": "1\n00:00:00,240 --> 00:00:06,560\nYesterday, December 7th, 1941,\n\n2\n00:00:07,380 --> 00:00:12,810\na state which will live in infamy.\n",
"srtUrl": "https://api.apify.com/v2/key-value-stores/STORE_ID/records/001-1941_Roosevelt_speech_pearlharbor_p1.srt",
"wordCount": 30,
"partsTotal": 1,
"partsTranscribed": 1,
"error": null
}

The transcript above is the model's literal output for the sample clip.

API

One HTTP call runs the Actor and returns the results directly:

curl -X POST "https://api.apify.com/v2/acts/zenomastro~audio-video-to-text/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"mediaUrls": ["https://example.com/interview.mp3"], "outputFormats": ["text", "segments"]}'

For long files start the run asynchronously and fetch the dataset when it finishes. The same works from the Apify Python/JavaScript clients, Make, n8n, Zapier, LangChain and LlamaIndex integrations, or on a schedule.

Use with AI agents (MCP)

Claude, ChatGPT, Cursor, VS Code, n8n AI agents and any MCP client can run this Actor through the official Apify MCP server. Add https://mcp.apify.com?tools=zenomastro/audio-video-to-text to your client, then ask for example:

"Transcribe this podcast episode https://example.com/episode.mp3 and give me SRT subtitles and a 5-bullet summary."

The agent fills the input from the field descriptions and receives clean JSON with the transcript, timestamps and subtitle links it can reason over.

Why use this Actor?

Transcribe audio and video files from direct URLs (mp3, wav, m4a, ogg, mp4, mov, webm) with Whisper large-v3-turbo. Get text, SRT, VTT, timestamped segments and RAG chunks. $0.004 per audio minute, platform usage included.

Features

  • Audio or video file URLs — Direct links to audio or video files, one per line (mp3, wav, m4a, ogg, opus, flac, webm, mp4, mov, mkv). Links must be public and point straight to the file. YouTube and social-media page links are not supported; use the YouTube Transcript Actor for those.
  • Spoken language — Language spoken in the audio. Auto-detect works well for most recordings; choosing the language explicitly can improve accuracy and avoids wrong detection on short or noisy clips.
  • Task — Transcribe keeps the original language. Translate to English returns the transcript translated into English.
  • Output formats — What to include in each result row. SRT and VTT subtitles are also saved as files in the key-value store of the run, with direct links. RAG chunks split the transcript into timestamped pieces for embeddings.
  • RAG chunk size (characters) — Maximum characters per chunk when the RAG chunks format is selected. Chunks follow segment boundaries and keep start and end times.
  • Maximum duration per file (minutes) — Files longer than this are reported as TOO_LONG and are not charged. The hard cap is 180 minutes per file.
  • Parallel transcriptions — How many files are processed, and how many audio parts are transcribed, at the same time. Lower values are gentler on the transcription service.
  • Skip silence (voice activity detection) — Remove silent and non-speech sections before transcribing. Helps with long pauses, music or background noise, where the model can otherwise invent text.
  • Context prompt — Optional text that helps the model with names, product terms or spelling, for example a list of speaker names or technical vocabulary. Maximum 800 characters.

Use cases

  • Podcast and interview transcription.
  • Meeting and call recording transcripts.
  • Subtitle files for videos and courses.
  • Embedding and vector search preparation from audio.

Example input

{
"mediaUrls": [
"https://upload.wikimedia.org/wikipedia/commons/4/46/1941_Roosevelt_speech_pearlharbor_p1.ogg"
],
"language": "auto",
"task": "transcribe",
"outputFormats": [
"text",
"srt"
],
"chunkSizeChars": 1000,
"maxDurationMinutes": 60
}

Pricing & cost control

The primary event costs $0.004000 per audio minute transcribed (about $4.00 per 1,000 successful primary events). Only successful primary events are intentionally billed by this Actor; summary/status rows add context without adding primary-event charges.

Use the bounded input limits and filters to keep both event charges and platform usage predictable.

FAQ

What is this Actor for?
It is designed for podcast and interview transcription, meeting and call recording transcripts, subtitle files for videos and courses.

Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.