Audio & Video to Text — AI Transcript API (Whisper) avatar

Audio & Video to Text — AI Transcript API (Whisper)

Pricing

from $6.00 / 1,000 audio minute (fast)s

Go to Apify Store
Audio & Video to Text — AI Transcript API (Whisper)

Audio & Video to Text — AI Transcript API (Whisper)

Transcribe audio & video to text in bulk with OpenAI Whisper AI. Paste file links, Google Drive/Dropbox or podcast RSS feeds → transcripts with timestamps + SRT/VTT subtitles. 99 languages, auto-detect, translate to English. From $0.006/min, no API key.

Pricing

from $6.00 / 1,000 audio minute (fast)s

Rating

0.0

(0)

Developer

Leandro Zanatta

Leandro Zanatta

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

12 hours ago

Last modified

Categories

Share

Turn any audio or video file into text with OpenAI Whisper, in bulk, without an API key or a GPU. Paste links to MP3, MP4, WAV, M4A and other media files, share links from Google Drive or Dropbox, or whole podcast RSS feeds. You get clean transcripts with timestamps and ready-to-use SRT and VTT subtitles.

  • 🌍 99 languages, with automatic language detection
  • 🇬🇧 Translate to English from any supported language
  • ⏱️ Segment and word-level timestamps
  • 🎬 SRT / VTT subtitle files for every transcript
  • 📡 Podcast transcription straight from RSS feeds
  • 📦 Bulk transcription: hundreds of files in one run
  • 💸 From $0.006 per minute, and you only pay for audio that was transcribed successfully

What does this speech-to-text Actor do?

It downloads each media file, runs Whisper speech recognition on it and saves the result to a dataset: full transcript text, detected language, duration, timestamped segments and links to SRT/VTT subtitle files. Export as JSON, CSV, Excel or HTML, or fetch it through the API.

Typical uses:

  • Podcast transcripts: show notes, blog posts, SEO articles and searchable archives
  • Video subtitles and captions for courses, webinars, YouTube uploads and social clips
  • Meeting, interview and lecture transcription
  • AI and LLM pipelines: feed spoken content into RAG, summarisation, chatbots and AI agents
  • Content repurposing: turn audio and video into text, quotes and social posts

How to transcribe audio or video to text

  1. Paste one or more file URLs into Audio / video URLs, or podcast feeds into Podcast RSS feeds.
  2. Choose a Quality (see below). Leave Language empty to auto-detect it.
  3. Click Start. When the run finishes, open the Output tab to see the transcripts and download the subtitles.

Quality levels and pricing

You pay per started minute of transcribed audio. Failed files are free.

QualityWhisper modelPrice per minute1-hour podcastBest for
Fastbase$0.006$0.36Clear speech, large volumes
Balanced (default)small$0.015$0.90Podcasts, interviews, meetings
Accuratelarge-v3-turbo$0.07$4.20Accents, noisy audio, names and jargon, non-English

There is also a small start fee per run ($0.00125 per GB of memory, about $0.005 with the default 4 GB). Set Max duration per file or a maximum cost per run to stay within budget: the Actor stops transcribing when the limit is reached.

Supported inputs

  • Audio: MP3, M4A, AAC, WAV, FLAC, OGG, OPUS, WMA
  • Video: MP4, MOV, WEBM, MKV, AVI (the audio track is transcribed)
  • Links: any direct, publicly downloadable URL, plus public Google Drive and Dropbox share links
  • Podcasts: any standard podcast RSS feed; the newest N episodes are transcribed

Input example

{
"mediaUrls": ["https://github.com/openai/whisper/raw/main/tests/jfk.flac"],
"rssFeeds": ["https://feeds.npr.org/510289/podcast.xml"],
"maxEpisodesPerFeed": 2,
"quality": "balanced",
"language": "",
"translateToEnglish": false,
"wordTimestamps": false
}

Output example

{
"sourceUrl": "https://github.com/openai/whisper/raw/main/tests/jfk.flac",
"language": "en",
"languageProbability": 0.98,
"durationSeconds": 11.0,
"transcribedSeconds": 11.0,
"truncated": false,
"text": "And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country.",
"srtUrl": "https://api.apify.com/v2/key-value-stores/.../records/transcript-0001.srt",
"vttUrl": "https://api.apify.com/v2/key-value-stores/.../records/transcript-0001.vtt",
"segments": [
{
"start": 0.0,
"end": 11.0,
"text": "And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country.",
"words": [{ "start": 0.0, "end": 0.52, "word": "And" }, { "start": 0.52, "end": 0.82, "word": "so" }]
}
]
}

words is only included when Word-level timestamps is on.

Use it as a transcription API

Call the Actor from your code with the Apify API or client libraries:

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("adorable_partial/whisper-audio-video-transcriber").call(
run_input={"mediaUrls": ["https://example.com/interview.mp3"], "quality": "balanced"}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item["text"])
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('adorable_partial/whisper-audio-video-transcriber').call({
mediaUrls: ['https://example.com/interview.mp3'],
quality: 'balanced',
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].text);

It also works with Make, Zapier, n8n, webhooks and scheduled runs, and as a tool for AI agents (Claude, ChatGPT, Cursor…) through the Apify MCP server.

FAQ

Which languages are supported? All 99 Whisper languages, including English, Spanish, Portuguese, French, German, Italian, Hindi, Japanese, Chinese, Arabic and Russian. The language is detected automatically, or you can set it.

Do I need an OpenAI API key? No. Whisper runs inside the Actor, and you only pay the per-minute price.

Can it transcribe YouTube, TikTok or Instagram links? No. It needs a direct link to a media file (or a podcast RSS feed). Only process content you have the right to use.

How long can a file be? Up to 4 GB per file. Use Max duration per file to transcribe only the first N minutes.

How accurate is it? Accurate mode (large-v3-turbo) is close to the best open speech recognition models available. Fast and Balanced trade some accuracy for a lower price and are usually enough for clear speech.

How fast is it? Fast mode handles about 4 minutes of audio per minute of runtime at the default 4 GB of memory. For Accurate mode on long files, give the run 8–16 GB of memory (more CPU cores).

Is speaker diarization included? Not yet. If you need speaker labels, open an issue.

Tips

  • Set Language when you know it: detection is skipped and short clips come out more accurate.
  • Use Translate to English to get English subtitles for foreign-language video.
  • Found a bug or need a feature? Open an issue in the Issues tab.