Audio & Video to Text — Whisper Large v3, SRT/VTT, $0.008/min avatar

Audio & Video to Text — Whisper Large v3, SRT/VTT, $0.008/min

Pricing

$8.00 / 1,000 audio minute transcribeds

Go to Apify Store
Audio & Video to Text — Whisper Large v3, SRT/VTT, $0.008/min

Audio & Video to Text — Whisper Large v3, SRT/VTT, $0.008/min

Transcribe audio and video files (MP3, MP4, WAV, M4A, podcasts, meetings, interviews) to text with Whisper large-v3 in 99 languages. Timestamps, SRT and VTT subtitles, translation to English. Hours of audio in minutes. Pay per audio minute.

Pricing

$8.00 / 1,000 audio minute transcribeds

Rating

0.0

(0)

Developer

Robin

Robin

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Audio & Video to Text — Whisper Large v3 Transcriber

Transcribe audio and video files to text with OpenAI's Whisper large-v3 model: podcasts, meetings, interviews, lectures, voice notes, YouTube-style videos you own, customer calls. 99 languages, word-level timestamps, ready-to-use SRT and VTT subtitles, AI subtitle translation into 40+ languages, and translation to English — for $0.008 per audio minute (≈ $0.48 per hour).

  • Fast: one hour of audio is usually done in 1–3 minutes.
  • Accurate: Whisper large-v3 (turbo), the strongest open speech model.
  • Cheap and fair: billed per started minute of audio, only for files that succeed.
  • Any format: MP3, MP4, M4A, WAV, OGG, OPUS, WebM, MOV, MKV, FLAC, AAC… (ffmpeg under the hood). Video is fine — the audio track is extracted.
  • Long files: up to 2 GB per file. Long recordings are split at silences so no words are cut.
  • Bulk & podcasts: paste many URLs, or a podcast RSS feed to transcribe its newest episodes.
  • Real subtitles, not raw segments: cues are built from word timings — max 2 lines × 42 characters and ~6 s on screen, split at sentence and clause boundaries. One-line short captions for TikTok / Reels / Shorts with one setting.
  • Subtitles in other languages: add es, fr, de, ja… and get translated SRT/VTT timed to the original speech. Translated sentence by sentence by a large language model, so word order differences (German, Japanese…) stay correct.
  • Share links work: public Google Drive and Dropbox file links are converted automatically.

Output

Each file gives one result row:

{
"sourceUrl": "https://example.com/episode-42.mp3",
"status": "ok",
"language": "english",
"durationSeconds": 1867.3,
"billedMinutes": 32,
"text": "Welcome back to the show…",
"txtUrl": "https://api.apify.com/v2/key-value-stores/…/records/transcript-00001.txt",
"srtUrl": "https://api.apify.com/v2/key-value-stores/…/records/transcript-00001.srt",
"vttUrl": "https://api.apify.com/v2/key-value-stores/…/records/transcript-00001.vtt",
"subtitleCues": 412,
"segments": [
{ "start": 0.0, "end": 3.2, "text": "Welcome back to the show." }
]
}

With Include word-level timestamps on, each row also gets "words": [{ "word": "Welcome", "start": 0.0, "end": 0.42 }, …].

Export all results as JSON, CSV or Excel, or download the .srt / .vtt subtitle files directly.

With Subtitle translations set, each row also gets:

"translations": [
{ "language": "fr", "status": "ok", "subtitleCues": 398, "text": "Bienvenue dans l'émission…",
"srtUrl": "…/transcript-00001.fr.srt", "vttUrl": "…/transcript-00001.fr.vtt", "txtUrl": "…/transcript-00001.fr.txt" }
]

Subtitle example (SRT, default settings)

1
00:00:19,260 --> 00:00:23,540
El idioma castellano llegó a Venezuela
en la conquista española llevada a cabo
2
00:00:23,540 --> 00:00:25,900
desde los primeros años del siglo XVI.

How to use

  1. Paste one or more audio/video URLs (direct file links, Drive/Dropbox share links, or podcast RSS feeds).
  2. Optional: set the language code (e.g. en, es, de) for best accuracy — or leave empty to auto-detect.
  3. Optional: choose Translate to English, or list subtitle translation languages (e.g. es, fr, de).
  4. Optional: add a vocabulary hint with names and jargon so they're spelled right.
  5. Optional: adjust subtitle line length and lines per cue (e.g. 1 line × 25 characters for vertical videos).
  6. Click Start.

Input example

{
"mediaUrls": [
"https://example.com/interview.mp4",
"https://feeds.example.com/podcast.xml"
],
"language": "en",
"task": "transcribe",
"prompt": "Apify, Kubernetes, Dr. Nguyen",
"maxEpisodesPerFeed": 3,
"subtitleLanguages": ["es", "fr"],
"subtitleMaxCharsPerLine": 42,
"subtitleMaxLines": "2",
"includeWords": false
}

Pricing

$0.008 per started minute of audio (see the Pricing tab). A 30-minute podcast costs about $0.24; a 1-hour meeting about $0.48. Each extra subtitle translation language is billed like one more transcription of the file: a 10-minute video with subtitles in 2 extra languages = 30 minutes = $0.24. Files that fail (broken link, no audio) and translations that fail are not charged. You can set a maximum cost per run in the run options.

Use cases

  • Podcast show notes, blog posts and SEO content from episodes
  • Meeting and interview notes, research and journalism
  • Subtitles (SRT/VTT) for YouTube, Vimeo, courses and social videos you publish — upload the file directly
  • Call-center and sales-call analysis pipelines
  • Feeding transcripts into ChatGPT/Claude/RAG for summaries and search
  • Automations with Make, Zapier, n8n, or AI agents via the Apify MCP server

Limits & notes

  • Max 2 GB per file. Files must be publicly downloadable (no login pages).
  • YouTube/TikTok page URLs are not supported — use a direct media file link.
  • Speaker labels (diarization) are not included yet — ask on the Issues tab if you need them.
  • Only transcribe content you have the right to process.
  • Audio is sent to a speech-recognition API for processing and is not stored by this Actor beyond your run's storage.

Feedback

Missing a feature (speaker labels, word-level timestamps, summaries)? Open an issue — it gets answered.