Audio & Video to Text - Speech to Text Transcription, SRT
Pricing
Pay per event
Audio & Video to Text - Speech to Text Transcription, SRT
Whisper transcription: audio to text and video to text from files, Drive/Dropbox links and podcast RSS feeds, with timestamps and SRT/VTT subtitles. 90+ languages, optional speaker labels. $0.006/min.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Yukai Lin
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 hours ago
Last modified
Categories
Share
What does Audio & Video Transcriber do?
It turns audio files, videos and podcast episodes into text with timestamps, readable paragraphs, and ready-to-use SRT and VTT subtitle files. It uses OpenAI's open Whisper large-v3-turbo model, detects the language automatically and supports 90+ languages.
- 🎙️ Podcasts, interviews, meetings, lectures, webinars, audiobooks, voice notes
- 🎬 Video to text: MP4, MOV, MKV, WEBM and more. The audio track is extracted automatically
- 📡 Podcast RSS feeds: paste a feed (or an Apple Podcasts show link) and the newest episodes are transcribed, with episode title and publish date
- 📝 Publisher transcripts, no audio minutes: when a podcast publishes its own transcript in the feed (Podcasting 2.0
<podcast:transcript>: VTT, SRT, JSON, HTML or text), it is used instead of transcribing, with the same output andtranscriptSource: "publisher" - 🗓️ Feed filters: only episodes published since / until a date, or whose title contains (or does not contain) a text
- 📤 File upload: upload recordings from your computer in Apify Console (
audioFiles) - 🌐 Web pages with one player: a page with exactly one audio or video file (e.g. a news, lecture or archive.org page); the file is found and transcribed
- 🔗 Google Drive, Dropbox and OneDrive share links work if the file is shared publicly
- ⏱️ Timestamped segments, paragraphs and full plain text in one result
- 🗣️ Speaker labels (
"speakerLabels": true): who said what, as Speaker 1, Speaker 2… on every segment, paragraph and subtitle line, plus the number of speakers - 🔠 Word-level timestamps (
"wordTimestamps": true): start and end time of every word, for video editing and precise subtitles (included in the regular price) - 🌍 Translated subtitles in one step (
"subtitleLanguages": ["Spanish", "German"]): SRT and VTT in up to 5 extra languages, timings and speaker labels kept - 🔁 Only new episodes (
"newEpisodesOnly": true): scheduled runs skip episodes and files that were already transcribed, and only take episodes released after the newest one already done, so a day without a new episode costs nothing (backfillOlderEpisodesworks through the back catalogue instead) - 🧾 SRT and VTT subtitles saved for every file (no extra charge), cut into subtitle-sized cues: at most 7 seconds and two lines of 42 characters, timed by word
- 🧠 Optional AI summary and chapters (
"summarize": true), with no API key needed - 🌍 Automatic language detection, or set the language yourself
- 🔤 Vocabulary hints: give names and terms (e.g. "Apify, Kubernetes") so they are spelled correctly
- 📏 Large files: up to 2 GB per file, any length. Long recordings are cut in pauses and transcribed in parallel
- 🔕 Silent files are free: files with no speech are reported but not charged
- ⚡ Fast: in our tests an 8-minute MP4 video took 33 seconds and a 26-minute MP3 took 109 seconds
How much does it cost?
| Event | Price |
|---|---|
| Audio minute | $0.006 per minute ($0.36 per hour) |
| Audio minute with speaker labels (optional) | $0.015 per minute ($0.90 per hour), instead of the regular audio minute, only when speakerLabels is on |
| Translated subtitles (optional) | $0.002 per minute per language (e.g. a 30-minute episode in Spanish and German = $0.12), only for languages that succeed |
| AI summary and chapters (optional) | $0.01 per file, only when summarize is on |
- No start fee. You pay only for transcribed audio.
- SRT and VTT subtitles, paragraphs and timestamps are included.
- Billed per started minute of audio (a 500-second video is 9 minutes).
- Word-level timestamps are included in the regular $0.006 per minute (Whisper word timings). With speaker labels on, words come from nova-3 and the speaker-label price applies, not both.
- Translated subtitles (
subtitleLanguages, up to 5 languages): SRT + VTT per language with the original timings and speaker labels, made right after the transcript. A language that fails is free and reported intranslatedSubtitles. - Not charged: files that cannot be downloaded or decoded, files with no speech, files skipped by your limits (
maxMinutesPerFile,maxTotalMinutes), files skipped bynewEpisodesOnly, episodes removed by the feed filters, and the audio minutes of episodes that use the publisher's own transcript (withsummarizeon, their AI summary is still $0.01). - Higher Apify plans get volume discounts (Silver 10%, Gold and above 20%).
Control your cost
- What is charged: the minutes of each transcribed file (rounded up per file) and, with
summarizeon, one AI summary per file. Minutes are charged only after the transcript row has been saved. - What is free: lines that are not a URL, duplicate links, YouTube / TikTok / Instagram / Spotify page links, web pages without exactly one media file, files that cannot be downloaded or decoded, files without speech, files skipped by your limits, and publisher transcripts (
billedMinutes: 0). Every row hascharged: true/false, and failed rows say "(not charged)". - Plan at the start: the run's status message and log start with the cost plan (audio minutes known from podcast feeds, summaries, and how many files are measured when downloaded); it is also saved as
SUMMARY.costPlan. - Caps: Max minutes per file and Max total minutes per run (Advanced settings) skip files before they are transcribed, using the duration read from the file (or from the podcast feed, before downloading).
- Max charge per run: set Maximum cost per run in the run options (or
maxTotalChargeUsdin the API). A file is only started when the limit still covers all its minutes; when it does not, the run stops cleanly:SUMMARY.statusisLIMIT_REACHEDandSUMMARY.notProcessedlists the files that were not started (count and up to 100 inputs). The AI summary is only made when the limit also covers it. - If Apify restarts the run (server migration or Resurrect), items already finished are skipped and not charged again (
SUMMARY.resumedSkipped). - Time limit per file (
fileTimeoutSecs, default 1,800 seconds): a file whose download, conversion and transcription take longer gets a row witherrorType: "timeout"and is not charged, and the run goes on with the next file. Raise it for recordings of several hours. - Files in parallel (
fileConcurrency, default 3, up to 10): several files are downloaded and transcribed at the same time, so podcast batches finish sooner. Rows are saved in the order files finish;inputIndexis each file's position in the input. Your max charge per run and Max total minutes per run still hold: before a file is transcribed, its cost is reserved together with the files already in progress, so files running side by side can never go over the limit. With little run memory fewer files run at once (about 3 at the default 1,024 MB); the time a file waits for budget held by other files does not count againstfileTimeoutSecs.
Price comparison (checked September 2026)
Per-minute prices of other transcription Actors on Apify Store, Free plan unless noted; "not stated" means we did not find it in that Actor's input or README:
| Actor | Price | Speaker labels |
|---|---|---|
| This Actor | $0.006 per minute, no start fee (Gold+ $0.0048) | $0.015 per minute instead of $0.006 |
| sian.agency / INCREDIBLY-FAST-audio-transcriber | $0.30 per minute ($0.005 per second) on the Free plan (Bronze $0.015, Gold $0.009), plus $0.005 per run | extra $0.001 per second on the Free plan, $0.0001 per second on paid plans |
| memo23 / video-audio-transcriber | $0.048 per minute (billed per second, $0.0008 per second), plus $0.005 per run | yes |
| amanatools / whisper-transcriber | $0.008 per minute, plus $0.005 per file | not stated |
| steadyfetch / media-transcriber | $0.003 per minute | not stated |
| kaz_kakyo / audio-transcriber | $0.01 per minute; chapters $0.01 per file | yes (Deepgram nova-3) |
| viralanalyzer / video-transcriber | $0.015 per minute | not stated |
| hgservices / speech-to-text | $0.042 per minute (Gold $0.024), plus $0.09 per YouTube video | yes |
| tictechid / vanzi-universal-transcriber | $0.0025 per second ($0.15 per minute), plus $0.01 per run | not stated |
| parseforge / audio-transcriber | $0.44 per file, plus $0.012 per run | not stated |
We are not the cheapest per minute: steadyfetch is lower. Some of these Actors also accept YouTube or TikTok links, which this Actor does not.
Supported input
| Input | Details |
|---|---|
| Audio files | MP3, WAV, M4A, AAC, OGG, OPUS, FLAC, AIFF, WMA, AMR and other formats ffmpeg can read |
| Video files | MP4, MOV, MKV, WEBM, AVI and more. Only the audio track is used |
| Size | Up to 2 GB per file, any length |
| Share links | Google Drive, Dropbox, OneDrive / SharePoint (the file must be shared publicly) |
| Uploads | audioFiles: files uploaded in Apify Console (stored in a key-value store of your account). In API calls, pass file URLs in urls |
| Web pages | A page with exactly one <audio> / <video> file, or one og:audio / og:video / JSON-LD contentUrl file link; the found file is in mediaUrl |
| Podcast feeds | Any podcast RSS feed, or an Apple Podcasts show or episode link. Publisher transcripts are used when the feed has them; date and title filters are optional |
| Dataset | The datasetId of another Actor run; the media URL field is detected automatically or set with datasetField |
Links must be downloadable without logging in. YouTube, TikTok, Instagram, Facebook, X, Vimeo, Loom, SoundCloud and Spotify page links are web pages, not files: each comes back as a free row with errorType: "unsupported_platform" and a hint. Download the media first (for example with a downloader Actor from Apify Store) and pass the file URL, or that Actor's dataset via datasetId. For Vimeo and Loom, use the file behind the video's Download button when the owner allows downloads; direct file links on these hosts, such as a Vimeo "video file link" (player.vimeo.com/external/... or .../progressive_redirect/...), are transcribed normally.
Web pages: for any other web page, the page is read and its media file is looked up: the one <audio> or <video> element (its src or its first playable <source>; several sources of one element are one file), or else one direct file named by og:audio, og:video, twitter:player:stream or JSON-LD contentUrl. Only direct files count: streaming playlists (HLS .m3u8, DASH .mpd) and player pages are not used. When the page has no such file (errorType: "no_media") or several ("multiple_media", listed in mediaCandidates), the row is free; pass the file you want as a direct URL.
A line that is not a URL does not stop the run: it becomes one free row with errorType: "invalid_input" and the other files are processed. Use urls (a plain list of strings) in API calls; the request-list field audioUrls from earlier versions still works.
How to use it
- Paste audio or video URLs (or pages with one player), upload files, and/or add podcast feeds (optionally with date and title filters).
- Optional: set the language (e.g.
en,es,ja), vocabulary hints, and turn on AI summary and chapters, speaker labels or word-level timestamps. - Optional (Advanced settings): cap the cost with Max minutes per file and Max total minutes per run; for schedules, turn on Only new episodes with a monitor name.
- Click Start. Each file becomes one result with the text, segments, paragraphs and links to its SRT/VTT files.
Input example
{"urls": ["https://example.com/meeting-recording.mp4"],"podcastFeeds": ["https://feeds.npr.org/510289/podcast.xml"],"maxEpisodesPerFeed": 3,"language": "en","vocabulary": "Apify, Cloudflare, TidyTools","summarize": true,"maxMinutesPerFile": 90,"maxTotalMinutes": 300}
Output example (real result, shortened)
An 8-minute MP4 video with "summarize": true:
{"url": "https://archive.org/download/gerald-ford-inaugural-address-august-9-1974-720p/Gerald%20Ford%20inaugural%20address_%20August%209%2C%201974%20%28720p%29.mp4","success": true,"format": "mp4","fileSizeBytes": 48964530,"durationSeconds": 500.1,"billedMinutes": 9,"language": "en","wordCount": 876,"text": "Mr. Chief Justice, my dear friends, my fellow Americans, the oath that I have taken is the same oath that was taken by George Washington...","segments": [{ "start": 1.36, "end": 20.4, "text": "Mr. Chief Justice, my dear friends, my fellow Americans, the oath that I have taken is the same oath that was taken by George Washington and by every president under the Constitution." },{ "start": 20.4, "end": 31.32, "text": "But I assume the presidency under extraordinary circumstances never before experienced by Americans." }],"paragraphs": [{ "start": 1.36, "end": 39.68, "text": "Mr. Chief Justice, my dear friends, ... This is an hour of history that troubles our minds and hurts our hearts." },{ "start": 41.44, "end": 73.64, "text": "Therefore, I feel it is my first duty to make an unprecedented compact with my countrymen. ..." }],"summary": "The president assumes office under extraordinary circumstances ... and asks for prayers for Richard Nixon and his family.","chapters": [{ "startSeconds": 0, "start": "0:00", "title": "Introduction and Compact with the Nation" },{ "startSeconds": 304, "start": "5:04", "title": "Message of Hope and Unity" }],"srtUrl": "https://api.apify.com/v2/key-value-stores/.../records/0000-Gerald-20Ford-...mp4.srt","vttUrl": "https://api.apify.com/v2/key-value-stores/.../records/0000-Gerald-20Ford-...mp4.vtt","processingSeconds": 33}
With "speakerLabels": true (a real 74-minute Changelog & Friends episode: 3 speakers found, 13,069 words, 195 seconds of processing, 74 minutes billed at $0.015):
{"success": true,"model": "nova-3","durationSeconds": 4390.7,"billedMinutes": 74,"speakers": 3,"segments": [{ "start": 16.14, "end": 23.34, "speaker": 1, "text": "Welcome to Changelog and Friends, a weekly talk show about thinking outside the Dropbox." }],"paragraphs": [{ "start": 16.14, "end": 34.87, "speaker": 1, "text": "Welcome to Changelog and Friends, a weekly talk show about thinking outside the Dropbox. Thanks as always to our partners at fly.io, ..." },{ "start": 43.38, "end": 69.23, "speaker": 2, "text": "This is the year we almost break the database. Let me explain. Where do agents actually store their stuff? ..." }],"speakerTranscript": "Speaker 1 [0:16]: Welcome to Changelog and Friends, ...\n\nSpeaker 2 [0:43]: This is the year we almost break the database. ..."}
With "wordTimestamps": true, every segment also gets its words (real result, a 5-minute news episode):
{ "start": 4.16, "end": 7.52, "speaker": 1, "text": "What's up, friends? Adam here. Big week here at ChangeLog.","words": [{ "word": "What's", "start": 4.16, "end": 4.4, "speaker": 1 }, { "word": "up,", "start": 4.4, "end": 4.56, "speaker": 1 }, { "word": "friends?", "start": 4.56, "end": 4.96, "speaker": 1 }] }
In the SRT file a line starts with Speaker 2: when the speaker changes; the VTT file uses <v Speaker 2> voice tags. Files longer than 50 minutes are sent in overlapping pieces, and the speakers of each piece are matched by the words in the overlap, so Speaker 1 stays Speaker 1 for the whole file.
A podcast episode from an RSS feed also carries the episode details:
{"url": "https://tracking.swap.fm/track/.../default.mp3?...","podcastTitle": "Planet Money","episodeTitle": "Middlegarchs are the new Oligarchs","pubDate": "2026-09-25T21:46:44.000Z","episodeUrl": "https://www.npr.org/2026/09/25/nx-s1-5981194/stealthy-wealthy-everywhere-millionaires-pass-through","feedUrl": "https://feeds.npr.org/510289/podcast.xml","success": true,"format": "mp3","durationSeconds": 2065.2,"billedMinutes": 35,"wordCount": 5505}
Publisher transcripts (no audio minutes)
Many podcasts put their own transcript in the feed (<podcast:transcript>, Podcasting 2.0). With usePublisherTranscripts on (the default), such an episode is not transcribed: the publisher's file is downloaded and converted to the same output. Timed formats (VTT, SRT, JSON) give segments with timestamps, paragraphs and our SRT/VTT files; HTML and plain text give the text and paragraphs only (timestamps: false, no subtitle files). Speaker names from the transcript are kept as speakerName on each segment. No audio minutes are charged (billedMinutes: 0); the AI summary, if you turn it on, is charged as usual.
The transcript is checked first: it must have at least 20 words, cover at least half of the episode's duration (from the feed) without running far past it, and have a plausible number of words per minute. A transcript that fails the check, or cannot be downloaded, is ignored and the episode is transcribed normally; the row then says why in publisherTranscriptError. Set "usePublisherTranscripts": false to always transcribe the audio yourself.
Real result (Podnews Daily, whose feed has a VTT transcript per episode; with "summarize": true, so only the $0.01 summary was charged):
{"podcastTitle": "Podnews Daily - podcast industry and podcasting news","episodeTitle": "Was Inception Point AI just a PR stunt?","pubDate": "2026-09-29T10:30:00.000Z","success": true,"transcriptSource": "publisher","transcriptUrl": "https://podnews.net/audio/podnews260929.mp3.vtt","transcriptFormat": "vtt","timestamps": true,"model": "publisher","language": "en","durationSeconds": 273,"billedMinutes": 0,"wordCount": 642,"segments": [{ "start": 0.56, "end": 4.96, "text": "From Singapore Airport, the latest from podnews.net with The Podglomerate.", "speakerName": "James Cridland" }],"summary": "Inception Point AI's CEO Jeanine Wright has suggested that the company's thousands of AI-made podcasts may have been perceived as a publicity stunt, ...","summaryCharged": true}
In the same test, a 63-minute Buzzcast episode (12,067 words, 810 segments, 10 chapters) was also taken from the publisher's transcript: 0 audio minutes instead of 64.
Feed filters
{"podcastFeeds": ["https://podnews.net/rss"],"episodesSince": "2026-09-20","episodesUntil": "2026-09-28","titleIncludes": ["patreon", "spotify*video"],"titleExcludes": ["weekly"],"maxEpisodesPerFeed": 1}
episodesSince: a date (YYYY-MM-DD) or a relative time such as30 days,2 weeks,6 months.episodesUntil: a date; the whole day counts.titleIncludes: keep an episode when its title contains one of the texts;titleExcludes: drop it when its title contains one. Case does not matter;*matches anything (spotify*video).- Filters are applied first, then
newEpisodesOnly, thenmaxEpisodesPerFeed. With a date filter, episodes without a publish date are dropped. A link to one Apple Podcasts episode is not filtered. - Filtered episodes get no row and are not charged;
SUMMARY.episodesFilteredcounts them. In our test above, 150 of 151 episodes were filtered and the one match ("Patreon removes per-creation memberships", 4.6 minutes) was transcribed.
Web page with one player
{ "urls": ["https://archive.org/details/nasa_tv-ScienceCast_-_Record-Setting_Asteroid_Flyby"] }
Real result (shortened): the page's og:video file was found and transcribed, 262 seconds, 5 minutes billed:
{"url": "https://archive.org/details/nasa_tv-ScienceCast_-_Record-Setting_Asteroid_Flyby","mediaUrl": "https://archive.org/download/nasa_tv-ScienceCast_-_Record-Setting_Asteroid_Flyby/ScienceCast_-_Record-Setting_Asteroid_Flyby.mp4","success": true,"transcriptSource": "asr","format": "mp4","billedMinutes": 5,"segments": [{ "start": 3.7, "end": 8.48, "text": "Record-setting asteroid flyby, presented by Science at NASA." }]}
Results that are not charged
Every row has success: true or false. Rows with success: false are free and carry an errorType:
| errorType | Meaning |
|---|---|
no_text | No speech in the audio (silence or music only); the row also has noSpeech: true. This is a valid answer, not a failure: it does not make the run PARTIAL_RESULTS and is counted in SUMMARY.noSpeech |
too_large | Longer than maxMinutesPerFile, or larger than 2 GB |
limit_reached | Would go over maxTotalMinutes or your Apify spending limit |
unsupported | Not an audio or video file (e.g. a web page), no audio track, or a language that speaker labels do not support (e.g. Chinese) |
unsupported_platform | A YouTube, TikTok, Instagram, Facebook, X, Vimeo, Loom, SoundCloud or Spotify page link (not a file); the error says how to get the file |
no_media | A web page without an audio or video file we can use (no player, or only a streaming playlist or player page) |
multiple_media | A web page with several audio or video files; they are listed in mediaCandidates |
invalid_input | A line that is not a URL |
not_found, blocked, network, timeout, http_error | The file could not be downloaded |
The run's key-value store also has a SUMMARY record with status (SUCCESS, PARTIAL_RESULTS, FAILED, NO_RESULTS or LIMIT_REACHED), files transcribed, files without speech (noSpeech, counted as successes), minutes billed and more.
Use with AI agents (MCP)
Connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/audio-transcriber) to Claude, Cursor or any MCP client, then ask e.g. "Transcribe the last 2 episodes of this podcast feed and give me the chapters".
Minimal input:
{ "urls": ["https://example.com/episode.mp3"] }
Failed or skipped items are not charged and carry an errorType, so an agent can see at a glance what worked.
Scheduled podcast transcription (only new episodes)
{"podcastFeeds": ["https://changelog.com/news/feed"],"maxEpisodesPerFeed": 1,"newEpisodesOnly": true,"monitorName": "changelog-news"}
Run it on an Apify schedule (e.g. weekly). The first run transcribes the newest episode; each later run transcribes only episodes it has not transcribed before (in our test, the second run skipped "Bitwarden CLI compromised" and took the next one). Runs with the same monitorName share the list, which is kept in a named key-value store (audio-transcriber-<monitorName>). When there is nothing new, the run ends without charges. Failed files are not remembered, so they are tried again next time.
Chaining with other Actors
Set datasetId to the dataset of another run (for example a podcast or video scraper). The media URL field is found automatically (audioUrl, mediaUrl, videoUrl, url, ...), or set it with datasetField, e.g. media.audioUrl.
Subtitles in other languages: pass the srtUrl or vttUrl of a result to our Bulk Text & JSON Translator (subtitleUrls). It translates the subtitles cue by cue, keeps every timestamp, and returns new SRT/VTT files in any language, at $1.50 per million characters. Speech has about 1,000 characters per minute, so that is roughly $0.0015 per audio minute and language (in our test, 20 minutes of subtitles into Spanish and Traditional Chinese cost $0.059).
Limitations
- Transcription only: the turbo model does not translate. Use our Bulk Text & JSON Translator on the text or the subtitle files if you need another language.
- Speaker labels are numbers (Speaker 1, 2…), not names. Short interjections ("Yeah.") are sometimes given to the wrong speaker, and speaker labels are not available for every language (Chinese is not). When a speaker change is detected a few words late ("…Welcome to | the podcast. Thank you"), those words are moved back to the end of the previous speaker's sentence.
- Segments that Whisper marks as probably not speech (its no-speech probability) are dropped, together with typical filler it invents over music or silence ("Thanks for watching!"); the count is in
segmentsFiltered. - Podcast hosts often insert ads when the file is downloaded; those minutes are part of the file and are billed.
maxMinutesPerFileuses the duration from the feed for podcast episodes. - Long recordings are cut into pieces of about 90 seconds at the quietest moment nearby; a word at a cut can occasionally be lost. Whisper sometimes skips several seconds of speech inside a piece: every stretch of 5 seconds or more without text is transcribed again on its own, and what is found there is added (
SUMMARY.segmentsRecovered; no extra charge). - Publisher transcripts are the publisher's own text: their wording, timing and speaker names are used as they are (only checked for being complete and matching the episode length). They describe the publisher's version of the file, which may differ slightly from the file you download when ads are inserted.
- Uploaded files are read from your key-value store by the Actor; if a record cannot be read (errorType
blocked), upload it again with the file upload field or pass a public or signed link to the file. - Accuracy depends on audio quality, accents and background noise.
FAQ
Which speech to text model does it use? OpenAI's open Whisper large-v3-turbo model, with automatic language detection and 90+ languages. Speaker labels, when turned on, come from nova-3.
How much does audio to text cost? $0.006 per minute ($0.36 per hour), no start fee. SRT and VTT subtitles, paragraphs and timestamps are included; files with no speech are free.
Can it transcribe new podcast episodes automatically? Yes. Paste the podcast RSS feed, turn on newEpisodesOnly and run it on a schedule: only episodes released since the last run are transcribed, so a day without a new episode costs nothing.
Can I turn a video into text and subtitles? Yes. MP4, MOV, MKV, WEBM and more: the audio track is extracted automatically, and every file gets text, timestamps and SRT/VTT subtitle files.
Support
Open an issue in the Issues tab with the audio URL. Issues are checked regularly.