Audio & Podcast Transcriber – MP3, WAV, RSS to Text, SRT
Pricing
Pay per usage
Audio & Podcast Transcriber – MP3, WAV, RSS to Text, SRT
Transcribe audio files, video files and podcast RSS feeds to text, timed segments, SRT and VTT with the detected language. Open-source Whisper models (base, small) on Apify, a minute cap you set, new-episodes-only mode.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Tinlark
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 hours ago
Last modified
Categories
Share
Turn audio files, video files and podcast feeds into text. Give the Actor direct links to media files, or a podcast RSS feed and the number of latest episodes. It returns one row per file or episode: the transcript, timed segments, ready SRT and WebVTT subtitles, and the detected language.
It runs open-source Whisper models on Apify, in two sizes: base (fast, cheapest) and small (slower, usually more accurate). No third-party transcription API is called. Your cost is set by minutes of audio, and you set a hard cap on them.
Use it when you have a list of recordings or a podcast to follow and want the text in a dataset, in a spreadsheet or in an AI pipeline, without running Whisper yourself.
What you get
- Transcript as one text field, per file or episode.
- Timed segments: a list of
{start, end, text}with times in seconds, for search, quotes and chapter marks. - SRT and WebVTT subtitle files as text, ready to save next to a video.
- Detected language and Whisper's confidence in it, or the language you set.
- Episode data for feeds: title, guid, publication date and the feed link, on every row.
- Optional translation to English with Whisper's built-in translation task.
- Error rows for anything that fails (blocked link, 404, not audio, robots.txt), with the reason and what to do. They are free and never stop the run.
Use cases
- Searchable archives of interviews, calls, lectures and webinars you recorded.
- Following a podcast: a scheduled run with New episodes only transcribes just the episodes that appeared since the last run.
- Subtitles for your own videos: ask for
srtorvtt. - Text for RAG, summarising or topic search: the dataset is plain JSON. Pair it with Document to Markdown by Tinlark to get documents and audio into the same text form.
How to use it
- Paste direct links to audio or video files under Audio or video file links, and/or podcast RSS links under Podcast RSS feeds.
- Pick the model (
baseorsmall) and the output formats. Set Max total minutes to the most audio you want to pay for in this run. - Start the run. Rows appear in the dataset as files finish. For a podcast, add a schedule and switch on New episodes only.
The prefilled input transcribes a 2-minute public-domain recording and finishes in under a minute.
Input
| Field | What it does | Default |
|---|---|---|
Audio or video file links (mediaUrls) | Direct http(s) links to mp3, m4a, wav, ogg, flac, mp4, webm and other formats FFmpeg reads | prefilled example |
Podcast RSS feeds (podcastFeedUrls) | RSS or Atom feed links; the feed host's robots.txt is checked first | none |
Model (model) | base or small | base |
Spoken language (language) | auto, or a code such as en, de, fr, es, ja | auto |
Output formats (outputFormats) | any of text, segments, srt, vtt | text, segments |
Episodes per feed (episodesPerFeed) | The latest N episodes of each feed, 1 to 50 | 3 |
New episodes only (newEpisodesOnly) | Skip episodes an earlier run with the same state name already transcribed | off |
State name (stateKey) | Name of the memory used by New episodes only | default |
Max total minutes (maxTotalMinutes) | Hard cap on audio minutes in this run | 120 |
Max minutes per file (maxMinutesPerFile) | Longer files are cut here, with a warning | 180 |
Translate to English (translateToEnglish) | Transcript and segments in English whatever is spoken | off |
Example input:
{"podcastFeedUrls": ["https://librivox.org/rss/1078"],"mediaUrls": ["https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg"],"episodesPerFeed": 2,"model": "small","language": "auto","outputFormats": ["text", "segments", "srt"],"maxTotalMinutes": 60}
Output
One row per transcribed file or episode. A shortened row from the prefilled run (the segment list and text are cut here):
{"recordType": "transcript","status": "ok","sourceUrl": "https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg","feedUrl": null,"episodeTitle": null,"durationSec": 133.2,"language": "en","languageProbability": 0.993,"model": "base","text": "This is a Libravox Recording. All Libravox recordings are in the public domain. For more information or to volunteer, please visit Libravox.org. ...","segments": [{"start": 1.62, "end": 7.46, "text": "This is a Libravox Recording. All Libravox recordings are in the public domain."},{"start": 7.46, "end": 12.46, "text": "For more information or to volunteer, please visit Libravox.org."}],"srt": null,"vtt": null,"wordCount": 302,"billedMinutes": 3,"processingSec": 18.3,"warnings": [],"error": null}
In this run the recording's spoken name came out as "Libravox" (it is LibriVox), and a web address later in the text as "americanafonic": proper names and unusual words can be misspelled, more so with the smaller model. Read the output of anything where exact names matter.
Feed episodes also fill feedUrl, episodeTitle, episodeGuid and publishedAt. An error row has status: "error", an errorCode and an error text. Codes: unsupported-source (YouTube and social links), not-found, access-denied, too-large, not-media (the link is a web page or a feed), unreadable-media, no-audio-track, robots-disallowed, robots-unavailable, not-a-feed, budget-reached, timeout, download-failed, blocked-address and others. Rows are written as files finish; sourceIndex is a running number. A summary with counts, billed minutes and per-feed numbers is stored under the SUMMARY key.
Podcast feeds and new episodes only
Each feed is read newest episode first, and the first Episodes per feed episodes with an audio or video enclosure are used. With New episodes only on:
- The first run for a state name transcribes the latest episodes and remembers them.
- Later runs with the same state name transcribe only episodes not remembered yet. If there is nothing new, the run ends with the message "No new episodes since the last run" and no rows.
- An episode is remembered only after it was transcribed. Failed or budget-skipped episodes are tried again next time.
- The memory is a record in a key-value store named
audio-podcast-transcriber-statein your Apify account. Use different state names for independent schedules. - Only the latest Episodes per feed episodes are looked at. If a feed publishes more than that between two runs, raise the number.
Cost control
- Max total minutes is a hard cap. A file that would pass it is cut at the cap (the row says so), and the files after it get a free
budget-reachederror row. - Apify's Maximum cost per run also applies once pricing is on: the Actor reads the remaining budget before each file and stops cleanly.
- A started minute counts as a minute. Each row shows
billedMinutes.
Speed and platform cost
Measured on Apify at 4096 MB (the default), int8 on CPU, with both models inside the image so runs do not wait for a download:
| Model | Speed on a 13 min 47 s MP3 | Apify platform cost per audio hour |
|---|---|---|
| base | 108 s total run; transcription about 9 times faster than real time | about $0.10 |
| small | 313 s total run; transcription about 2.7 times faster than real time | about $0.30 |
Start-up takes about 15 to 20 seconds per run. Speed depends on the audio and on how busy Apify's hardware is: a 133-second clip took 12 to 18 s with base and about 32 to 39 s with small in Tinlark's test runs. The default memory is 4096 MB. At 2048 MB the same 13-minute file took 251 s instead of 111 s at about the same cost, so 4096 MB is the better choice. Decoding a 3-hour file needs about 1.1 GB of memory on top of the model: for files longer than the default 180 minutes, raise the memory (8192 MB is the maximum). The Actor uses two CPU threads.
Pricing
Free during launch (until 31 October 2026). You pay only Apify's own platform usage for your runs, roughly the figures above.
From 1 November 2026: pay per event, per started audio minute. Prices fall with your Apify plan:
| Event | Free plan | Bronze | Silver | Gold |
|---|---|---|---|---|
| Audio minute, base model | $0.020 | $0.012 | $0.010 | $0.008 |
| Audio minute, small model | $0.030 | $0.020 | $0.017 | $0.014 |
So an hour of audio costs $0.72 with base on the Bronze plan, and $1.20 with small. Failed files are never charged. The Store page shows the live prices once they are on.
Limits
- Direct links to media files and podcast feeds only. YouTube, Instagram, TikTok, Facebook and X links are not supported and give an
unsupported-sourceerror row, also when a link redirects there. - Files up to 1 GB, at most 600 minutes per file (180 by default). Longer audio is cut with a warning.
- One download per host at a time. Feed hosts' robots.txt is checked before a feed is fetched; media files are fetched exactly as linked.
- Files must be public: no login, no cookies, no signed-in pages. Private network addresses are refused.
- Speech is transcribed as spoken: no speaker names, no punctuation guarantees, no word timings. The model is open-source Whisper; Tinlark has not measured word error rates, so none are claimed. Noisy, accented or overlapping speech is harder, and
smallusually does better there thanbase. - Music-only or silent files return an empty transcript and a warning ("No speech was detected").
- Long recordings can contain repeated or invented passages where the speech is unclear: check the output when it matters.
Use with AI agents (MCP)
Apify's MCP server can expose this Actor as a tool, so an AI agent can transcribe a link and read the dataset. Use it when an agent needs the words of an audio or video file at a URL. Through the API:
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_TOKEN>")run = client.actor("tinlark/audio-podcast-transcriber").call(run_input={"mediaUrls": ["https://example.com/interview.mp3"],"model": "base","maxTotalMinutes": 30,})for row in client.dataset(run["defaultDatasetId"]).iterate_items():print(row["language"], row["text"][:200])
FAQ
Which formats work? Anything FFmpeg can decode: mp3, m4a, wav, ogg, flac, mp4, webm and more. The Actor reads the audio track; video files work the same way.
base or small? Start with base. Try small when names, accents or noise matter. It is about three times slower and costs 1.5 to 1.75 times more per minute.
Why set the language? auto listens to the first seconds. A recording that starts with music can be detected wrongly. If you know the language, set it.
Does it keep my audio? Files are downloaded into the run, transcribed inside the run and deleted when it ends. Nothing is sent to a third-party AI service. The transcript is in your dataset.
Can it transcribe a YouTube video? No. Use a direct file link or a podcast feed.
How do I get SRT or VTT? Add srt or vtt to Output formats. The subtitle file is a text field in the row; save it with the .srt or .vtt extension.
Disclaimers and legality. Process only recordings you have the right to use. You are responsible for lawful use of the audio and of the output, including copyright, consent and the personal data that speech can contain. The Actor fetches exactly the links you give it, and the episodes in feeds you give it, with a User-Agent that names it and gives a contact address. It does not search, crawl or log in anywhere, and it refuses YouTube and social-media links. Do not use it to get around access controls or the terms of a site. Transcription is automatic and contains errors: do not rely on it alone for legal, medical or financial decisions.
Support
Something wrong or missing? Open an issue on this Actor's Issues tab with your input (the link or feed, the model) and what you expected. Related: Document to Markdown by Tinlark turns PDF, Word, Excel and scans into Markdown for the same kind of pipeline.