Audio & Podcast Transcriber – MP3, WAV, RSS to Text, SRT avatar

Audio & Podcast Transcriber – MP3, WAV, RSS to Text, SRT

Pricing

Pay per usage

Go to Apify Store
Audio & Podcast Transcriber – MP3, WAV, RSS to Text, SRT

Audio & Podcast Transcriber – MP3, WAV, RSS to Text, SRT

Transcribe audio files, video files and podcast RSS feeds to text, timed segments, SRT and VTT with the detected language. Open-source Whisper models (base, small) on Apify, a minute cap you set, new-episodes-only mode.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Tinlark

Tinlark

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

Turn audio files, video files and podcast feeds into text. Give the Actor direct links to media files, or a podcast RSS feed and the number of latest episodes. It returns one row per file or episode: the transcript, timed segments, ready SRT and WebVTT subtitles, and the detected language.

It runs open-source Whisper models on Apify, in two sizes: base (fast, cheapest) and small (slower, usually more accurate). No third-party transcription API is called. Your cost is set by minutes of audio, and you set a hard cap on them.

Use it when you have a list of recordings or a podcast to follow and want the text in a dataset, in a spreadsheet or in an AI pipeline, without running Whisper yourself.

What you get

  • Transcript as one text field, per file or episode.
  • Timed segments: a list of {start, end, text} with times in seconds, for search, quotes and chapter marks.
  • SRT and WebVTT subtitle files as text, ready to save next to a video.
  • Detected language and Whisper's confidence in it, or the language you set.
  • Episode data for feeds: title, guid, publication date and the feed link, on every row.
  • Optional translation to English with Whisper's built-in translation task.
  • Error rows for anything that fails (blocked link, 404, not audio, robots.txt), with the reason and what to do. They are free and never stop the run.

Use cases

  • Searchable archives of interviews, calls, lectures and webinars you recorded.
  • Following a podcast: a scheduled run with New episodes only transcribes just the episodes that appeared since the last run.
  • Subtitles for your own videos: ask for srt or vtt.
  • Text for RAG, summarising or topic search: the dataset is plain JSON. Pair it with Document to Markdown by Tinlark to get documents and audio into the same text form.

How to use it

  1. Paste direct links to audio or video files under Audio or video file links, and/or podcast RSS links under Podcast RSS feeds.
  2. Pick the model (base or small) and the output formats. Set Max total minutes to the most audio you want to pay for in this run.
  3. Start the run. Rows appear in the dataset as files finish. For a podcast, add a schedule and switch on New episodes only.

The prefilled input transcribes a 2-minute public-domain recording and finishes in under a minute.

Input

FieldWhat it doesDefault
Audio or video file links (mediaUrls)Direct http(s) links to mp3, m4a, wav, ogg, flac, mp4, webm and other formats FFmpeg readsprefilled example
Podcast RSS feeds (podcastFeedUrls)RSS or Atom feed links; the feed host's robots.txt is checked firstnone
Model (model)base or smallbase
Spoken language (language)auto, or a code such as en, de, fr, es, jaauto
Output formats (outputFormats)any of text, segments, srt, vtttext, segments
Episodes per feed (episodesPerFeed)The latest N episodes of each feed, 1 to 503
New episodes only (newEpisodesOnly)Skip episodes an earlier run with the same state name already transcribedoff
State name (stateKey)Name of the memory used by New episodes onlydefault
Max total minutes (maxTotalMinutes)Hard cap on audio minutes in this run120
Max minutes per file (maxMinutesPerFile)Longer files are cut here, with a warning180
Translate to English (translateToEnglish)Transcript and segments in English whatever is spokenoff

Example input:

{
"podcastFeedUrls": ["https://librivox.org/rss/1078"],
"mediaUrls": ["https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg"],
"episodesPerFeed": 2,
"model": "small",
"language": "auto",
"outputFormats": ["text", "segments", "srt"],
"maxTotalMinutes": 60
}

Output

One row per transcribed file or episode. A shortened row from the prefilled run (the segment list and text are cut here):

{
"recordType": "transcript",
"status": "ok",
"sourceUrl": "https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg",
"feedUrl": null,
"episodeTitle": null,
"durationSec": 133.2,
"language": "en",
"languageProbability": 0.993,
"model": "base",
"text": "This is a Libravox Recording. All Libravox recordings are in the public domain. For more information or to volunteer, please visit Libravox.org. ...",
"segments": [
{"start": 1.62, "end": 7.46, "text": "This is a Libravox Recording. All Libravox recordings are in the public domain."},
{"start": 7.46, "end": 12.46, "text": "For more information or to volunteer, please visit Libravox.org."}
],
"srt": null,
"vtt": null,
"wordCount": 302,
"billedMinutes": 3,
"processingSec": 18.3,
"warnings": [],
"error": null
}

In this run the recording's spoken name came out as "Libravox" (it is LibriVox), and a web address later in the text as "americanafonic": proper names and unusual words can be misspelled, more so with the smaller model. Read the output of anything where exact names matter.

Feed episodes also fill feedUrl, episodeTitle, episodeGuid and publishedAt. An error row has status: "error", an errorCode and an error text. Codes: unsupported-source (YouTube and social links), not-found, access-denied, too-large, not-media (the link is a web page or a feed), unreadable-media, no-audio-track, robots-disallowed, robots-unavailable, not-a-feed, budget-reached, timeout, download-failed, blocked-address and others. Rows are written as files finish; sourceIndex is a running number. A summary with counts, billed minutes and per-feed numbers is stored under the SUMMARY key.

Podcast feeds and new episodes only

Each feed is read newest episode first, and the first Episodes per feed episodes with an audio or video enclosure are used. With New episodes only on:

  • The first run for a state name transcribes the latest episodes and remembers them.
  • Later runs with the same state name transcribe only episodes not remembered yet. If there is nothing new, the run ends with the message "No new episodes since the last run" and no rows.
  • An episode is remembered only after it was transcribed. Failed or budget-skipped episodes are tried again next time.
  • The memory is a record in a key-value store named audio-podcast-transcriber-state in your Apify account. Use different state names for independent schedules.
  • Only the latest Episodes per feed episodes are looked at. If a feed publishes more than that between two runs, raise the number.

Cost control

  • Max total minutes is a hard cap. A file that would pass it is cut at the cap (the row says so), and the files after it get a free budget-reached error row.
  • Apify's Maximum cost per run also applies once pricing is on: the Actor reads the remaining budget before each file and stops cleanly.
  • A started minute counts as a minute. Each row shows billedMinutes.

Speed and platform cost

Measured on Apify at 4096 MB (the default), int8 on CPU, with both models inside the image so runs do not wait for a download:

ModelSpeed on a 13 min 47 s MP3Apify platform cost per audio hour
base108 s total run; transcription about 9 times faster than real timeabout $0.10
small313 s total run; transcription about 2.7 times faster than real timeabout $0.30

Start-up takes about 15 to 20 seconds per run. Speed depends on the audio and on how busy Apify's hardware is: a 133-second clip took 12 to 18 s with base and about 32 to 39 s with small in Tinlark's test runs. The default memory is 4096 MB. At 2048 MB the same 13-minute file took 251 s instead of 111 s at about the same cost, so 4096 MB is the better choice. Decoding a 3-hour file needs about 1.1 GB of memory on top of the model: for files longer than the default 180 minutes, raise the memory (8192 MB is the maximum). The Actor uses two CPU threads.

Pricing

Free during launch (until 31 October 2026). You pay only Apify's own platform usage for your runs, roughly the figures above.

From 1 November 2026: pay per event, per started audio minute. Prices fall with your Apify plan:

EventFree planBronzeSilverGold
Audio minute, base model$0.020$0.012$0.010$0.008
Audio minute, small model$0.030$0.020$0.017$0.014

So an hour of audio costs $0.72 with base on the Bronze plan, and $1.20 with small. Failed files are never charged. The Store page shows the live prices once they are on.

Limits

  • Direct links to media files and podcast feeds only. YouTube, Instagram, TikTok, Facebook and X links are not supported and give an unsupported-source error row, also when a link redirects there.
  • Files up to 1 GB, at most 600 minutes per file (180 by default). Longer audio is cut with a warning.
  • One download per host at a time. Feed hosts' robots.txt is checked before a feed is fetched; media files are fetched exactly as linked.
  • Files must be public: no login, no cookies, no signed-in pages. Private network addresses are refused.
  • Speech is transcribed as spoken: no speaker names, no punctuation guarantees, no word timings. The model is open-source Whisper; Tinlark has not measured word error rates, so none are claimed. Noisy, accented or overlapping speech is harder, and small usually does better there than base.
  • Music-only or silent files return an empty transcript and a warning ("No speech was detected").
  • Long recordings can contain repeated or invented passages where the speech is unclear: check the output when it matters.

Use with AI agents (MCP)

Apify's MCP server can expose this Actor as a tool, so an AI agent can transcribe a link and read the dataset. Use it when an agent needs the words of an audio or video file at a URL. Through the API:

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("tinlark/audio-podcast-transcriber").call(run_input={
"mediaUrls": ["https://example.com/interview.mp3"],
"model": "base",
"maxTotalMinutes": 30,
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
print(row["language"], row["text"][:200])

FAQ

Which formats work? Anything FFmpeg can decode: mp3, m4a, wav, ogg, flac, mp4, webm and more. The Actor reads the audio track; video files work the same way.

base or small? Start with base. Try small when names, accents or noise matter. It is about three times slower and costs 1.5 to 1.75 times more per minute.

Why set the language? auto listens to the first seconds. A recording that starts with music can be detected wrongly. If you know the language, set it.

Does it keep my audio? Files are downloaded into the run, transcribed inside the run and deleted when it ends. Nothing is sent to a third-party AI service. The transcript is in your dataset.

Can it transcribe a YouTube video? No. Use a direct file link or a podcast feed.

How do I get SRT or VTT? Add srt or vtt to Output formats. The subtitle file is a text field in the row; save it with the .srt or .vtt extension.

Disclaimers and legality. Process only recordings you have the right to use. You are responsible for lawful use of the audio and of the output, including copyright, consent and the personal data that speech can contain. The Actor fetches exactly the links you give it, and the episodes in feeds you give it, with a User-Agent that names it and gives a contact address. It does not search, crawl or log in anywhere, and it refuses YouTube and social-media links. Do not use it to get around access controls or the terms of a site. Transcription is automatic and contains errors: do not rely on it alone for legal, medical or financial decisions.

Support

Something wrong or missing? Open an issue on this Actor's Issues tab with your input (the link or feed, the model) and what you expected. Related: Document to Markdown by Tinlark turns PDF, Word, Excel and scans into Markdown for the same kind of pipeline.