# Audio & Podcast Transcriber – MP3, WAV, RSS to Text, SRT (`tinlark/audio-podcast-transcriber`) Actor

Transcribe audio files, video files and podcast RSS feeds to text, timed segments, SRT and VTT with the detected language. Open-source Whisper models (base, small) on Apify, a minute cap you set, new-episodes-only mode.

- **URL**: https://apify.com/tinlark/audio-podcast-transcriber.md
- **Developed by:** [Tinlark](https://apify.com/tinlark) (community)
- **Categories:** AI, Automation, For creators
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Audio & Podcast Transcriber – MP3, WAV, RSS to Text, SRT

Turn audio files, video files and podcast feeds into text. Give the Actor direct links to media files, or a podcast RSS feed and the number of latest episodes. It returns one row per file or episode: the transcript, timed segments, ready SRT and WebVTT subtitles, and the detected language.

It runs open-source Whisper models on Apify, in two sizes: **base** (fast, cheapest) and **small** (slower, usually more accurate). No third-party transcription API is called. Your cost is set by minutes of audio, and you set a hard cap on them.

Use it when you have a list of recordings or a podcast to follow and want the text in a dataset, in a spreadsheet or in an AI pipeline, without running Whisper yourself.

### What you get

- **Transcript** as one text field, per file or episode.
- **Timed segments**: a list of `{start, end, text}` with times in seconds, for search, quotes and chapter marks.
- **SRT and WebVTT** subtitle files as text, ready to save next to a video.
- **Detected language** and Whisper's confidence in it, or the language you set.
- **Episode data** for feeds: title, guid, publication date and the feed link, on every row.
- **Optional translation to English** with Whisper's built-in translation task.
- **Error rows** for anything that fails (blocked link, 404, not audio, robots.txt), with the reason and what to do. They are free and never stop the run.

### Use cases

- Searchable archives of interviews, calls, lectures and webinars you recorded.
- Following a podcast: a scheduled run with **New episodes only** transcribes just the episodes that appeared since the last run.
- Subtitles for your own videos: ask for `srt` or `vtt`.
- Text for RAG, summarising or topic search: the dataset is plain JSON. Pair it with [Document to Markdown](https://apify.com/tinlark/document-to-markdown) by Tinlark to get documents and audio into the same text form.

### How to use it

1. Paste direct links to audio or video files under **Audio or video file links**, and/or podcast RSS links under **Podcast RSS feeds**.
2. Pick the model (`base` or `small`) and the output formats. Set **Max total minutes** to the most audio you want to pay for in this run.
3. Start the run. Rows appear in the dataset as files finish. For a podcast, add a schedule and switch on **New episodes only**.

The prefilled input transcribes a 2-minute public-domain recording and finishes in under a minute.

### Input

| Field | What it does | Default |
|---|---|---|
| Audio or video file links (`mediaUrls`) | Direct http(s) links to mp3, m4a, wav, ogg, flac, mp4, webm and other formats FFmpeg reads | prefilled example |
| Podcast RSS feeds (`podcastFeedUrls`) | RSS or Atom feed links; the feed host's robots.txt is checked first | none |
| Model (`model`) | `base` or `small` | `base` |
| Spoken language (`language`) | `auto`, or a code such as `en`, `de`, `fr`, `es`, `ja` | `auto` |
| Output formats (`outputFormats`) | any of `text`, `segments`, `srt`, `vtt` | `text`, `segments` |
| Episodes per feed (`episodesPerFeed`) | The latest N episodes of each feed, 1 to 50 | 3 |
| New episodes only (`newEpisodesOnly`) | Skip episodes an earlier run with the same state name already transcribed | off |
| State name (`stateKey`) | Name of the memory used by New episodes only | `default` |
| Max total minutes (`maxTotalMinutes`) | Hard cap on audio minutes in this run | 120 |
| Max minutes per file (`maxMinutesPerFile`) | Longer files are cut here, with a warning | 180 |
| Translate to English (`translateToEnglish`) | Transcript and segments in English whatever is spoken | off |

Example input:

```json
{
    "podcastFeedUrls": ["https://librivox.org/rss/1078"],
    "mediaUrls": ["https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg"],
    "episodesPerFeed": 2,
    "model": "small",
    "language": "auto",
    "outputFormats": ["text", "segments", "srt"],
    "maxTotalMinutes": 60
}
```

### Output

One row per transcribed file or episode. A shortened row from the prefilled run (the segment list and text are cut here):

```json
{
    "recordType": "transcript",
    "status": "ok",
    "sourceUrl": "https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg",
    "feedUrl": null,
    "episodeTitle": null,
    "durationSec": 133.2,
    "language": "en",
    "languageProbability": 0.993,
    "model": "base",
    "text": "This is a Libravox Recording. All Libravox recordings are in the public domain. For more information or to volunteer, please visit Libravox.org. ...",
    "segments": [
        {"start": 1.62, "end": 7.46, "text": "This is a Libravox Recording. All Libravox recordings are in the public domain."},
        {"start": 7.46, "end": 12.46, "text": "For more information or to volunteer, please visit Libravox.org."}
    ],
    "srt": null,
    "vtt": null,
    "wordCount": 302,
    "billedMinutes": 3,
    "processingSec": 18.3,
    "warnings": [],
    "error": null
}
```

In this run the recording's spoken name came out as "Libravox" (it is LibriVox), and a web address later in the text as "americanafonic": proper names and unusual words can be misspelled, more so with the smaller model. Read the output of anything where exact names matter.

Feed episodes also fill `feedUrl`, `episodeTitle`, `episodeGuid` and `publishedAt`. An error row has `status: "error"`, an `errorCode` and an `error` text. Codes: `unsupported-source` (YouTube and social links), `not-found`, `access-denied`, `too-large`, `not-media` (the link is a web page or a feed), `unreadable-media`, `no-audio-track`, `robots-disallowed`, `robots-unavailable`, `not-a-feed`, `budget-reached`, `timeout`, `download-failed`, `blocked-address` and others. Rows are written as files finish; `sourceIndex` is a running number. A summary with counts, billed minutes and per-feed numbers is stored under the `SUMMARY` key.

### Podcast feeds and new episodes only

Each feed is read newest episode first, and the first **Episodes per feed** episodes with an audio or video enclosure are used. With **New episodes only** on:

- The first run for a state name transcribes the latest episodes and remembers them.
- Later runs with the same state name transcribe only episodes not remembered yet. If there is nothing new, the run ends with the message "No new episodes since the last run" and no rows.
- An episode is remembered only after it was transcribed. Failed or budget-skipped episodes are tried again next time.
- The memory is a record in a key-value store named `audio-podcast-transcriber-state` in your Apify account. Use different state names for independent schedules.
- Only the latest **Episodes per feed** episodes are looked at. If a feed publishes more than that between two runs, raise the number.

### Cost control

- **Max total minutes** is a hard cap. A file that would pass it is cut at the cap (the row says so), and the files after it get a free `budget-reached` error row.
- Apify's **Maximum cost per run** also applies once pricing is on: the Actor reads the remaining budget before each file and stops cleanly.
- A started minute counts as a minute. Each row shows `billedMinutes`.

### Speed and platform cost

Measured on Apify at 4096 MB (the default), `int8` on CPU, with both models inside the image so runs do not wait for a download:

| Model | Speed on a 13 min 47 s MP3 | Apify platform cost per audio hour |
|---|---|---|
| base | 108 s total run; transcription about 9 times faster than real time | about $0.10 |
| small | 313 s total run; transcription about 2.7 times faster than real time | about $0.30 |

Start-up takes about 15 to 20 seconds per run. Speed depends on the audio and on how busy Apify's hardware is: a 133-second clip took 12 to 18 s with base and about 32 to 39 s with small in Tinlark's test runs. The default memory is 4096 MB. At 2048 MB the same 13-minute file took 251 s instead of 111 s at about the same cost, so 4096 MB is the better choice. Decoding a 3-hour file needs about 1.1 GB of memory on top of the model: for files longer than the default 180 minutes, raise the memory (8192 MB is the maximum). The Actor uses two CPU threads.

### Pricing

**Free during launch (until 31 October 2026).** You pay only Apify's own platform usage for your runs, roughly the figures above.

From 1 November 2026: pay per event, **per started audio minute**. Prices fall with your Apify plan:

| Event | Free plan | Bronze | Silver | Gold |
|---|---|---|---|---|
| Audio minute, base model | $0.020 | $0.012 | $0.010 | $0.008 |
| Audio minute, small model | $0.030 | $0.020 | $0.017 | $0.014 |

So an hour of audio costs $0.72 with base on the Bronze plan, and $1.20 with small. Failed files are never charged. The Store page shows the live prices once they are on.

### Limits

- Direct links to media files and podcast feeds only. **YouTube, Instagram, TikTok, Facebook and X links are not supported** and give an `unsupported-source` error row, also when a link redirects there.
- Files up to 1 GB, at most 600 minutes per file (180 by default). Longer audio is cut with a warning.
- One download per host at a time. Feed hosts' robots.txt is checked before a feed is fetched; media files are fetched exactly as linked.
- Files must be public: no login, no cookies, no signed-in pages. Private network addresses are refused.
- Speech is transcribed as spoken: no speaker names, no punctuation guarantees, no word timings. The model is open-source Whisper; Tinlark has not measured word error rates, so none are claimed. Noisy, accented or overlapping speech is harder, and `small` usually does better there than `base`.
- Music-only or silent files return an empty transcript and a warning ("No speech was detected").
- Long recordings can contain repeated or invented passages where the speech is unclear: check the output when it matters.

### Use with AI agents (MCP)

Apify's MCP server can expose this Actor as a tool, so an AI agent can transcribe a link and read the dataset. Use it when an agent needs the words of an audio or video file at a URL. Through the API:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("tinlark/audio-podcast-transcriber").call(run_input={
    "mediaUrls": ["https://example.com/interview.mp3"],
    "model": "base",
    "maxTotalMinutes": 30,
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(row["language"], row["text"][:200])
```

### FAQ

**Which formats work?** Anything FFmpeg can decode: mp3, m4a, wav, ogg, flac, mp4, webm and more. The Actor reads the audio track; video files work the same way.

**base or small?** Start with `base`. Try `small` when names, accents or noise matter. It is about three times slower and costs 1.5 to 1.75 times more per minute.

**Why set the language?** `auto` listens to the first seconds. A recording that starts with music can be detected wrongly. If you know the language, set it.

**Does it keep my audio?** Files are downloaded into the run, transcribed inside the run and deleted when it ends. Nothing is sent to a third-party AI service. The transcript is in your dataset.

**Can it transcribe a YouTube video?** No. Use a direct file link or a podcast feed.

**How do I get SRT or VTT?** Add `srt` or `vtt` to **Output formats**. The subtitle file is a text field in the row; save it with the `.srt` or `.vtt` extension.

**Disclaimers and legality.** Process only recordings you have the right to use. You are responsible for lawful use of the audio and of the output, including copyright, consent and the personal data that speech can contain. The Actor fetches exactly the links you give it, and the episodes in feeds you give it, with a User-Agent that names it and gives a contact address. It does not search, crawl or log in anywhere, and it refuses YouTube and social-media links. Do not use it to get around access controls or the terms of a site. Transcription is automatic and contains errors: do not rely on it alone for legal, medical or financial decisions.

### Support

Something wrong or missing? Open an issue on this Actor's Issues tab with your input (the link or feed, the model) and what you expected. Related: [Document to Markdown](https://apify.com/tinlark/document-to-markdown) by Tinlark turns PDF, Word, Excel and scans into Markdown for the same kind of pipeline.

# Actor input Schema

## `mediaUrls` (type: `array`):

Direct links (http or https) to audio or video files: mp3, m4a, wav, ogg, flac, mp4, webm and other formats FFmpeg can read. The files must be public and you must have the right to process them. Links to YouTube, Instagram, TikTok, Facebook and X are not supported.

## `podcastFeedUrls` (type: `array`):

Links to podcast RSS (or Atom) feeds. The Actor transcribes the latest episodes of each feed (see 'Episodes per feed') and adds the episode title, guid and publication date to every row. The feed host's robots.txt is checked first.

## `model` (type: `string`):

base is fast and cheapest. small is slower and usually more accurate, especially with accents, noise and non-English speech. Both are open-source Whisper models running on Apify, no third-party API.

## `language` (type: `string`):

auto detects the language from the first seconds of audio. If you know it, enter the two-letter code (en, de, fr, es, it, pt, nl, ja, zh ...): it is faster and avoids wrong detection when a recording starts with music.

## `outputFormats` (type: `array`):

Fields to fill in each row. text: the full transcript. segments: a list of {start, end, text} with times in seconds. srt, vtt: ready subtitle files as text.

## `episodesPerFeed` (type: `integer`):

How many of the latest episodes of each feed to look at (newest first). With 'New episodes only', set it at least as high as the number of episodes the feed publishes between two runs.

## `newEpisodesOnly` (type: `boolean`):

For scheduled runs: transcribe only episodes that an earlier run with the same state name has not transcribed yet. The first run transcribes the latest episodes and remembers them. Episodes that failed are tried again next time. Applies to podcast feeds only.

## `stateKey` (type: `string`):

Name under which 'New episodes only' remembers finished episodes, in a key-value store of your account. Use different names for independent schedules. Letters, digits, dot, dash, underscore.

## `maxTotalMinutes` (type: `integer`):

Hard cap on audio minutes transcribed in this run. When it is reached, the current file is cut at the cap and remaining files get a free error row (budget-reached). A started minute counts as a minute. Set the run's 'Maximum cost per run' as well if you want a cap in dollars.

## `maxMinutesPerFile` (type: `integer`):

Longer files are cut after this many minutes and the row gets a warning. Files are also limited to 1 GB.

## `translateToEnglish` (type: `boolean`):

Use Whisper's built-in translation task: the transcript and segments are in English whatever language is spoken. Quality is lower than for transcription in the original language.

## Actor input object example

```json
{
  "mediaUrls": [
    "https://example.com/interview.mp3",
    "https://example.com/lecture.mp4"
  ],
  "podcastFeedUrls": [
    "https://feeds.example.com/my-podcast.xml"
  ],
  "model": "base",
  "language": "auto",
  "outputFormats": [
    "text",
    "segments"
  ],
  "episodesPerFeed": 3,
  "newEpisodesOnly": false,
  "stateKey": "default",
  "maxTotalMinutes": 120,
  "maxMinutesPerFile": 180,
  "translateToEnglish": false
}
```

# Actor output Schema

## `transcripts` (type: `string`):

Overview of the transcribed files and episodes.

## `content` (type: `string`):

Transcript text, SRT and VTT per file.

## `errors` (type: `string`):

Files, feeds and episodes that could not be transcribed, with the reason.

## `summary` (type: `string`):

Counts, billed minutes, per-feed numbers and the list of errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mediaUrls": [
        "https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg"
    ],
    "language": "auto",
    "outputFormats": [
        "text",
        "segments"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tinlark/audio-podcast-transcriber").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mediaUrls": ["https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg"],
    "language": "auto",
    "outputFormats": [
        "text",
        "segments",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("tinlark/audio-podcast-transcriber").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mediaUrls": [
    "https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg"
  ],
  "language": "auto",
  "outputFormats": [
    "text",
    "segments"
  ]
}' |
apify call tinlark/audio-podcast-transcriber --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tinlark/audio-podcast-transcriber"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/H6UayWI7BE00XIrp9/builds/5chnhVVmpG9QPAFBR/openapi.json
