# Audio & Video Transcriber (Whisper) with Podcast RSS (`rod_analytics/media-transcriber`) Actor

Transcribe audio, video and podcast RSS episodes with Whisper on Groq or OpenAI using your own API key. Get text, SRT and VTT subtitles and timestamped JSON segments.

- **URL**: https://apify.com/rod\_analytics/media-transcriber.md
- **Developed by:** [Rod Services](https://apify.com/rod_analytics) (community)
- **Categories:** AI, Videos
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Audio & Video Transcriber (Whisper) do?

**Audio & Video Transcriber** turns **podcasts, meeting recordings, interviews, lectures, webinars and videos** into text with **OpenAI Whisper**. You get a clean transcript, **SRT and WebVTT subtitles**, and **JSON segments with timestamps** for every file.

Paste direct file links, upload a file, or give it a **podcast RSS feed** and it transcribes the latest episodes. Long recordings are split into chunks automatically, so a 3 hour podcast works the same as a 3 minute voice memo.

It runs on **your own Groq or OpenAI API key** (bring your own key). Groq runs **Whisper large v3 turbo** at about **$0.04 per audio hour**, so a 1 hour podcast costs you about 4 cents at Groq plus **$0.003 per minute** here.

As an Apify Actor you also get an API, scheduling, webhooks, integrations with Make, Zapier, n8n and LangChain, and run monitoring. AI agents can call it through the Apify MCP server.

### Why use this Whisper transcription tool?

- **Podcast transcription from RSS.** Point it at any podcast feed. It picks the newest N episodes and adds episode title and publish date.
- **Subtitles in SRT and VTT.** Ready for YouTube Studio uploads, video editors, HTML5 players and accessibility.
- **Meeting recordings and interviews.** MP4, MOV, MKV and WEBM files work. Only the audio track is used.
- **Timestamps for search and RAG.** Every segment has start and end seconds. Link quotes back to the exact moment.
- **AI agents and LLM pipelines.** Feed transcripts into summarizers, show notes generators, chatbots or vector databases.
- **Long files handled for you.** ffmpeg converts audio to mono 16 kHz and splits it into chunks under the 25 MB provider limit, with overlap so no words are lost at the cuts.
- **Translate to English.** One switch gives an English transcript of speech in about 100 languages.
- **Cheap and fast.** You pay the provider directly at their rates. There is no markup on the model.

### How to transcribe audio or a podcast

1. Get an API key. Groq keys are free to create at [console.groq.com/keys](https://console.groq.com/keys). OpenAI keys are at [platform.openai.com/api-keys](https://platform.openai.com/api-keys).
2. Open the **Input** tab and paste the key into **API key**. It is stored encrypted.
3. Add links in **Audio or video URLs**, upload a file, or paste a **Podcast RSS feed URL**.
4. Optional: set a **Language hint** such as `en`, `de` or `lt`, and pick **Output formats**.
5. Click **Start**. Transcripts appear in the **Output** tab. Subtitle files are in the key-value store.
6. Download the dataset as JSON, CSV or Excel, or call the API from your own code.

Want to check your links first? Turn on **Dry run**. It downloads, converts and splits the audio without calling the provider and without the per minute charge.

#### Running without an API key

If you start the Actor without a key, it does not download anything. It finishes successfully and writes one dataset item that explains a key is required. This is also what happens with the example input, so you can try the Actor safely.

### Input

All fields are on the **Input** tab. The main ones:

| Field | What it does |
| --- | --- |
| `audioUrls` | Direct links to MP3, M4A, WAV, FLAC, OGG, OPUS, AAC, MP4, MOV, MKV, WEBM and more. Google Drive and Dropbox share links work. |
| `rssFeedUrl`, `maxEpisodes` | Podcast RSS or Atom feed, and how many of the newest episodes to transcribe. |
| `uploadedFile`, `keyValueStoreRecords` | Files uploaded in Console, or records in a key-value store. |
| `provider`, `apiKey`, `model` | `groq` or `openai`, your key, and the model. `auto` picks the best fit. |
| `language` | ISO 639-1 language hint. Empty means auto detect. |
| `translateToEnglish` | Return an English translation. |
| `outputFormats` | `text`, `srt`, `vtt`, `json` files saved to the key-value store. |
| `maxDurationMinutes` | Transcribe at most this many minutes per file. Protects you from surprise costs. |
| `dryRun` | Download and split only. No key needed. |
| `prompt` | Names, brands and jargon that help the model spell them right. |

Example input:

```json
{
    "provider": "groq",
    "apiKey": "YOUR_GROQ_KEY",
    "rssFeedUrl": "https://librivox.org/rss/389",
    "maxEpisodes": 3,
    "language": "en",
    "outputFormats": ["text", "srt", "vtt", "json"]
}
```

#### Models

| Provider | Model | Timestamps | Translation | Provider price |
| --- | --- | --- | --- | --- |
| Groq | `whisper-large-v3-turbo` (default) | Yes | No, switches to v3 | ~$0.04 per hour |
| Groq | `whisper-large-v3` | Yes | Yes | ~$0.111 per hour |
| OpenAI | `whisper-1` | Yes | Yes | $0.006 per minute |
| OpenAI | `gpt-4o-mini-transcribe` | No, estimated | No | $0.003 per minute |
| OpenAI | `gpt-4o-transcribe` | No, estimated | No | $0.006 per minute |

With `auto` on OpenAI, the Actor uses `whisper-1` when you ask for SRT, VTT or JSON, because gpt-4o transcribe models return no timestamps. Check current prices on the provider sites.

### Output

One dataset item per file. The example below is shortened. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

```json
{
    "sourceUrl": "https://www.archive.org/download/gettysburg_shurtagal_librivox/Gettysburg_Address_Lincoln_64kb.mp3",
    "sourceType": "rss",
    "title": "Gettysburg Address",
    "feedTitle": "Gettysburg Address, The by Abraham Lincoln (1809 - 1865)",
    "publishedAt": null,
    "durationSeconds": 100.34,
    "language": "en",
    "text": "Four score and seven years ago our fathers brought forth on this continent a new nation...",
    "segments": [
        { "id": 0, "start": 0.0, "end": 6.2, "text": "Four score and seven years ago" },
        { "id": 1, "start": 6.2, "end": 11.8, "text": "our fathers brought forth on this continent a new nation," }
    ],
    "wordCount": 272,
    "srtUrl": "https://api.apify.com/v2/key-value-stores/.../records/srt-0000-Gettysburg-Address.srt",
    "vttUrl": "https://api.apify.com/v2/key-value-stores/.../records/vtt-0000-Gettysburg-Address.vtt",
    "provider": "groq",
    "model": "whisper-large-v3-turbo",
    "billedMinutes": 2,
    "warnings": [],
    "error": null
}
```

#### Data fields

| Field | Description |
| --- | --- |
| `sourceUrl` | Media URL or RSS enclosure URL. |
| `title`, `feedTitle`, `publishedAt` | Episode title and date from RSS, or the media title tag, or the file name. |
| `durationSeconds`, `transcribedSeconds` | Length of the file and of the part that was transcribed. |
| `language` | ISO 639-1 code, detected or from your hint. |
| `text` | Full transcript. |
| `segments` | `{id, start, end, text}` with times in seconds. |
| `srtUrl`, `vttUrl`, `txtUrl`, `jsonUrl` | Files in the key-value store. |
| `wordCount` | Words in the transcript. |
| `provider`, `model` | What transcribed the file. |
| `timestampsApproximate` | `true` when the model gives no timestamps and times are estimated. |
| `billedMinutes` | Audio minutes charged for this file. |
| `warnings`, `error` | What went wrong or was changed, in plain words. |

The **Overview** view shows one row per file. The **Segments** view shows one row per timestamped segment.

### How much does it cost to transcribe audio?

This Actor uses **pay per event** pricing:

- **$0.003 per audio minute** transcribed, rounded up per file.
- A small **start fee** per run.
- **Dry runs, rejected links and failed files are free** of the per minute charge.

The provider bills its own part to your key. Examples with Groq whisper-large-v3-turbo:

| Audio | This Actor | Groq (approx.) | Total |
| --- | --- | --- | --- |
| 1 hour podcast | $0.18 | $0.04 | about $0.22 |
| 10 episodes of 45 min | $0.90 | $0.30 | about $1.20 |
| 100 hours of meetings | $12.00 | $4.00 | about $16 |

Set **Maximum cost per run** in the run options. The Actor stops taking new minutes when it is reached, and cuts a long file short with a warning.

### Tips and advanced options

- **Give a language hint.** It avoids wrong language detection on short clips and music intros.
- **Use the vocabulary prompt** for names, product names and acronyms.
- **Rate limits.** Groq free keys have hourly and daily audio limits. Check them on your Groq limits page. The Actor retries 429 responses with backoff and honours `retry-after`. Lower **Parallel API requests** if you see many retries, or use a paid Groq tier.
- **Exact subtitles.** Use Groq models or OpenAI `whisper-1`. gpt-4o transcribe models return text only, so cue times are spread by text length.
- **Chunk length.** 10 minutes is a good default. Chunks overlap by 5 seconds and the transcript is stitched at the middle of the overlap.
- **Memory.** 1 GB is enough. Audio is processed on disk, not in memory.

### FAQ, disclaimers and support

#### Can I transcribe YouTube, TikTok, Instagram, Facebook or Vimeo links?

No. Their terms of service forbid downloading media with third-party tools, so these page links are rejected with a message. Use a direct link to a file you own or may use, upload the file, or use the podcast RSS feed.

#### Is my API key safe?

The key is a secret input. Apify stores it encrypted. The Actor sends it only to the provider API you picked, never logs it, and never writes it to the dataset. We recommend a separate key for this Actor with a spending limit, and revoking it when you are done. You are responsible for using your key within your provider's terms.

#### Do you store my audio?

Audio is downloaded into the run container, converted, sent to the provider and deleted when the file is done. Transcripts are saved in your own run storage. The provider processes the audio under its own data policy.

#### Known limitations

- Speaker labels (diarization) are not included yet.
- Word counts for Chinese and Japanese count characters.
- Files behind logins or DRM cannot be downloaded.

#### Feedback

Found a bug or need a feature? Open an issue on the **Issues** tab. Custom pipelines such as speaker labels, summaries or delivery to your storage are available on request.

# Actor input Schema

## `audioUrls` (type: `array`):

Direct links to media files: MP3, M4A, WAV, FLAC, OGG, OPUS, AAC, MP4, MOV, MKV, WEBM and more. Google Drive and Dropbox share links are converted to direct downloads. YouTube, TikTok, Instagram, Facebook and Vimeo page links are rejected because their terms forbid downloading.

## `rssFeedUrl` (type: `string`):

A podcast RSS or Atom feed. The Actor takes the audio enclosures of the latest episodes and adds the episode title and publish date to the output.

## `maxEpisodes` (type: `integer`):

How many of the newest feed episodes to transcribe.

## `uploadedFile` (type: `string`):

Upload one audio or video file from your computer. Apify stores it in a key-value store and passes its URL to the Actor.

## `keyValueStoreRecords` (type: `array`):

Keys of media records in this run's default key-value store, for example when another Actor calls this one. Use `storeId/key` for other stores. For your own stores the record URL with its signature in Audio or video URLs is more reliable.

## `provider` (type: `string`):

Speech to text provider for your own API key. Groq runs Whisper large v3 turbo very fast and cheap. OpenAI offers gpt-4o-mini-transcribe and whisper-1.

## `apiKey` (type: `string`):

Your Groq key (console.groq.com/keys) or OpenAI key (platform.openai.com/api-keys). It is stored encrypted and used only to call the provider from your run. The provider bills transcription to your account. Tip: create a separate key for this Actor with a spending limit, and revoke it when you are done.

## `model` (type: `string`):

`auto` picks whisper-large-v3-turbo on Groq. On OpenAI it picks whisper-1 when you ask for SRT, VTT or JSON timestamps, otherwise gpt-4o-mini-transcribe. Translation uses whisper-large-v3 on Groq and whisper-1 on OpenAI. gpt-4o models return no timestamps, so their subtitle times are estimated.

## `language` (type: `string`):

ISO 639-1 code of the spoken language, for example `en`, `de`, `es`, `lt`. Leave empty to detect it automatically. A hint improves accuracy and speed.

## `translateToEnglish` (type: `boolean`):

Return an English translation instead of the original language transcript.

## `outputFormats` (type: `array`):

Files saved to the key-value store for each transcript. The dataset always has plain text and timestamped segments.

## `prompt` (type: `string`):

Optional context for the model: names, brands, jargon or spelling style. Whisper uses about the first 224 tokens.

## `maxDurationMinutes` (type: `integer`):

Transcribe at most this many minutes of each file. Longer files are cut and a warning is added. Protects you from surprise costs.

## `dryRun` (type: `boolean`):

Download, convert and split the audio into chunks, then stop. No API key needed, no provider calls, no per minute charge. Use it to check that your files and feed work.

## `chunkMinutes` (type: `integer`):

Long audio is split into chunks of this length with a 5 second overlap. Chunks are mono 16 kHz and stay under 20 MB, below the 25 MB provider limit.

## `concurrency` (type: `integer`):

How many chunks are sent to the provider at the same time. Lower it if your key hits rate limits. 429 responses are retried with backoff.

## `maxFileSizeMb` (type: `integer`):

Skip source files larger than this. Video files can be large; only the audio track is used.

## Actor input object example

```json
{
  "audioUrls": [
    "https://www.archive.org/download/gettysburg_shurtagal_librivox/Gettysburg_Address_Lincoln_64kb.mp3"
  ],
  "rssFeedUrl": "https://librivox.org/rss/389",
  "maxEpisodes": 3,
  "keyValueStoreRecords": [],
  "provider": "groq",
  "model": "auto",
  "translateToEnglish": false,
  "outputFormats": [
    "text",
    "srt",
    "vtt",
    "json"
  ],
  "maxDurationMinutes": 180,
  "dryRun": false,
  "chunkMinutes": 10,
  "concurrency": 3,
  "maxFileSizeMb": 1000
}
```

# Actor output Schema

## `transcripts` (type: `string`):

No description

## `overview` (type: `string`):

No description

## `files` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "audioUrls": [
        "https://www.archive.org/download/gettysburg_shurtagal_librivox/Gettysburg_Address_Lincoln_64kb.mp3"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("rod_analytics/media-transcriber").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "audioUrls": ["https://www.archive.org/download/gettysburg_shurtagal_librivox/Gettysburg_Address_Lincoln_64kb.mp3"] }

# Run the Actor and wait for it to finish
run = client.actor("rod_analytics/media-transcriber").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "audioUrls": [
    "https://www.archive.org/download/gettysburg_shurtagal_librivox/Gettysburg_Address_Lincoln_64kb.mp3"
  ]
}' |
apify call rod_analytics/media-transcriber --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,rod_analytics/media-transcriber"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/c6TBESlhFCg25j821/builds/E8HSSUsLufZuBY0FY/openapi.json
