# Whisper Transcriber (`adam-frank/whisper-transcriber`) Actor

Transcribes audio and video from direct media URLs using faster-whisper large-v3 on a GPU, with optional speaker diarization. Returns plain text, SRT, VTT and speaker-labelled segments, priced per transcribed minute.

- **URL**: https://apify.com/adam-frank/whisper-transcriber.md
- **Developed by:** [Adam Schepis](https://apify.com/adam-frank) (community)
- **Categories:** AI, Videos
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$20.00 / 1,000 transcribed minutes

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Whisper Transcriber

Transcribes audio and video from **direct media URLs** using **faster-whisper large-v3 on a GPU**, with optional speaker diarization. Returns plain text, SRT, VTT and speaker-labelled segments. Priced per transcribed minute, with a 3-minute minimum per run.

### Direct media links only — no YouTube

This Actor accepts a **direct link to an audio or video file**: a podcast enclosure MP3, or a hosted MP4, WAV, M4A, FLAC or OGG.

It does **not** download from YouTube, Vimeo, SoundCloud, Spotify, TikTok, Instagram, Apple Podcasts or any other streaming site, and it never will. Pulling media off those platforms breaks their terms of service. Passing one of those links returns a clear error instead of a transcript, and costs you nothing.

If you have a YouTube video you own, export the audio or video file, host it somewhere reachable, and pass that URL.

### What you get

- **large-v3, not tiny.** Most transcription Actors run a small Whisper model on CPU. This one runs the full `faster-whisper large-v3` model on a GPU, so accents, crosstalk and technical vocabulary survive.
- **Speaker diarization.** Turn it on and every segment is labelled `SPEAKER_00`, `SPEAKER_01`, … — usable for interviews, panels and meetings.
- **Four output shapes from one run:** plain text, SRT subtitles, WebVTT captions, and timed segments with speaker labels.
- **Pay per transcribed minute**, rounded up, with a 3-minute minimum per run. A file that fails to download or fails to transcribe is not charged.

### Input

| Field | Type | Required | Description |
|---|---|---|---|
| `mediaUrls` | array of URLs | yes | Direct links to audio/video files. Streaming-site links are rejected. |
| `language` | string | no (default `auto`) | ISO 639-1 code, or `auto` to detect. Naming the language is faster and more accurate on short or noisy clips. |
| `diarization` | boolean | no (default `false`) | Label who is speaking. Adds word-level alignment, so runs take noticeably longer. |
| `maxAudioMinutes` | integer | no (default 120) | Per-file safety cap. A file needing longer is abandoned and not charged. |
| `maxFileSizeMb` | integer | no (default 500) | Files above this are skipped before any GPU time is spent. |
| `maxCostUsd` | number | no (default 10) | Hard ceiling on what the run may charge. |

Example input:

```json
{
    "mediaUrls": [
        { "url": "https://cdn.example.com/podcast/episode-42.mp3" }
    ],
    "language": "en",
    "diarization": true,
    "maxCostUsd": 5
}
```

### Output example

One record per media URL:

```json
{
    "url": "https://cdn.example.com/podcast/episode-42.mp3",
    "status": "ok",
    "model": "faster-whisper large-v3 (WhisperX)",
    "language": "en",
    "languageProbability": 0.9971,
    "durationSeconds": 11.86,
    "billedMinutes": 1,
    "diarization": true,
    "speakers": ["SPEAKER_00", "SPEAKER_01"],
    "transcript": "SPEAKER_00: Welcome back to the show, today we are talking about margins.\n\nSPEAKER_01: Thanks for having me, it is a topic I think about a lot.",
    "srt": "1\n00:00:00,031 --> 00:00:03,512\n[SPEAKER_00] Welcome back to the show, today we are talking about margins.\n",
    "vtt": "WEBVTT\n\n00:00:00.031 --> 00:00:03.512\n[SPEAKER_00] Welcome back to the show, today we are talking about margins.\n",
    "segments": [
        { "start": 0.031, "end": 3.512, "text": "Welcome back to the show, today we are talking about margins.", "speaker": "SPEAKER_00" }
    ],
    "gpuSeconds": 2.41,
    "queueSeconds": 1.8,
    "error": null,
    "errorMessage": null,
    "scrapedAt": "2026-09-16T12:00:00.000Z"
}
```

`billedMinutes` is what that one file contributed; a run that transcribed anything is billed at least 3 minutes in total.

A rejected or failed file produces the same record with `status: "error"`, an `error` code (`unsupported-source`, `not-direct-media`, `file-too-large`, `invalid-url`, `unreachable`, or a backend job status) and an `errorMessage` explaining it. Download results as JSON, CSV, Excel, or via the Apify API.

### Pricing

This Actor uses [pay-per-event pricing](https://docs.apify.com/platform/actors/publishing/monetize/pay-per-event). You are charged for the `transcribed-minute` event, once per minute of transcribed audio, rounded up per file. **Runs are billed a 3-minute minimum.** The exact per-event price is shown on the Actor's Store page before you run it.

| Event | When it's charged |
|---|---|
| `transcribed-minute` | Once per minute of audio, after the transcript is saved to your dataset. A run that transcribed at least one file is billed at least 3 of them. |

What that means in practice:

- **Nothing is charged until the transcript is in your dataset.** If the download fails, the backend errors, or the job runs past `maxAudioMinutes`, you get an error record and no charge. A run where every file failed is billed nothing at all — the 3-minute minimum does not apply to it.
- **Runs are billed a 3-minute minimum.** Starting a GPU worker costs real money whatever the clip length, so a run totalling less than 3 minutes of audio is billed as 3. The minimum is per run, not per file: a run of three 1-minute clips is billed 3 minutes, the same as a single 3-minute file. Batch short clips into one run rather than running each on its own.
- **Minutes are measured to the end of the last spoken segment**, so trailing silence is not billed.
- **Each file rounds up independently.** Five 30-second clips are billed 5 minutes; one 4-minute file is billed 4.
- `maxCostUsd` and Apify's own **Maximum cost per run** both cap a run, minimum included: neither is ever exceeded to reach the 3-minute floor.

### What the buyer should know

- **Formats:** anything FFmpeg decodes — MP3, M4A, M4B, WAV, FLAC, OGG/OGA, OPUS, AAC, MP4, MOV, MKV, WEBM, AVI, 3GP, AMR. Video is accepted; only the audio track is used.
- **Max file size:** 500 MB by default, adjustable to 2 GB via `maxFileSizeMb`. Oversized files are skipped before any GPU time is spent.
- **Max length:** `maxAudioMinutes` defaults to 120 and accepts up to 300. The backend also enforces its own 15-minute GPU execution ceiling per file, which is ample for multi-hour audio at large-v3's speed but will cut off a pathological file.
- **Languages:** large-v3 handles roughly 100 languages. The `language` dropdown lists the common ones; detection is automatic by default. Transcription is in the spoken language — this Actor does not translate.
- **Cold starts:** the GPU backend scales to zero when idle, so the first file in a run can wait roughly 20–60 seconds for a worker. `queueSeconds` in each record reports the actual wait. Subsequent files in the same run reuse the warm worker.
- **Files are processed one at a time**, in the order given, so a long list takes proportionally longer.
- **The URL must be reachable by the transcription backend**, which fetches the file itself with a plain HTTP client. Hosts that block non-browser user agents (Wikimedia Commons is one) will fail even though the link looks fine in a browser. Podcast CDNs, S3/R2/GCS buckets, and GitHub raw links all work.
- **Diarization** labels speakers as `SPEAKER_00`, `SPEAKER_01`, …; it does not identify them by name.

### Why this Actor

The transcription Actors already on Store run Whisper `tiny`, `base` or `small` on CPU and still charge $0.015–$0.048 per minute. This one runs the full `large-v3` model on a GPU and adds speaker diarization, at a price in the middle of that range. You get materially better transcripts for the same money, with no API key to manage and no GPU to rent.

### Notes

- Transcription runs on a dedicated GPU serverless endpoint (NVIDIA 24 GB class) running [WhisperX](https://github.com/m-bain/whisperX) with `faster-whisper large-v3`. Word-level alignment is enabled only when diarization is on, since that is the only thing that needs it.
- Built with the [Apify SDK for JavaScript](https://docs.apify.com/sdk/js/).

### Reference docs used to build this Actor

- Apify SDK for JS: https://docs.apify.com/sdk/js/
- Pay-per-event monetization overview: https://docs.apify.com/platform/actors/publishing/monetize/pay-per-event
- Pay-per-event SDK guide (`Actor.charge`, `ChargingManager`): https://docs.apify.com/sdk/js/docs/concepts/pay-per-event
- `.actor/actor.json` reference: https://docs.apify.com/platform/actors/development/actor-definition/actor-json
- Runpod serverless job operations: https://docs.runpod.io/serverless/endpoints/job-operations
- WhisperX worker image: https://github.com/kodxana/whisperx-worker\_v2

# Actor input Schema

## `mediaUrls` (type: `array`):

Direct links to audio or video FILES — a podcast enclosure MP3, or a hosted MP4/WAV/M4A/FLAC/OGG. YouTube, Vimeo, SoundCloud, Spotify and other streaming sites are NOT supported and are rejected with an explanatory message: downloading from them breaks their terms of service. A link that returns a web page rather than a media file is rejected too.

## `language` (type: `string`):

Language spoken in the media, as an ISO 639-1 code. Leave on "Detect automatically" unless detection is getting it wrong — naming the language is faster and more accurate on short or noisy clips.

## `diarization` (type: `boolean`):

Label each segment with who is speaking (SPEAKER\_00, SPEAKER\_01, …). Useful for interviews, panels and meetings. It adds word-level alignment and a speaker-clustering pass, so runs take noticeably longer; leave it off for single-speaker audio.

## `maxAudioMinutes` (type: `integer`):

Safety cap on how long a single file may take to transcribe. A file that needs longer is abandoned and NOT charged. Raise it for long-form audio such as full-length podcasts or lectures.

## `maxFileSizeMb` (type: `integer`):

Files larger than this are skipped before any GPU time is spent, so you are never charged for them.

## `maxCostUsd` (type: `number`):

Hard ceiling on what this run may charge. The run stops once transcribed minutes reach it, so a long list of files cannot produce a surprise bill. Runs are billed a 3-minute minimum, so the smallest useful ceiling is the price of 3 minutes. Apify's own "Maximum cost per run" setting applies on top of this.

## Actor input object example

```json
{
  "mediaUrls": [
    {
      "url": "https://archive.org/download/spc277_2607_librivox/spc277_april_pac_64kb.mp3"
    },
    {
      "url": "https://raw.githubusercontent.com/runpod-workers/sample-inputs/main/audio/gettysburg.wav"
    }
  ],
  "language": "auto",
  "diarization": false,
  "maxAudioMinutes": 120,
  "maxFileSizeMb": 500,
  "maxCostUsd": 10
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset containing one transcript record per media file

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mediaUrls": [
        {
            "url": "https://archive.org/download/spc277_2607_librivox/spc277_april_pac_64kb.mp3"
        },
        {
            "url": "https://raw.githubusercontent.com/runpod-workers/sample-inputs/main/audio/gettysburg.wav"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("adam-frank/whisper-transcriber").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "mediaUrls": [
        { "url": "https://archive.org/download/spc277_2607_librivox/spc277_april_pac_64kb.mp3" },
        { "url": "https://raw.githubusercontent.com/runpod-workers/sample-inputs/main/audio/gettysburg.wav" },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("adam-frank/whisper-transcriber").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mediaUrls": [
    {
      "url": "https://archive.org/download/spc277_2607_librivox/spc277_april_pac_64kb.mp3"
    },
    {
      "url": "https://raw.githubusercontent.com/runpod-workers/sample-inputs/main/audio/gettysburg.wav"
    }
  ]
}' |
apify call adam-frank/whisper-transcriber --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,adam-frank/whisper-transcriber"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/5rzEBDJWBMU0tgmvT/builds/IitsPa7InfSznzX7A/openapi.json
