# Audio & Video to Text - Whisper Transcription API (`zenomastro/audio-video-to-text`) Actor

Transcribe audio and video files from direct URLs (mp3, wav, m4a, ogg, mp4, mov, webm) with Whisper large-v3-turbo. Get text, SRT, VTT, timestamped segments and RAG chunks. $0.004 per audio minute, platform usage included.

- **URL**: https://apify.com/zenomastro/audio-video-to-text.md
- **Developed by:** [Rosario Vitale](https://apify.com/zenomastro) (community)
- **Categories:** AI, Videos, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.20 / 1,000 audio minute transcribeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Audio & Video to Text do?

Audio & Video to Text turns recordings into text. Give it direct links to audio or video files, and it returns one row per file with the transcript, timestamped segments, SRT and WebVTT subtitles and RAG-ready chunks. It runs the Whisper large-v3-turbo speech recognition model, detects the spoken language automatically and can also translate speech to English.

No API key and no model hosting are needed. The Actor downloads each file, extracts the audio track with ffmpeg, converts it to compact 16 kHz mono mp3, cuts long recordings at natural pauses into parts of about 10 minutes, transcribes the parts in parallel and merges everything back into one transcript with correct timestamps. Supported formats include mp3, wav, m4a, ogg, opus, flac, webm, mp4, mov and mkv, up to 2 GB and 180 minutes per file.

### Who is it for?

- **Podcasters and media teams** who need transcripts, show notes and subtitle files for episodes and interviews.
- **Developers and data teams** who feed call recordings, meetings or lectures into search, analytics or LLM pipelines.
- **AI and RAG builders** who want timestamped chunks ready for embeddings and a vector database.
- **Researchers, journalists and educators** who transcribe interviews, archives and recorded courses in bulk.
- **Video editors and creators** who need SRT or WebVTT subtitle files from their own media files.

### Output fields

| Field | Description |
|---|---|
| `url` | The media link that was processed. |
| `status` | `COMPLETE`, `PARTIAL`, `VALID_EMPTY`, `INVALID_INPUT`, `UPSTREAM_FAILED` or `TOO_LONG` (see below). |
| `durationSeconds` | Length of the audio in seconds. |
| `billedMinutes` | Started audio minutes that were charged for this file. |
| `language` | Detected (or selected) spoken language code, for example `en`. |
| `text` | Full transcript as plain text. |
| `segments` | List of `{start, end, text}` with times in seconds from the start of the file. |
| `srt`, `vtt` | Subtitles as text, when the format is selected. |
| `srtUrl`, `vttUrl` | Direct links to the `.srt` and `.vtt` files saved in the key-value store of the run. |
| `chunks` | RAG chunks: `{index, text, startSeconds, endSeconds, charCount}` limited to the chosen chunk size. |
| `wordCount`, `task`, `model`, `partsTotal`, `partsTranscribed`, `error`, `processedAt` | Details and diagnostics. |

Only the formats you select in **Output formats** are filled, the other fields stay `null`.

**Row status values:**

- `COMPLETE`: the whole file was transcribed. Charged per started audio minute.
- `PARTIAL`: some parts of the file were transcribed, others failed or the run's maximum charge was reached. Charged only for the minutes that were transcribed; `error` lists what is missing.
- `VALID_EMPTY`: the file was processed but no speech was found. Free.
- `INVALID_INPUT`: not a public direct media link, a web page, a file without audio, a YouTube or social-media link, or a private network address. Free.
- `TOO_LONG`: the file is longer than **Maximum duration per file** (hard cap 180 minutes) or larger than 2 GB. Free.
- `UPSTREAM_FAILED`: the file could not be downloaded (for example HTTP 404) or the transcription service was unavailable. Free.

### Pricing at a glance

| | Price per 1,000 audio minutes |
|---|---:|
| **This Actor** | **$4.00** |
| Median of 12 audio and video transcription Actors in Apify Store with a per-minute price (October 2026) | $17.50 |
| Cheapest comparable Actor found (October 2026) | $3.00 |

That is about **$0.24 per hour of audio**, roughly 75% below the median per-minute price for this task in Apify Store.

**Cost example:** a 45-minute podcast episode is 45 × $0.004 = **$0.18**. A 90-second clip counts as 2 started minutes ($0.008). Files with no speech, invalid or too long files and failed downloads cost nothing, and you pay only for the parts that were actually transcribed.

**Subscriber discounts:** on a paid Apify plan you pay less per minute: Bronze −10%, Silver −15%, Gold and higher −20% (**$3.20 per 1,000 minutes** on Gold).

**Apify platform usage (compute, storage, data transfer) is included.** You pay only the audio-minute price plus a small run start fee of $0.0001 per GB of run memory.

Use the run's *maximum charge* setting to cap spending. The Actor stops cleanly when the limit is reached and returns what was transcribed so far.

### How to use

1. Open the Actor and paste direct links to your audio or video files into **Audio or video file URLs**, one per line. The prefilled example is a short public-domain speech recording.
2. Optionally choose the **Spoken language**, switch **Task** to translate to English, and select the **Output formats** you need.
3. Click **Start**. A short clip finishes in a few seconds; a one-hour file typically takes a few minutes.
4. Open the **Output** tab to read the transcripts, or export as JSON, CSV, Excel or HTML. Download the `.srt` and `.vtt` files from the links in each row.

#### Good to know

- **Direct media links only.** The link must return the audio or video file itself (for example `https://example.com/episode.mp3`). Web pages, YouTube, TikTok, Instagram, Facebook, X, Vimeo and similar links are rejected as `INVALID_INPUT`. For YouTube videos use the [YouTube Transcript Actor](https://apify.com/zenomastro/youtube-transcript-reliable).
- **Public files only.** Private, local and reserved network addresses are blocked, and links that need a login or cookies are not supported.
- **Shared daily capacity.** Transcription runs on a hosted model with a daily capacity limit. If it is used up during your run, the remaining files come back as `UPSTREAM_FAILED` and are not charged; run them again later.
- **Privacy.** Audio is processed in temporary storage, deleted after each file, and sent in short parts to the hosted Whisper model for transcription.
- **Accuracy.** Quality depends on the recording. Choosing the spoken language, adding a context prompt with names and terms, and enabling *Skip silence* helps with noisy or music-heavy audio.

#### Input example

```json
{
  "mediaUrls": ["https://upload.wikimedia.org/wikipedia/commons/4/46/1941_Roosevelt_speech_pearlharbor_p1.ogg"],
  "language": "auto",
  "task": "transcribe",
  "outputFormats": ["text", "srt"],
  "maxDurationMinutes": 60
}
```

### Output example

```json
{
  "recordType": "transcript",
  "url": "https://upload.wikimedia.org/wikipedia/commons/4/46/1941_Roosevelt_speech_pearlharbor_p1.ogg",
  "status": "COMPLETE",
  "durationSeconds": 26.1,
  "billedMinutes": 1,
  "language": "en",
  "task": "transcribe",
  "model": "whisper-large-v3-turbo",
  "text": "Yesterday, December 7th, 1941, a state which will live in infamy. The States of America was suddenly and deliberately attacked by naval and air forces of the Empire of Japan.",
  "segments": [
    { "start": 0.24, "end": 6.56, "text": "Yesterday, December 7th, 1941," },
    { "start": 7.38, "end": 12.81, "text": "a state which will live in infamy." }
  ],
  "srt": "1\n00:00:00,240 --> 00:00:06,560\nYesterday, December 7th, 1941,\n\n2\n00:00:07,380 --> 00:00:12,810\na state which will live in infamy.\n",
  "srtUrl": "https://api.apify.com/v2/key-value-stores/STORE_ID/records/001-1941_Roosevelt_speech_pearlharbor_p1.srt",
  "wordCount": 30,
  "partsTotal": 1,
  "partsTranscribed": 1,
  "error": null
}
```

The transcript above is the model's literal output for the sample clip.

### API

One HTTP call runs the Actor and returns the results directly:

```bash
curl -X POST "https://api.apify.com/v2/acts/zenomastro~audio-video-to-text/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"mediaUrls": ["https://example.com/interview.mp3"], "outputFormats": ["text", "segments"]}'
```

For long files start the run asynchronously and fetch the dataset when it finishes. The same works from the Apify Python/JavaScript clients, Make, n8n, Zapier, LangChain and LlamaIndex integrations, or on a schedule.

### Use with AI agents (MCP)

Claude, ChatGPT, Cursor, VS Code, n8n AI agents and any MCP client can run this Actor through the official Apify MCP server. Add `https://mcp.apify.com?tools=zenomastro/audio-video-to-text` to your client, then ask for example:

> *"Transcribe this podcast episode https://example.com/episode.mp3 and give me SRT subtitles and a 5-bullet summary."*

The agent fills the input from the field descriptions and receives clean JSON with the transcript, timestamps and subtitle links it can reason over.

### Why use this Actor?

Transcribe audio and video files from direct URLs (mp3, wav, m4a, ogg, mp4, mov, webm) with Whisper large-v3-turbo. Get text, SRT, VTT, timestamped segments and RAG chunks. $0.004 per audio minute, platform usage included.

### Features

- **Audio or video file URLs** — Direct links to audio or video files, one per line (mp3, wav, m4a, ogg, opus, flac, webm, mp4, mov, mkv). Links must be public and point straight to the file. YouTube and social-media page links are not supported; use the YouTube Transcript Actor for those.
- **Spoken language** — Language spoken in the audio. Auto-detect works well for most recordings; choosing the language explicitly can improve accuracy and avoids wrong detection on short or noisy clips.
- **Task** — Transcribe keeps the original language. Translate to English returns the transcript translated into English.
- **Output formats** — What to include in each result row. SRT and VTT subtitles are also saved as files in the key-value store of the run, with direct links. RAG chunks split the transcript into timestamped pieces for embeddings.
- **RAG chunk size (characters)** — Maximum characters per chunk when the RAG chunks format is selected. Chunks follow segment boundaries and keep start and end times.
- **Maximum duration per file (minutes)** — Files longer than this are reported as TOO\_LONG and are not charged. The hard cap is 180 minutes per file.
- **Parallel transcriptions** — How many files are processed, and how many audio parts are transcribed, at the same time. Lower values are gentler on the transcription service.
- **Skip silence (voice activity detection)** — Remove silent and non-speech sections before transcribing. Helps with long pauses, music or background noise, where the model can otherwise invent text.
- **Context prompt** — Optional text that helps the model with names, product terms or spelling, for example a list of speaker names or technical vocabulary. Maximum 800 characters.

### Use cases

- Podcast and interview transcription.
- Meeting and call recording transcripts.
- Subtitle files for videos and courses.
- Embedding and vector search preparation from audio.

### Related tools

- [YouTube Transcript Scraper - Subtitles, SRT & RAG](https://apify.com/zenomastro/youtube-transcript-reliable)
- [Text Splitter for RAG - LLM Chunking API](https://apify.com/zenomastro/text-splitter-for-llm)
- [Image to Text OCR - Extract Text from Images & PDFs](https://apify.com/zenomastro/image-ocr-text-extractor)

### Example input

```json
{
  "mediaUrls": [
    "https://upload.wikimedia.org/wikipedia/commons/4/46/1941_Roosevelt_speech_pearlharbor_p1.ogg"
  ],
  "language": "auto",
  "task": "transcribe",
  "outputFormats": [
    "text",
    "srt"
  ],
  "chunkSizeChars": 1000,
  "maxDurationMinutes": 60
}
```

### Pricing & cost control

The primary event costs **$0.004000 per audio minute transcribed** (about **$4.00 per 1,000** successful primary events).
Only successful primary events are intentionally billed by this Actor; summary/status rows add context without adding primary-event charges.

Use the bounded input limits and filters to keep both event charges and platform usage predictable.

### FAQ

**What is this Actor for?**\
It is designed for podcast and interview transcription, meeting and call recording transcripts, subtitle files for videos and courses.

**Can I run it on a schedule?**\
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

**How do I control cost and run size?**\
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

# Changelog

This Actor's version history is a separate document: https://apify.com/zenomastro/audio-video-to-text/changelog.md

# Actor input Schema

## `mediaUrls` (type: `array`):

Direct links to audio or video files, one per line (mp3, wav, m4a, ogg, opus, flac, webm, mp4, mov, mkv). Links must be public and point straight to the file. YouTube and social-media page links are not supported; use the YouTube Transcript Actor for those.

## `language` (type: `string`):

Language spoken in the audio. Auto-detect works well for most recordings; choosing the language explicitly can improve accuracy and avoids wrong detection on short or noisy clips.

## `task` (type: `string`):

Transcribe keeps the original language. Translate to English returns the transcript translated into English.

## `outputFormats` (type: `array`):

What to include in each result row. SRT and VTT subtitles are also saved as files in the key-value store of the run, with direct links. RAG chunks split the transcript into timestamped pieces for embeddings.

## `chunkSizeChars` (type: `integer`):

Maximum characters per chunk when the RAG chunks format is selected. Chunks follow segment boundaries and keep start and end times.

## `maxDurationMinutes` (type: `integer`):

Files longer than this are reported as TOO\_LONG and are not charged. The hard cap is 180 minutes per file.

## `maxConcurrency` (type: `integer`):

How many files are processed, and how many audio parts are transcribed, at the same time. Lower values are gentler on the transcription service.

## `vadFilter` (type: `boolean`):

Remove silent and non-speech sections before transcribing. Helps with long pauses, music or background noise, where the model can otherwise invent text.

## `initialPrompt` (type: `string`):

Optional text that helps the model with names, product terms or spelling, for example a list of speaker names or technical vocabulary. Maximum 800 characters.

## Actor input object example

```json
{
  "mediaUrls": [
    "https://upload.wikimedia.org/wikipedia/commons/4/46/1941_Roosevelt_speech_pearlharbor_p1.ogg"
  ],
  "language": "auto",
  "task": "transcribe",
  "outputFormats": [
    "text",
    "srt"
  ],
  "chunkSizeChars": 1000,
  "maxDurationMinutes": 60,
  "maxConcurrency": 2,
  "vadFilter": false
}
```

# Actor output Schema

## `overview` (type: `string`):

Main fields of every transcript, ready to preview or export as CSV, Excel or JSON.

## `results` (type: `string`):

Every transcript with all fields, for APIs, AI agents and integrations.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mediaUrls": [
        "https://upload.wikimedia.org/wikipedia/commons/4/46/1941_Roosevelt_speech_pearlharbor_p1.ogg"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("zenomastro/audio-video-to-text").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "mediaUrls": ["https://upload.wikimedia.org/wikipedia/commons/4/46/1941_Roosevelt_speech_pearlharbor_p1.ogg"] }

# Run the Actor and wait for it to finish
run = client.actor("zenomastro/audio-video-to-text").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mediaUrls": [
    "https://upload.wikimedia.org/wikipedia/commons/4/46/1941_Roosevelt_speech_pearlharbor_p1.ogg"
  ]
}' |
apify call zenomastro/audio-video-to-text --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zenomastro/audio-video-to-text"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/0Pl2RMg4pcdJtDENv/builds/35uhaxQvsy9XQtlR0/openapi.json
