# Audio & Video to Text: Whisper Transcription & SRT Subtitles (`spokentext/audio-video-to-text`) Actor

Turn any audio or video file into text. Paste a file link (MP3, MP4, WAV, M4A...), a Google Drive or Dropbox share link, or upload a file. Get the transcript, timestamps and SRT subtitles in 99+ languages. Pay per minute.

- **URL**: https://apify.com/spokentext/audio-video-to-text.md
- **Developed by:** [clement](https://apify.com/spokentext) (community)
- **Categories:** AI, For creators, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 audio minute transcribeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Audio & Video to Text: Whisper Transcription & SRT Subtitles

Turn any audio or video file into accurate text. Paste a file link, a **Google Drive** or **Dropbox** share link, or **upload a file**, and get the transcript as plain text, timestamped segments and SRT subtitles.

- **Any common format.** MP3, M4A, WAV, OGG, FLAC, AAC, MP4, MOV, WEBM and more. For video, the audio track is transcribed.
- **Pay per audio minute.** No subscription, and nothing charged for files that fail.
- **99+ languages** with automatic language detection (Whisper large-v3).
- **Fast.** One hour of audio is typically ready in about a minute.
- **Long recordings handled.** Files are split and stitched back together automatically. Recordings up to about 2 hours are supported, in files up to 1.5 GB.

### What you can provide

| Input | Example |
|---|---|
| Direct file link | `https://example.com/interview.mp3` |
| Google Drive share link | `https://drive.google.com/file/d/FILE_ID/view` |
| Dropbox share link | `https://www.dropbox.com/scl/fi/.../meeting.mp4?rlkey=...` |
| File upload | Use the **Or upload a file** field |

Share links must be set to "anyone with the link can view". You can mix several links in one run.

### How to use

1. Get a link to your file: a direct file URL, or a Google Drive or Dropbox link shared with "anyone with the link". Or skip links and use **Or upload a file**.
2. Paste the link into **Audio or video file URLs**. Add as many as you like.
3. Optionally set the language or turn on SRT subtitles.
4. Click **Start**. When the run finishes, open the **Output** tab and download the transcripts as JSON, CSV or Excel.

### Output

One dataset item per file:

```json
{
    "inputUrl": "https://example.com/interview.mp3",
    "status": "ok",
    "fileName": "interview.mp3",
    "language": "English",
    "durationSeconds": 1842,
    "transcribedMinutes": 31,
    "truncated": false,
    "text": "Thanks for joining me today. Let's start with...",
    "segments": [{ "start": 0, "end": 3.4, "text": "Thanks for joining me today." }],
    "srt": "1\n00:00:00,000 --> 00:00:03,400\nThanks for joining me today.\n"
}
```

`segments` is included by default; `srt` is included when **Include SRT subtitles** is on. Files that cannot be processed produce an item with `"status": "error"` and an `error` message explaining why.

### Pricing

You pay for each minute of audio transcribed, rounded to the nearest minute (minimum one minute per file). Set **Maximum minutes per file** or the run's maximum charge to cap spending: the Actor stops before exceeding either, and marks a cut file with `"truncated": true`.

### Use cases

- Transcribe interviews, meetings, lectures, calls and voice memos.
- Create SRT subtitles for videos.
- Turn recordings into searchable notes, or feed them to an LLM for summaries and analysis.
- Batch-process an archive of recordings through the API.

### Limits

- **Links must point to a file, not a web page.** YouTube, TikTok and other video-page links are not supported.
- For podcasts, use [Spotify Podcast Transcript](https://apify.com/spokentext/spotify-podcast-transcript), which accepts Spotify, Apple Podcasts and RSS links.
- No speaker labels: the transcript does not say who is speaking.
- Private files that require a login cannot be read. Upload them instead.

### FAQ

#### How do I convert an audio file to text?

Paste a link to the file (MP3, M4A, WAV and more) or upload it, then start the run. The transcript appears in the Output tab, with timestamps for every sentence.

#### How do I transcribe a video to text?

The same way. For MP4, MOV, WEBM and other video files the Actor reads the audio track and ignores the picture, so large videos are handled quickly.

#### How long does it take and how much does it cost?

A 50-minute recording takes about 40 seconds and costs $0.15. Price is $0.003 per audio minute, with nothing charged for files that fail.

#### How accurate is the transcript?

It uses Whisper large-v3. Clear speech comes out very accurately; names, jargon, heavy background noise and people talking over each other can contain mistakes. The transcript does not label speakers.

#### Which languages are supported?

More than 99, detected automatically. For short or noisy audio, setting the language code (en, fr, es, de...) improves accuracy.

#### Can I transcribe long recordings?

Yes, up to about 2 hours per recording. The file is split into parts, transcribed and stitched back together with continuous timestamps. Cut longer recordings into parts of under 2 hours first.

#### Can I transcribe a YouTube or TikTok link?

No. Links must point to a file. Download or export the recording first, then upload it or share it through Google Drive or Dropbox.

#### Can I use it from my own code, Make, Zapier or n8n?

Yes. Call it through the Apify API (see the example below), or connect it with Apify's integrations for Make, Zapier and n8n. Every run returns the same JSON.

#### What happens to my file?

The file is downloaded during the run, its audio is sent to our speech-to-text provider (Groq) to be transcribed, and it is deleted when the run ends. Transcripts are stored only in your own Apify account.

### Run it from the API

```bash
curl -X POST "https://api.apify.com/v2/acts/spokentext~audio-video-to-text/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "urls": ["https://example.com/interview.mp3"], "includeSrt": true }'
```

# Actor input Schema

## `urls` (type: `array`):

Direct links to audio or video files (mp3, m4a, wav, ogg, flac, mp4, mov, webm...), or Google Drive and Dropbox share links. Share links must be open to anyone with the link.

## `file` (type: `string`):

Upload one audio or video file from your computer instead of pasting a link.

## `language` (type: `string`):

Two-letter language code of the audio (en, fr, es, de...). Leave empty to detect it automatically. Setting it improves accuracy on short or noisy audio.

## `includeSegments` (type: `boolean`):

Add a list of timestamped segments (start, end, text) to each result.

## `includeSrt` (type: `boolean`):

Add the transcript as an SRT subtitle file string, ready to load into a video editor or player.

## `maxMinutesPerFile` (type: `integer`):

Longer files are cut at this length, so a single run can never cost more than you expect.

## Actor input object example

```json
{
  "urls": [
    "https://api.apify.com/v2/key-value-stores/kyR7mGkrVQ6H3xQe8/records/jfk-we-choose-the-moon.mp3?signature=1IsTRMMS6FEWC0SVKQvL6"
  ],
  "includeSegments": true,
  "includeSrt": false,
  "maxMinutesPerFile": 300
}
```

# Actor output Schema

## `transcripts` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://api.apify.com/v2/key-value-stores/kyR7mGkrVQ6H3xQe8/records/jfk-we-choose-the-moon.mp3?signature=1IsTRMMS6FEWC0SVKQvL6"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("spokentext/audio-video-to-text").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://api.apify.com/v2/key-value-stores/kyR7mGkrVQ6H3xQe8/records/jfk-we-choose-the-moon.mp3?signature=1IsTRMMS6FEWC0SVKQvL6"] }

# Run the Actor and wait for it to finish
run = client.actor("spokentext/audio-video-to-text").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://api.apify.com/v2/key-value-stores/kyR7mGkrVQ6H3xQe8/records/jfk-we-choose-the-moon.mp3?signature=1IsTRMMS6FEWC0SVKQvL6"
  ]
}' |
apify call spokentext/audio-video-to-text --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,spokentext/audio-video-to-text"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/9CeMMqpCRgnEcJfvx/builds/tFd7ysdBmDLSm6oE8/openapi.json
