# Audio & Video to Text Transcriber: Whisper, SRT/VTT (`glistening_film/audio-video-transcriber`) Actor

- **URL**: https://apify.com/glistening\_film/audio-video-transcriber.md
- **Developed by:** [Yodesla](https://apify.com/glistening_film) (community)
- **Categories:** AI, Videos
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $12.00 / 1,000 audio minute (tiny/base model)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Audio & Video to Text Transcriber: Whisper, SRT/VTT (per minute)

Turn **public audio and video files** into clean text by URL, up to 20 files per run, with
OpenAI Whisper ([faster-whisper](https://github.com/SYSTRAN/faster-whisper), int8 on CPU).
Get the full transcript, timestamped segments, and ready-to-use **SRT or VTT** subtitle
files. **$0.012 per transcribed minute (tiny/base models) or $0.025 (small model), no start fee.** Built for podcast/newsletter
transcription, meeting and interview notes, subtitle generation, and RAG ingestion of
audio/video content.

### Use cases

- Transcribing podcasts, lectures, interviews and voicemails to text
- Generating SRT/VTT subtitles for videos
- Timestamped segment data for search, chaptering and RAG ingestion
- Meeting and call notes in an automation pipeline (mp3, m4a, wav, ogg, flac, mp4, webm, mov)

### What you get

For each file, one dataset item:

- `status`: `ok`, `download_failed`, `unsupported_format`, `invalid_file`, or `no_speech`
- `durationSec`: seconds actually transcribed (long files are truncated, see below)
- `language` + `languageProbability`: the language Whisper detected (or your forced one)
- `text`: the full transcript (when `text` is in the output formats)
- `segments[]`: `{start, end, text}` per speech segment (when requested)
- `srt` / `vtt`: complete subtitle files with correct timestamps (when requested)
- `billedMinutes`: whole minutes rounded up, charged as `audio-minute` events
- `truncated`: true when the file was longer than what was transcribed
- `processingMs`, and `error` with a plain-language reason when something fails

The file type is detected from its **content (magic bytes)**, never the URL extension: a
`.mp3` that is really an MP4 is transcribed as an MP4. Audio longer than **180 minutes** is
truncated and flagged `truncated`.

### Pricing: you only pay for minutes that worked

**One event per whole minute (rounded up) transcribed from a file that came back `ok`:
`audio-minute` ($0.012) for the tiny/base models, `audio-minute-small` ($0.025) for the
more accurate small model.** A 90-second clip bills 2 minutes. Failed downloads, unsupported or invalid
files, and files with no speech (`no_speech`) are **reported in the output but never
charged**. There is no start fee. The run checks your spending limit before each file,
transcribes at most the minutes your remaining limit covers (the rest of the file is
skipped), and stops cleanly when the limit is reached.

### Accuracy: Whisper on CPU

The engine is OpenAI Whisper running as faster-whisper in int8 on CPU — strong on clear
speech in most languages, weaker on heavy music, overlap and strong accents. Model sizes:
**tiny** (fastest), **base** (default, good balance), **small** (most accurate). Language is
auto-detected by default; pass an ISO 639-1 code to force one.

### Limits (by design)

- Max 20 URLs per run, max 200 MB per file, max 180 minutes transcribed per file (truncated
  and flagged beyond that).
- Supported inputs: mp3, m4a, wav, ogg, flac, mp4, webm, mov (audio track of the video).
  Anything else is reported as `unsupported_format` and not charged.
- Only public `http(s)` links; local and private-network addresses are refused.
- CPU inference on the base model runs roughly 0.5-1x realtime per core; a 60-minute file
  takes a while by design.

### FAQ

**Do I pay for files with no speech or that fail?** No. Failed downloads, unsupported or
invalid files, and `no_speech` results are reported but never charged; you pay only for
whole minutes transcribed from files that came back `ok`, with no start fee.

**What happens if a file is longer than my remaining limit?** It is truncated to the
minutes your limit covers, those minutes are billed, and the run stops cleanly.

**Does it work on video?** Yes — the audio track of mp4, webm or mov files is extracted and
transcribed with the system ffmpeg.

**Which languages are supported?** Whatever Whisper supports (99+ languages), auto-detected
by default or forced with the `language` field.

### Input example

```json
{
  "urls": ["https://example.com/podcast-ep1.mp3"],
  "language": "",
  "model": "base",
  "outputFormats": ["text", "srt", "vtt", "segments"]
}
```

# Actor input Schema

## `urls` (type: `array`):

Public http(s) links to audio or video files (mp3, m4a, wav, ogg, flac, mp4, webm, mov; max 20 per run, max 200 MB each). The file type is detected from its content (magic bytes), not the URL extension. Only files with status ok are charged.

## `language` (type: `string`):

ISO 639-1 code of the audio (e.g. en, de, fr). Leave empty to auto-detect the language.

## `model` (type: `string`):

Whisper model size. Bigger is more accurate and slower on CPU: tiny (fastest), base (default, good balance), small (most accurate).

## `outputFormats` (type: `array`):

Which outputs to include in each result: plain text, SRT subtitles, WebVTT subtitles, and/or the timestamped segment list.

## Actor input object example

```json
{
  "urls": [
    "https://upload.wikimedia.org/wikipedia/commons/a/a1/Hello_world_said_by_eSpeakNG.ogg"
  ],
  "language": "",
  "model": "base",
  "outputFormats": [
    "text",
    "segments"
  ]
}
```

# Actor output Schema

## `results` (type: `string`):

One dataset item per input URL: status (ok, download\_failed, unsupported\_format, invalid\_file or no\_speech), transcribed duration in seconds, detected language with probability, full text, timestamped segments, SRT and VTT subtitle files, billed minutes and error reason. Only items with status ok are charged, ceil(minutes transcribed) audio-minute events each.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://upload.wikimedia.org/wikipedia/commons/a/a1/Hello_world_said_by_eSpeakNG.ogg"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("glistening_film/audio-video-transcriber").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://upload.wikimedia.org/wikipedia/commons/a/a1/Hello_world_said_by_eSpeakNG.ogg"] }

# Run the Actor and wait for it to finish
run = client.actor("glistening_film/audio-video-transcriber").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://upload.wikimedia.org/wikipedia/commons/a/a1/Hello_world_said_by_eSpeakNG.ogg"
  ]
}' |
apify call glistening_film/audio-video-transcriber --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,glistening_film/audio-video-transcriber"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/iuCDa4LUJGreT9Gxe/builds/VnfecqlTlpi17gRJ2/openapi.json
