# Speech to Text & Audio Transcription: $0.0025/min, No Login (`conserving_celerytop/audio-podcast-transcription`) Actor

$0.0025 per audio minute in English. Transcribe audio and video files and podcast RSS episodes to text with timestamps and SRT or VTT subtitles. High-accuracy English $0.003 and about 100 other languages $0.01 per minute. No API key or login. Failed or silent files are free.

- **URL**: https://apify.com/conserving\_celerytop/audio-podcast-transcription.md
- **Developed by:** [Don Mangu](https://apify.com/conserving_celerytop) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 audio minute (english, standard)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Audio and Podcast Transcription

Audio and Podcast Transcription turns audio and video files into text. Give it direct links to your own recordings, or podcast RSS feeds, and it returns one row per file with the full transcript, the spoken language, the length, and timestamped segments. It also saves SRT and VTT subtitle files for each transcript. It costs $0.0025 per audio minute in English ($0.003 with high accuracy) and $0.01 per minute in about 100 other languages, with no API key and no login.

The speech recognition runs inside the Actor with open-weights models. Your audio is not sent to an outside transcription service, and nothing is kept after the run apart from the results in your own Apify storage.

### What the audio to text transcription returns

- **Transcript**: `text`, the full transcript, and `wordCount`.
- **Timestamped segments**: `segments`, a list of `start` and `end` (in seconds) and `text`, one per spoken phrase. Turn off **Include timestamped segments** to leave them out.
- **Subtitles**: `srtUrl` and `vttUrl` link to SubRip (.srt) and WebVTT (.vtt) files in the run's key-value store, ready for video players and editors.
- **File details**: `language`, `durationSeconds`, `transcribedSeconds`, `speechSeconds`, `audioCodec`, `fileSizeBytes` and the final `fileUrl` after redirects.
- **Podcast details** for feed episodes: `feedTitle`, `episodeTitle` and `episodePublishedAt`, as the feed publishes them.
- **Status for every file**: `status` is `ok`, or says why there is no transcript (for example `not_found`, `not_media`, `no_speech`, `file_too_large`), with a plain-language `error`.

### How to transcribe a podcast or audio file, step by step

1. Open the Actor and go to the **Input** tab.
2. In **Audio and video file links**, paste direct links to files, one per line. MP3, M4A, AAC, WAV, FLAC, OGG, OPUS, MP4, MOV, WEBM and MKV work. A link to a web page that plays audio is not a file link: open the page, find the download or file link, and use that.
3. To transcribe a podcast, paste its RSS feed link in **Podcast RSS feeds** and set **Episodes per feed** (the newest episodes are taken first).
4. Pick the **Language**. For English, pick **English accuracy**: Standard (lowest price) or High (fewer errors on real-world recordings). For any other language, pick it, or pick **Detect automatically**.
5. Choose the **Subtitle files** you want (SRT, VTT or both).
6. Click **Start**. The example input transcribes an 11-second clip in a few seconds.
7. Open the **Output** tab. The **Overview** view shows each transcript with its subtitle links; the **Segments** view shows one row per timestamped segment; the **Podcast episodes** view shows feed and episode details. Download as JSON, CSV or Excel, or read the results through the Apify API.

### How much does audio transcription cost?

You pay per minute of audio transcribed, counted for each file and rounded up to the next whole minute:

| Setting | Event | Price per audio minute |
| --- | --- | --- |
| English, Standard accuracy | `audio-minute` | $0.0025 |
| English, High accuracy | `audio-minute-high-accuracy` | $0.003 |
| Any other language, or Detect automatically | `audio-minute-multilingual` | $0.01 |

Files that could not be downloaded or decoded, links that are not audio or video, and files with no speech return a row with the reason and are **not charged**.

**Worked example:** a podcast feed with 4 English episodes of 32, 45, 51 and 38 minutes, plus one broken link, costs (32 + 45 + 51 + 38) x $0.0025 = 166 x $0.0025 = **$0.42** at Standard accuracy, or 166 x $0.003 = $0.50 at High accuracy. The broken link costs nothing. The same 166 minutes in Spanish cost 166 x $0.01 = $1.66.

Set **Maximum minutes per file** to transcribe only the start of long files, and set a spending limit on the run: the Actor transcribes only the minutes that fit, and the row says the file was cut (`truncated` is `true`).

### Input example

```json
{
    "audioUrls": ["https://raw.githubusercontent.com/openai/whisper/main/tests/jfk.flac"],
    "rssFeedUrls": [],
    "episodesPerFeed": 1,
    "maxFiles": 1,
    "language": "en",
    "englishAccuracy": "high",
    "subtitleFormats": ["srt", "vtt"],
    "includeSegments": true,
    "maxMinutesPerFile": 240,
    "maxFileSizeMb": 1024
}
```

| Field | Meaning |
| --- | --- |
| `audioUrls` | Direct links to audio or video files. |
| `rssFeedUrls` | Podcast RSS feed links. |
| `episodesPerFeed` | Newest episodes to transcribe from each feed (1 to 100). |
| `maxFiles` | Most files to transcribe in the run (1 to 1,000). File links come first, then feed episodes. |
| `language` | `en` for English, another language code, or `auto` to detect it. |
| `englishAccuracy` | `standard` (lowest price) or `high` (English only). |
| `subtitleFormats` | `srt`, `vtt`, both, or none. |
| `includeSegments` | Add timestamped segments to each row. |
| `maxMinutesPerFile` | Transcribe at most this many minutes of each file (1 to 720). |
| `maxFileSizeMb` | Largest file to download, in MB (1 to 4,096). |

### Output example

A row from a test run with the example input, High accuracy (the 11-second clip is the "Ask not" passage of the 1961 US presidential inaugural address):

```json
{
    "inputUrl": "https://raw.githubusercontent.com/openai/whisper/main/tests/jfk.flac",
    "fileUrl": "https://raw.githubusercontent.com/openai/whisper/main/tests/jfk.flac",
    "source": "url",
    "status": "ok",
    "language": "en",
    "durationSeconds": 11.0,
    "transcribedSeconds": 11.0,
    "speechSeconds": 8.02,
    "billedMinutes": 1,
    "truncated": false,
    "text": "And so my fellow Americans. Ask not. What your country can do for you? Ask what you can do for your country.",
    "wordCount": 22,
    "segments": [
        { "start": 0.29, "end": 2.29, "text": "And so my fellow Americans." },
        { "start": 3.27, "end": 4.5, "text": "Ask not." },
        { "start": 5.38, "end": 7.8, "text": "What your country can do for you?" },
        { "start": 8.17, "end": 10.55, "text": "Ask what you can do for your country." }
    ],
    "srtUrl": "https://api.apify.com/v2/key-value-stores/<store id>/records/transcript-0001.srt",
    "vttUrl": "https://api.apify.com/v2/key-value-stores/<store id>/records/transcript-0001.vtt",
    "model": "English, high accuracy",
    "audioCodec": "flac",
    "charged": true,
    "error": null
}
```

The SRT file for the same clip:

```
1
00:00:00,290 --> 00:00:02,290
And so my fellow Americans.

2
00:00:03,270 --> 00:00:04,500
Ask not.
```

### Use cases

- **Podcasters**: show notes, blog posts and searchable archives from your episodes; subtitles for video versions.
- **Content and marketing teams**: quotes and clips from webinars, interviews and recorded talks you own.
- **Researchers**: text from interview recordings and public-domain speech collections.
- **Developers**: a transcription step in a pipeline, called through the Apify API or scheduled to pick up new podcast episodes.

### Related Actors

Other Actors from the same developer turn public sources into clean datasets: company job boards, SEC filings and Form D funding rounds, public tenders, weather observations and website technology lookups. Search the Apify Store for them if your pipeline needs more data next to your transcripts.

### FAQ

**Is it legal to transcribe audio with this Actor?**
Transcribe files you own or have the right to use: your own recordings, public-domain audio, and podcast episodes for your own listening, research or accessibility needs. The Actor downloads only the files and feeds you list, reads robots.txt on every host first and skips files it disallows, and never logs in. Republishing someone else's podcast transcript can need the owner's permission; check the podcast's terms.

**Does it work with YouTube, Spotify, Apple Podcasts or social media links?**
No. It needs a direct link to an audio or video file, or a podcast RSS feed. Page links from video and social platforms are not supported. For a podcast, use its RSS feed: the Actor reads the episode file links from it.

**Which languages are supported?**
English has its own models (Standard and High accuracy). Other languages (about 100, including Spanish, French, German, Portuguese, Italian, Dutch, Polish, Russian, Hebrew, Arabic, Hindi, Japanese, Korean and Chinese) use the multilingual model. Pick the language when you know it; **Detect automatically** finds it from the audio.

**How accurate is it, and Standard or High?**
On a clean 11-minute English recording, both English settings got about 95 to 97 words in 100 right, not counting punctuation and capital letters. On short real-world recordings, High accuracy was clearly better: on 5 test clips it made no errors that changed the meaning, while Standard misheard words in 2 of them (for example "As not" for "Ask not"). Use Standard for clear, single-speaker audio and when price matters most; use High for interviews, phone-quality audio and anything you will publish. Accuracy drops with background music, crosstalk, heavy accents and poor microphones. There is no speaker labelling (who said what).

**How long does it take?**
At the default 4 GB memory, an hour of English audio takes about 3 minutes of processing at Standard accuracy and about 5 to 6 minutes at High accuracy; other languages take about 17 to 20 minutes per hour. Downloading adds a little. More memory gives more CPU and a faster run; the price per minute stays the same.

**Why does a file show `robots_disallowed` or `blocked`?**
The host's robots.txt does not allow that address, or the server refused the download (HTTP 401 or 403). The Actor does not retry or work around it, and the row is free. After two refusals from the same host, the rest of that host's files in the run are skipped.

**What are the limits?**
Up to 1,000 files per run, files up to 4,096 MB, and up to 720 minutes per file. Very long files are best split across runs with **Maximum minutes per file**.

**Is my audio stored?**
Downloaded files are deleted as soon as each one is transcribed. Only the transcript rows, subtitle files and run statistics stay, in your own Apify storage.

# Actor input Schema

## `audioUrls` (type: `array`):

Enter direct links to audio or video files, one per line: MP3, M4A, AAC, WAV, FLAC, OGG, OPUS, MP4, MOV, WEBM or MKV. Use your own files or a podcast episode's file link.

## `rssFeedUrls` (type: `array`):

Enter podcast RSS feed links, one per line. The newest episodes of each feed are transcribed (see Episodes per feed).

## `episodesPerFeed` (type: `integer`):

Transcribe this many of the newest episodes from each podcast RSS feed.

## `maxFiles` (type: `integer`):

Transcribe at most this many files in this run, file links first, then feed episodes. Duplicates count once.

## `language` (type: `string`):

Choose the spoken language. English uses the English models. Any other choice, or Detect automatically, uses the multilingual model (about 100 languages), which costs more per minute.

## `englishAccuracy` (type: `string`):

Standard is the lowest price. High uses a larger English model with fewer errors and costs more per minute. Applies when Language is English.

## `subtitleFormats` (type: `array`):

Save these subtitle files for each transcript in the key-value store. Each dataset row links to them.

## `includeSegments` (type: `boolean`):

Add the segments list (start, end, text) to each dataset row.

## `maxMinutesPerFile` (type: `integer`):

Transcribe at most the first this many minutes of each file. You pay only for the minutes transcribed.

## `maxFileSizeMb` (type: `integer`):

Download files up to this size. Larger files return a row with status file\_too\_large and are free.

## Actor input object example

```json
{
  "audioUrls": [
    "https://raw.githubusercontent.com/openai/whisper/main/tests/jfk.flac"
  ],
  "rssFeedUrls": [
    "https://example.com/podcast/feed.xml"
  ],
  "episodesPerFeed": 1,
  "maxFiles": 1,
  "language": "en",
  "englishAccuracy": "high",
  "subtitleFormats": [
    "srt",
    "vtt"
  ],
  "includeSegments": true,
  "maxMinutesPerFile": 240,
  "maxFileSizeMb": 1024
}
```

# Actor output Schema

## `overview` (type: `string`):

inputUrl, episodeTitle, status, language, durationSeconds, billedMinutes, wordCount, text, srtUrl, vttUrl and error.

## `segments` (type: `string`):

One row per segment: start, end and text.

## `subtitles` (type: `string`):

Key-value store records transcript-0001.srt, transcript-0001.vtt and so on.

## `stats` (type: `string`):

JSON with files planned and done, statuses, audio minutes billed, seconds transcribed, CPU seconds, requests, retries and peak memory.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "audioUrls": [
        "https://raw.githubusercontent.com/openai/whisper/main/tests/jfk.flac"
    ],
    "maxFiles": 1,
    "language": "en",
    "englishAccuracy": "high",
    "subtitleFormats": [
        "srt",
        "vtt"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("conserving_celerytop/audio-podcast-transcription").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "audioUrls": ["https://raw.githubusercontent.com/openai/whisper/main/tests/jfk.flac"],
    "maxFiles": 1,
    "language": "en",
    "englishAccuracy": "high",
    "subtitleFormats": [
        "srt",
        "vtt",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("conserving_celerytop/audio-podcast-transcription").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "audioUrls": [
    "https://raw.githubusercontent.com/openai/whisper/main/tests/jfk.flac"
  ],
  "maxFiles": 1,
  "language": "en",
  "englishAccuracy": "high",
  "subtitleFormats": [
    "srt",
    "vtt"
  ]
}' |
apify call conserving_celerytop/audio-podcast-transcription --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,conserving_celerytop/audio-podcast-transcription"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/1phrSwVYhpn48erFN/builds/meC590CtXvAr0OGG6/openapi.json
