# Audio & Video to Text - Speech to Text Transcription, SRT (`tidytools/audio-transcriber`) Actor

Whisper transcription: audio to text and video to text from files, Drive/Dropbox links and podcast RSS feeds, with timestamps and SRT/VTT subtitles. 90+ languages, optional speaker labels. $0.006/min.

- **URL**: https://apify.com/tidytools/audio-transcriber.md
- **Developed by:** [Yukai Lin](https://apify.com/tidytools) (community)
- **Categories:** AI, Videos, Automation
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Audio & Video Transcriber do?

It turns **audio files, videos and podcast episodes** into **text with timestamps**, readable **paragraphs**, and ready-to-use **SRT and VTT subtitle files**. It uses OpenAI's open **Whisper large-v3-turbo** model, detects the language automatically and supports 90+ languages.

- 🎙️ **Podcasts, interviews, meetings, lectures, webinars, audiobooks, voice notes**
- 🎬 **Video to text**: MP4, MOV, MKV, WEBM and more. The audio track is extracted automatically
- 📡 **Podcast RSS feeds**: paste a feed (or an Apple Podcasts show link) and the newest episodes are transcribed, with episode title and publish date
- 📝 **Publisher transcripts, no audio minutes**: when a podcast publishes its own transcript in the feed (Podcasting 2.0 `<podcast:transcript>`: VTT, SRT, JSON, HTML or text), it is used instead of transcribing, with the same output and `transcriptSource: "publisher"`
- 🗓️ **Feed filters**: only episodes published since / until a date, or whose title contains (or does not contain) a text
- 📤 **File upload**: upload recordings from your computer in Apify Console (`audioFiles`)
- 🌐 **Web pages with one player**: a page with exactly one audio or video file (e.g. a news, lecture or archive.org page); the file is found and transcribed
- 🔗 **Google Drive, Dropbox and OneDrive share links** work if the file is shared publicly
- ⏱️ **Timestamped segments**, **paragraphs** and full plain text in one result
- 🗣️ **Speaker labels** (`"speakerLabels": true`): who said what, as Speaker 1, Speaker 2… on every segment, paragraph and subtitle line, plus the number of speakers
- 🔠 **Word-level timestamps** (`"wordTimestamps": true`): start and end time of every word, for video editing and precise subtitles (included in the regular price)
- 🌍 **Translated subtitles in one step** (`"subtitleLanguages": ["Spanish", "German"]`): SRT and VTT in up to 5 extra languages, timings and speaker labels kept
- 🔁 **Only new episodes** (`"newEpisodesOnly": true`): scheduled runs skip episodes and files that were already transcribed, and only take episodes released after the newest one already done, so a day without a new episode costs nothing (`backfillOlderEpisodes` works through the back catalogue instead)
- 🧾 **SRT and VTT subtitles** saved for every file (no extra charge), cut into subtitle-sized cues: at most 7 seconds and two lines of 42 characters, timed by word
- 🧠 **Optional AI summary and chapters** (`"summarize": true`), with no API key needed
- 🌍 **Automatic language detection**, or set the language yourself
- 🔤 **Vocabulary hints**: give names and terms (e.g. "Apify, Kubernetes") so they are spelled correctly
- 📏 **Large files**: up to 2 GB per file, any length. Long recordings are cut in pauses and transcribed in parallel
- 🔕 **Silent files are free**: files with no speech are reported but not charged
- ⚡ **Fast**: in our tests an 8-minute MP4 video took 33 seconds and a 26-minute MP3 took 109 seconds

### How much does it cost?

| Event | Price |
|---|---|
| Audio minute | **$0.006 per minute** ($0.36 per hour) |
| Audio minute with speaker labels (optional) | **$0.015 per minute** ($0.90 per hour), **instead of** the regular audio minute, only when `speakerLabels` is on |
| Translated subtitles (optional) | **$0.002 per minute per language** (e.g. a 30-minute episode in Spanish and German = $0.12), only for languages that succeed |
| AI summary and chapters (optional) | **$0.01 per file**, only when `summarize` is on |

- **No start fee.** You pay only for transcribed audio.
- **SRT and VTT subtitles, paragraphs and timestamps are included.**
- Billed per started minute of audio (a 500-second video is 9 minutes).
- **Word-level timestamps are included in the regular $0.006 per minute** (Whisper word timings). With speaker labels on, words come from nova-3 and the speaker-label price applies, not both.
- **Translated subtitles** (`subtitleLanguages`, up to 5 languages): SRT + VTT per language with the original timings and speaker labels, made right after the transcript. A language that fails is free and reported in `translatedSubtitles`.
- **Not charged:** files that cannot be downloaded or decoded, files with no speech, files skipped by your limits (`maxMinutesPerFile`, `maxTotalMinutes`), files skipped by `newEpisodesOnly`, episodes removed by the feed filters, and the audio minutes of episodes that use the publisher's own transcript (with `summarize` on, their AI summary is still $0.01).
- Higher Apify plans get volume discounts (Silver 10%, Gold and above 20%).

#### Control your cost

- **What is charged**: the minutes of each transcribed file (rounded up per file) and, with `summarize` on, one AI summary per file. Minutes are charged only after the transcript row has been saved.
- **What is free**: lines that are not a URL, duplicate links, YouTube / TikTok / Instagram / Spotify page links, web pages without exactly one media file, files that cannot be downloaded or decoded, files without speech, files skipped by your limits, and publisher transcripts (`billedMinutes: 0`). Every row has `charged: true/false`, and failed rows say "(not charged)".
- **Plan at the start**: the run's status message and log start with the cost plan (audio minutes known from podcast feeds, summaries, and how many files are measured when downloaded); it is also saved as `SUMMARY.costPlan`.
- **Caps**: *Max minutes per file* and *Max total minutes per run* (Advanced settings) skip files before they are transcribed, using the duration read from the file (or from the podcast feed, before downloading).
- **Max charge per run**: set *Maximum cost per run* in the run options (or `maxTotalChargeUsd` in the API). A file is only started when the limit still covers all its minutes; when it does not, the run stops cleanly: `SUMMARY.status` is `LIMIT_REACHED` and `SUMMARY.notProcessed` lists the files that were not started (count and up to 100 inputs). The AI summary is only made when the limit also covers it.
- **If Apify restarts the run** (server migration or Resurrect), items already finished are skipped and not charged again (`SUMMARY.resumedSkipped`).
- **Time limit per file** (`fileTimeoutSecs`, default 1,800 seconds): a file whose download, conversion and transcription take longer gets a row with `errorType: "timeout"` and is not charged, and the run goes on with the next file. Raise it for recordings of several hours.
- **Files in parallel** (`fileConcurrency`, default 3, up to 10): several files are downloaded and transcribed at the same time, so podcast batches finish sooner. Rows are saved in the order files finish; `inputIndex` is each file's position in the input. Your max charge per run and *Max total minutes per run* still hold: before a file is transcribed, its cost is reserved together with the files already in progress, so files running side by side can never go over the limit. With little run memory fewer files run at once (about 3 at the default 1,024 MB); the time a file waits for budget held by other files does not count against `fileTimeoutSecs`.

#### Price comparison (checked September 2026)

Per-minute prices of other transcription Actors on Apify Store, Free plan unless noted; "not stated" means we did not find it in that Actor's input or README:

| Actor | Price | Speaker labels |
|---|---|---|
| **This Actor** | **$0.006 per minute**, no start fee (Gold+ $0.0048) | $0.015 per minute instead of $0.006 |
| sian.agency / INCREDIBLY-FAST-audio-transcriber | $0.30 per minute ($0.005 per second) on the Free plan (Bronze $0.015, Gold $0.009), plus $0.005 per run | extra $0.001 per second on the Free plan, $0.0001 per second on paid plans |
| memo23 / video-audio-transcriber | $0.048 per minute (billed per second, $0.0008 per second), plus $0.005 per run | yes |
| amanatools / whisper-transcriber | $0.008 per minute, plus $0.005 per file | not stated |
| steadyfetch / media-transcriber | $0.003 per minute | not stated |
| kaz\_kakyo / audio-transcriber | $0.01 per minute; chapters $0.01 per file | yes (Deepgram nova-3) |
| viralanalyzer / video-transcriber | $0.015 per minute | not stated |
| hgservices / speech-to-text | $0.042 per minute (Gold $0.024), plus $0.09 per YouTube video | yes |
| tictechid / vanzi-universal-transcriber | $0.0025 per second ($0.15 per minute), plus $0.01 per run | not stated |
| parseforge / audio-transcriber | $0.44 per file, plus $0.012 per run | not stated |

We are not the cheapest per minute: steadyfetch is lower. Some of these Actors also accept YouTube or TikTok links, which this Actor does not.

### Supported input

| Input | Details |
|---|---|
| Audio files | MP3, WAV, M4A, AAC, OGG, OPUS, FLAC, AIFF, WMA, AMR and other formats ffmpeg can read |
| Video files | MP4, MOV, MKV, WEBM, AVI and more. Only the audio track is used |
| Size | Up to **2 GB** per file, any length |
| Share links | Google Drive, Dropbox, OneDrive / SharePoint (the file must be shared publicly) |
| Uploads | `audioFiles`: files uploaded in Apify Console (stored in a key-value store of your account). In API calls, pass file URLs in `urls` |
| Web pages | A page with exactly one `<audio>` / `<video>` file, or one `og:audio` / `og:video` / JSON-LD `contentUrl` file link; the found file is in `mediaUrl` |
| Podcast feeds | Any podcast RSS feed, or an Apple Podcasts show or episode link. Publisher transcripts are used when the feed has them; date and title filters are optional |
| Dataset | The `datasetId` of another Actor run; the media URL field is detected automatically or set with `datasetField` |

Links must be downloadable without logging in. YouTube, TikTok, Instagram, Facebook, X, Vimeo, Loom, SoundCloud and Spotify page links are web pages, not files: each comes back as a free row with `errorType: "unsupported_platform"` and a hint. Download the media first (for example with a downloader Actor from Apify Store) and pass the file URL, or that Actor's dataset via `datasetId`. For Vimeo and Loom, use the file behind the video's Download button when the owner allows downloads; direct file links on these hosts, such as a Vimeo "video file link" (`player.vimeo.com/external/...` or `.../progressive_redirect/...`), are transcribed normally.

**Web pages**: for any other web page, the page is read and its media file is looked up: the one `<audio>` or `<video>` element (its `src` or its first playable `<source>`; several sources of one element are one file), or else one direct file named by `og:audio`, `og:video`, `twitter:player:stream` or JSON-LD `contentUrl`. Only direct files count: streaming playlists (HLS `.m3u8`, DASH `.mpd`) and player pages are not used. When the page has no such file (`errorType: "no_media"`) or several (`"multiple_media"`, listed in `mediaCandidates`), the row is free; pass the file you want as a direct URL.

A line that is not a URL does not stop the run: it becomes one free row with `errorType: "invalid_input"` and the other files are processed. Use `urls` (a plain list of strings) in API calls; the request-list field `audioUrls` from earlier versions still works.

### How to use it

1. Paste **audio or video URLs** (or pages with one player), **upload files**, and/or add **podcast feeds** (optionally with date and title filters).
2. Optional: set the **language** (e.g. `en`, `es`, `ja`), **vocabulary hints**, and turn on **AI summary and chapters**, **speaker labels** or **word-level timestamps**.
3. Optional (Advanced settings): cap the cost with **Max minutes per file** and **Max total minutes per run**; for schedules, turn on **Only new episodes** with a **monitor name**.
4. Click **Start**. Each file becomes one result with the text, segments, paragraphs and links to its SRT/VTT files.

#### Input example

```json
{
    "urls": ["https://example.com/meeting-recording.mp4"],
    "podcastFeeds": ["https://feeds.npr.org/510289/podcast.xml"],
    "maxEpisodesPerFeed": 3,
    "language": "en",
    "vocabulary": "Apify, Cloudflare, TidyTools",
    "summarize": true,
    "maxMinutesPerFile": 90,
    "maxTotalMinutes": 300
}
```

#### Output example (real result, shortened)

An 8-minute MP4 video with `"summarize": true`:

```json
{
    "url": "https://archive.org/download/gerald-ford-inaugural-address-august-9-1974-720p/Gerald%20Ford%20inaugural%20address_%20August%209%2C%201974%20%28720p%29.mp4",
    "success": true,
    "format": "mp4",
    "fileSizeBytes": 48964530,
    "durationSeconds": 500.1,
    "billedMinutes": 9,
    "language": "en",
    "wordCount": 876,
    "text": "Mr. Chief Justice, my dear friends, my fellow Americans, the oath that I have taken is the same oath that was taken by George Washington...",
    "segments": [
        { "start": 1.36, "end": 20.4, "text": "Mr. Chief Justice, my dear friends, my fellow Americans, the oath that I have taken is the same oath that was taken by George Washington and by every president under the Constitution." },
        { "start": 20.4, "end": 31.32, "text": "But I assume the presidency under extraordinary circumstances never before experienced by Americans." }
    ],
    "paragraphs": [
        { "start": 1.36, "end": 39.68, "text": "Mr. Chief Justice, my dear friends, ... This is an hour of history that troubles our minds and hurts our hearts." },
        { "start": 41.44, "end": 73.64, "text": "Therefore, I feel it is my first duty to make an unprecedented compact with my countrymen. ..." }
    ],
    "summary": "The president assumes office under extraordinary circumstances ... and asks for prayers for Richard Nixon and his family.",
    "chapters": [
        { "startSeconds": 0, "start": "0:00", "title": "Introduction and Compact with the Nation" },
        { "startSeconds": 304, "start": "5:04", "title": "Message of Hope and Unity" }
    ],
    "srtUrl": "https://api.apify.com/v2/key-value-stores/.../records/0000-Gerald-20Ford-...mp4.srt",
    "vttUrl": "https://api.apify.com/v2/key-value-stores/.../records/0000-Gerald-20Ford-...mp4.vtt",
    "processingSeconds": 33
}
```

With `"speakerLabels": true` (a real 74-minute Changelog & Friends episode: 3 speakers found, 13,069 words, 195 seconds of processing, 74 minutes billed at $0.015):

```json
{
    "success": true,
    "model": "nova-3",
    "durationSeconds": 4390.7,
    "billedMinutes": 74,
    "speakers": 3,
    "segments": [
        { "start": 16.14, "end": 23.34, "speaker": 1, "text": "Welcome to Changelog and Friends, a weekly talk show about thinking outside the Dropbox." }
    ],
    "paragraphs": [
        { "start": 16.14, "end": 34.87, "speaker": 1, "text": "Welcome to Changelog and Friends, a weekly talk show about thinking outside the Dropbox. Thanks as always to our partners at fly.io, ..." },
        { "start": 43.38, "end": 69.23, "speaker": 2, "text": "This is the year we almost break the database. Let me explain. Where do agents actually store their stuff? ..." }
    ],
    "speakerTranscript": "Speaker 1 [0:16]: Welcome to Changelog and Friends, ...\n\nSpeaker 2 [0:43]: This is the year we almost break the database. ..."
}
```

With `"wordTimestamps": true`, every segment also gets its words (real result, a 5-minute news episode):

```json
{ "start": 4.16, "end": 7.52, "speaker": 1, "text": "What's up, friends? Adam here. Big week here at ChangeLog.",
  "words": [{ "word": "What's", "start": 4.16, "end": 4.4, "speaker": 1 }, { "word": "up,", "start": 4.4, "end": 4.56, "speaker": 1 }, { "word": "friends?", "start": 4.56, "end": 4.96, "speaker": 1 }] }
```

In the SRT file a line starts with `Speaker 2: ` when the speaker changes; the VTT file uses `<v Speaker 2>` voice tags. Files longer than 50 minutes are sent in overlapping pieces, and the speakers of each piece are matched by the words in the overlap, so Speaker 1 stays Speaker 1 for the whole file.

A podcast episode from an RSS feed also carries the episode details:

```json
{
    "url": "https://tracking.swap.fm/track/.../default.mp3?...",
    "podcastTitle": "Planet Money",
    "episodeTitle": "Middlegarchs are the new Oligarchs",
    "pubDate": "2026-09-25T21:46:44.000Z",
    "episodeUrl": "https://www.npr.org/2026/09/25/nx-s1-5981194/stealthy-wealthy-everywhere-millionaires-pass-through",
    "feedUrl": "https://feeds.npr.org/510289/podcast.xml",
    "success": true,
    "format": "mp3",
    "durationSeconds": 2065.2,
    "billedMinutes": 35,
    "wordCount": 5505
}
```

#### Publisher transcripts (no audio minutes)

Many podcasts put their own transcript in the feed (`<podcast:transcript>`, Podcasting 2.0). With `usePublisherTranscripts` on (the default), such an episode is not transcribed: the publisher's file is downloaded and converted to the same output. Timed formats (VTT, SRT, JSON) give segments with timestamps, paragraphs and our SRT/VTT files; HTML and plain text give the text and paragraphs only (`timestamps: false`, no subtitle files). Speaker names from the transcript are kept as `speakerName` on each segment. No audio minutes are charged (`billedMinutes: 0`); the AI summary, if you turn it on, is charged as usual.

The transcript is checked first: it must have at least 20 words, cover at least half of the episode's duration (from the feed) without running far past it, and have a plausible number of words per minute. A transcript that fails the check, or cannot be downloaded, is ignored and the episode is transcribed normally; the row then says why in `publisherTranscriptError`. Set `"usePublisherTranscripts": false` to always transcribe the audio yourself.

Real result (Podnews Daily, whose feed has a VTT transcript per episode; with `"summarize": true`, so only the $0.01 summary was charged):

```json
{
    "podcastTitle": "Podnews Daily - podcast industry and podcasting news",
    "episodeTitle": "Was Inception Point AI just a PR stunt?",
    "pubDate": "2026-09-29T10:30:00.000Z",
    "success": true,
    "transcriptSource": "publisher",
    "transcriptUrl": "https://podnews.net/audio/podnews260929.mp3.vtt",
    "transcriptFormat": "vtt",
    "timestamps": true,
    "model": "publisher",
    "language": "en",
    "durationSeconds": 273,
    "billedMinutes": 0,
    "wordCount": 642,
    "segments": [
        { "start": 0.56, "end": 4.96, "text": "From Singapore Airport, the latest from podnews.net with The Podglomerate.", "speakerName": "James Cridland" }
    ],
    "summary": "Inception Point AI's CEO Jeanine Wright has suggested that the company's thousands of AI-made podcasts may have been perceived as a publicity stunt, ...",
    "summaryCharged": true
}
```

In the same test, a 63-minute Buzzcast episode (12,067 words, 810 segments, 10 chapters) was also taken from the publisher's transcript: 0 audio minutes instead of 64.

#### Feed filters

```json
{
    "podcastFeeds": ["https://podnews.net/rss"],
    "episodesSince": "2026-09-20",
    "episodesUntil": "2026-09-28",
    "titleIncludes": ["patreon", "spotify*video"],
    "titleExcludes": ["weekly"],
    "maxEpisodesPerFeed": 1
}
```

- `episodesSince`: a date (`YYYY-MM-DD`) or a relative time such as `30 days`, `2 weeks`, `6 months`. `episodesUntil`: a date; the whole day counts.
- `titleIncludes`: keep an episode when its title contains one of the texts; `titleExcludes`: drop it when its title contains one. Case does not matter; `*` matches anything (`spotify*video`).
- Filters are applied first, then `newEpisodesOnly`, then `maxEpisodesPerFeed`. With a date filter, episodes without a publish date are dropped. A link to one Apple Podcasts episode is not filtered.
- Filtered episodes get **no row** and are not charged; `SUMMARY.episodesFiltered` counts them. In our test above, 150 of 151 episodes were filtered and the one match ("Patreon removes per-creation memberships", 4.6 minutes) was transcribed.

#### Web page with one player

```json
{ "urls": ["https://archive.org/details/nasa_tv-ScienceCast_-_Record-Setting_Asteroid_Flyby"] }
```

Real result (shortened): the page's `og:video` file was found and transcribed, 262 seconds, 5 minutes billed:

```json
{
    "url": "https://archive.org/details/nasa_tv-ScienceCast_-_Record-Setting_Asteroid_Flyby",
    "mediaUrl": "https://archive.org/download/nasa_tv-ScienceCast_-_Record-Setting_Asteroid_Flyby/ScienceCast_-_Record-Setting_Asteroid_Flyby.mp4",
    "success": true,
    "transcriptSource": "asr",
    "format": "mp4",
    "billedMinutes": 5,
    "segments": [{ "start": 3.7, "end": 8.48, "text": "Record-setting asteroid flyby, presented by Science at NASA." }]
}
```

#### Results that are not charged

Every row has `success: true` or `false`. Rows with `success: false` are free and carry an `errorType`:

| errorType | Meaning |
|---|---|
| `no_text` | No speech in the audio (silence or music only); the row also has `noSpeech: true`. This is a valid answer, not a failure: it does not make the run `PARTIAL_RESULTS` and is counted in `SUMMARY.noSpeech` |
| `too_large` | Longer than `maxMinutesPerFile`, or larger than 2 GB |
| `limit_reached` | Would go over `maxTotalMinutes` or your Apify spending limit |
| `unsupported` | Not an audio or video file (e.g. a web page), no audio track, or a language that speaker labels do not support (e.g. Chinese) |
| `unsupported_platform` | A YouTube, TikTok, Instagram, Facebook, X, Vimeo, Loom, SoundCloud or Spotify page link (not a file); the error says how to get the file |
| `no_media` | A web page without an audio or video file we can use (no player, or only a streaming playlist or player page) |
| `multiple_media` | A web page with several audio or video files; they are listed in `mediaCandidates` |
| `invalid_input` | A line that is not a URL |
| `not_found`, `blocked`, `network`, `timeout`, `http_error` | The file could not be downloaded |

The run's key-value store also has a `SUMMARY` record with `status` (`SUCCESS`, `PARTIAL_RESULTS`, `FAILED`, `NO_RESULTS` or `LIMIT_REACHED`), files transcribed, files without speech (`noSpeech`, counted as successes), minutes billed and more.

### Use with AI agents (MCP)

Connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/audio-transcriber) to Claude, Cursor or any MCP client, then ask e.g. "Transcribe the last 2 episodes of this podcast feed and give me the chapters".

Minimal input:

```json
{ "urls": ["https://example.com/episode.mp3"] }
```

Failed or skipped items are not charged and carry an `errorType`, so an agent can see at a glance what worked.

#### Scheduled podcast transcription (only new episodes)

```json
{
    "podcastFeeds": ["https://changelog.com/news/feed"],
    "maxEpisodesPerFeed": 1,
    "newEpisodesOnly": true,
    "monitorName": "changelog-news"
}
```

Run it on an Apify schedule (e.g. weekly). The first run transcribes the newest episode; each later run transcribes only episodes it has not transcribed before (in our test, the second run skipped "Bitwarden CLI compromised" and took the next one). Runs with the same `monitorName` share the list, which is kept in a named key-value store (`audio-transcriber-<monitorName>`). When there is nothing new, the run ends without charges. Failed files are not remembered, so they are tried again next time.

### Chaining with other Actors

Set `datasetId` to the dataset of another run (for example a podcast or video scraper). The media URL field is found automatically (`audioUrl`, `mediaUrl`, `videoUrl`, `url`, ...), or set it with `datasetField`, e.g. `media.audioUrl`.

**Subtitles in other languages:** pass the `srtUrl` or `vttUrl` of a result to our [Bulk Text & JSON Translator](https://apify.com/tidytools/web-page-translator) (`subtitleUrls`). It translates the subtitles cue by cue, keeps every timestamp, and returns new SRT/VTT files in any language, at $1.50 per million characters. Speech has about 1,000 characters per minute, so that is roughly $0.0015 per audio minute and language (in our test, 20 minutes of subtitles into Spanish and Traditional Chinese cost $0.059).

### Limitations

- **Transcription only**: the turbo model does not translate. Use our *Bulk Text & JSON Translator* on the text or the subtitle files if you need another language.
- Speaker labels are numbers (Speaker 1, 2…), not names. Short interjections ("Yeah.") are sometimes given to the wrong speaker, and speaker labels are not available for every language (Chinese is not). When a speaker change is detected a few words late ("…Welcome to | the podcast. Thank you"), those words are moved back to the end of the previous speaker's sentence.
- Segments that Whisper marks as probably not speech (its no-speech probability) are dropped, together with typical filler it invents over music or silence ("Thanks for watching!"); the count is in `segmentsFiltered`.
- Podcast hosts often insert ads when the file is downloaded; those minutes are part of the file and are billed. `maxMinutesPerFile` uses the duration from the feed for podcast episodes.
- Long recordings are cut into pieces of about 90 seconds at the quietest moment nearby; a word at a cut can occasionally be lost. Whisper sometimes skips several seconds of speech inside a piece: every stretch of 5 seconds or more without text is transcribed again on its own, and what is found there is added (`SUMMARY.segmentsRecovered`; no extra charge).
- Publisher transcripts are the publisher's own text: their wording, timing and speaker names are used as they are (only checked for being complete and matching the episode length). They describe the publisher's version of the file, which may differ slightly from the file you download when ads are inserted.
- Uploaded files are read from your key-value store by the Actor; if a record cannot be read (errorType `blocked`), upload it again with the file upload field or pass a public or signed link to the file.
- Accuracy depends on audio quality, accents and background noise.

### FAQ

**Which speech to text model does it use?** OpenAI's open Whisper large-v3-turbo model, with automatic language detection and 90+ languages. Speaker labels, when turned on, come from nova-3.

**How much does audio to text cost?** $0.006 per minute ($0.36 per hour), no start fee. SRT and VTT subtitles, paragraphs and timestamps are included; files with no speech are free.

**Can it transcribe new podcast episodes automatically?** Yes. Paste the podcast RSS feed, turn on `newEpisodesOnly` and run it on a schedule: only episodes released since the last run are transcribed, so a day without a new episode costs nothing.

**Can I turn a video into text and subtitles?** Yes. MP4, MOV, MKV, WEBM and more: the audio track is extracted automatically, and every file gets text, timestamps and SRT/VTT subtitle files.

### Support

Open an issue in the **Issues** tab with the audio URL. Issues are checked regularly.

# Actor input Schema

## `urls` (type: `array`):

Main input (fill this, or `audioFiles` / `podcastFeeds` / `audioUrls` / `datasetId`). Links to audio or video files (MP3, WAV, M4A, AAC, OGG, OPUS, FLAC, MP4, MOV, MKV, WEBM...), up to 2 GB each, one per line. Public Google Drive, Dropbox and OneDrive share links work, and so does a web page with exactly one audio or video player. YouTube, TikTok, Instagram, Facebook, X, Vimeo, Loom, SoundCloud or Spotify page links are not files: they return a free error row explaining how to get the file.

## `audioFiles` (type: `array`):

Upload recordings from your computer in Apify Console (they are stored in a key-value store of your account and read from there). Same formats and 2 GB limit as the URLs. In API calls, pass file URLs in "urls" instead.

## `podcastFeeds` (type: `array`):

Podcast RSS feed URLs or Apple Podcasts show links. The newest episodes are transcribed (see the next field). Each result includes the episode title and publish date.

## `audioUrls` (type: `array`):

Same as "Audio or video file URLs", in Apify's request list format (also accepts a link to a text file with URLs). Kept for existing integrations; new integrations should use "urls".

## `maxEpisodesPerFeed` (type: `integer`):

How many of the newest episodes to transcribe from each feed.

## `usePublisherTranscripts` (type: `boolean`):

Many podcasts publish their own transcript in the feed (Podcasting 2.0 <podcast:transcript>: VTT, SRT, JSON, HTML or text). When an episode has one, it is downloaded and converted to the same output (text, segments, SRT/VTT for timed formats) with transcriptSource "publisher", and no audio minutes are charged. A transcript that is empty or does not match the episode length is ignored and the audio is transcribed as usual.

## `episodesSince` (type: `string`):

Only feed episodes published on or after this date (YYYY-MM-DD), or within e.g. "30 days". Applied before "Episodes per podcast feed". Filtered episodes get no row; SUMMARY.episodesFiltered counts them.

## `episodesUntil` (type: `string`):

Only feed episodes published on or before this date (YYYY-MM-DD, the whole day counts).

## `titleIncludes` (type: `array`):

Only feed episodes whose title contains one of these texts (case-insensitive). \* matches anything, e.g. "interview\*ceo".

## `titleExcludes` (type: `array`):

Skip feed episodes whose title contains one of these texts (case-insensitive), e.g. "trailer", "rerun", "best of".

## `language` (type: `string`):

ISO 639-1 code such as en, es, de, ja or zh. Leave empty to detect automatically.

## `vocabulary` (type: `string`):

Names, brands or terms that appear in the audio, e.g. "Apify, Cloudflare, Kubernetes". Helps spell them correctly.

## `saveSubtitles` (type: `boolean`):

Also save subtitle files for every transcript (no extra charge).

## `summarize` (type: `boolean`):

Add a short summary and timestamped chapters to each transcript. Charged as a separate event per file (see Pricing). Recordings shorter than 2 minutes get a summary only.

## `speakerLabels` (type: `boolean`):

Label each segment and paragraph with Speaker 1, Speaker 2… and count the speakers. Uses Deepgram nova-3 instead of Whisper and is billed as "audio minute with speaker labels" ($0.015/min) instead of the regular audio minute. Some languages (e.g. Chinese) are not available; those files fail with errorType "unsupported" and are not charged.

## `wordTimestamps` (type: `boolean`):

Add a words array \[{word, start, end}] to every segment (for video editing and precise subtitle alignment). Included in the regular $0.006/min price. With speaker labels on, the words come from nova-3 and the speaker-label price applies (not charged twice).

## `subtitleLanguages` (type: `array`):

Also make SRT + VTT subtitles in these languages (up to 5), e.g. Spanish, German, Traditional Chinese, ja. Timings and speaker labels are kept. $0.002 per audio minute per language, charged only for languages that succeed.

## `datasetId` (type: `string`):

A dataset from another Actor run, e.g. a podcast or video scraper. Its media URLs are added to the list above.

## `datasetField` (type: `string`):

Field that holds the media URL, e.g. audioUrl or media.url. Leave empty to detect it (audioUrl, mediaUrl, videoUrl, url...).

## `maxMinutesPerFile` (type: `integer`):

Skip files longer than this. Skipped files get a result row with errorType "too\_large" and are not charged. Empty = no limit.

## `maxTotalMinutes` (type: `integer`):

Stop charging for new files once this many audio minutes were transcribed. Files that would go over it get a row with errorType "limit\_reached" and are not charged. Empty = no limit (your Apify spending limit still applies).

## `maxConcurrency` (type: `integer`):

Long files are split into 90-second pieces that are transcribed in parallel. More is faster.

## `fileConcurrency` (type: `integer`):

How many files (podcast episodes, URLs) are downloaded and transcribed at the same time. More is faster for batches; rows then arrive in the order files finish (inputIndex gives the input position). Your max charge per run and maxTotalMinutes are still respected: the cost of every file in progress is reserved before it is transcribed. With little run memory fewer files run at once (about 3 at 1024 MB).

## `fileTimeoutSecs` (type: `integer`):

A file whose download, conversion and transcription take longer than this is reported with errorType "timeout" and not charged; the run goes on with the next file. Raise it for recordings of several hours. 0 = no limit.

## `proxyConfiguration` (type: `object`):

Only used for requests sent from Apify's network (file downloads and podcast feeds). Apify proxy usage is billed to your Apify account.

## `newEpisodesOnly` (type: `boolean`):

Remember what was transcribed (podcast episode GUID or file URL) and skip it in later runs with the same monitor name. From podcast feeds, only episodes published after the newest one already transcribed are taken, so a run on a day without a new episode transcribes (and charges) nothing. Failed files are retried next time.

## `backfillOlderEpisodes` (type: `boolean`):

With "Only new episodes": also take older episodes that were never transcribed (up to Max episodes per feed per run), to work through a back catalogue over several runs. Off by default, so scheduled runs only pick up new releases.

## `monitorName` (type: `string`):

Runs with the same name share the list of transcribed files, e.g. one name per podcast schedule. Letters, digits and dashes.

## Actor input object example

```json
{
  "urls": [
    "https://webcapture-api.yukailin.workers.dev/samples/speech-sample.wav"
  ],
  "maxEpisodesPerFeed": 3,
  "usePublisherTranscripts": true,
  "saveSubtitles": true,
  "summarize": false,
  "speakerLabels": false,
  "wordTimestamps": false,
  "subtitleLanguages": [],
  "maxConcurrency": 3,
  "fileConcurrency": 3,
  "fileTimeoutSecs": 1800,
  "newEpisodesOnly": false,
  "backfillOlderEpisodes": false
}
```

# Actor output Schema

## `transcripts` (type: `string`):

No description

## `full` (type: `string`):

No description

## `subtitles` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://webcapture-api.yukailin.workers.dev/samples/speech-sample.wav"
    ],
    "subtitleLanguages": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("tidytools/audio-transcriber").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://webcapture-api.yukailin.workers.dev/samples/speech-sample.wav"],
    "subtitleLanguages": [],
}

# Run the Actor and wait for it to finish
run = client.actor("tidytools/audio-transcriber").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://webcapture-api.yukailin.workers.dev/samples/speech-sample.wav"
  ],
  "subtitleLanguages": []
}' |
apify call tidytools/audio-transcriber --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tidytools/audio-transcriber"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3W2KtIbvHitDH8GCc/builds/TnRaxI672XBp7vbES/openapi.json
