Whisper Transcriber avatar

Whisper Transcriber

Pricing

$20.00 / 1,000 transcribed minutes

Go to Apify Store
Whisper Transcriber

Whisper Transcriber

Transcribes audio and video from direct media URLs using faster-whisper large-v3 on a GPU, with optional speaker diarization. Returns plain text, SRT, VTT and speaker-labelled segments, priced per transcribed minute.

Pricing

$20.00 / 1,000 transcribed minutes

Rating

0.0

(0)

Developer

Adam Schepis

Adam Schepis

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Transcribes audio and video from direct media URLs using faster-whisper large-v3 on a GPU, with optional speaker diarization. Returns plain text, SRT, VTT and speaker-labelled segments. Priced per transcribed minute, with a 3-minute minimum per run.

This Actor accepts a direct link to an audio or video file: a podcast enclosure MP3, or a hosted MP4, WAV, M4A, FLAC or OGG.

It does not download from YouTube, Vimeo, SoundCloud, Spotify, TikTok, Instagram, Apple Podcasts or any other streaming site, and it never will. Pulling media off those platforms breaks their terms of service. Passing one of those links returns a clear error instead of a transcript, and costs you nothing.

If you have a YouTube video you own, export the audio or video file, host it somewhere reachable, and pass that URL.

What you get

  • large-v3, not tiny. Most transcription Actors run a small Whisper model on CPU. This one runs the full faster-whisper large-v3 model on a GPU, so accents, crosstalk and technical vocabulary survive.
  • Speaker diarization. Turn it on and every segment is labelled SPEAKER_00, SPEAKER_01, … — usable for interviews, panels and meetings.
  • Four output shapes from one run: plain text, SRT subtitles, WebVTT captions, and timed segments with speaker labels.
  • Pay per transcribed minute, rounded up, with a 3-minute minimum per run. A file that fails to download or fails to transcribe is not charged.

Input

FieldTypeRequiredDescription
mediaUrlsarray of URLsyesDirect links to audio/video files. Streaming-site links are rejected.
languagestringno (default auto)ISO 639-1 code, or auto to detect. Naming the language is faster and more accurate on short or noisy clips.
diarizationbooleanno (default false)Label who is speaking. Adds word-level alignment, so runs take noticeably longer.
maxAudioMinutesintegerno (default 120)Per-file safety cap. A file needing longer is abandoned and not charged.
maxFileSizeMbintegerno (default 500)Files above this are skipped before any GPU time is spent.
maxCostUsdnumberno (default 10)Hard ceiling on what the run may charge.

Example input:

{
"mediaUrls": [
{ "url": "https://cdn.example.com/podcast/episode-42.mp3" }
],
"language": "en",
"diarization": true,
"maxCostUsd": 5
}

Output example

One record per media URL:

{
"url": "https://cdn.example.com/podcast/episode-42.mp3",
"status": "ok",
"model": "faster-whisper large-v3 (WhisperX)",
"language": "en",
"languageProbability": 0.9971,
"durationSeconds": 11.86,
"billedMinutes": 1,
"diarization": true,
"speakers": ["SPEAKER_00", "SPEAKER_01"],
"transcript": "SPEAKER_00: Welcome back to the show, today we are talking about margins.\n\nSPEAKER_01: Thanks for having me, it is a topic I think about a lot.",
"srt": "1\n00:00:00,031 --> 00:00:03,512\n[SPEAKER_00] Welcome back to the show, today we are talking about margins.\n",
"vtt": "WEBVTT\n\n00:00:00.031 --> 00:00:03.512\n[SPEAKER_00] Welcome back to the show, today we are talking about margins.\n",
"segments": [
{ "start": 0.031, "end": 3.512, "text": "Welcome back to the show, today we are talking about margins.", "speaker": "SPEAKER_00" }
],
"gpuSeconds": 2.41,
"queueSeconds": 1.8,
"error": null,
"errorMessage": null,
"scrapedAt": "2026-09-16T12:00:00.000Z"
}

billedMinutes is what that one file contributed; a run that transcribed anything is billed at least 3 minutes in total.

A rejected or failed file produces the same record with status: "error", an error code (unsupported-source, not-direct-media, file-too-large, invalid-url, unreachable, or a backend job status) and an errorMessage explaining it. Download results as JSON, CSV, Excel, or via the Apify API.

Pricing

This Actor uses pay-per-event pricing. You are charged for the transcribed-minute event, once per minute of transcribed audio, rounded up per file. Runs are billed a 3-minute minimum. The exact per-event price is shown on the Actor's Store page before you run it.

EventWhen it's charged
transcribed-minuteOnce per minute of audio, after the transcript is saved to your dataset. A run that transcribed at least one file is billed at least 3 of them.

What that means in practice:

  • Nothing is charged until the transcript is in your dataset. If the download fails, the backend errors, or the job runs past maxAudioMinutes, you get an error record and no charge. A run where every file failed is billed nothing at all — the 3-minute minimum does not apply to it.
  • Runs are billed a 3-minute minimum. Starting a GPU worker costs real money whatever the clip length, so a run totalling less than 3 minutes of audio is billed as 3. The minimum is per run, not per file: a run of three 1-minute clips is billed 3 minutes, the same as a single 3-minute file. Batch short clips into one run rather than running each on its own.
  • Minutes are measured to the end of the last spoken segment, so trailing silence is not billed.
  • Each file rounds up independently. Five 30-second clips are billed 5 minutes; one 4-minute file is billed 4.
  • maxCostUsd and Apify's own Maximum cost per run both cap a run, minimum included: neither is ever exceeded to reach the 3-minute floor.

What the buyer should know

  • Formats: anything FFmpeg decodes — MP3, M4A, M4B, WAV, FLAC, OGG/OGA, OPUS, AAC, MP4, MOV, MKV, WEBM, AVI, 3GP, AMR. Video is accepted; only the audio track is used.
  • Max file size: 500 MB by default, adjustable to 2 GB via maxFileSizeMb. Oversized files are skipped before any GPU time is spent.
  • Max length: maxAudioMinutes defaults to 120 and accepts up to 300. The backend also enforces its own 15-minute GPU execution ceiling per file, which is ample for multi-hour audio at large-v3's speed but will cut off a pathological file.
  • Languages: large-v3 handles roughly 100 languages. The language dropdown lists the common ones; detection is automatic by default. Transcription is in the spoken language — this Actor does not translate.
  • Cold starts: the GPU backend scales to zero when idle, so the first file in a run can wait roughly 20–60 seconds for a worker. queueSeconds in each record reports the actual wait. Subsequent files in the same run reuse the warm worker.
  • Files are processed one at a time, in the order given, so a long list takes proportionally longer.
  • The URL must be reachable by the transcription backend, which fetches the file itself with a plain HTTP client. Hosts that block non-browser user agents (Wikimedia Commons is one) will fail even though the link looks fine in a browser. Podcast CDNs, S3/R2/GCS buckets, and GitHub raw links all work.
  • Diarization labels speakers as SPEAKER_00, SPEAKER_01, …; it does not identify them by name.

Why this Actor

The transcription Actors already on Store run Whisper tiny, base or small on CPU and still charge $0.015–$0.048 per minute. This one runs the full large-v3 model on a GPU and adds speaker diarization, at a price in the middle of that range. You get materially better transcripts for the same money, with no API key to manage and no GPU to rent.

Notes

  • Transcription runs on a dedicated GPU serverless endpoint (NVIDIA 24 GB class) running WhisperX with faster-whisper large-v3. Word-level alignment is enabled only when diarization is on, since that is the only thing that needs it.
  • Built with the Apify SDK for JavaScript.

Reference docs used to build this Actor