Whisper Transcriber
Pricing
$20.00 / 1,000 transcribed minutes
Whisper Transcriber
Transcribes audio and video from direct media URLs using faster-whisper large-v3 on a GPU, with optional speaker diarization. Returns plain text, SRT, VTT and speaker-labelled segments, priced per transcribed minute.
Pricing
$20.00 / 1,000 transcribed minutes
Rating
0.0
(0)
Developer
Adam Schepis
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Share
Transcribes audio and video from direct media URLs using faster-whisper large-v3 on a GPU, with optional speaker diarization. Returns plain text, SRT, VTT and speaker-labelled segments. Priced per transcribed minute, with a 3-minute minimum per run.
Direct media links only — no YouTube
This Actor accepts a direct link to an audio or video file: a podcast enclosure MP3, or a hosted MP4, WAV, M4A, FLAC or OGG.
It does not download from YouTube, Vimeo, SoundCloud, Spotify, TikTok, Instagram, Apple Podcasts or any other streaming site, and it never will. Pulling media off those platforms breaks their terms of service. Passing one of those links returns a clear error instead of a transcript, and costs you nothing.
If you have a YouTube video you own, export the audio or video file, host it somewhere reachable, and pass that URL.
What you get
- large-v3, not tiny. Most transcription Actors run a small Whisper model on CPU. This one runs the full
faster-whisper large-v3model on a GPU, so accents, crosstalk and technical vocabulary survive. - Speaker diarization. Turn it on and every segment is labelled
SPEAKER_00,SPEAKER_01, … — usable for interviews, panels and meetings. - Four output shapes from one run: plain text, SRT subtitles, WebVTT captions, and timed segments with speaker labels.
- Pay per transcribed minute, rounded up, with a 3-minute minimum per run. A file that fails to download or fails to transcribe is not charged.
Input
| Field | Type | Required | Description |
|---|---|---|---|
mediaUrls | array of URLs | yes | Direct links to audio/video files. Streaming-site links are rejected. |
language | string | no (default auto) | ISO 639-1 code, or auto to detect. Naming the language is faster and more accurate on short or noisy clips. |
diarization | boolean | no (default false) | Label who is speaking. Adds word-level alignment, so runs take noticeably longer. |
maxAudioMinutes | integer | no (default 120) | Per-file safety cap. A file needing longer is abandoned and not charged. |
maxFileSizeMb | integer | no (default 500) | Files above this are skipped before any GPU time is spent. |
maxCostUsd | number | no (default 10) | Hard ceiling on what the run may charge. |
Example input:
{"mediaUrls": [{ "url": "https://cdn.example.com/podcast/episode-42.mp3" }],"language": "en","diarization": true,"maxCostUsd": 5}
Output example
One record per media URL:
{"url": "https://cdn.example.com/podcast/episode-42.mp3","status": "ok","model": "faster-whisper large-v3 (WhisperX)","language": "en","languageProbability": 0.9971,"durationSeconds": 11.86,"billedMinutes": 1,"diarization": true,"speakers": ["SPEAKER_00", "SPEAKER_01"],"transcript": "SPEAKER_00: Welcome back to the show, today we are talking about margins.\n\nSPEAKER_01: Thanks for having me, it is a topic I think about a lot.","srt": "1\n00:00:00,031 --> 00:00:03,512\n[SPEAKER_00] Welcome back to the show, today we are talking about margins.\n","vtt": "WEBVTT\n\n00:00:00.031 --> 00:00:03.512\n[SPEAKER_00] Welcome back to the show, today we are talking about margins.\n","segments": [{ "start": 0.031, "end": 3.512, "text": "Welcome back to the show, today we are talking about margins.", "speaker": "SPEAKER_00" }],"gpuSeconds": 2.41,"queueSeconds": 1.8,"error": null,"errorMessage": null,"scrapedAt": "2026-09-16T12:00:00.000Z"}
billedMinutes is what that one file contributed; a run that transcribed anything is billed at least 3 minutes in total.
A rejected or failed file produces the same record with status: "error", an error code (unsupported-source, not-direct-media, file-too-large, invalid-url, unreachable, or a backend job status) and an errorMessage explaining it. Download results as JSON, CSV, Excel, or via the Apify API.
Pricing
This Actor uses pay-per-event pricing. You are charged for the transcribed-minute event, once per minute of transcribed audio, rounded up per file. Runs are billed a 3-minute minimum. The exact per-event price is shown on the Actor's Store page before you run it.
| Event | When it's charged |
|---|---|
transcribed-minute | Once per minute of audio, after the transcript is saved to your dataset. A run that transcribed at least one file is billed at least 3 of them. |
What that means in practice:
- Nothing is charged until the transcript is in your dataset. If the download fails, the backend errors, or the job runs past
maxAudioMinutes, you get an error record and no charge. A run where every file failed is billed nothing at all — the 3-minute minimum does not apply to it. - Runs are billed a 3-minute minimum. Starting a GPU worker costs real money whatever the clip length, so a run totalling less than 3 minutes of audio is billed as 3. The minimum is per run, not per file: a run of three 1-minute clips is billed 3 minutes, the same as a single 3-minute file. Batch short clips into one run rather than running each on its own.
- Minutes are measured to the end of the last spoken segment, so trailing silence is not billed.
- Each file rounds up independently. Five 30-second clips are billed 5 minutes; one 4-minute file is billed 4.
maxCostUsdand Apify's own Maximum cost per run both cap a run, minimum included: neither is ever exceeded to reach the 3-minute floor.
What the buyer should know
- Formats: anything FFmpeg decodes — MP3, M4A, M4B, WAV, FLAC, OGG/OGA, OPUS, AAC, MP4, MOV, MKV, WEBM, AVI, 3GP, AMR. Video is accepted; only the audio track is used.
- Max file size: 500 MB by default, adjustable to 2 GB via
maxFileSizeMb. Oversized files are skipped before any GPU time is spent. - Max length:
maxAudioMinutesdefaults to 120 and accepts up to 300. The backend also enforces its own 15-minute GPU execution ceiling per file, which is ample for multi-hour audio at large-v3's speed but will cut off a pathological file. - Languages: large-v3 handles roughly 100 languages. The
languagedropdown lists the common ones; detection is automatic by default. Transcription is in the spoken language — this Actor does not translate. - Cold starts: the GPU backend scales to zero when idle, so the first file in a run can wait roughly 20–60 seconds for a worker.
queueSecondsin each record reports the actual wait. Subsequent files in the same run reuse the warm worker. - Files are processed one at a time, in the order given, so a long list takes proportionally longer.
- The URL must be reachable by the transcription backend, which fetches the file itself with a plain HTTP client. Hosts that block non-browser user agents (Wikimedia Commons is one) will fail even though the link looks fine in a browser. Podcast CDNs, S3/R2/GCS buckets, and GitHub raw links all work.
- Diarization labels speakers as
SPEAKER_00,SPEAKER_01, …; it does not identify them by name.
Why this Actor
The transcription Actors already on Store run Whisper tiny, base or small on CPU and still charge $0.015–$0.048 per minute. This one runs the full large-v3 model on a GPU and adds speaker diarization, at a price in the middle of that range. You get materially better transcripts for the same money, with no API key to manage and no GPU to rent.
Notes
- Transcription runs on a dedicated GPU serverless endpoint (NVIDIA 24 GB class) running WhisperX with
faster-whisper large-v3. Word-level alignment is enabled only when diarization is on, since that is the only thing that needs it. - Built with the Apify SDK for JavaScript.
Reference docs used to build this Actor
- Apify SDK for JS: https://docs.apify.com/sdk/js/
- Pay-per-event monetization overview: https://docs.apify.com/platform/actors/publishing/monetize/pay-per-event
- Pay-per-event SDK guide (
Actor.charge,ChargingManager): https://docs.apify.com/sdk/js/docs/concepts/pay-per-event .actor/actor.jsonreference: https://docs.apify.com/platform/actors/development/actor-definition/actor-json- Runpod serverless job operations: https://docs.runpod.io/serverless/endpoints/job-operations
- WhisperX worker image: https://github.com/kodxana/whisperx-worker_v2