Whisper Transcriber โ€” Video Links & Files to Text avatar

Whisper Transcriber โ€” Video Links & Files to Text

Pricing

from $6.30 / 1,000 minute transcribeds

Go to Apify Store
Whisper Transcriber โ€” Video Links & Files to Text

Whisper Transcriber โ€” Video Links & Files to Text

Whisper large v3 transcription, no API key: paste a YouTube, TikTok, Vimeo or SoundCloud link, or an MP3, WAV, M4A, FLAC, OGG, MP4 or MOV file, and get clean text with timecodes in 90+ languages. Export data, run via API, schedule runs, or integrate with AI workflows.

Pricing

from $6.30 / 1,000 minute transcribeds

Rating

5.0

(2)

Developer

Matvey

Matvey

Maintained by Community

Actor stats

0

Bookmarked

24

Total users

20

Monthly active users

3 days ago

Last modified

Share

Whisper Transcriber runs Whisper large v3 on your audio and video and gives back clean text with timecodes โ€” in 90+ languages, with no API key to set up and no model to choose. MP3, M4A, WAV, FLAC, OGG, MP4, MOV, WEBM: anything with sound. Files of any length are split, transcribed and stitched back together, so a three-hour recording is one input and one row of output.

What it does

What you give itWhat you get back
An MP3 or WAV linkTranscript with timed passages
A video fileThe same โ€” the video track is discarded
A three-hour podcastOne transcript, timings continuous across parts
A recording in any languageText in that language, or translated to English
diarize: trueEvery passage labelled with who is speaking
subtitleFormat: srtA finished SRT or WebVTT file
chunkForRag: trueChunks with timecodes, ready to embed

Two recognition settings: accurate (Whisper large v3, the default here) and fast (Whisper large v3 turbo, several times quicker, same price per minute). A vocabulary hint helps with names and jargon the model has not met.

What data does it return?

FieldExample
source, fileNamehttps://example.com/episode-12.mp3 ยท episode-12.mp3
durationSeconds, durationMinutes461.05 ยท 7.68
language, model, translatedToEnglishEnglish ยท accurate ยท false
transcript, wordCount, charCountthe full text ยท 1665 ยท 9218
segments[{"start": 0, "duration": 4.56, "text": "โ€ฆ", "speaker": "Speaker 1"}]
speakerCount, speakers3 ยท ["Speaker 1", "Speaker 2", "Speaker 3"]
chunks[{"index": 0, "startTimecode": "00:00:00", "text": "โ€ฆ"}]
subtitlesa complete SRT or WebVTT file
status, errorCode, errorMessageok, or why a file failed

How much does it cost?

EventPriceWhen it is charged
Minute transcribed$0.009Per started minute, using the built-in key
Minute with your own key$0.004Per started minute when you supply a Groq key
Add-on: Speaker labels$0.006Per started minute, only when speaker labels are on
File processed$0.002Per file downloaded and prepared
Video link resolved$0.09Per YouTube, TikTok, Instagram, X, Vimeo or other video page whose audio has to be pulled through a residential exit node. Direct file links never pay it

A failed file is returned as an error row and is never charged. Larger monthly plans get 10โ€“30 % off.

Bulk export

This Actor is built for bulk jobs โ€” put hundreds of links into one run, or call it from the API on a schedule. There is no fee per run: you pay per started minute of audio plus $0.002 per file, and $0.09 only for a video page (not a direct file link) whose audio has to be fetched through a residential exit node. Files that fail are returned as error rows and are free.

โฌ‡๏ธ Input

{
"urls": ["https://example.com/episode-12.mp3"],
"quality": "accurate",
"includeSegments": true
}

Speaker labels for a two-person interview:

{ "urls": ["https://example.com/interview.m4a"], "diarize": true, "speakerCount": 2 }

Every input field

FieldWhat it does
urlsPublic links to audio or video files. Several at a time is fine.
fileAn uploaded file instead of a link.
qualityfast for speed, accurate for difficult audio and non-English speech.
languageName the spoken language instead of letting the model guess โ€” faster and safer on short clips.
translateToEnglishWrite the text in English whatever was spoken.
diarize, speakerCountSplit the transcript by speaker. Give the count when you know it.
vocabularyHintNames, brands and jargon, so they come out spelled your way.
includeSegmentsTime-stamped passages alongside the full text.
subtitleFormatsrt or vtt to get a ready subtitle file.
apiKeyYour own Groq key โ€” cuts the per-minute price by more than half.
maxConcurrencyFiles processed in parallel.
maxFileSizeMb, timeoutPerFileSecsGuard rails for very large or very slow downloads.

โฌ†๏ธ Output

{
"source": "https://example.com/episode-12.mp3",
"fileName": "episode-12.mp3",
"durationMinutes": 7.68,
"language": "English",
"transcript": "Welcome back to the showโ€ฆ",
"wordCount": 1665,
"segments": [{ "start": 0, "duration": 4.56, "text": "Welcome back to the show" }],
"status": "ok"
}

Use cases

Whisper without the setup

No GPU, no model download, no ffmpeg pipeline. A URL in, a transcript out.

Multilingual archives

Whisper's strength is breadth: name the language or let it detect, and get the text in the original or translated to English.

Batch transcription

A list of URLs processed in parallel, one row per file, one dataset at the end.

Subtitles from the same pass

Ask for srt or vtt and the subtitle file is built from the timings Whisper already produced. Captions follow the usual reading norm: at most two lines of 42 characters each, split at commas and sentence ends.

Research and RAG pipelines

Time-stamped segments and optional chunks, so a spoken source can be cited as precisely as a written one.

๐Ÿค– For AI agents and LLM apps

Compact reference for agents calling this Actor through the Apify MCP server or the Apify API (lergassy/whisper-transcriber).

Purpose: run OpenAI Whisper large v3 on any audio or video URL and return the transcript as data.

Minimal input:

{ "urls": ["https://example.com/audio.mp3"] }

Behaviours an agent should know:

  • Billing is per started minute of audio, not per row. A 40-minute file costs the same whether you keep the segments or not.
  • durationMinutes on the row tells you what the run actually cost.
  • Speaker labels are an extra per-minute charge. Turn diarize on only when who-said-what matters.
  • Passing speakerCount when you know it makes the split noticeably cleaner.
  • language is detected automatically, but naming it removes the guess on short or noisy clips.
  • A file that cannot be downloaded or decoded returns a status: "error" row with the reason, free of charge โ€” check status before reading transcript.
  • Video files work: the audio track is extracted and the picture discarded.

Use via API

Run this Actor from your own code or pipeline. Get your token in Apify Console โ†’ Settings โ†’ API & Integrations. Every input field below has the same name as in the JSON tab of the input form.

Python

from apify_client import ApifyClient
client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("lergassy/whisper-transcriber").call(run_input={
"urls": [
"https://raw.githubusercontent.com/lergassy/apify-actor-assets/main/audio-samples/sample-1min.m4a"
],
"quality": "accurate",
"translateToEnglish": False,
"diarize": False,
"includeSegments": True,
"chunkForRag": False,
"chunkSize": 1200,
"chunkOverlapSeconds": 0,
"subtitleFormat": "none",
"maxConcurrency": 3,
"maxFileSizeMb": 500,
"timeoutPerFileSecs": 600
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)

JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });
const run = await client.actor('lergassy/whisper-transcriber').call({
"urls": [
"https://raw.githubusercontent.com/lergassy/apify-actor-assets/main/audio-samples/sample-1min.m4a"
],
"quality": "accurate",
"translateToEnglish": false,
"diarize": false,
"includeSegments": true,
"chunkForRag": false,
"chunkSize": 1200,
"chunkOverlapSeconds": 0,
"subtitleFormat": "none",
"maxConcurrency": 3,
"maxFileSizeMb": 500,
"timeoutPerFileSecs": 600
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

cURL (runs the Actor and returns the results in one call)

curl -X POST "https://api.apify.com/v2/acts/lergassy~whisper-transcriber/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \
-H "Content-Type: application/json" \
-d '{"urls": ["https://raw.githubusercontent.com/lergassy/apify-actor-assets/main/audio-samples/sample-1min.m4a"], "quality": "accurate", "translateToEnglish": false, "diarize": false, "includeSegments": true, "chunkForRag": false, "chunkSize": 1200, "chunkOverlapSeconds": 0, "subtitleFormat": "none", "maxConcurrency": 3, "maxFileSizeMb": 500, "timeoutPerFileSecs": 600}'

Limitations, stated plainly

  • Audio quality decides accuracy. Overlapping speech, heavy accents and phone-line compression cost accuracy no model recovers fully.
  • Speaker labels are labels, not identities. You get "Speaker 1" and "Speaker 2", never names โ€” mapping them to people is your step.
  • Timings are passage-level, not word-level.
  • Very large files need time. Raise timeoutPerFileSecs for multi-hour recordings, and expect a long run.
  • Links must be public. A file behind a login or an expiring signed URL cannot be fetched.

Integrations

Runs from the Apify API and the Python and JavaScript clients, and from n8n, Make or Zapier through the Apify app. Point a webhook at the dataset to push transcripts into Notion, Google Docs or your own database, or let an agent call it through the Apify MCP server.

โ“ FAQ

Do I need an API key or an account anywhere?

No. Recognition runs with a built-in key. Supply your own Groq key only if you want the cheaper per-minute rate.

Which languages are supported?

Around 90, including Russian, Indonesian, Spanish, German, French, Japanese and Chinese. Use accurate for non-English audio.

Can it handle a three-hour file?

Yes. Long files are processed in parts and the timings stay continuous across them.

Does it work with video?

Yes. Give it an MP4, MOV or WEBM link and the audio track is pulled out for you.

What happens if a file is unreachable?

You get an error row with the reason and no charge for that file. The other files in the run are still transcribed.

How accurate is it?

It runs Whisper large v3. On clean speech it is close to a careful human first pass; on overlapping or noisy audio it is not.

Can I get speaker names?

No โ€” you get "Speaker 1", "Speaker 2" and so on. Naming them is your step.

How do I make it spell product names correctly?

Put them in vocabularyHint. The hint is given to the model before it starts.

You might also like

ActorWhat it does
Speech to TextThe same engine, framed as a speech-to-text API
SRT Subtitle GeneratorReady subtitle files from audio or video
YouTube Transcript ScraperCaptions straight from YouTube, no transcription needed
Document Text ExtractorThe same idea for PDFs and Word files

Also known as

People look for this Actor as a Whisper API, Whisper large v3 transcription, OpenAI Whisper online, Whisper speech to text, audio transcription API and a Whisper batch transcriber.

Notes

The recognition key is built in โ€” there is nothing to sign up for. If you already pay for a Groq key, paste it and the per-minute price drops by 60 %.

Files come from public URLs or from an upload. Very large files are limited by maxFileSizeMb (500 MB by default) and each file by timeoutPerFileSecs.