Audio & Video to Text Transcriber: Whisper, SRT/VTT
Pricing
from $12.00 / 1,000 audio minute (tiny/base model)s
Audio & Video to Text Transcriber: Whisper, SRT/VTT
Audio & Video to Text Transcriber: Whisper, SRT/VTT (per minute)
Turn public audio and video files into clean text by URL, up to 20 files per run, with OpenAI Whisper (faster-whisper, int8 on CPU). Get the full transcript, timestamped segments, and ready-to-use SRT or VTT subtitle files. $0.012 per transcribed minute (tiny/base models) or $0.025 (small model), no start fee. Built for podcast/newsletter transcription, meeting and interview notes, subtitle generation, and RAG ingestion of audio/video content.
Use cases
- Transcribing podcasts, lectures, interviews and voicemails to text
- Generating SRT/VTT subtitles for videos
- Timestamped segment data for search, chaptering and RAG ingestion
- Meeting and call notes in an automation pipeline (mp3, m4a, wav, ogg, flac, mp4, webm, mov)
What you get
For each file, one dataset item:
status:ok,download_failed,unsupported_format,invalid_file, orno_speechdurationSec: seconds actually transcribed (long files are truncated, see below)language+languageProbability: the language Whisper detected (or your forced one)text: the full transcript (whentextis in the output formats)segments[]:{start, end, text}per speech segment (when requested)srt/vtt: complete subtitle files with correct timestamps (when requested)billedMinutes: whole minutes rounded up, charged asaudio-minuteeventstruncated: true when the file was longer than what was transcribedprocessingMs, anderrorwith a plain-language reason when something fails
The file type is detected from its content (magic bytes), never the URL extension: a
.mp3 that is really an MP4 is transcribed as an MP4. Audio longer than 180 minutes is
truncated and flagged truncated.
Pricing: you only pay for minutes that worked
One event per whole minute (rounded up) transcribed from a file that came back ok:
audio-minute ($0.012) for the tiny/base models, audio-minute-small ($0.025) for the
more accurate small model. A 90-second clip bills 2 minutes. Failed downloads, unsupported or invalid
files, and files with no speech (no_speech) are reported in the output but never
charged. There is no start fee. The run checks your spending limit before each file,
transcribes at most the minutes your remaining limit covers (the rest of the file is
skipped), and stops cleanly when the limit is reached.
Accuracy: Whisper on CPU
The engine is OpenAI Whisper running as faster-whisper in int8 on CPU — strong on clear speech in most languages, weaker on heavy music, overlap and strong accents. Model sizes: tiny (fastest), base (default, good balance), small (most accurate). Language is auto-detected by default; pass an ISO 639-1 code to force one.
Limits (by design)
- Max 20 URLs per run, max 200 MB per file, max 180 minutes transcribed per file (truncated and flagged beyond that).
- Supported inputs: mp3, m4a, wav, ogg, flac, mp4, webm, mov (audio track of the video).
Anything else is reported as
unsupported_formatand not charged. - Only public
http(s)links; local and private-network addresses are refused. - CPU inference on the base model runs roughly 0.5-1x realtime per core; a 60-minute file takes a while by design.
FAQ
Do I pay for files with no speech or that fail? No. Failed downloads, unsupported or
invalid files, and no_speech results are reported but never charged; you pay only for
whole minutes transcribed from files that came back ok, with no start fee.
What happens if a file is longer than my remaining limit? It is truncated to the minutes your limit covers, those minutes are billed, and the run stops cleanly.
Does it work on video? Yes — the audio track of mp4, webm or mov files is extracted and transcribed with the system ffmpeg.
Which languages are supported? Whatever Whisper supports (99+ languages), auto-detected
by default or forced with the language field.
Input example
{"urls": ["https://example.com/podcast-ep1.mp3"],"language": "","model": "base","outputFormats": ["text", "srt", "vtt", "segments"]}