Audio Transcription - Speech to Text, SRT, Diarization
Pricing
from $10.00 / 1,000 audio minute (zero setup)s
Audio Transcription - Speech to Text, SRT, Diarization
Transcribe audio/video file URLs via API, MCP, or schedule — fast, accurate speech to text and video to text. Whisper alternative for meetings, podcast mp3/mp4 files, SRT subtitles, speaker diarization, summaries. Zero setup $0.010/min or BYOK $0.004/min. One JSON row per file.
Pricing
from $10.00 / 1,000 audio minute (zero setup)s
Rating
0.0
(0)
Developer
Heim AI
Maintained by CommunityActor stats
1
Bookmarked
251
Total users
72
Monthly active users
5 days ago
Last modified
Categories
Share
Audio Transcriber — Fast, Accurate Speech-to-Text
URL in → transcript out. Pass direct audio/video file URLs; get one JSON dataset row per file. Google Drive / Dropbox share links and Apple Podcasts episode links are resolved to the media file automatically. No separate speech-to-text account required. API callers still need an Apify API token. Built for MCP agents, API clients, and scheduled pipelines (that is how almost all usage runs today).
| Actor id | kaz_kakyo/audio-transcriber |
| Minimal input | { "audioUrls": ["https://dpgr.am/spacewalk.wav"] } |
| Cost | $0.010/min zero-setup · $0.004/min BYOK · $0.00005/run start |
| Output | Dataset rows with type: "transcript" or type: "error" |
Call it (MCP / API / schedule)
MCP (agents)
{"actor": "kaz_kakyo/audio-transcriber","input": {"audioUrls": ["https://dpgr.am/spacewalk.wav"]}}
Optional extras agents usually want:
{"audioUrls": ["https://dpgr.am/spacewalk.wav"],"diarize": true,"summarize": true,"includeSrt": true}
After the run, read the default dataset. Every row has a type discriminator — filter on "transcript"; treat "error" as per-file failure. Bad/unsupported URLs become error rows and the run still SUCCEEDS (including all-failed batches) so agent mistakes do not look like platform outages. Missing provider credentials, upstream auth/credit errors, or a failure to save output after charging can fail the run.
API / apify-client
Install apify-client and set APIFY_TOKEN in your environment. This public 26-second sample should return one transcript, billed as one minute: $0.01005 including the start event at 512 MB, with optional chapters off. It has one speaker; use your own permitted interview to try multiple-speaker labels.
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('kaz_kakyo/audio-transcriber').call({ audioUrls: ['https://dpgr.am/spacewalk.wav'], diarize: true },{ maxTotalChargeUsd: 1.0 }, // hard budget for this run);const { items } = await client.dataset(run.defaultDatasetId).listItems();const transcripts = items.filter((i) => i.type === 'transcript');const errors = items.filter((i) => i.type !== 'transcript');console.log({ runId: run.id, transcripts, errors });
Same shape via REST: POST /v2/acts/kaz_kakyo~audio-transcriber/runs with your token, then poll or attach a webhook.
Process new recordings repeatedly
Use a saved Task for fixed options, and pass new media URLs from your CMS, RSS feed or file uploader. A new run can transcribe and bill the same URL again; scheduling a fixed list does not provide cross-run deduplication.
- Record a stable source ID and the Apify run ID in a durable ledger before retrying uncertain work.
- Set
maxTotalChargeUsdon every run. This caps each run, not monthly spending or all retries combined. - Read every dataset row. Save transcripts and route error/skip rows to a review destination. Also handle FAILED/TIMED-OUT runs.
- If a request times out, inspect the existing run and dataset before starting another paid run.
Long runs checkpoint completed URLs for migration/resume within the same run. This is not an across-run ledger.
Podcast transcription guide and n8n starter — runnable API example, cost calculations, and explicit replay/restart limits.
For actor developers — use this as your transcription backend
Building an actor that touches audio or video (podcast tools, meeting pipelines, media scrapers)? Call this actor from yours instead of integrating a speech-to-text vendor: no key to manage, no audio handling, and your user's account pays the per-minute event directly — you never proxy the cost.
import { Actor } from 'apify';const run = await Actor.call('kaz_kakyo/audio-transcriber',{ audioUrls: mediaUrls, diarize: true },{ maxTotalChargeUsd: 2.0 },);const { items } = await Actor.apifyClient.dataset(run.defaultDatasetId).listItems();const transcripts = items.filter((i) => i.type === 'transcript');
Contract stability promise: the input fields and the type: "transcript" | "error"
output shape documented on this page are a frozen public contract — fields are only
ever added, never renamed or removed, so a nested integration does not break on our
releases. Batch-safe by design: your users' bad URLs come back as error rows, never
as a failed run inside your actor.
Output contract
One dataset item per input URL (plus skipped/invalid rows). Success shape:
{"type": "transcript","url": "https://dpgr.am/spacewalk.wav","transcript": "Full smart-formatted text…","durationSeconds": 204.3,"minutesBilled": 4,"model": "nova-3","language": "en","confidence": 0.97,"summary": "…","speakerTranscript": "Speaker 0: …\nSpeaker 1: …","srt": "1\n00:00:00,000 --> …","utterances": [{ "start": 0.0, "end": 3.2, "speaker": 0, "text": "…" }],"words": [{ "word": "Hello", "start": 0.0, "end": 0.4, "confidence": 0.99, "speaker": 0 }]}
| Field | When present |
|---|---|
transcript, durationSeconds, minutesBilled, language, confidence, model | always on success |
summary | summarize: true (English audio) |
speakerTranscript | diarize: true |
srt | includeSrt: true |
utterances / words | respective toggles |
resolvedUrl | input was a share/episode link — the direct media URL actually transcribed (url stays your input) |
wordsUrl / utterancesUrl / srtUrl | rare — oversized payloads spilled to the key-value store |
Failure / skip row (never charged):
{ "type": "error", "url": "https://…", "error": "…" }
Download the dataset as JSON, CSV, Excel, or HTML from Console or the dataset API.
Why this one
- Two pricing options. Managed transcription is $0.01/min; BYOK is $0.004/min actor fee plus your own provider bill. Both add the start event.
- Nova-3 by default — or
nova-2/whisper-large. Diarization, SRT, summaries, keyterm boosting. - HTTP-only. The speech-to-text service fetches your URL; the actor does not download media, so no proxy/compute surcharge.
- Batch-safe for agents. Bad links become
type: "error"rows; URL/decode mistakes do not fail the run — inspect run failures separately from per-file errors.
Pricing
| Event | Price | When |
|---|---|---|
| Audio minute (zero-setup) | $0.010 | No key — transcription included |
| Audio minute (BYOK) | $0.004 | deepgramApiKey set — you pay your provider at cost |
| Actor start | $0.00005 | Per run |
Minutes round up per file. At 512 MB and with optional chapters off: 90 seconds costs $0.02005 managed. A 60-minute recording costs $0.60005 managed, or $0.24005 actor fees plus your own provider bill with BYOK. Splitting a recording into files can increase rounded minutes.
Optional chapter generation adds $0.01 per generated chapter set plus your own LLM usage. Start-event pricing can scale with run memory; check the Pricing tab for your settings. BYOK savings depend on your provider plan.
Input rules agents must follow
audioUrls(required) — directhttpslinks to files (mp3,wav,m4a,flac,ogg,opus,mp4,mov,webm,mkv, ≤2 GB). Google Drive / Dropbox share links and Apple Podcasts episode links (the URL with?i=…) are resolved to the file automatically. Not YouTube / TikTok / Spotify / SoundCloud / Vimeo page URLs — those cannot be resolved.deepgramApiKey— optional; encrypted; sent only to the transcription provider.model—nova-3(default),nova-2,whisper-large.language/detectLanguage— BCP-47 or auto-detect;multi+ nova-3 for code-switching.- Toggles —
diarize,smartFormat,paragraphs,summarize,includeSrt,includeUtterances,includeWords(see Input tab). keyterms— nova-3 only; boost product names / jargon / speaker names.
Limits: 500 files/run, ~10 min upstream processing per file, concurrency 4 (2 for Whisper). Silent audio that decodes is still billed by duration.
See the Input tab for the full schema. See the API tab for run/dataset endpoints.
FAQ
Do I need my own speech-to-text account? No. Bring a key only for the $0.004/min rate.
What languages? Nova-3: 30+. whisper-large: 90+ for rarer languages.
Is my audio stored? The actor never downloads or stores media — the speech-to-text service fetches the URL; only transcript JSON lands in your dataset.
Why did my URL fail? It was not a direct, publicly reachable (or signed) media file, and not a resolvable share link (Drive/Dropbox/Apple Podcasts episode). Resolve other page URLs with a downloader first, then call this actor. The run still succeeds with type: "error" rows — check the dataset, not the run status. (Drive note: files past ~100 MB hit Drive's virus-scan page — host very large files elsewhere or use a signed URL.)
How do I keep costs predictable on a schedule? Set maxTotalChargeUsd on the run/task. Prefer BYOK once volume is steady.
If this saved you time, a Store review on the actor page helps a solo dev. Hit a problem? Open an issue.
Telemetry
When telemetry is configured, the actor sends usage and charge-summary events to the developer: a pseudonymous caller hash, paying-plan flag, run ID, origin, timestamps and event counts. No media URLs, input content or transcripts are included.