Audio Transcription - Speech to Text, SRT, Diarization avatar

Audio Transcription - Speech to Text, SRT, Diarization

Pricing

from $10.00 / 1,000 audio minute (zero setup)s

Go to Apify Store
Audio Transcription - Speech to Text, SRT, Diarization

Audio Transcription - Speech to Text, SRT, Diarization

Transcribe audio/video file URLs via API, MCP, or schedule — fast, accurate speech to text and video to text. Whisper alternative for meetings, podcast mp3/mp4 files, SRT subtitles, speaker diarization, summaries. Zero setup $0.010/min or BYOK $0.004/min. One JSON row per file.

Pricing

from $10.00 / 1,000 audio minute (zero setup)s

Rating

0.0

(0)

Developer

Heim AI

Heim AI

Maintained by Community

Actor stats

1

Bookmarked

251

Total users

72

Monthly active users

5 days ago

Last modified

Share

Audio Transcriber — Fast, Accurate Speech-to-Text

URL in → transcript out. Pass direct audio/video file URLs; get one JSON dataset row per file. Google Drive / Dropbox share links and Apple Podcasts episode links are resolved to the media file automatically. No separate speech-to-text account required. API callers still need an Apify API token. Built for MCP agents, API clients, and scheduled pipelines (that is how almost all usage runs today).

Actor idkaz_kakyo/audio-transcriber
Minimal input{ "audioUrls": ["https://dpgr.am/spacewalk.wav"] }
Cost$0.010/min zero-setup · $0.004/min BYOK · $0.00005/run start
OutputDataset rows with type: "transcript" or type: "error"

Call it (MCP / API / schedule)

MCP (agents)

{
"actor": "kaz_kakyo/audio-transcriber",
"input": {
"audioUrls": ["https://dpgr.am/spacewalk.wav"]
}
}

Optional extras agents usually want:

{
"audioUrls": ["https://dpgr.am/spacewalk.wav"],
"diarize": true,
"summarize": true,
"includeSrt": true
}

After the run, read the default dataset. Every row has a type discriminator — filter on "transcript"; treat "error" as per-file failure. Bad/unsupported URLs become error rows and the run still SUCCEEDS (including all-failed batches) so agent mistakes do not look like platform outages. Missing provider credentials, upstream auth/credit errors, or a failure to save output after charging can fail the run.

API / apify-client

Install apify-client and set APIFY_TOKEN in your environment. This public 26-second sample should return one transcript, billed as one minute: $0.01005 including the start event at 512 MB, with optional chapters off. It has one speaker; use your own permitted interview to try multiple-speaker labels.

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('kaz_kakyo/audio-transcriber').call(
{ audioUrls: ['https://dpgr.am/spacewalk.wav'], diarize: true },
{ maxTotalChargeUsd: 1.0 }, // hard budget for this run
);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
const transcripts = items.filter((i) => i.type === 'transcript');
const errors = items.filter((i) => i.type !== 'transcript');
console.log({ runId: run.id, transcripts, errors });

Same shape via REST: POST /v2/acts/kaz_kakyo~audio-transcriber/runs with your token, then poll or attach a webhook.

Process new recordings repeatedly

Use a saved Task for fixed options, and pass new media URLs from your CMS, RSS feed or file uploader. A new run can transcribe and bill the same URL again; scheduling a fixed list does not provide cross-run deduplication.

  1. Record a stable source ID and the Apify run ID in a durable ledger before retrying uncertain work.
  2. Set maxTotalChargeUsd on every run. This caps each run, not monthly spending or all retries combined.
  3. Read every dataset row. Save transcripts and route error/skip rows to a review destination. Also handle FAILED/TIMED-OUT runs.
  4. If a request times out, inspect the existing run and dataset before starting another paid run.

Long runs checkpoint completed URLs for migration/resume within the same run. This is not an across-run ledger.

Podcast transcription guide and n8n starter — runnable API example, cost calculations, and explicit replay/restart limits.

For actor developers — use this as your transcription backend

Building an actor that touches audio or video (podcast tools, meeting pipelines, media scrapers)? Call this actor from yours instead of integrating a speech-to-text vendor: no key to manage, no audio handling, and your user's account pays the per-minute event directly — you never proxy the cost.

import { Actor } from 'apify';
const run = await Actor.call(
'kaz_kakyo/audio-transcriber',
{ audioUrls: mediaUrls, diarize: true },
{ maxTotalChargeUsd: 2.0 },
);
const { items } = await Actor.apifyClient.dataset(run.defaultDatasetId).listItems();
const transcripts = items.filter((i) => i.type === 'transcript');

Contract stability promise: the input fields and the type: "transcript" | "error" output shape documented on this page are a frozen public contract — fields are only ever added, never renamed or removed, so a nested integration does not break on our releases. Batch-safe by design: your users' bad URLs come back as error rows, never as a failed run inside your actor.

Output contract

One dataset item per input URL (plus skipped/invalid rows). Success shape:

{
"type": "transcript",
"url": "https://dpgr.am/spacewalk.wav",
"transcript": "Full smart-formatted text…",
"durationSeconds": 204.3,
"minutesBilled": 4,
"model": "nova-3",
"language": "en",
"confidence": 0.97,
"summary": "…",
"speakerTranscript": "Speaker 0: …\nSpeaker 1: …",
"srt": "1\n00:00:00,000 --> …",
"utterances": [{ "start": 0.0, "end": 3.2, "speaker": 0, "text": "…" }],
"words": [{ "word": "Hello", "start": 0.0, "end": 0.4, "confidence": 0.99, "speaker": 0 }]
}
FieldWhen present
transcript, durationSeconds, minutesBilled, language, confidence, modelalways on success
summarysummarize: true (English audio)
speakerTranscriptdiarize: true
srtincludeSrt: true
utterances / wordsrespective toggles
resolvedUrlinput was a share/episode link — the direct media URL actually transcribed (url stays your input)
wordsUrl / utterancesUrl / srtUrlrare — oversized payloads spilled to the key-value store

Failure / skip row (never charged):

{ "type": "error", "url": "https://…", "error": "…" }

Download the dataset as JSON, CSV, Excel, or HTML from Console or the dataset API.

Why this one

  • Two pricing options. Managed transcription is $0.01/min; BYOK is $0.004/min actor fee plus your own provider bill. Both add the start event.
  • Nova-3 by default — or nova-2 / whisper-large. Diarization, SRT, summaries, keyterm boosting.
  • HTTP-only. The speech-to-text service fetches your URL; the actor does not download media, so no proxy/compute surcharge.
  • Batch-safe for agents. Bad links become type: "error" rows; URL/decode mistakes do not fail the run — inspect run failures separately from per-file errors.

Pricing

EventPriceWhen
Audio minute (zero-setup)$0.010No key — transcription included
Audio minute (BYOK)$0.004deepgramApiKey set — you pay your provider at cost
Actor start$0.00005Per run

Minutes round up per file. At 512 MB and with optional chapters off: 90 seconds costs $0.02005 managed. A 60-minute recording costs $0.60005 managed, or $0.24005 actor fees plus your own provider bill with BYOK. Splitting a recording into files can increase rounded minutes.

Optional chapter generation adds $0.01 per generated chapter set plus your own LLM usage. Start-event pricing can scale with run memory; check the Pricing tab for your settings. BYOK savings depend on your provider plan.

Input rules agents must follow

  • audioUrls (required) — direct https links to files (mp3, wav, m4a, flac, ogg, opus, mp4, mov, webm, mkv, ≤2 GB). Google Drive / Dropbox share links and Apple Podcasts episode links (the URL with ?i=…) are resolved to the file automatically. Not YouTube / TikTok / Spotify / SoundCloud / Vimeo page URLs — those cannot be resolved.
  • deepgramApiKey — optional; encrypted; sent only to the transcription provider.
  • model — nova-3 (default), nova-2, whisper-large.
  • language / detectLanguage — BCP-47 or auto-detect; multi + nova-3 for code-switching.
  • Toggles — diarize, smartFormat, paragraphs, summarize, includeSrt, includeUtterances, includeWords (see Input tab).
  • keyterms — nova-3 only; boost product names / jargon / speaker names.

Limits: 500 files/run, ~10 min upstream processing per file, concurrency 4 (2 for Whisper). Silent audio that decodes is still billed by duration.

See the Input tab for the full schema. See the API tab for run/dataset endpoints.

FAQ

Do I need my own speech-to-text account? No. Bring a key only for the $0.004/min rate.

What languages? Nova-3: 30+. whisper-large: 90+ for rarer languages.

Is my audio stored? The actor never downloads or stores media — the speech-to-text service fetches the URL; only transcript JSON lands in your dataset.

Why did my URL fail? It was not a direct, publicly reachable (or signed) media file, and not a resolvable share link (Drive/Dropbox/Apple Podcasts episode). Resolve other page URLs with a downloader first, then call this actor. The run still succeeds with type: "error" rows — check the dataset, not the run status. (Drive note: files past ~100 MB hit Drive's virus-scan page — host very large files elsewhere or use a signed URL.)

How do I keep costs predictable on a schedule? Set maxTotalChargeUsd on the run/task. Prefer BYOK once volume is steady.


If this saved you time, a Store review on the actor page helps a solo dev. Hit a problem? Open an issue.

Telemetry

When telemetry is configured, the actor sends usage and charge-summary events to the developer: a pseudonymous caller hash, paying-plan flag, run ID, origin, timestamps and event counts. No media URLs, input content or transcripts are included.