Transcribe Video to Text & Audio to Text — 99+ Languages
Pricing
from $0.15 / 1,000 processed audio seconds
Transcribe Video to Text & Audio to Text — 99+ Languages
Transcribe video to text and audio to text in bulk on Apify. 99+ languages, word-level timestamps, speaker diarization, SRT/VTT export. Try free.
Pricing
from $0.15 / 1,000 processed audio seconds
Rating
5.0
(2)
Developer
SIÁN OÜ
Maintained by CommunityActor stats
6
Bookmarked
485
Total users
18
Monthly active users
3 days ago
Last modified
Categories
Share
🎙️ Transcribe audio and video to text
📋 Overview
Turn recordings into searchable transcripts, timestamped segments and SRT/VTT subtitles. Use direct media URLs or uploaded files on either tier. YouTube, TikTok and Instagram links require a paid Apify account.
Open the actor · SIÁN Agency Store
🔎 What is the Audio & Video Transcriber — and when should you use it?
The Audio & Video Transcriber turns audio and video recordings into clean, structured rows you can filter, export and feed straight into a spreadsheet, database or AI agent. No account, no portal API key, no browser automation to maintain.
Use it when you need: searchable transcripts, timestamped quotes, speaker labels or subtitle files. Direct files can include word timings; caption-derived rows may have only line timings.
Use something else when: you need Facebook transcripts or a live microphone feed. Use Facebook AI Transcript Extractor for Facebook video and Reel links. This actor processes accessible recordings, not live audio.
🤖 Use with AI agents
Already connected to the Apify MCP server? Just ask for this Actor by name: sian.agency/INCREDIBLY-FAST-audio-transcriber
Your agent can pay for its own runs. This Actor is eligible for agentic payments, so an agent can discover it, run it and settle the bill over x402 (USDC on Base) or Skyfire — without an Apify account or API token of its own. Billing is the same either way: accepted audio seconds, rounded up per file, plus enabled speaker-label or translation seconds and the platform start fee. EU processing uses its own audio-second event. Error rows incur no result charge; the start fee is separate. Current rates appear in Pricing.
Otherwise copy this prompt into Claude, ChatGPT, Cursor or any MCP-enabled assistant:
I want transcripts and subtitle files from recorded audio or video using the Apify Actor `sian.agency/INCREDIBLY-FAST-audio-transcriber`.Use it when I need: searchable transcripts, timestamped quotes, speaker labels or subtitle files. Direct files can include word timings; caption-derived rows may have only line timings.Don't use it when: you need Facebook transcripts or a live microphone feed — use facebook-ai-transcript-extractor instead.How to call it: pass audioUrls for direct media links, or audioFiles for uploads. Choose language or auto-detect. Set speakerDiarization or translateToEnglish only when needed.Start with this input:{"audioUrls": ["https://raw.githubusercontent.com/rara-cyber/podcast-test-file/main/KchSfOdwCJE-56eb7ccd6dc030fa3cc280ab26925c4e234.opus"],"language": "auto","speakerDiarization": false,"translateToEnglish": false}Ask me which recordings to process, the language if known, and whether speaker labels or English translation are needed, then run the Actor and summarise the results as a table.
Things you can ask your agent for:
- Transcribe a podcast recording and copy its SRT captions.
- Label the speakers in an interview and review each timestamped quote.
- Turn recorded lecture audio into searchable text.
Machine-readable API, MCP config and OpenAPI definition for this Actor are published at apify.com/sian.agency/INCREDIBLY-FAST-audio-transcriber.md.
✨ Features
- Language selection or automatic detection.
- Timestamped segments and, for engine transcripts, word timings.
- Optional speaker labels and English translation.
- SRT/WebVTT text for caption exports.
🚀 Getting started
- Add a direct media URL to Audio/Video URLs, or upload a file in Upload Audio File or Video to Text. Clear the example URL when uploading your own file; both input lists are processed together.
- Choose the spoken language or leave Auto-detect selected. Enable speaker labels or English translation only when you need them.
- Run the actor, open Audio transcripts, and review the text against the recording. Processing Report includes a copyable transcript, subtitles, settings and accepted charges when the platform exposes counts and prices.
The default input contains a short sample recording. Replace it with your recording, or run it first to see the output format.
⚡ Quick start
{"audioUrls": ["https://example.com/interview.mp3"],"language": "auto","speakerDiarization": false,"translateToEnglish": false,"useEuServers": false}
📥 Input configuration and limits
| Field | What it does |
|---|---|
audioUrls | Direct media URLs, or paid-tier YouTube, TikTok and Instagram links |
audioFiles | Uploaded audio/video files; combined with the URL list |
language | Source language from the dropdown, or auto for detection |
speakerDiarization | Add speaker labels where available; billed per audio second |
translateToEnglish | Return English text for non-English speech; billed per audio second |
useEuServers | Request EU-region transcription processing; uses the EU base event |
Supported direct formats include MP3, WAV, FLAC, AAC, OPUS, OGG, M4A, MP4, MPEG, MOV and WebM. URLs must resolve to accessible media, not a web page containing a player.
The FREE tier processes one direct file per run. Files exceeding 5 MB or 60 seconds are rejected before transcription. The PAID tier supports multiple files, up to 1 GB each, with up to 10 concurrent files. Platform extraction can fail when content is private, removed, unavailable or lacks usable captions/media.
EU processing selects the transcription region. It does not change the location or retention settings of your Apify dataset and key-value store.
📤 Output fields
Successful rows contain the following fields. Fields unavailable for a source can be null or empty; error rows contain success: false, error and processedAt, with the source URL when known.
| Field | Meaning |
|---|---|
transcript | Recognized text; translated English text when translation is enabled |
detected_language | Language code, or null when unavailable |
duration | Audio duration in seconds |
segments | Text segments with numeric start/end seconds, language, speaker and words |
segments[].words[] | Word text and start/end seconds; speaker can be null |
srt, vtt | Subtitle file contents; save as .srt or .vtt |
speakers, languages | Unique labels/codes present in the segments |
inputUrl, mediaUrl | Submitted source and resolved media URL |
sourcePlatform | direct, youtube, tiktok or instagram |
transcriptSource | engine or captions |
fileSizeMB | File size when known, otherwise null |
mediaTitle, mediaAuthor | Source title/author when available, otherwise null |
success, processedAt, metadata | Outcome, processing time and run/tier information |
Engine-derived rows can include word timings. Caption-derived rows may have only line timings, empty word arrays and no speaker labels. Speaker labels such as SPEAKER_00 distinguish detected speakers; they do not identify people. English translation appears in transcript; a separate original-language transcript is not returned.
Text, timing and speaker labels can contain mistakes, especially with noise, overlapping speech or music. Review the recording before publishing quotes or captions. Processing time depends on recording length, accessibility, enabled options and service load.
💼 Use cases
- Search podcast and lecture recordings by text, then locate passages with timestamps.
- Review interviews and meetings with optional speaker labels.
- Export SRT or WebVTT captions for a video editor or player.
- Feed transcripts into a search index or a summarization workflow.
For Facebook video/Reel links, use Facebook AI Transcript Extractor. This actor accepts recordings, not a live microphone stream.
💰 Billing
Current event rates appear in the actor's Pricing section. The processing tier and your account's event rates determine the bill.
Audio is billed in whole seconds, rounded up per successful file. EU processing replaces the regular audio-second event with its EU event. Speaker diarization and English translation add their own per-second events when enabled. A platform-resolution fallback can have its own event. The platform start fee is separate and can apply even when a run returns an error. Error rows incur no transcription, speaker-label or translation result charge.
The report reads accepted event counts and runtime unit prices from the platform. If either is unavailable, it omits the money statement and directs you to the run's billing. Splitting a batch into multiple runs can incur multiple start fees.
🔗 Integration examples
curl -X POST 'https://api.apify.com/v2/acts/sian.agency~INCREDIBLY-FAST-audio-transcriber/runs' \-H 'Authorization: Bearer YOUR_APIFY_TOKEN' \-H 'Content-Type: application/json' \-d '{"audioUrls":["https://example.com/recording.mp3"],"language":"auto"}'
Read the completed run's dataset as JSON, CSV or Excel. Subtitle fields are file contents, not download URLs; save their text with the appropriate extension.
❓ FAQ
Does the EU option change Apify storage? It selects the transcription processing region. Manage dataset and key-value-store retention separately in Apify.
Why is a speaker or source title null? The source may not provide it, or the requested processing may not produce speaker labels. An empty/null optional field does not mean the transcript failed; check success and error.
🛠️ Troubleshooting
For a rejected file, check its size, duration and account tier. For inaccessible media, use a direct download URL or upload the file. If a YouTube row contains captions without word timings or speaker labels, use the original media file for those features. For transcription-service errors, retry a small accessible recording first and contact support if the problem persists.
⚖️ Legal: content and storage
Process recordings you have permission to use. Audio is sent for transcription, and results remain in your Apify dataset/key-value store according to your storage settings. Review those settings before processing sensitive recordings. The EU option alone does not establish compliance with a privacy law.
🤝 Support
Report an issue · Leave a review · Email support
Built by SIÁN Agency. Browse more automation tools.