Bulk Audio & Video Transcriber with Subtitles
Pricing
$2.50 / 1,000 media transcribeds
Bulk Audio & Video Transcriber with Subtitles
Transcribe audio and video from direct links, uploads, YouTube, TikTok, or Instagram in bulk, with timestamped text plus SRT and VTT subtitle files for every item.
Transcribe audio and video from direct links, uploads, YouTube, TikTok, or Instagram in bulk, with timestamped text plus SRT and VTT subtitle files for every item.
Best for: audio transcription, video to text, bulk transcription.
Input and output: Start with any public audio or video URL the Actor can fetch.
What can Bulk Media Transcriber do?
Add direct media links, supported social links, or public file URLs. The Actor processes distinct items in parallel, saves one result row per item, and charges only for successful transcriptions.
| What you get | Features |
|---|---|
| 📝 Text and timestamped segments for each successful item | 📦 Batch input with a set concurrency limit |
| 🎞️ Optional TXT, SRT, and VTT files | 🔁 One residential proxy retry for eligible social download blocks |
Who this is for
- Transcribe an audio or video archive
- Create subtitle files for public media
- Turn a media list into searchable text
What you get back
| Field | Type | What you get | Example |
|---|---|---|---|
index | integer | Position of this item in the input list. | 1 |
status | string | Whether this item completed its requested analysis. | succeeded |
source_type | string | Kind of source that was processed. | youtube |
source_url | string | Public source page for this record. | https://www.youtube.com/shorts/fwBIZRq-vzY |
language | string | Language detected or supplied for transcription. | en |
duration_seconds | number | Length of the media in seconds. | 45.8026875 |
word_count | integer | Number of words in the transcript. | 186 |
text | string | Complete transcript text for the media item. | This cooking trick recently changed my life. If you just drop an egg to crack it, instead of tapping it while holding it, you'll never get shells in the resu... |
segments | array | Timestamped pieces of the transcript. | [{"start":0,"end":3.58,"text":"This cooking trick recently changed my life. If you just drop an egg to crack it,"},{"start":3.66,"end":6.52,"text":"instead o... |
error | null | Problem details when this item does not complete. | null |
processing_seconds | number | Time spent processing this item. | 23.444 |
acquisition_attempts | integer | Number of attempts to fetch the media. | 2 |
acquisition_retry | string | Fallback used after an initial fetch failure. | residential_proxy |
failure_stage | null | Step that failed when the item did not complete. | null |
billing_event | string | Pay per event charge recorded for this item. | media-transcribed |
billing_event_count | integer | Number of charged successful events for this item. | 1 |
billing_status | string | Whether the successful item was charged. | charged |
title | string | Title reported for the source item. | 5 life-changing Linux tips |
creator | string | Creator name reported by the media source. | Fireship |
resolved_url | string | Final public media page used after source resolution. | https://www.youtube.com/watch?v=fwBIZRq-vzY |
bytes | integer | Downloaded media size in bytes. | 2632183 |
media_filename | string | Temporary media filename used during processing. | media.mp4 |
The run also links to its dataset and any files named in the Actor output.
What you need to provide
| Field | Type | Required | What it does | Example |
|---|---|---|---|---|
urls | array | No | Instagram reel/post, TikTok, YouTube, or direct downloadable audio/video URLs. Each successful URL costs $0.0025. | [] |
mediaUrls | array | No | Direct public HTTP(S) links to audio or video files. | ["https://raw.githubusercontent.com/ggerganov/whisper.cpp/927cfce34f31707e17f2bff35c349... |
uploadedFiles | array | No | Public or signed URLs for files uploaded to an Apify key-value store or another object store. API callers may also send objects with url and name. | `` |
language | string | No | Optional ISO-639-1 code such as en, es, or hi. Leave blank for automatic detection. | `` |
prompt | string | No | Optional names, acronyms, spellings, or context that may improve transcription. | `` |
transcription | object | No | Optional override for the built-in GPU Whisper service. Most users should leave this unchanged. To bring your own OpenAI-compatible endpoint, supply model, baseUrl, and apiKey; omitted values use the Actor's secure defaults. | {"model":"whisper-1"} |
includeSubtitles | boolean | No | Save timestamped SRT and WebVTT subtitle files for every successful item. | true |
concurrency | integer | No | Number of media items processed concurrently. Lower this if source sites throttle downloads. | 3 |
maxItems | integer | No | Safety limit for the total number of URLs accepted in one run. | 100 |
maxFileSizeMb | integer | No | Reject direct or uploaded media larger than this many megabytes before transcription. | 24 |
Quick start
- Open the Actor in Apify Console.
- Click Try for free or Create a task.
- Replace the sample values with your own input.
- Click Start.
- Open the dataset and the named output files when the run ends.
Pricing
media-transcribed: $0.0025 per media item transcribed.- Example: 100 short clips cost $0.25. Files that cannot be downloaded or have no speech are not charged. Files that cannot be downloaded or have no speech are not charged.
- You pay only for successful results. Failed or skipped items are not charged.
- Normal Apify compute and proxy costs may also apply.
Limits and honest notes
- Direct and uploaded files over
maxFileSizeMbare rejected before transcription. - A social platform may block a download. Eligible blocks get one residential proxy retry.
- A mixed batch ends with
partialstatus. An all failed batch raises an error after it saves the failed rows and summary.
Code and API
The examples below use the same values as the Apify Console sample.
Input JSON
{"urls": [],"mediaUrls": ["https://raw.githubusercontent.com/ggerganov/whisper.cpp/927cfce34f31707e17f2bff35c349632fb9e2c3a/samples/jfk.wav"],"transcription": {"model": "whisper-1"},"includeSubtitles": true,"concurrency": 3,"maxItems": 100,"maxFileSizeMb": 24}
curl
curl -X POST "https://api.apify.com/v2/acts/physealabs~bulk-media-transcriber/runs?token=$APIFY_TOKEN" \-H "Content-Type: application/json" \-d @input.json
Python
from apify_client import ApifyClientclient = ApifyClient("YOUR_APIFY_TOKEN")run = client.actor("physealabs/bulk-media-transcriber").call(run_input={'urls': [], 'mediaUrls': ['https://raw.githubusercontent.com/ggerganov/whisper.cpp/927cfce34f31707e17f2bff35c349632fb9e2c3a/samples/jfk.wav'], 'transcription': {'model': 'whisper-1'}, 'includeSubtitles': True, 'concurrency': 3, 'maxItems': 100, 'maxFileSizeMb': 24})items = client.dataset(run["defaultDatasetId"]).list_items().items
Node.js
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('physealabs/bulk-media-transcriber').call({"urls": [], "mediaUrls": ["https://raw.githubusercontent.com/ggerganov/whisper.cpp/927cfce34f31707e17f2bff35c349632fb9e2c3a/samples/jfk.wav"], "transcription": {"model": "whisper-1"}, "includeSubtitles": true, "concurrency": 3, "maxItems": 100, "maxFileSizeMb": 24});const { items } = await client.dataset(run.defaultDatasetId).listItems();
You can call this Actor from an agent or LLM tool that can send HTTP requests to the Apify API. Keep the Apify token in a secret store.
FAQ
Do I need a transcription API key?
No. The Actor has a default Whisper service. You can supply another compatible endpoint in transcription.
Do failed items cost money?
No. The media-transcribed event is charged only after a successful transcription.
Can it make subtitles?
Yes. Keep includeSubtitles set to true to save SRT and VTT files.