YouTube Transcripts - Bulk, Multi-Language & RAG-ready
Pricing
from $2.40 / 1,000 transcripts
YouTube Transcripts - Bulk, Multi-Language & RAG-ready
Bulk YouTube transcripts: videos, playlists and channels in ONE run - the others take one at a time. Language preference that favours human-written subtitles, RAG-ready chunks with timestamps, SRT and VTT, full metadata. Videos without subtitles are never charged.
Pricing
from $2.40 / 1,000 transcripts
Rating
0.0
(0)
Developer
victor
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
YouTube Transcripts — Bulk, Multi-Language & RAG-ready
Give it a list of YouTube videos, playlists and channels — mixed, in a single run — and get back clean transcripts with timestamps, metadata, and optional chunks ready to drop straight into a vector index.
Most transcript Actors take one video, or one channel, per run. If you have 2,000 URLs to process, that is 2,000 runs and 2,000 lots of start-up overhead. This one takes the whole list at once.
What you get
| Mixed batch input | Videos, playlists, channels and @handles in one run |
| Language preference chain | Ask for ["es","en"] and a human-written Spanish subtitle wins; a human-written English one beats an auto-generated Spanish one |
| RAG-ready chunks | Text split at sentence boundaries where the captions have punctuation, at caption-line boundaries where they do not — never mid-word. Each chunk carries the second it starts and ends, and says whether it closed on a sentence |
| SRT and WebVTT | Subtitle files, not just JSON |
| Honest failures | Every video that has no transcript says exactly why — and is not charged |
| Metadata included | Title, channel, duration, view count, thumbnail |
Input
Paste URLs in any of these shapes — they all work, and you can mix them:
https://www.youtube.com/watch?v=dQw4w9WgXcQhttps://youtu.be/dQw4w9WgXcQhttps://www.youtube.com/shorts/xxxxxxxxxxxhttps://www.youtube.com/playlist?list=PL...https://www.youtube.com/channel/UCYO_jab_esuFRV4b17AJtAwhttps://www.youtube.com/@3blue1brown@veritasiumdQw4w9WgXcQ
Use maxVideosPerSource to take, say, the 20 most recent videos of each channel, and maxVideos as a hard cap on the whole run.
Output
One record per video:
{"video_id": "dQw4w9WgXcQ","url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","success": true,"title": "Rick Astley - Never Gonna Give You Up","channel": "Rick Astley","video_duration_seconds": 213,"view_count": 1657482913,"language": "en","is_auto_generated": false,"available_languages": ["de", "en", "es", "fr", "ja", "pt"],"segment_count": 61,"character_count": 1834,"transcript": "We're no strangers to love…","segments": [{ "text": "We're no strangers to love", "start": 18.64, "duration": 3.2 }],"chunks": [{ "chunk_index": 0, "text": "…", "start": 18.64, "end": 74.1,"char_count": 987, "ends_sentence": true }]}
When a video has no usable transcript you still get a record, with success: false and a plain reason:
{ "video_id": "…", "success": false, "error": "this video has no subtitle track at all" }
So a run over 1,000 videos gives you the 940 that worked and a clear list of the 60 that did not — instead of silently returning fewer rows than you asked for.
Why chunks matter
If you are building search or Q&A over video, raw caption segments are the wrong shape: each one is two seconds of half a sentence. You end up writing a chunker that merges them sensibly and keeps the timestamps so you can link an answer back to the exact moment.
Turn on includeChunks and that is already done.
One honest detail: auto-generated captions often contain no punctuation at all, sometimes for the whole video. Waiting for a full stop would then produce one enormous useless block, so chunks close at a caption-line boundary instead — never mid-word — and ends_sentence tells you which happened. You know what you are indexing.
About the proxy
YouTube rate-limits transcript requests by IP — in our testing it starts refusing after roughly 80 videos from one address. For anything beyond a handful of videos you need a proxy, and residential is the option that holds up. The default input is already set to Apify Residential Proxy.
This is not a quirk of this Actor; it is the reason every transcript Actor either uses proxies or fails on larger runs.
Pricing
$3.00 per 1,000 transcripts — 40% below the market leader's $5.00, and it drops further on paid Apify plans:
| Your Apify plan | Price per 1,000 |
|---|---|
| Free | $3.00 |
| Bronze | $2.80 |
| Silver | $2.60 |
| Gold and above | $2.40 |
Videos without subtitles are not charged. If a run finds 940 transcripts out of 1,000 URLs, you pay for 940 — the 60 failures cost you nothing and each one tells you why.
Limits, stated plainly
- Only videos that have subtitles, manual or auto-generated. There is no speech-to-text here; nothing can invent a transcript for a video that has none.
- Private, deleted, age-restricted and members-only videos cannot be read.
- Channels and playlists return their most recent videos; very long back-catalogues may not be returned in full.