YouTube Transcript Scraper: Captions, Timestamps & Metadata
Pricing
$4.00 / 1,000 transcript returneds
YouTube Transcript Scraper: Captions, Timestamps & Metadata
Get YouTube video transcripts (manual or auto-generated captions) as timestamped segments and clean plain text, with title, channel, duration, views and publish date. Paste video, channel or playlist URLs. No browser, no API key. Videos without captions are free.
Pricing
$4.00 / 1,000 transcript returneds
Rating
0.0
(0)
Developer
Changefeeds Tools
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Share
Paste YouTube video links and get back each video's transcript: the captions as timestamped segments and as one clean plain-text string, plus the video's title, channel, duration, view count and publish date. Built for "youtube transcript" and "transcript for LLM" jobs: feeding videos to an LLM or a RAG index, summarising talks and podcasts, content research, SEO briefs, subtitles analysis, quote search.
- Works with manual captions and YouTube's auto-generated captions, in any language the video has.
- Plain HTTP, no browser: in our tests a 3.5-hour video's transcript (5,700+ caption lines) arrived in under 2 seconds.
- Videos that have no transcript are not charged. They get a status row that says why (no captions, private, age-restricted, removed, blocked).
- No YouTube account, no API key, no cookies.
Input
| Field | Default | Notes |
|---|---|---|
videos | required | Watch URLs, youtu.be links, /shorts/, /live/, /embed/ URLs or bare 11-character ids. Also a channel URL (https://www.youtube.com/@name, /channel/UC…) for its latest uploads, or a playlist URL (/playlist?list=…) for its first videos. Duplicates are fetched once. |
language | en | Preferred caption language (en, es, de, pt-BR, ja…). See "Which caption track?" below. |
outputFormat | both | segments (timestamped lines), text (one string), or both. |
includeTimestampsInText | false | Write the text as [hh:mm:ss] line rows instead of one paragraph. |
includeDescription | false | Add the video description to each row. |
maxVideos | 1000 | Cap on videos per run (after expanding channels and playlists). |
maxVideosPerChannel | 10 | For channel/playlist URLs: how many videos to take, up to 30. |
proxyConfiguration | Apify residential | YouTube bot-checks cloud IPs, so requests go through residential proxy by default. Included in the price. |
{"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ","https://youtu.be/jNQXAC9IVRw","kCc8FmEb1nY"],"language": "en","outputFormat": "both"}
Spanish transcripts with timestamps, from a video and the first two videos of a playlist:
{"videos": ["https://www.youtube.com/watch?v=kJQP7kiw5Fk","https://www.youtube.com/playlist?list=PLZHQObOWTQDPD3MizzM2xVFitgF8hE_ab"],"language": "es","maxVideosPerChannel": 2,"outputFormat": "text","includeTimestampsInText": true}
Output
One transcript row per video with captions. This is a real row from a run on
2026-09-29, with segments cut to three entries:
{"type": "transcript","video_id": "jNQXAC9IVRw","url": "https://www.youtube.com/watch?v=jNQXAC9IVRw","title": "Me at the zoo","channel_name": "jawed","channel_id": "UC4QobU6STFB0P71PMvOGN5A","duration_seconds": 19,"view_count": 438445209,"published_at": "2005-04-23T20:31:52-07:00","language": "en","language_name": "English","is_generated": false,"requested_language_matched": true,"available_languages": [{ "language": "en", "name": "English", "is_generated": false },{ "language": "de", "name": "German", "is_generated": false }],"transcript_text": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say","segments": [{ "start": 1.2, "duration": 2.16, "text": "All right, so here we are, in front of the elephants" },{ "start": 5.318, "duration": 2.656, "text": "the cool thing about these guys is that they have really..." },{ "start": 7.974, "duration": 4.642, "text": "really really long trunks" }],"word_count": 39,"segment_count": 6,"thumbnail": "https://i.ytimg.com/vi/jNQXAC9IVRw/hqdefault.jpg","fetched_at": "2026-09-30T04:34:23.525Z"}
With includeTimestampsInText: true the text looks like this (Spanish
captions of a 3Blue1Brown video from the same test):
[00:00:10] El elemento fundamental y raíz de todo el álgebra lineal es el vector.
startanddurationare seconds (millisecond precision). Auto-generated captions often overlap the next line; the values are YouTube's own.is_generated: truemeans YouTube's automatic speech recognition: no punctuation or capitals, and occasional wrong words.published_atis best effort; it isnullon the rare video where YouTube does not report it.- Rows are written as each video finishes, so their order can differ from the
input order. A run summary (counts per status, HTTP requests) is saved as
OUTPUTin the key-value store.
Videos without a transcript get a free video row instead:
{"type": "video","input": "https://www.youtube.com/watch?v=EColTNIbOko","video_id": "EColTNIbOko","url": "https://www.youtube.com/watch?v=EColTNIbOko","status": "no_captions","error": "This video has no captions (manual or auto-generated).","title": "Forests from Above (No Sound) — 10 Hours Screensaver of 4K UHD Drone Aerials","available_languages": [],"fetched_at": "2026-09-30T04:34:22.976Z"}
status | Meaning |
|---|---|
no_captions | The video has no manual or auto-generated captions (common for music-only, silent and live videos). |
unavailable | Removed, never existed, members-only, region-blocked, or a premiere that has not started. |
private | The video is private. |
age_restricted | YouTube requires a signed-in adult account; this actor does not sign in. |
blocked | YouTube refused the request (bot check or rate limit) even after fresh proxy IPs. Try again later. |
invalid | The input is not a YouTube video, channel or playlist link. |
error | Anything else; the error field has the details. |
Which caption track?
For language: "en" the actor takes, in order: a manual English track (exact
code first, so en before en-GB), then the auto-generated English track.
If the video has no English track at all, it returns the first manual track in
another language (or the first track) and sets
requested_language_matched: falseavailable_languages. The actor does not machine-translate.
Pricing
$4 per 1,000 transcripts ($0.004 per transcript row), residential proxy included. Status rows
(no captions, unavailable, private, age-restricted, blocked, invalid) are free.
The run stops cleanly at the maximum total charge you set; it never fetches a
video it could not bill within that limit. If a run delivers nothing, it
leaves one free run_info row explaining why.
Limits (honest)
- Only videos that have captions. There is no speech-to-text here: if a
video has neither manual nor auto-generated captions, you get a free
no_captionsrow, not a transcript. - YouTube can block some requests, especially from data-centre IPs and
at high volume. Blocked videos come back as free
blockedrows, never as empty transcripts. Requests go through Apify residential proxy by default and blocked videos are retried through two fresh proxy IPs; proxy traffic is included in the price, you are not billed for it. - Channel and playlist URLs cover the first page YouTube lists: up to 30 of a channel's latest regular uploads (not Shorts or live streams) or a playlist's first videos. For more, list the video URLs directly.
- Age-restricted and private videos need a signed-in account and are not supported.
- This actor reads YouTube's own caption endpoints, which YouTube can change without notice. If that happens, affected videos fail as free rows until the actor is updated.
Troubleshooting: youtube-transcript-api RequestBlocked / IpBlocked
If you're running the open-source youtube-transcript-api on a cloud server and it works on your laptop but throws RequestBlocked or IpBlocked errors in production, the issue is not your code — it's YouTube's IP-based bot detection.
Why it happens
YouTube maintains IP-range data for AWS, GCP, Azure, Railway, Heroku and other datacenters, and treats requests from those ranges far more aggressively than residential ISP ranges. As the library's own maintainer states in the README: "YouTube has started blocking most IPs that are known to belong to cloud providers (like AWS, Google Cloud Platform, Azure, etc.), which means you will most likely run into RequestBlocked or IpBlocked exceptions when deploying your code to any cloud solutions."
Free fixes
- Add retries and backoff to your requests and cache results so you never re-fetch a transcript you already have.
- Route through rotating residential proxies. The library ships first-class proxy support via
youtube_transcript_api.proxies:
from youtube_transcript_api import YouTubeTranscriptApifrom youtube_transcript_api.proxies import GenericProxyConfigytt_api = YouTubeTranscriptApi(proxy_config=GenericProxyConfig(http_url="http://user:pass@your-proxy.org:port",https_url="https://user:pass@your-proxy.org:port",))
- Avoid hammering one video or channel. Space out requests and use rotating proxy IPs to reduce how often you trip the check.
Keep your proxy config current and re-test whenever YouTube adjusts its detection — the pattern evolves periodically.
This actor handles it
This actor uses residential proxy by default, rotates IPs on any block, and retries through fresh proxies. Blocked, private, or caption-less videos are not charged.
FAQ
Can I use the transcript with ChatGPT, Claude or another LLM?
Yes; that is the main use. Take transcript_text (one clean paragraph) or
turn on includeTimestampsInText if you want the model to cite times.
Does it work for Shorts and live streams?
Shorts, yes, when they have captions. Live streams usually get captions only
after they end; until then they return no_captions.
Does it download the video or audio? No. It only reads the caption text and the video's public metadata.
Can I get SRT or VTT subtitle files?
Not directly; the segments array has everything an SRT needs (start,
duration, text) and converts in a few lines of code.
How fast is it? Three videos run in parallel. In our tests (2026-09-29) a transcript took 1 to 2 seconds, including multi-hour videos; ten mixed videos took about 8 seconds.
Is this allowed? It reads publicly available captions the way YouTube's own apps do, without signing in. You are responsible for how you use the transcripts (copyright and YouTube's terms apply to the content).