YouTube Transcript Scraper
Pricing
Pay per event
YouTube Transcript Scraper
Get the transcript (captions) of any public YouTube video, Short or live replay: timed segments, plain text, SRT or VTT, in the language you choose (manual or auto-generated, optional YouTube translation), plus video metadata. LLM-ready output, pay only per transcript returned.
What does YouTube Transcript Scraper do?
YouTube Transcript Scraper gets the transcript (captions / subtitles) of public YouTube videos, Shorts and live-stream replays, and returns it as clean, LLM-ready data: timed segments, plain text, SRT or WebVTT. It works as a simple YouTube transcript API: send a list of video URLs or IDs, get one dataset item per video.
- ✅ Manual (human) captions and auto-generated captions, in the language you choose, with a clear fallback when that language doesn't exist
- ✅ Optional translation with YouTube's own machine translation (
translateTo) - ✅ Video metadata in the same item: title, channel, duration, publish date, view count
- ✅ You only pay for transcripts returned (plus a small start fee per run). Videos without captions, private or removed videos are reported as errors and never charged
- ❌ Does not download video/audio, does not transcribe audio itself, does not scrape comments
Why use YouTube Transcript Scraper?
- RAG and AI pipelines: feed talks, podcasts, lectures and tutorials into a vector database or an LLM for summaries, Q&A and search.
- Content repurposing: turn videos into blog posts, newsletters, threads and show notes.
- Research and monitoring: analyze what channels, brands or competitors say in their videos.
- Subtitles: get SRT or VTT files ready for an editor or a player.
It runs on the Apify platform, so you get an API, scheduling, webhooks, integrations (Make, Zapier, n8n, LangChain, LlamaIndex...), and proxy rotation built in. Each video is retried with fresh sessions and, when needed, residential IPs, so runs don't fail silently.
What data can YouTube Transcript Scraper extract?
| Field | Type | Description |
|---|---|---|
videoId, url | string | Video ID and canonical watch URL |
title, channelName, channelId | string | Public video and channel metadata |
durationSeconds, viewCount | integer | Length and views at scrape time |
publishDate | string | ISO 8601 publish date |
language, languageName | string | Caption track that was used |
isAutoGenerated | boolean | true for YouTube's automatic speech recognition |
translatedTo | string / null | Target language of YouTube's translation |
availableLanguages | array | All caption tracks on the video: {code, name, isAutoGenerated} |
segments / text / srt / vtt | array / string | The transcript, in the format you chose |
wordCount | integer | Words in the transcript (CJK characters count individually) |
error, errorMessage | string | Only on videos without a transcript (not charged) |
HTML entities are decoded (' → ') and sound tags such as [Music] or [Applause] are kept as they appear on YouTube.
How to scrape YouTube transcripts
- Click Try for free and open the Input tab.
- Paste video URLs or IDs into Videos (
watch?v=,youtu.be/,shorts/,live/andembed/links all work). - Set Preferred language (for example
en,pt-BR,es,ja), and optionally Translate to. - Pick an Output format:
textis best for LLMs,segmentskeeps timestamps,srt/vttare subtitle files. - Click Start and download the results as JSON, CSV, Excel or HTML, or read them through the API.
How much does it cost to scrape YouTube transcripts?
This Actor uses pay-per-event pricing:
| Event | Price | When |
|---|---|---|
apify-actor-start | US$ 0.005 | Once per run |
transcript | US$ 0.005 | Per video that returned a transcript |
translation | US$ 0.05 | Extra, per video whose transcript was translated with translateTo |
Examples: 1,000 videos in one run cost US$ 5.005. One video in one run costs US$ 0.01. One translated video costs US$ 0.06 (start + transcript + translation).
⚠️ Translation is slower and more expensive. YouTube throttles translated captions heavily, so each translated video takes about 1 to 3 minutes instead of a few seconds, and costs US$ 0.05 on top of the transcript. Leave
translateToempty unless you need it. If you only want a transcript in a language that already exists on the video, uselanguageinstead: it's free of the translation fee.
You are not charged for videos that have no captions, are private, removed, age-restricted, or fail for any other reason. A failed translation (translation_failed) is not charged either. Platform usage (compute, proxy) is included. You can cap the spend of a run with the Maximum cost per run setting; the Actor stops cleanly when it is reached.
Input
See the Input tab for every option. Example:
{"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ","https://youtu.be/NNnIGh9g6fA","https://www.youtube.com/shorts/9dMFxnpEkKc"],"language": "en","includeAutoGenerated": true,"outputFormat": "text","includeVideoMetadata": true}
Language fallback. For each video the Actor picks, in this order: a manual track in language → an auto-generated track in language → any manual track → any auto-generated track. pt also matches pt-BR, en matches en-US, and so on. The item always says which track was used (language, isAutoGenerated) and lists every track in availableLanguages. Turn Include auto-generated captions off to get only human-made captions.
Translation. translateTo asks YouTube for its machine translation of the chosen track (for example "translateTo": "pt"). If YouTube doesn't offer that language for the video, the item is an error translation_unavailable (not charged). If the track is already in that language, no translation is done.
Output
You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. Example item with outputFormat: "segments" (segments shortened):
{"videoId": "dQw4w9WgXcQ","url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","channelName": "Rick Astley","channelId": "UCuAXFkgsw1L7xaCfnd5JJOw","durationSeconds": 213,"publishDate": "2009-10-24T23:57:33-07:00","viewCount": 1821342774,"isLiveContent": false,"language": "en","languageName": "English","isAutoGenerated": false,"translatedTo": null,"availableLanguages": [{ "code": "en", "name": "English", "isAutoGenerated": false },{ "code": "en", "name": "English (auto-generated)", "isAutoGenerated": true },{ "code": "de-DE", "name": "German (Germany)", "isAutoGenerated": false },{ "code": "ja", "name": "Japanese", "isAutoGenerated": false },{ "code": "pt-BR", "name": "Portuguese (Brazil)", "isAutoGenerated": false },{ "code": "es-419", "name": "Spanish (Latin America)", "isAutoGenerated": false }],"segments": [{ "start": 1.36, "duration": 1.68, "text": "[♪♪♪]" },{ "start": 18.64, "duration": 3.24, "text": "♪ We're no strangers to love ♪" },{ "start": 22.64, "duration": 4.32, "text": "♪ You know the rules and so do I ♪" }],"wordCount": 366,"scrapedAt": "2026-09-29T17:14:02.592Z"}
With outputFormat: "text" the item has text (one string) instead of segments; with srt or vtt it has an srt or vtt string.
A video without a transcript (never charged):
{"videoId": "ywX-RVl5FFM","url": "https://www.youtube.com/watch?v=ywX-RVl5FFM","error": "no_captions","errorMessage": "This video has no captions.","scrapedAt": "2026-09-29T17:20:30.101Z"}
Error codes: no_captions, unavailable (removed or doesn't exist), geo_restricted (the uploader limits the video to some countries and it could not be read from any of them), private, age_restricted, members_only, live_not_finished (live now, upcoming or premiere), translation_unavailable, translation_failed, no_track_for_filter (only auto captions exist and you turned them off), invalid_input, blocked (YouTube refused all retries), not_processed (maximum cost per run reached).
The dataset has two views: Transcripts and Errors (not charged). Run statistics are stored in the RUN_STATS key-value record.
Tips
- For LLMs, use
outputFormat: "text": it is the most compact. - Auto-generated segments are clipped so they don't overlap (YouTube's own auto lines overlap by design).
- Translated transcripts take longer (1 to 3 minutes each) and carry the extra
translationcharge: YouTube throttles translation requests much more than original captions. - Batch many videos in one run: the start fee is charged once per run.
- The run only fails when every video failed; otherwise failed videos are listed as error items.
Limits
- Only public videos. Private, members-only and age-restricted videos (which require a signed-in account) are reported as errors; the Actor never logs in.
- Videos that are live right now or upcoming have no transcript yet (
live_not_finished). Replays work once YouTube has processed them. - Transcripts come from YouTube's captions. If a video has no captions at all (typical for music, ambience or silent videos), there is nothing to return.
- Region-locked videos: when the uploader limits a video to some countries, the Actor automatically retries from a residential IP in one of the allowed countries.
publishDateis best-effort and may benullfor a few videos.
FAQ and support
Is it legal to scrape YouTube transcripts? This Actor only reads publicly available captions and public video metadata. Our Actors are ethical and do not extract any private user data, such as email addresses, gender, or location. They only extract what the user has chosen to share publicly. We therefore believe that our Actors, when used for ethical purposes by Apify users, are safe. However, you should be aware that your results could contain personal data. Personal data is protected by the GDPR in the European Union and by other regulations around the world. You should not scrape personal data unless you have a legitimate reason to do so. If you're unsure whether your reason is legitimate, consult your lawyers. Respect the copyright of the content you process.
Can I use it from code or an AI agent? Yes: see the API tab for ready-made calls (HTTP, JavaScript, Python, CLI), or use it through the Apify MCP server.
Found a problem or need a feature? Open an issue in the Issues tab.