YouTube Transcript Scraper + AI Summary & Chapters
Pricing
from $5.00 / 1,000 transcript fetcheds
YouTube Transcript Scraper + AI Summary & Chapters
Get YouTube transcripts (manual or auto-generated subtitles, any language) with timestamps from video URLs, playlists or channels, as text, SRT, VTT or segments. Optional AI summary, chapters, key insights and translation. 98%+ success rate. Pay per transcript.
Pricing
from $5.00 / 1,000 transcript fetcheds
Rating
0.0
(0)
Developer
Maxence Bernerd
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
YouTube Transcript Scraper gets the transcript of any public YouTube video with timestamps: manual or auto-generated subtitles, in any language, from video URLs, ids, playlists or whole channels. Every row carries the transcript as timestamped segments, plain text with punctuation and paragraphs, SRT and VTT, plus the video metadata (title, channel, duration, views, tags, publish date). Turn on the AI layer to add a three-level summary, YouTube-ready chapters, key insights (entities, quotes, topics, sentiment, calls to action) and a full translation. Built for reliability: 98%+ success rate measured on a mixed sample (long videos, Shorts, non-English, live replays), videos without captions are reported cleanly instead of failing the run, and you pay $0.005 per transcript, nothing for videos where no transcript exists.
What data does YouTube Transcript Scraper extract?
One dataset row per video, flat enough for CSV and complete enough for AI agents and RAG pipelines:
| Field | Description |
|---|---|
videoId, videoUrl, source | Stable YouTube id, canonical URL, and where the video came from (input, playlist:<id>, channel:<handle>) |
status, errorReason, errorMessage | ok, no_transcript, unavailable, restricted or error, with a machine-readable reason (see the edge-case table below) |
title, channelName, channelId, channelUrl, durationSeconds, viewCount, description, tags, thumbnailUrl, isLive | Video metadata, always included |
publishedAt, likeCount, category, isUnlisted | Extended metadata (with includeMetadata) |
transcriptLanguage, transcriptType | Language of the transcript and whether it is manual (uploaded by the creator) or auto (YouTube speech recognition) |
availableLanguages, translatableLanguages | Every caption track the video offers (code, name, type) and the languages YouTube can machine-translate to |
segments | [{ start, end, text }] in seconds, one entry per caption line |
text | Continuous transcript with punctuation and paragraphs (reconstructed when the auto track has none) |
srt, vtt | Ready-to-use subtitle files |
wordCount | Number of words in the transcript |
summary | oneLine, paragraph, keyPoints (with enableSummary) |
chapters, chaptersText | 4 to 15 timestamped chapters as objects and in YouTube description format (with enableChapters) |
insights | Entities, timestamped quotes, questions, topics, sentiment, calls to action, links and promotions (with enableInsights) |
translation | Full translation of text (with translateTo) |
aiBillingUnits, scrapedAt | 20-minute blocks billed for the AI events, and the scrape timestamp |
No account is used, no cookies, no personal data: only public captions and public metadata.
How to get YouTube transcripts
- Paste video URLs (any form:
watch?v=,youtu.be, Shorts, live replays, embeds), playlist URLs or channel handles (@mkbhd) in the input. - Optionally pick a caption language, the output formats, and the AI options you want.
- Click Start. Videos are processed in parallel; a 50-video channel takes about a minute.
- Read the results in the Output tab (views: Transcripts, Summaries, Chapters, Metadata) or download them as JSON, CSV, Excel or through the API.
Input examples
Transcripts of three videos, original language, every format:
{"videoUrls": ["https://www.youtube.com/watch?v=9Ff4S1FJdRM","https://youtu.be/dQw4w9WgXcQ","https://www.youtube.com/shorts/R6yNUnRXZ64"]}
The 50 most popular videos of a channel, English captions only, with summaries and chapters:
{"channelUrls": ["@mkbhd"],"maxVideosPerChannel": 50,"channelSort": "popular","language": "en","outputFormat": "text","enableSummary": true,"enableChapters": true}
A whole playlist in SRT, plus insights and a French translation of each transcript:
{"playlistUrls": ["https://www.youtube.com/playlist?list=PLFgquLnL59alCl_2TQvOiD5Vgm1hCaGSI"],"maxVideosPerChannel": 100,"outputFormat": "srt","enableInsights": true,"translateTo": "fr","summaryLanguage": "fr"}
Input fields: videoUrls, videoIds, playlistUrls, channelUrls, maxVideosPerChannel (default 50), channelSort (latest | popular), includeShorts, language (empty = original language), preferManual (default true), fallbackToAutoGenerated (default true), outputFormat (all | segments | text | srt | vtt), includeMetadata, enableSummary, summaryLanguage, enableChapters, enableInsights, translateTo, maxConcurrency, proxyConfiguration.
Output example
A real row (transcript fields shortened):
{"videoId": "9Ff4S1FJdRM","videoUrl": "https://www.youtube.com/watch?v=9Ff4S1FJdRM","source": "input","status": "ok","errorReason": null,"title": "Is AI Turning Us All into the Same Person? | Sandra Matz | TED","channelName": "TED","channelId": "UCAuUUnT6oDeKwE6v1NGQxug","channelUrl": "https://www.youtube.com/channel/UCAuUUnT6oDeKwE6v1NGQxug","publishedAt": "2026-09-28T15:00:04.000Z","durationSeconds": 817,"viewCount": 76770,"likeCount": 656,"category": "People & Blogs","tags": ["TEDTalk", "TEDTalks", "TED Talk", "TED Talks", "TED"],"thumbnailUrl": "https://i.ytimg.com/vi_webp/9Ff4S1FJdRM/sddefault.webp","transcriptLanguage": "en","transcriptType": "manual","availableLanguages": [{ "code": "ar", "name": "Arabic", "type": "manual" },{ "code": "en", "name": "English", "type": "manual" },{ "code": "es", "name": "Spanish", "type": "manual" }],"translatableLanguages": ["ar", "zh-Hant", "nl", "en", "fr", "de", "…"],"segments": [{ "start": 4.668, "end": 8.046, "text": "People worry about AI for all kinds of reasons." },{ "start": 8.046, "end": 10.84, "text": "It's polarizing, it spreads misinformation," },{ "start": 10.84, "end": 12.717, "text": "it's coming for our jobs." }],"text": "People worry about AI for all kinds of reasons. It's polarizing, it spreads misinformation, it's coming for our jobs. And those are all good reasons to be nervous. But what really keeps me up at night is something else. …","srt": "1\n00:00:04,668 --> 00:00:08,046\nPeople worry about AI for all kinds of reasons.\n\n2\n…","vtt": "WEBVTT\n\n00:00:04.668 --> 00:00:08.046\nPeople worry about AI for all kinds of reasons.\n\n…","wordCount": 2157,"summary": null,"chapters": null,"insights": null,"translation": null,"scrapedAt": "2026-09-29T18:41:59.575Z"}
A video without captions is a valid row, not a failed run, and it is not charged:
{"videoId": "Ct6BUPvE2sM","status": "no_transcript","errorReason": "no_captions","errorMessage": "This video has no captions (neither manual nor auto-generated).","title": "PIKOTARO - PPAP (Pen Pineapple Apple Pen) (Long Version) (Official Video) [Ultra Records]","availableLanguages": []}
Edge cases: what you get for each kind of video
| Case | status | errorReason | Charged |
|---|---|---|---|
| Video with manual or auto captions | ok | — | yes |
| Shorts, live replays, premieres already aired | ok | — | yes |
| Video without any captions | no_transcript | no_captions | no |
Requested language not available (other languages listed in availableLanguages) | no_transcript | language_unavailable | no |
Only auto captions and fallbackToAutoGenerated is off | no_transcript | no_captions | no |
| Live stream in progress | unavailable | live_in_progress | no |
| Upcoming premiere or scheduled stream | unavailable | upcoming | no |
| Deleted video, wrong id | unavailable | video_not_found | no |
| Private video | restricted | private | no |
| Age-restricted video | restricted | age_restricted | no |
| Members-only video | restricted | members_only | no |
| Region-blocked video | restricted | region_blocked | no |
| YouTube blocked every retry (rare, retried on 4 fresh proxy sessions first) | error | blocked_by_youtube | no |
Restrictions are reported, never bypassed: the Actor uses no account and no cookies.
How much does it cost to get YouTube transcripts?
Pay per event, no subscription, no compute charges on top:
| Event | Price | When it is charged |
|---|---|---|
transcript-fetched | $0.005 | Per video whose transcript was obtained (status: ok) |
metadata-fetched | $0.0005 | Per video, with includeMetadata, when the extended fields (publish date, likes, category) were obtained |
summary-generated | $0.02 per 20-minute block | Per video with a summary |
chapters-generated | $0.02 per 20-minute block | Per video with chapters |
insights-extracted | $0.03 per 20-minute block | Per video with insights |
translation-generated | $0.04 per 1,000 words | Per transcript translated (source words, rounded up) |
| Actor start | $0.00005 | Per run |
Videos up to 20 minutes count as one block; a 50-minute video counts as 3 blocks (the transcript itself is $0.005 whatever the length). Nothing is charged for videos without transcript, for AI results the model failed to produce, or for videos skipped because the run budget was reached. Set Maximum cost per run in the run options: the Actor stops cleanly at the limit and tells you in the log.
Real cost for typical 10-minute videos:
| Videos | Transcripts only | + extended metadata | + summary and chapters | + insights and everything above |
|---|---|---|---|---|
| 100 | $0.50 | $0.55 | $4.55 | $7.55 |
| 1,000 | $5 | $5.50 | $45.50 | $75.50 |
| 10,000 | $50 | $55 | $455 | $755 |
A full translation of a 10-minute video (about 1,500 words) adds $0.08; of a 90-minute podcast (about 14,000 words) $0.56.
AI summary, chapters and insights
Each AI feature runs on Claude and returns validated JSON. Long videos are processed in 20-minute windows whose results are merged, so a 3-hour podcast gets one coherent summary and a single chapter list. Output language follows summaryLanguage (default: the transcript language); quotes stay in the original language.
Real output for the TED talk above (enableSummary):
{"oneLine": "AI systems optimized for engagement narrow human preferences and suppress exploration, risking a flattened human experience unless companies redesign incentives to encourage discovery.","paragraph": "AI recommendation systems are trained to optimize for short-term engagement and exploitation of known preferences rather than exploration and discovery. Studies show that when people rely on AI for guidance, their preferences become more normative and less diverse, their creative output becomes less unique, and their choices converge toward sameness. An experiment with ChatGPT recommending Baskin-Robbins flavors 100 times resulted in 96 recommendations of just two popular flavors. …","keyPoints": ["ChatGPT recommended same two ice cream flavors 96 of 100 times, showing algorithmic preference for popularity over diversity","AI systems trained to optimize for short-term engagement and exploitation, not exploration or discovery","People using AI guidance show less diverse preferences, less unique creative output, and more conformist choices","Companies should implement adjustment dials letting users control exploration versus exploitation levels in recommendations"]}
Chapters (enableChapters), as chaptersText, ready to paste in the video description:
0:00 AI Makes Us Boring1:07 Exploitation vs. Exploration Trade-off2:39 How AI Optimizes for Safety5:19 The Flattening of Human Experience6:53 AI Narrows Your Taste9:27 Reclaiming Risk and Discovery10:30 The Dial: Balancing Exploration12:27 Rewarding AI for Smart Risks
Insights (enableInsights) on a 30-minute Spanish vlog with Tony Hawk, output in English:
{"entities": { "people": ["Tony Hawk", "Rubius"], "brands": ["PlayStation", "Gamecube", "Switch"], "products": ["Tony Hawk Pro Skater 3+4", "motorized skateboard"], "places": ["San Diego", "Spain"] },"quotes": [{ "start": 105, "timestamp": "1:45", "text": "if you never try, you never get it" }],"questions": ["How did the creator get Tony Hawk to respond to his DM?", "Can the creator win against Tony Hawk in the Tony Hawk Pro Skater video game?"],"topics": ["skateboarding", "childhood nostalgia", "video game competition", "recovery from injury"],"sentiment": "positive","callsToAction": ["Subscribe to the channel", "Like the video"],"linksAndPromotions": ["Tony Hawk Pro Skater 3+4 game", "Tony Hawk's skateboard signed by Tony Hawk"]}
Translation (translateTo) returns { language, text, wordCount, units } with the paragraphs of text preserved.
Reliability: 98%+ success rate, measured
On a test set of 31 real videos covering every edge case (multilingual manual captions, French, Spanish, Japanese, German, Portuguese, Hindi and Korean videos, two podcasts over 3 hours, a 1-hour live replay, four Shorts, videos without captions, a members-only video, two live streams, a private video, an age-restricted video, two deleted ids), the Actor returned the expected status for 31 videos out of 31, with zero errors, both from a residential connection and from the Apify platform. The same set run through the most-used transcript Actor on the Store produced identical caption content on every video that has captions, but failed the run on the private, age-restricted, live and deleted videos.
How it stays reliable:
- Transcripts come straight from YouTube's caption endpoints through residential proxy sessions, one sticky session per video, rotated as soon as YouTube pushes back (rate limit, captcha, bot check). Blocks are retried on fresh sessions before a video is ever marked as
error. - Two independent client identities are tried before concluding that a video has no captions.
- The requested language is never silently replaced by another one.
- Every failure is a row with a reason: the run succeeds, your pipeline decides what to retry.
How does it compare?
Prices and figures as published on Apify Store on 29 September 2026:
| This Actor | pintostudio/youtube-transcript-scraper | starvibe/youtube-video-transcript | supreme_coder/youtube-transcript-scraper | |
|---|---|---|---|---|
| Price per transcript | $0.005 | $0.01 | $0.005 | $0.001 |
| Videos per run | unlimited (URLs, ids, playlists, channels) | 1 | 1 video or 1 channel | many |
| Failure handling | one row per video with a reason, never charged | run fails on private, age-restricted, live or deleted videos (about 15 % of runs) | — | error codes |
| Formats | segments, text, SRT, VTT | segments | text | JSON, text, SRT, VTT |
| Original-language detection | yes (dubbed videos handled) | English by default | yes | languages list |
| Punctuated text and paragraphs | yes | no | — | no |
| AI summary, chapters, insights | yes | no | no | no |
| LLM translation | yes | no | no | YouTube machine translation |
| Metadata | 14 fields | none | yes | yes |
Use cases
- Content creators and editors — chapters, summaries and quotes for descriptions, newsletters and social posts, in the language of your audience.
- Marketing and competitive intelligence — transcribe a competitor's channel, extract promoted products, calls to action and sponsors, track topics over time.
- Research and journalism — searchable, citable transcripts with timestamps from long interviews, lectures and hearings.
- RAG pipelines and AI agents — clean text with stable ids, ready to embed; see the LangChain example below and the MCP setup.
- Localization — SRT/VTT files and full translations of a whole playlist in one run.
- Monitoring — schedule the Actor on a channel and get a weekly digest of what was said.
Using the Actor from code
Run it and read the results with the Apify API:
curl -X POST "https://api.apify.com/v2/acts/maxencebernerd~youtube-transcript-scraper/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \-H "Content-Type: application/json" \-d '{"videoUrls": ["https://www.youtube.com/watch?v=9Ff4S1FJdRM"], "outputFormat": "text"}'
JavaScript (npm install apify-client):
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });const run = await client.actor('maxencebernerd/youtube-transcript-scraper').call({channelUrls: ['@mkbhd'],maxVideosPerChannel: 20,enableSummary: true,enableChapters: true,});const { items } = await client.dataset(run.defaultDatasetId).listItems();for (const video of items.filter((v) => v.status === 'ok')) {console.log(video.title, '\n', video.chaptersText, '\n', video.summary.oneLine);}
Python (pip install apify-client):
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_TOKEN>")run = client.actor("maxencebernerd/youtube-transcript-scraper").call(run_input={"playlistUrls": ["https://www.youtube.com/playlist?list=PLFgquLnL59alCl_2TQvOiD5Vgm1hCaGSI"],"maxVideosPerChannel": 100,"outputFormat": "text",})for video in client.dataset(run["defaultDatasetId"]).iterate_items():if video["status"] == "ok":print(video["title"], video["wordCount"], "words")
MCP: use it from Claude, Cursor or any AI agent
The Actor is exposed through the Apify MCP server. Add it to your MCP client (example for Claude Desktop):
{"mcpServers": {"apify": {"url": "https://mcp.apify.com?tools=maxencebernerd/youtube-transcript-scraper","headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }}}}
Then ask your agent: "Get the transcript and chapters of the last 5 videos of @mkbhd and tell me which products he recommends." The agent calls the Actor, waits for the dataset and reads the rows. The input schema is self-describing, so agents pick the right options on their own.
Load a YouTube channel into a RAG index (LangChain)
from apify_client import ApifyClientfrom langchain_core.documents import Documentfrom langchain_text_splitters import RecursiveCharacterTextSplitterfrom langchain_community.vectorstores import FAISSfrom langchain_openai import OpenAIEmbeddings # or any embeddings you useclient = ApifyClient("<YOUR_APIFY_TOKEN>")run = client.actor("maxencebernerd/youtube-transcript-scraper").call(run_input={"channelUrls": ["@hubermanlab"],"maxVideosPerChannel": 50,"outputFormat": "text",})documents = [Document(page_content=video["text"],metadata={"source": video["videoUrl"], "title": video["title"], "channel": video["channelName"], "published": video["publishedAt"]},)for video in client.dataset(run["defaultDatasetId"]).iterate_items()if video["status"] == "ok"]chunks = RecursiveCharacterTextSplitter(chunk_size=1500, chunk_overlap=200).split_documents(documents)index = FAISS.from_documents(chunks, OpenAIEmbeddings())print(index.similarity_search("what does he say about morning sunlight?", k=3))
Keep segments instead of text when you need timestamped citations: each chunk can carry the start of its first segment and link to videoUrl?t=<seconds>. The same dataset loads through the LangChain and LlamaIndex Apify dataset loaders (examples/langchain_rag.py in the repository has the complete script).
Integrations
Every run can trigger Slack, Zapier, Make, n8n, Google Sheets, Google Drive, Notion, webhooks or any other Apify integration. Schedule the Actor from the Schedules tab to transcribe a channel's new videos every day or week.
FAQ
What about videos without subtitles? They come back with status: no_transcript and errorReason: no_captions, without charge. The Actor only reads captions that exist on YouTube (manual or auto-generated); it does not run speech recognition on the audio.
Which languages are supported? Every language YouTube offers captions in. Leave language empty to get the original spoken language (detected from the video's audio track, so dubbed videos and videos with English subtitles switched on by default still return the spoken language). Set language to force a track; if the video does not have it, you get language_unavailable and the list of available tracks, never a transcript in another language.
Manual vs auto-generated captions: what is the difference? Manual tracks are uploaded by the creator (or a translator): accurate, punctuated, sometimes in several languages. Auto tracks are produced by YouTube's speech recognition: available on most videos, occasionally wrong on names and technical terms. transcriptType tells you which one you got; preferManual (default) picks the manual track when both exist.
Are Shorts supported? Yes, by URL and from the channel's Shorts tab (includeShorts). They are processed like any other video.
What about live streams? A stream in progress has no transcript yet (live_in_progress); once YouTube has processed the recording, the replay works like a normal video.
How fast is it and is there a volume limit? About one video per second per run at the default concurrency of 5; a 1,000-video channel takes roughly 15 minutes, less with maxConcurrency up to 20. There is no hard limit on the number of videos per run; keep runs under a few thousand videos so that a failure does not cost you a long rerun, and use Maximum cost per run as a safety net.
How do I use it in a RAG pipeline? Use outputFormat: "text" for embeddings, or segments when you want timestamped citations. See the LangChain example above.
Is it legal? The Actor reads publicly available captions and metadata, the same data any visitor sees on YouTube, without logging in or bypassing any restriction. You are responsible for how you use the transcripts (captions are the creators' content). This Actor is not affiliated with YouTube or Google.
What if YouTube blocks the run? Blocks are retried on fresh residential sessions; if a video still cannot be fetched it is reported as error / blocked_by_youtube and not charged, so you can rerun only those ids. The failure rate is monitored and the Actor is updated when YouTube changes something.
Can I request a feature? Yes, through the Issues tab. Planned: comparison of several videos, scheduled channel watch with a weekly summary, export to Notion and Google Docs.