YouTube Transcript Scraper + AI Summary & Chapters avatar

YouTube Transcript Scraper + AI Summary & Chapters

Pricing

from $5.00 / 1,000 transcript fetcheds

Go to Apify Store
YouTube Transcript Scraper + AI Summary & Chapters

YouTube Transcript Scraper + AI Summary & Chapters

Get YouTube transcripts (manual or auto-generated subtitles, any language) with timestamps from video URLs, playlists or channels, as text, SRT, VTT or segments. Optional AI summary, chapters, key insights and translation. 98%+ success rate. Pay per transcript.

Pricing

from $5.00 / 1,000 transcript fetcheds

Rating

0.0

(0)

Developer

Maxence Bernerd

Maxence Bernerd

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Categories

Share

YouTube Transcript Scraper gets the transcript of any public YouTube video with timestamps: manual or auto-generated subtitles, in any language, from video URLs, ids, playlists or whole channels. Every row carries the transcript as timestamped segments, plain text with punctuation and paragraphs, SRT and VTT, plus the video metadata (title, channel, duration, views, tags, publish date). Turn on the AI layer to add a three-level summary, YouTube-ready chapters, key insights (entities, quotes, topics, sentiment, calls to action) and a full translation. Built for reliability: 98%+ success rate measured on a mixed sample (long videos, Shorts, non-English, live replays), videos without captions are reported cleanly instead of failing the run, and you pay $0.005 per transcript, nothing for videos where no transcript exists.

What data does YouTube Transcript Scraper extract?

One dataset row per video, flat enough for CSV and complete enough for AI agents and RAG pipelines:

FieldDescription
videoId, videoUrl, sourceStable YouTube id, canonical URL, and where the video came from (input, playlist:<id>, channel:<handle>)
status, errorReason, errorMessageok, no_transcript, unavailable, restricted or error, with a machine-readable reason (see the edge-case table below)
title, channelName, channelId, channelUrl, durationSeconds, viewCount, description, tags, thumbnailUrl, isLiveVideo metadata, always included
publishedAt, likeCount, category, isUnlistedExtended metadata (with includeMetadata)
transcriptLanguage, transcriptTypeLanguage of the transcript and whether it is manual (uploaded by the creator) or auto (YouTube speech recognition)
availableLanguages, translatableLanguagesEvery caption track the video offers (code, name, type) and the languages YouTube can machine-translate to
segments[{ start, end, text }] in seconds, one entry per caption line
textContinuous transcript with punctuation and paragraphs (reconstructed when the auto track has none)
srt, vttReady-to-use subtitle files
wordCountNumber of words in the transcript
summaryoneLine, paragraph, keyPoints (with enableSummary)
chapters, chaptersText4 to 15 timestamped chapters as objects and in YouTube description format (with enableChapters)
insightsEntities, timestamped quotes, questions, topics, sentiment, calls to action, links and promotions (with enableInsights)
translationFull translation of text (with translateTo)
aiBillingUnits, scrapedAt20-minute blocks billed for the AI events, and the scrape timestamp

No account is used, no cookies, no personal data: only public captions and public metadata.

How to get YouTube transcripts

  1. Paste video URLs (any form: watch?v=, youtu.be, Shorts, live replays, embeds), playlist URLs or channel handles (@mkbhd) in the input.
  2. Optionally pick a caption language, the output formats, and the AI options you want.
  3. Click Start. Videos are processed in parallel; a 50-video channel takes about a minute.
  4. Read the results in the Output tab (views: Transcripts, Summaries, Chapters, Metadata) or download them as JSON, CSV, Excel or through the API.

Input examples

Transcripts of three videos, original language, every format:

{
"videoUrls": [
"https://www.youtube.com/watch?v=9Ff4S1FJdRM",
"https://youtu.be/dQw4w9WgXcQ",
"https://www.youtube.com/shorts/R6yNUnRXZ64"
]
}

The 50 most popular videos of a channel, English captions only, with summaries and chapters:

{
"channelUrls": ["@mkbhd"],
"maxVideosPerChannel": 50,
"channelSort": "popular",
"language": "en",
"outputFormat": "text",
"enableSummary": true,
"enableChapters": true
}

A whole playlist in SRT, plus insights and a French translation of each transcript:

{
"playlistUrls": ["https://www.youtube.com/playlist?list=PLFgquLnL59alCl_2TQvOiD5Vgm1hCaGSI"],
"maxVideosPerChannel": 100,
"outputFormat": "srt",
"enableInsights": true,
"translateTo": "fr",
"summaryLanguage": "fr"
}

Input fields: videoUrls, videoIds, playlistUrls, channelUrls, maxVideosPerChannel (default 50), channelSort (latest | popular), includeShorts, language (empty = original language), preferManual (default true), fallbackToAutoGenerated (default true), outputFormat (all | segments | text | srt | vtt), includeMetadata, enableSummary, summaryLanguage, enableChapters, enableInsights, translateTo, maxConcurrency, proxyConfiguration.

Output example

A real row (transcript fields shortened):

{
"videoId": "9Ff4S1FJdRM",
"videoUrl": "https://www.youtube.com/watch?v=9Ff4S1FJdRM",
"source": "input",
"status": "ok",
"errorReason": null,
"title": "Is AI Turning Us All into the Same Person? | Sandra Matz | TED",
"channelName": "TED",
"channelId": "UCAuUUnT6oDeKwE6v1NGQxug",
"channelUrl": "https://www.youtube.com/channel/UCAuUUnT6oDeKwE6v1NGQxug",
"publishedAt": "2026-09-28T15:00:04.000Z",
"durationSeconds": 817,
"viewCount": 76770,
"likeCount": 656,
"category": "People & Blogs",
"tags": ["TEDTalk", "TEDTalks", "TED Talk", "TED Talks", "TED"],
"thumbnailUrl": "https://i.ytimg.com/vi_webp/9Ff4S1FJdRM/sddefault.webp",
"transcriptLanguage": "en",
"transcriptType": "manual",
"availableLanguages": [
{ "code": "ar", "name": "Arabic", "type": "manual" },
{ "code": "en", "name": "English", "type": "manual" },
{ "code": "es", "name": "Spanish", "type": "manual" }
],
"translatableLanguages": ["ar", "zh-Hant", "nl", "en", "fr", "de", "…"],
"segments": [
{ "start": 4.668, "end": 8.046, "text": "People worry about AI for all kinds of reasons." },
{ "start": 8.046, "end": 10.84, "text": "It's polarizing, it spreads misinformation," },
{ "start": 10.84, "end": 12.717, "text": "it's coming for our jobs." }
],
"text": "People worry about AI for all kinds of reasons. It's polarizing, it spreads misinformation, it's coming for our jobs. And those are all good reasons to be nervous. But what really keeps me up at night is something else. …",
"srt": "1\n00:00:04,668 --> 00:00:08,046\nPeople worry about AI for all kinds of reasons.\n\n2\n…",
"vtt": "WEBVTT\n\n00:00:04.668 --> 00:00:08.046\nPeople worry about AI for all kinds of reasons.\n\n…",
"wordCount": 2157,
"summary": null,
"chapters": null,
"insights": null,
"translation": null,
"scrapedAt": "2026-09-29T18:41:59.575Z"
}

A video without captions is a valid row, not a failed run, and it is not charged:

{
"videoId": "Ct6BUPvE2sM",
"status": "no_transcript",
"errorReason": "no_captions",
"errorMessage": "This video has no captions (neither manual nor auto-generated).",
"title": "PIKOTARO - PPAP (Pen Pineapple Apple Pen) (Long Version) (Official Video) [Ultra Records]",
"availableLanguages": []
}

Edge cases: what you get for each kind of video

CasestatuserrorReasonCharged
Video with manual or auto captionsok—yes
Shorts, live replays, premieres already airedok—yes
Video without any captionsno_transcriptno_captionsno
Requested language not available (other languages listed in availableLanguages)no_transcriptlanguage_unavailableno
Only auto captions and fallbackToAutoGenerated is offno_transcriptno_captionsno
Live stream in progressunavailablelive_in_progressno
Upcoming premiere or scheduled streamunavailableupcomingno
Deleted video, wrong idunavailablevideo_not_foundno
Private videorestrictedprivateno
Age-restricted videorestrictedage_restrictedno
Members-only videorestrictedmembers_onlyno
Region-blocked videorestrictedregion_blockedno
YouTube blocked every retry (rare, retried on 4 fresh proxy sessions first)errorblocked_by_youtubeno

Restrictions are reported, never bypassed: the Actor uses no account and no cookies.

How much does it cost to get YouTube transcripts?

Pay per event, no subscription, no compute charges on top:

EventPriceWhen it is charged
transcript-fetched$0.005Per video whose transcript was obtained (status: ok)
metadata-fetched$0.0005Per video, with includeMetadata, when the extended fields (publish date, likes, category) were obtained
summary-generated$0.02 per 20-minute blockPer video with a summary
chapters-generated$0.02 per 20-minute blockPer video with chapters
insights-extracted$0.03 per 20-minute blockPer video with insights
translation-generated$0.04 per 1,000 wordsPer transcript translated (source words, rounded up)
Actor start$0.00005Per run

Videos up to 20 minutes count as one block; a 50-minute video counts as 3 blocks (the transcript itself is $0.005 whatever the length). Nothing is charged for videos without transcript, for AI results the model failed to produce, or for videos skipped because the run budget was reached. Set Maximum cost per run in the run options: the Actor stops cleanly at the limit and tells you in the log.

Real cost for typical 10-minute videos:

VideosTranscripts only+ extended metadata+ summary and chapters+ insights and everything above
100$0.50$0.55$4.55$7.55
1,000$5$5.50$45.50$75.50
10,000$50$55$455$755

A full translation of a 10-minute video (about 1,500 words) adds $0.08; of a 90-minute podcast (about 14,000 words) $0.56.

AI summary, chapters and insights

Each AI feature runs on Claude and returns validated JSON. Long videos are processed in 20-minute windows whose results are merged, so a 3-hour podcast gets one coherent summary and a single chapter list. Output language follows summaryLanguage (default: the transcript language); quotes stay in the original language.

Real output for the TED talk above (enableSummary):

{
"oneLine": "AI systems optimized for engagement narrow human preferences and suppress exploration, risking a flattened human experience unless companies redesign incentives to encourage discovery.",
"paragraph": "AI recommendation systems are trained to optimize for short-term engagement and exploitation of known preferences rather than exploration and discovery. Studies show that when people rely on AI for guidance, their preferences become more normative and less diverse, their creative output becomes less unique, and their choices converge toward sameness. An experiment with ChatGPT recommending Baskin-Robbins flavors 100 times resulted in 96 recommendations of just two popular flavors. …",
"keyPoints": [
"ChatGPT recommended same two ice cream flavors 96 of 100 times, showing algorithmic preference for popularity over diversity",
"AI systems trained to optimize for short-term engagement and exploitation, not exploration or discovery",
"People using AI guidance show less diverse preferences, less unique creative output, and more conformist choices",
"Companies should implement adjustment dials letting users control exploration versus exploitation levels in recommendations"
]
}

Chapters (enableChapters), as chaptersText, ready to paste in the video description:

0:00 AI Makes Us Boring
1:07 Exploitation vs. Exploration Trade-off
2:39 How AI Optimizes for Safety
5:19 The Flattening of Human Experience
6:53 AI Narrows Your Taste
9:27 Reclaiming Risk and Discovery
10:30 The Dial: Balancing Exploration
12:27 Rewarding AI for Smart Risks

Insights (enableInsights) on a 30-minute Spanish vlog with Tony Hawk, output in English:

{
"entities": { "people": ["Tony Hawk", "Rubius"], "brands": ["PlayStation", "Gamecube", "Switch"], "products": ["Tony Hawk Pro Skater 3+4", "motorized skateboard"], "places": ["San Diego", "Spain"] },
"quotes": [{ "start": 105, "timestamp": "1:45", "text": "if you never try, you never get it" }],
"questions": ["How did the creator get Tony Hawk to respond to his DM?", "Can the creator win against Tony Hawk in the Tony Hawk Pro Skater video game?"],
"topics": ["skateboarding", "childhood nostalgia", "video game competition", "recovery from injury"],
"sentiment": "positive",
"callsToAction": ["Subscribe to the channel", "Like the video"],
"linksAndPromotions": ["Tony Hawk Pro Skater 3+4 game", "Tony Hawk's skateboard signed by Tony Hawk"]
}

Translation (translateTo) returns { language, text, wordCount, units } with the paragraphs of text preserved.

Reliability: 98%+ success rate, measured

On a test set of 31 real videos covering every edge case (multilingual manual captions, French, Spanish, Japanese, German, Portuguese, Hindi and Korean videos, two podcasts over 3 hours, a 1-hour live replay, four Shorts, videos without captions, a members-only video, two live streams, a private video, an age-restricted video, two deleted ids), the Actor returned the expected status for 31 videos out of 31, with zero errors, both from a residential connection and from the Apify platform. The same set run through the most-used transcript Actor on the Store produced identical caption content on every video that has captions, but failed the run on the private, age-restricted, live and deleted videos.

How it stays reliable:

  • Transcripts come straight from YouTube's caption endpoints through residential proxy sessions, one sticky session per video, rotated as soon as YouTube pushes back (rate limit, captcha, bot check). Blocks are retried on fresh sessions before a video is ever marked as error.
  • Two independent client identities are tried before concluding that a video has no captions.
  • The requested language is never silently replaced by another one.
  • Every failure is a row with a reason: the run succeeds, your pipeline decides what to retry.

How does it compare?

Prices and figures as published on Apify Store on 29 September 2026:

This Actorpintostudio/youtube-transcript-scraperstarvibe/youtube-video-transcriptsupreme_coder/youtube-transcript-scraper
Price per transcript$0.005$0.01$0.005$0.001
Videos per rununlimited (URLs, ids, playlists, channels)11 video or 1 channelmany
Failure handlingone row per video with a reason, never chargedrun fails on private, age-restricted, live or deleted videos (about 15 % of runs)—error codes
Formatssegments, text, SRT, VTTsegmentstextJSON, text, SRT, VTT
Original-language detectionyes (dubbed videos handled)English by defaultyeslanguages list
Punctuated text and paragraphsyesno—no
AI summary, chapters, insightsyesnonono
LLM translationyesnonoYouTube machine translation
Metadata14 fieldsnoneyesyes

Use cases

  • Content creators and editors — chapters, summaries and quotes for descriptions, newsletters and social posts, in the language of your audience.
  • Marketing and competitive intelligence — transcribe a competitor's channel, extract promoted products, calls to action and sponsors, track topics over time.
  • Research and journalism — searchable, citable transcripts with timestamps from long interviews, lectures and hearings.
  • RAG pipelines and AI agents — clean text with stable ids, ready to embed; see the LangChain example below and the MCP setup.
  • Localization — SRT/VTT files and full translations of a whole playlist in one run.
  • Monitoring — schedule the Actor on a channel and get a weekly digest of what was said.

Using the Actor from code

Run it and read the results with the Apify API:

curl -X POST "https://api.apify.com/v2/acts/maxencebernerd~youtube-transcript-scraper/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
-H "Content-Type: application/json" \
-d '{"videoUrls": ["https://www.youtube.com/watch?v=9Ff4S1FJdRM"], "outputFormat": "text"}'

JavaScript (npm install apify-client):

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('maxencebernerd/youtube-transcript-scraper').call({
channelUrls: ['@mkbhd'],
maxVideosPerChannel: 20,
enableSummary: true,
enableChapters: true,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const video of items.filter((v) => v.status === 'ok')) {
console.log(video.title, '\n', video.chaptersText, '\n', video.summary.oneLine);
}

Python (pip install apify-client):

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("maxencebernerd/youtube-transcript-scraper").call(run_input={
"playlistUrls": ["https://www.youtube.com/playlist?list=PLFgquLnL59alCl_2TQvOiD5Vgm1hCaGSI"],
"maxVideosPerChannel": 100,
"outputFormat": "text",
})
for video in client.dataset(run["defaultDatasetId"]).iterate_items():
if video["status"] == "ok":
print(video["title"], video["wordCount"], "words")

MCP: use it from Claude, Cursor or any AI agent

The Actor is exposed through the Apify MCP server. Add it to your MCP client (example for Claude Desktop):

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com?tools=maxencebernerd/youtube-transcript-scraper",
"headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }
}
}
}

Then ask your agent: "Get the transcript and chapters of the last 5 videos of @mkbhd and tell me which products he recommends." The agent calls the Actor, waits for the dataset and reads the rows. The input schema is self-describing, so agents pick the right options on their own.

Load a YouTube channel into a RAG index (LangChain)

from apify_client import ApifyClient
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import FAISS
from langchain_openai import OpenAIEmbeddings # or any embeddings you use
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("maxencebernerd/youtube-transcript-scraper").call(run_input={
"channelUrls": ["@hubermanlab"],
"maxVideosPerChannel": 50,
"outputFormat": "text",
})
documents = [
Document(
page_content=video["text"],
metadata={"source": video["videoUrl"], "title": video["title"], "channel": video["channelName"], "published": video["publishedAt"]},
)
for video in client.dataset(run["defaultDatasetId"]).iterate_items()
if video["status"] == "ok"
]
chunks = RecursiveCharacterTextSplitter(chunk_size=1500, chunk_overlap=200).split_documents(documents)
index = FAISS.from_documents(chunks, OpenAIEmbeddings())
print(index.similarity_search("what does he say about morning sunlight?", k=3))

Keep segments instead of text when you need timestamped citations: each chunk can carry the start of its first segment and link to videoUrl?t=<seconds>. The same dataset loads through the LangChain and LlamaIndex Apify dataset loaders (examples/langchain_rag.py in the repository has the complete script).

Integrations

Every run can trigger Slack, Zapier, Make, n8n, Google Sheets, Google Drive, Notion, webhooks or any other Apify integration. Schedule the Actor from the Schedules tab to transcribe a channel's new videos every day or week.

FAQ

What about videos without subtitles? They come back with status: no_transcript and errorReason: no_captions, without charge. The Actor only reads captions that exist on YouTube (manual or auto-generated); it does not run speech recognition on the audio.

Which languages are supported? Every language YouTube offers captions in. Leave language empty to get the original spoken language (detected from the video's audio track, so dubbed videos and videos with English subtitles switched on by default still return the spoken language). Set language to force a track; if the video does not have it, you get language_unavailable and the list of available tracks, never a transcript in another language.

Manual vs auto-generated captions: what is the difference? Manual tracks are uploaded by the creator (or a translator): accurate, punctuated, sometimes in several languages. Auto tracks are produced by YouTube's speech recognition: available on most videos, occasionally wrong on names and technical terms. transcriptType tells you which one you got; preferManual (default) picks the manual track when both exist.

Are Shorts supported? Yes, by URL and from the channel's Shorts tab (includeShorts). They are processed like any other video.

What about live streams? A stream in progress has no transcript yet (live_in_progress); once YouTube has processed the recording, the replay works like a normal video.

How fast is it and is there a volume limit? About one video per second per run at the default concurrency of 5; a 1,000-video channel takes roughly 15 minutes, less with maxConcurrency up to 20. There is no hard limit on the number of videos per run; keep runs under a few thousand videos so that a failure does not cost you a long rerun, and use Maximum cost per run as a safety net.

How do I use it in a RAG pipeline? Use outputFormat: "text" for embeddings, or segments when you want timestamped citations. See the LangChain example above.

Is it legal? The Actor reads publicly available captions and metadata, the same data any visitor sees on YouTube, without logging in or bypassing any restriction. You are responsible for how you use the transcripts (captions are the creators' content). This Actor is not affiliated with YouTube or Google.

What if YouTube blocks the run? Blocks are retried on fresh residential sessions; if a video still cannot be fetched it is reported as error / blocked_by_youtube and not charged, so you can rerun only those ids. The failure rate is monitored and the Actor is updated when YouTube changes something.

Can I request a feature? Yes, through the Issues tab. Planned: comparison of several videos, scheduled channel watch with a weekly summary, export to Notion and Google Docs.