YouTube Scraper — Videos, Comments & Transcripts for AI
Pricing
from $2.68 / 1,000 videos
YouTube Scraper — Videos, Comments & Transcripts for AI
Scrape YouTube channels, playlists, videos and search results: full metadata, top comments and timestamped transcripts in one run. Built for AI, RAG and LLM pipelines. No API key or login needed.
Pricing
from $2.68 / 1,000 videos
Rating
0.0
(0)
Developer
Eonix Pvt Ltd
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
YouTube Transcript, Comments & Video Scraper for AI and RAG
Turn any YouTube channel, playlist, search or video into clean, AI-ready data: full video details, the top comments, and complete transcripts with timestamps — all in one run, with no API key and no login. Every video comes back as a single tidy record, so you can drop it straight into ChatGPT, Claude, a vector database or a spreadsheet.
Headline use case: build a RAG knowledge base from any YouTube channel. Paste a channel link, press Start, and a few minutes later you have every talk, lecture or tutorial as searchable text with timestamps that link back to the exact second in the video.
What you get for every video
- Video details — title, description, channel, publish date, duration, views, likes, comment count, tags, category, thumbnail, Short / live flags
- Transcript — the full captions as timestamped segments plus one
fullTextstring and a word count, ready for chunking and embeddings - Comments — the top comments with author, likes, reply count, and pinned / creator-hearted flags
- Clear status fields —
transcriptStatusandcommentsStatustell you exactly why something is missing (no captions, age-restricted, comments turned off…) instead of silently returning nothing
Use cases
- AI agents & RAG — feed transcripts of an entire channel into a vector store and let your chatbot answer questions with citations to the exact minute of the video.
- Market research — read what thousands of viewers actually say about a product, competitor or topic; spot recurring questions and complaints in the comments.
- Lead generation & creator research — find the channels and videos that dominate a niche, how often they publish and how engaged their audience is.
- Content monitoring — schedule a daily run with “Uploaded after: 1 day” to get every new video (and what people say about it) from the channels you follow.
- Content repurposing — turn talks and podcasts into blog posts, newsletters, show notes or training data.
Sample output
A real record from a test run (trimmed for readability — real records include up to your comment limit and the full transcript):
{"videoId": "SVTPv4sI_Jc","url": "https://www.youtube.com/watch?v=SVTPv4sI_Jc","title": "The CIA's new tech doesn't make sense","channelName": "Veritasium","publishedAt": "2026-05-03T19:26:56.000Z","durationSeconds": 1278,"viewCount": 3140584,"likeCount": 81902,"commentCount": 6008,"tags": ["veritasium", "science", "physics"],"isShort": false,"commentsStatus": "OK","transcriptStatus": "OK","comments": [{"author": "@veritasium","text": "Get all sides of every story at https://ground.news/Ve - and read the news with a data-driven approach to spot media bias for yourself. Subscribe through our link for 40% off the unlimited access Vantage Plan.","likeCount": 306,"replyCount": 64,"isPinned": true}],"transcript": {"language": "en","isAutoGenerated": false,"segments": [{"start": 0.035,"duration": 2.045,"text": "- Could the CIA really track your heartbeat"},{"start": 2.08,"duration": 1.68,"text": "from kilometers away?"}],"fullText": "- Could the CIA really track your heartbeat from kilometers away? On April 3rd, 2026, Iranian forces shot down an American fighter plane just over Isfahan. Insi…","wordCount": 3698}}
How much does it cost?
You pay only for what you get — no subscription and no separate compute bill:
| Event | Price | When it's charged |
|---|---|---|
| Actor start | $0.01 | Once per run |
| Video scraped | $0.002 | Each video in your results |
| Comment scraped | $0.0005 | Each comment included |
| Transcript scraped | $0.005 | Each video that comes back with a transcript |
Worked examples
- The default run (20 latest videos, 100 comments each, transcripts): $0.01 + 20 × $0.002 + 2,000 × $0.0005 + 20 × $0.005 = about $1.15.
- 100 videos + 100 comments each + transcripts: $0.01 + $0.20 + $5.00 + $0.50 = about $5.71.
- 1,000 videos, transcripts only (comments off): $0.01 + $2.00 + $5.00 = about $7.01.
- 1,000 videos + 100 comments each + transcripts: $0.01 + $2.00 + $50.00 + $5.00 = about $57.01.
Comments are the biggest cost driver. If you only need text for AI, switch Scrape comments off or lower Max comments per video.
You are never charged for a video that has no transcript (transcriptStatus other than OK), for videos that are skipped, or for anything beyond the Maximum cost per run you set in Apify — the scraper stops cleanly when your budget is reached and the run's STATS record shows "budgetReached": true.
Input
| Field | What it does | Default |
|---|---|---|
Start URLs (startUrls) | Channel (/@handle, /@handle/shorts, /@handle/streams, /channel/UC…), playlist, video (watch?v=, youtu.be, /shorts/) or search-results URLs | https://www.youtube.com/@veritasium |
Search queries (searchQueries) | Search terms, one per line | empty |
Max videos per source (maxVideos) | Videos to take from each channel, playlist or search | 20 |
Sort videos by (sortBy) | newest or popular (see FAQ for how this applies to search) | newest |
Uploaded after (uploadedAfter) | Only videos published on/after a date (2026-01-31) or within a period (30 days, 2 weeks) | no filter |
Scrape comments (scrapeComments) | Include top comments | on |
Max comments per video (maxCommentsPerVideo) | Comment limit per video (0 = none) | 100 |
Scrape transcripts (scrapeTranscripts) | Include the transcript | on |
Transcript languages (transcriptLanguages) | Preferred languages in order, e.g. en, es, pt-BR | ["en"] |
Proxy configuration (proxyConfiguration) | Apify Proxy settings; switch to RESIDENTIAL if you see blocks | Apify Proxy |
Max request retries (maxRequestRetries) | Retries per request, each with a fresh proxy session | 5 |
Max concurrency (maxConcurrency) | Videos processed in parallel | 10 |
Output fields
| Field | Type | Notes |
|---|---|---|
videoId, url, title, description | string | |
channelId, channelName, channelUrl | string | |
publishedAt | ISO-8601 date | exact publish time |
durationSeconds, viewCount, likeCount | number | |
commentCount | number | exact total when comments are scraped, otherwise YouTube's rounded figure |
tags | string[] | |
category, thumbnailUrl | string | |
isShort, isLive, isAgeRestricted | boolean | |
comments[] | array | commentId, author, authorChannelId, isAuthorChannelOwner, text, likeCount, publishedAt, publishedTimeText, replyCount, isPinned, isHearted |
commentsStatus | string | OK, NOT_REQUESTED, DISABLED, UNAVAILABLE, BUDGET_LIMIT, ERROR |
transcript | object or null | language, languageName, isAutoGenerated, isTranslated, segments[{start, duration, text}], fullText, wordCount (times in seconds) |
transcriptStatus | string | OK, NOT_REQUESTED, NO_CAPTIONS, LANGUAGE_NOT_AVAILABLE, AGE_RESTRICTED, UNAVAILABLE, BUDGET_LIMIT, ERROR |
sourceType, sourceInput | string | which of your inputs produced this video |
scrapedAt | ISO-8601 date |
Missing values are always null (never empty strings), numbers are real numbers, and each video appears only once per run even if several of your inputs contain it.
How to use it from code and no-code tools
Replace <username>/youtube-channel-comments-transcripts below with the Actor ID shown on this page, and APIFY_TOKEN with your token from Apify Console → Settings → API & Integrations.
REST API
curl -X POST "https://api.apify.com/v2/acts/<username>~youtube-channel-comments-transcripts/run-sync-get-dataset-items?token=$APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"startUrls":[{"url":"https://www.youtube.com/@veritasium"}],"maxVideos":10,"scrapeComments":false}'
Python
from apify_client import ApifyClientclient = ApifyClient("APIFY_TOKEN")run = client.actor("<username>/youtube-channel-comments-transcripts").call(run_input={"startUrls": [{"url": "https://www.youtube.com/@veritasium"}],"maxVideos": 10,"maxCommentsPerVideo": 20,})for video in client.dataset(run["defaultDatasetId"]).iterate_items():print(video["title"], video["transcriptStatus"], len(video["comments"]))
Node.js
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('<username>/youtube-channel-comments-transcripts').call({searchQueries: ['retrieval augmented generation explained'],maxVideos: 10,sortBy: 'popular',});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items.map((v) => v.title));
LangChain (build a RAG index from a channel)
from langchain_apify import ApifyWrapperfrom langchain_core.documents import Documentfrom langchain_text_splitters import RecursiveCharacterTextSplitterloader = ApifyWrapper().call_actor(actor_id="<username>/youtube-channel-comments-transcripts",run_input={"startUrls": [{"url": "https://www.youtube.com/@veritasium"}], "maxVideos": 50, "scrapeComments": False},dataset_mapping_function=lambda v: Document(page_content=(v.get("transcript") or {}).get("fullText") or v.get("description") or "",metadata={"source": v["url"], "title": v["title"], "published": v["publishedAt"]},),)chunks = RecursiveCharacterTextSplitter(chunk_size=1500, chunk_overlap=150).split_documents(loader.load())# → pass `chunks` to any vector store (Chroma, Pinecone, pgvector…)
Tip: to cite the exact moment in a video, chunk on transcript.segments instead and store url + "&t=" + int(segment["start"]) as the source.
LlamaIndex
from llama_index.core import Document, VectorStoreIndexfrom llama_index.readers.apify import ApifyActorreader = ApifyActor("APIFY_TOKEN")documents = reader.load_data(actor_id="<username>/youtube-channel-comments-transcripts",run_input={"startUrls": [{"url": "https://www.youtube.com/@veritasium"}], "maxVideos": 50, "scrapeComments": False},dataset_mapping_function=lambda v: Document(text=(v.get("transcript") or {}).get("fullText") or "",metadata={"url": v["url"], "title": v["title"]},),)index = VectorStoreIndex.from_documents(documents)print(index.as_query_engine().query("What did they say about the Enigma machine?"))
Make, n8n and Zapier
- Make: add the Apify → Run an Actor module, pick this Actor, paste your input JSON, then use Apify → Get Dataset Items to loop over videos.
- n8n: use the official Apify node → Run Actor and get dataset, select this Actor and map
transcript.fullTextinto your next step (e.g. an OpenAI or vector-store node). - Zapier: use the Apify app → Run Actor action, then Find Last Run Dataset Items.
MCP (Claude, Cursor, VS Code and other AI agents)
Let your AI assistant call this scraper as a tool through the Apify MCP server. Add this server URL to your MCP client (you sign in with your Apify account):
https://mcp.apify.com?tools=actors,docs,<username>/youtube-channel-comments-transcripts
Then just ask: “Get the transcripts of the last 10 videos on @veritasium and summarize the main ideas.”
FAQ
Do I need a YouTube API key or account? No. The scraper reads the same public data the YouTube website shows to a logged-out visitor.
Which transcript do I get? For each language in Transcript languages (in order) it takes human-made captions first, then auto-generated ones. If none of your languages exist, it asks YouTube to machine-translate into your first language; if YouTube refuses the translation it returns the original-language transcript, clearly marked by language and isTranslated: false.
How does “Sort by” work for search queries? In 2025 YouTube removed the “upload date” sort from search. popular prioritizes popular videos; newest uses YouTube's standard relevance order. To get recent results for a search, combine it with Uploaded after.
Why is commentCount sometimes a round number? When comments are scraped, it's the exact total YouTube shows. When comments are off, YouTube only exposes a rounded figure (e.g. 5.2K → 5200).
How exact are comment dates? YouTube only shows relative times for comments (“3 days ago”), so publishedAt on comments is an approximation; the original text is kept in publishedTimeText. Video publishedAt is exact.
Are replies included? Not in this version — you get top-level comments with their replyCount.
Why did I get fewer videos than “Max videos”? The channel/playlist/search simply has fewer (matching) videos, some were unavailable, your Uploaded after filter excluded them, or your Maximum cost per run was reached (check STATS → budgetReached).
The run log shows “blocked” errors. YouTube occasionally challenges datacenter IPs. The scraper retries with new sessions automatically; if blocks persist, set the proxy group to RESIDENTIAL.
Limitations
- Top-level comments only (no reply threads); comments are in YouTube's “Top comments” order.
- Age-restricted videos return full metadata but no transcript or comments (YouTube requires a signed-in adult account for those).
- Live streams in progress have no transcript, and their live chat is not collected.
- Private, members-only and removed videos are skipped and counted in
STATS.videosUnavailable. - Data is returned in English locale (
hl=en, US region) regardless of the transcript language you choose.
Is it legal to scrape YouTube?
This Actor only collects publicly available data that anyone can see without logging in — it does not bypass logins, paywalls or age gates. That said, you are responsible for how you use the data: comply with YouTube's Terms of Service and with data-protection laws such as the GDPR and CCPA. Comments contain personal data (usernames and what people wrote); only collect what you need, have a legitimate purpose, don't use it for spam or profiling individuals, and respect copyright when re-publishing transcripts. If in doubt, consult a lawyer. Read more in Apify's guide: Is web scraping legal?
Support
Found a bug or need a feature (reply threads, more fields, other sort orders)? Open an issue on the Actor's Issues tab — we usually reply within one business day.