YouTube Scraper — Videos, Comments & Transcripts for AI avatar

YouTube Scraper — Videos, Comments & Transcripts for AI

Pricing

from $2.68 / 1,000 videos

Go to Apify Store
YouTube Scraper — Videos, Comments & Transcripts for AI

YouTube Scraper — Videos, Comments & Transcripts for AI

Scrape YouTube channels, playlists, videos and search results: full metadata, top comments and timestamped transcripts in one run. Built for AI, RAG and LLM pipelines. No API key or login needed.

Pricing

from $2.68 / 1,000 videos

Rating

0.0

(0)

Developer

Eonix Pvt Ltd

Eonix Pvt Ltd

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

YouTube Transcript, Comments & Video Scraper for AI and RAG

Turn any YouTube channel, playlist, search or video into clean, AI-ready data: full video details, the top comments, and complete transcripts with timestamps — all in one run, with no API key and no login. Every video comes back as a single tidy record, so you can drop it straight into ChatGPT, Claude, a vector database or a spreadsheet.

Headline use case: build a RAG knowledge base from any YouTube channel. Paste a channel link, press Start, and a few minutes later you have every talk, lecture or tutorial as searchable text with timestamps that link back to the exact second in the video.

What you get for every video

  • Video details — title, description, channel, publish date, duration, views, likes, comment count, tags, category, thumbnail, Short / live flags
  • Transcript — the full captions as timestamped segments plus one fullText string and a word count, ready for chunking and embeddings
  • Comments — the top comments with author, likes, reply count, and pinned / creator-hearted flags
  • Clear status fields — transcriptStatus and commentsStatus tell you exactly why something is missing (no captions, age-restricted, comments turned off…) instead of silently returning nothing

Use cases

  • AI agents & RAG — feed transcripts of an entire channel into a vector store and let your chatbot answer questions with citations to the exact minute of the video.
  • Market research — read what thousands of viewers actually say about a product, competitor or topic; spot recurring questions and complaints in the comments.
  • Lead generation & creator research — find the channels and videos that dominate a niche, how often they publish and how engaged their audience is.
  • Content monitoring — schedule a daily run with “Uploaded after: 1 day” to get every new video (and what people say about it) from the channels you follow.
  • Content repurposing — turn talks and podcasts into blog posts, newsletters, show notes or training data.

Sample output

A real record from a test run (trimmed for readability — real records include up to your comment limit and the full transcript):

{
"videoId": "SVTPv4sI_Jc",
"url": "https://www.youtube.com/watch?v=SVTPv4sI_Jc",
"title": "The CIA's new tech doesn't make sense",
"channelName": "Veritasium",
"publishedAt": "2026-05-03T19:26:56.000Z",
"durationSeconds": 1278,
"viewCount": 3140584,
"likeCount": 81902,
"commentCount": 6008,
"tags": ["veritasium", "science", "physics"],
"isShort": false,
"commentsStatus": "OK",
"transcriptStatus": "OK",
"comments": [
{
"author": "@veritasium",
"text": "Get all sides of every story at https://ground.news/Ve - and read the news with a data-driven approach to spot media bias for yourself. Subscribe through our link for 40% off the unlimited access Vantage Plan.",
"likeCount": 306,
"replyCount": 64,
"isPinned": true
}
],
"transcript": {
"language": "en",
"isAutoGenerated": false,
"segments": [
{
"start": 0.035,
"duration": 2.045,
"text": "- Could the CIA really track your heartbeat"
},
{
"start": 2.08,
"duration": 1.68,
"text": "from kilometers away?"
}
],
"fullText": "- Could the CIA really track your heartbeat from kilometers away? On April 3rd, 2026, Iranian forces shot down an American fighter plane just over Isfahan. Insi…",
"wordCount": 3698
}
}

How much does it cost?

You pay only for what you get — no subscription and no separate compute bill:

EventPriceWhen it's charged
Actor start$0.01Once per run
Video scraped$0.002Each video in your results
Comment scraped$0.0005Each comment included
Transcript scraped$0.005Each video that comes back with a transcript

Worked examples

  • The default run (20 latest videos, 100 comments each, transcripts): $0.01 + 20 × $0.002 + 2,000 × $0.0005 + 20 × $0.005 = about $1.15.
  • 100 videos + 100 comments each + transcripts: $0.01 + $0.20 + $5.00 + $0.50 = about $5.71.
  • 1,000 videos, transcripts only (comments off): $0.01 + $2.00 + $5.00 = about $7.01.
  • 1,000 videos + 100 comments each + transcripts: $0.01 + $2.00 + $50.00 + $5.00 = about $57.01.

Comments are the biggest cost driver. If you only need text for AI, switch Scrape comments off or lower Max comments per video.

You are never charged for a video that has no transcript (transcriptStatus other than OK), for videos that are skipped, or for anything beyond the Maximum cost per run you set in Apify — the scraper stops cleanly when your budget is reached and the run's STATS record shows "budgetReached": true.

Input

FieldWhat it doesDefault
Start URLs (startUrls)Channel (/@handle, /@handle/shorts, /@handle/streams, /channel/UC…), playlist, video (watch?v=, youtu.be, /shorts/) or search-results URLshttps://www.youtube.com/@veritasium
Search queries (searchQueries)Search terms, one per lineempty
Max videos per source (maxVideos)Videos to take from each channel, playlist or search20
Sort videos by (sortBy)newest or popular (see FAQ for how this applies to search)newest
Uploaded after (uploadedAfter)Only videos published on/after a date (2026-01-31) or within a period (30 days, 2 weeks)no filter
Scrape comments (scrapeComments)Include top commentson
Max comments per video (maxCommentsPerVideo)Comment limit per video (0 = none)100
Scrape transcripts (scrapeTranscripts)Include the transcripton
Transcript languages (transcriptLanguages)Preferred languages in order, e.g. en, es, pt-BR["en"]
Proxy configuration (proxyConfiguration)Apify Proxy settings; switch to RESIDENTIAL if you see blocksApify Proxy
Max request retries (maxRequestRetries)Retries per request, each with a fresh proxy session5
Max concurrency (maxConcurrency)Videos processed in parallel10

Output fields

FieldTypeNotes
videoId, url, title, descriptionstring
channelId, channelName, channelUrlstring
publishedAtISO-8601 dateexact publish time
durationSeconds, viewCount, likeCountnumber
commentCountnumberexact total when comments are scraped, otherwise YouTube's rounded figure
tagsstring[]
category, thumbnailUrlstring
isShort, isLive, isAgeRestrictedboolean
comments[]arraycommentId, author, authorChannelId, isAuthorChannelOwner, text, likeCount, publishedAt, publishedTimeText, replyCount, isPinned, isHearted
commentsStatusstringOK, NOT_REQUESTED, DISABLED, UNAVAILABLE, BUDGET_LIMIT, ERROR
transcriptobject or nulllanguage, languageName, isAutoGenerated, isTranslated, segments[{start, duration, text}], fullText, wordCount (times in seconds)
transcriptStatusstringOK, NOT_REQUESTED, NO_CAPTIONS, LANGUAGE_NOT_AVAILABLE, AGE_RESTRICTED, UNAVAILABLE, BUDGET_LIMIT, ERROR
sourceType, sourceInputstringwhich of your inputs produced this video
scrapedAtISO-8601 date

Missing values are always null (never empty strings), numbers are real numbers, and each video appears only once per run even if several of your inputs contain it.

How to use it from code and no-code tools

Replace <username>/youtube-channel-comments-transcripts below with the Actor ID shown on this page, and APIFY_TOKEN with your token from Apify Console → Settings → API & Integrations.

REST API

curl -X POST "https://api.apify.com/v2/acts/<username>~youtube-channel-comments-transcripts/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"startUrls":[{"url":"https://www.youtube.com/@veritasium"}],"maxVideos":10,"scrapeComments":false}'

Python

from apify_client import ApifyClient
client = ApifyClient("APIFY_TOKEN")
run = client.actor("<username>/youtube-channel-comments-transcripts").call(run_input={
"startUrls": [{"url": "https://www.youtube.com/@veritasium"}],
"maxVideos": 10,
"maxCommentsPerVideo": 20,
})
for video in client.dataset(run["defaultDatasetId"]).iterate_items():
print(video["title"], video["transcriptStatus"], len(video["comments"]))

Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('<username>/youtube-channel-comments-transcripts').call({
searchQueries: ['retrieval augmented generation explained'],
maxVideos: 10,
sortBy: 'popular',
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((v) => v.title));

LangChain (build a RAG index from a channel)

from langchain_apify import ApifyWrapper
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter
loader = ApifyWrapper().call_actor(
actor_id="<username>/youtube-channel-comments-transcripts",
run_input={"startUrls": [{"url": "https://www.youtube.com/@veritasium"}], "maxVideos": 50, "scrapeComments": False},
dataset_mapping_function=lambda v: Document(
page_content=(v.get("transcript") or {}).get("fullText") or v.get("description") or "",
metadata={"source": v["url"], "title": v["title"], "published": v["publishedAt"]},
),
)
chunks = RecursiveCharacterTextSplitter(chunk_size=1500, chunk_overlap=150).split_documents(loader.load())
# → pass `chunks` to any vector store (Chroma, Pinecone, pgvector…)

Tip: to cite the exact moment in a video, chunk on transcript.segments instead and store url + "&t=" + int(segment["start"]) as the source.

LlamaIndex

from llama_index.core import Document, VectorStoreIndex
from llama_index.readers.apify import ApifyActor
reader = ApifyActor("APIFY_TOKEN")
documents = reader.load_data(
actor_id="<username>/youtube-channel-comments-transcripts",
run_input={"startUrls": [{"url": "https://www.youtube.com/@veritasium"}], "maxVideos": 50, "scrapeComments": False},
dataset_mapping_function=lambda v: Document(
text=(v.get("transcript") or {}).get("fullText") or "",
metadata={"url": v["url"], "title": v["title"]},
),
)
index = VectorStoreIndex.from_documents(documents)
print(index.as_query_engine().query("What did they say about the Enigma machine?"))

Make, n8n and Zapier

  • Make: add the Apify → Run an Actor module, pick this Actor, paste your input JSON, then use Apify → Get Dataset Items to loop over videos.
  • n8n: use the official Apify node → Run Actor and get dataset, select this Actor and map transcript.fullText into your next step (e.g. an OpenAI or vector-store node).
  • Zapier: use the Apify app → Run Actor action, then Find Last Run Dataset Items.

MCP (Claude, Cursor, VS Code and other AI agents)

Let your AI assistant call this scraper as a tool through the Apify MCP server. Add this server URL to your MCP client (you sign in with your Apify account):

https://mcp.apify.com?tools=actors,docs,<username>/youtube-channel-comments-transcripts

Then just ask: “Get the transcripts of the last 10 videos on @veritasium and summarize the main ideas.”

FAQ

Do I need a YouTube API key or account? No. The scraper reads the same public data the YouTube website shows to a logged-out visitor.

Which transcript do I get? For each language in Transcript languages (in order) it takes human-made captions first, then auto-generated ones. If none of your languages exist, it asks YouTube to machine-translate into your first language; if YouTube refuses the translation it returns the original-language transcript, clearly marked by language and isTranslated: false.

How does “Sort by” work for search queries? In 2025 YouTube removed the “upload date” sort from search. popular prioritizes popular videos; newest uses YouTube's standard relevance order. To get recent results for a search, combine it with Uploaded after.

Why is commentCount sometimes a round number? When comments are scraped, it's the exact total YouTube shows. When comments are off, YouTube only exposes a rounded figure (e.g. 5.2K → 5200).

How exact are comment dates? YouTube only shows relative times for comments (“3 days ago”), so publishedAt on comments is an approximation; the original text is kept in publishedTimeText. Video publishedAt is exact.

Are replies included? Not in this version — you get top-level comments with their replyCount.

Why did I get fewer videos than “Max videos”? The channel/playlist/search simply has fewer (matching) videos, some were unavailable, your Uploaded after filter excluded them, or your Maximum cost per run was reached (check STATS → budgetReached).

The run log shows “blocked” errors. YouTube occasionally challenges datacenter IPs. The scraper retries with new sessions automatically; if blocks persist, set the proxy group to RESIDENTIAL.

Limitations

  • Top-level comments only (no reply threads); comments are in YouTube's “Top comments” order.
  • Age-restricted videos return full metadata but no transcript or comments (YouTube requires a signed-in adult account for those).
  • Live streams in progress have no transcript, and their live chat is not collected.
  • Private, members-only and removed videos are skipped and counted in STATS.videosUnavailable.
  • Data is returned in English locale (hl=en, US region) regardless of the transcript language you choose.

This Actor only collects publicly available data that anyone can see without logging in — it does not bypass logins, paywalls or age gates. That said, you are responsible for how you use the data: comply with YouTube's Terms of Service and with data-protection laws such as the GDPR and CCPA. Comments contain personal data (usernames and what people wrote); only collect what you need, have a legitimate purpose, don't use it for spam or profiling individuals, and respect copyright when re-publishing transcripts. If in doubt, consult a lawyer. Read more in Apify's guide: Is web scraping legal?

Support

Found a bug or need a feature (reply threads, more fields, other sort orders)? Open an issue on the Actor's Issues tab — we usually reply within one business day.