YouTube Transcript & Video Analytics Scraper
Pricing
$4.00 / 1,000 video processeds
YouTube Transcript & Video Analytics Scraper
Pull the full transcript (captions) of any public YouTube video plus video and channel analytics - no login, no API key. Returns full-text transcript + timestamped segments, views, likes, channel and publish date. Pure HTTP, great for AI/content pipelines.
YouTube Transcript & Analytics Scraper
Pull the full transcript of any public YouTube video — plus rich video and channel analytics — with no login and no YouTube Data API key. Built for AI/content pipelines: feed real spoken content straight into summarizers, RAG stores, translation or repurposing workflows, alongside honest view/like/upload metadata.
What it does
- Transcript (the premium feature) — reads the video's actual caption track (manual or
auto-generated), in your preferred language when available, and returns both a clean
full-text
transcriptand timestampedtranscriptSegments({start, dur, text}). Two independent retrieval paths are tried automatically for every video: (1) YouTube's direct caption (timedtext) API in three response formats, and (2) as a fallback, the same internalget_transcriptcall YouTube's own website makes to render the transcript panel. - Video metadata — title, channel, view count, like count, upload/publish date, category, duration, description, keywords, thumbnails, family-safe flag, available countries, and the list of all caption languages the video actually has.
- Channel scraping — give it a channel URL or
@handleand it lists the channel's recent videos (title, view count text, published-time text, duration, thumbnail) and — if you want — fetches the full metadata + transcript for each of those videos too.
Input
{"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"channelUrls": ["https://www.youtube.com/@MrBeast"],"languages": ["en"],"maxVideos": 30,"includeTranscript": true,"proxyConfiguration": { "useApifyProxy": true },"timeoutSecs": 30}
videoUrls— watch URLs,youtu.belinks,/shorts/links, or bare 11-char video IDs.channelUrls— channel URLs (/@handle,/channel/UC...,/c/Name,/user/Name) or bare@handle.- At least one of
videoUrls/channelUrlsis required. languages— ordered language-code preference for the transcript (e.g.["en","pt"]). A real (human-made) caption track in one of these languages is preferred; then an auto-generated one in these languages; and if neither exists the actor honestly falls back to whichever caption track the video actually has, reporting exactly which one viatranscriptLanguage/transcriptIsAutoGenerated— it never invents transcript text.maxVideos— how many recent videos to pull per channel (first page of the Videos tab).includeTranscript— setfalseto skip transcripts and only get metadata/listings (faster, cheaper on proxy bandwidth).proxyConfiguration— Apify Proxy; datacenter proxy works fine for YouTube, no need for residential.
Output
One dataset item per video, plus one summary item per channel.
Video item (from videoUrls, or expanded from a channel when includeTranscript=true) —
metadata fields are always real when the video loads; the transcript* fields reflect
whichever of the two retrieval paths above actually succeeded for that video at request
time (see "Known limitation" below):
{"type": "video","videoId": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","author": "Rick Astley","viewCount": 1820040849,"likeCount": null,"publishDate": "2009-10-24T23:57:33-07:00","category": "Music","keywords": ["rick astley", "..."],"channelSubscriberCountText": "4.54M subscribers","availableCaptionLanguages": [{ "languageCode": "en", "name": "English", "isAutoGenerated": false },{ "languageCode": "en", "name": "English (auto-generated)", "isAutoGenerated": true },{ "languageCode": "pt-BR", "name": "Portuguese (Brazil)", "isAutoGenerated": false }],"has_transcript": true,"transcript": "We're no strangers to love. You know the rules and so do I...","transcriptSegments": [{ "start": 18.32, "dur": 3.02, "text": "We're no strangers to love" }],"transcriptLanguage": "en","transcriptIsAutoGenerated": false,"transcriptSource": "timedtext","transcriptError": null}
When neither retrieval path succeeds for a given video (see limitation below), you instead honestly get
"has_transcript": false, "transcript": null, "transcriptSource": null, "transcriptError": "innertube: HTTP 400 fetching ..."Channel item:
{"type": "channel","channelUrl": "https://www.youtube.com/@MrBeast","title": "MrBeast","subscriberCountText": "410M subscribers","videoCountText": "700 videos"}
If a video has no captions at all (or YouTube would not serve them to this request),
has_transcript is false, transcript is null, and transcriptError explains exactly
why — the actor never fabricates transcript text. availableCaptionLanguages is always
returned honestly from the video's own player data, independent of whether the transcript
body itself could be fetched, so you can always see which languages exist even on a run
where the text couldn't be pulled.
If a video/channel fails to load (blocked, region-locked, removed, not found), you get an
honest { input, error } item instead of partial/guessed data.
Known limitation (read before relying on 100% transcript coverage)
YouTube has been progressively tightening anti-bot checks on caption-serving endpoints.
In testing, both retrieval paths above can currently return an empty/400 FAILED_PRECONDITION
response for some videos even though the video genuinely has captions (confirmed via
availableCaptionLanguages) — this reproduces identically with or without a proxy, so it is
YouTube-side request validation, not a proxy quality issue, and not specific to this actor's
code path. When this happens the item is still fully honest: has_transcript: false and a
transcriptError describing which mechanism failed and why, never invented text. Metadata,
view/like counts, dates, and channel listings are unaffected and reliable regardless.
If YouTube's enforcement eases (it fluctuates) or for videos it doesn't trigger on, the
actor returns the real transcript automatically — no input change needed.
Notes
- Source: YouTube's own public, server-rendered pages (
ytInitialPlayerResponse/ytInitialData) and YouTube's own caption/transcript endpoints — no private API, no login. - Every HTTP request uses a fresh proxy session and retries (up to 4 attempts) if YouTube responds with a block (403/429 or an "unusual traffic" page).
- Channel scraping reads only the first page of the Videos tab (no pagination token
yet) — the returned
channelItem.notesfield says so honestly when the channel has more videos than were returned. Both of YouTube's current channel-grid JSON layouts (classicvideoRendererand the newerlockupViewModel) are supported. likeCountis best-effort (parsed from the page's own accessibility label) and isnullwhen YouTube doesn't expose it in a reliably parseable way — never guessed.
⭐ Enjoying this Actor?
A quick rating/review helps others find it. Want multi-page channel pagination or playlist support added? Open a ticket on the Issues tab.