YouTube Transcript & Video Analytics Scraper avatar

YouTube Transcript & Video Analytics Scraper

Pricing

$4.00 / 1,000 video processeds

Go to Apify Store
YouTube Transcript & Video Analytics Scraper

YouTube Transcript & Video Analytics Scraper

Pull the full transcript (captions) of any public YouTube video plus video and channel analytics - no login, no API key. Returns full-text transcript + timestamped segments, views, likes, channel and publish date. Pure HTTP, great for AI/content pipelines.

Pricing

$4.00 / 1,000 video processeds

Rating

0.0

(0)

Developer

Bruno

Bruno

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

YouTube Transcript & Analytics Scraper

Pull the full transcript of any public YouTube video — plus rich video and channel analytics — with no login and no YouTube Data API key. Built for AI/content pipelines: feed real spoken content straight into summarizers, RAG stores, translation or repurposing workflows, alongside honest view/like/upload metadata.

What it does

  • Transcript (the premium feature) — reads the video's actual caption track (manual or auto-generated), in your preferred language when available, and returns both a clean full-text transcript and timestamped transcriptSegments ({start, dur, text}). Two independent retrieval paths are tried automatically for every video: (1) YouTube's direct caption (timedtext) API in three response formats, and (2) as a fallback, the same internal get_transcript call YouTube's own website makes to render the transcript panel.
  • Video metadata — title, channel, view count, like count, upload/publish date, category, duration, description, keywords, thumbnails, family-safe flag, available countries, and the list of all caption languages the video actually has.
  • Channel scraping — give it a channel URL or @handle and it lists the channel's recent videos (title, view count text, published-time text, duration, thumbnail) and — if you want — fetches the full metadata + transcript for each of those videos too.

Input

{
"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"channelUrls": ["https://www.youtube.com/@MrBeast"],
"languages": ["en"],
"maxVideos": 30,
"includeTranscript": true,
"proxyConfiguration": { "useApifyProxy": true },
"timeoutSecs": 30
}
  • videoUrls — watch URLs, youtu.be links, /shorts/ links, or bare 11-char video IDs.
  • channelUrls — channel URLs (/@handle, /channel/UC..., /c/Name, /user/Name) or bare @handle.
  • At least one of videoUrls / channelUrls is required.
  • languages — ordered language-code preference for the transcript (e.g. ["en","pt"]). A real (human-made) caption track in one of these languages is preferred; then an auto-generated one in these languages; and if neither exists the actor honestly falls back to whichever caption track the video actually has, reporting exactly which one via transcriptLanguage / transcriptIsAutoGenerated — it never invents transcript text.
  • maxVideos — how many recent videos to pull per channel (first page of the Videos tab).
  • includeTranscript — set false to skip transcripts and only get metadata/listings (faster, cheaper on proxy bandwidth).
  • proxyConfiguration — Apify Proxy; datacenter proxy works fine for YouTube, no need for residential.

Output

One dataset item per video, plus one summary item per channel.

Video item (from videoUrls, or expanded from a channel when includeTranscript=true) — metadata fields are always real when the video loads; the transcript* fields reflect whichever of the two retrieval paths above actually succeeded for that video at request time (see "Known limitation" below):

{
"type": "video",
"videoId": "dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"author": "Rick Astley",
"viewCount": 1820040849,
"likeCount": null,
"publishDate": "2009-10-24T23:57:33-07:00",
"category": "Music",
"keywords": ["rick astley", "..."],
"channelSubscriberCountText": "4.54M subscribers",
"availableCaptionLanguages": [
{ "languageCode": "en", "name": "English", "isAutoGenerated": false },
{ "languageCode": "en", "name": "English (auto-generated)", "isAutoGenerated": true },
{ "languageCode": "pt-BR", "name": "Portuguese (Brazil)", "isAutoGenerated": false }
],
"has_transcript": true,
"transcript": "We're no strangers to love. You know the rules and so do I...",
"transcriptSegments": [{ "start": 18.32, "dur": 3.02, "text": "We're no strangers to love" }],
"transcriptLanguage": "en",
"transcriptIsAutoGenerated": false,
"transcriptSource": "timedtext",
"transcriptError": null
}

When neither retrieval path succeeds for a given video (see limitation below), you instead honestly get

"has_transcript": false, "transcript": null, "transcriptSource": null, "transcriptError": "innertube: HTTP 400 fetching ..."
— with every metadata field above still fully populated and real.

Channel item:

{
"type": "channel",
"channelUrl": "https://www.youtube.com/@MrBeast",
"title": "MrBeast",
"subscriberCountText": "410M subscribers",
"videoCountText": "700 videos"
}

If a video has no captions at all (or YouTube would not serve them to this request), has_transcript is false, transcript is null, and transcriptError explains exactly why — the actor never fabricates transcript text. availableCaptionLanguages is always returned honestly from the video's own player data, independent of whether the transcript body itself could be fetched, so you can always see which languages exist even on a run where the text couldn't be pulled. If a video/channel fails to load (blocked, region-locked, removed, not found), you get an honest { input, error } item instead of partial/guessed data.

Known limitation (read before relying on 100% transcript coverage)

YouTube has been progressively tightening anti-bot checks on caption-serving endpoints. In testing, both retrieval paths above can currently return an empty/400 FAILED_PRECONDITION response for some videos even though the video genuinely has captions (confirmed via availableCaptionLanguages) — this reproduces identically with or without a proxy, so it is YouTube-side request validation, not a proxy quality issue, and not specific to this actor's code path. When this happens the item is still fully honest: has_transcript: false and a transcriptError describing which mechanism failed and why, never invented text. Metadata, view/like counts, dates, and channel listings are unaffected and reliable regardless. If YouTube's enforcement eases (it fluctuates) or for videos it doesn't trigger on, the actor returns the real transcript automatically — no input change needed.

Notes

  • Source: YouTube's own public, server-rendered pages (ytInitialPlayerResponse / ytInitialData) and YouTube's own caption/transcript endpoints — no private API, no login.
  • Every HTTP request uses a fresh proxy session and retries (up to 4 attempts) if YouTube responds with a block (403/429 or an "unusual traffic" page).
  • Channel scraping reads only the first page of the Videos tab (no pagination token yet) — the returned channelItem.notes field says so honestly when the channel has more videos than were returned. Both of YouTube's current channel-grid JSON layouts (classic videoRenderer and the newer lockupViewModel) are supported.
  • likeCount is best-effort (parsed from the page's own accessibility label) and is null when YouTube doesn't expose it in a reliably parseable way — never guessed.

⭐ Enjoying this Actor?

A quick rating/review helps others find it. Want multi-page channel pagination or playlist support added? Open a ticket on the Issues tab.