Scrape YouTube channels, playlists, videos and search results: full metadata, top comments and timestamped transcripts in one run. Built for AI, RAG and LLM pipelines. No API key or login needed.
All notable changes to this Actor are documented here. Format: Keep a Changelog ; versions follow the Apify major.minor build versions.
[1.0] — 2026-09-27
Added
Scrape YouTube channels (Videos, Shorts and Live tabs), playlists, single videos and search (URLs or plain search terms).
Per video: full metadata, top-level comments (Top order, pinned/hearted flags, exact total), and transcripts with timestamped segments, fullText and wordCount.
uploadedAfter (absolute or relative dates), sortBy (newest / popular), transcript language preference list.
Run statistics (STATS key-value record, logged every 30 s) with categorized error counters and budgetReached.
Dataset views (Overview, Transcripts) and output schema.
Design decisions
HTTP only, no browser. All data comes from YouTube's InnerTube API (/browse, /search, /next, /player, /navigation/resolve_url) using the WEB client context and the API key / client version / visitor data read from the live page's ytcfg at start-up (with built-in fallbacks if that page is blocked).
Captions via the ANDROID/IOS InnerTube clients. Since 2025, caption URLs returned to the WEB client require a browser-generated proof-of-origin (PO) token and return empty bodies without one. The mobile clients' caption URLs do not, so they are used only to obtain caption tracks; all metadata still comes from the WEB client.
Machine-translation fallback. When none of the requested transcript languages exist, YouTube's translation (tlang) is attempted; because YouTube rate-limits translation aggressively, a failed translation falls back to the original-language track (isTranslated: false).
Search sort. YouTube removed the upload-date sort from search in 2025. sortBy: popular maps to "Prioritize: Popularity"; sortBy: newest uses relevance order for search (channels still sort newest-first). Documented in the input schema and README.
maxVideos instead of maxResults. The repo convention asks for maxResults (default 50); this Actor's specification defines maxVideos (default 20, per source) as the result limit, so maxResults is intentionally not added to avoid two overlapping limits.
Proxy default.proxyConfiguration defaults to { useApifyProxy: true } (datacenter) per repo convention; blocked/rate-limited requests rotate sessions, and the README tells users to switch to RESIDENTIAL if blocks persist.
Budget safety. Budget is reserved before each video's expensive work (comments, transcript) so concurrent handlers can never overspend; items are pushed first and charged immediately after, only for what was emitted. Free events (non-PPE runs) are never charged, so they can't be mistaken for an exhausted budget.
Quota refill. If a queued video turns out to be unavailable or older than uploadedAfter (by its exact publish date), its slot is refilled from the listing's backlog / next page, except for newest-first channel listings where everything after is older anyway.
Final-attempt degradation. On a request's last retry, a blocked comments/transcript sub-request is recorded as ERROR instead of discarding the whole video.
Comment dates are derived from YouTube's relative times and are approximate (publishedTimeText keeps the original).
Known limitations
Replies are not collected (only replyCount).
Age-restricted videos: metadata only (no transcript/comments).