YouTube Transcript Scraper — Channel & Video Subtitles
Pricing
$3.00 / 1,000 transcript scrapeds
YouTube Transcript Scraper — Channel & Video Subtitles
Scrape full transcripts, timestamped segments, and subtitles from YouTube channels, playlists, and single videos. Pure HTTP Innertube engine, zero headless browser, ultra-fast, multi-language & translation support.
Pricing
$3.00 / 1,000 transcript scrapeds
Rating
0.0
(0)
Developer
Kashif Ali
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
16 hours ago
Last modified
Categories
Share
🎬 YouTube Transcript Scraper
Extract full video transcripts, timestamped segments, and SRT/VTT subtitles from any YouTube channel, playlist, or individual video.
Built with a pure HTTP Innertube RPC engine — zero heavy browser overhead, runs on 256 MB RAM, and extracts at lightning speed.
🚀 Why Use This YouTube Transcript Scraper?
Existing transcript scrapers on the market frequently fail:
- ❌ Incomplete Channel Scrapes: Most scrapers trust the channel's curated Videos tab and silently miss uploads a channel has hidden from it — plus Shorts and past streams.
- ❌ Fragile to Blocks: One IP block or a burst of 404s and the run dies mid-channel.
🌟 What Makes This Scraper Elite:
- ✅ Truly Complete Channel Coverage: Discovery merges four sources — curated Videos tab, Shorts tab, Streams tab, and the channel's full uploads playlist via the Innertube browse API — then dedupes. If a channel has it, you get it.
- ✅ Target Type Control: Paste links and let auto-detect classify them, or explicitly declare "this is a channel / playlist / single video" — mislabeled inputs are corrected or clearly rejected with a reason, never silently mis-scraped.
- ✅ Complete Playlist Extraction: All videos in any public playlist, resolved via the Innertube browse API first (immune to page soft-blocks) with an HTML fallback, preserving playlist order.
- ✅ Single Video Support: Instant transcript extraction for any video URL (
/watch,youtu.be,/shorts). - ✅ Pure HTTP Architecture: Powered by lightweight RPC calls (256 MB RAM, < 1.5s per video) — 93% cheaper in compute costs than browser-based scrapers.
- ✅ Multi-Language & Auto-Translation: Prioritize manual captions, fall back to auto-generated speech-to-text (ASR), and translate transcripts to any target language.
- ✅ LLM & RAG-Ready Outputs: Emits clean prose text (no mid-sentence line breaks), timestamped segments (
start,duration,text), and standard.srt/.vttsubtitle files. - ✅ Consolidated Channel Archive: Option to generate a single Markdown knowledge base file saved directly to your Key-Value store.
📥 Input Options
| Parameter | Type | Default | Description |
|---|---|---|---|
startUrls | Array | (Required) | Paste one or more YouTube links — channels, playlists, and single videos can be mixed freely. Accepted: @handle, /channel/UC..., /c/Name, /user/Name, /playlist?list=..., /watch?v=..., youtu.be/..., /shorts/..., /live/..., /embed/..., bare 11-char IDs. |
targetType | Select | auto | "Auto-detect from URL" (recommended) or force channel / playlist / video. Forcing helps when a watch?v=…&list=… link should expand as a playlist instead of scraping one video. |
maxVideosPerUrl | Integer | 0 | Cap per link; 0 scrapes every video. Channel discovery merges the curated Videos tab, Shorts, streams, and the full uploads playlist — so you always get the complete set, even when a channel hides older videos from its tab (e.g. Starter Story's tab shows 186 of 524 uploads). |
startDate | String | (Optional) | For channel/playlist URLs, only discover videos published on or after this date (inclusive, YYYY-MM-DD). |
endDate | String | (Optional) | For channel/playlist URLs, only discover videos published on or before this date (inclusive, YYYY-MM-DD). |
fetchVideoMetadata | Boolean | true | Fetch per-video metadata (description, views, publish date, thumbnail, keywords, category, available caption languages). Disable for maximum speed on very large runs. |
includeShorts | Boolean | true | When scraping a channel, also scrape transcripts for YouTube Shorts. |
includeLiveStreams | Boolean | true | When scraping a channel, also scrape transcripts for past live streams. |
preferredLanguages | Array | ["en"] | Priority list of ISO language codes (e.g. ["en", "es", "de"]). Falls back to available tracks if not found. |
preferAutoGenerated | Boolean | true | Automatically fall back to YouTube's auto-generated speech recognition when manual subtitles are absent. |
translateTo | String | "" | Optional ISO language code to translate the transcript into (e.g. "es" for Spanish). |
outputFormats | Select | "all" | Choose "all" (text + segments + SRT + VTT), "plainTextOnly", "segmentsOnly", or "srtOnly". |
generateChannelSummaryFile | Boolean | false | Consolidates all transcripts into a single TRANSCRIPTS_<channel>.md file in the Key-Value store. |
proxyConfiguration | Object | { "useApifyProxy": true } | Apify proxy configuration to bypass rate limits on massive batch runs. Keep enabled — without a proxy, YouTube frequently blocks transcript fetching (IP_BLOCKED). |
maxConcurrency | Integer | 10 | Number of concurrent video transcripts to fetch simultaneously (1 to 50). |
Resilience: Videos that fail due to blocks or transient errors are automatically re-attempted in retry rounds with fresh proxy sessions and growing backoff before being reported as
IP_BLOCKED/ERROR, so blocked videos are recovered whenever possible.
📤 Output Format
Each scraped video is saved as a structured record in the default Apify Dataset:
{"videoId": "dQw4w9WgXcQ","videoUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Music Video)","channelName": "Rick Astley","channelId": "UCuAXFkgsw1L7xaCfnd5JJOw","channelUrl": "https://www.youtube.com/channel/UCuAXFkgsw1L7xaCfnd5JJOw","videoType": "VIDEO","hasTranscript": true,"language": "English","languageCode": "en","isAutoGenerated": false,"isTranslation": false,"text": "We're no strangers to love. You know the rules and so do I. A full commitment's what I'm thinking of...","textWordCount": 384,"textCharacterCount": 2048,"durationSeconds": 213,"description": "The official video...","viewCount": 1818112471,"publishDate": "2009-10-24T23:57:33-07:00","thumbnail": "https://i.ytimg.com/vi/dQw4w9WgXcQ/maxresdefault.jpg","keywords": ["rick astley", "never gonna give you up"],"category": "Music","availableLanguages": ["en", "de-DE", "ja", "pt-BR", "es-419"],"segmentCount": 61,"segments": [{"start": 18.64,"duration": 3.24,"text": "We're no strangers to love"},{"start": 22.64,"duration": 4.32,"text": "You know the rules and so do I"}],"srt": "1\n00:00:18,640 --> 00:00:21,880\nWe're no strangers to love\n\n2\n00:00:22,640 --> 00:00:26,960\nYou know the rules and so do I","vtt": "WEBVTT\n\n00:00:18.640 --> 00:00:21.880\nWe're no strangers to love\n\n00:00:22.640 --> 00:00:26.960\nYou know the rules and so do I","status": "SUCCESS","errorMessage": null,"scrapedAt": "2026-09-20T16:30:00.000Z"}
If a video does not have subtitles (e.g., ambient music or captions disabled by the creator), the actor records hasTranscript: false with the matching status without failing the run. Only when none of the provided URLs resolve to at least one video does the run fail explicitly.
Status Reference
Every dataset record carries exactly one status:
| Status | Meaning | Fails the run? |
|---|---|---|
SUCCESS | Transcript extracted successfully. | No |
NO_TRANSCRIPT_AVAILABLE | No caption tracks exist, or none match preferredLanguages while preferAutoGenerated: false. | No |
TRANSCRIPTS_DISABLED | The creator disabled captions for this video. | No |
VIDEO_UNAVAILABLE | Video is private, deleted, or unavailable. | No |
AGE_RESTRICTED | Video is age-restricted; transcripts require authentication. | No |
IP_BLOCKED | YouTube blocked the request (after one automatic retry). Enable Apify RESIDENTIAL proxies for large runs. | No |
ERROR | Unexpected retrieval failure — see errorMessage. | No |
Notes on Output
channelUrl/channelIdare populated from channel metadata whenever the video is discovered via a channel or playlist; for single-video targets the actor enriches them via oEmbed when available.durationSecondscomes from video metadata when known, otherwise it is approximated from the last transcript timestamp.segments,srt, andvttare included according to theoutputFormatssetting (allemits everything;plainTextOnlyomits all three).- Cue tags such as
[Music],[Applause], and[Laughter]are stripped from text, segments, SRT, and VTT. availableLanguageslists every caption language YouTube offers for the video — re-run with one of them viapreferredLanguagesif the default pick was wrong.- Set
fetchVideoMetadata: falseto skip the metadata round-trip per video (text/segments/SRT/VTT are unaffected; metadata fields becomenull).
💻 Integrations & API Usage
Python (Apify Client)
from apify_client import ApifyClientclient = ApifyClient("<YOUR_API_TOKEN>")# Scrape an entire YouTube channelrun_input = {"startUrls": ["https://www.youtube.com/@veritasium"],"maxVideosPerUrl": 50,"includeShorts": True,"outputFormats": "all"}run = client.actor("DataSiphon/youtube-transcript-scraper").call(run_input=run_input)# Fetch dataset itemsfor item in client.dataset(run["defaultDatasetId"]).iterate_items():print(f"[{item['title']}] ({item['language']}): {item['textWordCount']} words")
Node.js (Apify Client)
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });const run = await client.actor('DataSiphon/youtube-transcript-scraper').call({// Complete uploads playlist of the "Google for Developers" channel.// Tip: replace the "UC" prefix of any channel ID with "UU" to get its uploads playlist.startUrls: ['https://www.youtube.com/playlist?list=UU_x5XG1OV2P6uZZ5FSM9Ttw'],preferredLanguages: ['en'],outputFormats: 'all',});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(`Scraped ${items.length} video transcripts.`);
cURL
curl -X POST "https://api.apify.com/v2/acts/DataSiphon~youtube-transcript-scraper/runs?token=<YOUR_API_TOKEN>" \-H "Content-Type: application/json" \-d '{"startUrls": ["https://www.youtube.com/@hubermanlab"],"maxVideosPerUrl": 10}'
❓ Frequently Asked Questions
Can it scrape private or unlisted videos?
It can scrape unlisted videos if you provide the direct URL. Private videos or videos behind a member-only paywall cannot be scraped without authentication cookies.
What if a channel has thousands of videos?
Set maxVideosPerUrl: 0 to scrape without limits. The Innertube continuation loop will paginate through thousands of uploads seamlessly while consuming under 256 MB of RAM.
Are auto-generated subtitles supported?
Yes. If manual human subtitles were not uploaded by the creator, the scraper automatically grabs YouTube's speech-to-text transcript (preferAutoGenerated: true).
How does translation work?
Set translateTo: "es" (or any valid language code). The scraper uses YouTube's server-side translation track to return the transcript in your target language.
📄 License & Terms
This actor is intended for personal research, accessibility, and AI training applications. Please respect YouTube's Terms of Service and copyright guidelines for content reuse.
Pricing & cost estimator
Pay-per-event: $3.00 per 1,000 transcripts (videos without a transcript or blocked videos are not billed). You are charged only for rows that contain data; empty or failed targets are free. A 100-item run typically costs well under $0.50 in total. Use maxItems / the per-link cap to control spend.
Example output (real run)
{"videoId": "P14HA83uNJE","videoUrl": "https://www.youtube.com/watch?v=P14HA83uNJE","title": "a video to watch if you're ambitious and in your 20s or 30s","channelName": "Alex Hormozi","channelId": "UCUyDOdBWhC1MCxEjC46d-zw","channelUrl": "https://www.youtube.com/channel/UCUyDOdBWhC1MCxEjC46d-zw","description": "Download your free scaling roadmap here: https://www.acquisition.com/roadmap?el=yt-alex-ma…","viewCount": 471126,"publishDate": "2026-09-17T06:00:39-07:00","thumbnail": "https://i.ytimg.com/vi_webp/P14HA83uNJE/maxresdefault.webp","keywords": ["Alex Hormozi","Alex Hormozi Business Tips"],"category": "Entertainment","availableLanguages": ["en"],"videoType": "VIDEO","hasTranscript": true,"language": "English (auto-generated)","languageCode": "en","isAutoGenerated": true,"isTranslation": false,"text": "Just got off the phone with a friend of mine who's competing in a fitness thing and she di…","textWordCount": 1796,"textCharacterCount": 9412,"segmentCount": 282,"durationSeconds": 531,"status": "SUCCESS","scrapedAt": "2026-10-01T06:51:40.175354+00:00","segments": [{"start": 0,"duration": 3.24,"text": "Just got off the phone with a friend of"},{"start": 1,"duration": 5.48,"text": "mine who's competing in a fitness thing"}],"srt": "1\n00:00:00,000 --> 00:00:03,240\nJust got off the phone with a friend of\n\n2\n00:00:01,000 --…","vtt": "WEBVTT\n\n\n00:00:00.000 --> 00:00:03.240\nJust got off the phone with a friend of\n\n00:00:01.0…"}
Limitations
Videos with captions disabled return a free row with hasTranscript: false. YouTube sometimes blocks cheap IPs; blocked videos are retried through a residential proxy automatically (the only extra cost).
No login, no cookies
This actor only reads public YouTube watch pages and caption tracks. It never asks for a cookie, password or account and does not access private data. Make sure your use of the scraped data complies with the source site's terms and applicable law (GDPR/CCPA for personal data).
YouTube transcript scraper — FAQ & use cases
How do I get a YouTube transcript with an API? Run this Actor via the Apify API with video, playlist or channel URLs and fetch the dataset as JSON, CSV or Excel.
Can it scrape subtitles and captions in bulk? Yes — whole channels, playlists and Shorts, with maxVideosPerUrl to control cost.
Which formats? Clean plain text, timestamped segments (start, duration, text), SRT and VTT; translate to any language with translateTo.
Does it cache transcripts? Yes — repeat requests are served from a cache (default maxCacheAgeDays: 90; set 0 for always-fresh). Cached rows carry cached, fetched_at and cacheAgeDays; a 30-video re-run finishes in seconds instead of a minute.
Is it good for LLM / RAG? Yes — plain-text output and a Markdown channel knowledge base file.
Pricing: pay per event, $3 per 1,000 transcripts; videos without captions or blocked videos are free. Blocked videos are retried through a residential proxy automatically.
Use cases: AI and RAG datasets · content repurposing · SEO and keyword research from video text · accessibility and subtitles · research and archiving.
YouTube transcript API, subtitle downloader & SRT/VTT
A reliable YouTube transcript API and subtitle downloader: retrieve transcripts and subtitles for YouTube videos, with timed transcript segments in JSON, plus SRT and VTT captions in bulk — extract transcripts from videos, channels and playlists, then export.
Related scrapers by the same author
YouTube Email Scraper · Reddit Scraper · LinkedIn Jobs Scraper · Airbnb Scraper · Amazon Product Scraper