YouTube Transcript Scraper
Pricing
from $2.40 / 1,000 videos
YouTube Transcript Scraper
Extract YouTube video transcripts, captions, and subtitles in JSON format with segments and timestamps. Supports multiple languages and auto-generated captions.
Pricing
from $2.40 / 1,000 videos
Rating
0.0
(0)
Developer
Murat Uzun
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
2 hours ago
Last modified
Categories
Share
What does YouTube Transcript Scraper do?
YouTube Transcript Scraper extracts video transcripts, captions, and subtitles from YouTube videos in JSON format with segment-level timestamps. It supports multiple languages, auto-generated captions, and exports data as structured JSON with the full transcript text and optional segments. No YouTube Data API key required.
Why use YouTube Transcript Scraper?
- Extract transcripts at scale — Scrape hundreds of videos' captions in a single run, perfect for content analysis, accessibility, and archival
- Structured transcript data — Get segment-level timestamps alongside plain text, enabling precise video-to-text references
- Multilingual — Prefer captions in any language; fall back to auto-generated captions when manual ones aren't available
- No API quota limits — Works directly with YouTube's innertube API, no YouTube Data API key needed
- Accessible output — Download transcripts as JSON, CSV, Excel, or HTML for immediate use in workflows and integrations
How to use YouTube Transcript Scraper
- Open the Actor — Navigate to the Actor page on Apify Store
- Paste video URLs — Enter one or more YouTube video URLs or IDs in the Video URLs field
- Full URL:
https://www.youtube.com/watch?v=dQw4w9WgXcQ - Shorts:
https://www.youtube.com/shorts/dQw4w9WgXcQ - youtu.be:
https://youtu.be/dQw4w9WgXcQ - Bare ID:
dQw4w9WgXcQ
- Full URL:
- Configure options — Choose language preference, auto-generated caption inclusion, and segment details
- Run the Actor — Click Start and wait for the results
- Download results — Export the transcript dataset as JSON, CSV, Excel, or HTML
Input
The Actor accepts the following input fields in the Input tab:
| Field | Type | Default | Description |
|---|---|---|---|
| videoUrls | string[] | ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"] | YouTube video URLs or IDs to extract transcripts from |
| language | string | "en" | Preferred two-letter language code (e.g., 'en', 'es', 'de') |
| includeAutoGenerated | boolean | true | Whether to include auto-generated captions when manual ones aren't available |
| includeSegments | boolean | true | Include segment data with timestamps; if false, only full transcript text is returned |
| proxyMode | string | "auto" | Proxy strategy: 'auto' (direct first, proxy on block), 'always' (always proxy), 'never' (direct only) |
| proxyConfiguration | object | {"useApifyProxy":true,"apifyProxyGroups":["RESIDENTIAL"]} | Apify Proxy settings for residential IP fallback |
Example input
{"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ","jNQXAC9IVRw"],"language": "en","includeAutoGenerated": true,"includeSegments": true,"proxyMode": "auto"}
Output
Each video produces one row in the dataset with transcript text, metadata, and optional segments. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.
Example output (single row)
{"videoId": "dQw4w9WgXcQ","url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up","channelName": "Rick Astley","channelId": "UCuAXFkgsw1L7xaCfnd5J1vQ","durationSeconds": 212,"language": "en","languageName": "English","isAutoGenerated": false,"availableLanguages": ["en", "es", "de"],"text": "Never gonna give you up, never gonna let you down, never gonna run around and desert you...","segments": [{ "start": 0.0, "dur": 2.5, "text": "Never gonna give you up" },{ "start": 2.5, "dur": 2.3, "text": "never gonna let you down" },{ "start": 4.8, "dur": 2.7, "text": "never gonna run around" }],"segmentCount": 87,"scrapedAt": "2024-09-28T12:30:45.123Z"}
Data fields
| Field | Type | Description |
|---|---|---|
videoId | string | 11-character YouTube video ID |
url | string | Full YouTube video URL |
title | string | Video title |
channelName | string | Name of the uploading channel |
channelId | string | YouTube channel ID (starts with 'UC') |
durationSeconds | integer | Video length in seconds |
language | string | ISO 639-1 language code of the extracted captions |
languageName | string | Human-readable language name (e.g., "English") |
isAutoGenerated | boolean | Whether captions were auto-generated (vs. manually created) |
availableLanguages | string[] | Array of all language codes available for this video |
text | string | Complete transcript as a single plain-text string |
segments | object[] | Array of transcript segments (only if includeSegments is true) with {start, dur, text} |
segmentCount | integer | Total number of transcript segments |
error | string | Error message if transcript extraction failed (e.g., "no captions") |
scrapedAt | string | ISO 8601 timestamp when the transcript was extracted |
Cost estimation
Pricing: $0.004 per video
The cost is determined by:
- Platform startup: ~$0.0001 per run
- Per-video player API call: ~$0.00015 (direct) or negligible (Apify datacenter)
- Per-video caption fetch: ~$0.00006 via residential proxy (direct fetch returns 429)
- Total per video: ~$0.004 (accounts for platform overhead and per-event billing)
Example costs
- 10 videos: ~$0.04
- 100 videos: ~$0.40
- 1,000 videos: ~$4.00
The Actor uses direct requests by default and only switches to residential proxy if YouTube blocks the datacenter IP, keeping costs minimal when possible.
Tips & advanced options
Language preferences
- The
languagefield specifies your preferred caption language (default:"en") - If that language isn't available, the Actor falls back to:
- First manual (non-auto-generated) caption track in any language
- First auto-generated track (if
includeAutoGeneratedis true) - Any available caption track
Segments vs. plain text
- With segments (default): Each transcript segment includes
starttime (in seconds),durduration, andtext— useful for video chapters, timed annotations, or player sync - Without segments (
includeSegments: false): Only the full concatenated transcript text, reducing output size
Proxy strategy
proxyMode: "auto"(default): Try direct; if YouTube blocks datacenter IPs (HTTP 429/403), automatically switch to residential proxy for all remaining videos in the run — minimizes costproxyMode: "always": Always use residential proxy from the start (costs ~$0.00006/video more)proxyMode: "never": Never use proxy; if direct requests get blocked, the video fails — useful for testing
FAQ
Videos with no captions
If a video has no captions (transcripts disabled), the Actor records an error row with "error": "no captions". These rows are not charged — only successful transcripts count toward billing.
Private or age-restricted videos
The Actor can only extract transcripts from public videos with captions. Age-restricted videos may fail if captions aren't publicly available. Unlisted videos work fine if the URL is known.
Auto-generated vs. manual captions
- Manual captions: Created by the channel owner or community; generally more accurate
- Auto-generated captions: Created by YouTube's speech-to-text engine; useful for videos without manual captions
Set includeAutoGenerated: false to only extract manually-created transcripts.
I got HTTP 429 "too many requests"
This happens when YouTube rate-limits datacenter IPs. The Actor automatically escalates to residential proxy in auto mode (default). If you're running hundreds of videos locally, use proxyMode: "always" to avoid repeated fallbacks.
Output columns show as "null"
The Actor gracefully handles missing data:
- Videos without the video title in captions show
title: null - Videos with no auto-generated captions show
isAutoGenerated: null
All fields except scrapedAt are nullable in the schema.
Large transcripts take a long time
Long videos (2+ hours) with dense captions can take 10-20 seconds per video due to segment parsing. This is normal. Consider batching into multiple runs if you have many long videos.
Limitations
- Captions only: The Actor only extracts captions/transcripts. It does not extract video metadata (views, likes, comments) — use YouTube Video Details Scraper for that
- Public videos only: Private videos return an error unless the URL is accessible to you
- No chat history: Does not extract YouTube Live chat messages, only video transcripts
- Language codes: Limited to YouTube's available caption languages; some videos may have captions in languages not listed in
availableLanguages
Related Actors
Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:
Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.
Website & domain intelligence
- Email Extractor — Website Contact & Social Finder — e-mails, phones and social profiles per domain
- Tech Stack Detector — Wappalyzer & BuiltWith Alternative — CMS, e-commerce, analytics, pixels and payments per domain
- Domain DNS & Email Security Checker — SPF, DKIM, DMARC, MX provider, registrar and domain age
- Domain Security Audit (TLS, HTTP headers, redirects, robots) — TLS expiry, security headers, redirect chain, robots and llms.txt
- Subdomain Finder (Certificate Transparency) — every subdomain seen in CT logs, with a live DNS check
- Bulk Core Web Vitals & PageSpeed Audit — Lighthouse scores, LCP, CLS, INP and top fixes per URL
- On-Page SEO Audit — title, meta, headings, links, images and schema issues per page
- Sitemap URL Extractor & Change Monitor — every sitemap URL, or new and removed pages between runs
- Wayback Machine Snapshot & Page Change Tracker — how a page changed over time, or every archived snapshot
- Bulk Domain WHOIS & RDAP Lookup — registrar, dates, status and nameservers per domain
- Web Scraper — CSS Selector & Data Extractor — pull any CSS selector off any page, one row per URL
- Website Screenshot Generator — full-page or viewport PNG/JPEG screenshots of any URL
Content for AI, LLMs and RAG
- AI Web Search & Read: Google results as clean Markdown — a query turned into clean Markdown from the top search results
- Website to Markdown — Content Crawler for LLM & RAG — any site as clean Markdown per page, no browser
- Article & News Extractor (clean text, author, date, markdown) — clean article text, author, date and Markdown per URL
- Structured Data & JSON-LD Extractor (Schema.org, Open Graph) — Schema.org and Open Graph data from any page
- Google News Scraper (RSS search by keyword, topic, site) — news results by keyword, topic or site
- Press Release Monitor: PR Newswire, BusinessWire, GlobeNewswire — PR Newswire, Business Wire and GlobeNewswire releases
Search, video and social
- YouTube Shorts Scraper — Shorts from channels, hashtags and searches with view counts
- Pinterest Pins Scraper — latest pins of public Pinterest profiles and boards
- Google Search Results Scraper — SERP API — organic SERP results per keyword and country
- YouTube Comments Scraper — Comments & Replies — comments and replies with likes, no API key
- YouTube Channel Latest Videos (RSS, no API key) — the latest 15 videos of any channel from RSS
- YouTube Channel Scraper (videos, shorts, live) — a channel's full video, shorts and stream list
- YouTube Search Results Scraper (videos, channels, no API key) — videos, channels and playlists per query
- YouTube Video Details Scraper (views, likes, description, tags) — views, likes, description, tags and chapters per video
- Apple Podcasts Lookup & Episodes Scraper — podcast metadata and episodes from iTunes and RSS
- Bluesky Post, Search & Profile Scraper — posts, profiles, followers and threads from the AT Protocol API
- Telegram Channel Posts Scraper — posts, views and media flags from any public channel
- Substack Publication & Posts Scraper — archive, authors and paywall status per publication
- Google Play Reviews Scraper — reviews, ratings, replies and app versions per app
- App Store Reviews Scraper — iOS reviews and ratings per app and country
- Google Trends Scraper — interest over time, by region, and related queries per keyword
- Google Ads Transparency Scraper — ads any advertiser runs on Google, with format and dates
- Keyword Suggestions Scraper (Google, YouTube, Amazon, Bing) — autocomplete keyword ideas from four search engines
- Bilibili Scraper (Videos, Search, Popular) — Chinese video platform: views, likes, coins, danmaku, uploader
- Mastodon Scraper (Hashtags, Accounts, Trending) — public fediverse posts by hashtag, account or trending
- Meetup Events Scraper (Search by Keyword & City) — upcoming events with RSVPs, fees, venues and groups
- Eventbrite Scraper (Events by Keyword & City) — events by keyword and city with venue, dates and organizer
Leads, jobs and company data
- Career Site Jobs API (Greenhouse, Lever, Ashby, Workday +1) — company domains in, their open jobs out, ATS detected automatically
- Workday Jobs Scraper — jobs with full descriptions from any Workday career site
- Google Maps Scraper — businesses with phone, website, address, rating and coordinates per search
- LinkedIn Jobs Scraper — job titles, companies, locations and full descriptions from LinkedIn job search
- Company 360: full company profile from a domain — one row per domain: contacts, tech, security, hiring and company facts
- Hiring Signals Scraper (Greenhouse, Lever, Ashby, Workable) — open jobs and hiring velocity from 10 public ATS boards
- Y Combinator Companies & Founders Scraper — YC startups by batch, industry and hiring status
- Wikidata Entity & Company Enrichment (facts, IDs, links) — HQ, founders, employees, revenue and social IDs per company
- Email Validator & Verifier — Bulk Email Check — syntax, MX, disposable, role and free-provider checks
- OpenStreetMap POI Extractor (Overpass API: shops, amenities) — shops and amenities by radius, bbox or area
- Stock, Crypto & FX Quotes — one row per symbol from Yahoo, Binance and ECB rates
- Remote Jobs Aggregator (RemoteOK, WWR, Hacker News) — one clean row per remote job, de-duplicated across feeds
- Greenhouse Jobs Scraper — jobs with descriptions from any Greenhouse job board
- Lever Jobs Scraper — jobs with descriptions from any Lever careers page
- Ashby Jobs Scraper — jobs, salaries and descriptions from any Ashby job board
- SmartRecruiters Jobs Scraper — jobs with descriptions from any SmartRecruiters company
- Seek Jobs Scraper (Australia & New Zealand) — Seek job ads with salary, work type and location
- Dice Jobs Scraper — US tech jobs from Dice with salary and remote flag
- AutoScout24 Scraper — European car listings with price, mileage and seller
- Rightmove Scraper — UK property for sale or rent with price and agent
- Wellfound Jobs Scraper (AngelList Startup Jobs) — startup jobs with salary and equity ranges, company size and stage
- Yandex Maps Scraper (Places, Ratings, Phones) — businesses in Russia, Türkiye and the CIS with phones, ratings, hours
- Craigslist Scraper (Listings, Prices, Locations) — listings in any area and category with price, date and coordinates
- JobStreet Scraper (Malaysia, Singapore, PH, ID + JobsDB) — JobStreet and JobsDB jobs in 6 Asian countries with parsed salaries
- InfoJobs Scraper (Spain Jobs, Salaries, Companies) — Spanish jobs with salary range, contract type and full description
- Redfin Scraper (Homes for Sale, Prices, Details) — US homes for sale or sold from any Redfin search, with price and details
- Kleinanzeigen Scraper (Ads, Prices, Locations) — German classifieds with price, VB flag, ZIP, city and seller type
Developer, app and research data
- npm, PyPI & Crates.io Package Health Checker — releases, downloads, deprecation and a health score
- GitHub Repository Health & Activity Report — stars, commits, contributors and risk flags per repo
- VS Code Marketplace Extension Scraper (installs, ratings) — installs, ratings and versions per extension
- Chrome Web Store Extension Scraper (installs, ratings) — users, rating, version and developer per extension
- Google Play Scraper — apps, ratings, installs, developer contact and reviews
- App Store (iOS) App Metadata, Ratings & Top Charts Lookup — ratings, price, version and charts per app
- CrossRef DOI & Citation Metadata Lookup — papers, authors, journals and citation counts
- FDA Recalls & Adverse Events Monitor (openFDA) — food, drug and device recalls from openFDA
- iCal / ICS Calendar Feed to Events Extractor — any public calendar feed as event rows
- Shopify Store Products Scraper — catalog, prices, variants and stock per store
- Hacker News Search & Front Page Scraper — stories, comments and points by query or front page
- GitHub Trending Repositories Scraper — trending repos and developers by language and period
- Stack Overflow & Stack Exchange Q&A Scraper — questions, answers and scores by query, tag or site
- Bulk Image Downloader — download image URLs to storage with size, dimensions and a ZIP
- Google Flights Scraper (Prices, Airlines, Stops) — flight prices, airlines, times, stops and CO2 by route and date
- Google Hotels Scraper (Prices, Ratings, Reviews) — hotel prices per night, stars, rating and reviews by city and dates
- AliExpress Scraper (Search Products & Prices) — AliExpress search results with USD price, discount and rank
- Lazada Scraper (Products, Prices, Sold, Ratings) — Lazada products in 6 countries with price, rating, units sold and seller
Support
Found a bug or want to request a feature? Visit the Issues tab or contact support through Apify. For custom solutions or bulk transcription needs, reach out to the team.