YouTube Transcript Scraper avatar

YouTube Transcript Scraper

Pricing

from $2.40 / 1,000 videos

Go to Apify Store
YouTube Transcript Scraper

YouTube Transcript Scraper

Extract YouTube video transcripts, captions, and subtitles in JSON format with segments and timestamps. Supports multiple languages and auto-generated captions.

Pricing

from $2.40 / 1,000 videos

Rating

0.0

(0)

Developer

Murat Uzun

Murat Uzun

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

2 hours ago

Last modified

Share

What does YouTube Transcript Scraper do?

YouTube Transcript Scraper extracts video transcripts, captions, and subtitles from YouTube videos in JSON format with segment-level timestamps. It supports multiple languages, auto-generated captions, and exports data as structured JSON with the full transcript text and optional segments. No YouTube Data API key required.

Why use YouTube Transcript Scraper?

  • Extract transcripts at scale — Scrape hundreds of videos' captions in a single run, perfect for content analysis, accessibility, and archival
  • Structured transcript data — Get segment-level timestamps alongside plain text, enabling precise video-to-text references
  • Multilingual — Prefer captions in any language; fall back to auto-generated captions when manual ones aren't available
  • No API quota limits — Works directly with YouTube's innertube API, no YouTube Data API key needed
  • Accessible output — Download transcripts as JSON, CSV, Excel, or HTML for immediate use in workflows and integrations

How to use YouTube Transcript Scraper

  1. Open the Actor — Navigate to the Actor page on Apify Store
  2. Paste video URLs — Enter one or more YouTube video URLs or IDs in the Video URLs field
    • Full URL: https://www.youtube.com/watch?v=dQw4w9WgXcQ
    • Shorts: https://www.youtube.com/shorts/dQw4w9WgXcQ
    • youtu.be: https://youtu.be/dQw4w9WgXcQ
    • Bare ID: dQw4w9WgXcQ
  3. Configure options — Choose language preference, auto-generated caption inclusion, and segment details
  4. Run the Actor — Click Start and wait for the results
  5. Download results — Export the transcript dataset as JSON, CSV, Excel, or HTML

Input

The Actor accepts the following input fields in the Input tab:

FieldTypeDefaultDescription
videoUrlsstring[]["https://www.youtube.com/watch?v=dQw4w9WgXcQ"]YouTube video URLs or IDs to extract transcripts from
languagestring"en"Preferred two-letter language code (e.g., 'en', 'es', 'de')
includeAutoGeneratedbooleantrueWhether to include auto-generated captions when manual ones aren't available
includeSegmentsbooleantrueInclude segment data with timestamps; if false, only full transcript text is returned
proxyModestring"auto"Proxy strategy: 'auto' (direct first, proxy on block), 'always' (always proxy), 'never' (direct only)
proxyConfigurationobject{"useApifyProxy":true,"apifyProxyGroups":["RESIDENTIAL"]}Apify Proxy settings for residential IP fallback

Example input

{
"videoUrls": [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"jNQXAC9IVRw"
],
"language": "en",
"includeAutoGenerated": true,
"includeSegments": true,
"proxyMode": "auto"
}

Output

Each video produces one row in the dataset with transcript text, metadata, and optional segments. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

Example output (single row)

{
"videoId": "dQw4w9WgXcQ",
"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up",
"channelName": "Rick Astley",
"channelId": "UCuAXFkgsw1L7xaCfnd5J1vQ",
"durationSeconds": 212,
"language": "en",
"languageName": "English",
"isAutoGenerated": false,
"availableLanguages": ["en", "es", "de"],
"text": "Never gonna give you up, never gonna let you down, never gonna run around and desert you...",
"segments": [
{ "start": 0.0, "dur": 2.5, "text": "Never gonna give you up" },
{ "start": 2.5, "dur": 2.3, "text": "never gonna let you down" },
{ "start": 4.8, "dur": 2.7, "text": "never gonna run around" }
],
"segmentCount": 87,
"scrapedAt": "2024-09-28T12:30:45.123Z"
}

Data fields

FieldTypeDescription
videoIdstring11-character YouTube video ID
urlstringFull YouTube video URL
titlestringVideo title
channelNamestringName of the uploading channel
channelIdstringYouTube channel ID (starts with 'UC')
durationSecondsintegerVideo length in seconds
languagestringISO 639-1 language code of the extracted captions
languageNamestringHuman-readable language name (e.g., "English")
isAutoGeneratedbooleanWhether captions were auto-generated (vs. manually created)
availableLanguagesstring[]Array of all language codes available for this video
textstringComplete transcript as a single plain-text string
segmentsobject[]Array of transcript segments (only if includeSegments is true) with {start, dur, text}
segmentCountintegerTotal number of transcript segments
errorstringError message if transcript extraction failed (e.g., "no captions")
scrapedAtstringISO 8601 timestamp when the transcript was extracted

Cost estimation

Pricing: $0.004 per video

The cost is determined by:

  • Platform startup: ~$0.0001 per run
  • Per-video player API call: ~$0.00015 (direct) or negligible (Apify datacenter)
  • Per-video caption fetch: ~$0.00006 via residential proxy (direct fetch returns 429)
  • Total per video: ~$0.004 (accounts for platform overhead and per-event billing)

Example costs

  • 10 videos: ~$0.04
  • 100 videos: ~$0.40
  • 1,000 videos: ~$4.00

The Actor uses direct requests by default and only switches to residential proxy if YouTube blocks the datacenter IP, keeping costs minimal when possible.

Tips & advanced options

Language preferences

  • The language field specifies your preferred caption language (default: "en")
  • If that language isn't available, the Actor falls back to:
    1. First manual (non-auto-generated) caption track in any language
    2. First auto-generated track (if includeAutoGenerated is true)
    3. Any available caption track

Segments vs. plain text

  • With segments (default): Each transcript segment includes start time (in seconds), dur duration, and text — useful for video chapters, timed annotations, or player sync
  • Without segments (includeSegments: false): Only the full concatenated transcript text, reducing output size

Proxy strategy

  • proxyMode: "auto" (default): Try direct; if YouTube blocks datacenter IPs (HTTP 429/403), automatically switch to residential proxy for all remaining videos in the run — minimizes cost
  • proxyMode: "always": Always use residential proxy from the start (costs ~$0.00006/video more)
  • proxyMode: "never": Never use proxy; if direct requests get blocked, the video fails — useful for testing

FAQ

Videos with no captions

If a video has no captions (transcripts disabled), the Actor records an error row with "error": "no captions". These rows are not charged — only successful transcripts count toward billing.

Private or age-restricted videos

The Actor can only extract transcripts from public videos with captions. Age-restricted videos may fail if captions aren't publicly available. Unlisted videos work fine if the URL is known.

Auto-generated vs. manual captions

  • Manual captions: Created by the channel owner or community; generally more accurate
  • Auto-generated captions: Created by YouTube's speech-to-text engine; useful for videos without manual captions

Set includeAutoGenerated: false to only extract manually-created transcripts.

I got HTTP 429 "too many requests"

This happens when YouTube rate-limits datacenter IPs. The Actor automatically escalates to residential proxy in auto mode (default). If you're running hundreds of videos locally, use proxyMode: "always" to avoid repeated fallbacks.

Output columns show as "null"

The Actor gracefully handles missing data:

  • Videos without the video title in captions show title: null
  • Videos with no auto-generated captions show isAutoGenerated: null

All fields except scrapedAt are nullable in the schema.

Large transcripts take a long time

Long videos (2+ hours) with dense captions can take 10-20 seconds per video due to segment parsing. This is normal. Consider batching into multiple runs if you have many long videos.

Limitations

  • Captions only: The Actor only extracts captions/transcripts. It does not extract video metadata (views, likes, comments) — use YouTube Video Details Scraper for that
  • Public videos only: Private videos return an error unless the URL is accessible to you
  • No chat history: Does not extract YouTube Live chat messages, only video transcripts
  • Language codes: Limited to YouTube's available caption languages; some videos may have captions in languages not listed in availableLanguages

Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:

Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.

Website & domain intelligence

Content for AI, LLMs and RAG

Search, video and social

Leads, jobs and company data

Developer, app and research data

Support

Found a bug or want to request a feature? Visit the Issues tab or contact support through Apify. For custom solutions or bulk transcription needs, reach out to the team.