YouTube Transcript Scraper avatar

YouTube Transcript Scraper

Pricing

Pay per event

Go to Apify Store
YouTube Transcript Scraper

YouTube Transcript Scraper

Get the transcript (captions) of any public YouTube video, Short or live replay: timed segments, plain text, SRT or VTT, in the language you choose (manual or auto-generated, optional YouTube translation), plus video metadata. LLM-ready output, pay only per transcript returned.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Atalaia

Atalaia

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

What does YouTube Transcript Scraper do?

YouTube Transcript Scraper gets the transcript (captions / subtitles) of public YouTube videos, Shorts and live-stream replays, and returns it as clean, LLM-ready data: timed segments, plain text, SRT or WebVTT. It works as a simple YouTube transcript API: send a list of video URLs or IDs, get one dataset item per video.

  • ✅ Manual (human) captions and auto-generated captions, in the language you choose, with a clear fallback when that language doesn't exist
  • ✅ Optional translation with YouTube's own machine translation (translateTo)
  • ✅ Video metadata in the same item: title, channel, duration, publish date, view count
  • ✅ You only pay for transcripts returned (plus a small start fee per run). Videos without captions, private or removed videos are reported as errors and never charged
  • ❌ Does not download video/audio, does not transcribe audio itself, does not scrape comments

Why use YouTube Transcript Scraper?

  • RAG and AI pipelines: feed talks, podcasts, lectures and tutorials into a vector database or an LLM for summaries, Q&A and search.
  • Content repurposing: turn videos into blog posts, newsletters, threads and show notes.
  • Research and monitoring: analyze what channels, brands or competitors say in their videos.
  • Subtitles: get SRT or VTT files ready for an editor or a player.

It runs on the Apify platform, so you get an API, scheduling, webhooks, integrations (Make, Zapier, n8n, LangChain, LlamaIndex...), and proxy rotation built in. Each video is retried with fresh sessions and, when needed, residential IPs, so runs don't fail silently.

What data can YouTube Transcript Scraper extract?

FieldTypeDescription
videoId, urlstringVideo ID and canonical watch URL
title, channelName, channelIdstringPublic video and channel metadata
durationSeconds, viewCountintegerLength and views at scrape time
publishDatestringISO 8601 publish date
language, languageNamestringCaption track that was used
isAutoGeneratedbooleantrue for YouTube's automatic speech recognition
translatedTostring / nullTarget language of YouTube's translation
availableLanguagesarrayAll caption tracks on the video: {code, name, isAutoGenerated}
segments / text / srt / vttarray / stringThe transcript, in the format you chose
wordCountintegerWords in the transcript (CJK characters count individually)
error, errorMessagestringOnly on videos without a transcript (not charged)

HTML entities are decoded (' → ') and sound tags such as [Music] or [Applause] are kept as they appear on YouTube.

How to scrape YouTube transcripts

  1. Click Try for free and open the Input tab.
  2. Paste video URLs or IDs into Videos (watch?v=, youtu.be/, shorts/, live/ and embed/ links all work).
  3. Set Preferred language (for example en, pt-BR, es, ja), and optionally Translate to.
  4. Pick an Output format: text is best for LLMs, segments keeps timestamps, srt/vtt are subtitle files.
  5. Click Start and download the results as JSON, CSV, Excel or HTML, or read them through the API.

How much does it cost to scrape YouTube transcripts?

This Actor uses pay-per-event pricing:

EventPriceWhen
apify-actor-startUS$ 0.005Once per run
transcriptUS$ 0.005Per video that returned a transcript
translationUS$ 0.05Extra, per video whose transcript was translated with translateTo

Examples: 1,000 videos in one run cost US$ 5.005. One video in one run costs US$ 0.01. One translated video costs US$ 0.06 (start + transcript + translation).

⚠️ Translation is slower and more expensive. YouTube throttles translated captions heavily, so each translated video takes about 1 to 3 minutes instead of a few seconds, and costs US$ 0.05 on top of the transcript. Leave translateTo empty unless you need it. If you only want a transcript in a language that already exists on the video, use language instead: it's free of the translation fee.

You are not charged for videos that have no captions, are private, removed, age-restricted, or fail for any other reason. A failed translation (translation_failed) is not charged either. Platform usage (compute, proxy) is included. You can cap the spend of a run with the Maximum cost per run setting; the Actor stops cleanly when it is reached.

Input

See the Input tab for every option. Example:

{
"videos": [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"https://youtu.be/NNnIGh9g6fA",
"https://www.youtube.com/shorts/9dMFxnpEkKc"
],
"language": "en",
"includeAutoGenerated": true,
"outputFormat": "text",
"includeVideoMetadata": true
}

Language fallback. For each video the Actor picks, in this order: a manual track in language → an auto-generated track in language → any manual track → any auto-generated track. pt also matches pt-BR, en matches en-US, and so on. The item always says which track was used (language, isAutoGenerated) and lists every track in availableLanguages. Turn Include auto-generated captions off to get only human-made captions.

Translation. translateTo asks YouTube for its machine translation of the chosen track (for example "translateTo": "pt"). If YouTube doesn't offer that language for the video, the item is an error translation_unavailable (not charged). If the track is already in that language, no translation is done.

Output

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. Example item with outputFormat: "segments" (segments shortened):

{
"videoId": "dQw4w9WgXcQ",
"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"channelName": "Rick Astley",
"channelId": "UCuAXFkgsw1L7xaCfnd5JJOw",
"durationSeconds": 213,
"publishDate": "2009-10-24T23:57:33-07:00",
"viewCount": 1821342774,
"isLiveContent": false,
"language": "en",
"languageName": "English",
"isAutoGenerated": false,
"translatedTo": null,
"availableLanguages": [
{ "code": "en", "name": "English", "isAutoGenerated": false },
{ "code": "en", "name": "English (auto-generated)", "isAutoGenerated": true },
{ "code": "de-DE", "name": "German (Germany)", "isAutoGenerated": false },
{ "code": "ja", "name": "Japanese", "isAutoGenerated": false },
{ "code": "pt-BR", "name": "Portuguese (Brazil)", "isAutoGenerated": false },
{ "code": "es-419", "name": "Spanish (Latin America)", "isAutoGenerated": false }
],
"segments": [
{ "start": 1.36, "duration": 1.68, "text": "[♪♪♪]" },
{ "start": 18.64, "duration": 3.24, "text": "♪ We're no strangers to love ♪" },
{ "start": 22.64, "duration": 4.32, "text": "♪ You know the rules and so do I ♪" }
],
"wordCount": 366,
"scrapedAt": "2026-09-29T17:14:02.592Z"
}

With outputFormat: "text" the item has text (one string) instead of segments; with srt or vtt it has an srt or vtt string.

A video without a transcript (never charged):

{
"videoId": "ywX-RVl5FFM",
"url": "https://www.youtube.com/watch?v=ywX-RVl5FFM",
"error": "no_captions",
"errorMessage": "This video has no captions.",
"scrapedAt": "2026-09-29T17:20:30.101Z"
}

Error codes: no_captions, unavailable (removed or doesn't exist), geo_restricted (the uploader limits the video to some countries and it could not be read from any of them), private, age_restricted, members_only, live_not_finished (live now, upcoming or premiere), translation_unavailable, translation_failed, no_track_for_filter (only auto captions exist and you turned them off), invalid_input, blocked (YouTube refused all retries), not_processed (maximum cost per run reached).

The dataset has two views: Transcripts and Errors (not charged). Run statistics are stored in the RUN_STATS key-value record.

Tips

  • For LLMs, use outputFormat: "text": it is the most compact.
  • Auto-generated segments are clipped so they don't overlap (YouTube's own auto lines overlap by design).
  • Translated transcripts take longer (1 to 3 minutes each) and carry the extra translation charge: YouTube throttles translation requests much more than original captions.
  • Batch many videos in one run: the start fee is charged once per run.
  • The run only fails when every video failed; otherwise failed videos are listed as error items.

Limits

  • Only public videos. Private, members-only and age-restricted videos (which require a signed-in account) are reported as errors; the Actor never logs in.
  • Videos that are live right now or upcoming have no transcript yet (live_not_finished). Replays work once YouTube has processed them.
  • Transcripts come from YouTube's captions. If a video has no captions at all (typical for music, ambience or silent videos), there is nothing to return.
  • Region-locked videos: when the uploader limits a video to some countries, the Actor automatically retries from a residential IP in one of the allowed countries.
  • publishDate is best-effort and may be null for a few videos.

FAQ and support

Is it legal to scrape YouTube transcripts? This Actor only reads publicly available captions and public video metadata. Our Actors are ethical and do not extract any private user data, such as email addresses, gender, or location. They only extract what the user has chosen to share publicly. We therefore believe that our Actors, when used for ethical purposes by Apify users, are safe. However, you should be aware that your results could contain personal data. Personal data is protected by the GDPR in the European Union and by other regulations around the world. You should not scrape personal data unless you have a legitimate reason to do so. If you're unsure whether your reason is legitimate, consult your lawyers. Respect the copyright of the content you process.

Can I use it from code or an AI agent? Yes: see the API tab for ready-made calls (HTTP, JavaScript, Python, CLI), or use it through the Apify MCP server.

Found a problem or need a feature? Open an issue in the Issues tab.