YouTube Transcript Scraper - Captions & Subtitles for AI/RAG avatar

YouTube Transcript Scraper - Captions & Subtitles for AI/RAG

Pricing

from $4.50 / 1,000 transcripts

Go to Apify Store
YouTube Transcript Scraper - Captions & Subtitles for AI/RAG

YouTube Transcript Scraper - Captions & Subtitles for AI/RAG

Get YouTube transcripts and captions as plain text, timestamped segments, SRT or VTT, with video metadata. Videos, channels and playlists; language preference, auto-generated fallback and translation. Pay only per transcript.

Pricing

from $4.50 / 1,000 transcripts

Rating

0.0

(0)

Developer

Connor Prussin

Connor Prussin

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

19 hours ago

Last modified

Categories

Share

YouTube Transcript Scraper: captions and subtitles for AI and RAG

YouTube Transcript Scraper extracts the transcript (captions / subtitles) of any public YouTube video as clean plain text, timestamped segments, SRT or WebVTT, together with the video's metadata. Paste video links, or a whole channel or playlist, and get one tidy JSON record per video that you can drop straight into an LLM prompt, a vector database or a spreadsheet.

  • ✅ Videos, Shorts, live replays, channels and playlists. Channel and playlist URLs are expanded to their newest videos.
  • ✅ Language control. Pick preferred languages in order; human-made captions win over auto-generated ones. Optional fallback to any language, and optional YouTube machine translation into a target language.
  • ✅ Four output formats: plain text, [{start, duration, text}] segments, SRT and VTT.
  • ✅ Metadata included: title, channel, publish date, duration, views, likes, category, keywords, description, thumbnail.
  • ✅ Pay only for transcripts you get. Videos with no captions, private or deleted videos cost nothing.
  • ✅ Residential proxies built in and included in the price, with automatic retries when YouTube pushes back.

What can I use YouTube transcripts for?

  • RAG and AI assistants: index a channel's talks, lectures or podcasts in a vector store and answer questions with citations (segments carry timestamps, so you can link to the exact second).
  • Summaries and notes: feed the transcript field to ChatGPT, Claude or Gemini to summarize videos, extract action items or write blog posts and show notes.
  • LLM fine-tuning and evaluation datasets built from spoken-language content.
  • Content research and SEO: find what competitors say, mine keywords and topics, repurpose your own videos into articles.
  • Subtitles: download SRT/VTT to re-edit, burn in or translate captions.
  • Market and academic research: analyze what is said across many videos, channels or languages.

How does the YouTube transcript scraper work?

  1. Each URL is parsed. Channel and playlist URLs are expanded to up to maxVideosPerSource of their newest videos.
  2. For every video the actor asks YouTube's own player API which caption tracks exist, picks the best track for your language settings and downloads it.
  3. Captions are cleaned (HTML entities, formatting tags and line breaks removed) and returned in the formats you chose.
  4. If YouTube rate-limits a request, the actor retries with a new residential IP and a different app client.

The actor does not download audio or run speech-to-text. If a video has no captions at all (neither human-made nor auto-generated), you get a free item with errorCode: "noCaptions".

How do I choose which YouTube videos to get transcripts for?

FieldDescriptionDefault
urlsVideo URLs or IDs, channel URLs (@handle, /channel/UC…, /videos, /shorts, /streams), playlist URLstwo sample videos
maxVideosPerSourceNewest videos to take per channel or playlist20
languagesPreferred caption languages, in order["en"]
allowAutoGeneratedUse auto-generated captions when no human-made track matchestrue
fallbackToAnyLanguageIf no preferred language exists, return the default track instead of an errortrue
translateToLanguage code to machine-translate into (YouTube's translation), e.g. ennone
formatsAny of text, segments, srt, vtt["text", "segments"]
includeMetadataAdd channel, publish date, duration, views, description and moretrue
proxyConfigurationProxy settingsApify residential

Example: the 50 newest long-form videos of a channel, English (or translated to English), text and SRT:

{
"urls": ["https://www.youtube.com/@TED/videos"],
"maxVideosPerSource": 50,
"languages": ["en"],
"translateTo": "en",
"formats": ["text", "srt"]
}

What data do you get for each YouTube transcript?

One dataset item per video (shortened):

{
"videoId": "dQw4w9WgXcQ",
"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"channelName": "Rick Astley",
"channelId": "UCuAXFkgsw1L7xaCfnd5JJOw",
"channelUrl": "http://www.youtube.com/@RickAstleyYT",
"publishDate": "2009-10-24T23:57:33-07:00",
"durationSec": 213,
"viewCount": 1821336715,
"likeCount": 19430714,
"category": "Music",
"keywords": ["rick astley", "Never Gonna Give You Up", "..."],
"description": "The official video for “Never Gonna Give You Up” by Rick Astley. ...",
"thumbnailUrl": "https://i.ytimg.com/vi/dQw4w9WgXcQ/sddefault.jpg",
"language": "en",
"languageName": "English",
"isAutoGenerated": false,
"isTranslated": false,
"availableLanguages": [
{ "code": "en", "name": "English", "isAutoGenerated": false },
{
"code": "en",
"name": "English (auto-generated)",
"isAutoGenerated": true
},
{ "code": "de-DE", "name": "German (Germany)", "isAutoGenerated": false }
],
"transcript": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ ...",
"wordCount": 487,
"segments": [
{ "start": 1.36, "duration": 1.68, "text": "[♪♪♪]" },
{
"start": 18.64,
"duration": 3.24,
"text": "♪ We're no strangers to love ♪"
}
],
"source": "https://youtu.be/dQw4w9WgXcQ",
"error": null,
"errorCode": null
}

With formats including srt or vtt, the item also has an srt or vtt string, ready to save as a subtitle file.

Videos without a transcript are still returned, free of charge, with transcript: null and an explanation:

errorCodeMeaning
noCaptionsThe video has no captions of any kind
languageNotAvailableNo track in your languages and fallback is off (availableLanguages lists what exists)
translationNotAvailableYouTube doesn't offer translation for this track
videoUnavailableDeleted, wrong ID or region-blocked
privateVideoPrivate video
loginRequiredAge-restricted or members-only
blockedYouTube kept refusing the request after several retries on new IPs
requestFailedNetwork or unexpected error

How much does it cost to download YouTube transcripts?

Pay per event, no subscription:

EventPrice
Transcript (one video with a transcript)$0.0045 ($4.50 per 1,000)
Actor start$0.00005

Proxy and compute costs are included. Videos that return an error are not charged. If you set a maximum cost per run, the actor stops cleanly when it is reached.

YouTube transcripts FAQ

Does it work for auto-generated captions? Yes. Auto-generated (speech recognition) tracks are used when no human-made captions exist in your languages. isAutoGenerated tells you which one you got.

Can I get a transcript in a language the video doesn't have? Set translateTo. If the video has captions in that language they're used as-is; otherwise YouTube's own machine translation is returned and isTranslated is true.

What about videos without any captions? YouTube has no transcript to give, so you get a free item with errorCode: "noCaptions". This actor does not do speech-to-text.

How many videos can I process? There is no fixed limit. Large channels are paginated; set maxVideosPerSource to how many of the newest videos you need.

Why residential proxies? YouTube often answers requests from cloud servers with "Sign in to confirm you're not a bot". Residential IPs avoid most of that. You can switch to your own proxy or no proxy in the Advanced section, but expect more blocked items without one.

Private, age-restricted or members-only videos? These need a signed-in account, which this actor does not use, so they return an error item and are not charged.

Can AI agents use it? Yes, through the Apify API, the Apify MCP server or any Apify integration (Make, Zapier, n8n, LangChain, LlamaIndex).

Is this legal? The actor only reads publicly available captions, the same data YouTube shows in its "Show transcript" panel. Transcripts are the creators' content: respect copyright and YouTube's Terms of Service in how you use them.

Disclaimer: This actor is not affiliated with, endorsed by or sponsored by YouTube or Google.