YouTube Transcript Scraper - Subtitles & Captions to Text avatar

YouTube Transcript Scraper - Subtitles & Captions to Text

Pricing

from $0.77 / 1,000 transcripts

Go to Apify Store
YouTube Transcript Scraper - Subtitles & Captions to Text

YouTube Transcript Scraper - Subtitles & Captions to Text

Get YouTube transcripts as plain text, timed segments, SRT and VTT, from links, video ids or a keyword search. Typed or auto-generated captions, in the language you ask for when the video has it, with title, channel and views. No login or API key. $0.80 per 1,000 transcripts.

Pricing

from $0.77 / 1,000 transcripts

Rating

0.0

(0)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

1

Monthly active users

2 days ago

Last modified

Share

Scrape YouTube transcripts as text, timed segments, SRT and VTT: paste watch, Shorts, live, embed or youtu.be links or bare 11 character ids, or just type a search, and each video comes back as one row holding the full plain text, the timed segments, a ready-to-save SRT and a ready-to-save WebVTT. No API key, no cookies, no login.

The honest limit: it reads caption tracks the video already has. It does not listen to the audio, so a video nobody captioned and that YouTube never auto-captioned comes back as NO_TRANSCRIPT, and no setting changes that.

InputYouTube links, video ids, or search keywords
OutputOne row per video: full text, timed segments, SRT, WebVTT, the language used and every language available, title, channel, views and length
Ceiling400 videos per run
SpeedRoughly 20 to 30 transcripts a minute on measured cloud runs
Account neededNone from you
Price$0.80 per 1,000 transcripts, flat on every plan. The free plan's $5 a month covers about 6,250

๐Ÿ” What YouTube Transcript Scraper does

For each video it finds the caption tracks, picks one using languagePreferences in order, and pulls the cues. An exact match wins, then a base-language match, so pt-BR falls back to pt. If none of your preferences exist it takes another available track rather than failing the video, and detectedLanguage tells you what it actually used.

Manual captions are preferred over auto-generated ones when both exist. isAutoGenerated says which you got, and availableLanguages lists everything the video had, so you can re-run for a different one.

Type something into searchKeywords instead and it searches YouTube, takes the top maxVideosPerSearch videos from each search, and handles them like pasted links. Shorts, channels and playlists are left out of the results. Each row it finds carries searchKeyword and searchRank.

One charge covers the whole row: the plain text, the segments, the SRT, the VTT, the language list and the metadata. There is no per-format or per-language extra. A video that turns up in more than one place, links or searches, is processed once.

๐Ÿ“‹ What data you get from each YouTube transcript

What you getField
The whole transcript as one stringtext
Every cue with its start, duration and end in secondssegments, segmentCount
Subtitle files, ready to savesrt, vtt
The language you got and the one you asked for firstdetectedLanguage, requestedLanguage
Machine-made track or a person'sisAutoGenerated
Every caption track the video carriedavailableLanguages
Video metadatavideoId, url, title, author, channelId, viewCount, durationSeconds, thumbnailUrl
Which search found it, and where it rankedsearchKeyword, searchRank

โ–ถ๏ธ How to scrape YouTube transcripts

  1. Open YouTube Transcript Scraper and click Try for free.
  2. Press Start with the input empty to see the free sample row.
  3. Paste links or ids into Video URLs or IDs, or type what you want into Search keywords.
  4. Set Preferred transcript languages if English is not what you want.
  5. Download the dataset as JSON, or pull srt out of a row and save it as a subtitle file.

๐Ÿ’ฐ How much does it cost to scrape YouTube transcripts?

$0.80 per 1,000 transcripts, the same on every Apify plan. On the free plan, the $5 Apify gives you each month covers about 6,250 transcripts.

You pay per transcript returned. The sample row, diagnostic rows and duplicate ids are never billed, so a batch of 100 videos where 12 have no captions bills 88 transcripts, not 100.

๐Ÿ“ฅ What you give it

{
"videoUrls": [
"https://www.youtube.com/watch?v=M7lc1UVf-VE",
"https://youtu.be/dQw4w9WgXcQ",
"aqz-KE-bpKQ"
],
"languagePreferences": ["fr-CA", "fr", "en"]
}

Or search instead of pasting:

{
"searchKeywords": ["sourdough for beginners", "no knead bread recipe"],
"maxVideosPerSearch": 10
}
FieldDefaultWhat it is
videoUrls[]Watch, Shorts, live, embed and youtu.be links, or bare ids. Up to 200. The form opens with one example link.
videoIds[]A second list for raw ids, handy when your pipeline stores them separately. Up to 200.
searchKeywords[]Searches to run, one per line. Up to 20. Works on its own or next to links.
maxVideosPerSearch10How many top results to take from each search, 1 to 100.
languagePreferences["en"]BCP 47 codes in priority order, such as en, es, pt-BR. Up to 20.
maxConcurrency2Videos handled at once, 1 to 10. Low on purpose: it keeps the run steady and cheap.
retries2Retries for a video that did not answer, 0 to 5.
requestTimeoutSecs45Total time allowed per video, 10 to 180 seconds.
proxyConfigurationoffLeave it off. A normal run needs nothing here.

๐Ÿ“ค What you get back

A real row from a real run. text, segments, srt and vtt are cut short here, because the full row is about 20 KB of the same transcript in four shapes:

{
"ok": true,
"_sample": false,
"_diagnostic": false,
"videoId": "M7lc1UVf-VE",
"url": "https://www.youtube.com/watch?v=M7lc1UVf-VE",
"title": "YouTube Developers Live: Embedded Web Player Customization",
"author": "Google for Developers",
"channelId": "UC_x5XG1OV2P6uZZ5FSM9Ttw",
"viewCount": 1656845,
"durationSeconds": 1344,
"detectedLanguage": "en",
"requestedLanguage": "en",
"isAutoGenerated": false,
"segmentCount": 466,
"text": "JEFF POSNICK: Hey, everybody. Welcome to this week's show of YouTube Developers Live. ...",
"segments": [
{ "text": "JEFF POSNICK: Hey, everybody.", "start": 10.349, "duration": 1, "end": 11.349 },
{ "text": "Welcome to this week's show of YouTube Developers Live.", "start": 11.349, "duration": 2.681, "end": 14.03 }
],
"srt": "1\n00:00:10,349 --> 00:00:11,349\nJEFF POSNICK: Hey, everybody. ...",
"vtt": "WEBVTT\n\n00:00:10.349 --> 00:00:11.349\nJEFF POSNICK: Hey, everybody. ...",
"availableLanguages": [
{ "languageCode": "en", "languageName": "English", "isAutoGenerated": false },
{ "languageCode": "en", "languageName": "English (auto-generated)", "isAutoGenerated": true }
]
}
FieldHow to read it
segmentsstart, duration and end in seconds, for anything timestamp driven.
srt, vttSave the string straight to .srt or .vtt.
detectedLanguage, requestedLanguageCompare them to spot a fallback.
isAutoGeneratedtrue when the track is machine-made, which is where the punctuation and names get loose.
title, author, channelId, viewCount, thumbnailUrlAny of them can be null when YouTube did not return it.
searchKeyword, searchRankOnly on rows that came from a search.

๐Ÿงพ Reading the output

Three kinds of row can land in your dataset, and the flags are how you tell them apart.

RowHow to spot itBilled
A transcriptok: true and _sample: falseyes
The sample row_sample: true, videoId SAMPLE00000no
A diagnosticok: false, _diagnostic: true, an errorCodeno

The Overview table has no _sample column, so the free sample row reads there as a perfectly good transcript. Filter on _sample == false before counting anything.

The error codes are the useful part of a diagnostic row, and each carries a hint:

errorCodeWhat happened
NO_TRANSCRIPTThe video has no caption track in any language.
VIDEO_UNAVAILABLEPrivate, removed, or the id does not exist.
LOGIN_REQUIREDAge-gated or members-only, so a signed-out reader cannot open it.
INVALID_VIDEOThe string you passed is not a YouTube link or an 11 character id.
RATE_LIMITEDThat video would not answer this run. Re-running usually clears it.
NO_SEARCH_RESULTSA search found no videos. Try fewer or different words.
SEARCH_FAILEDA search could not be run. Try again in a few minutes.
EXTRACTION_FAILEDTracks were found but no cues came back.
TIME_LIMITThe run got close to its time limit and stopped early. Every transcript delivered before that is in your dataset; give the run more time, or ask for fewer videos, to get the rest. If none came back at all, the run is marked as failed.
RUN_ERRORThe run itself failed, written as one row so the reason survives.

A batch where every video fails still finishes as a successful run, with diagnostics and nothing else. Count ok: true rows rather than trusting the run status.

๐Ÿ’ก What people use it for

  • Turning a talk or interview into text you can search, quote and paste into notes.
  • Feeding transcripts to a model for summaries, without paying a speech-to-text bill.
  • Making .srt files for re-uploads, or as the starting point for a translation pass.
  • Finding the exact moment a phrase was said, using segments to jump to the second.

From a topic to what the top videos say and what viewers ask, in two runs:

  1. Type a few phrases into Search keywords here.
  2. Sort the rows on viewCount and read the first segments of the top videos to see how they open.
  3. Paste their url into the Video URLs box of YouTube Comments Scraper to see the questions viewers left.

๐Ÿšง What it does not do

  • No speech to text. Published caption tracks only, auto-generated ones included. Nothing captioned means nothing to return.
  • No translation. It fetches tracks that exist. If the video has no French track, asking for French gets you another language, not a translation.
  • One video per entry. Channel and playlist links are not expanded here. YouTube Transcript Bulk Scraper does that.
  • durationSeconds can be the transcript's length, not the video's, when YouTube returned no metadata for that video. Sanity check it against the last segment's end if it matters.
  • Metadata can be null. Title, author, view count and thumbnail come from YouTube's own response and are sometimes absent, while the transcript itself is fine.
  • Auto-generated tracks have no reliable punctuation or speaker names, and get proper nouns wrong. isAutoGenerated is how you know which you are reading.
  • Rows are large, roughly 18 to 24 KB each, because text, segments, SRT and VTT hold the same words four ways. A 400 video run is a chunky dataset.
  • Search takes YouTube's own top results. There is no filter for upload date, length or views, and a video without captions still uses up one of the places you asked for.
  • 400 videos per run in total, links and search results together.

๐Ÿงญ Which YouTube scraper do you need?

If you wantUse
What a video says, as text, SRT and VTT, for links you already haveThis one
Transcripts for a whole channel or playlist in one runYouTube Transcript Bulk Scraper
The video, its audio or its subtitles as a fileYouTube Downloader Pro
Comments under a videoYouTube Comments Scraper
Videos to run this on, by keyword or a channel's newest uploadsYouTube Scraper
Speech from a video file that has no captions, with your own OpenAI keyVideo & Audio Transcriber
The same job on TikTokTikTok Transcript Scraper

โ“ Questions people ask

Do I need a YouTube Data API key?

No. No key, no cookies, no OAuth, no quota.

What if the video has no captions?

You get an uncharged NO_TRANSCRIPT row and the batch carries on. It does not transcribe the audio.

Which language do I get?

The first one in languagePreferences that exists, an exact match before a base-language one. Otherwise another available track, named in detectedLanguage.

Can I get the SRT file directly?

The srt field is the file's contents. Write the string to a .srt and it plays.

Why is one video missing from my results?

Look for its diagnostic row. LOGIN_REQUIRED and VIDEO_UNAVAILABLE are permanent for a signed-out reader, RATE_LIMITED is worth retrying.

Can I call it from code or connect it to an AI assistant?

Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/youtube-transcript-scraper. Either way the run happens on your Apify account at the same price.

This reads publicly available captions. What you then do with somebody else's words is your call, and copyright still applies. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

๐Ÿ†˜ If something breaks

Open the Issues tab on the actor page. Include the video link and the run ID, and the errorCode on the diagnostic row if you got one.