YouTube Transcript Scraper avatar

YouTube Transcript Scraper

Pricing

from $17.00 / 1,000 single video's transcripts

Go to Apify Store
YouTube Transcript Scraper

YouTube Transcript Scraper

YouTube Transcript Scraper: get transcripts of up to 10 YouTube videos or Shorts per run, in any language, as text, timestamped text, SRT or VTT. $0.02 each.

Pricing

from $17.00 / 1,000 single video's transcripts

Rating

0.0

(0)

Developer

Monster Leads

Monster Leads

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

YouTube Transcript Scraper gets the full transcript of any YouTube video, Short or finished live stream. Paste in links or video IDs and you get clean text, timestamped lines, timed segments, or ready-to-use SRT and WebVTT subtitle files, in any language the video has. Each run takes up to 10 videos and costs $0.02 per transcript.

You don't need a YouTube API key, a Google account or a browser. Proxies and retries are handled for you, so there is nothing to set up. You pay only for transcripts you receive.

  • ✅ Transcripts from videos, Shorts and live stream replays
  • ✅ Creator captions and auto-generated captions
  • ✅ Any language, with your own order of preference
  • ✅ 5 output formats: plain text, timestamped text, JSON segments, SRT, VTT
  • ✅ Video title, channel, length and word count with every transcript
  • ✅ Up to 10 videos per run at a flat $0.02 per transcript
  • ✅ No charge for videos that have no transcript

What can you do with a YouTube transcript scraper?

  • Feed AI and LLMs. Summarize videos, build RAG knowledge bases, or train on spoken content with ChatGPT, Claude or your own models.
  • Repurpose content. Turn videos into blog posts, newsletters, show notes and social posts.
  • Research and analysis. Run keyword, sentiment and topic analysis across whole channels or niches.
  • SEO. Find the exact words competitors use in their videos.
  • Subtitles. Download SRT or VTT files to edit, translate or re-upload.
  • Accessibility and archiving. Keep a searchable text copy of video content.

How to scrape YouTube transcripts

  1. Click Try for free at the top of this page.
  2. Paste up to 10 YouTube video URLs or IDs into Video URLs or IDs.
  3. Optionally, pick the transcript language(s) and the format.
  4. Click Start. When the run finishes, download the transcripts as JSON, CSV, Excel, XML or HTML, or fetch them through the API.

Input options

Every option has a default, so an empty input {} works: it fetches the transcript of Rick Astley's "Never Gonna Give You Up". In practice you'll always set videoIds.

OptionTypeDefaultWhat it does
videoIdsarray of stringsRick Astley, dQw4w9WgXcQThe videos to scrape, one per entry. Maximum 10.
languagesarray of strings["en"]Languages you want, in order of preference.
includeAutoGeneratedbooleantrueAllow YouTube's speech-recognition captions.
fallbackToAnyLanguagebooleantrueIf none of your languages exist, return whatever language the video has.
outputFormatstring"segments"segments, text, timestamped, srt or vtt.

videoIds — the videos to scrape (max 10)

You can pass up to 10 videos per run. A run with more than 10 entries is rejected before it starts, and you are not charged. For more videos, start another run or call the API once per batch of 10.

Any form of YouTube link works, and so does the bare 11-character ID:

"videoIds": [
"dQw4w9WgXcQ",
"https://www.youtube.com/watch?v=jNQXAC9IVRw",
"https://youtu.be/9bZkp7q19f0",
"https://www.youtube.com/shorts/tPEE9ZwTmy0",
"https://www.youtube.com/live/VIDEO_ID",
"https://www.youtube.com/embed/VIDEO_ID"
]

Extra URL parameters such as &t=30s or &list=... are ignored. A video listed twice is fetched and charged only once.

languages — which transcript language you want

This is a list of ISO language codes such as en, es, de, fr, pt-BR or ja, in order of preference. For each video, the scraper goes down your list and uses the first language the video has.

"languages": ["es", "en"]

This means Spanish if the video has it, otherwise English.

  • Regional codes match each other. en finds en-GB or en-US, and pt-BR finds pt. An exact match is tried first.
  • Within each language, creator-uploaded captions are always preferred over auto-generated ones.
  • Check the availableLanguages field of a result to see every language the video offers.

includeAutoGenerated — allow auto-generated captions

YouTube has two kinds of captions: ones the creator uploaded, and ones YouTube makes with speech recognition.

ValueResult
true (default)If a language has only auto-generated captions, use them. Most videos have only these.
falseUse creator-uploaded captions only. Videos without them are skipped and not charged.
"includeAutoGenerated": false

fallbackToAnyLanguage — what to do if your language isn't there

ValueResult
true (default)Return the transcript in a language the video does have. Check the language field to see which one.
falseSkip the video. Nothing is stored and you are not charged. The run log lists the languages it had.
"fallbackToAnyLanguage": false

Set this to false when you need a specific language and a different one would be useless to you.

outputFormat — how the transcript is written

This sets how the transcript field of each result is written. There are five choices: segments (default), text, timestamped, srt and vtt.

"outputFormat": "srt"

See Transcript formats explained below for what each one returns, with a real example.

Full input example

{
"videoIds": [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"https://youtu.be/jNQXAC9IVRw"
],
"languages": ["en", "es"],
"includeAutoGenerated": true,
"fallbackToAnyLanguage": false,
"outputFormat": "timestamped"
}

Output

You get one dataset item per video, up to 10 per run. For example:

{
"videoId": "dQw4w9WgXcQ",
"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"channel": "Rick Astley",
"channelId": "UCuAXFkgsw1L7xaCfnd5JJOw",
"lengthSeconds": 213,
"language": "en",
"languageName": "English",
"isAutoGenerated": false,
"availableLanguages": [
{ "language": "en", "name": "English", "isAutoGenerated": false },
{ "language": "en", "name": "English (auto-generated)", "isAutoGenerated": true },
{ "language": "de-DE", "name": "German (Germany)", "isAutoGenerated": false }
],
"segmentCount": 61,
"wordCount": 487,
"transcript": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ ...",
"segments": [
{ "start": 1.36, "duration": 1.68, "text": "[♪♪♪]" },
{ "start": 18.64, "duration": 3.24, "text": "♪ We're no strangers to love ♪" }
],
"format": "segments"
}
FieldDescription
videoIdThe 11-character YouTube video ID
urlThe video's watch URL
titleVideo title
channel, channelIdChannel name and ID
lengthSecondsVideo length in seconds
languageLanguage code of the transcript you got, such as en or de-DE
languageNameLanguage name as YouTube shows it
isAutoGeneratedtrue if this is YouTube's speech-recognition transcript
availableLanguagesEvery caption track the video has
segmentCountNumber of caption lines
wordCountNumber of words in the transcript
transcriptThe transcript in your chosen outputFormat
segmentsTimed lines. Only present when outputFormat is segments
formatThe outputFormat used

You can download the dataset as JSON, CSV, Excel, XML, HTML or RSS from the run page or through the API.

Transcript formats explained

The format only changes the transcript field (and whether a segments field is added). Every other field, like title, language and wordCount, is the same in all five formats.

FormatInput valuetranscript fieldExtra fieldBest for
Text + timed segmentssegments (default)Plain textsegments: every line with its timeApps that need both text and timing
Plain texttextPlain text—AI and LLMs, search, word counts, summaries
Timestamped texttimestampedOne line per caption, with [m:ss] time—Reading, show notes, quoting with time links
SRT subtitlessrtA complete .srt file—Video editors, re-uploading captions to YouTube
WebVTT subtitlesvttA complete .vtt file—HTML5 <video> players, web apps

All examples below are real output for Rick Astley - Never Gonna Give You Up, shortened to the first few lines. In the JSON, line breaks inside transcript are written as \n. They become real line breaks when you print the value or save it to a file.

1. Text + timed segments (segments), the default

You get two things: the whole transcript as plain text in transcript, and a segments list with every caption line and its timing. start is when the line appears and duration is how long it stays on screen, both in seconds.

Input

{ "videoIds": ["dQw4w9WgXcQ"], "outputFormat": "segments" }

Returns

{
"videoId": "dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"language": "en",
"segmentCount": 61,
"wordCount": 487,
"transcript": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ ♪ A full commitment's what I'm thinking of ♪ ...",
"segments": [
{ "start": 1.36, "duration": 1.68, "text": "[♪♪♪]" },
{ "start": 18.64, "duration": 3.24, "text": "♪ We're no strangers to love ♪" },
{ "start": 22.64, "duration": 4.32, "text": "♪ You know the rules and so do I ♪" },
{ "start": 27.04, "duration": 4, "text": "♪ A full commitment's what I'm thinking of ♪" }
],
"format": "segments"
}

In the Apify Console, the Overview tab shows only the plain text. Open the Timed segments tab to see one row per caption line:

Video IDTitleStart (s)Duration (s)Text
dQw4w9WgXcQRick Astley - Never Gonna Give You Up…1.361.68[♪♪♪]
dQw4w9WgXcQRick Astley - Never Gonna Give You Up…18.643.24♪ We're no strangers to love ♪
dQw4w9WgXcQRick Astley - Never Gonna Give You Up…22.644.32♪ You know the rules and so do I ♪

Use it when you need to know when something was said: jumping to a moment in the video, cutting clips, syncing text with playback, or making chapters. It is the most complete format, since you can build any of the others from segments.

2. Plain text (text)

The whole transcript as one paragraph, with caption lines joined by spaces. There are no times and no line breaks.

Input

{ "videoIds": ["dQw4w9WgXcQ"], "outputFormat": "text" }

Returns

{
"videoId": "dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"language": "en",
"wordCount": 487,
"transcript": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ ♪ A full commitment's what I'm thinking of ♪ ♪ You wouldn't get this from any other guy ♪ ...",
"format": "text"
}

Use it when you only care about the words: pasting into ChatGPT or Claude, summaries, RAG and search indexes, keyword or sentiment analysis, or turning a video into a blog post. It is also the smallest output.

3. Timestamped text (timestamped)

One line per caption, each starting with the time it appears in the video. Times are m:ss, or h:mm:ss for videos over an hour.

Input

{ "videoIds": ["dQw4w9WgXcQ"], "outputFormat": "timestamped" }

Returns

{
"videoId": "dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"language": "en",
"wordCount": 487,
"transcript": "[0:01] [♪♪♪]\n[0:18] ♪ We're no strangers to love ♪\n[0:22] ♪ You know the rules and so do I ♪\n[0:27] ♪ A full commitment's what I'm thinking of ♪\n...",
"format": "timestamped"
}

When printed or saved

[0:01] [♪♪♪]
[0:18] ♪ We're no strangers to love ♪
[0:22] ♪ You know the rules and so do I ♪
[0:27] ♪ A full commitment's what I'm thinking of ♪
[0:31] ♪ You wouldn't get this from any other guy ♪

Use it when people will read the transcript: show notes, meeting or lecture notes, quoting a video with the time, or giving an AI a transcript it can cite by timestamp.

4. SRT subtitles (srt)

A complete, ready-to-use SubRip (.srt) subtitle file. Each caption has a number, a start and end time (hh:mm:ss,mmm), and its text. Captions never overlap, so they don't stack on screen.

Input

{ "videoIds": ["dQw4w9WgXcQ"], "outputFormat": "srt" }

Returns

{
"videoId": "dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"language": "en",
"wordCount": 487,
"transcript": "1\n00:00:01,360 --> 00:00:03,040\n[♪♪♪]\n\n2\n00:00:18,640 --> 00:00:21,880\n♪ We're no strangers to love ♪\n\n3\n00:00:22,640 --> 00:00:26,960\n♪ You know the rules and so do I ♪\n...",
"format": "srt"
}

When saved as video.srt

1
00:00:01,360 --> 00:00:03,040
[♪♪♪]
2
00:00:18,640 --> 00:00:21,880
♪ We're no strangers to love ♪
3
00:00:22,640 --> 00:00:26,960
♪ You know the rules and so do I ♪

Use it when you want subtitles for a video file: Premiere Pro, DaVinci Resolve, Final Cut, CapCut, VLC, or re-uploading captions to YouTube or other platforms. You can also translate the text and keep the timing.

5. WebVTT subtitles (vtt)

A complete WebVTT (.vtt) subtitle file. It is the web version of SRT: it starts with a WEBVTT header, uses a dot in times (hh:mm:ss.mmm), and has no cue numbers.

Input

{ "videoIds": ["dQw4w9WgXcQ"], "outputFormat": "vtt" }

Returns

{
"videoId": "dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"language": "en",
"wordCount": 487,
"transcript": "WEBVTT\n\n00:00:01.360 --> 00:00:03.040\n[♪♪♪]\n\n00:00:18.640 --> 00:00:21.880\n♪ We're no strangers to love ♪\n\n00:00:22.640 --> 00:00:26.960\n♪ You know the rules and so do I ♪\n...",
"format": "vtt"
}

When saved as video.vtt

WEBVTT
00:00:01.360 --> 00:00:03.040
[♪♪♪]
00:00:18.640 --> 00:00:21.880
♪ We're no strangers to love ♪
00:00:22.640 --> 00:00:26.960
♪ You know the rules and so do I ♪

Use it when you play video on a website: the HTML5 <track> element, Video.js, Plyr, JW Player and most web players read VTT directly.

Saving SRT or VTT as a file

The subtitle file is the value of transcript. Save it as-is with the right extension:

from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("ecommerce_leads/youtube-transcript-scraper").call(
run_input={"videoIds": ["dQw4w9WgXcQ"], "outputFormat": "srt"}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
with open(f"{item['videoId']}.srt", "w", encoding="utf-8") as f:
f.write(item["transcript"])

Use the YouTube Transcript Scraper through the API

You can run the scraper from your own code with the Apify API. The request body is the same JSON as the input examples above, with up to 10 videos per call. Replace YOUR_APIFY_TOKEN with the token from Console → Settings → Integrations.

cURL

This runs the scraper and returns the transcripts in one request:

curl -X POST \
"https://api.apify.com/v2/acts/ecommerce_leads~youtube-transcript-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"videoIds": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"languages": ["en"],
"outputFormat": "text"
}'

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("ecommerce_leads/youtube-transcript-scraper").call(run_input={
"videoIds": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"languages": ["en"],
"outputFormat": "text",
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item["title"], item["language"])
print(item["transcript"])

JavaScript / Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('ecommerce_leads/youtube-transcript-scraper').call({
videoIds: ['https://www.youtube.com/watch?v=dQw4w9WgXcQ'],
languages: ['en'],
outputFormat: 'text',
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => console.log(item.title, item.transcript));

You can also connect the scraper to Make, Zapier, n8n, Google Sheets, Slack and other tools through Apify integrations, or give it to AI agents through the Apify MCP server.

How much does it cost to scrape YouTube transcripts?

This scraper uses pay per event pricing: $0.02 per transcript stored, and nothing else. There is no start fee.

Videos in the runTranscripts returnedYou pay
11$0.02
1010$0.20
107 (3 had no captions)$0.14
100$0.00

A run takes at most 10 videos, so a single run never costs more than $0.20.

You are not charged for videos that return no transcript: private, removed or age-restricted videos, videos without captions, or videos skipped because fallbackToAnyLanguage or includeAutoGenerated is false. A run that returns nothing costs nothing.

Limitations

  • 10 videos per run. Longer lists are rejected. Split them into runs of 10.
  • Only captions the video has. YouTube's automatic translation is not available. A video with only English captions can't be returned in Spanish.
  • Age-restricted videos need a signed-in account, so they are skipped.
  • Live streams still in progress usually don't have a transcript yet.
  • Auto-generated transcripts come from YouTube's speech recognition. In many languages they have no punctuation, and they can get names and jargon wrong.

FAQ

Do I need a YouTube API key?

No. The YouTube Transcript Scraper works without a YouTube Data API key or any Google account.

Can I get transcripts for YouTube Shorts?

Yes. Paste the /shorts/ URL, or just the video ID.

How many videos can I scrape in one run?

Up to 10. Each video is one dataset item. For more, start several runs, or loop over batches of 10 through the API.

How much does one transcript cost?

$0.02 per transcript returned. Videos with no transcript are free, and a run of 10 videos costs at most $0.20.

How do I get a transcript in a specific language only?

Set languages to that language, for example ["de"], and set fallbackToAnyLanguage to false. Videos without German are skipped and not charged.

How do I download YouTube subtitles as an SRT file?

Set outputFormat to srt. The transcript field then holds a complete SRT file you can save as .srt. For HTML5 players, use vtt.

Why did I get a different language than I asked for?

fallbackToAnyLanguage is on by default, so a video without your language comes back in one it has. Check the language field, or set fallbackToAnyLanguage to false.

How can I tell whether a transcript was written by a person?

Look at isAutoGenerated. false means the creator uploaded the captions. To get only those, set includeAutoGenerated to false.

This scraper only collects publicly available captions. Some transcripts may still be copyrighted, so check that your use complies with YouTube's terms and the laws that apply to you. If you're unsure, ask a lawyer.

Feedback

Found a bug or need a feature? Open an issue in the Issues tab and we'll get back to you.