YouTube Transcript Scraper
Pricing
from $17.00 / 1,000 single video's transcripts
YouTube Transcript Scraper
YouTube Transcript Scraper: get transcripts of up to 10 YouTube videos or Shorts per run, in any language, as text, timestamped text, SRT or VTT. $0.02 each.
Pricing
from $17.00 / 1,000 single video's transcripts
Rating
0.0
(0)
Developer
Monster Leads
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
YouTube Transcript Scraper gets the full transcript of any YouTube video, Short or finished live stream. Paste in links or video IDs and you get clean text, timestamped lines, timed segments, or ready-to-use SRT and WebVTT subtitle files, in any language the video has. Each run takes up to 10 videos and costs $0.02 per transcript.
You don't need a YouTube API key, a Google account or a browser. Proxies and retries are handled for you, so there is nothing to set up. You pay only for transcripts you receive.
- ✅ Transcripts from videos, Shorts and live stream replays
- ✅ Creator captions and auto-generated captions
- ✅ Any language, with your own order of preference
- ✅ 5 output formats: plain text, timestamped text, JSON segments, SRT, VTT
- ✅ Video title, channel, length and word count with every transcript
- ✅ Up to 10 videos per run at a flat $0.02 per transcript
- ✅ No charge for videos that have no transcript
What can you do with a YouTube transcript scraper?
- Feed AI and LLMs. Summarize videos, build RAG knowledge bases, or train on spoken content with ChatGPT, Claude or your own models.
- Repurpose content. Turn videos into blog posts, newsletters, show notes and social posts.
- Research and analysis. Run keyword, sentiment and topic analysis across whole channels or niches.
- SEO. Find the exact words competitors use in their videos.
- Subtitles. Download SRT or VTT files to edit, translate or re-upload.
- Accessibility and archiving. Keep a searchable text copy of video content.
How to scrape YouTube transcripts
- Click Try for free at the top of this page.
- Paste up to 10 YouTube video URLs or IDs into Video URLs or IDs.
- Optionally, pick the transcript language(s) and the format.
- Click Start. When the run finishes, download the transcripts as JSON, CSV, Excel, XML or HTML, or fetch them through the API.
Input options
Every option has a default, so an empty input {} works: it fetches the transcript of Rick Astley's "Never Gonna Give You Up". In practice you'll always set videoIds.
| Option | Type | Default | What it does |
|---|---|---|---|
videoIds | array of strings | Rick Astley, dQw4w9WgXcQ | The videos to scrape, one per entry. Maximum 10. |
languages | array of strings | ["en"] | Languages you want, in order of preference. |
includeAutoGenerated | boolean | true | Allow YouTube's speech-recognition captions. |
fallbackToAnyLanguage | boolean | true | If none of your languages exist, return whatever language the video has. |
outputFormat | string | "segments" | segments, text, timestamped, srt or vtt. |
videoIds — the videos to scrape (max 10)
You can pass up to 10 videos per run. A run with more than 10 entries is rejected before it starts, and you are not charged. For more videos, start another run or call the API once per batch of 10.
Any form of YouTube link works, and so does the bare 11-character ID:
"videoIds": ["dQw4w9WgXcQ","https://www.youtube.com/watch?v=jNQXAC9IVRw","https://youtu.be/9bZkp7q19f0","https://www.youtube.com/shorts/tPEE9ZwTmy0","https://www.youtube.com/live/VIDEO_ID","https://www.youtube.com/embed/VIDEO_ID"]
Extra URL parameters such as &t=30s or &list=... are ignored. A video listed twice
is fetched and charged only once.
languages — which transcript language you want
This is a list of ISO language codes such as en, es, de, fr, pt-BR or ja,
in order of preference. For each video, the scraper goes down your list and uses the
first language the video has.
"languages": ["es", "en"]
This means Spanish if the video has it, otherwise English.
- Regional codes match each other.
enfindsen-GBoren-US, andpt-BRfindspt. An exact match is tried first. - Within each language, creator-uploaded captions are always preferred over auto-generated ones.
- Check the
availableLanguagesfield of a result to see every language the video offers.
includeAutoGenerated — allow auto-generated captions
YouTube has two kinds of captions: ones the creator uploaded, and ones YouTube makes with speech recognition.
| Value | Result |
|---|---|
true (default) | If a language has only auto-generated captions, use them. Most videos have only these. |
false | Use creator-uploaded captions only. Videos without them are skipped and not charged. |
"includeAutoGenerated": false
fallbackToAnyLanguage — what to do if your language isn't there
| Value | Result |
|---|---|
true (default) | Return the transcript in a language the video does have. Check the language field to see which one. |
false | Skip the video. Nothing is stored and you are not charged. The run log lists the languages it had. |
"fallbackToAnyLanguage": false
Set this to false when you need a specific language and a different one would be
useless to you.
outputFormat — how the transcript is written
This sets how the transcript field of each result is written. There are five
choices: segments (default), text, timestamped, srt and vtt.
"outputFormat": "srt"
See Transcript formats explained below for what each one returns, with a real example.
Full input example
{"videoIds": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ","https://youtu.be/jNQXAC9IVRw"],"languages": ["en", "es"],"includeAutoGenerated": true,"fallbackToAnyLanguage": false,"outputFormat": "timestamped"}
Output
You get one dataset item per video, up to 10 per run. For example:
{"videoId": "dQw4w9WgXcQ","url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","channel": "Rick Astley","channelId": "UCuAXFkgsw1L7xaCfnd5JJOw","lengthSeconds": 213,"language": "en","languageName": "English","isAutoGenerated": false,"availableLanguages": [{ "language": "en", "name": "English", "isAutoGenerated": false },{ "language": "en", "name": "English (auto-generated)", "isAutoGenerated": true },{ "language": "de-DE", "name": "German (Germany)", "isAutoGenerated": false }],"segmentCount": 61,"wordCount": 487,"transcript": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ ...","segments": [{ "start": 1.36, "duration": 1.68, "text": "[♪♪♪]" },{ "start": 18.64, "duration": 3.24, "text": "♪ We're no strangers to love ♪" }],"format": "segments"}
| Field | Description |
|---|---|
videoId | The 11-character YouTube video ID |
url | The video's watch URL |
title | Video title |
channel, channelId | Channel name and ID |
lengthSeconds | Video length in seconds |
language | Language code of the transcript you got, such as en or de-DE |
languageName | Language name as YouTube shows it |
isAutoGenerated | true if this is YouTube's speech-recognition transcript |
availableLanguages | Every caption track the video has |
segmentCount | Number of caption lines |
wordCount | Number of words in the transcript |
transcript | The transcript in your chosen outputFormat |
segments | Timed lines. Only present when outputFormat is segments |
format | The outputFormat used |
You can download the dataset as JSON, CSV, Excel, XML, HTML or RSS from the run page or through the API.
Transcript formats explained
The format only changes the transcript field (and whether a segments field is
added). Every other field, like title, language and wordCount, is the same in
all five formats.
| Format | Input value | transcript field | Extra field | Best for |
|---|---|---|---|---|
| Text + timed segments | segments (default) | Plain text | segments: every line with its time | Apps that need both text and timing |
| Plain text | text | Plain text | — | AI and LLMs, search, word counts, summaries |
| Timestamped text | timestamped | One line per caption, with [m:ss] time | — | Reading, show notes, quoting with time links |
| SRT subtitles | srt | A complete .srt file | — | Video editors, re-uploading captions to YouTube |
| WebVTT subtitles | vtt | A complete .vtt file | — | HTML5 <video> players, web apps |
All examples below are real output for
Rick Astley - Never Gonna Give You Up,
shortened to the first few lines. In the JSON, line breaks inside transcript are
written as \n. They become real line breaks when you print the value or save it to a
file.
1. Text + timed segments (segments), the default
You get two things: the whole transcript as plain text in transcript, and a
segments list with every caption line and its timing. start is when the line
appears and duration is how long it stays on screen, both in seconds.
Input
{ "videoIds": ["dQw4w9WgXcQ"], "outputFormat": "segments" }
Returns
{"videoId": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","language": "en","segmentCount": 61,"wordCount": 487,"transcript": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ ♪ A full commitment's what I'm thinking of ♪ ...","segments": [{ "start": 1.36, "duration": 1.68, "text": "[♪♪♪]" },{ "start": 18.64, "duration": 3.24, "text": "♪ We're no strangers to love ♪" },{ "start": 22.64, "duration": 4.32, "text": "♪ You know the rules and so do I ♪" },{ "start": 27.04, "duration": 4, "text": "♪ A full commitment's what I'm thinking of ♪" }],"format": "segments"}
In the Apify Console, the Overview tab shows only the plain text. Open the Timed segments tab to see one row per caption line:
| Video ID | Title | Start (s) | Duration (s) | Text |
|---|---|---|---|---|
| dQw4w9WgXcQ | Rick Astley - Never Gonna Give You Up… | 1.36 | 1.68 | [♪♪♪] |
| dQw4w9WgXcQ | Rick Astley - Never Gonna Give You Up… | 18.64 | 3.24 | ♪ We're no strangers to love ♪ |
| dQw4w9WgXcQ | Rick Astley - Never Gonna Give You Up… | 22.64 | 4.32 | ♪ You know the rules and so do I ♪ |
Use it when you need to know when something was said: jumping to a moment in
the video, cutting clips, syncing text with playback, or making chapters. It is the
most complete format, since you can build any of the others from segments.
2. Plain text (text)
The whole transcript as one paragraph, with caption lines joined by spaces. There are no times and no line breaks.
Input
{ "videoIds": ["dQw4w9WgXcQ"], "outputFormat": "text" }
Returns
{"videoId": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","language": "en","wordCount": 487,"transcript": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ ♪ A full commitment's what I'm thinking of ♪ ♪ You wouldn't get this from any other guy ♪ ...","format": "text"}
Use it when you only care about the words: pasting into ChatGPT or Claude, summaries, RAG and search indexes, keyword or sentiment analysis, or turning a video into a blog post. It is also the smallest output.
3. Timestamped text (timestamped)
One line per caption, each starting with the time it appears in the video. Times are
m:ss, or h:mm:ss for videos over an hour.
Input
{ "videoIds": ["dQw4w9WgXcQ"], "outputFormat": "timestamped" }
Returns
{"videoId": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","language": "en","wordCount": 487,"transcript": "[0:01] [♪♪♪]\n[0:18] ♪ We're no strangers to love ♪\n[0:22] ♪ You know the rules and so do I ♪\n[0:27] ♪ A full commitment's what I'm thinking of ♪\n...","format": "timestamped"}
When printed or saved
[0:01] [♪♪♪][0:18] ♪ We're no strangers to love ♪[0:22] ♪ You know the rules and so do I ♪[0:27] ♪ A full commitment's what I'm thinking of ♪[0:31] ♪ You wouldn't get this from any other guy ♪
Use it when people will read the transcript: show notes, meeting or lecture notes, quoting a video with the time, or giving an AI a transcript it can cite by timestamp.
4. SRT subtitles (srt)
A complete, ready-to-use SubRip (.srt) subtitle file. Each caption has a number, a
start and end time (hh:mm:ss,mmm), and its text. Captions never overlap, so they
don't stack on screen.
Input
{ "videoIds": ["dQw4w9WgXcQ"], "outputFormat": "srt" }
Returns
{"videoId": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","language": "en","wordCount": 487,"transcript": "1\n00:00:01,360 --> 00:00:03,040\n[♪♪♪]\n\n2\n00:00:18,640 --> 00:00:21,880\n♪ We're no strangers to love ♪\n\n3\n00:00:22,640 --> 00:00:26,960\n♪ You know the rules and so do I ♪\n...","format": "srt"}
When saved as video.srt
100:00:01,360 --> 00:00:03,040[♪♪♪]200:00:18,640 --> 00:00:21,880♪ We're no strangers to love ♪300:00:22,640 --> 00:00:26,960♪ You know the rules and so do I ♪
Use it when you want subtitles for a video file: Premiere Pro, DaVinci Resolve, Final Cut, CapCut, VLC, or re-uploading captions to YouTube or other platforms. You can also translate the text and keep the timing.
5. WebVTT subtitles (vtt)
A complete WebVTT (.vtt) subtitle file. It is the web version of SRT: it starts
with a WEBVTT header, uses a dot in times (hh:mm:ss.mmm), and has no cue numbers.
Input
{ "videoIds": ["dQw4w9WgXcQ"], "outputFormat": "vtt" }
Returns
{"videoId": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","language": "en","wordCount": 487,"transcript": "WEBVTT\n\n00:00:01.360 --> 00:00:03.040\n[♪♪♪]\n\n00:00:18.640 --> 00:00:21.880\n♪ We're no strangers to love ♪\n\n00:00:22.640 --> 00:00:26.960\n♪ You know the rules and so do I ♪\n...","format": "vtt"}
When saved as video.vtt
WEBVTT00:00:01.360 --> 00:00:03.040[♪♪♪]00:00:18.640 --> 00:00:21.880♪ We're no strangers to love ♪00:00:22.640 --> 00:00:26.960♪ You know the rules and so do I ♪
Use it when you play video on a website: the HTML5 <track> element, Video.js,
Plyr, JW Player and most web players read VTT directly.
Saving SRT or VTT as a file
The subtitle file is the value of transcript. Save it as-is with the right
extension:
from apify_client import ApifyClientclient = ApifyClient("YOUR_APIFY_TOKEN")run = client.actor("ecommerce_leads/youtube-transcript-scraper").call(run_input={"videoIds": ["dQw4w9WgXcQ"], "outputFormat": "srt"})for item in client.dataset(run["defaultDatasetId"]).iterate_items():with open(f"{item['videoId']}.srt", "w", encoding="utf-8") as f:f.write(item["transcript"])
Use the YouTube Transcript Scraper through the API
You can run the scraper from your own code with the
Apify API. The request body is the same JSON as the
input examples above, with up to 10 videos per call. Replace YOUR_APIFY_TOKEN with the token from
Console → Settings → Integrations.
cURL
This runs the scraper and returns the transcripts in one request:
curl -X POST \"https://api.apify.com/v2/acts/ecommerce_leads~youtube-transcript-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"videoIds": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"languages": ["en"],"outputFormat": "text"}'
Python
from apify_client import ApifyClientclient = ApifyClient("YOUR_APIFY_TOKEN")run = client.actor("ecommerce_leads/youtube-transcript-scraper").call(run_input={"videoIds": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"languages": ["en"],"outputFormat": "text",})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item["title"], item["language"])print(item["transcript"])
JavaScript / Node.js
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });const run = await client.actor('ecommerce_leads/youtube-transcript-scraper').call({videoIds: ['https://www.youtube.com/watch?v=dQw4w9WgXcQ'],languages: ['en'],outputFormat: 'text',});const { items } = await client.dataset(run.defaultDatasetId).listItems();items.forEach((item) => console.log(item.title, item.transcript));
You can also connect the scraper to Make, Zapier, n8n, Google Sheets, Slack and other tools through Apify integrations, or give it to AI agents through the Apify MCP server.
How much does it cost to scrape YouTube transcripts?
This scraper uses pay per event pricing: $0.02 per transcript stored, and nothing else. There is no start fee.
| Videos in the run | Transcripts returned | You pay |
|---|---|---|
| 1 | 1 | $0.02 |
| 10 | 10 | $0.20 |
| 10 | 7 (3 had no captions) | $0.14 |
| 10 | 0 | $0.00 |
A run takes at most 10 videos, so a single run never costs more than $0.20.
You are not charged for videos that return no transcript: private, removed or
age-restricted videos, videos without captions, or videos skipped because
fallbackToAnyLanguage or includeAutoGenerated is false. A run that returns
nothing costs nothing.
Limitations
- 10 videos per run. Longer lists are rejected. Split them into runs of 10.
- Only captions the video has. YouTube's automatic translation is not available. A video with only English captions can't be returned in Spanish.
- Age-restricted videos need a signed-in account, so they are skipped.
- Live streams still in progress usually don't have a transcript yet.
- Auto-generated transcripts come from YouTube's speech recognition. In many languages they have no punctuation, and they can get names and jargon wrong.
FAQ
Do I need a YouTube API key?
No. The YouTube Transcript Scraper works without a YouTube Data API key or any Google account.
Can I get transcripts for YouTube Shorts?
Yes. Paste the /shorts/ URL, or just the video ID.
How many videos can I scrape in one run?
Up to 10. Each video is one dataset item. For more, start several runs, or loop over batches of 10 through the API.
How much does one transcript cost?
$0.02 per transcript returned. Videos with no transcript are free, and a run of 10 videos costs at most $0.20.
How do I get a transcript in a specific language only?
Set languages to that language, for example ["de"], and set
fallbackToAnyLanguage to false. Videos without German are skipped and not charged.
How do I download YouTube subtitles as an SRT file?
Set outputFormat to srt. The transcript field then holds a complete SRT file you
can save as .srt. For HTML5 players, use vtt.
Why did I get a different language than I asked for?
fallbackToAnyLanguage is on by default, so a video without your language comes back
in one it has. Check the language field, or set fallbackToAnyLanguage to false.
How can I tell whether a transcript was written by a person?
Look at isAutoGenerated. false means the creator uploaded the captions. To get only
those, set includeAutoGenerated to false.
Is scraping YouTube transcripts legal?
This scraper only collects publicly available captions. Some transcripts may still be copyrighted, so check that your use complies with YouTube's terms and the laws that apply to you. If you're unsure, ask a lawyer.
Feedback
Found a bug or need a feature? Open an issue in the Issues tab and we'll get back to you.