YouTube Transcript: Captions & Subtitles to Text, SRT, VTT
Pricing
from $1.50 / 1,000 transcripts
YouTube Transcript: Captions & Subtitles to Text, SRT, VTT
Get the transcript of any public YouTube video with captions as plain text, timestamped segments, SRT or VTT. Paste a whole list of videos in one run. Manual or auto captions in your language. No API key, no login. Videos without captions are never charged. Works in Claude and ChatGPT via MCP.
Pricing
from $1.50 / 1,000 transcripts
Rating
0.0
(0)
Developer
The Mine Works
Maintained by CommunityActor stats
0
Bookmarked
18
Total users
5
Monthly active users
a day ago
Last modified
Categories
Share
From The Mine Works, makers of Threads Scraper and B2B Leads Finder, with over 140,000 runs across 170+ public actors.
Paste a list of YouTube videos and get each one's transcript back as plain text, timestamped segments, an SRT subtitle file or a WebVTT file, in one run. It reads the captions YouTube shows in its own player, uploaded or auto-generated, in the language you ask for. No YouTube API key, no account, no OAuth, and it works on videos you do not own.
Why choose this actor?
- You pay only for transcripts you receive. A video with no captions, a private or deleted video, a bad link, a duplicate or a request YouTube refused costs nothing, and the row says which of those happened. In our 2 October test runs, 7 videos refused by YouTube on the default proxy came back as clear rows and none of them was charged.
- A residential switch for the videos YouTube refuses. On the default datacenter proxy YouTube often answers "Sign in to confirm you're not a bot": in our runs from 26 September to 2 October, 16 of 53 videos returned a transcript there. Through Apify residential proxy, 9 of 12 did. Turn on
useResidentialProxyonly when you need it; each transcript it delivers is charged as 2 events. - Four formats at one price, no API key.
fullTextfor AI and search, timedsegments, and ready-to-save SRT and VTT files. Picking more formats does not change the price, and no Google account or quota is involved.
Part of The Mine Works Social media and video family: Threads Scraper, Reddit Scraper, Threads Search Scraper, Instagram Profile Scraper, Instagram Followers & Following, Reddit Search Scraper.
Try it in one minute
Paste this into the JSON tab of the input page and press Start. Three videos usually take well under a minute.
{"videoUrls": ["https://www.youtube.com/watch?v=UF8uR6Z6KLc","https://youtu.be/dQw4w9WgXcQ","https://www.youtube.com/watch?v=8S0FDjFBj8o"],"outputFormats": ["text", "segments"],"language": "en"}
videoUrls takes watch URLs, youtu.be short links, /shorts/, /embed/ and /live/ URLs, or bare 11-character video IDs such as dQw4w9WgXcQ, one per line. If a row comes back with "needsResidentialProxy": true, YouTube asked that request to sign in. That row is free: run the same video again with "useResidentialProxy": true.
Apify's free plan includes $5 of credit every month, which covers about 2,000 transcripts at today's Free price ($0.0025 each, no start fee), or about 990 from 12 October 2026 ($0.005 each plus $0.005 per run, in runs of 100). With useResidentialProxy on, about half as many.
Copy to your AI assistant
themineworks/youtube-transcript-scraper on Apify. Returns the transcript of public YouTube videos that have captions (uploaded or auto-generated), one row per video, as fullText, timestamped segments, SRT or VTT. Call ApifyClient("TOKEN").actor("themineworks/youtube-transcript-scraper").call(run_input={...}), then client.dataset(run["defaultDatasetId"]).list_items().items. Required: videoUrls (array of watch, youtu.be, shorts, embed or live URLs, or 11-character video IDs). Optional: outputFormats (array of text, segments, srt, vtt; default ["text","segments"]), language (caption language code, default "en"), includeTimestamps (default true), useResidentialProxy (default false; when true the actor fetches through Apify residential proxy and each delivered transcript is charged as 2 transcript-scraped events), proxy (default Apify datacenter proxy). Only rows with status "ok" are charged. Rows with needsResidentialProxy true were refused by YouTube's bot check on the datacenter proxy, are free, and can be rerun with useResidentialProxy true. Rows with _type "summary", "stopped" or "info" are run reports and are never billed. Full spec: GET https://api.apify.com/v2/acts/themineworks~youtube-transcript-scraper/builds/default (Bearer TOKEN), which returns inputSchema and readme. Token: https://console.apify.com/account/integrations?fpr=ymnoit&utm_source=apify-readme&utm_medium=referral
Key features
- Four output formats per row.
text(onefullTextstring),segments(timed lines, each{ start, dur, text }in seconds),srtandvtt, in any combination. In SRT and VTT each line ends no later than the next one starts, so YouTube's overlapping auto captions do not show as two stacked, repeating lines. - Your language first. An exact match for
language, then a regional variant (enmatchesen-USanden-GB), then another track the video offers, uploaded captions before automatic ones. Thelanguagefield on each row says which track you got, andisAutoGeneratedsays whether YouTube made it. - Two proxy modes, chosen by you. The default Apify datacenter proxy costs one event per transcript.
useResidentialProxyswitches every request of the run to Apify residential proxy and charges 2 events per transcript delivered. Nothing switches to residential on its own. - Refusals you can act on. A video YouTube refuses on the datacenter proxy comes back with
reason: "bot-check"andneedsResidentialProxy: true, free, and the run'ssummaryrow counts them invideosNeedingResidentialProxy. - Long lists without surprises. Duplicate links are fetched and charged once. After 5 request-level failures in a row, or close to the run's time limit, or when your maximum charge is spent, the run stops cleanly and lists every unprocessed video in a
stoppedrow, uncharged, ready to paste into a new run. - One proof-of-origin token per run. YouTube requires a token on caption requests; the actor creates it once per run in a lightweight simulated page (no real browser), so 512 MB of memory is enough.
How to use it
Basic: one video
{"videoUrls": ["https://www.youtube.com/watch?v=UF8uR6Z6KLc"]}
On 2 October this video (Steve Jobs' 2005 Stanford Commencement Address) returned 244 timed lines and 12,131 characters of uploaded English captions.
Several videos with subtitle files
{"videoUrls": ["https://www.youtube.com/watch?v=UF8uR6Z6KLc","https://www.youtube.com/watch?v=8S0FDjFBj8o","https://www.youtube.com/shorts/VIDEO_ID","dQw4w9WgXcQ"],"outputFormats": ["text", "srt", "vtt"]}
Each ok row carries the full SRT and VTT files as strings. Save them as .srt or .vtt for your player or editor.
Rerun the videos YouTube refused, on residential
After a default run, collect the url of every row with needsResidentialProxy: true and run them again with the switch on:
{"videoUrls": ["https://www.youtube.com/watch?v=iG9CE55wbtY","https://www.youtube.com/watch?v=UF8uR6Z6KLc"],"useResidentialProxy": true}
Both were refused on the datacenter proxy in most of our 2 October runs (iG9CE55wbtY in 3 of 4, UF8uR6Z6KLc in 4 of 4). With the switch on, UF8uR6Z6KLc came back as a full transcript, charged as 2 events. Doing it in two steps means you pay the residential price only for the videos that need it.
Transcripts for RAG and AI agents
{"videoUrls": ["https://www.youtube.com/watch?v=8S0FDjFBj8o"],"outputFormats": ["text", "segments"],"includeTimestamps": true}
Embed fullText in chunks, or embed segments so each answer can cite the second it comes from (start).
Spanish captions for a list of talks
{"videoUrls": ["https://www.youtube.com/watch?v=iG9CE55wbtY"],"language": "es","outputFormats": ["text", "vtt"]}
If a video has no Spanish track you get another one it offers, and language on the row tells you which. The actor never translates.
New videos every week
Keep a saved task with the week's video links and run it on a schedule (Console, Schedules, Create new, pick a time such as 0 6 * * 1). Scheduled runs are billed exactly like manual ones, per transcript delivered.
Input parameters
| Parameter | Type | Default | What it does |
|---|---|---|---|
videoUrls | array of strings | one example video | Videos to transcribe, one per line: watch URLs, youtu.be links, /shorts/, /embed/ and /live/ URLs, or bare 11-character IDs. Duplicates are fetched and charged once. Channel and playlist URLs are refused (free) |
outputFormats | array | ["text", "segments"] | Any of text, segments, srt, vtt. Same price whatever you pick |
language | string | en | Caption language to prefer, such as en, es, hi, fr or pt. Falls back to another track the video offers, never translates |
includeTimestamps | boolean | true | Keeps the segments array when segments is selected. SRT and VTT always carry their own timings |
useResidentialProxy | boolean | false | When it is on, the actor fetches through Apify RESIDENTIAL proxy and each delivered transcript is charged as TWO transcript-scraped events. That is about $10 per 1,000 transcripts on the Free plan from 12 October 2026, and $5 per 1,000 now. Videos without a transcript are still free. The proxy field is ignored while it is on (a country set there is kept) |
proxy | object | Apify Proxy (datacenter) | Leave it on the default. Videos refused there are marked needsResidentialProxy; use useResidentialProxy for them. A residential group picked here instead is charged at the single price, and the cost guard stops such a run after about 100 videos |
Calling it from code with input written for another tool? If videoUrls is empty, the actor also reads urls, videoIds, startUrls (strings or { "url": ... } objects), youtube_url and videoUrl.
Run time. A run has 300 seconds by default. On the datacenter proxy most videos take a few seconds each. Our 11 video residential run took 206 seconds (6 to 51 seconds per video, plus the one-time token setup), so for more than about 15 videos with useResidentialProxy on, raise the timeout under Run options. If a run gets close to its limit, it stops about 15 seconds early and lists the videos it did not reach, uncharged.
What data do you get?
One row per video you listed, plus one summary row at the end of every run.
Transcript rows (status: "ok", charged): videoId, url, title, language, isAutoGenerated, segmentCount, charCount, and the formats you asked for: fullText, segments, srt, vtt.
Rows without a transcript (free): status is no-caption-track or error, reason says why, and error explains it in a sentence. Reasons you will see:
bot-check: YouTube asked the request to sign in to confirm it is not a bot. On the datacenter proxy the row also hasneedsResidentialProxy: true.no-tracks-listed: no caption track came back. Either the video has none, or YouTube withheld them from this request; one response cannot tell which, and the row says so.requested-language-unavailable: the video has tracks but none usable;availableLanguageslists what it has.video-unavailable: private, deleted, age-restricted or members-only.invalid-video-id: a channel, playlist or other link with no video ID.caption-fetch-failed,caption-track-empty,watch-page-http-error,watch-page-fetch-failed,no-player-data: the request failed; usually worth a retry later.
Run rows (never charged): a stopped row (_type: "stopped") when the run ended early, with reason (transport-failing, run-timeout, charge-limit-reached, cost-limit or residential-proxy-unavailable), message, videosNotAttempted and notAttempted; a summary row (_type: "summary") with videosRequested, transcriptsScraped, noCaptionTrack, errored, duplicatesSkipped, chargedFor, chargedEvents, useResidentialProxy, eventsPerTranscript and videosNeedingResidentialProxy; and an info row when transcripts were delivered.
Fields with no value are left out rather than sent as null. The dataset has two views: Transcripts (one line per video) and Full text.
Stable fields for automations
Present on every ok row today (url, status and scrapedAt are on every video row of any kind):
| Field | Meaning |
|---|---|
videoId | 11-character YouTube video ID |
url | Canonical watch URL |
title | Video title |
language | Language code of the caption track used |
isAutoGenerated | true for YouTube's automatic captions |
segmentCount | Number of timed lines |
charCount | Length of the full transcript text |
status | ok on a charged transcript |
scrapedAt | ISO time the row was written |
fullText | Whole transcript, when text is selected (default) |
segments | Timed lines, when segments is selected (default) |
These names will not change. New fields may be added next to them.
Output examples
A transcript row from a 2 October run with useResidentialProxy on (text and segments, trimmed):
{"videoId": "UF8uR6Z6KLc","url": "https://www.youtube.com/watch?v=UF8uR6Z6KLc","title": "Steve Jobs' 2005 Stanford Commencement Address","language": "en","isAutoGenerated": false,"segmentCount": 244,"charCount": 12131,"fullText": "This program is brought to you by Stanford University. Please visit us at stanford.edu ...","segments": [{ "start": 7.47, "dur": 2.899, "text": "This program is brought to you by Stanford University." },{ "start": 10.47, "dur": 3.973, "text": "Please visit us at stanford.edu" }],"status": "ok","scrapedAt": "2026-10-02T11:22:02.978Z"}
With srt and vtt selected, the same two lines arrive as "srt": "1\n00:00:07,470 --> 00:00:10,369\nThis program is brought to you by Stanford University.\n\n2\n00:00:10,470 --> 00:00:14,443\nPlease visit us at stanford.edu\n..." and "vtt": "WEBVTT\n\n00:00:07.470 --> 00:00:10.369\nThis program is brought to you by Stanford University.\n\n...".
The same video on the default datacenter proxy earlier that day, refused by YouTube and not charged:
{"videoId": "UF8uR6Z6KLc","url": "https://www.youtube.com/watch?v=UF8uR6Z6KLc","status": "error","reason": "bot-check","needsResidentialProxy": true,"error": "YouTube asked this request to sign in to confirm it is not a bot (Sign in to confirm you’re not a bot). This happens to many videos on the default datacenter proxy and says nothing about the video itself. Not charged. To get this transcript, run the video again with useResidentialProxy turned on: the actor then fetches through Apify residential proxy and each delivered transcript is charged as 2 transcript-scraped events.","scrapedAt": "2026-10-02T11:21:00.091Z"}
A playlist link, refused up front and free:
{"url": "https://www.youtube.com/playlist?list=PLx0sYbCqOb8TBPRdmBHs5Iftvv9TPboYG","status": "error","reason": "invalid-video-id","error": "Could not find a YouTube video ID in this input. Pass a video URL (watch, youtu.be, shorts, embed or live) or an 11-character video ID. Channel and playlist URLs are not supported.","scrapedAt": "2026-10-02T11:21:00.647Z"}
The summary row of the residential run (one transcript, 2 events; trimmed):
{"_type": "summary","status": "summary","videosRequested": 1,"duplicatesSkipped": 0,"transcriptsScraped": 1,"noCaptionTrack": 0,"errored": 0,"videosNotAttempted": 0,"chargedFor": 1,"chargedEvents": 2,"useResidentialProxy": true,"eventsPerTranscript": 2,"videosNeedingResidentialProxy": 0,"outputFormats": ["text", "segments"],"scrapedAt": "2026-10-02T11:22:02.987Z"}
Pricing
Pay per event: one transcript-scraped event for each transcript saved to your dataset, two with useResidentialProxy on. Prices change on 12 October 2026; both are shown.
Until 11 October 2026 (current price, no start fee)
| Per transcript | Free | Bronze (Starter) | Silver (Scale) | Gold and above (Business) |
|---|---|---|---|---|
| Default proxy, 1 event | $0.0025 | $0.002 | $0.00175 | $0.0015 |
| Per 1,000, default proxy | $2.50 | $2.00 | $1.75 | $1.50 |
useResidentialProxy on, 2 events | $0.005 | $0.004 | $0.0035 | $0.003 |
| Per 1,000, residential | $5.00 | $4.00 | $3.50 | $3.00 |
From 12 October 2026 (already scheduled)
| Per transcript | Free | Bronze (Starter) | Silver (Scale) | Gold and above (Business) |
|---|---|---|---|---|
| Default proxy, 1 event | $0.005 | $0.0042 | $0.0036 | $0.003 |
| Per 1,000, default proxy | $5.00 | $4.20 | $3.60 | $3.00 |
useResidentialProxy on, 2 events | $0.01 | $0.0084 | $0.0072 | $0.006 |
| Per 1,000, residential | $10.00 | $8.40 | $7.20 | $6.00 |
Start fee, apify-actor-start | $0.005 per GB of run memory, minimum one event | same | same | same |
The start fee, exactly. From 12 October, Apify's apify-actor-start event is charged once when a run starts, at $0.005 for each GB of memory, with a minimum of one event. This actor runs on 512 MB by default, so a default run pays one event: $0.005. Before 12 October there is no start fee.
Worked examples on the Free plan. 100 transcripts on the default proxy: $0.25 today, $0.505 from 12 October. The same 100 with useResidentialProxy on: $0.50 today, $1.005 from 12 October.
Never charged: videos without a caption track, private or unavailable videos, bot-check rows, failed requests, invalid links, duplicate links in the same run, videos a stopped run did not reach, and the summary, stopped and info rows. The charge fires only after an ok row is saved. With useResidentialProxy on, a video is started only if the maximum charge you set for the run still covers its 2 events.
The Pricing tab on this page always shows the rate for your own plan; if it and these tables ever differ, the Pricing tab is right.
FAQ
Do I need a YouTube API key or a Google account? No. The actor reads the public caption tracks YouTube serves to every viewer of a video. There is no API key, OAuth, login, cookie or daily quota, and it never signs in to YouTube. YouTube's own Data API only lets you download captions for videos on channels you own.
Why do some rows say "Sign in to confirm you're not a bot"?
YouTube shows that check to many requests that come from datacenter IP addresses, video by video. In our runs from 26 September to 2 October, 16 of 53 videos returned a transcript on the default datacenter proxy. Those rows have reason: "bot-check" and needsResidentialProxy: true, and they are never charged.
When should I turn on useResidentialProxy?
When the videos you need came back with needsResidentialProxy: true, or when you would rather pay more per transcript than rerun. Residential traffic costs us about as much per video as the whole default price, so with the switch on each delivered transcript is charged as 2 events: $5 per 1,000 on the Free plan now and about $10 per 1,000 from 12 October 2026. In residential runs, 9 of 12 videos returned a transcript; the others failed at the request level and were free.
How much do 1,000 transcripts cost?
On the default proxy: $2.50 on the Free plan and $1.50 on Gold and above today; from 12 October 2026, $5.00 and $3.00 plus $0.005 per run. Double the per transcript price with useResidentialProxy on. Only ok rows count.
Am I charged for videos that have no captions?
No. The actor reads existing captions and does not transcribe audio, so a video with no caption track returns no-caption-track and is free. The same goes for private, deleted and age-restricted videos, bad links, duplicates and refused requests.
How many videos can I put in one run?
There is no fixed cap on the list. A run has 300 seconds by default, which suits a few dozen videos on the default proxy and about 15 on residential; raise the timeout under Run options for longer lists. A run that nears its limit stops cleanly and lists the rest in a stopped row, uncharged.
How fresh are the transcripts? Each run fetches the captions live from YouTube at that moment, so an edited or newly added caption track shows up on the next run. Nothing is cached between runs.
Can I choose the caption language?
Yes, with language (en, es, hi, de, ja and so on). You get an exact match first, then a regional variant, then another available track, and the row's language field says which. It does not translate.
Does it work with Shorts and live stream replays?
Yes, if the video has captions. /shorts/ and /live/ URLs are read like watch URLs. Many Shorts have no captions at all; those come back as free no-caption-track rows. To transcribe the audio of Shorts without captions, use our Video Transcriber.
Can I get transcripts for a whole channel or playlist?
Not directly. Channel and playlist URLs are refused with invalid-video-id, free. List the channel's videos with our YouTube Channel Scraper, then paste the video URLs here.
What export formats are there?
From the run's Storage tab as JSON, CSV, Excel, XML or HTML, or through the Apify API. Skip rows that have a _type field if you want transcripts only.
Can I run it on a schedule? Yes. Save your input as a task and add it under Schedules in Apify Console. Scheduled runs are billed exactly like manual ones, and the schedule itself is free.
Can I use it from Claude, ChatGPT or another AI assistant?
- Connector URL:
https://mcp.apify.com/?tools=themineworks/youtube-transcript-scraper. - Claude: Settings > Connectors > Add custom connector, paste the URL, sign in with Apify.
- ChatGPT: developer mode, add an MCP connector with the URL, sign in with Apify.
- Cursor or VS Code: add it as an HTTP MCP server with that URL.
- Claude Code:
claude mcp add -t http youtube-transcript-scraper "https://mcp.apify.com/?tools=themineworks/youtube-transcript-scraper".
Is it legal to use this data? The actor collects only captions and titles that YouTube shows to anyone watching a public video, and it never signs in. Transcripts are usually the creator's copyrighted work. You are responsible for how you use them, including YouTube's terms, copyright and data protection laws such as GDPR and CCPA. This is general information, not legal advice.
Integrations
- Google Sheets: export a run to a sheet, one row per video, with
fullTextin a column. - Make, Zapier and n8n: use the Apify app or node to start a run when a new video link arrives and send the transcript on to a document, a CRM or Slack.
- Webhooks: have Apify call your URL when a run succeeds, then fetch the dataset.
- API and client libraries: start runs and read datasets from Python, JavaScript or any HTTP client. The "Copy to your AI assistant" block above has the exact call.
- MCP clients: Claude, ChatGPT, Cursor, VS Code and Claude Code can call the actor through
https://mcp.apify.com. - Next step: get a channel's video list with our YouTube Channel Scraper, or episode transcripts from podcast feeds with our Podcast Episode Scraper.
More from The Mine Works
Social media and video
- Threads Scraper
- Reddit Scraper
- Threads Search Scraper
- Instagram Profile Scraper
- Instagram Followers & Following
- Reddit Search Scraper
- Twitter / X Scraper
- Xiaohongshu (RED) Scraper
- Telegram Channel Scraper
- Telegram Channel Finder
- Pinterest Profile Scraper
- TikTok Scraper
Leads and business directories
Marketing, SEO and reviews
Real estate
Science, health and government data
Jobs and hiring
Company and business data
E-commerce and marketplaces
Food and local services
Developer and AI tools
More tools
Support
Found a bug or need a field we do not return yet? Open an issue on the Issues tab of this page and we will reply there. To ask for a new data source, email dmineworks@gmail.com. A guide and FAQ for this actor also live at themineworks.com.
YouTube Transcript turns a list of public YouTube videos into text, timed segments, SRT and VTT, charges only for transcripts delivered, and fetches the videos YouTube refuses on the default proxy through a residential switch you control.

