YouTube Transcript & Subtitle Scraper
Pricing
from $1.00 / 1,000 transcript fetcheds
YouTube Transcript & Subtitle Scraper
Extract transcripts and subtitles from YouTube videos in bulk using video, playlist, channel URLs, or keyword search. Returns timed transcript segments, plain text, SRT, and WebVTT subtitle files, with optional auto-translation to other languages.
Pricing
from $1.00 / 1,000 transcript fetcheds
Rating
0.0
(0)
Developer
Abot API
Maintained by CommunityActor stats
0
Bookmarked
63
Total users
12
Monthly active users
4 days ago
Last modified
Categories
Share
YouTube Transcript Scraper: Captions, Timestamps & Subtitles
Turn any YouTube video into readable text. Give this actor a video link, a playlist link, a channel link, or a search keyword, and it returns each video's transcript as timed segments, a single block of plain text, and ready to use SRT and WebVTT subtitle files, with optional translation into another language. Export the results to JSON, CSV or Excel, or pull them straight into your app through the API.
Why This Scraper?
- Four ways in. Mix direct video links or IDs, playlist links, and channel links (
@handle,/channel/UC...,/c/...,/user/...) in one run, or switch to search mode to find videos by keyword. - Three formats per video. Every record carries the transcript as timed
segments, a single plain texttranscriptfield, andsrtplusvttstrings you can save straight to disk. - Language control. List your preferred languages in priority order, or translate the result into another language YouTube supports.
- Handles both caption types. Works with human written captions and auto generated ones, and falls back to whatever the video offers when none of your preferred languages are available.
- Time windows on demand. Trim every transcript to a specific second range, and pay only for the trimmed text.
- Built for large pulls. Resume an interrupted playlist or channel crawl, or schedule the actor to run on a recurring basis and get only what's new, updated, or gone since last time.
- Pay only for transcripts you get. Failed, captions disabled, or incremental mode suppressed videos are never billed.
Use Cases
- Content repurposing: turn a video's spoken words into blog posts, show notes, or social captions without watching it end to end.
- Research and note taking: skim long lectures, podcasts, or interviews as searchable text instead of sitting through the video.
- Localization and accessibility: generate subtitle files or translated transcripts for viewers who need them.
- SEO and content indexing: feed video transcripts into your own search tool or content database.
- Media monitoring: track new uploads from a channel or keyword and pull their transcripts automatically on a schedule.
Data You Get
Sample shape: values are illustrative placeholders, not from a live record.
| Field | Example |
|---|---|
videoId | "EXAMPLE_ID1" |
videoUrl | "https://www.youtube.com/watch?v=EXAMPLE_ID1" |
videoTitle | "Sample Video Title" |
channelName | "Sample Channel" |
channelUrl | "https://www.youtube.com/@SampleChannel" |
durationSeconds | 213 |
source | "python tutorial" (the URL or search query that produced this video) |
sourceType | "search" (also "video", "playlist", "channel") |
language | "English" |
languageCode | "en" |
isGenerated | false (auto generated caption when true) |
isTranslated | false |
translatedTo | null |
charCount | 5234 |
segmentCount | 142 |
transcript | "Hello and welcome to this sample transcript..." |
segments | [{ "text": "Hello and welcome", "start": 0.0, "duration": 1.84 }, ...] |
srt | "1\n00:00:00,000 --> 00:00:01,840\nHello and welcome\n\n..." |
vtt | "WEBVTT\n\n00:00:00.000 --> 00:00:01.840\nHello and welcome\n\n..." |
trimmedStart / trimmedDuration | 0 / 0 (the time window applied to this video; 0 means no trim) |
success | true |
error | null (a short message when a transcript could not be fetched) |
fetchedAt | "2026-07-30T08:13:20Z" (when this row was fetched, ISO 8601 UTC) |
changeType, changedFields, firstSeenAt and lastSeenAt are added only when Incremental mode is on; a normal run's rows keep the shape above without them.
How to Use
- Pick a mode:
url(paste video, playlist, or channel links) orsearch(find videos by keyword). - Fill in the URLs or search terms for that mode, and set your preferred languages.
- Optionally trim the time window, cap how many videos to process, or turn on translation.
- Click Start, then download the dataset as JSON, CSV, or Excel, or read it through the API.
Fetch one video's transcript:
{"mode": "url","videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"languages": ["en"]}
Search YouTube by keyword and transcribe the top results:
{"mode": "search","searchQueries": ["langgraph tutorial", "apify actor development"],"maxVideosPerSource": 5,"languages": ["en"]}
Expand a playlist and a channel, capped per source, translated to English:
{"mode": "url","videoUrls": ["https://www.youtube.com/playlist?list=PLxxxxxxxxxxxxxxxx","https://www.youtube.com/@SomeChannel"],"maxVideosPerSource": 25,"maxVideos": 100,"languages": ["en"],"translateToLanguage": "en"}
Trim every transcript to a specific time window:
{"mode": "url","videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"startSec": 30,"durationSec": 30,"languages": ["en"]}
Run it from your code
Python:
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_TOKEN>")run = client.actor("abotapi/youtube-transcript-scraper").call(run_input={"mode": "url","videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"languages": ["en"],})for video in client.dataset(run["defaultDatasetId"]).iterate_items():print(video["videoTitle"], video["charCount"])
JavaScript:
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });const run = await client.actor('abotapi/youtube-transcript-scraper').call({mode: 'url',videoUrls: ['https://www.youtube.com/watch?v=dQw4w9WgXcQ'],languages: ['en'],});const { items } = await client.dataset(run.defaultDatasetId).listItems();
Or connect it to Make, Zapier, n8n, Google Sheets or webhooks from the Integrations tab.
Resume and recurring updates
- Resume (
resumeFromRunId) continues one interrupted crawl. Paste its run or dataset ID and this run skips every video already collected there, so a large playlist or channel pull picks up where it left off without re fetching, or re billing, the same videos. - Incremental mode (
incrementalMode) is for scheduling this actor on the samevideoUrls/searchQueriessetup again and again (daily, weekly) and getting only what changed. Every video is classifiedNEW,UPDATED,UNCHANGED,REAPPEARED, orEXPIRED. By default onlyNEW,UPDATED, andREAPPEAREDare returned and billed; turn onemitUnchangedoremitExpiredto also get those.stateKeynames or shares one monitoring campaign's saved state. EXPIREDis only ever produced inmode: search, and only once a run has scanned every tracked query to its end. Inmode: urlthe tracked set is whatever links you pasted, so a missing one simply wasn't pasted this run, not confirmed gone.
Send results into your apps (MCP connectors)
Optionally pipe the scraped transcripts into the apps you already use, via Model Context Protocol (MCP) connectors. This is an extra delivery step after the scrape: the Apify dataset is never changed.
What gets written to the connector: a condensed, human readable summary of each record, not the full JSON. Each item becomes one entry with a title and its key fields flattened to plain text. The complete record always stays in the Apify dataset.
- Authorize a connector once under Apify → Settings → Integrations (Notion, Linear, Airtable, or Apify).
- Select it in the "Pipe results into your apps" input field. (If the picker is empty, you haven't authorized a connector yet.)
- For Notion, also set
notionParentPageUrlto the page where items should be created.
The connection is mediated by Apify's MCP proxy, so this actor never sees your third party credentials. Leave the field empty to skip.
Input Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
mode | string | "url" | "url" (read videoUrls) or "search" (read searchQueries). The other field is ignored. |
videoUrls | array | [] | Used when mode = url. Watch URLs, youtu.be URLs, shorts URLs, playlist URLs, channel URLs (@handle, /channel/, /c/, /user/), or 11 character video IDs. Playlists and channels are expanded to their videos. |
searchQueries | array | [] | Used when mode = search. Keywords to search YouTube; top results per query, up to Max videos per source, are fetched. |
startSec | integer | 0 | Trim every transcript to start at this second. 0 keeps from the beginning of each video. |
durationSec | integer | 0 | Trim every transcript to this many seconds from startSec. 0 keeps until the end of each video. |
maxVideosPerSource | integer | 10 | How many videos to take from each playlist, channel, or search query. |
maxVideos | integer | 0 | Hard cap on the total number of videos processed across all sources. 0 means no overall cap. |
languages | array | ["en"] | Preferred transcript language codes in priority order. The first available one is used; if none are available, any transcript YouTube offers is returned. |
translateToLanguage | string | "" | Optional language code to translate the transcript into using YouTube's auto translation. Empty keeps the original language. |
preserveFormatting | boolean | false | Keep inline formatting tags (such as italics) in the transcript text instead of stripping them. |
resumeFromRunId | string | (none) | Continue one interrupted run: paste a previous run ID or dataset ID and this run skips videos already collected there. |
incrementalMode | boolean | false | Turn on for recurring monitoring of the same setup. Classifies each video as NEW, UPDATED, UNCHANGED, REAPPEARED, or EXPIRED. |
emitUnchanged | boolean | false | Incremental mode only. Also return (and bill) videos whose transcript has not changed since the last run, marked UNCHANGED. |
emitExpired | boolean | false | Incremental mode only, mode: search only. Also return videos no longer found, marked EXPIRED. Never billed. |
stateKey | string | (none) | Optional name for this monitoring campaign's saved state. Leave empty to derive one automatically from the setup. |
proxyConfiguration | object | Apify Proxy, RESIDENTIAL | Connection settings. YouTube blocks most datacenter and cloud IP ranges from fetching transcripts, so the RESIDENTIAL group is strongly recommended. |
mcpConnectors | array | (none) | Optional: send a summary of each record to apps you authorized under Integrations. |
notionParentPageUrl | string | (none) | Notion connector only: page under which records are created. |
maxNotifyListings | integer | 50 | Cap on items written to each connector per run. Does not affect the dataset. |
Output Example
Sample shape: values are illustrative placeholders, not from a live video.
{"videoId": "EXAMPLE_ID1","videoUrl": "https://www.youtube.com/watch?v=EXAMPLE_ID1","videoTitle": "Sample Video Title","channelName": "Sample Channel","channelUrl": "https://www.youtube.com/@SampleChannel","durationSeconds": 213,"source": "https://www.youtube.com/watch?v=EXAMPLE_ID1","sourceType": "video","language": "English","languageCode": "en","isGenerated": false,"isTranslated": false,"translatedTo": null,"charCount": 5234,"segmentCount": 142,"transcript": "Hello and welcome to this sample transcript.\nThis is the second line of the transcript.","segments": [{ "text": "Hello and welcome to this sample transcript.", "start": 0.0, "duration": 2.32 },{ "text": "This is the second line of the transcript.", "start": 2.32, "duration": 2.08 }],"srt": "1\n00:00:00,000 --> 00:00:02,320\nHello and welcome to this sample transcript.\n","vtt": "WEBVTT\n\n00:00:00.000 --> 00:00:02.320\nHello and welcome to this sample transcript.\n","trimmedStart": 0,"trimmedDuration": 0,"success": true,"error": null,"fetchedAt": "2026-07-30T08:13:20Z"}
With incrementalMode on, rows also carry changeType, changedFields, firstSeenAt, and lastSeenAt.
Plan Requirement
This actor needs a live connection to YouTube for every transcript fetch. The RESIDENTIAL proxy group is selected by default and is strongly recommended: YouTube blocks most datacenter and cloud IP ranges, so without it many videos come back with an error instead of a transcript. Keep RESIDENTIAL selected under Connection for reliable results.
FAQ
How much does it cost?
You pay per run, plus a small fee per keyword search that resolves, plus a length based unit per fetched transcript, roughly one unit per 4,000 characters of transcript text. Failed or captions disabled videos are never billed, and if you trim a transcript with startSec/durationSec you only pay for the trimmed text. The Pricing tab on this actor's Store page shows the current rates.
Is it legal to scrape YouTube transcripts?
This actor reads publicly available captions attached to a video, the same ones YouTube's own player can display. You are responsible for how you use the results: check YouTube's Terms of Service and the copyright rules that apply to the video's spoken content in your jurisdiction before republishing or redistributing a transcript, especially at scale or for commercial use.
What's the difference between Resume and Incremental mode?
Resume (resumeFromRunId) continues one specific interrupted run, so a large playlist or channel pull that got cut off keeps going without re fetching (and re billing) videos it already collected. Incremental mode is for scheduling the same videoUrls/searchQueries setup again and again and getting only what's new or changed since the last run. They solve different problems and are not normally combined.
Why did I get an auto generated transcript instead of the creator's captions?
The actor prefers a human written transcript over an auto generated one whenever both exist in one of your Preferred transcript languages. If the creator did not upload captions in any of your preferred languages, the actor falls back to whatever the video offers, which may be auto generated.
Why did my run fail instead of returning an empty or partial dataset?
The run only fails outright in two cases: a resumeFromRunId that cannot be resolved to a run or dataset you own, or turning on Incremental mode together with resumeFromRunId when that monitoring setup already has saved state from an earlier run. Everything else, including empty videoUrls/searchQueries, a video with disabled captions, a private or region locked video, or a temporary block from YouTube, still finishes the run and adds a dataset row explaining what happened, with success set to false, so a partial run always gives you every transcript that could be read.
Can I get only new or changed transcripts on a schedule?
Yes. Schedule the actor from the Schedules tab and turn on Incremental mode. Each run then classifies every video, and only new, updated, and reappeared rows are returned and billed by default.
Can I use it with AI agents or MCP?
Yes. Call it from any Apify integration or MCP client, and use the connector field to send a summary of each transcript to Notion, Linear, or Airtable at the end of each run.
🔗 Want more video data?
Pair this actor with these related scrapers from the same team:
| 📱 ViewStats Scraper Scrape YouTube channel analytics from ViewStats. Look up channels by handle, URL, or... | 📱 Universal Media Extractor Extract videos, audio, and metadata from 1000+ websites including YouTube, TikTok... |
| 📱 Lemon8 Media Scraper Lightweight Lemon8 scraper for extracting image and video URLs from posts, with optional... | 🎬 Twitch ALL IN ONE URL From $1/1K. Scrape Twitch data at scale, including channels, live streams, clips, VODs... |
| 📱 Tiktok Live Recorder Record TikTok live streams to MP4 with full metadata, all stream quality URLs, and... | 🤖 Suno Scraper Collect public Suno music data: songs with lyrics, style tags, model version, play and... |
💬 Support & custom scrapers
- 🐞 Found a bug or a missing field? Open a ticket on the Issues tab. We usually reply within hours.
- 🛠️ Need another site, extra fields or a private build? Email abotapi@proton.me or message Telegram @abotapi.
- ⭐ Enjoying it? A quick review on the actor page helps other users find it.