YouTube Transcript & Subtitle Scraper avatar

YouTube Transcript & Subtitle Scraper

Pricing

from $1.00 / 1,000 transcript fetcheds

Go to Apify Store
YouTube Transcript & Subtitle Scraper

YouTube Transcript & Subtitle Scraper

Extract transcripts and subtitles from YouTube videos in bulk using video, playlist, channel URLs, or keyword search. Returns timed transcript segments, plain text, SRT, and WebVTT subtitle files, with optional auto-translation to other languages.

Pricing

from $1.00 / 1,000 transcript fetcheds

Rating

0.0

(0)

Developer

Abot API

Abot API

Maintained by Community

Actor stats

0

Bookmarked

63

Total users

12

Monthly active users

4 days ago

Last modified

Share

YouTube Transcript Scraper: Captions, Timestamps & Subtitles

Turn any YouTube video into readable text. Give this actor a video link, a playlist link, a channel link, or a search keyword, and it returns each video's transcript as timed segments, a single block of plain text, and ready to use SRT and WebVTT subtitle files, with optional translation into another language. Export the results to JSON, CSV or Excel, or pull them straight into your app through the API.

Why This Scraper?

  • Four ways in. Mix direct video links or IDs, playlist links, and channel links (@handle, /channel/UC..., /c/..., /user/...) in one run, or switch to search mode to find videos by keyword.
  • Three formats per video. Every record carries the transcript as timed segments, a single plain text transcript field, and srt plus vtt strings you can save straight to disk.
  • Language control. List your preferred languages in priority order, or translate the result into another language YouTube supports.
  • Handles both caption types. Works with human written captions and auto generated ones, and falls back to whatever the video offers when none of your preferred languages are available.
  • Time windows on demand. Trim every transcript to a specific second range, and pay only for the trimmed text.
  • Built for large pulls. Resume an interrupted playlist or channel crawl, or schedule the actor to run on a recurring basis and get only what's new, updated, or gone since last time.
  • Pay only for transcripts you get. Failed, captions disabled, or incremental mode suppressed videos are never billed.

Use Cases

  • Content repurposing: turn a video's spoken words into blog posts, show notes, or social captions without watching it end to end.
  • Research and note taking: skim long lectures, podcasts, or interviews as searchable text instead of sitting through the video.
  • Localization and accessibility: generate subtitle files or translated transcripts for viewers who need them.
  • SEO and content indexing: feed video transcripts into your own search tool or content database.
  • Media monitoring: track new uploads from a channel or keyword and pull their transcripts automatically on a schedule.

Data You Get

Sample shape: values are illustrative placeholders, not from a live record.

FieldExample
videoId"EXAMPLE_ID1"
videoUrl"https://www.youtube.com/watch?v=EXAMPLE_ID1"
videoTitle"Sample Video Title"
channelName"Sample Channel"
channelUrl"https://www.youtube.com/@SampleChannel"
durationSeconds213
source"python tutorial" (the URL or search query that produced this video)
sourceType"search" (also "video", "playlist", "channel")
language"English"
languageCode"en"
isGeneratedfalse (auto generated caption when true)
isTranslatedfalse
translatedTonull
charCount5234
segmentCount142
transcript"Hello and welcome to this sample transcript..."
segments[{ "text": "Hello and welcome", "start": 0.0, "duration": 1.84 }, ...]
srt"1\n00:00:00,000 --> 00:00:01,840\nHello and welcome\n\n..."
vtt"WEBVTT\n\n00:00:00.000 --> 00:00:01.840\nHello and welcome\n\n..."
trimmedStart / trimmedDuration0 / 0 (the time window applied to this video; 0 means no trim)
successtrue
errornull (a short message when a transcript could not be fetched)
fetchedAt"2026-07-30T08:13:20Z" (when this row was fetched, ISO 8601 UTC)

changeType, changedFields, firstSeenAt and lastSeenAt are added only when Incremental mode is on; a normal run's rows keep the shape above without them.

How to Use

  1. Pick a mode: url (paste video, playlist, or channel links) or search (find videos by keyword).
  2. Fill in the URLs or search terms for that mode, and set your preferred languages.
  3. Optionally trim the time window, cap how many videos to process, or turn on translation.
  4. Click Start, then download the dataset as JSON, CSV, or Excel, or read it through the API.

Fetch one video's transcript:

{
"mode": "url",
"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"languages": ["en"]
}

Search YouTube by keyword and transcribe the top results:

{
"mode": "search",
"searchQueries": ["langgraph tutorial", "apify actor development"],
"maxVideosPerSource": 5,
"languages": ["en"]
}

Expand a playlist and a channel, capped per source, translated to English:

{
"mode": "url",
"videoUrls": [
"https://www.youtube.com/playlist?list=PLxxxxxxxxxxxxxxxx",
"https://www.youtube.com/@SomeChannel"
],
"maxVideosPerSource": 25,
"maxVideos": 100,
"languages": ["en"],
"translateToLanguage": "en"
}

Trim every transcript to a specific time window:

{
"mode": "url",
"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"startSec": 30,
"durationSec": 30,
"languages": ["en"]
}

Run it from your code

Python:

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("abotapi/youtube-transcript-scraper").call(run_input={
"mode": "url",
"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"languages": ["en"],
})
for video in client.dataset(run["defaultDatasetId"]).iterate_items():
print(video["videoTitle"], video["charCount"])

JavaScript:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('abotapi/youtube-transcript-scraper').call({
mode: 'url',
videoUrls: ['https://www.youtube.com/watch?v=dQw4w9WgXcQ'],
languages: ['en'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();

Or connect it to Make, Zapier, n8n, Google Sheets or webhooks from the Integrations tab.

Resume and recurring updates

  • Resume (resumeFromRunId) continues one interrupted crawl. Paste its run or dataset ID and this run skips every video already collected there, so a large playlist or channel pull picks up where it left off without re fetching, or re billing, the same videos.
  • Incremental mode (incrementalMode) is for scheduling this actor on the same videoUrls/searchQueries setup again and again (daily, weekly) and getting only what changed. Every video is classified NEW, UPDATED, UNCHANGED, REAPPEARED, or EXPIRED. By default only NEW, UPDATED, and REAPPEARED are returned and billed; turn on emitUnchanged or emitExpired to also get those. stateKey names or shares one monitoring campaign's saved state.
  • EXPIRED is only ever produced in mode: search, and only once a run has scanned every tracked query to its end. In mode: url the tracked set is whatever links you pasted, so a missing one simply wasn't pasted this run, not confirmed gone.

Send results into your apps (MCP connectors)

Optionally pipe the scraped transcripts into the apps you already use, via Model Context Protocol (MCP) connectors. This is an extra delivery step after the scrape: the Apify dataset is never changed.

What gets written to the connector: a condensed, human readable summary of each record, not the full JSON. Each item becomes one entry with a title and its key fields flattened to plain text. The complete record always stays in the Apify dataset.

  1. Authorize a connector once under Apify → Settings → Integrations (Notion, Linear, Airtable, or Apify).
  2. Select it in the "Pipe results into your apps" input field. (If the picker is empty, you haven't authorized a connector yet.)
  3. For Notion, also set notionParentPageUrl to the page where items should be created.

The connection is mediated by Apify's MCP proxy, so this actor never sees your third party credentials. Leave the field empty to skip.

Input Parameters

ParameterTypeDefaultDescription
modestring"url""url" (read videoUrls) or "search" (read searchQueries). The other field is ignored.
videoUrlsarray[]Used when mode = url. Watch URLs, youtu.be URLs, shorts URLs, playlist URLs, channel URLs (@handle, /channel/, /c/, /user/), or 11 character video IDs. Playlists and channels are expanded to their videos.
searchQueriesarray[]Used when mode = search. Keywords to search YouTube; top results per query, up to Max videos per source, are fetched.
startSecinteger0Trim every transcript to start at this second. 0 keeps from the beginning of each video.
durationSecinteger0Trim every transcript to this many seconds from startSec. 0 keeps until the end of each video.
maxVideosPerSourceinteger10How many videos to take from each playlist, channel, or search query.
maxVideosinteger0Hard cap on the total number of videos processed across all sources. 0 means no overall cap.
languagesarray["en"]Preferred transcript language codes in priority order. The first available one is used; if none are available, any transcript YouTube offers is returned.
translateToLanguagestring""Optional language code to translate the transcript into using YouTube's auto translation. Empty keeps the original language.
preserveFormattingbooleanfalseKeep inline formatting tags (such as italics) in the transcript text instead of stripping them.
resumeFromRunIdstring(none)Continue one interrupted run: paste a previous run ID or dataset ID and this run skips videos already collected there.
incrementalModebooleanfalseTurn on for recurring monitoring of the same setup. Classifies each video as NEW, UPDATED, UNCHANGED, REAPPEARED, or EXPIRED.
emitUnchangedbooleanfalseIncremental mode only. Also return (and bill) videos whose transcript has not changed since the last run, marked UNCHANGED.
emitExpiredbooleanfalseIncremental mode only, mode: search only. Also return videos no longer found, marked EXPIRED. Never billed.
stateKeystring(none)Optional name for this monitoring campaign's saved state. Leave empty to derive one automatically from the setup.
proxyConfigurationobjectApify Proxy, RESIDENTIALConnection settings. YouTube blocks most datacenter and cloud IP ranges from fetching transcripts, so the RESIDENTIAL group is strongly recommended.
mcpConnectorsarray(none)Optional: send a summary of each record to apps you authorized under Integrations.
notionParentPageUrlstring(none)Notion connector only: page under which records are created.
maxNotifyListingsinteger50Cap on items written to each connector per run. Does not affect the dataset.

Output Example

Sample shape: values are illustrative placeholders, not from a live video.

{
"videoId": "EXAMPLE_ID1",
"videoUrl": "https://www.youtube.com/watch?v=EXAMPLE_ID1",
"videoTitle": "Sample Video Title",
"channelName": "Sample Channel",
"channelUrl": "https://www.youtube.com/@SampleChannel",
"durationSeconds": 213,
"source": "https://www.youtube.com/watch?v=EXAMPLE_ID1",
"sourceType": "video",
"language": "English",
"languageCode": "en",
"isGenerated": false,
"isTranslated": false,
"translatedTo": null,
"charCount": 5234,
"segmentCount": 142,
"transcript": "Hello and welcome to this sample transcript.\nThis is the second line of the transcript.",
"segments": [
{ "text": "Hello and welcome to this sample transcript.", "start": 0.0, "duration": 2.32 },
{ "text": "This is the second line of the transcript.", "start": 2.32, "duration": 2.08 }
],
"srt": "1\n00:00:00,000 --> 00:00:02,320\nHello and welcome to this sample transcript.\n",
"vtt": "WEBVTT\n\n00:00:00.000 --> 00:00:02.320\nHello and welcome to this sample transcript.\n",
"trimmedStart": 0,
"trimmedDuration": 0,
"success": true,
"error": null,
"fetchedAt": "2026-07-30T08:13:20Z"
}

With incrementalMode on, rows also carry changeType, changedFields, firstSeenAt, and lastSeenAt.

Plan Requirement

This actor needs a live connection to YouTube for every transcript fetch. The RESIDENTIAL proxy group is selected by default and is strongly recommended: YouTube blocks most datacenter and cloud IP ranges, so without it many videos come back with an error instead of a transcript. Keep RESIDENTIAL selected under Connection for reliable results.

FAQ

How much does it cost?

You pay per run, plus a small fee per keyword search that resolves, plus a length based unit per fetched transcript, roughly one unit per 4,000 characters of transcript text. Failed or captions disabled videos are never billed, and if you trim a transcript with startSec/durationSec you only pay for the trimmed text. The Pricing tab on this actor's Store page shows the current rates.

This actor reads publicly available captions attached to a video, the same ones YouTube's own player can display. You are responsible for how you use the results: check YouTube's Terms of Service and the copyright rules that apply to the video's spoken content in your jurisdiction before republishing or redistributing a transcript, especially at scale or for commercial use.

What's the difference between Resume and Incremental mode?

Resume (resumeFromRunId) continues one specific interrupted run, so a large playlist or channel pull that got cut off keeps going without re fetching (and re billing) videos it already collected. Incremental mode is for scheduling the same videoUrls/searchQueries setup again and again and getting only what's new or changed since the last run. They solve different problems and are not normally combined.

Why did I get an auto generated transcript instead of the creator's captions?

The actor prefers a human written transcript over an auto generated one whenever both exist in one of your Preferred transcript languages. If the creator did not upload captions in any of your preferred languages, the actor falls back to whatever the video offers, which may be auto generated.

Why did my run fail instead of returning an empty or partial dataset?

The run only fails outright in two cases: a resumeFromRunId that cannot be resolved to a run or dataset you own, or turning on Incremental mode together with resumeFromRunId when that monitoring setup already has saved state from an earlier run. Everything else, including empty videoUrls/searchQueries, a video with disabled captions, a private or region locked video, or a temporary block from YouTube, still finishes the run and adds a dataset row explaining what happened, with success set to false, so a partial run always gives you every transcript that could be read.

Can I get only new or changed transcripts on a schedule?

Yes. Schedule the actor from the Schedules tab and turn on Incremental mode. Each run then classifies every video, and only new, updated, and reappeared rows are returned and billed by default.

Can I use it with AI agents or MCP?

Yes. Call it from any Apify integration or MCP client, and use the connector field to send a summary of each transcript to Notion, Linear, or Airtable at the end of each run.

🔗 Want more video data?

Pair this actor with these related scrapers from the same team:

📱 ViewStats Scraper
Scrape YouTube channel analytics from ViewStats. Look up channels by handle, URL, or...
📱 Universal Media Extractor
Extract videos, audio, and metadata from 1000+ websites including YouTube, TikTok...
📱 Lemon8 Media Scraper
Lightweight Lemon8 scraper for extracting image and video URLs from posts, with optional...
🎬 Twitch ALL IN ONE URL
From $1/1K. Scrape Twitch data at scale, including channels, live streams, clips, VODs...
📱 Tiktok Live Recorder
Record TikTok live streams to MP4 with full metadata, all stream quality URLs, and...
🤖 Suno Scraper
Collect public Suno music data: songs with lyrics, style tags, model version, play and...

👉 Browse all abotapi scrapers

💬 Support & custom scrapers

  • 🐞 Found a bug or a missing field? Open a ticket on the Issues tab. We usually reply within hours.
  • 🛠️ Need another site, extra fields or a private build? Email abotapi@proton.me or message Telegram @abotapi.
  • ⭐ Enjoying it? A quick review on the actor page helps other users find it.