YouTube Transcript Extractor avatar

YouTube Transcript Extractor

Pricing

from $8.00 / 1,000 transcript delivereds

Go to Apify Store
YouTube Transcript Extractor

YouTube Transcript Extractor

Extract public YouTube transcripts from videos, Shorts, playlists and channels.

Pricing

from $8.00 / 1,000 transcript delivereds

Rating

0.0

(0)

Developer

Aitor Sanchez-Mansilla

Aitor Sanchez-Mansilla

Maintained by Community

Actor stats

1

Bookmarked

1

Total users

0

Monthly active users

17 hours ago

Last modified

Share

YouTube Transcript Extractor: bulk transcripts from videos, playlists and channels

YouTube Transcript Extractor status

Get the transcript of every video in a channel or playlist in one run. Paste YouTube video, Shorts, playlist or channel URLs and get back clean text, timestamped segments or SRT/VTT subtitles, one row per video, in the language the video was recorded in. Rerun the same channel later and only the new videos are fetched and charged.

No API key, no cookies, no per-video clicking. Paste URLs, get transcripts.

Why this YouTube transcript extractor

  • Whole channels and playlists, not one video at a time. A channel URL means every upload (Videos, Shorts and Live tabs); a playlist URL means every entry. Mix them with single videos in the same run. Tested on a 394-episode podcast playlist: 5 minutes, one dataset.
  • The right language, automatically. YouTube now auto-dubs many videos, and most extractors hand you whichever caption track comes first, often a machine translation. This one reads which language the video was actually recorded in and returns that transcript. Ask for English, Spanish or any of 66 languages when you want a translation instead.
  • Creator captions before auto-captions. When a video has subtitles uploaded by the creator (punctuated, accurate) and YouTube's automatic ones, you get the uploaded ones by default. Flip it if you prefer the automatic track.
  • Never pay twice for the same video. Give the job a name and rerun it on a schedule: videos already delivered under that name come back as free unchanged rows and only new uploads are fetched and charged. Keep a knowledge base current for the cost of the new videos.
  • One row per video, always. Every URL you paste and every video in a channel produces a row. When there is no transcript you get a plain reason (no_captions, private, unavailable, language_not_available…) instead of a missing line. A list of 500 always comes back as 500 rows.
  • Pay only for delivered transcripts. Videos without captions, private or removed videos, blocked attempts and already-delivered videos cost nothing. Set Max total charge on the run and it stops cleanly at your budget.
  • Clean text. HTML entities decoded, Unicode normalized, one caption per line. Timestamps in integer milliseconds. Every row carries a SHA-256 of the transcript so you can detect changes.
  • Every caption track listed. Each row includes availableTracks (language, name, uploaded or automatic), so you can see what else the video offers and ask for it.

Use cases

  • RAG and AI knowledge bases: index a creator's or a company's back catalogue, then schedule a weekly run so new videos flow in automatically.
  • Research and monitoring: pull transcripts of a playlist or a set of videos for analysis, summarization, quote checking or topic tracking.
  • Content repurposing: turn videos into blog posts, newsletters or show notes from plain text; get SRT/VTT for reuploads and translation.
  • Podcast and interview archives: full-text search across hundreds of episodes.
  • Compliance and archiving: keep a dated, hashed record of what a video's captions said.

How to extract YouTube transcripts in bulk

  1. Paste one or more URLs into YouTube URLs: videos, Shorts, playlists and channels, in any mix.
  2. Leave Transcript language on Original language of the video, or pick languages in order of preference.
  3. Click Start. Each video becomes a row in the dataset; open the Issues view to see everything that could not be delivered and why.

Minimal input, one video:

{
"startUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"]
}

A whole channel plus a playlist, English if available, otherwise the original, with subtitle files, run weekly under the name my-channel:

{
"startUrls": [
"https://www.youtube.com/@somechannel",
"https://www.youtube.com/playlist?list=PLxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
],
"preferredLanguages": ["en", "original"],
"outputFormats": ["text", "segments", "srt"],
"jobName": "my-channel"
}

Only a channel's videos from 2025 onwards, plain text only:

{
"startUrls": ["https://www.youtube.com/@somechannel/videos"],
"outputFormats": ["text"],
"publishedAfter": "2025-01-01"
}

Input

FieldWhat it doesDefault
YouTube URLs (startUrls)Video, Shorts, playlist or channel URLs, or bare video IDs. Channels accept @handle, /channel/UC…, /c/… and /user/…. Add /videos, /shorts or /streams to a channel URL to take one tab only.required
Transcript language (preferredLanguages)Languages in order of preference, chosen from a list of 66 plus original. original is the language the video was recorded in, as reported by YouTube. ["en", "original"] means English when the video has it, otherwise the original. A language also matches its regional variants (es covers es-419). Without original, videos that have none of your languages come back as language_not_available.["original"]
Which captions to prefer (captionPreference)manual_first: captions uploaded by the creator, automatic ones as fallback. auto_first: the other way round.manual_first
What to include in each row (outputFormats)Any of text (one block, best for reading and LLMs), segments (timestamped lines in milliseconds), srt and vtt (subtitle files).text, segments
Max videos per run (maxTotalVideos)Safety limit across everything you pasted. Channels and playlists are read newest first until it is reached; videos beyond it become truncated rows you can pick up in the next run.10000
Published after / before (publishedAfter, publishedBefore)Leave empty to take all videos, or keep only a date range (YYYY-MM-DD), using YouTube's exact publish date.none
Remember delivered videos under this name (jobName)A short name for the job, e.g. my-channel. Rerun with the same name and already-delivered videos are skipped for free; use a new name to start from scratch.none

Output

One dataset row per video:

{
"videoId": "dQw4w9WgXcQ",
"canonicalUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"status": "success",
"errorCode": null,
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"channelName": "Rick Astley",
"channelId": "UCuAXFkgsw1L7xaCfnd5JJOw",
"durationSeconds": 213,
"publishedAt": "2009-10-24T23:57:33-07:00",
"sources": [{ "sourceUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ", "sourceType": "video", "sourceId": "dQw4w9WgXcQ", "position": null, "tab": null }],
"languageCode": "en",
"trackKind": "manual",
"trackName": "English",
"matchedLanguage": "en",
"matchKind": "original",
"availableTracks": [
{ "languageCode": "en", "name": "English", "kind": "manual", "isTranslatable": true },
{ "languageCode": "en", "name": "English (auto-generated)", "kind": "asr", "isTranslatable": true }
],
"text": "[♪♪♪]\n♪ We're no strangers to love ♪\n♪ You know the rules and so do I ♪",
"segments": [
{ "startMs": 1360, "endMs": 3040, "durationMs": 1680, "text": "[♪♪♪]" },
{ "startMs": 18640, "endMs": 21880, "durationMs": 3240, "text": "♪ We're no strangers to love ♪" }
],
"segmentCount": 61,
"transcriptChars": 1834,
"transcriptSha256": "3c4f…",
"retrievedAt": "2026-09-23T10:15:00.000Z"
}
FieldMeaning
videoId, canonicalUrl, title, channelName, channelId, durationSeconds, publishedAtVideo identity and metadata.
statussuccess, or why not: no_captions, language_not_available, private, unavailable, blocked_suspect, invalid_url, failed, truncated, skipped, unchanged.
errorCode, errorMessageA stable code and a plain-language reason when the transcript was not delivered.
sourcesEvery URL you pasted that led to this video, with its position in the playlist or channel.
languageCode, trackKind, trackName, matchedLanguage, matchKindThe caption track delivered: manual = uploaded by the creator, asr = automatic. matchKind says whether it matched a requested language exactly, by base language, or as the original.
availableTracksEvery caption track the video offers.
text, segments, srt, vttThe transcript in the formats you asked for.
segmentCount, transcriptChars, transcriptSha256Size and a stable hash of the transcript, also filled on unchanged rows.
syncWhen you named the job: isNew, isUnchanged and the previously delivered hash.
stopReasonOn truncated rows: why the run stopped early (max_total_videos, charge_limit_reached, blocked_storm, listing_stalled).

A video without captions looks like this and is not charged:

{
"videoId": "abcdefghijk",
"status": "no_captions",
"errorCode": "NO_CAPTIONS",
"errorMessage": "The video is public but has no caption tracks (manual or automatic).",
"availableTracks": []
}

Keep a channel in sync automatically

  1. Run the channel once with a name in Remember delivered videos under this name (for example my-channel).
  2. In Apify, create a Schedule for the Actor with the same input: daily or weekly, depending on how often the channel uploads.
  3. Every scheduled run lists the channel again, returns free unchanged rows for what you already have and fetches only the new videos.

From there, plug the dataset into whatever consumes it: the Apify API or client libraries (Python, JavaScript), the Apify MCP server so an AI agent can call the Actor directly, or integrations such as Make, Zapier, n8n, Google Sheets, Airtable and webhooks that fire when a run finishes.

Use it in real time (API and MCP)

Need one transcript right now, without starting a run and reading a dataset? The Actor also answers single requests over HTTP and as an MCP server, so an app or an AI agent gets the transcript in the same response.

curl -H "Authorization: Bearer <APIFY_TOKEN>" \
"https://aitorsm--youtube-transcript-extractor.apify.actor/transcript?url=https://youtu.be/pSO7TXMVl48&language=original&format=text"

You get the same row as in the dataset. The fields you use most: status, text (or segments, srt, vtt with format=), languageCode, trackKind and availableTracks. Add format=all for every format, captionPreference=auto_first to prefer automatic captions, and repeat language= to give an order of preference. GET /tracks?url=... lists the caption languages and the original language for free.

A delivered transcript costs the same event as in a batch run (see Pricing). Errors are free and come with a clear status code: 400 for a bad request (playlists and channels included: use the batch run for those), 404 for no captions or an unavailable video, 422 when your languages do not exist (the tracks it does have are in the body), 503 with Retry-After when YouTube blocked the attempt. Real-time requests run in a Standby run on your account, so its compute time (including the idle period before it shuts down) is billed to you on top.

Add it to an AI agent. Through the Apify MCP server at https://mcp.apify.com, load the tool aitorsm/youtube-transcript-extractor, or connect straight to https://aitorsm--youtube-transcript-extractor.apify.actor/mcp (Streamable HTTP, with your Apify token as a Bearer header). Two tools are exposed: get_transcript and list_caption_tracks (free).

Pricing

You pay per delivered transcript (a row with status success) plus Apify's small Actor start fee. Every other row is free: no captions, private or removed videos, language not available, blocked or failed attempts, unchanged videos in a rerun and truncated rows.

Price per 1,000 delivered transcripts: $8 on the Free plan, $7 on Bronze, $6 on Silver, $5 on Gold and above. The price on this page is the one that applies to your account.

Max total charge on the run is your spending limit. The Actor keeps delivering while there is budget and work left, stops starting new videos when the ones in progress could exhaust it, and writes the rest as truncated rows so the next run can continue. For a big job set both Max videos per run and Max total charge high enough (30,000 videos × $0.008 = $240 on the Free plan).

Each transcript is charged at most once per run. If the confirmation of a charge is lost in transit, the Actor neither retries (that could bill you twice) nor assumes it failed; the run's SUMMARY record lists such videos under unresolvedCharges.

Limits and good to know

  • Public videos only. Private, members-only, age-verified and removed videos are reported as private or unavailable.
  • Transcripts come from the captions YouTube publishes, uploaded or automatic. A video with no caption track at all returns no_captions; audio is not transcribed with speech recognition.
  • Channels and playlists are listed newest first. Playlists larger than YouTube's own listing limit cannot be fully enumerated.
  • blocked_suspect means YouTube refused the request from every network route available to the run, even after retrying. Those rows are free; run again later. If 50 videos in a row are blocked the run stops early (blocked_storm) rather than burning through your list for nothing.
  • Very long jobs (tens of thousands of videos) take hours: give the run a long timeout, or name the job so each run continues where the previous one stopped.
  • YouTube changes its pages regularly. If a listing or caption format changes, rows report failed with a stable code rather than silently returning wrong data.

Privacy

The Actor reads only publicly available YouTube pages and caption files. It uses no Google account, cookies or API keys and keeps no copy of transcripts outside your own Apify storage. When you name a job, a small state record (video IDs, transcript hashes, delivery times) is kept in a key-value store in your account so later runs can skip delivered videos; delete that store to forget it. Transcripts may contain personal data spoken in videos; handle them according to your own obligations.

FAQ

How do I get the transcript of an entire YouTube channel?

Paste the channel URL (https://www.youtube.com/@handle) into YouTube URLs and start the run. Every video on the Videos, Shorts and Live tabs becomes a row with its transcript. Add /videos to the URL to skip Shorts and live streams, or set Published after to take only recent uploads.

Can I extract transcripts from a YouTube playlist?

Yes. Paste the playlist URL (https://www.youtube.com/playlist?list=…) and every entry is transcribed, newest first. A watch URL with a list= parameter is treated as a single video; use the playlist URL for the whole list.

Which language will the transcript be in?

By default, the language the video was recorded in, even when YouTube offers auto-dubbed versions. Choose languages in order of preference to get a translation when it exists: ["en", "original"] returns English if the video has an English track, otherwise the original. en also accepts en-US and en-GB. If you leave original out and the video has none of your languages, the row says language_not_available and lists the tracks it does have.

Does it use the creator's subtitles or YouTube's automatic captions?

Creator-uploaded captions first when they exist in your language, automatic captions otherwise. Uploaded captions are punctuated and usually more accurate; automatic ones exist for almost every video. Set Which captions to prefer to Automatic captions first to reverse the order. Each row tells you which one you got (trackKind: manual or asr).

How do I convert a YouTube transcript to text, SRT or VTT?

Pick the formats in What to include in each row. text is the transcript as plain text; segments adds timestamps in milliseconds; srt and vtt are ready-to-use subtitle files. You can request several at once and download the dataset as JSON, CSV or Excel.

How do I rerun a channel without paying for the same videos again?

Type a name in Remember delivered videos under this name (for example my-channel) and keep using it. The first run delivers everything within your limits; later runs return already-delivered videos as free unchanged rows and charge only the new ones. unchanged rows count toward Max videos per run. Use a different name to start over.

Why do I get "IP blocked" or "rate limit" errors with the YouTube transcript API, and does this Actor have the same problem?

YouTube blocks caption requests from cloud IP ranges, which is why open-source transcript libraries fail from servers. This Actor routes requests through Apify's proxy network and retries through a fresh route when YouTube pushes back. The rare video that still fails is reported as blocked_suspect, is not charged, and usually succeeds on a later run.

Why did a video come back as no_captions?

The video is public but YouTube offers no caption track for it, neither uploaded nor automatic. This happens with music-only videos, very new uploads whose automatic captions are still processing, and videos where the creator disabled captions. It is not charged.

Does it work with Shorts and live-stream recordings?

Yes. Shorts URLs work directly, and channel listings include the Videos, Shorts and Live tabs unless you add /videos, /shorts or /streams to the channel URL.

Can I use it from an AI agent or via API?

Yes. The Actor has a full input and output schema, so it works through the Apify API, the Python and JavaScript clients, and the Apify MCP server: an agent can pass URLs and read the dataset back without any glue code.

Can an AI agent call this directly?

Yes. The real-time endpoint exposes get_transcript and list_caption_tracks as MCP tools, so an agent gets the transcript inside the tool call instead of starting a run and fetching a dataset; see "Use it in real time (API and MCP)".

  • Long Tail Keyword Generator: find the questions people search on a topic, then transcribe the videos that answer them.
  • Keyword Search Volume: rank the topics your transcripts cover by real Google search volume before you write or film.
  • AI Visibility Tracker: check whether ChatGPT, Perplexity or AI Overviews mention your brand on the topics these videos cover.