YouTube Transcript Scraper - Channels & Playlists avatar

YouTube Transcript Scraper - Channels & Playlists

Pricing

Pay per event

Go to Apify Store
YouTube Transcript Scraper - Channels & Playlists

YouTube Transcript Scraper - Channels & Playlists

Bulk YouTube transcripts from videos, whole channels and playlists: publish-date window, language priority with fallback, uploaded vs auto captions, video metadata and only-new-since-last-run.

Pricing

Pay per event

Rating

0.0

(0)

Developer

datagrit

datagrit

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

11 hours ago

Last modified

Share

What does YouTube Transcript Scraper - Channels & Playlists do?

It pulls YouTube transcripts in bulk: paste any mix of single videos, whole channels and playlists, and get one row per video with the full transcript text, timestamped caption lines and the video's metadata (title, channel, exact publish date, duration, views, likes, category). It is built for people who need many transcripts at once: researchers, content and SEO teams, and anyone feeding YouTube content into an LLM, a RAG index or a search tool.

Why use it?

  • Whole channels and playlists in one run. Channel links (@handle, /channel/UC..., /c/..., /user/...), playlist links and video links can be mixed in one list. Set how many videos to take per channel or playlist; there is no fixed cap per run.
  • Publish-date window. Keep only videos published after and/or before a date, or within a period such as "7 days" or "3 months". Every video is checked against its exact publish date, and channel listing stops once it reaches older videos.
  • Language control. Give caption languages in order of preference (en, es, pt-BR...), choose uploaded or auto-generated captions first, or accept only one kind. Turn the language fallback on or off.
  • Only new uploads. For scheduled runs, the Actor remembers per channel and per settings which videos it already read and stops the listing at the first of them, so a daily run returns only that day's uploads, never the older backlog.
  • Honest rows. A video without a usable transcript gets a free row with the reason (no captions, other languages only, unavailable, blocked, age-restricted, upcoming), so you always know what happened to every video.
  • Channel content choice. Long-form videos, Shorts, live stream replays or all uploads.

Example output

titlechannelNamepublishedAtdurationSecondslanguagecaptionTypewordCount
The Right to Life, Liberty — and Free Time | Danielle Roberts | TEDTED2026-09-29T15:00:05.000Z233enmanual547
{
"found": true,
"reason": null,
"videoId": "1qmF_znXxrE",
"title": "The Right to Life, Liberty — and Free Time | Danielle Roberts | TED",
"channelName": "TED",
"channelId": "UCAuUUnT6oDeKwE6v1NGQxug",
"publishedAt": "2026-09-29T15:00:05.000Z",
"durationSeconds": 233,
"viewCount": 59691,
"likeCount": 568,
"category": "People & Blogs",
"language": "en",
"languageName": "English",
"captionType": "manual",
"isLanguageFallback": false,
"availableLanguages": ["ar", "my", "zh-CN", "zh-TW", "en", "fr", "iw", "hi", "pt-BR", "sr", "es", "vi"],
"transcriptText": "So it's exciting and ironic that I'm giving a TED Talk about time in three minutes. ...",
"wordCount": 547,
"segmentCount": 73,
"segments": [{ "start": 4.292, "duration": 2.711, "text": "So it's exciting and ironic" }],
"input": "https://www.youtube.com/@TED",
"sourceType": "channel",
"sourceTitle": "TED",
"sourceUrl": "https://www.youtube.com/watch?v=1qmF_znXxrE",
"scrapedAt": "2026-10-01T08:00:00.000Z"
}

A video without a transcript in the requested language (here French, with fallback off) gets a free row:

{
"found": false,
"reason": "languageNotAvailable",
"message": "The video has no captions in the requested languages and language fallback is off.",
"videoId": "9Ff4S1FJdRM",
"title": "Is AI Turning Us All into the Same Person? | Sandra Matz | TED",
"publishedAt": "2026-09-28T15:00:04.000Z",
"availableLanguages": ["ar", "my", "en", "iw", "pt-BR", "es"],
"transcriptText": null
}

How much does it cost?

You pay per transcript delivered. Pricing depends on your Apify plan: a small fee when a run starts, then a price per transcript that is lower on paid plans. The Apify free plan includes monthly credit you can use to try it. Rows for videos without a transcript, duplicates and the status row are never charged. You can set a maximum spend on the run: the Actor stops when the limit is reached. It reads YouTube's public web endpoints over plain HTTP (no browser, no video download), so runs are fast.

Input

  • Videos, channels or playlists – one link or ID per line. A channel link ending in /shorts or /streams reads that tab. A watch link with &list= is read as that single video; use the /playlist?list= link for the whole playlist.
  • Maximum transcripts – total limit for the run (only delivered transcripts count).
  • Maximum videos per channel or playlist – channels newest first; 0 means no per-source limit.
  • Channel content – long-form videos, Shorts, live stream replays or all uploads.
  • Published on or after / Published before – a date (2026-09-01) or a period back from the run (7 days, 2 weeks, 3 months).
  • Title keywords – keep only videos whose title contains one of the words.
  • Caption languages, Caption type, Fall back to another language – which caption track to return.
  • Include timestamped segments – on by default; the plain text is always included.
  • Only videos new since my last run and Only-new memory name – for schedules.
  • Proxy configuration – Apify Proxy (datacenter) by default.

How it works

For every channel the Actor reads the channel's uploads list page by page (100 videos per page), newest first. For every video it reads the public video details (publish date, metadata) and the caption track list, picks the track that matches your languages and caption type, and downloads that caption file. That is three small requests per video, plus one listing request per 100 videos of a channel or playlist.

The Actor reads only information that YouTube shows to any visitor without signing in: public video pages, captions and channel lists. It does not log in, does not use cookies of an account, does not solve CAPTCHAs and does not download video files. Transcripts are the work of their creators; how you use them (for example quoting, analysis, or training) is your responsibility under copyright law and YouTube's terms. This description is not legal advice.

FAQ

Why do some videos have no transcript? Many videos have no captions at all (music, very new uploads, some Shorts). Such videos get a free row with reason noCaptions. The status message counts every reason.

Why does a run end without transcripts but green? When none of the videos has captions in your languages (with fallback off) or of the caption type you asked for, every video gets a free row with the reason and the run succeeds. A run fails only when videos do have a matching caption track and YouTube still returns no text.

What does "blocked" mean? YouTube sometimes answers cloud IP addresses with a "Sign in to confirm you're not a bot" check or an empty caption file. The Actor then switches to a new proxy session and retries a few times. If YouTube still refuses, the video gets a free row with reason blocked. If no transcript at all could be read although videos have captions, the run fails with a message instead of finishing green; run it again with the RESIDENTIAL proxy group. When more than half of the videos in a run are blocked, the status message warns and gives the count.

How many videos can one run read? There is no fixed cap: the run stops at Maximum transcripts or your spending limit. The Actor reads at most 50 listing pages (about 5,000 videos) per channel or playlist, and says so in the status message when it stops there. YouTube lists only the newest Shorts of a channel (for @TED: 100 of 520), and the status message says when a list is shorter than the channel's total.

How often can I run it? As often as you like. For a daily or weekly feed of new uploads, schedule it with Only videos new since my last run on: the first run returns the newest videos up to the per-channel limit, every later run only the uploads published since then (a run with nothing new returns one free row "No new videos since your last run"). A relative date such as "7 days" works too, but returns the same video again in each run that covers it.

Can "only new" miss a video? Yes, in two cases. A video that becomes public later but keeps an older publish date (for example a private upload made public weeks later) sits below the last video read in the channel list, so the next run stops before it. Run once without "only new" and with a publish-date window to catch such videos. Also, on a channel a video that had no captions when it was read is not read again, even if captions are added later. Videos in a temporary state (an upcoming premiere, an empty caption file, a YouTube bot check) are not remembered and are read again by the next run, as long as they are newer than the last video read.

Are timestamps included? Yes, each caption line has a start time and duration in seconds.

Can it translate transcripts? No. It returns caption tracks that exist on YouTube (uploaded or auto-generated), in the language you ask for when the video has it.

Other Actors from the same publisher cover job postings with salaries, company registers and public procurement data.