# YouTube Transcript Scraper - Channels & Playlists (`datagrit/youtube-channel-playlist-transcripts`) Actor

Bulk YouTube transcripts from videos, whole channels and playlists: publish-date window, language priority with fallback, uploaded vs auto captions, video metadata and only-new-since-last-run.

- **URL**: https://apify.com/datagrit/youtube-channel-playlist-transcripts.md
- **Developed by:** [datagrit](https://apify.com/datagrit) (community)
- **Categories:** Videos, AI, Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does YouTube Transcript Scraper - Channels & Playlists do?

It pulls YouTube transcripts in bulk: paste any mix of single videos, whole channels and playlists, and get one row per video with the full transcript text, timestamped caption lines and the video's metadata (title, channel, exact publish date, duration, views, likes, category). It is built for people who need many transcripts at once: researchers, content and SEO teams, and anyone feeding YouTube content into an LLM, a RAG index or a search tool.

### Why use it?

- **Whole channels and playlists in one run.** Channel links (@handle, /channel/UC..., /c/..., /user/...), playlist links and video links can be mixed in one list. Set how many videos to take per channel or playlist; there is no fixed cap per run.
- **Publish-date window.** Keep only videos published after and/or before a date, or within a period such as "7 days" or "3 months". Every video is checked against its exact publish date, and channel listing stops once it reaches older videos.
- **Language control.** Give caption languages in order of preference (en, es, pt-BR...), choose uploaded or auto-generated captions first, or accept only one kind. Turn the language fallback on or off.
- **Only new uploads.** For scheduled runs, the Actor remembers per channel and per settings which videos it already read and stops the listing at the first of them, so a daily run returns only that day's uploads, never the older backlog.
- **Honest rows.** A video without a usable transcript gets a free row with the reason (no captions, other languages only, unavailable, blocked, age-restricted, upcoming), so you always know what happened to every video.
- **Channel content choice.** Long-form videos, Shorts, live stream replays or all uploads.

### Example output

| title | channelName | publishedAt | durationSeconds | language | captionType | wordCount |
|---|---|---|---|---|---|---|
| The Right to Life, Liberty — and Free Time | Danielle Roberts | TED | TED | 2026-09-29T15:00:05.000Z | 233 | en | manual | 547 |

```json
{
  "found": true,
  "reason": null,
  "videoId": "1qmF_znXxrE",
  "title": "The Right to Life, Liberty — and Free Time | Danielle Roberts | TED",
  "channelName": "TED",
  "channelId": "UCAuUUnT6oDeKwE6v1NGQxug",
  "publishedAt": "2026-09-29T15:00:05.000Z",
  "durationSeconds": 233,
  "viewCount": 59691,
  "likeCount": 568,
  "category": "People & Blogs",
  "language": "en",
  "languageName": "English",
  "captionType": "manual",
  "isLanguageFallback": false,
  "availableLanguages": ["ar", "my", "zh-CN", "zh-TW", "en", "fr", "iw", "hi", "pt-BR", "sr", "es", "vi"],
  "transcriptText": "So it's exciting and ironic that I'm giving a TED Talk about time in three minutes. ...",
  "wordCount": 547,
  "segmentCount": 73,
  "segments": [{ "start": 4.292, "duration": 2.711, "text": "So it's exciting and ironic" }],
  "input": "https://www.youtube.com/@TED",
  "sourceType": "channel",
  "sourceTitle": "TED",
  "sourceUrl": "https://www.youtube.com/watch?v=1qmF_znXxrE",
  "scrapedAt": "2026-10-01T08:00:00.000Z"
}
```

A video without a transcript in the requested language (here French, with fallback off) gets a free row:

```json
{
  "found": false,
  "reason": "languageNotAvailable",
  "message": "The video has no captions in the requested languages and language fallback is off.",
  "videoId": "9Ff4S1FJdRM",
  "title": "Is AI Turning Us All into the Same Person? | Sandra Matz | TED",
  "publishedAt": "2026-09-28T15:00:04.000Z",
  "availableLanguages": ["ar", "my", "en", "iw", "pt-BR", "es"],
  "transcriptText": null
}
```

### How much does it cost?

You pay per transcript delivered. Pricing depends on your Apify plan: a small fee when a run starts, then a price per transcript that is lower on paid plans. The Apify free plan includes monthly credit you can use to try it. Rows for videos without a transcript, duplicates and the status row are never charged. You can set a maximum spend on the run: the Actor stops when the limit is reached. It reads YouTube's public web endpoints over plain HTTP (no browser, no video download), so runs are fast.

### Input

- **Videos, channels or playlists** – one link or ID per line. A channel link ending in /shorts or /streams reads that tab. A watch link with \&list= is read as that single video; use the /playlist?list= link for the whole playlist.
- **Maximum transcripts** – total limit for the run (only delivered transcripts count).
- **Maximum videos per channel or playlist** – channels newest first; 0 means no per-source limit.
- **Channel content** – long-form videos, Shorts, live stream replays or all uploads.
- **Published on or after / Published before** – a date (2026-09-01) or a period back from the run (7 days, 2 weeks, 3 months).
- **Title keywords** – keep only videos whose title contains one of the words.
- **Caption languages, Caption type, Fall back to another language** – which caption track to return.
- **Include timestamped segments** – on by default; the plain text is always included.
- **Only videos new since my last run** and **Only-new memory name** – for schedules.
- **Proxy configuration** – Apify Proxy (datacenter) by default.

### How it works

For every channel the Actor reads the channel's uploads list page by page (100 videos per page), newest first. For every video it reads the public video details (publish date, metadata) and the caption track list, picks the track that matches your languages and caption type, and downloads that caption file. That is three small requests per video, plus one listing request per 100 videos of a channel or playlist.

### Is it legal to scrape YouTube transcripts?

The Actor reads only information that YouTube shows to any visitor without signing in: public video pages, captions and channel lists. It does not log in, does not use cookies of an account, does not solve CAPTCHAs and does not download video files. Transcripts are the work of their creators; how you use them (for example quoting, analysis, or training) is your responsibility under copyright law and YouTube's terms. This description is not legal advice.

### FAQ

**Why do some videos have no transcript?** Many videos have no captions at all (music, very new uploads, some Shorts). Such videos get a free row with reason `noCaptions`. The status message counts every reason.

**Why does a run end without transcripts but green?** When none of the videos has captions in your languages (with fallback off) or of the caption type you asked for, every video gets a free row with the reason and the run succeeds. A run fails only when videos do have a matching caption track and YouTube still returns no text.

**What does "blocked" mean?** YouTube sometimes answers cloud IP addresses with a "Sign in to confirm you're not a bot" check or an empty caption file. The Actor then switches to a new proxy session and retries a few times. If YouTube still refuses, the video gets a free row with reason `blocked`. If no transcript at all could be read although videos have captions, the run fails with a message instead of finishing green; run it again with the RESIDENTIAL proxy group. When more than half of the videos in a run are blocked, the status message warns and gives the count.

**How many videos can one run read?** There is no fixed cap: the run stops at Maximum transcripts or your spending limit. The Actor reads at most 50 listing pages (about 5,000 videos) per channel or playlist, and says so in the status message when it stops there. YouTube lists only the newest Shorts of a channel (for @TED: 100 of 520), and the status message says when a list is shorter than the channel's total.

**How often can I run it?** As often as you like. For a daily or weekly feed of new uploads, schedule it with **Only videos new since my last run** on: the first run returns the newest videos up to the per-channel limit, every later run only the uploads published since then (a run with nothing new returns one free row "No new videos since your last run"). A relative date such as "7 days" works too, but returns the same video again in each run that covers it.

**Can "only new" miss a video?** Yes, in two cases. A video that becomes public later but keeps an older publish date (for example a private upload made public weeks later) sits below the last video read in the channel list, so the next run stops before it. Run once without "only new" and with a publish-date window to catch such videos. Also, on a channel a video that had no captions when it was read is not read again, even if captions are added later. Videos in a temporary state (an upcoming premiere, an empty caption file, a YouTube bot check) are not remembered and are read again by the next run, as long as they are newer than the last video read.

**Are timestamps included?** Yes, each caption line has a start time and duration in seconds.

**Can it translate transcripts?** No. It returns caption tracks that exist on YouTube (uploaded or auto-generated), in the language you ask for when the video has it.

### Related Actors

Other Actors from the same publisher cover job postings with salaries, company registers and public procurement data.

# Changelog

This Actor's version history is a separate document: https://apify.com/datagrit/youtube-channel-playlist-transcripts/changelog.md

# Actor input Schema

## `urls` (type: `array`):

Any mix of YouTube links and IDs, one per line: video links (watch, youtu.be, Shorts, live, embed) or 11-character video IDs; channel links (https://www.youtube.com/@TED, /channel/UC..., /c/..., /user/...) or bare @handles; playlist links (https://www.youtube.com/playlist?list=PL...) or playlist IDs. A channel link ending in /shorts or /streams reads that tab. A watch link that also carries \&list= is read as that single video; use the /playlist?list= link for the whole playlist. If the list is empty, the Actor runs a small free example (3 videos from @TED) and says so in the status message.

## `maxItems` (type: `integer`):

Stop after this many transcripts in total across all sources. Only delivered transcripts count and are charged; rows for videos without captions are free and do not count.

## `maxVideosPerSource` (type: `integer`):

How many videos to take from each channel or playlist (channels newest first, playlists in playlist order). Videos without captions count toward this limit; videos outside the publish-date window, without the title keywords or already delivered by an earlier run do not. 0 = no per-source limit (the run still stops at Maximum transcripts and reads at most 50 listing pages, about 5,000 videos, per source).

## `channelContent` (type: `string`):

Which uploads to read from a channel: long-form videos (the channel's Videos tab), Shorts, live stream replays, or all uploads mixed. A channel link ending in /shorts or /streams overrides this for that channel. YouTube lists only the newest Shorts of a channel (measured: 100 of 520 for @TED); the run status says when a list is shorter than the channel's total.

## `publishedAfter` (type: `string`):

Only videos published on or after this date (UTC). Use a date like 2026-09-01 or a period back from the run start like 7 days, 2 weeks, 3 months or 1 year. Checked against each video's exact publish date. On channels the listing stops once it reaches older videos. Leave empty for no lower limit.

## `publishedBefore` (type: `string`):

Only videos published before this date (UTC, the day itself excluded), same format as above. Leave empty for no upper limit. On a channel the newer videos still have to be listed first, so a far-back window reads more listing pages.

## `titleKeywords` (type: `array`):

Only videos whose title contains at least one of these words or phrases (case-insensitive substring match). Leave empty to take every video.

## `languages` (type: `array`):

Caption language codes in order of preference, for example en, es, pt-BR. The first language the video has wins; en also matches en-US and en-GB, and pt-BR also matches pt. Leave empty to take the video's default caption track (usually its spoken language).

## `captionType` (type: `string`):

Uploaded (manual) captions are written by the channel; auto-generated captions come from YouTube speech recognition. Choose which to prefer within each language, or accept only one kind.

## `languageFallback` (type: `boolean`):

When a video has none of the caption languages above, return its default caption track instead (the row says isLanguageFallback: true). Turn off to get a free row with reason languageNotAvailable instead.

## `includeSegments` (type: `boolean`):

Add the caption lines with start time and duration in seconds (segments). The full plain text (transcriptText) is always included.

## `onlyNewSinceLastRun` (type: `boolean`):

For scheduled runs. On a channel the Actor remembers every video an earlier run with the same settings read (with or without a transcript) and stops the listing (newest first) at the first of them, so each run returns only the uploads published since the last run; the first run returns the newest videos up to the per-channel limit. Videos in a temporary state are not remembered and are read again next time: upcoming premieres and scheduled live streams, empty caption files, YouTube bot checks, failed requests and videos without a publish date while a date window is set. On a channel they are read again only while they are newer than the stop point. On playlists and single videos, videos whose transcript was already delivered are skipped and videos without a transcript are tried again. The memory is kept per source and per combination of channel content, date window, title keywords, languages and caption settings (plus the optional memory name).

## `stateKey` (type: `string`):

Optional name that keeps a separate only-new memory, for example one per schedule or per client. Runs with the same sources, settings and name share the memory.

## `proxyConfiguration` (type: `object`):

YouTube is read through Apify Proxy (datacenter) by default. The Actor switches to a new proxy session when YouTube asks for a sign-in bot check, rate-limits, or returns an empty caption file. If a run fails with a block message, switch to the RESIDENTIAL group.

## Actor input object example

```json
{
  "urls": [
    "https://www.youtube.com/@TED",
    "https://www.youtube.com/playlist?list=PLUCPc-R61w-s"
  ],
  "maxItems": 10,
  "maxVideosPerSource": 5,
  "channelContent": "videos",
  "publishedAfter": "",
  "publishedBefore": "",
  "titleKeywords": [],
  "languages": [
    "en"
  ],
  "captionType": "manualFirst",
  "languageFallback": true,
  "includeSegments": true,
  "onlyNewSinceLastRun": false,
  "stateKey": "",
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One row per video: transcript text, timestamped segments and video metadata, or the reason there is no transcript.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.youtube.com/@TED",
        "https://www.youtube.com/playlist?list=PLUCPc-R61w-s"
    ],
    "maxItems": 10,
    "maxVideosPerSource": 5,
    "channelContent": "videos",
    "publishedAfter": "",
    "publishedBefore": "",
    "titleKeywords": [],
    "languages": [
        "en"
    ],
    "captionType": "manualFirst",
    "languageFallback": true,
    "includeSegments": true,
    "onlyNewSinceLastRun": false,
    "stateKey": "",
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagrit/youtube-channel-playlist-transcripts").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://www.youtube.com/@TED",
        "https://www.youtube.com/playlist?list=PLUCPc-R61w-s",
    ],
    "maxItems": 10,
    "maxVideosPerSource": 5,
    "channelContent": "videos",
    "publishedAfter": "",
    "publishedBefore": "",
    "titleKeywords": [],
    "languages": ["en"],
    "captionType": "manualFirst",
    "languageFallback": True,
    "includeSegments": True,
    "onlyNewSinceLastRun": False,
    "stateKey": "",
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("datagrit/youtube-channel-playlist-transcripts").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.youtube.com/@TED",
    "https://www.youtube.com/playlist?list=PLUCPc-R61w-s"
  ],
  "maxItems": 10,
  "maxVideosPerSource": 5,
  "channelContent": "videos",
  "publishedAfter": "",
  "publishedBefore": "",
  "titleKeywords": [],
  "languages": [
    "en"
  ],
  "captionType": "manualFirst",
  "languageFallback": true,
  "includeSegments": true,
  "onlyNewSinceLastRun": false,
  "stateKey": "",
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call datagrit/youtube-channel-playlist-transcripts --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagrit/youtube-channel-playlist-transcripts"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/c8flcUB1xxySw2yvT/builds/s2OjtTm8Jl1WuMFfg/openapi.json
