# YouTube Transcript Scraper (`lakex/youtube-transcript-scraper`) Actor

Get the transcript of any YouTube video, with timestamps, language, title and channel. Manual captions first, auto-generated as fallback. Pay only for transcripts delivered.

- **URL**: https://apify.com/lakex/youtube-transcript-scraper.md
- **Developed by:** [Seth Lake](https://apify.com/lakex) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Transcript Scraper

Paste YouTube video, channel or playlist links and get each video's full transcript, with timestamps, language, title and channel, as one clean row. **You pay only for transcripts delivered**: never for bot checks, retries, channel and playlist listing, videos without captions, removed videos or failed runs.

### Who it's for

- **AI and RAG builders** feeding video content into LLMs, search indexes and agents.
- **Creators and marketers** repurposing videos into posts, show notes and summaries, or studying what competitors say.
- **Researchers and analysts** working with talks, lectures, interviews and news at scale.

### What you get

One row per video:

| Field                                                 | Example                                                       |
| ----------------------------------------------------- | ------------------------------------------------------------- |
| `transcript`                                          | the full text, ready for an LLM or a document                 |
| `segments`                                            | `start` and `duration` in seconds, and `text`, per line       |
| `language`, `languageName`, `isAutoGenerated`         | en, English, false                                            |
| `availableLanguages`                                  | every caption track the video has                             |
| `wordCount`                                           | 1843                                                          |
| `title`, `channelName`, `channelId`                   | the video's title and channel                                 |
| `lengthSeconds`, `viewCount`, `publishDate`, `isLive` | 213, 1234567, null (not provided by YouTube's app API), false |

Manual (creator-uploaded) captions are used first, YouTube's auto-generated captions as fallback. The row tells you which one you got. Download as JSON, CSV or Excel, or use the API.

### How to use it

1. Paste links into **Video, channel or playlist URLs**, one per line:
   - **Videos:** `youtube.com/watch?v=…`, `youtu.be/…`, `/shorts/…`, `/live/…`, or the 11-character ID.
   - **Channels:** `youtube.com/@name`, `/channel/UC…`, `/c/…` or `/user/…`. You get the channel's uploads, newest first.
   - **Playlists:** `youtube.com/playlist?list=…`.
2. Set **Preferred language** (default `en`).
3. Set **Max videos** to cap the cost of the run (across all links), and **Max videos per channel or playlist** (default 100). Click **Start**.

A video that shows up twice (say, in a playlist and as its own link) is fetched and charged once.

#### Input example

```json
{
    "videoUrls": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://www.youtube.com/@3blue1brown",
        "https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_67000Dx_ZCJB-3pi"
    ],
    "language": "en",
    "includeTimestamps": true,
    "maxItems": 100,
    "maxVideosPerSource": 50
}
```

#### Output example

Shortened; numbers are illustrative.

```json
{
    "videoId": "dQw4w9WgXcQ",
    "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "title": "Rick Astley - Never Gonna Give You Up (Official Video)",
    "channelName": "Rick Astley",
    "channelId": "UCuAXFkgsw1L7xaCfnd5JJOw",
    "lengthSeconds": 213,
    "viewCount": 1234567,
    "publishDate": null,
    "isLive": false,
    "language": "en",
    "languageName": "English",
    "isAutoGenerated": false,
    "availableLanguages": [
        { "code": "en", "name": "English", "isAutoGenerated": false },
        { "code": "en", "name": "English", "isAutoGenerated": true }
    ],
    "transcript": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ …",
    "wordCount": 13,
    "segments": [
        { "start": 1.36, "duration": 1.68, "text": "[♪♪♪]" },
        { "start": 18.64, "duration": 3.24, "text": "♪ We're no strangers to love ♪" }
    ],
    "input": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "scrapedAt": "2026-09-30T12:00:00.000Z"
}
```

`input` is the link the video came from: the video link itself, or the channel or playlist link. Missing values are `null`, never left out, so every row has the same columns. Turn off **Include timestamped segments** to get plain text only (`segments` is then `null`).

### Pricing

**$5 per 1,000 transcripts delivered** ($0.005 per video), less on higher Apify plans: $4 on Silver, $3 on Gold and above. Proxies and compute are included, and there is no start fee.

- You're charged only when a transcript row lands in your dataset.
- Bot checks, retries, listing channels and playlists, videos without captions and removed or private videos are free.
- A run that delivers nothing costs nothing.
- **Max videos** and Apify's **Maximum cost per run** both cap your spend. If a run reaches your spending limit, it stops cleanly, keeps every transcript delivered so far, and tells you how many were left.

### What's not charged, and where it's listed

Every input that didn't produce a row is listed in the run's **RUN\_SUMMARY** record (Storage → Key-value store), with the reason. RUN\_SUMMARY also lists each channel and playlist with its title and how many videos were found.

| Reason          | Meaning                                                                                                                       |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| `no_transcript` | The video has no captions and no auto-generated transcript.                                                                   |
| `unavailable`   | Removed, private, age-restricted or region-locked (checked twice), or a channel or playlist that doesn't exist or is private. |
| `blocked`       | YouTube kept showing a bot check after every retry. Try again later.                                                          |
| `failed`        | Network or page error after every retry.                                                                                      |
| `invalid_input` | Not a YouTube video, channel or playlist link.                                                                                |
| `not_attempted` | The run stopped first (Max videos or your spending limit).                                                                    |

### Known limits

- **No search** yet: paste video, channel or playlist links.
- Channels give all uploads (videos, Shorts and past live streams), newest first. Mixes, Watch Later and Liked videos aren't supported.
- **No translation** yet: you get a language the video actually has.
- **No speech-to-text.** Videos without any captions are skipped (and free), not transcribed.
- Age-restricted and members-only videos need a login, so they're skipped.
- YouTube changes its site often. If something looks wrong, open an issue and we'll fix it fast.

### Advanced settings

- **Proxy**: Apify US residential by default, included in the price. YouTube blocks datacenter IPs, so leave this unless you know you need something else.
- **Max parallel videos** (default 10) and **Retries per video** (default 5).

### Is it legal?

This Actor collects only public information that anyone can see on YouTube without logging in: captions, titles and channel names. No personal data, no accounts. YouTube's terms restrict automated access, so check that your use case is allowed where you are, and consult a lawyer if unsure.

### Changelog

- **Price** (2026-09-30): lowered to $5 per 1,000 ($4 Silver, $3 Gold and above).
- **0.2** (2026-09-30): Channel and playlist links, with **Max videos per channel or playlist**. Listing is free; you still pay only for transcripts delivered.
- **0.1** (2026-09-30): First version. Video links and IDs, manual and auto-generated transcripts with timestamps, pay only for transcripts delivered.

### Development

```bash
npm install
npm test            # parser tests + crawl test against a local fake YouTube
npx apify run       # local run; residential proxy only works on the platform, so locally set
                    # "proxyConfiguration": { "useApifyProxy": false } in the input
```

Code: `src/youtube.ts` (pages and responses → data, pure), `src/scraper.ts` (crawl, retries, fallbacks, charging), `src/main.ts` (input, summary, status message). Tests are in `test/`. Channel and playlist fixtures are trimmed real responses (`test/fixtures/trim-real.mjs`); the rest are synthetic.

# Actor input Schema

## `videoUrls` (type: `array`):

One per line. Videos: youtube.com/watch?v=…, youtu.be/…, /shorts/…, /live/…, or the 11-character ID. Channels: youtube.com/@name, /channel/UC…, /c/… or /user/… (all uploads, newest first). Playlists: youtube.com/playlist?list=….

## `language` (type: `string`):

Language code like en, es or de. We pick a manual transcript in this language first, then an auto-generated one, then English, then whatever the video has. The row says which language you got.

## `includeTimestamps` (type: `boolean`):

Add the transcript as segments with start time and duration (in seconds). The plain full text is always included.

## `maxItems` (type: `integer`):

Stop after this many transcripts are delivered, across all links. You're charged per transcript delivered, so this caps the cost of a run.

## `maxVideosPerSource` (type: `integer`):

For each channel or playlist link, take at most this many videos (newest first for channels). Listing channels and playlists is free; you pay only for transcripts delivered.

## `proxyConfiguration` (type: `object`):

YouTube blocks datacenter IPs. The default (Apify US residential proxy) is included in the price, so most users should leave it as is.

## `maxConcurrency` (type: `integer`):

How many videos to fetch at once.

## `maxRequestRetries` (type: `integer`):

How many times to retry a video on a new IP when YouTube shows a bot check. Retries are free for you.

## Actor input object example

```json
{
  "videoUrls": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://youtu.be/jNQXAC9IVRw",
    "iG9CE55wbtY"
  ],
  "language": "en",
  "includeTimestamps": true,
  "maxItems": 20,
  "maxVideosPerSource": 100,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  },
  "maxConcurrency": 10,
  "maxRequestRetries": 5
}
```

# Actor output Schema

## `transcripts` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoUrls": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://youtu.be/jNQXAC9IVRw",
        "iG9CE55wbtY"
    ],
    "language": "en",
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "US"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("lakex/youtube-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "videoUrls": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://youtu.be/jNQXAC9IVRw",
        "iG9CE55wbtY",
    ],
    "language": "en",
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "US",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("lakex/youtube-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoUrls": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://youtu.be/jNQXAC9IVRw",
    "iG9CE55wbtY"
  ],
  "language": "en",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}' |
apify call lakex/youtube-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,lakex/youtube-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wIQJT7FlSj5VCiLPH/builds/LZzVKwnGbzQdzIZbl/openapi.json
