# YouTube Transcript Scraper (`cairnware/youtube-transcript-scraper`) Actor

Get YouTube video and Shorts transcripts as plain text, timed segments or SRT, with title, channel, views and publish date. Pick languages in priority order. Videos without captions are never charged.

- **URL**: https://apify.com/cairnware/youtube-transcript-scraper.md
- **Developed by:** [Cairnware](https://apify.com/cairnware) (community)
- **Categories:** Videos, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$4.00 / 1,000 transcripts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Transcript Scraper

Get the **transcript (captions / subtitles) of any YouTube video or Short** as clean plain text, **timestamped segments** or an **SRT subtitle file**, together with the video's title, channel, duration, views, likes and publish date. One row per video, ready for JSON, CSV, Excel, Google Sheets, the Apify API or your LLM pipeline.

Built for runs you do not have to babysit:

- **You only pay for transcripts delivered.** Videos without captions, unavailable or private videos, invalid links and failures are returned as free rows with a clear reason. No start fee.
- **Every input is accounted for.** Each link you enter appears in the output with `status` `ok`, `no_transcript` or `error`. The run ends with a summary such as *"48 of 50 videos have a transcript; 2 have no captions in the requested languages (not charged)"*.
- **Smart retries.** When YouTube blocks or rate-limits a request, it is retried with exponential backoff on a **fresh residential IP**, switching between YouTube's app clients.
- **Language control.** Give your preferred languages in order; human-made captions are preferred over auto-generated ones, and you can fall back to the video's original language.
- **Respects your budget.** The Actor stops when your *maximum cost per run* is reached and lists the remaining videos as free, unscraped rows.

### What data you get

| Field | What it contains |
|---|---|
| `text` | The full transcript as one plain-text string (HTML entities decoded, line breaks joined). |
| `segments` | Every caption line with `start` and `duration` in seconds (on by default). |
| `srt` | The transcript as an SRT subtitle file (off by default). |
| `language`, `languageName`, `isAutoGenerated` | Which caption track you got, e.g. `en`, "English (auto-generated)", `true`. |
| `availableLanguages` | Every caption track the video has, so you can re-run in another language. |
| `title`, `channelName`, `channelId`, `channelUrl`, `durationSeconds`, `viewCount`, `isLiveContent` | Always included. |
| `publishDate`, `likeCount`, `category`, `description`, `keywords`, `thumbnailUrl` | With **Full video details** (on by default). |
| `wordCount`, `segmentCount`, `warnings`, `errorMessage`, `scrapedAt` | Run bookkeeping. |

### Use cases

- **AI and LLM workflows:** feed transcripts into summarisation, RAG, Q\&A or content repurposing pipelines.
- **Content research and SEO:** see what top videos in your niche actually say; turn videos into blog posts, show notes or quotes.
- **Market and brand research:** analyse product reviews, interviews and talks at scale.
- **Subtitles:** download SRT files for editing or translation workflows.
- **Education and accessibility:** create study notes and searchable text from lectures.

### How to use

1. Paste one or more **YouTube video or Shorts links** (or video IDs), one per line.
2. Optionally set your **preferred languages** (default `en`).
3. Click **Start**. Download the results as JSON, CSV, Excel or HTML, or read them through the Apify API.

All common link formats work: `youtube.com/watch?v=...`, `youtu.be/...`, `/shorts/...`, `/embed/...`, `/live/...`, `m.youtube.com`, `music.youtube.com`, and extra parameters like `&t=30s` or `&list=...` are ignored. Playlist, channel and search links are not supported yet: add the video links themselves.

### Input example

```json
{
    "videos": [
        "https://www.youtube.com/watch?v=jNQXAC9IVRw",
        "https://youtu.be/dQw4w9WgXcQ",
        "https://www.youtube.com/shorts/VIDEO_ID"
    ],
    "languages": ["en", "es"],
    "fallbackToAnyLanguage": true,
    "includeAutoGenerated": true,
    "includeTimestamps": true,
    "includeSrt": false,
    "includeMetadata": true,
    "proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}
```

| Field | Notes |
|---|---|
| `videos` | Required. Up to 10,000 per run. Duplicates are removed. |
| `languages` | Language codes in order of preference (`en`, `es`, `pt-BR`, `zh-Hans`...). `en` matches every English variant; `pt-BR` prefers Brazilian Portuguese, then other Portuguese. Empty = each video's original language. |
| `fallbackToAnyLanguage` | If none of your languages exists, return the original-language transcript (with a warning). Off = `no_transcript` row, not charged. |
| `includeAutoGenerated` | Use YouTube's automatic captions when the uploader added none (most videos only have these). |
| `includeTimestamps` | Add the `segments` array. |
| `includeSrt` | Add the `srt` field. |
| `includeMetadata` | Add publish date, likes, category, description, tags and thumbnail. |
| `maxRetries` | Retries per request before a video is marked `error` (default 6). |
| `maxConcurrency` | Videos processed in parallel (default 5). |

### Output example

Real output for "Me at the zoo", collected on 2026-10-03 (description and segments shortened):

```json
{
    "input": "jNQXAC9IVRw",
    "videoId": "jNQXAC9IVRw",
    "url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
    "status": "ok",
    "errorMessage": null,
    "title": "Me at the zoo",
    "channelName": "jawed",
    "channelId": "UC4QobU6STFB0P71PMvOGN5A",
    "channelUrl": "https://www.youtube.com/channel/UC4QobU6STFB0P71PMvOGN5A",
    "durationSeconds": 19,
    "viewCount": 439481354,
    "likeCount": 19982756,
    "publishDate": "2005-04-23T20:31:52-07:00",
    "category": "Film & Animation",
    "description": "...",
    "keywords": ["me at the zoo", "jawed karim", "first youtube video"],
    "thumbnailUrl": "https://i.ytimg.com/vi/jNQXAC9IVRw/hqdefault.jpg",
    "isLiveContent": false,
    "language": "en",
    "languageName": "English",
    "isAutoGenerated": false,
    "availableLanguages": [
        { "languageCode": "en", "languageName": "English", "isAutoGenerated": false },
        { "languageCode": "de", "languageName": "German", "isAutoGenerated": false }
    ],
    "text": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say",
    "wordCount": 39,
    "segmentCount": 6,
    "segments": [
        { "start": 1.2, "duration": 2.16, "text": "All right, so here we are, in front of the elephants" },
        { "start": 5.318, "duration": 2.656, "text": "the cool thing about these guys is that they have really..." },
        { "start": 7.974, "duration": 4.642, "text": "really really long trunks" }
    ],
    "srt": null,
    "warnings": [],
    "scrapedAt": "2026-10-03T10:25:02Z"
}
```

A video that has no captions in your languages (with fallback off) looks like this and is **not charged**:

```json
{
    "videoId": "9bZkp7q19f0",
    "status": "no_transcript",
    "errorMessage": "No captions in en. Available: ko (auto). Turn on \"Fall back to the original language\" or add one of these languages.",
    "title": "PSY - GANGNAM STYLE(강남스타일) M/V",
    "availableLanguages": [{ "languageCode": "ko", "languageName": "Korean (auto-generated)", "isAutoGenerated": true }],
    "text": null
}
```

The dataset has two views: **Overview** (one row per video) and **Timestamped lines** (one row per caption line, handy for spreadsheets).

### Pricing

**$4 per 1,000 transcripts** ($0.004 per video). **No start fee. You only pay for transcripts delivered.**

- Charged once per row with `status: "ok"`.
- `no_transcript` and `error` rows (no captions, unavailable, private, members-only, invalid link, failed after retries, skipped by your cost limit) are free.
- Residential proxy traffic and the metadata lookup are included in the price.
- Set a **maximum cost per run** in the run options; the Actor never goes over it.

Example: 500 videos a month ≈ $2.

### FAQ

**Why do some videos return `no_transcript`?**
The video has no caption track that matches your settings: the uploader added none and YouTube did not create automatic captions (common for music without speech, very short or very new videos), or captions exist only in other languages. `availableLanguages` lists what the video has. You are not charged.

**What is the difference between auto-generated and human-made captions?**
Human-made captions were uploaded by the creator and are usually accurate and punctuated. Auto-generated captions come from YouTube's speech recognition: they cover most videos but have no punctuation and some recognition errors. `isAutoGenerated` tells you which one you got.

**Can I get a translation into another language?**
Not yet. YouTube's machine translation of captions is heavily rate-limited, so the Actor returns the caption tracks that actually exist. Use `availableLanguages` to see them.

**In what order are the rows?**
In the order videos finish, which can differ from your input order when videos run in parallel. Each row has the `input` you gave and the `videoId`, so you can sort or join on them.

**Why is a residential proxy the default?**
YouTube blocks many datacenter IP ranges ("Sign in to confirm you're not a bot"). Residential IPs, with a new IP on every retry, keep the failure rate low. You can choose another proxy in *Proxy configuration*, but expect more retries and errors.

**Can I schedule it or call it from code?**
Yes. Use Apify Schedules for recurring runs, and the API tab of this Actor for ready-made examples in Python, JavaScript and cURL. Integrations exist for Google Sheets, Make, n8n, Zapier and webhooks.

### Limitations

- Only videos and Shorts. Playlist, channel and search links are not expanded yet (planned).
- Age-restricted, private and members-only videos cannot be read without signing in, so they return `error` (free). The Actor never signs in to YouTube.
- Live streams return captions only after the stream has ended and YouTube has processed them.
- Machine translation of captions is not offered (see FAQ).
- The Actor uses the same public endpoints as YouTube's apps. If YouTube changes them, results may break until the Actor is updated. Please report it and we will fix it.
- It does not download video or audio and does not collect comments or personal data. Transcripts are the creators' content: you are responsible for using them in line with YouTube's terms, copyright and the laws that apply to you.

### Support

Questions, bugs or feature requests: open an issue on the **Issues** tab or email **hello@cairnware.com**. Please include the run ID if something went wrong. Support replies may be AI-assisted.

Made by **Cairnware**.

# Actor input Schema

## `videos` (type: `array`):

Video links in any format (youtube.com/watch?v=..., youtu.be/..., /shorts/..., /embed/..., /live/..., m.youtube.com, music.youtube.com) or bare video IDs, one per line. Duplicates are removed.

## `languages` (type: `array`):

Caption languages in order of preference, as codes like en, es, de, pt-BR, zh-Hans. "en" also matches regional variants such as en-GB. For each language, captions written by the uploader are preferred over auto-generated ones. Leave empty to get each video's original language.

## `fallbackToAnyLanguage` (type: `boolean`):

If a video has no captions in your preferred languages, return its original-language transcript instead (the row gets a warning). Turn off to get status "no\_transcript" (not charged) for such videos.

## `includeAutoGenerated` (type: `boolean`):

Use YouTube's automatic (speech recognition) captions when the uploader did not provide any. Most videos only have these. Turn off to accept human-made captions only.

## `includeTimestamps` (type: `boolean`):

Add "segments": every caption line with its start time and duration in seconds. The full plain text is always included in "text".

## `includeSrt` (type: `boolean`):

Add the transcript as an SRT subtitle file in the "srt" field.

## `includeMetadata` (type: `boolean`):

Add publish date, like count, category, description, tags and thumbnail. Title, channel, duration and views are always included.

## `proxyConfiguration` (type: `object`):

YouTube blocks many datacenter IP addresses. Residential proxies give each retry a fresh IP and are strongly recommended; proxy traffic is included in the price.

## `maxRetries` (type: `integer`):

How many times a blocked or failed request is retried (exponential backoff, new proxy IP each time) before the video is returned with status "error".

## `maxConcurrency` (type: `integer`):

How many videos are processed at the same time.

## Actor input object example

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "languages": [
    "en"
  ],
  "fallbackToAnyLanguage": true,
  "includeAutoGenerated": true,
  "includeTimestamps": true,
  "includeSrt": false,
  "includeMetadata": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "maxRetries": 6,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `transcripts` (type: `string`):

One row per video: full transcript text (plus timestamped segments and SRT when enabled), language, caption type, title, channel, views, publish date and status. Only rows with status "ok" are charged.

## `segments` (type: `string`):

One row per caption line with start time and duration (when "Timestamped segments" is on, the default).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
    ],
    "languages": [
        "en"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("cairnware/youtube-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
    "languages": ["en"],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("cairnware/youtube-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "languages": [
    "en"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call cairnware/youtube-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cairnware/youtube-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/GKd9Tq0bf4srZtANe/builds/HKp1In993HTW7CgRk/openapi.json
