# YouTube Transcript & Subtitles API — Channels, Search & SRT (`cazadores/youtube-transcript-bulk`) Actor

Bulk YouTube transcript API: YouTube to text from videos, channels, playlists and search. Captions as text, timestamped segments, SRT/VTT or RAG chunks, with metadata and optional AI translation. Works as an MCP tool for AI agents. 99%+ success; you only pay for delivered transcripts.

- **URL**: https://apify.com/cazadores/youtube-transcript-bulk.md
- **Developed by:** [Juan Manuel D'Amico](https://apify.com/cazadores) (community)
- **Categories:** Videos, AI, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 transcripts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Transcript & Subtitles API — Bulk Channels, Playlists, Search & SRT

Turn YouTube to text: extract transcripts (captions / subtitles) in bulk from single videos, whole channels, playlists and search results — as plain text, timestamped segments, SRT/VTT subtitle files or RAG-ready chunks. **$3 per 1,000 transcripts, no start fee, and you only pay for successful transcripts.**

### Why this YouTube transcript API

| | This Actor | Typical alternatives |
|---|---|---|
| Price per 1,000 transcripts | **$3** | $5–10 |
| Start fee per run | **None** | Often charged |
| Failed videos (no captions, private, deleted) | **Free** | Often charged |
| Whole channels, playlists & YouTube search | **Yes, with pagination** | Often single videos only |
| Filter channels by Shorts / videos / lives and by date | **Yes** | Rare |
| Translation | **YouTube captions free, AI translation optional** | Rare |
| SRT / WebVTT subtitle files | **Yes** | Rare |
| RAG-ready chunks for LLMs | **Yes** | Rare |
| Video metadata (title, channel, date, views) | **Included** | Sometimes extra |

### What it does

- Accepts any YouTube URL format: `watch?v=`, `youtu.be/`, Shorts, `/embed/`, `/live/`, `m.youtube.com`, or bare 11-character video IDs.
- Takes channels (`@handle`, `/channel/UC…`, `/c/…`, `/user/…`) and playlists, and extracts the latest N videos (or all of them).
- Searches YouTube for keywords and extracts transcripts of the top results.
- Filters channels to regular videos, Shorts or past live streams only, and skips videos published before a date (skipped videos are free).
- Picks the best caption track for your preferred languages, preferring human-made captions over auto-generated ones.
- Optional translation to another language (see [Translation](#translation)).
- Removes duplicates, retries blocked requests with fresh residential IPs, and never fails the whole run because of a single video.

### Input example

```json
{
  "videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ", "https://youtu.be/9Ff4S1FJdRM"],
  "channels": ["https://www.youtube.com/@veritasium"],
  "playlists": ["https://www.youtube.com/playlist?list=PLFs4vir_WsTwEd-nJgVJCZPNL3HALHHpF"],
  "searchQueries": ["black holes explained"],
  "maxVideosPerSource": 50,
  "channelContentType": "videos",
  "publishedAfter": "2026-01-01",
  "languages": ["en", "es"],
  "translateTo": "",
  "outputFormat": "both",
  "chunkSeconds": 60,
  "subtitleFormats": ["srt"]
}
```

| Field | Default | Description |
|---|---|---|
| `videos` | – | Video URLs or IDs |
| `channels` | – | Channel URLs or handles |
| `playlists` | – | Playlist URLs or IDs |
| `searchQueries` | – | Keywords to search on YouTube (videos only) |
| `maxVideosPerSource` | 50 | Max videos per channel, playlist or search query (0 = all) |
| `channelContentType` | `all` | `all`, `videos`, `shorts` or `streams` (past live streams) |
| `publishedAfter` | – | Skip videos published before this date (`YYYY-MM-DD`); skipped videos are not charged |
| `languages` | `["en"]` | Preferred languages, in order. If none exists, the video's own language is returned |
| `preferManual` | `true` | Prefer human-made captions over auto-generated |
| `translateTo` | `""` | Language code to translate into (e.g. `es`) |
| `aiTranslation` | `false` | Translate with AI when YouTube can't (charged per 1,000 words, see [Pricing](#pricing)) |
| `outputFormat` | `both` | `text`, `segments` or `both` |
| `chunkSeconds` | 0 | If > 0, adds `chunks` of ~N seconds for embeddings / RAG |
| `subtitleFormats` | – | `srt` and/or `vtt`: adds ready-to-use subtitle files |
| `includeMetadata` | `true` | Title, channel, publish date, duration, views, description, tags, thumbnail |
| `maxConcurrency` | 10 | Videos processed in parallel |

### Output example

One dataset item per video:

```json
{
  "videoId": "abc123def45",
  "url": "https://www.youtube.com/watch?v=abc123def45",
  "status": "ok",
  "title": "How Black Holes Actually Work",
  "channelName": "Science Channel",
  "channelId": "UCxxxxxxxxxxxxxxxxxxxxxx",
  "publishedAt": "2026-05-01",
  "durationSeconds": 732,
  "viewCount": 1234567,
  "language": "en",
  "languageName": "English",
  "isAutoGenerated": false,
  "availableLanguages": ["en", "es", "pt-BR"],
  "wordCount": 1830,
  "text": "Imagine you could fall into a black hole. What would you actually see?…",
  "segments": [{ "start": 0.0, "duration": 3.2, "text": "Imagine you could fall into a black hole." }],
  "chunks": [{ "start": 0.0, "end": 60.4, "text": "Imagine you could fall into a black hole. What would you actually see?…" }]
}
```

`status` is one of:

- `ok` — transcript extracted (**the only charged status**)
- `no_transcript` — the video has no captions
- `unavailable` — private, deleted, age-restricted or region-blocked video (also used for channels/playlists that don't exist)
- `error` — invalid input or a failure after several retries; see the `error` field

A run summary (`ok`, `noTranscript`, `unavailable`, `errors`) is saved in the key-value store as `SUMMARY`.

### Use cases

- **RAG & chatbots**: feed transcripts of a whole channel into a vector database using `chunkSeconds`.
- **AI agents**: give an agent the ability to "watch" YouTube videos through the transcript API.
- **Content research & SEO**: analyze what top creators talk about, find keywords and topics.
- **Translated subtitles**: get transcripts in your audience's language.
- **Channel analysis**: combine transcripts with views, dates and durations.

### How to use it from code

#### API (cURL)

```bash
curl -X POST "https://api.apify.com/v2/acts/cazadores~youtube-transcript-bulk/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"]}'
```

#### Python

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_TOKEN")
run = client.actor("cazadores/youtube-transcript-bulk").call(run_input={
    "channels": ["https://www.youtube.com/@veritasium"],
    "maxVideosPerSource": 20,
    "outputFormat": "text",
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["title"], item.get("wordCount"))
```

#### JavaScript

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('cazadores/youtube-transcript-bulk').call({
    videos: ['https://youtu.be/dQw4w9WgXcQ'],
    chunkSeconds: 60,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].chunks);
```

#### As an MCP tool for AI agents

Add the Apify MCP server to Claude, Cursor or any MCP client and enable this Actor as a tool:

```
https://mcp.apify.com/?actors=cazadores/youtube-transcript-bulk
```

Your agent can then call it with a list of video URLs and read the transcripts directly.

### Pricing

Pay-per-event, no subscription and no start fee:

| Event | Price |
|---|---|
| Successful transcript (`status: ok`) | **$0.003** ($3 per 1,000) — lower on higher Apify plans |
| AI translation (only with `aiTranslation` on, and only when YouTube can't translate) | **$0.003 per started 1,000 words** |
| Videos without captions, unavailable videos, errors, videos skipped by date | **Free** |
| Listing channels, playlists & search, metadata, subtitle files, YouTube translation | **Included** |

Proxy (residential) and compute costs are included in the price. Set a maximum cost per run in the run options and the Actor stops cleanly when it's reached.

### Translation

When `translateTo` is set, the Actor:

1. Uses the video's own captions in that language when they exist (human-made or auto-generated).
2. Otherwise tries YouTube's automatic translation.
3. If YouTube doesn't provide a translation and `aiTranslation` is on, translates the transcript with AI line by line, so timestamps, chunks and SRT/VTT stay in sync with the video (`translationSource: "ai"`).
4. Otherwise returns the transcript in its original language with a `translationError` field (still a successful transcript).

The fields `translatedTo`, `translationSource` (`youtube_captions`, `youtube_auto_translate` or `ai`) and `originalLanguage` tell you what happened for each video.

### Known limitations

- Videos with no captions at all (common for Shorts, music and some live streams) return `no_transcript`. Speech-to-text for those videos is not included.
- Private, deleted, members-only, age-restricted and region-blocked videos return `unavailable`.
- YouTube's automatic translation is not available for every video; turn on `aiTranslation` to cover the rest (see [Translation](#translation)).
- Channel extraction uses the channel's uploads list, which includes regular videos, Shorts and past live streams.

### Related actors

- [TikTok Shop Scraper](https://apify.com/cazadores/tiktok-shop-scraper) — bulk TikTok Shop product data: price, units sold, stock per variant, reviews and shop stats.

### En español

**Extractor de transcripciones de YouTube en lote**: pegá URLs de videos, canales o playlists y obtené la transcripción (subtítulos) en texto, segmentos con tiempos o fragmentos listos para RAG. Cuesta **$3 cada 1.000 transcripciones**, sin costo de arranque, y **solo pagás las transcripciones exitosas**: los videos sin subtítulos, privados o borrados no se cobran. Incluye búsqueda por palabra clave, filtros por Shorts, videos o transmisiones y por fecha, archivos SRT/VTT, metadatos y traducción (gratis con los subtítulos de YouTube; con IA, opcional, a $0,003 cada 1.000 palabras).

# Actor input Schema

## `videos` (type: `array`):

YouTube video URLs (watch, youtu.be, Shorts, live, embed) or 11-character video IDs.

## `channels` (type: `array`):

Channel URLs (e.g. https://www.youtube.com/@handle). Transcripts of the channel's latest videos are extracted.

## `playlists` (type: `array`):

Playlist URLs. Transcripts of all videos in the playlist are extracted.

## `searchQueries` (type: `array`):

Search YouTube and extract transcripts of the top results (videos only). 'Max videos per source' applies to each query.

## `maxVideosPerSource` (type: `integer`):

Maximum number of videos taken from each channel or playlist. 0 = no limit.

## `channelContentType` (type: `string`):

Which uploads to take from channels.

## `publishedAfter` (type: `string`):

Only process videos published on or after this date (YYYY-MM-DD). Older videos are skipped and not charged.

## `languages` (type: `array`):

Transcript languages in order of preference (ISO codes like en, es, pt). The first available one is used.

## `preferManual` (type: `boolean`):

Use manually created captions when available instead of auto-generated ones.

## `translateTo` (type: `string`):

Optional language code (e.g. es). Uses YouTube's own captions in that language when they exist, otherwise YouTube auto-translation when available (best effort). If translation isn't possible, the original transcript is returned with a 'translationError' field.

## `aiTranslation` (type: `boolean`):

If YouTube has no captions or auto-translation in the requested language, translate the transcript with AI (timestamps preserved). Charged separately per 1,000 words translated; YouTube translations stay free.

## `outputFormat` (type: `string`):

text = full plain text, segments = timestamped segments, both = both.

## `chunkSeconds` (type: `integer`):

If greater than 0, adds a 'chunks' field grouping the transcript into blocks of this many seconds, ready for embeddings.

## `subtitleFormats` (type: `array`):

Also return the transcript as ready-to-use subtitle files (fields 'srt' and 'vtt').

## `includeMetadata` (type: `boolean`):

Adds title, channel, publish date, duration and view count.

## `maxConcurrency` (type: `integer`):

Number of videos processed in parallel.

## `proxyConfiguration` (type: `object`):

Proxy settings. The default (Apify residential) is included in the price and gives the best success rate; YouTube blocks most datacenter IPs.

## `failIfSuccessRateBelow` (type: `integer`):

For monitoring: mark the run as FAILED when fewer than this % of videos with captions succeed (0 = off).

## Actor input object example

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "maxVideosPerSource": 50,
  "channelContentType": "all",
  "languages": [
    "en"
  ],
  "preferManual": true,
  "translateTo": "",
  "aiTranslation": false,
  "outputFormat": "both",
  "chunkSeconds": 0,
  "includeMetadata": true,
  "maxConcurrency": 10,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "failIfSuccessRateBelow": 0
}
```

# Actor output Schema

## `transcripts` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("cazadores/youtube-transcript-bulk").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("cazadores/youtube-transcript-bulk").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call cazadores/youtube-transcript-bulk --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cazadores/youtube-transcript-bulk"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JwAH5POfTLnU2lpyU/builds/JZPLYO9NZcExIzgOu/openapi.json
