# YouTube Transcript API: Scrape Video Transcripts, No API Key (`openkrill/youtube-transcripts`) Actor

Fetch the transcript of any public YouTube video, Short or embed as clean text, timed JSON, SRT or WebVTT. Manual captions first, auto-generated as fallback, clear error per video. Uses Apify residential proxy. $4 per 1,000 transcripts, MCP-ready.

- **URL**: https://apify.com/openkrill/youtube-transcripts.md
- **Developed by:** [Nguyen Vu Trung Hieu](https://apify.com/openkrill) (community)
- **Categories:** Videos, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Transcript API: Scrape Video Transcripts, No API Key

Fetch the transcript of any public YouTube video, Short or embed as clean text, timed JSON segments, SRT or WebVTT.
Manual captions are used before auto-generated ones, and every video gets either a transcript or a clear error code instead of failing the run.
Pay per use: **$4 per 1,000 transcripts**, failed videos are free.

Use it to download YouTube transcripts in bulk, feed videos to ChatGPT, Claude or a RAG pipeline, or build a searchable transcript archive without a YouTube API key.

This Actor runs through **Apify residential proxy** on the Apify platform, because YouTube blocks most datacenter and cloud IPs.
The proxy is enabled by default in the input and its cost is included in the per-transcript price.

### Use it from an AI agent (MCP)

Call this Actor from any MCP client through the [Apify MCP server](https://docs.apify.com/integrations/mcp).
A hosted MCP server and HTTP API for the same data is being prepared at [api.openkrill.app](https://api.openkrill.app) (planned tool `get_transcript`, overview at [tools.openkrill.app](https://tools.openkrill.app)); the YouTube tool there is coming soon.

### What it does

- Accepts video IDs and every common link form: `watch?v=`, `youtu.be/`, `/shorts/`, `/embed/`, `/live/`, `m.` and `music.` hosts.
- Picks the best caption track by your language priority list, using human-made captions before auto-generated ones for each language; `en` also matches regional tracks such as `en-US` or `en-GB`, with an exact match winning.
- When none of your languages exist, returns the transcript in the video's original spoken language instead of failing (`languageFallback: true` marks it); set `fallbackToAnyLanguage` to `false` to get a `NoTranscriptFound` error instead.
- Optionally returns YouTube's own machine translation into a target language.
- Returns the video title and channel name with every transcript.
- Never fails the whole run because of one video: each video gets either a transcript or an error object with a machine-readable code and a `retryable` flag.
- Skips duplicate videos in the same run, so you are never charged twice for one transcript.
- Retries network and server errors with exponential backoff, and retries blocked requests from fresh proxy IPs when a proxy is configured.

### Input

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `videos` | string\[] | required | Video URLs or IDs, up to 1000 per run. |
| `languages` | string\[] | `["en"]` | Language codes in priority order, e.g. `["en", "en-GB", "de"]`. A code also matches its regional variants. |
| `fallbackToAnyLanguage` | boolean | `true` | When none of `languages` exist, return the original-language (or first available) track instead of an error. |
| `translateTo` | string | none | Language code to translate into when no track in that language exists. |
| `outputFormat` | `segments` | `text` | `srt` | `vtt` | `segments` | Shape of the transcript in each result. |
| `maxConcurrency` | integer 1-10 | `3` | Videos fetched in parallel. |
| `proxyConfiguration` | proxy | Apify residential | Leave the default on the Apify platform; see Limits. |

Example:

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://youtu.be/jNQXAC9IVRw",
    "https://www.youtube.com/shorts/xxxxxxxxxxx"
  ],
  "languages": ["en"],
  "outputFormat": "segments"
}
```

### Output

One dataset item per video.
A successful item (segments shortened):

```json
{
  "status": "ok",
  "input": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "videoId": "dQw4w9WgXcQ",
  "title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
  "channelName": "Rick Astley",
  "languageCode": "en",
  "language": "English",
  "isGenerated": false,
  "languageFallback": false,
  "translatedFrom": null,
  "format": "segments",
  "text": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ ...",
  "segments": [
    { "text": "[♪♪♪]", "start": 1.36, "duration": 1.68 },
    { "text": "♪ We're no strangers to love ♪", "start": 18.64, "duration": 3.24 }
  ],
  "segmentCount": 61,
  "fetchedAt": "2026-09-30T08:00:00.000Z"
}
```

With `outputFormat` `srt` or `vtt`, `segments` is replaced by `subtitles` holding the file content; with `text`, only `text` is returned.
`text` is always present.

A failed item:

```json
{
  "status": "error",
  "input": "https://www.youtube.com/shorts/xxxxxxxxxxx",
  "videoId": "xxxxxxxxxxx",
  "error": {
    "code": "VideoUnavailable",
    "message": "The video is no longer available: xxxxxxxxxxx",
    "retryable": false
  },
  "fetchedAt": "2026-09-30T08:00:00.000Z"
}
```

#### Error codes

| Code | Meaning | Retryable |
| --- | --- | --- |
| `InvalidVideoId` | The input is not a YouTube video ID or link. | no |
| `VideoUnavailable` | The video does not exist or was removed. | no |
| `VideoUnplayable` | YouTube will not play it (private, processing, region-locked); the message has YouTube's reason. | no |
| `AgeRestricted` | Age-gated; transcripts need a signed-in account, which this Actor does not use. | no |
| `TranscriptsDisabled` | The uploader turned captions off. | no |
| `NoTranscriptFound` | None of your languages exist and `fallbackToAnyLanguage` is `false`; `error.available` lists the ones that do. | no |
| `NotTranslatable` / `TranslationLanguageNotAvailable` | YouTube does not offer that translation; `error.available` lists options. | no |
| `RequestBlocked` / `IpBlocked` | YouTube blocked the request as automated traffic. | yes |
| `PoTokenRequired` | YouTube demanded a proof-of-origin token for this track. | no |
| `FailedToCreateConsentCookie` | The EU cookie consent page could not be passed. | no |
| `YouTubeRequestFailed` | Network or YouTube server error after retries. | yes |
| `YouTubeDataUnparsable` | YouTube changed its response format; we monitor for this and fix it. | no |

### Pricing

Pay per event: **$4 per 1,000 transcripts** ($0.004 per `transcript` event, charged only for items with `status: "ok"`).
Failed videos, invalid inputs and duplicates are free.
Residential proxy traffic is covered by that price; you are not billed for it separately.
On the Apify free plan a run processes the first 10 videos and skips the rest (the run log and status message say so); any paid Apify plan runs up to 1000 videos per run.

### FAQ

**Do I need a YouTube API key?**
No.

**Does it work for videos without captions?**
No. Only captions YouTube already has are returned; there is no speech-to-text.

**Can I get channel or playlist transcripts?**
Not directly: pass the video URLs or IDs.
Up to 1000 videos per run.

**Which formats are supported?**
`segments` (timed JSON plus text), `text`, `srt` and `vtt`.

### Limits

- YouTube blocks most datacenter and cloud IPs, so the Actor is meant to run on the Apify platform with the default residential proxy. Turning the proxy off (or running it elsewhere) typically produces `RequestBlocked` or `IpBlocked` on anything beyond a few videos.
- Translation uses YouTube's own machine translation. YouTube currently rate-limits translated caption downloads for anonymous clients, so `translateTo` often returns `RequestBlocked`; requesting the original language is reliable.
- Age-restricted, private and members-only videos are not supported, because they require signing in.
- On the Apify free plan, runs are capped at 10 videos; upgrade to any paid plan for up to 1000 per run.
- Only captions YouTube already has are returned; there is no speech-to-text for videos without captions.

### Credits

The caption discovery approach is a TypeScript port of [youtube-transcript-api](https://github.com/jdepoix/youtube-transcript-api) by Jonas Depoix (MIT License).
See `NOTICE` for the license text.

# Actor input Schema

## `videos` (type: `array`):

YouTube video URLs or 11-character video IDs. Accepts watch, youtu.be, Shorts, embed and live links. Up to 1000 per run; duplicates are fetched and charged once.

## `languages` (type: `array`):

Language codes in priority order, for example \["en", "en-GB", "de"]. For each language, human-made captions are used before auto-generated ones, and a code also matches its regional variants ("en" matches "en-US"; an exact match wins).

## `fallbackToAnyLanguage` (type: `boolean`):

When none of your languages exist, return the transcript in the video's original spoken language (or the first available track) instead of a NoTranscriptFound error. The result's languageFallback field is true when this happened. Ignored when Translate to is set.

## `translateTo` (type: `string`):

Optional language code. When set and the video has no track in this language, YouTube's own machine translation of the best available track is returned. YouTube currently rate-limits translated downloads heavily, so expect RequestBlocked errors for some videos.

## `outputFormat` (type: `string`):

segments: timed segments plus the full text. text: full text only. srt / vtt: subtitle file content plus the full text.

## `maxConcurrency` (type: `integer`):

How many videos to fetch in parallel.

## `proxyConfiguration` (type: `object`):

Defaults to Apify residential proxy, which is what makes YouTube requests succeed from the Apify cloud (YouTube blocks most datacenter IPs). Each request leaves from a new IP and blocked videos are retried automatically. Residential proxy usage is included in the per-transcript price.

## Actor input object example

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "languages": [
    "en"
  ],
  "fallbackToAnyLanguage": true,
  "outputFormat": "segments",
  "maxConcurrency": 3,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
    ],
    "languages": [
        "en"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("openkrill/youtube-transcripts").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
    "languages": ["en"],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("openkrill/youtube-transcripts").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "languages": [
    "en"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call openkrill/youtube-transcripts --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,openkrill/youtube-transcripts"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/b5JCtzYHisq43bUe0/builds/LMz6lZSE6WdrkNxWQ/openapi.json
