# YouTube Transcript Scraper (`sturdyscrape/youtube-transcript-scraper`) Actor

Get YouTube transcripts, captions and subtitles as plain text and timestamped segments, in the language you choose. Bulk URLs or IDs, Shorts included. Ready for AI, RAG and content workflows. Pay only for transcripts delivered.

- **URL**: https://apify.com/sturdyscrape/youtube-transcript-scraper.md
- **Developed by:** [Bernardo Henriques Sousa](https://apify.com/sturdyscrape) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.40 / 1,000 transcripts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does YouTube Transcript Scraper do?

YouTube Transcript Scraper gets the **transcript (captions and subtitles) of YouTube videos** and returns it as clean plain text and as timestamped segments. Paste video URLs or IDs, pick the languages you want, and get one row per video, ready for AI summaries, RAG pipelines, SEO content, subtitles, research or search.

- **Bulk:** hundreds of videos in one run, watch links, `youtu.be` links, Shorts, embeds, live replays or bare video IDs.
- **Language control:** choose languages in order of preference; captions written by a person win over auto-generated ones. If the video has none of them, you get the language it does have, and every available caption track is listed.
- **Reliable from the cloud:** YouTube blocks cloud servers, so the Actor routes every request through Apify Proxy and switches to residential IPs automatically when a request is blocked. You don't configure any proxy.
- **Fair pricing:** you pay only for transcripts delivered. Videos without captions, private videos and invalid links are reported in the dataset for free.

### Why scrape YouTube transcripts?

- **AI and LLM apps:** feed video content to ChatGPT, Claude or your own model for summaries, Q\&A and RAG knowledge bases.
- **Content repurposing:** turn videos into blog posts, newsletters, social posts and show notes.
- **SEO and research:** analyze what competitors and creators say, find keywords and quotes, study a topic across many videos.
- **Accessibility and subtitles:** reuse timestamped captions in your own player or editor.

### How to use YouTube Transcript Scraper

1. Click **Try for free**.
2. Paste one or more YouTube video URLs or IDs in **YouTube videos**.
3. Optional: set **Preferred languages** (for example `en`, `pt`, `es`).
4. Click **Start** and download the results as JSON, CSV, Excel or HTML, or get them through the API.

### Input

| Field | What it does | Default |
|---|---|---|
| `videos` | Video URLs or 11-character IDs | required |
| `languages` | Language codes in order of preference | `["en"]` |
| `fallbackToAnyLanguage` | If none of the preferred languages exists, return the language the video has | `true` |
| `includeSegments` | Also return timestamped segments | `true` |
| `maxConcurrency` | Videos fetched at the same time | `5` |

```json
{
    "videos": ["https://www.youtube.com/watch?v=aircAruvnKk", "dQw4w9WgXcQ"],
    "languages": ["en", "es"],
    "includeSegments": true
}
```

### Output

One item per video:

```json
{
    "videoId": "aircAruvnKk",
    "videoUrl": "https://www.youtube.com/watch?v=aircAruvnKk",
    "languageCode": "en",
    "language": "English",
    "isGenerated": false,
    "text": "This is a 3. It's sloppily written and rendered at an extremely low resolution of 28x28 pixels, but your brain has no trouble...",
    "wordCount": 3357,
    "segmentCount": 286,
    "durationSecs": 1105.6,
    "segments": [
        { "start": 4.22, "duration": 1.18, "text": "This is a 3." },
        { "start": 6.06, "duration": 4.653, "text": "It's sloppily written and rendered at an extremely low resolution of 28x28 pixels," }
    ],
    "availableLanguages": [
        { "languageCode": "ar", "language": "Arabic", "isGenerated": false },
        { "languageCode": "en", "language": "English", "isGenerated": false }
    ],
    "fetchedAt": "2026-10-10T20:51:59.030248+00:00",
    "error": null,
    "errorType": null
}
```

When a transcript can't be fetched, the item has `text: null` and explains why in `error` and `errorType` (for example `TranscriptsDisabled`, `NoTranscriptFound`, `VideoUnavailable`, `AgeRestricted`, `IpBlocked`). These items are free.

### How much does it cost?

The Actor uses pay-per-event pricing: you pay a small fee per **transcript delivered**, and nothing for failed videos. The current price is shown on the **Pricing** tab. Set a maximum cost per run in the run options and the Actor stops before going over it.

### Tips

- Auto-generated captions (`isGenerated: true`) are produced by YouTube's speech recognition and may contain mistakes.
- Videos with captions turned off by the creator have no transcript; the Actor can't create one.
- For very large batches, split them across several runs or raise **Max concurrency**.

### Is it legal to scrape YouTube transcripts?

The Actor reads only public captions that YouTube shows to any viewer, and stores no personal data. You are responsible for how you use the content: respect copyright and YouTube's terms, and check with a lawyer if you're unsure about your use case.

### Feedback

Found a video that fails, or need another field? Open an issue on the **Issues** tab and it will be looked at.

# Actor input Schema

## `videos` (type: `array`):

Video URLs or 11-character video IDs. Works with watch, youtu.be, Shorts, embed and live links. Duplicates are fetched once.

## `languages` (type: `array`):

Language codes in order of preference (for example en, pt, es). The first one the video has wins; captions written by a person are preferred over auto-generated ones.

## `fallbackToAnyLanguage` (type: `boolean`):

If the video has none of the preferred languages, return the transcript it does have instead of an error.

## `includeSegments` (type: `boolean`):

Also return the transcript split into segments with start time and duration in seconds. The full text is always returned.

## `maxConcurrency` (type: `integer`):

How many videos are fetched at the same time.

## Actor input object example

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://youtu.be/aircAruvnKk"
  ],
  "languages": [
    "en"
  ],
  "fallbackToAnyLanguage": true,
  "includeSegments": true,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://youtu.be/aircAruvnKk"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("sturdyscrape/youtube-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://youtu.be/aircAruvnKk",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("sturdyscrape/youtube-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://youtu.be/aircAruvnKk"
  ]
}' |
apify call sturdyscrape/youtube-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,sturdyscrape/youtube-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ousmBTFEjAl0rEkjd/builds/YCmiqhzPWIfPZtf94/openapi.json
