# YouTube Transcripts & Metadata Scraper (`casaucao/youtube-transcript-scraper`) Actor

Extract reliable batch transcripts and metadata from YouTube videos, Shorts, channels and playlists. Multi-language, AI-ready output (text, markdown, SRT, VTT). Built for RAG, LLM and agent pipelines.

- **URL**: https://apify.com/casaucao/youtube-transcript-scraper.md
- **Developed by:** [Eric Z. Casaucao](https://apify.com/casaucao) (community)
- **Categories:** Videos, Social media, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 transcript extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**YouTube transcript API, captions and subtitles** — extract clean, timestamped transcripts and
metadata from YouTube videos, Shorts, channels and playlists, reliably and at scale. Output as
text, markdown, **SRT** or **VTT**, ready for **RAG, LLM and AI agent** pipelines.

### What does YouTube Transcript Scraper do?

It turns any public [YouTube](https://www.youtube.com) video into structured text: the full
**transcript segmented by timestamp**, plus video **metadata**. It works on single videos, full
**URL batches**, **Shorts**, entire **playlists** and the **latest videos of a channel**, in the
languages you choose — and it can **translate** captions when a language is missing.

It does **not** transcribe audio: it extracts captions/subtitles that already exist on the video
(auto-generated or uploaded). It's a reliable **YouTube transcript API alternative** that needs no
Google API key and no OAuth.

### Why scrape YouTube transcripts?

Video is where a huge amount of knowledge lives, but it isn't searchable, embeddable or
queryable. Transcripts make it all three.

- **Feed RAG and LLM pipelines.** Drop transcripts straight into a vector store, a LangChain or
  LlamaIndex loader, or an agent's context window. Each row is AI-ready with `segments`,
  `fullText` and `markdown`.
- **Repurpose video into text.** Turn webinars, interviews and tutorials into blog posts,
  newsletters, subtitles and clip scripts.
- **Research and analyse at scale.** Pull hundreds of videos on a topic and run topic
  classification, entity extraction, sentiment or search over the text.
- **Publish transcripts for SEO and accessibility.** Make video content indexable and provide a
  text alternative for viewers who need it.
- **Monitor channels and playlists.** Run on a schedule to keep a keyword or knowledge base fresh.

#### Built for reliability and scale

- Batch thousands of URLs in one run, with configurable concurrency.
- Per-item `status`/`error` — one bad video never breaks the run.
- Retries with exponential backoff and Apify Proxy support to avoid IP blocks.
- Multiple output formats: `segments`, `text`, `markdown`, `SRT`, `VTT`.

### What data can it extract?

Every video produces one dataset row:

| Field | Type | Description |
|---|---|---|
| `videoId` | string | YouTube video ID. |
| `title` | string | Video title. |
| `channelName` / `channelId` | string | Publishing channel. |
| `publishDate` | string | Publication date. |
| `durationSec` | number | Duration in seconds. |
| `viewCount` | number | View count at run time. |
| `tags` | array | Video keywords/tags. |
| `description` | string | Video description. |
| `thumbnailUrl` | string | Max-resolution thumbnail. |
| `language` | string | Transcript language (ISO 639-1). |
| `isGenerated` | boolean | Whether captions are auto-generated. |
| `translatedFrom` | string | Source language if translated. |
| `segments` | array | `{start, end, text}` timestamped segments. |
| `fullText` | string | Full transcript as one string. |
| `markdown` | string | Transcript as markdown with the video link. |
| `srt` / `vtt` | string | Subtitle formats (when requested). |
| `status` / `error` | string | Per-item outcome. |

### How to scrape YouTube transcripts (step-by-step)

1. Open the Actor in [Apify Console](https://console.apify.com) or call it via API.
2. Add **Video URLs** (watch/`youtu.be`/Shorts), or add **Channel URLs** / **Playlist URLs** to
   expand them automatically.
3. Set **Language priority** (e.g. `["en", "es"]`).
4. Choose the **Output format** (`segments`, `text`, `markdown`, `srt`, `vtt`).
5. Start the run. Watch the log and results in the **Output** tab.
6. Download the dataset as JSON, CSV or Excel, or read it via the API. Schedule the run to keep
   your data fresh.

### How much does it cost to scrape YouTube?

You pay **per event**, not per compute time. Prices:

- **$3 per 1,000 transcripts** (`transcript-extracted`)
- **$0.50 per 1,000 videos** of metadata (`metadata-enriched`)

Failed videos and videos without captions are **not charged**. On Apify's free plan you get
monthly credits to test, and this Actor caps free-plan runs to the first few videos, so you can
validate the output before paying.

Example: 1,000 videos with transcripts and metadata ≈ **$3.50**.

### Input

See the **Input** tab for full configuration. Key fields:

| Field | Type | Default | Description |
|---|---|---|---|
| `videoUrls` | array | — | Video/Shorts URLs (watch, youtu.be, shorts). |
| `channelUrls` | array | — | Channel URLs, expanded to their latest videos. |
| `playlistUrls` | array | — | Playlist URLs, expanded to their videos. |
| `languages` | array | `["en"]` | ISO 639-1 codes in priority order. |
| `includeAutoGenerated` | boolean | `true` | Fall back to auto-generated captions. |
| `outputFormat` | string | `segments` | `segments`, `text`, `markdown`, `srt`, `vtt`. |
| `includeMetadata` | boolean | `true` | Collect video metadata. |
| `maxItems` | integer | `0` | Max videos. 0 = unlimited (free plan capped). |
| `maxRetries` | integer | `3` | Per-request retries. |
| `concurrency` | integer | `5` | Parallel video workers. |
| `proxyConfiguration` | object | Apify Proxy | Apify Proxy (auto) by default; set Residential at high volume if blocked. |

### Output

One JSON object per video. Example:

```json
{
  "videoId": "dQw4w9WgXcQ",
  "title": "Rick Astley - Never Gonna Give You Up (Official Video)",
  "channelName": "Rick Astley",
  "publishDate": "2009-10-25",
  "durationSec": 213,
  "viewCount": 1820610046,
  "language": "en",
  "isGenerated": false,
  "segments": [{"start": 1.36, "end": 3.04, "text": "We're no strangers to love"}],
  "fullText": "We're no strangers to love ...",
  "status": "ok",
  "error": null
}
```

### Integrate (API / SDK / MCP)

```python
from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("youtube-transcript-scraper").call(run_input={
    "videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
    "languages": ["en", "es"],
    "outputFormat": "srt",
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["videoId"], item["status"])
```

This Actor is callable as a **tool for AI agents** through the Apify **MCP server**
(`https://mcp.apify.com`), so an agent can fetch a transcript on demand — no glue code.

### FAQ

#### Can I download subtitles as SRT or VTT?

Yes. Set `outputFormat` to `srt` or `vtt` and each row includes subtitles ready to use.

#### Does it work with YouTube Shorts and playlists?

Yes. Shorts URLs are accepted and playlist/channel URLs are expanded into their videos
automatically.

#### Can I get transcripts in another language?

Yes. Pass a language priority list; if the language isn't available, the transcript is
translated when possible (the `translatedFrom` field tells you the source language).

#### Does it transcribe videos without captions?

No. It extracts existing captions/subtitles; it does not perform speech-to-text.

#### How many videos can I process per run?

There is no fixed limit — batch as many URLs as you need, subject to `maxItems` and your plan.

### Disclaimers & support

Use this Actor only for lawful purposes and content you are permitted to process. Transcripts may
contain copyrighted material or personal data; keep source attribution and comply with YouTube's
terms and your local laws. For bug reports or feature requests, open an issue on the Actor's page.

# Changelog

This Actor's version history is a separate document: https://apify.com/casaucao/youtube-transcript-scraper/changelog.md

# Actor input Schema

## `videoUrls` (type: `array`):

YouTube video/Shorts URLs (watch, youtu.be, shorts).

## `channelUrls` (type: `array`):

Channel URLs; expanded to their latest videos.

## `playlistUrls` (type: `array`):

Playlist URLs; expanded to their videos.

## `languages` (type: `array`):

ISO 639-1 codes in priority order. If a requested language is absent, the transcript is translated to the first language.

## `includeAutoGenerated` (type: `boolean`):

Fall back to auto-generated (ASR) captions when no manual captions match the requested languages.

## `outputFormat` (type: `string`):

Primary transcript representation. `segments` also returns fullText and markdown; `srt`/`vtt` add subtitles.

## `includeMetadata` (type: `boolean`):

Collect per-video metadata: title, channel, publish date, duration, views, tags, description, thumbnail.

## `maxItems` (type: `integer`):

Maximum number of videos to process. 0 = unlimited. Free plan is capped.

## `maxRetries` (type: `integer`):

Maximum retries per HTTP request when blocked or rate-limited.

## `concurrency` (type: `integer`):

Number of videos processed in parallel.

## `proxyConfiguration` (type: `object`):

Proxy settings. Apify Proxy (auto) by default; set a RESIDENTIAL group at high volume if YouTube blocks requests.

## Actor input object example

```json
{
  "languages": [
    "en"
  ],
  "includeAutoGenerated": true,
  "outputFormat": "segments",
  "includeMetadata": true,
  "maxItems": 0,
  "maxRetries": 3,
  "concurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset of videos with their transcripts and metadata.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("casaucao/youtube-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("casaucao/youtube-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call casaucao/youtube-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,casaucao/youtube-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/k3xeFO6nkDOjD5AfH/builds/AZWTNGh1KvsdrdOdX/openapi.json
