# YouTube Transcript Scraper - Captions & Subtitles API (`tenfoldfleet/youtube-transcript-scraper`) Actor

Get the transcript of any YouTube video or download subtitles for a whole list: timestamped caption segments and full text in your language, manual or auto-generated, with title, channel, duration and views. Shorts and live links work. $3 per 1,000 videos; no captions or failures are free.

- **URL**: https://apify.com/tenfoldfleet/youtube-transcript-scraper.md
- **Developed by:** [Tenfold Fleet](https://apify.com/tenfoldfleet) (community)
- **Categories:** Videos, AI, Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 transcript extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does YouTube Transcript Scraper do?

**YouTube Transcript Scraper** downloads the **transcript (captions / subtitles) of any [YouTube](https://www.youtube.com) video** in bulk. Give it video URLs or IDs and it returns **timestamped caption segments** and the **full transcript text**, plus the video's **title, channel, duration, view count and description**.

It works with every link form: `youtube.com/watch?v=`, `youtu.be/`, **YouTube Shorts**, embeds, live-stream recordings and bare 11-character video IDs. Pick your **preferred languages** (e.g. `en`, `es`, `de`), choose whether **auto-generated captions** are allowed, and get clean JSON you can feed straight into an LLM, a RAG pipeline or a spreadsheet.

It is a fast **YouTube transcript API** that runs on plain HTTP requests (no browser), so it's cheap and quick: **$3 per 1,000 transcripts**, and you **only pay when a non-empty transcript is returned**.

### Why use a YouTube transcript scraper?

- 🤖 **AI and LLM workflows**: summarize videos, answer questions about them, or build a RAG knowledge base from a channel's back catalogue.
- ✍️ **Content repurposing**: turn videos into blog posts, newsletters, show notes and social threads.
- 🔎 **Research and monitoring**: search what creators, competitors or brands actually say on camera.
- 🎓 **Education and accessibility**: get readable text and subtitles for lectures and tutorials.
- 📈 **SEO and marketing**: mine video transcripts for keywords and topics.

Runs on the Apify platform, so you get an **API**, scheduling, webhooks and integrations with Make, Zapier, n8n, LangChain and LlamaIndex. It also works as an **MCP tool for AI agents** through the [Apify MCP server](https://mcp.apify.com).

### What data can it extract?

| Field | Example |
|---|---|
| `input` | the URL or ID exactly as you entered it |
| `videoId`, `url` | `dQw4w9WgXcQ`, `https://www.youtube.com/watch?v=dQw4w9WgXcQ` |
| `title`, `channel`, `channelId` | video title and channel name / ID |
| `durationSec`, `viewCount` | `213`, `1821858729` |
| `description` | the video description |
| `language`, `languageName` | `en`, `English` |
| `isAutoGenerated` | `true` for YouTube's speech-recognition captions |
| `availableLanguages` | every caption track on the video: `[{ "languageCode": "de-DE", "name": "German (Germany)", "isAutoGenerated": false }]` |
| `segments` | `[{ "start": 18.64, "dur": 3.24, "text": "We're no strangers to love" }]` (seconds) |
| `fullText` | the whole transcript as one string |
| `segmentCount` | number of caption segments |
| `error` | why a video failed (no captions, private, removed...), otherwise `null` |

### How to get the transcript of a YouTube video

1. Click **Try for free**.
2. Paste one or more YouTube links or video IDs into **YouTube video URLs or IDs** (one per line).
3. Set **Preferred languages** (default `en`). If a video has no captions in those languages, you get its original-language transcript.
4. Click **Start**, then download the results as **JSON, CSV, Excel or HTML**, or read them through the API.

#### Call it from the API

```bash
curl -X POST "https://api.apify.com/v2/acts/tenfoldfleet~youtube-transcript-scraper/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"], "languages": ["en"], "outputFormat": "text"}'
```

Python (`pip install "apify-client>=3"`):

```python
from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("tenfoldfleet/youtube-transcript-scraper").call(run_input={"videos": ["https://youtu.be/jNQXAC9IVRw"]})
for video in client.dataset(run.default_dataset_id).iterate_items():
    print(video["title"], video.get("language"), video.get("error") or (video.get("fullText") or "")[:300])
```

On apify-client 2.x, `call()` returns a dict: use `run["defaultDatasetId"]` instead.

### How much does it cost?

You pay **$0.003 per transcript** (**$3 per 1,000 videos**). Videos **without captions**, **private or removed videos**, invalid links and any other failures are **free**: they appear in the dataset with an `error` field and are never charged. Apify's free plan includes monthly credit, so you can try it on hundreds of videos at no cost. Set a **maximum cost per run** and the Actor stops before it goes over that amount.

### Input

```json
{
    "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://youtu.be/jNQXAC9IVRw",
        "https://www.youtube.com/shorts/aqz-KE-bpKQ",
        "kJQP7kiw5Fk"
    ],
    "languages": ["en", "es"],
    "includeAutoGenerated": true,
    "outputFormat": "both"
}
```

- `outputFormat`: `both` (segments + full text), `segments` or `text`.
- `languages`: leave empty to always get the video's original language.
- `strictLanguage`: set to `true` to get a free error item instead of another language's transcript when none of your `languages` exists.

### Output

One item per video. You can download the dataset as **JSON, CSV, Excel, XML or HTML**.

```json
{
    "input": "https://youtu.be/jNQXAC9IVRw",
    "videoId": "jNQXAC9IVRw",
    "url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
    "title": "Me at the zoo",
    "channel": "jawed",
    "channelId": "UC4QobU6STFB0P71PMvOGN5A",
    "durationSec": 19,
    "viewCount": 438706382,
    "language": "en",
    "languageName": "English",
    "isAutoGenerated": false,
    "availableLanguages": [{ "languageCode": "en", "name": "English", "isAutoGenerated": false }, { "languageCode": "de", "name": "German", "isAutoGenerated": false }],
    "segmentCount": 6,
    "segments": [
        { "start": 1.2, "dur": 2.16, "text": "All right, so here we are, in front of the elephants" }
    ],
    "fullText": "All right, so here we are, in front of the elephants ...",
    "error": null
}
```

### Tips

- **Shorts, embeds and live-stream links** are accepted as they are. Many Shorts and live recordings have no captions, though; those come back as free error items. Streams that are live right now are skipped for free.
- Use `outputFormat: "text"` for LLM prompts: smaller output, faster downloads.
- **Large lists**: the Actor handles about 4 videos per second, so the default 1-hour timeout covers roughly 10,000 videos. For bigger lists, raise the run timeout or split the list. If a run times out, **resurrect** it: videos already in the dataset are skipped, never charged twice.
- Region variants match automatically: asking for `pt` also finds `pt-BR`, and `en` finds `en-GB`.
- Turn **Include auto-generated captions** off if you only want creator-uploaded subtitles.

### FAQ

#### Does it work for videos without subtitles?

Only if YouTube has generated automatic captions, which it does for most spoken videos. If a video has no captions at all (music without lyrics, very new uploads, some live streams), you get a free item with an `error` explaining why.

#### Can I get the transcript in a specific language?

Yes. List language codes in **Preferred languages** in order. The Actor picks human-made captions first, then auto-generated ones, and falls back to the original language if none of your languages exist. `availableLanguages` shows every track the video has.

#### Can I get every transcript from a channel or playlist?

Not directly yet. Paste the URLs of the individual videos. A channel or playlist URL returns a free error item.

#### Do I need a YouTube account, API key or cookies?

No. It reads publicly available caption data. No login, no YouTube Data API quota.

#### Can it scrape private, members-only or age-restricted videos?

No. Those need a signed-in account, so they return a free error item.

#### Is it legal to scrape YouTube transcripts?

The Actor only reads publicly available caption data. You are responsible for how you use the transcripts: respect copyright and YouTube's Terms of Service, and don't republish creators' content without permission.

This Actor is not affiliated with, endorsed or sponsored by YouTube or Google.

### More tools from Tenfold Fleet

| Actor | Price |
|---|---|
| [ATS Jobs Scraper - Greenhouse, Lever, Ashby & 5 More](https://apify.com/tenfoldfleet/ats-jobs-scraper) | $2 per 1,000 (job posting) |
| [Website Contact Scraper - Emails, Phones & Socials](https://apify.com/tenfoldfleet/company-contact-finder) | $8 per 1,000 (website with contacts) |
| [Website Technology Detector - Wappalyzer Alternative](https://apify.com/tenfoldfleet/tech-stack-detector) | $10 per 1,000 (website analyzed) |
| [URL to Markdown - Web Page to LLM-Ready Text for AI Agents](https://apify.com/tenfoldfleet/url-to-markdown) | $1 per 1,000 (page converted) |

# Actor input Schema

## `videos` (type: `array`):

Videos to get transcripts for. Any form works: watch URLs, youtu.be links, Shorts, embeds, live recordings or bare 11-character video IDs. One video = one result. Channel and playlist URLs are not supported.

## `languages` (type: `array`):

Caption language codes in order of preference, e.g. en, es, de, pt-BR. If none is available, the video's original-language transcript is returned (turn on "Only my preferred languages" to skip those videos for free instead). Leave empty to always get the original language.

## `includeAutoGenerated` (type: `boolean`):

Use YouTube's automatic (speech recognition) captions when no human-made captions exist in your language. Turn off to get only creator-uploaded subtitles.

## `strictLanguage` (type: `boolean`):

When a video has none of your preferred languages, return a free error item instead of its transcript in another language.

## `outputFormat` (type: `string`):

Timestamped segments, plain full text, or both.

## `maxConcurrency` (type: `integer`):

How many videos to process in parallel. Limited to 1 per 10 MB of run memory (25 at the default 256 MB).

## Actor input object example

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://youtu.be/jNQXAC9IVRw",
    "kJQP7kiw5Fk"
  ],
  "languages": [
    "en"
  ],
  "includeAutoGenerated": true,
  "strictLanguage": false,
  "outputFormat": "both",
  "maxConcurrency": 20
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://youtu.be/jNQXAC9IVRw",
        "kJQP7kiw5Fk"
    ],
    "languages": [
        "en"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tenfoldfleet/youtube-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://youtu.be/jNQXAC9IVRw",
        "kJQP7kiw5Fk",
    ],
    "languages": ["en"],
}

# Run the Actor and wait for it to finish
run = client.actor("tenfoldfleet/youtube-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://youtu.be/jNQXAC9IVRw",
    "kJQP7kiw5Fk"
  ],
  "languages": [
    "en"
  ]
}' |
apify call tenfoldfleet/youtube-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tenfoldfleet/youtube-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/jz5LbnKZvh14axwgH/builds/d4Qy6EkBRFncr0qza/openapi.json
