# YouTube Transcript Scraper - Bulk Videos & Shorts to Text (`eiv/youtube-transcript-probe`) Actor

Get the transcript of any YouTube video or Short from its URL. Returns the full text with timed segments, language, title and channel. Export as plain text, timestamped text, SRT or VTT subtitles. Bulk URLs, manual and auto-generated captions. You pay only for transcripts delivered.

- **URL**: https://apify.com/eiv/youtube-transcript-probe.md
- **Developed by:** [Eimantas V](https://apify.com/eiv) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.50 / 1,000 transcript delivereds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Transcript Scraper

Get the transcript of any YouTube video or Short from its URL. Paste one link or a thousand, and you get the full text back with timestamps, the caption language, the video title and the channel. Download it as plain text, timestamped text, or ready-to-use **SRT** and **VTT** subtitle files.

No YouTube account, no API key, no browser extension. You pay only for transcripts that are actually delivered.

### What can you use it for?

- **Content repurposing** – turn videos into blog posts, newsletters, social posts or show notes.
- **AI and LLM workflows** – feed clean transcript text into ChatGPT, Claude or your own RAG pipeline for summaries, Q\&A and search.
- **Research and analysis** – study what is said across many videos or channels: topics, keywords, sentiment, brand mentions.
- **SEO** – publish transcripts alongside your videos to make them searchable.
- **Subtitles** – download SRT or VTT files to edit, translate or re-upload.
- **Accessibility and note-taking** – read a video instead of watching it.

### What data do you get?

For every video:

| Field | What it is |
|---|---|
| `text` | The full transcript, in the format you chose |
| `segments` | Every caption line with its start time and duration in seconds |
| `title`, `channel`, `channelId` | Who published the video and what it is called |
| `durationSeconds` | Length of the video |
| `language`, `languageName` | The language of the returned captions, e.g. `en` / `English` |
| `isAutoGenerated` | `true` for YouTube's automatic captions, `false` for captions added by the creator |
| `availableLanguages` | Every caption language the video has, so you know what else you can request |
| `status` | `ok`, `no_captions` or `failed` |
| `error`, `note` | Why a video failed, or a note such as "the language you asked for isn't available" |
| `subtitleFileUrl` | Download link for the `.srt` / `.vtt` file, when you choose a subtitle format |

### How to get a YouTube transcript

1. Click **Try for free** (or **Start**).
2. Paste one or more YouTube links into **Videos**. Regular videos, Shorts, `youtu.be` short links, embed links and plain video IDs all work.
3. Optionally, set a **Preferred language** (for example `en`, `de` or `es`) and a **Transcript format**.
4. Click **Start**. Most videos finish in 1–2 seconds.
5. Open the **Output** tab to read the transcripts, or export them as JSON, CSV, Excel, HTML or XML.

### Input

| Setting | Description | Default |
|---|---|---|
| **Videos** | YouTube video URLs or IDs, one per line | – |
| **Preferred language** | Caption language code such as `en`. Creator-made captions are preferred over automatic ones. If the video doesn't have that language, you get its default captions and a `note` saying so. | The video's own default |
| **Transcript format** | Plain text, text with timestamps, SRT subtitles or WebVTT subtitles | Plain text |
| **Include timed segments** | Adds the line-by-line `segments` list. Turn it off to keep large exports small. | On |
| **Videos at once** | How many videos are processed in parallel. Raise it for big lists – the price per video stays the same. | 5 |
| **Proxy groups** | Advanced. Leave as it is. | Residential |

Example input:

```json
{
    "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://www.youtube.com/watch?v=jNQXAC9IVRw"
    ],
    "language": "en",
    "outputFormat": "text"
}
```

### Output example

```json
{
    "url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
    "videoId": "jNQXAC9IVRw",
    "title": "Me at the zoo",
    "channel": "jawed",
    "channelId": "UC4QobU6STFB0P71PMvOGN5A",
    "durationSeconds": 19,
    "availableLanguages": [
        { "code": "en", "name": "English", "isAutoGenerated": false },
        { "code": "de", "name": "German", "isAutoGenerated": false }
    ],
    "status": "ok",
    "error": null,
    "note": null,
    "language": "en",
    "languageName": "English",
    "isAutoGenerated": false,
    "format": "text",
    "text": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say",
    "segments": [
        { "start": 1.2, "duration": 2.16, "text": "All right, so here we are, in front of the elephants" },
        { "start": 5.318, "duration": 2.656, "text": "the cool thing about these guys is that they have really..." },
        { "start": 7.974, "duration": 4.642, "text": "really really long trunks" }
    ]
}
```

### Transcript formats

**Plain text** – the whole transcript as one block, ideal for reading or for AI tools:

```
All right, so here we are, in front of the elephants the cool thing about these guys is...
```

**Timestamped text** – one line per caption, with the time it's spoken:

```
[0:01] All right, so here we are, in front of the elephants
[0:05] the cool thing about these guys is that they have really...
[0:07] really really long trunks
```

**SRT** and **VTT** – standard subtitle files that work in YouTube Studio, video editors and media players. Each file is also saved to the run's storage and linked from `subtitleFileUrl`, so you can download it directly.

```
1
00:00:01,200 --> 00:00:03,360
All right, so here we are, in front of the elephants
```

### How much does it cost?

You pay per transcript delivered – see the **Pricing** tab for the exact price on your plan. Higher Apify plans get a lower price per transcript.

You are **not** charged for:

- videos that have no captions,
- links that aren't YouTube videos,
- videos that couldn't be fetched.

Platform usage is included in the price. If you set a maximum cost per run, the Actor stops starting new videos once that limit is reached, so you never pay more than you set.

### Use it through the API

Run the Actor from your own code, or connect it to Make, Zapier, n8n, Google Sheets and more through Apify integrations.

**JavaScript**

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('eiv/youtube-transcript-scraper').call({
    videos: ['https://www.youtube.com/watch?v=dQw4w9WgXcQ'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].text);
```

**Python**

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("eiv/youtube-transcript-scraper").call(
    run_input={"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"]}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["text"])
```

### Good to know

- **The transcript comes from the video's captions.** Most videos have automatic captions, but some don't – for example music-only videos, very new uploads, or videos where the creator turned captions off. Those come back as `no_captions` and are free.
- **Automatic captions are as good as YouTube's speech recognition.** YouTube hides swear words in them as `[ __ ]`, and this Actor returns them exactly as YouTube provides them.
- **No translation.** You get the languages the video already has captions in. `availableLanguages` lists them.
- **Private, members-only and age-restricted videos** can't be transcribed and come back as `failed` with the reason.
- **Results arrive in the order they finish**, not always the order you entered them. Use the `url` or `videoId` field to match them up.

### FAQ

**Can I transcribe a whole channel or playlist?**
Not directly yet – paste the video links. Channel and playlist support is planned.

**Does it work with YouTube Shorts?**
Yes. Paste the Shorts link as it is.

**Can I get the transcript in a specific language?**
Yes, if the video has captions in that language. Set **Preferred language**. If it isn't available, you get the video's default captions and a `note` explaining why.

**Is it legal to scrape YouTube transcripts?**
This Actor only reads captions that YouTube shows publicly on the video page. It doesn't access private data. You are responsible for how you use the transcripts, including copyright – when in doubt, ask for legal advice.

**Something doesn't work?**
Open an issue on the **Issues** tab with the video link and the run ID, and it will be looked at.

# Actor input Schema

## `videos` (type: `array`):

YouTube video URLs or ids - watch, youtu.be, shorts and embed links all work. One output item per video.

## `language` (type: `string`):

Caption language code, e.g. "en" or "de". Manual captions are preferred over auto-generated ones in that language. Left empty, or when the video has no such captions, you get the video's own default track.

## `outputFormat` (type: `string`):

What the "text" field holds. SRT and VTT are also saved as files in the run's key-value store, linked from "subtitleFileUrl".

## `includeSegments` (type: `boolean`):

Adds "segments": every caption line with its start and duration in seconds. Turn off to keep the output of large runs small.

## `maxConcurrency` (type: `integer`):

How many videos are fetched in parallel. Higher is faster for big lists; the cost per video is the same.

## `groups` (type: `array`):

Tried in order for each video until one serves the transcript. Empty string means the account default (datacenter), which YouTube currently answers with a bot check. UNBLOCKER works too but costs ~15x more per video.

## Actor input object example

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://www.youtube.com/watch?v=jNQXAC9IVRw"
  ],
  "outputFormat": "text",
  "includeSegments": true,
  "maxConcurrency": 5,
  "groups": [
    "RESIDENTIAL",
    ""
  ]
}
```

# Actor output Schema

## `transcripts` (type: `string`):

One item per video: the transcript, its language, title, channel and timed segments.

## `subtitleFiles` (type: `string`):

SRT or VTT files, one per video, when that output format is chosen.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://www.youtube.com/watch?v=jNQXAC9IVRw"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("eiv/youtube-transcript-probe").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://www.youtube.com/watch?v=jNQXAC9IVRw",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("eiv/youtube-transcript-probe").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://www.youtube.com/watch?v=jNQXAC9IVRw"
  ]
}' |
apify call eiv/youtube-transcript-probe --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,eiv/youtube-transcript-probe"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/rMUcuxCStdIbkelkf/builds/vBfNtJqt85CcugErE/openapi.json
