# YouTube Transcript Scraper | Text + Timestamps, Any Language (`akatra/youtube-transcript-scraper`) Actor

Get the transcript of any YouTube video as clean text and timestamped lines: manual captions or auto-generated ones, in the language you choose, plus title, channel, length and views. Paste video URLs or IDs, including Shorts. You only pay for videos that have a transcript.

- **URL**: https://apify.com/akatra/youtube-transcript-scraper.md
- **Developed by:** [Akatra](https://apify.com/akatra) (community)
- **Categories:** Videos, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.80 / 1,000 video results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Transcript Scraper: text and timestamps, any language

Turn YouTube videos into text. Paste video URLs or IDs and get each video's transcript as **clean full text** and as **timestamped lines**, together with the title, channel, length and view count. It reads captions written by the uploader and YouTube's auto-generated captions, in the language you prefer.

**By default you only pay for videos that have a transcript.** Videos without captions, private or deleted videos are reported in the log and not charged. If you turn on "Also return videos without a transcript", those videos are added to the results as rows of their own and are charged like any other result.

**Pay per result** · **No login, no API key** · **Public data only** · **Tested daily** · **LLM-ready JSON for AI agents (MCP)**

### 🛠️ What does YouTube Transcript Scraper do?

- 📝 **Clean full text and timestamped lines** for every video, with title, channel, length and view count.
- 🌐 **Any language.** Reads captions written by the uploader and YouTube's auto-generated captions, in the language you prefer.
- 💸 **By default you pay only for videos that have a transcript.** Videos without captions, private or deleted videos are not charged.
- 🛡️ **Built to get through blocking.** Requests go out with a real browser fingerprint through rotating proxies, and a blocked request is retried from a fresh IP. A result that could not be completed is never charged.

#### What people use it for

- **AI and RAG pipelines.** Feed video content into LLMs, summarizers, vector stores and agents.
- **Content repurposing.** Turn talks, podcasts and tutorials into articles, newsletters and social posts.
- **Research and analysis.** Search and analyze what was said across hundreds of videos.
- **SEO.** Mine competitors' videos for topics, keywords and questions.
- **Subtitles and translation workflows.** Get timestamped lines ready to convert to SRT or VTT.

### 📊 What data can you extract?

Every result is one JSON object (one row in CSV or Excel) with these fields:

| Field | Meaning |
|---|---|
| `videoId`, `url` | The video |
| `title`, `channel`, `channelId`, `lengthSeconds`, `viewCount` | Basic video details |
| `isLiveContent` | `true` for a live stream or the recording of one |
| `transcript` | The whole transcript as one text |
| `segments` | Every caption line with its `start` time and `duration` in seconds (when timestamps are on) |
| `wordCount` | Number of words in the transcript |
| `language`, `languageName` | The language of the returned transcript |
| `isAutoGenerated` | `true` when the captions were generated by YouTube's speech recognition, `false` when the uploader provided them |
| `availableLanguages` | Every caption language the video has |
| `hasTranscript`, `unavailableReason` | Only relevant when you turn on "Also return videos without a transcript": why a video has none |

### 🚀 How to use YouTube Transcript Scraper

1. Create a free [Apify](https://apify.com) account and open this Actor.
2. Add your **videos**, one per line: full URLs, `youtu.be` links, Shorts links, embed links or plain video IDs.
3. Set your **preferred languages** in order, for example `en`, `es`. The first language the video has is used.
4. Decide whether to **use another language** when none of yours exist, and whether you need **timestamped lines**.
5. Run it and download the results as JSON, CSV or Excel, or connect them through the API.

### 📥 Input example (JSON)

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=8jPQjjsBbIc",
    "https://youtu.be/dQw4w9WgXcQ",
    "https://www.youtube.com/shorts/ZZ5LpwO-An4"
  ],
  "languages": ["en", "es"],
  "fallbackToAnyLanguage": true,
  "includeTimestamps": true,
  "includeVideosWithoutTranscript": false
}
```

### 📤 Sample output (JSON)

![Sample output table of YouTube Transcript Scraper](https://api.apify.com/v2/key-value-stores/GkJnrNz6YWKHQYnZr/records/youtube-transcript-scraper-output.png)

One result per video. The lists are shortened here.

```json
{
  "videoId": "dQw4w9WgXcQ",
  "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
  "channel": "Rick Astley",
  "channelId": "UCuAXFkgsw1L7xaCfnd5JJOw",
  "lengthSeconds": 213,
  "viewCount": 1822189128,
  "isLiveContent": false,
  "hasTranscript": true,
  "language": "en",
  "languageName": "English",
  "isAutoGenerated": false,
  "availableLanguages": [
    { "code": "en", "name": "English", "autoGenerated": false },
    { "code": "en", "name": "English (auto-generated)", "autoGenerated": true }
  ],
  "transcript": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ ...",
  "wordCount": 487,
  "segments": [
    { "start": 18.64, "duration": 3.24, "text": "♪ We're no strangers to love ♪" }
  ],
  "unavailableReason": null,
  "scrapedAt": "2026-10-02T00:30:00+00:00"
}
```

### 🤖 Use it with AI agents, MCP and the API

The output is clean JSON with stable field names, ready for LLM pipelines, RAG and agent tools without extra parsing.

- **AI agents and MCP (Claude, ChatGPT, Cursor and others).** Add the Apify MCP server with this Actor as a tool: `https://mcp.apify.com?tools=akatra/youtube-transcript-scraper`. The agent can then run it and read the results by itself.
- **API.** Start a run and get the results in a single call:

```bash
curl -X POST "https://api.apify.com/v2/acts/akatra~youtube-transcript-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"videos": ["https://www.youtube.com/watch?v=8jPQjjsBbIc", "https://youtu.be/dQw4w9WgXcQ", "https://www.youtube.com/shorts/ZZ5LpwO-An4"], "languages": ["en", "es"], "fallbackToAnyLanguage": true, "includeTimestamps": true, "includeVideosWithoutTranscript": false}'
```

- **No-code tools and SDKs.** Works with Make, n8n, Zapier and LangChain through Apify's integrations, and with the Apify clients for Python and JavaScript. Schedule it and send the results to a webhook, Google Sheets or your own database.

### 💰 Pricing: pay per result

You pay per result. By default a result is a video with a transcript; with "Also return videos without a transcript" on, every video you submit becomes a result. Nothing is charged for empty runs beyond a tiny start fee. See the price on the Pricing tab. The price is the same for a one-minute Short and a three-hour podcast. There is no subscription or rental fee. New to Apify? The free plan includes $5 of platform credit every month, enough for about 1,200 videos with this Actor, and no credit card is needed.

### 🚦 Run status and error messages

Every run ends with a plain status message, so automated workflows (API, Make, n8n, AI agents) can tell what happened without reading the log:

| Situation | What you get |
|---|---|
| Finished normally | `Saved N videos.` |
| No video has a transcript | The run succeeds with `No transcripts saved. The videos have no captions in the requested languages, or are unavailable.` You pay only the start fee. |
| Some videos have no transcript | `N videos have no transcript and were not charged.` (With 'Also return videos without a transcript' on, they are returned and charged.) |
| YouTube kept blocking some videos | `Skipped N videos whose details stayed blocked (not charged).` |
| Missing or unsupported input | The run fails at once and the log names the field to fix. Nothing is scraped. |
| Your maximum charge is reached | Stops cleanly with `Stopped at your maximum charge limit.` Everything saved so far stays in the dataset. |

### ℹ️ Good to know

- Captions written by the uploader are preferred over auto-generated ones in the same language. A request for `en` also accepts regional variants such as `en-US` or `en-GB`. Different scripts are never mixed up: `zh-Hans` does not return `zh-Hant`.
- One run handles up to 20,000 videos. Start another run for more.
- Videos with no captions at all cannot be transcribed by this scraper; it reads YouTube's captions and does not run speech recognition itself.
- Private, deleted and age-restricted videos, and live streams that are still running, have no transcript available without logging in.
- A video that stays blocked after all retries is skipped and not charged. The run's status message tells you how many were skipped, so you can run them again.
- If you set a maximum charge for a run, the scraper stops as soon as that limit is reached.
- Turn off timestamped lines when you only need the text: results get several times smaller.
- No personal data beyond the public channel name is collected.

### 📝 Changelog

- **2026-10-02** README restructured: data table, API and MCP examples, run status messages.
- **2026-10-02** Better language matching (writing system, manual captions first) and clearer reasons for unavailable videos, after an independent audit.
- **2026-10-02** First release: full text and timestamped lines, any language, videos without a transcript not charged.

### ⚖️ Legal and privacy

This scraper is an independent tool and is not affiliated with, endorsed by, or sponsored by YouTube or Google. "YouTube" is a trademark of its owner and is used here only to show which website the tool works with. Videos and their transcripts belong to their creators. Use the data in line with the laws that apply to you and the source site's terms.

The source site's name and logo are trademarks of their owner and appear here only to identify the website this tool works with. The Actor reads only pages that are public without logging in. How you store and use the data is your responsibility: follow the laws that apply to you, such as GDPR and CCPA.

### 🔗 More scrapers from Akatra

- [Indeed Jobs Scraper](https://apify.com/akatra/indeed-jobs-scraper): 💼 job ads from Indeed in 30 countries, with yearly salary and only-new mode.
- [Google News Scraper](https://apify.com/akatra/google-news-scraper): 📰 news by keyword or topic in 50 editions, with the real article URLs.
- [Google Trends Scraper](https://apify.com/akatra/google-trends-scraper): 📈 interest over time, by region, and rising related queries.
- [Google Ads Transparency Scraper](https://apify.com/akatra/google-ads-transparency-scraper): 🔎 every ad a company runs on Google, with regions and dates.

### 💬 Support

Found a bug or need a field that is missing? Open an issue on the Issues tab and we usually reply within a day.

# Actor input Schema

## `videos` (type: `array`):

YouTube video URLs or video IDs, one per line. Regular videos, Shorts, youtu.be links and embed links all work. Up to 20,000 videos per run.

## `languages` (type: `array`):

Language codes in order of preference, for example en, es, de. The first language the video has is used. Captions written by the uploader are preferred over auto-generated ones in the same language.

## `fallbackToAnyLanguage` (type: `boolean`):

When the video has no captions in your preferred languages, return its captions in whatever language it has. Turn off to skip such videos.

## `includeTimestamps` (type: `boolean`):

Add every caption line with its start time and duration. Turn off to get only the full text, which makes results much smaller.

## `includeVideosWithoutTranscript` (type: `boolean`):

By default videos that have no captions, or are unavailable, are left out and not charged. Turn on to get a row for every video, with the reason and the video details. These extra rows are charged like any other result.

## `proxyConfiguration` (type: `object`):

Apify Proxy is recommended. The scraper spreads requests over several IPs and switches to residential IPs only when YouTube keeps blocking a video.

## Actor input object example

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=8jPQjjsBbIc"
  ],
  "languages": [
    "en"
  ],
  "fallbackToAnyLanguage": true,
  "includeTimestamps": true,
  "includeVideosWithoutTranscript": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `transcripts` (type: `string`):

Every video with its transcript text, shown in the overview table.

## `transcriptsFull` (type: `string`):

Every video with all fields, including timestamped lines.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videos": [
        "https://www.youtube.com/watch?v=8jPQjjsBbIc"
    ],
    "languages": [
        "en"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("akatra/youtube-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "videos": ["https://www.youtube.com/watch?v=8jPQjjsBbIc"],
    "languages": ["en"],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("akatra/youtube-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videos": [
    "https://www.youtube.com/watch?v=8jPQjjsBbIc"
  ],
  "languages": [
    "en"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call akatra/youtube-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,akatra/youtube-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HboYKcBXhtz0IO9Bv/builds/C6FpttdataqhgA8cM/openapi.json
