# YouTube Transcript Scraper – Timestamped JSON for RAG (`leadsbrary/youtube-transcript-rag-feed`) Actor

Scrape YouTube transcripts with timestamps, chapters and video metadata as RAG-ready JSON chunks. Bulk video, playlist or channel URLs. No API key.

- **URL**: https://apify.com/leadsbrary/youtube-transcript-rag-feed.md
- **Developed by:** [Alexandre Manguis](https://apify.com/leadsbrary) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.56 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Transcript Scraper – Timestamped JSON for RAG

A **YouTube transcript scraper** built for RAG and LLM pipelines: it turns video, playlist or channel URLs into one flat JSON record per video with the full transcript, cue-level timestamped segments, embedding-sized text chunks with deep-link timestamp URLs, chapters and normalized video metadata. From $0.56 per 1,000 transcripts.

### What it does

The Actor resolves each input (video URL, playlist URL, channel URL or bare video id) to its video ids, then reads the same public watch-page payload and caption/`timedtext` endpoints a logged-out browser uses when you open YouTube's own transcript panel — no login, no API key, no paid subscription. For every video it returns the caption track (human-authored when available, auto-generated as a fallback) as a normalized transcript: full text, per-cue `start`/`end` timestamps, and merged RAG chunks sized to a configurable character budget with stable chunk ids and `&t=`-timestamp deep links for citation. Videos with no captions are returned as a successful, unbilled item with `transcriptStatus="no_captions"` instead of failing the run, so coverage gaps are explicit rather than hidden.

### Who it's for

- **RAG/LLM developers** who need video transcripts pre-chunked and timestamped for a vector database, with no custom parsing or chunking code to write.
- **AI agent and knowledge-base builders** who want citation-grade deep links back to the exact second a chunk came from.
- **Content, SEO and research teams** pulling spoken text and chapter titles from competitor or reference videos for topic analysis.
- **Pipeline operators** who need explicit per-video status (ok / no captions / unavailable) instead of a run that silently drops failures.

### Use cases

- Bulk-extract transcripts for a channel or playlist as embedding-sized chunks for a vector database.
- Build a citeable video knowledge base: every chunk carries `videoId`, `title`, `channelName` and a deep-link timestamp URL.
- Feed timestamped segments into summarization or clip-finding pipelines that map generated text back to a video position.
- Pull spoken text plus chapter titles from competitor videos for content/SEO research.
- Detect which videos in a list have no captions, so coverage gaps are visible instead of silent failures.

### Features

- **RAG chunks**, not just raw captions — `chunks[]` merges cues into blocks sized by `chunkMaxChars`, with configurable overlap, stable `chunkId`s and a `deepLinkUrl` per chunk.
- **Chapter-aware chunking** (`respectChapters`) — chunks never span two chapters when chapters are detected, and each chunk is tagged with its `chapterTitle`.
- **Cue-level `segments[]`** with numeric `start`/`end`/`duration` seconds, kept separate from chunks so you can disable one to keep items small.
- **Explicit coverage status** — every item has `transcriptStatus` (`ok`, `no_captions`, `language_unavailable`, `video_unavailable`, `error`); videos without captions are not billed.
- **Playlist and channel expansion** — pass a playlist or channel URL and it expands into its videos (capped by `maxResults`).
- **Language control** — preferred language priority list, with `strictLanguage` to fail explicitly instead of silently falling back.
- **Extra renderings on request** — SRT, WebVTT and chapter-headed Markdown alongside the JSON (`textFormats`).
- **Normalized video metadata** in the same record — title, channel, publish date, duration, view count, description, keywords — droppable in one flag (`includeVideoMetadata`) if you don't need it.
- Pay only for videos actually returning a transcript — no charge for `no_captions` or `video_unavailable` items.

### Input

| Field | Type | Required | Description |
|---|---|---|---|
| `videoUrls` | array of strings | No¹ | Video, playlist or channel URLs. Playlist/channel URLs expand into their videos (up to `maxResults`). |
| `videoIds` | array of strings | No¹ | Bare 11-character video ids, as an alternative to full URLs. |
| `languages` | array of strings | No | Preferred caption language codes in priority order. Default `["en"]`. |
| `strictLanguage` | boolean | No | If `true`, videos without a track in the requested languages get `transcriptStatus="language_unavailable"` instead of falling back. Default `false`. |
| `allowAutoGenerated` | boolean | No | Allow YouTube's auto-generated (ASR) captions when no human-authored track exists. Default `true`. |
| `includeSegments` | boolean | No | Include cue-level `segments[]`. Default `true`. |
| `includeChunks` | boolean | No | Include RAG `chunks[]`. Default `true`. |
| `chunkMaxChars` | integer | No | Target max characters per chunk (200–8000). Default `1200`. |
| `chunkOverlapChars` | integer | No | Overlap characters carried between chunks (0–2000). Default `100`. |
| `respectChapters` | boolean | No | Never let a chunk span two chapters. Default `true`. |
| `includeChapters` | boolean | No | Parse chapter markers from the description into `chapters[]`. Default `true`. |
| `includeVideoMetadata` | boolean | No | Include title, channel, publish date, duration, view count, description, keywords. Default `true`. |
| `textFormats` | array (`plain`|`srt`|`vtt`|`markdown`) | No | Extra text renderings to attach. Default `["plain"]`. |
| `maxResults` | integer | No | Safety cap on videos processed per run (`0` = no limit). Default `25`. |
| `maxConcurrency` | integer | No | Max parallel requests (1–50). Lower it if YouTube starts rate-limiting. Default `10`. |
| `proxyConfiguration` | object | No | Apify Proxy settings. Defaults to the **Residential** group — YouTube blocks the shared datacenter pool with a "Sign in to confirm you're not a bot" challenge. |

¹ At least one of `videoUrls` or `videoIds` must be supplied.

### Output

Each dataset row is one video. Key fields:

| Field | Type | Description |
|---|---|---|
| `videoId` / `videoUrl` | string | 11-character id and canonical watch URL. |
| `title` / `channelName` / `channelId` / `channelUrl` | string | null | Public video/channel identity. |
| `publishedAt` | string | null | ISO 8601 publish date. |
| `durationSeconds` / `viewCount` | number | null | Public video stats. |
| `description` / `keywords` | string/array | null | Public description (source of chapters) and tags. |
| `sourceInput` | string | The input URL/id this item came from (a playlist/channel URL for expanded videos). |
| `transcriptStatus` | string | `ok`, `no_captions`, `language_unavailable`, `video_unavailable` or `error` — always present. |
| `transcriptAvailable` | boolean | `true` when `transcriptStatus` is `ok`. |
| `language` / `isAutoGenerated` / `availableLanguages` | — | Caption track actually used, whether it's ASR, and all tracks offered. |
| `transcriptText` | string | null | Full transcript as one normalized string (when `textFormats` includes `plain`). |
| `segments` | array | null | Cue-level `[{index, start, end, duration, text}]`. |
| `chunks` | array | null | RAG chunks `[{chunkId, index, text, start, end, charCount, wordCount, chapterTitle, deepLinkUrl}]`. |
| `chapters` | array | null | Parsed `[{index, title, start, end}]`; `null` when the video defines none. |
| `srt` / `vtt` / `markdown` | string | null | Extra renderings (only when requested in `textFormats`). |
| `wordCount` / `charCount` / `chunkCount` | number | null | Full-transcript word/char count and number of chunks produced. |
| `error` | string | null | Reason when `transcriptStatus` is `error` or `video_unavailable`. |

Sample record (trimmed from an actual production run — full `segments`/`chunks` arrays and `transcriptText` normally run much longer; shown here as 2 segments and 1 chunk for readability):

```json
{
  "videoId": "arj7oStGLkU",
  "videoUrl": "https://www.youtube.com/watch?v=arj7oStGLkU",
  "sourceInput": "https://www.youtube.com/watch?v=arj7oStGLkU",
  "scrapedAt": "2026-09-30T15:51:07.117Z",
  "title": "Inside the Mind of a Master Procrastinator | Tim Urban | TED",
  "channelName": "TED",
  "channelId": "UCAuUUnT6oDeKwE6v1NGQxug",
  "channelUrl": "https://www.youtube.com/channel/UCAuUUnT6oDeKwE6v1NGQxug",
  "publishedAt": "2016-04-06T09:59:35-07:00",
  "durationSeconds": 844,
  "viewCount": 61963947,
  "description": "Tim Urban knows that procrastination doesn't make sense, but he's never been able to shake his habit of waiting until the last minute to get things done...",
  "keywords": ["TED Talk", "Tim Urban", "procrastination", "productivity"],
  "transcriptStatus": "ok",
  "transcriptAvailable": true,
  "availableLanguages": ["en", "es", "fr", "de", "ja", "pt-BR"],
  "error": null,
  "language": "en",
  "isAutoGenerated": false,
  "transcriptText": "So in college, I was a government major, which means I had to write a lot of papers...",
  "segments": [
    { "index": 0, "start": 12.645, "end": 14.015, "duration": 1.37, "text": "So in college," },
    { "index": 1, "start": 15.349, "end": 16.913, "duration": 1.564, "text": "I was a government major," }
  ],
  "chunks": [
    {
      "chunkId": "arj7oStGLkU-0",
      "index": 0,
      "text": "So in college, I was a government major, which means I had to write a lot of papers...",
      "start": 12.645,
      "end": 86.235,
      "charCount": 1183,
      "wordCount": 236,
      "chapterTitle": null,
      "deepLinkUrl": "https://www.youtube.com/watch?v=arj7oStGLkU&t=12s"
    }
  ],
  "chapters": null,
  "srt": null,
  "vtt": null,
  "markdown": null,
  "wordCount": 2277,
  "charCount": 12671,
  "chunkCount": 12
}
```

**Legal note:** transcript text is third-party copyrighted content owned by the video's creator. This Actor is a data feed for your own analysis, indexing and internal LLM use — you are responsible for how you store, display or republish any transcript text you retain.

### How to use it via the API

#### curl

```bash
curl "https://api.apify.com/v2/acts/YOUR_USERNAME~youtube-transcript-rag-feed/run-sync-get-dataset-items?token=YOUR_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "videoUrls": ["https://www.youtube.com/watch?v=arj7oStGLkU"],
    "languages": ["en"],
    "chunkMaxChars": 1200,
    "maxResults": 3
  }'
```

#### JavaScript (apify-client)

```javascript
import { ApifyClient } from "apify-client";

const client = new ApifyClient({ token: "YOUR_API_TOKEN" });

const run = await client.actor("YOUR_USERNAME/youtube-transcript-rag-feed").call({
  videoUrls: ["https://www.youtube.com/watch?v=arj7oStGLkU"],
  languages: ["en"],
  chunkMaxChars: 1200,
  maxResults: 3,
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

#### Python (apify-client)

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_API_TOKEN")

run = client.actor("YOUR_USERNAME/youtube-transcript-rag-feed").call(run_input={
    "videoUrls": ["https://www.youtube.com/watch?v=arj7oStGLkU"],
    "languages": ["en"],
    "chunkMaxChars": 1200,
    "strictLanguage": False,
    "maxResults": 3,
})

for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

### Pricing

Pay-per-event: **$0.56 per 1,000 transcripts** on the GOLD tier (event = one video successfully returning a transcript, `transcriptStatus="ok"`). That is 20% below the $0.70/1,000 charged by the mid-market transcript Actors with comparable reliability, and well below the $7/1,000 of the largest one. It is not the cheapest in the niche: at least one Actor charges $0.01/1,000 for a plain caption dump — if price is your only criterion, use that one; this Actor is priced for the chunked, chapter-aware, status-explicit output described above.

| Tier | Price per 1,000 transcripts |
|---|---|
| FREE | $0.84 |
| BRONZE | $0.73 |
| SILVER | $0.64 |
| GOLD / PLATINUM / DIAMOND | $0.56 |

Videos with no captions or that are unavailable are returned for transparency but never billed.

**Platform usage you also pay:** Apify charges you separately for platform resources this Actor uses while it runs, mainly the Residential proxy traffic needed because YouTube blocks the shared datacenter pool. Based on measured traffic, that adds roughly **$11 per 1,000 videos processed** (≈1.37 GB/1,000 at Apify's ~$8/GB residential rate) on top of the event price above. To reduce this: keep `includeSegments`/`includeChunks` limited to what you need, lower `maxConcurrency` if you don't need speed, and consider switching `proxyConfiguration` to the datacenter group for videos where YouTube's bot check isn't a problem for you (at the risk of more `video_unavailable`/blocked results).

### Troubleshooting

- **`transcriptStatus="no_captions"` for a video you know has captions**: check `availableLanguages` — the video may only offer tracks outside your `languages` list; widen the list or set `strictLanguage=false` to fall back to the default track.
- **`transcriptStatus="language_unavailable"`**: you set `strictLanguage=true` and none of your requested `languages` matched; either add the language shown in `availableLanguages` or drop `strictLanguage`.
- **`transcriptStatus="video_unavailable"`**: the video is private, deleted, age-restricted or region-blocked; the Actor never attempts a login or age-restriction bypass, so this is expected, not a bug.
- **Run returns 0 items**: check the run log for input validation errors — malformed or non-YouTube URLs are rejected before any request is made, with a clear message.
- **"Sign in to confirm you're not a bot" / unexpectedly high `video_unavailable`/`error` rate**: this happens when the Residential proxy pool serves an already-flagged exit IP; retry the run, and avoid manually switching to the datacenter proxy group unless you accept a higher block rate.
- **Fewer items than expected from a playlist/channel URL**: expansion is capped by `maxResults`; raise it (or set it to `0` for no limit) if you need the whole list.
- **Costs higher than expected**: most of the per-run cost is Residential proxy traffic, not the event price — see the Pricing section above.

### FAQ

**Does this require a YouTube account, cookies or API key?**
No. It only reads the public watch page and YouTube's own public caption endpoints, exactly what a logged-out browser loads when you open the transcript panel.

**Can I get transcripts for a whole playlist or channel?**
Yes — pass the playlist or channel URL in `videoUrls`; it expands into individual videos up to `maxResults`.

**Am I billed for videos with no captions?**
No. Only items with `transcriptStatus="ok"` are billed. `no_captions`, `language_unavailable`, `video_unavailable` and `error` items are returned for visibility but free.

**Can I republish the transcript text?**
Transcript text is the original creator's copyrighted content. This Actor delivers it to you as a data feed for your own use (indexing, analysis, internal LLM ingestion); you are responsible for how you use or display it further.

**Why are chunks and segments both provided?**
`segments` are the raw, often short caption cues; `chunks` merge them into embedding-sized blocks with overlap and stable ids for direct use in a vector database. Disable either with `includeSegments`/`includeChunks` if you only need one.

**Is this data allowed to be collected?**
This Actor reads only publicly visible, unauthenticated pages and endpoints, paces requests and never bypasses login, paywalls or access controls. YouTube's Terms of Service restrict automated access without permission — this is a known legal/business risk that is documented, not hidden; you are responsible for your own use of the collected data.

### Keywords

youtube transcript scraper, youtube transcript api, youtube transcript json, youtube captions scraper, youtube subtitles scraper, youtube transcript with timestamps, youtube rag data, youtube transcript for llm, youtube playlist transcript, youtube channel transcript, bulk youtube transcripts, video transcript chunks embeddings, youtube chapters extractor, youtube srt vtt export, youtube video metadata scraper

### Hashtags

\#youtube #transcript #captions #subtitles #rag #llm #ai #embeddings #vectordatabase #dataextraction #apify #developertools

# Changelog

This Actor's version history is a separate document: https://apify.com/leadsbrary/youtube-transcript-rag-feed/changelog.md

# Actor input Schema

## `videoUrls` (type: `array`):

YouTube video, playlist or channel URLs. Playlist and channel URLs are expanded into their videos (up to maxResults).

## `videoIds` (type: `array`):

11-character YouTube video ids, as an alternative to full URLs.

## `languages` (type: `array`):

Preferred caption language codes in priority order. The first available track wins; if none match, the default track is used unless strictLanguage is true.

## `strictLanguage` (type: `boolean`):

If true, videos with no caption track in the requested languages get transcriptStatus="language\_unavailable" instead of falling back to the default track.

## `allowAutoGenerated` (type: `boolean`):

Whether YouTube's automatic speech-recognition caption tracks may be used when no human-authored track exists.

## `includeSegments` (type: `boolean`):

Include the raw cue-level transcript segments with start/end seconds. Disable to keep dataset items small.

## `includeChunks` (type: `boolean`):

Include RAG chunks: cues merged into embedding-sized text blocks with stable ids, timestamps and deep-link URLs.

## `chunkMaxChars` (type: `integer`):

Target maximum characters per RAG chunk (chunks break on cue boundaries, so the real size may be slightly under).

## `chunkOverlapChars` (type: `integer`):

Characters of overlap carried from the end of one chunk into the next, to avoid cutting sentences across chunk boundaries.

## `respectChapters` (type: `boolean`):

When chapters are detected, never let a chunk span two chapters and tag each chunk with its chapter title.

## `includeChapters` (type: `boolean`):

Parse the video's chapter markers (from the public description) into a chapters array.

## `includeVideoMetadata` (type: `boolean`):

Include normalized video metadata (title, channel, publish date, duration, view count, description, keywords) in each item.

## `textFormats` (type: `array`):

Extra plain-text renderings to attach alongside the JSON.

## `maxResults` (type: `integer`):

Maximum number of videos to process in the run (safety cap on cost). 0 means no limit.

## `maxConcurrency` (type: `integer`):

Maximum parallel requests. Lower it if YouTube starts rate-limiting.

## `proxyConfiguration` (type: `object`):

Apify proxy settings. Defaults to Residential proxies: YouTube blocks the shared datacenter pool with a "Sign in to confirm you're not a bot" challenge. Datacenter proxies can still be selected manually to cut cost, at the risk of that block.

## Actor input object example

```json
{
  "videoUrls": [
    "https://www.youtube.com/watch?v=arj7oStGLkU"
  ],
  "videoIds": [],
  "languages": [
    "en"
  ],
  "strictLanguage": false,
  "allowAutoGenerated": true,
  "includeSegments": true,
  "includeChunks": true,
  "chunkMaxChars": 1200,
  "chunkOverlapChars": 100,
  "respectChapters": true,
  "includeChapters": true,
  "includeVideoMetadata": true,
  "textFormats": [
    "plain"
  ],
  "maxResults": 3,
  "maxConcurrency": 10,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoUrls": [
        "https://www.youtube.com/watch?v=arj7oStGLkU"
    ],
    "videoIds": [],
    "languages": [
        "en"
    ],
    "maxResults": 3,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("leadsbrary/youtube-transcript-rag-feed").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "videoUrls": ["https://www.youtube.com/watch?v=arj7oStGLkU"],
    "videoIds": [],
    "languages": ["en"],
    "maxResults": 3,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("leadsbrary/youtube-transcript-rag-feed").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoUrls": [
    "https://www.youtube.com/watch?v=arj7oStGLkU"
  ],
  "videoIds": [],
  "languages": [
    "en"
  ],
  "maxResults": 3,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call leadsbrary/youtube-transcript-rag-feed --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,leadsbrary/youtube-transcript-rag-feed"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/NLSAS7YswCjA1t71s/builds/CMdOcplQHvf4az3r7/openapi.json
