# YouTube Video Transcript Scraper: Captions, Subtitles, SRT (`hyperbach/youtube-transcript-scraper`) Actor

YouTube transcript extractor for any video, channel, playlist or search: the video transcript text and timed segments, each track labelled manual or auto-generated captions. YouTube subtitles as SRT, WebVTT, Markdown or chunks, with metadata, chapters, comments, translation. No login.

- **URL**: https://apify.com/hyperbach/youtube-transcript-scraper.md
- **Developed by:** [Hyperbach](https://apify.com/hyperbach) (community)
- **Categories:** Videos, AI, Social media
- **Stats:** 4 total users, 3 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 transcripts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Video Transcript Scraper: Captions, Subtitles, SRT

**YouTube transcripts in bulk, exactly as YouTube serves them, and labelled: the text and the timed segments of any video, channel, playlist or search, each row saying which track it is — the language, and whether the captions were written by people (manual) or by YouTube's speech recognition (auto-generated).** Ask for languages in order (`en`, `de-DE`, `pt-BR`); when a video lacks them, the row says which track came back instead. Every caption is kept as served: no cue dropped or merged, entities decoded, ♪ kept. Add SRT, WebVTT, paragraphs, RAG chunks or Markdown; every track of a video in one row; chapters, likes, category and comments; a machine translation next to the original, segment for segment; speech-to-text for videos without captions. **[Follow channels on a schedule](https://apify.com/hyperbach/youtube-transcript-scraper/examples/monitor-youtube-channels-new-video-transcripts) and get the transcripts of new videos only.** A video without a transcript comes back as a free row that says why. No login, no cookies, no API key, proxies handled for you.

### Why this transcript scraper

- **Labelled, not guessed.** Every row names its track: `language`, `track_kind` (manual or auto-generated) and `transcript_source`. In our test of 62 runs of other YouTube transcript actors on the same videos, 9 served the auto-generated track as if it were the only one; here a video's manual and auto-generated English are two different rows of truth, and `allTracks` puts every track of a video in one row. Ready-made: [Check which YouTube videos have human-made captions](https://apify.com/hyperbach/youtube-transcript-scraper/examples/check-which-youtube-videos-have-manual-captions).
- **Exact.** The segments are YouTube's own caption events with millisecond start and duration: 61 for the manual English track of the test video, 52 for its auto-generated track, 60 for German (Germany) — the same counts as YouTube's track. HTML entities decoded, music notes and line breaks inside a caption kept. 8 of the 62 rival runs dropped or merged captions.
- **Honest defaults.** Ask for `fr` on a video with no French track and you get the English one with `language_note` saying "no track in fr; returned en (manual), the video's spoken language" — never a different track passed off as the one asked. A video without captions, a private or deleted id, an age-restricted video: a free row with `status` and `reason`, never a silent gap.
- **Channels, playlists, searches.** A channel by `@handle`, link or id (all uploads, videos only, Shorts only, or live), a playlist, a search — with a count per source, a date range (a channel's walk stops at the first older video) and duration limits. `onlyNew` turns a scheduled task into a feed of new videos only. Ready-made: [Transcripts of the top YouTube videos for a search](https://apify.com/hyperbach/youtube-transcript-scraper/examples/youtube-search-video-transcripts).
- **The formats you feed into something else.** SRT and WebVTT with valid timecodes, paragraphs with start and end, RAG chunks cut at caption and chapter boundaries with the chapter's title, Markdown with a timestamp link per paragraph; optionally as files in the run's storage with their links on the row. Ready-made: [YouTube video transcript with timestamps, SRT and Markdown](https://apify.com/hyperbach/youtube-transcript-scraper/examples/youtube-video-transcript-with-timestamps-and-srt).
- **Machine translation, said as such.** `translateTo` adds a translation by an OpenAI model next to the original: the same number of segments with the same timing, labelled `machine (OpenAI gpt-6-luna)`. YouTube's own translation is not served to logged-out readers, so this Actor does not pretend to have it. A translation that misses a segment says which and costs nothing. Ready-made: [Translate YouTube transcripts into Spanish, segment by segment](https://apify.com/hyperbach/youtube-transcript-scraper/examples/translate-youtube-transcripts-to-spanish).
- **Fast and plain.** No browser: 258 videos from a channel, a playlist and a search in 45 seconds. A video without captions can be transcribed from its audio (OpenAI whisper-1, timed segments), charged per minute.

### Who it's for

- **AI and RAG builders** — clean transcript text and chunks with timestamps, chapter titles and the video's metadata, from whole channels or searches, ready to embed.
- **Researchers and analysts** — what was said across many videos, in the original language and, when needed, a labelled machine translation, with publish dates and view counts.
- **Content teams and editors** — subtitles in SRT or WebVTT, Markdown with clickable timestamps, and comments, from your own channel or others'.
- **Developers and agents** — one row shape, typed fields, a free labelled row for every video without a transcript, and a run summary to check a pipeline against.

### Quick start

**One video, the transcript**

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ]
}
```

**Several videos, German first, SRT and WebVTT**

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://youtu.be/jNQXAC9IVRw"
  ],
  "languages": [
    "de",
    "en"
  ],
  "outputFormats": [
    "srt",
    "vtt"
  ]
}
```

**The latest 50 videos of a channel, as RAG chunks**

```json
{
  "channels": [
    "@RickAstleyYT"
  ],
  "channelContent": "videos",
  "maxVideosPerSource": 50,
  "outputFormats": [
    "chunks"
  ]
}
```

**A search, auto-generated captions only, published in the last 30 days**

```json
{
  "searchQueries": [
    "python tutorial"
  ],
  "trackKind": "auto_only",
  "dateFrom": "30 days",
  "maxVideosPerSource": 20
}
```

**Every track of a video, plus a French machine translation**

```json
{
  "videos": [
    "dQw4w9WgXcQ"
  ],
  "allTracks": true,
  "translateTo": "fr"
}
```

### Output

One row per video: the transcript and its labels, the video's metadata, the formats and add-ons you asked for — or, for a video without a transcript, the same row with `status` and `reason` saying why (free):

| field | meaning |
|---|---|
| `status` | `ok` = the row carries a transcript (charged). Otherwise free, with `reason` saying why: `no_captions`, `no_matching_track`, `unavailable`, `private`, `age_restricted`, `live_offline`, `unplayable`, `login_required`, `bot_wall`, `track_empty`, `too_long_for_speech_to_text`, `speech_to_text_failed`, `invalid_id`, `source_error`, `error`. |
| `reason` | Why the row has no transcript, in words (null on `ok` rows). |
| `video_id` | The 11-character YouTube video id. Deduplicate on this. |
| `video_url` | The video's watch URL. |
| `title` | The video's title. |
| `channel_name` | The channel's name as YouTube shows it. |
| `channel_id` | The channel's id (`UC…`). |
| `channel_url` | The channel's URL. |
| `channel_handle` | The channel's `@handle`, when YouTube shows it on the video. |
| `channel_subscribers` | The channel's subscriber count as YouTube rounds it ("4.55M subscribers"). |
| `publish_date` | The day the video was published (YYYY-MM-DD). |
| `published_at` | When the video was published, with the time and YouTube's UTC offset (ISO 8601). |
| `duration_s` | The video's length in seconds. |
| `view_count` | Views, exact. |
| `like_count` | Likes, exact. |
| `category` | YouTube's category for the video ("Music", "Education"). |
| `keywords` | The video's tags as its uploader set them. |
| `description` | The video's description. |
| `thumbnail_url` | The largest thumbnail YouTube lists for the video. |
| `is_live_content` | True for a live stream or its replay. |
| `language` | The language code of the transcript on the row (`en`, `de-DE`, `pt-BR`). |
| `language_name` | The transcript's language in words ("German (Germany)"). |
| `requested_language` | The first language you asked for (`languages`), so a row can be checked against it. |
| `language_note` | Set when the row's track is not exactly the one asked: which one came back instead, and why. |
| `track_kind` | `manual` (captions written by people), `auto-generated` (YouTube's speech recognition) or `speech-to-text` (our model transcribed the audio, `transcribeMissing`). |
| `is_auto_generated` | False only for manual captions. |
| `transcript_source` | Where the transcript comes from, in words: "YouTube captions (manual)", "YouTube captions (auto-generated)" or "speech-to-text (OpenAI whisper-1)". |
| `segment_count` | How many timed segments the transcript has (YouTube's caption events with text). |
| `text` | The whole transcript as one text: the segments in order, a line break inside a caption turned into a space. |
| `word_count` | Words in `text`. |
| `char_count` | Characters in `text`. |
| `segments` | The timed segments: `start` and `duration` in seconds (millisecond precision) and `text` exactly as YouTube serves it. Off with `includeSegments: false`. |
| `srt` | The transcript as SRT subtitles (`outputFormats`). |
| `vtt` | The transcript as WebVTT subtitles (`outputFormats`). |
| `paragraphs` | Paragraphs with `start`, `end` and `text`: a new one at a pause of 2 seconds, a chapter start, or past about 700 characters at a sentence end (`outputFormats`). |
| `chunks` | RAG chunks of about `chunkSize` characters with `index`, `start`, `end`, `chapter` and `text`, cut at caption boundaries and chapter starts (`outputFormats`). |
| `markdown` | The transcript in Markdown: the title, a heading per chapter and a paragraph per paragraph, each led by a link to that moment of the video (`outputFormats`). |
| `chapters` | The video's chapters: `title`, `start`, `end` (seconds) and a `url` that opens the video there. |
| `tracks` | Every caption track the video has: `language`, `name`, `kind` (manual or auto-generated), `is_translatable`. |
| `translation_languages` | The languages YouTube's player offers to translate the video's captions into (`language`, `name`). YouTube serves those translations to signed-in viewers only; `translateTo` is our machine translation instead. |
| `other_tracks` | With `allTracks`: every other track of the video, each with `language`, `language_name`, `track_kind`, `is_auto_generated`, `transcript_source`, `segment_count`, `text` and `segments`. |
| `translation` | With `translateTo`: the machine translation next to the original — `language`, `language_name`, `source` ("machine (OpenAI gpt-6-luna)"), `model`, `status` (`complete`, `incomplete`, `failed`, `not_needed`), `source_language`, `source_chars`, `segment_count`, `missing_segments`, `text`, `segments` (the original's count and timing), `srt` / `vtt` when asked, `note`, `error`. |
| `speech_to_text_seconds` | With `transcribeMissing`: the seconds of audio our model transcribed (the row's `track_kind` is `speech-to-text`). |
| `comments` | With `maxComments`: top-level comments with `comment_id`, `text`, `author`, `author_channel_id`, `is_creator`, `is_verified`, `published` (as YouTube says it), `like_count`, `reply_count`. |
| `files` | With `saveFiles`: links to the files written to the run's storage, one per format (`text`, `srt`, …; `translation_…` for the translation). |
| `notes` | Anything the row could not carry in full, in words (a metadata call that failed, an incomplete translation). |
| `source_type` | What brought the video into the run: `video`, `channel`, `playlist` or `search`; `file` on the free row of a list-file link that gave no list. |
| `source` | The value you gave (the URL, id, handle or query) that brought the video into the run. |
| `source_position` | The video's place in your list, or in the channel, playlist or search it came from (1 = first). |
| `scraped_at` | When the row was read (UTC). |

Example record:

```json
{
  "status": "ok",
  "reason": null,
  "video_id": "jNQXAC9IVRw",
  "video_url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
  "title": "Me at the zoo",
  "channel_name": "jawed",
  "channel_id": "UC4QobU6STFB0P71PMvOGN5A",
  "channel_url": "https://www.youtube.com/channel/UC4QobU6STFB0P71PMvOGN5A",
  "channel_handle": "@jawed",
  "channel_subscribers": "6.65M subscribers",
  "publish_date": "2005-04-23",
  "published_at": "2005-04-23T20:31:52-07:00",
  "duration_s": 19,
  "view_count": 439491656,
  "like_count": 19983031,
  "category": "Film & Animation",
  "keywords": [
    "me at the zoo",
    "jawed karim",
    "first youtube video"
  ],
  "description": "Microplastics are accumulating in human brains at an alarming rate\nhttps://www.youtube.com/watch?v=0PT5c1z3LL8\n\n“Nanoplastics and Human Health” with Matthew J Campen, PhD, MSPH\nhtt …(truncated for display)",
  "thumbnail_url": "https://i.ytimg.com/vi/jNQXAC9IVRw/hqdefault.jpg?sqp=-oaymwEmCOADEOgC8quKqQMa8AEB-AG-AoAC8AGKAgwIABABGFUgWShlMA8=&rs=AOn4CLA9eLBatYv9WbkD4BbZ2Im-biSPTw",
  "is_live_content": false,
  "language": "en",
  "language_name": "English",
  "requested_language": null,
  "language_note": null,
  "track_kind": "manual",
  "is_auto_generated": false,
  "transcript_source": "YouTube captions (manual)",
  "segment_count": 6,
  "text": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say",
  "word_count": 39,
  "char_count": 217,
  "segments": [
    {
      "start": 1.2,
      "duration": 2.16,
      "text": "All right, so here we are, in front of the\nelephants"
    },
    {
      "start": 5.318,
      "duration": 2.656,
      "text": "the cool thing about these guys is that they\nhave really..."
    },
    {
      "start": 7.974,
      "duration": 4.642,
      "text": "really really long trunks"
    },
    {
      "start": 12.616,
      "duration": 1.751,
      "text": "and that's cool"
    },
    {
      "start": 14.421,
      "duration": 1.312,
      "text": "(baaaaaaaaaaahhh!!)"
    },
    {
      "start": 16.881,
      "duration": 2.0,
      "text": "and that's pretty much all there is to\nsay"
    }
  ],
  "srt": "1\n00:00:01,200 --> 00:00:03,360\nAll right, so here we are, in front of the\nelephants\n\n2\n00:00:05,318 --> 00:00:07,974\nthe cool thing about these guys is that they\nhave really...\n\n3\n00:00:07,974 --> 00:00:12,616\nreally really long trunks\n\n4\n00:00:12,616 --> 00:00:14,367\nand that's cool\n\n5\n00:00:14,421 --> 00:00:15,733\n(baaaaaaaaaaahhh!!)\n\n6\n00:00:16,881 --> 00:00:18,881\nand that's pretty much all there is to\nsay\n",
  "vtt": "WEBVTT\n\n00:00:01.200 --> 00:00:03.360\nAll right, so here we are, in front of the\nelephants\n\n00:00:05.318 --> 00:00:07.974\nthe cool thing about these guys is that they\nhave really...\n\n00:00:07.974 --> 00:00:12.616\nreally really long trunks\n\n00:00:12.616 --> 00:00:14.367\nand that's cool\n\n00:00:14.421 --> 00:00:15.733\n(baaaaaaaaaaahhh!!)\n\n00:00:16.881 --> 00:00:18.881\nand that's pretty much all there is to\nsay\n",
  "paragraphs": [
    {
      "start": 1.2,
      "end": 3.36,
      "text": "All right, so here we are, in front of the elephants"
    },
    {
      "start": 5.318,
      "end": 18.881,
      "text": "the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say"
    }
  ],
  "chunks": [
    {
      "index": 0,
      "start": 1.2,
      "end": 3.36,
      "chapter": "Intro",
      "text": "All right, so here we are, in front of the elephants"
    },
    {
      "index": 1,
      "start": 5.318,
      "end": 18.881,
      "chapter": "The cool thing",
      "text": "the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say"
    }
  ],
  "markdown": "# Me at the zoo\n\n*en · YouTube captions (manual)*\n\n## [0:00](https://youtu.be/jNQXAC9IVRw?t=0) Intro\n\n[0:01](https://youtu.be/jNQXAC9IVRw?t=1) All right, so here we are, in front of the elephants\n\n## [0:05](https://youtu.be/jNQXAC9IVRw?t=5) The cool thing\n\n[0:05](https://youtu.be/jNQXAC9IVRw?t=5) the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say\n",
  "chapters": [
    {
      "title": "Intro",
      "start": 0.0,
      "end": 5.0,
      "url": "https://youtu.be/jNQXAC9IVRw?t=0"
    },
    {
      "title": "The cool thing",
      "start": 5.0,
      "end": 17.0,
      "url": "https://youtu.be/jNQXAC9IVRw?t=5"
    },
    {
      "title": "End",
      "start": 17.0,
      "end": 19.0,
      "url": "https://youtu.be/jNQXAC9IVRw?t=17"
    }
  ],
  "tracks": [
    {
      "language": "en",
      "name": "English",
      "kind": "manual",
      "is_translatable": true
    },
    {
      "language": "de",
      "name": "German",
      "kind": "manual",
      "is_translatable": true
    }
  ],
  "translation_languages": null,
  "other_tracks": [
    {
      "language": "de",
      "language_name": "German",
      "track_kind": "manual",
      "is_auto_generated": false,
      "transcript_source": "YouTube captions (manual)",
      "segment_count": 4,
      "text": "Also hier sind wir vor den Elefanten. Das Coole an den Typen ist dass sie sehr, sehr, sehr, lange Rüssel haben. Und das ist cool. Ansonsten gibt es nicht wirklich viel zu sagen.",
      "segments": [
        {
          "start": 0.0,
          "duration": 5.067,
          "text": "Also hier sind wir vor den Elefanten."
        },
        {
          "start": 5.067,
          "duration": 7.4,
          "text": "Das Coole an den Typen ist dass sie sehr, sehr, sehr, lange Rüssel haben."
        },
        {
          "start": 12.467,
          "duration": 1.8,
          "text": "Und das ist cool."
        },
        {
          "start": 17.067,
          "duration": 3.933,
          "text": "Ansonsten gibt es nicht wirklich viel zu sagen."
        }
      ]
    }
  ],
  "translation": {
    "language": "fr",
    "language_name": "French",
    "source": "machine (OpenAI gpt-6-luna)",
    "model": "gpt-6-luna",
    "status": "complete",
    "source_language": "en",
    "source_chars": 217,
    "segment_count": 6,
    "missing_segments": null,
    "text": "Bon, nous voici donc devant les éléphants Ce qui est génial chez eux, c'est qu'ils ont vraiment… de très, très longues trompes et c'est génial (baaaaaaaaaaahhh!!) et c'est à peu près tout ce qu'il y a à dire",
    "segments": [
      {
        "start": 1.2,
        "duration": 2.16,
        "text": "Bon, nous voici donc devant les\néléphants"
      },
      {
        "start": 5.318,
        "duration": 2.656,
        "text": "Ce qui est génial chez eux, c'est qu'ils\nont vraiment…"
      },
      {
        "start": 7.974,
        "duration": 4.642,
        "text": "de très, très longues trompes"
      },
      {
        "start": 12.616,
        "duration": 1.751,
        "text": "et c'est génial"
      },
      {
        "start": 14.421,
        "duration": 1.312,
        "text": "(baaaaaaaaaaahhh!!)"
      },
      {
        "start": 16.881,
        "duration": 2.0,
        "text": "et c'est à peu près tout ce qu'il y a à\ndire"
      }
    ],
    "error": null,
    "srt": "1\n00:00:01,200 --> 00:00:03,360\nBon, nous voici donc devant les\néléphants\n\n2\n00:00:05,318 --> 00:00:07,974\nCe qui est génial chez eux, c'est qu'ils\nont vraiment…\n\n3\n00:00:07,974 --> 00:00:12,616\nde très, très longues trompes\n\n4\n00:00:12,616 --> 00:00:14,367\net c'est génial\n\n5\n00:00:14,421 --> 00:00:15,733\n(baaaaaaaaaaahhh!!)\n\n6\n00:00:16,881 --> 00:00:18,881\net c'est à peu près tout ce qu'il y a à\ndire\n",
    "vtt": "WEBVTT\n\n00:00:01.200 --> 00:00:03.360\nBon, nous voici donc devant les\néléphants\n\n00:00:05.318 --> 00:00:07.974\nCe qui est génial chez eux, c'est qu'ils\nont vraiment…\n\n00:00:07.974 --> 00:00:12.616\nde très, très longues trompes\n\n00:00:12.616 --> 00:00:14.367\net c'est génial\n\n00:00:14.421 --> 00:00:15.733\n(baaaaaaaaaaahhh!!)\n\n00:00:16.881 --> 00:00:18.881\net c'est à peu près tout ce qu'il y a à\ndire\n"
  },
  "speech_to_text_seconds": null,
  "comments": [
    {
      "comment_id": "UgzuC3zzpRZkjc5Qzsd4AaABAg",
      "text": "We're so honored that the first ever YouTube video was filmed here!",
      "author": "@SanDiegoZoo",
      "author_channel_id": "UCC5NfQ6Mf0dq_eEwv4P_hWA",
      "is_creator": false,
      "is_verified": true,
      "published": "6 years ago",
      "like_count": 4800000,
      "reply_count": 985
    },
    {
      "comment_id": "UgxMs_M_v_f-LRN4snt4AaABAg",
      "text": "I got recommended this at 3 Oktober 2026",
      "author": "@abizard6003",
      "author_channel_id": "UCGEbPUnlpt8BN0im2hLfA0w",
      "is_creator": false,
      "is_verified": false,
      "published": "3 weeks ago (edited)",
      "like_count": 37000,
      "reply_count": 992
    },
    {
      "comment_id": "UgzBlQPOlQOOIMt_TKB4AaABAg",
      "text": "Is anyone here today?",
      "author": "@7_or-r6",
      "author_channel_id": "UC5NIhJcX0587B0dhK8vqvqw",
      "is_creator": false,
      "is_verified": false,
      "published": "1 month ago",
      "like_count": 41000,
      "reply_count": 991
    }
  ],
  "files": null,
  "notes": null,
  "source_type": "video",
  "source": "https://youtu.be/jNQXAC9IVRw",
  "source_position": 1,
  "scraped_at": "2026-10-03T11:33:50Z"
}
```

### Pricing

**Pay only for the transcripts a run delivers**, with no start fee and no platform usage billed to you. Prices fall with your Apify plan.

| per 1,000 | Free | Starter | Scale | Business |
|---|---|---|---|---|
| Transcript (`transcript`) | $5.00 | $4.25 | $3.50 | $2.50 |
| Extra caption track on that row (`extra_track`) | +$1.00 | +$0.85 | +$0.70 | +$0.50 |
| Comment (`comment`) | +$0.50 | +$0.425 | +$0.35 | +$0.25 |

| per unit | Free | Starter | Scale | Business |
|---|---|---|---|---|
| Machine translation, per 1,000 characters (`translation`) | +$0.003 | +$0.00255 | +$0.0021 | +$0.0015 |
| Speech-to-text, per minute (`speech_to_text`) | +$0.012 | +$0.0105 | +$0.0095 | +$0.009 |

A transcript is **$5.00 per 1,000 videos** on the Free plan and $2.50 on Business. `+` fees are charged **only on a row that carries the thing**: an extra track per further caption track delivered with `allTracks`; a comment per comment delivered; a machine translation per started 1,000 characters of the original, and only when every segment came back; speech-to-text per started minute of audio transcribed. A video without a transcript (no captions, unavailable, private, age-restricted, a speech-to-text download YouTube refuses) is a labelled row and free; so are saved files and a run that failed. `maxItems` is exact; before the first request the log says the most the run can cost at your plan's prices. Enterprise plans have their own rates — the Actor's Pricing tab shows the price for your plan.

### Usage patterns

- **A channel feed** — Save a task with a channel, `onlyNew` and the formats you need, and schedule it daily: each run reads the channel's newest `maxVideosPerSource` videos and delivers and charges only those it has not delivered before — a day without a new upload delivers nothing. A video without captions is not remembered: it is read again on each run, as a free row, because its captions can come later. Apify's integrations send the rows to a sheet, a webhook or Slack. Ready-made: [Monitor YouTube channels: transcripts of new videos only](https://apify.com/hyperbach/youtube-transcript-scraper/examples/monitor-youtube-channels-new-video-transcripts).
- **A corpus for retrieval** — `outputFormats: ["chunks"]` with `chunkSize` gives chunks with start, end and chapter title; `includeSegments: false` keeps rows small; the metadata (channel, publish time, views, category) filters the corpus. Ready-made: [YouTube channel transcripts of the last 60 days as RAG chunks](https://apify.com/hyperbach/youtube-transcript-scraper/examples/youtube-channel-transcripts-last-60-days-rag-chunks).
- **Subtitles** — `outputFormats: ["srt", "vtt"]` with `saveFiles: true` writes `<videoId>.srt` and `<videoId>.vtt` to the run's storage with their links on the row. With `translateTo`, the translation gets its own files. Ready-made: [Download a YouTube playlist's subtitles as SRT and VTT files](https://apify.com/hyperbach/youtube-transcript-scraper/examples/download-youtube-playlist-subtitles-srt-vtt-files).
- **Many languages** — `languages` is an order of preference; `allTracks` returns every caption track a video has, manual and auto-generated, each labelled; `translation_languages` lists the languages YouTube itself offers for the video.

### Input configuration

| field | type | default | what it does |
|---|---|---|---|
| `videos` | `array` | `[]` | YouTube videos in any form: watch, youtu.be, Shorts, live, embed and music.youtube.com links, or bare 11-character ids. One run takes many; repeats of the same video are read once. A channel or playlist link pasted here is read as one. A link to a text file, a CSV or a published Google Sheet with one URL or id per line is read as a list. |
| `channels` | `array` | `[]` | Channels to read the latest videos of: `@handle`, a channel link, or a `UC…` channel id. Newest first, up to `maxVideosPerSource` each. |
| `playlists` | `array` | `[]` | Playlists: a playlist link or its id (`PL…`). In playlist order, up to `maxVideosPerSource` each. |
| `searchQueries` | `array` | `[]` | YouTube searches: the videos the search returns, in YouTube's order, up to `maxVideosPerSource` each. |
| `maxVideosPerSource` | `integer` | `50` | How many videos to take from each channel, playlist and search (videos without captions count; videos outside your dates or durations do not). Videos you list one by one are all read. |
| `channelContent` | `all` / `videos` / `shorts` / `live` | `"all"` | Which of a channel's uploads to read. |
| `dateFrom` | `string` | `""` | Only videos published on or after this date: `2026-09-01`, or a span back from today such as `30 days`, `2 weeks`, `6 months`. Applies to channels, playlists and searches; a channel's walk stops at the first video older than this. |
| `dateTo` | `string` | `""` | Only videos published on or before this date (same forms as `dateFrom`). Applies to channels, playlists and searches. |
| `minDurationSeconds` | `integer` | `0` | Skip shorter videos from channels, playlists and searches. 0 = no limit. |
| `maxDurationSeconds` | `integer` | `0` | Skip longer videos from channels, playlists and searches. 0 = no limit. |
| `languages` | `array` | `[]` | Transcript languages to look for, in order of preference: `en`, `de`, `pt-BR`, `es-419`. A code without a region also matches the regional track (`de` finds `de-DE`). When the video has none of them, the row carries the video's spoken language and `language_note` says so. Empty = the video's spoken language. |
| `trackKind` | `manual_first` / `auto_first` / `manual_only` / `auto_only` | `"manual_first"` | YouTube has captions written by people (manual) and its own speech recognition (auto-generated); every row says which it is (`track_kind`). The two "only" choices return a free labelled row when the video lacks that kind. |
| `allTracks` | `boolean` | `false` | Also put every other caption track of the video on the row (`other_tracks`: each language, manual and auto-generated, with its text and segments). With `trackKind` manual only or auto-generated only, the other tracks are of that kind only. Charged per extra track. |
| `outputFormats` | `array` | `[]` | Formats added to each row next to the plain text and the timed segments, which every row has. |
| `chunkSize` | `integer` | `1000` | About how long each RAG chunk is. Chunks end at caption boundaries and at chapter starts, never inside a caption. |
| `includeSegments` | `boolean` | `true` | Keep the timed segments (`segments`: start, duration, text) on each row. Off = text and formats only, for smaller rows. |
| `saveFiles` | `boolean` | `false` | Also write the text and each extra format as files in the run's key-value store (`<videoId>.txt`, `.srt`, `.vtt`, `.md`, …), with their links on the row (`files`). Free. |
| `translateTo` | `string` | `""` | A language code (`fr`, `de`, `ja`, `pt-BR`, `zh-Hans`): each transcript is also translated by an OpenAI model and delivered next to the original (`translation`), segment for segment with the same timing, labelled `machine (OpenAI gpt-6-luna)`. It is not YouTube's own translation, which YouTube does not serve logged out. Charged per started 1,000 characters of the original, only when every segment came back; a translation with a missing segment says which and costs nothing. |
| `transcribeMissing` | `boolean` | `false` | For a video without any caption track, transcribe its audio with OpenAI whisper-1 (timed segments), labelled `speech-to-text (OpenAI whisper-1)`. Up to about 105 minutes of audio per video. Charged per started minute. |
| `maxComments` | `integer` | `0` | Top-level comments to add to each row (`comments`: text, author, likes, replies, when). 0 = none. Charged per comment. |
| `maxItems` | `integer` | `0` | The most transcripts this run delivers (and charges). 0 = no limit. Free rows (no captions, unavailable) do not count. |
| `onlyNew` | `boolean` | `false` | Skip videos this task delivered a transcript for in an earlier run: run a channel daily and get only its new videos. A video without captions is read again on each run (a free row): its captions can come later. Turns the memory on. |
| `stateStoreName` | `string` | `""` | Name of the memory between runs. Empty = derived from your videos, channels, playlists and searches, so the same task always meets its own memory. |

### Errors

**A run does not fail because of your input.** When an input cannot be used — a date that is not a date, an id in the wrong form, two settings that contradict each other — the run ends **Succeeded**, charges nothing, and says what to change:

- the status message starts with `INPUT REJECTED`;
- the dataset holds one row, `{"error": true, "code": "…", "message": "…"}`, and no results;
- the key-value store record `ERROR` holds the same object.

From code, check `error` on the first row before reading results.

A search that matches nothing is not an error: the dataset is empty and nothing is charged. A run that ends **Failed** is a fault on our side, never your input; it charges nothing, and we are alerted.

### FAQ

**Is the auto-generated track marked?**

Yes, on every row: `track_kind` is `manual` or `auto-generated`, `is_auto_generated` is true or false, and `transcript_source` says it in words. `trackKind` chooses which one you prefer, or only one kind.

**Why not YouTube's own translations?**

YouTube does not serve its translated captions reliably to logged-out readers: we measured every route, including the browser and its tokens; YouTube refuses almost every request (a rival that relies on them got 1 translated track in 5 tries on our test). `translateTo` is a machine translation by an OpenAI model instead, labelled as such and kept next to the original.

**What is never charged?**

A video without captions (unless you ask for speech-to-text and it is transcribed), an unavailable, private or age-restricted video, a value that is not a YouTube video, a track YouTube served empty, a channel that does not resolve, a translation with a missing segment, and any run that failed.

**Are the timestamps exact?**

They are YouTube's: each segment's `start` and `duration` come from the caption track in milliseconds. Auto-generated tracks overlap (a caption stays on screen while the next one starts); that is kept, and SRT and WebVTT allow it.

**Does it work for Shorts and live streams?**

Yes: a Short or a finished live stream is an ordinary video id; `channelContent: "shorts"` or `"live"` reads only those from a channel. A live stream that is still running or has not started comes back as a free `live_offline` row.

**One row per segment?**

Each row keeps its segments in one list. For one dataset row per segment, export the dataset with Apify's `unwind=segments` option (for example `https://api.apify.com/v2/datasets/<datasetId>/items?format=csv&unwind=segments`): each segment becomes a row with the video's fields next to it.

**Can an AI agent use it (MCP)?**

Yes. Apify's MCP server (`https://mcp.apify.com`) offers Store Actors as tools to Claude, Cursor and other MCP clients: add this Actor to the server's tool list and the agent sends `videos` and gets the rows back. The synchronous API (`run-sync-get-dataset-items`) does the same in one HTTP call.

**How does maxVideosPerSource count?**

It counts the videos a source delivers after the date and duration filters. A video already in the run from another list (or typed in `videos`) is delivered once and does not count again, so a channel then delivers its next video. With `onlyNew`, a video delivered by an earlier run still takes its place in the count.

**Age-restricted videos?**

YouTube serves them to signed-in viewers only, and this Actor does not sign in. They come back as free `age_restricted` rows with the reason.

### Integration

#### JavaScript

```javascript
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('hyperbach/youtube-transcript-scraper').call({"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"]});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
```

#### Python

```python
from apify_client import ApifyClient
client = ApifyClient('YOUR_TOKEN')
run = client.actor('hyperbach/youtube-transcript-scraper').call(run_input={'videos': ['https://www.youtube.com/watch?v=dQw4w9WgXcQ']})
items = client.dataset(run['defaultDatasetId']).list_items().items
```

#### CLI

```bash
apify call hyperbach/youtube-transcript-scraper --input '{"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"]}'
```

#### REST

```bash
curl -X POST "https://api.apify.com/v2/acts/hyperbach~youtube-transcript-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H 'Content-Type: application/json' -d '{"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"]}'
```

### Support

apify@hyperbach.com

*This page is generated from the Actor's schemas and a live sample — it cannot describe a field the Actor does not have.*

# Actor input Schema

## `videos` (type: `array`):

YouTube videos in any form: watch, youtu.be, Shorts, live, embed and music.youtube.com links, or bare 11-character ids. One run takes many; repeats of the same video are read once. A channel or playlist link pasted here is read as one. A link to a text file, a CSV or a published Google Sheet with one URL or id per line is read as a list.

## `channels` (type: `array`):

Channels to read the latest videos of: `@handle`, a channel link, or a `UC…` channel id. Newest first, up to `maxVideosPerSource` each.

## `playlists` (type: `array`):

Playlists: a playlist link or its id (`PL…`). In playlist order, up to `maxVideosPerSource` each.

## `searchQueries` (type: `array`):

YouTube searches: the videos the search returns, in YouTube's order, up to `maxVideosPerSource` each.

## `maxVideosPerSource` (type: `integer`):

How many videos to take from each channel, playlist and search (videos without captions count; videos outside your dates or durations do not). Videos you list one by one are all read.

## `channelContent` (type: `string`):

Which of a channel's uploads to read.

## `dateFrom` (type: `string`):

Only videos published on or after this date: `2026-09-01`, or a span back from today such as `30 days`, `2 weeks`, `6 months`. Applies to channels, playlists and searches; a channel's walk stops at the first video older than this.

## `dateTo` (type: `string`):

Only videos published on or before this date (same forms as `dateFrom`). Applies to channels, playlists and searches.

## `minDurationSeconds` (type: `integer`):

Skip shorter videos from channels, playlists and searches. 0 = no limit.

## `maxDurationSeconds` (type: `integer`):

Skip longer videos from channels, playlists and searches. 0 = no limit.

## `languages` (type: `array`):

Transcript languages to look for, in order of preference: `en`, `de`, `pt-BR`, `es-419`. A code without a region also matches the regional track (`de` finds `de-DE`). When the video has none of them, the row carries the video's spoken language and `language_note` says so. Empty = the video's spoken language.

## `trackKind` (type: `string`):

YouTube has captions written by people (manual) and its own speech recognition (auto-generated); every row says which it is (`track_kind`). The two "only" choices return a free labelled row when the video lacks that kind.

## `allTracks` (type: `boolean`):

Also put every other caption track of the video on the row (`other_tracks`: each language, manual and auto-generated, with its text and segments). With `trackKind` manual only or auto-generated only, the other tracks are of that kind only. Charged per extra track.

## `outputFormats` (type: `array`):

Formats added to each row next to the plain text and the timed segments, which every row has.

## `chunkSize` (type: `integer`):

About how long each RAG chunk is. Chunks end at caption boundaries and at chapter starts, never inside a caption.

## `includeSegments` (type: `boolean`):

Keep the timed segments (`segments`: start, duration, text) on each row. Off = text and formats only, for smaller rows.

## `saveFiles` (type: `boolean`):

Also write the text and each extra format as files in the run's key-value store (`<videoId>.txt`, `.srt`, `.vtt`, `.md`, …), with their links on the row (`files`). Free.

## `translateTo` (type: `string`):

A language code (`fr`, `de`, `ja`, `pt-BR`, `zh-Hans`): each transcript is also translated by an OpenAI model and delivered next to the original (`translation`), segment for segment with the same timing, labelled `machine (OpenAI gpt-6-luna)`. It is not YouTube's own translation, which YouTube does not serve logged out. Charged per started 1,000 characters of the original, only when every segment came back; a translation with a missing segment says which and costs nothing.

## `transcribeMissing` (type: `boolean`):

For a video without any caption track, transcribe its audio with OpenAI whisper-1 (timed segments), labelled `speech-to-text (OpenAI whisper-1)`. Up to about 105 minutes of audio per video. Charged per started minute.

## `maxComments` (type: `integer`):

Top-level comments to add to each row (`comments`: text, author, likes, replies, when). 0 = none. Charged per comment.

## `maxItems` (type: `integer`):

The most transcripts this run delivers (and charges). 0 = no limit. Free rows (no captions, unavailable) do not count.

## `onlyNew` (type: `boolean`):

Skip videos this task delivered a transcript for in an earlier run: run a channel daily and get only its new videos. A video without captions is read again on each run (a free row): its captions can come later. Turns the memory on.

## `stateStoreName` (type: `string`):

Name of the memory between runs. Empty = derived from your videos, channels, playlists and searches, so the same task always meets its own memory.

## Actor input object example

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "channels": [],
  "playlists": [],
  "searchQueries": [],
  "maxVideosPerSource": 50,
  "channelContent": "all",
  "dateFrom": "",
  "dateTo": "",
  "minDurationSeconds": 0,
  "maxDurationSeconds": 0,
  "languages": [],
  "trackKind": "manual_first",
  "allTracks": false,
  "outputFormats": [],
  "chunkSize": 1000,
  "includeSegments": true,
  "saveFiles": false,
  "translateTo": "",
  "transcribeMissing": false,
  "maxComments": 0,
  "maxItems": 0,
  "onlyNew": false,
  "stateStoreName": ""
}
```

# Actor output Schema

## `results` (type: `string`):

All scraped records in the default dataset. One row per video: the transcript and its labels, the video's metadata, the formats and add-ons you asked for — or, for a video without a transcript, the same row with `status` and `reason` saying why (free):

## `runSummary` (type: `string`):

How the run went: transcripts charged, free rows by status, per channel, playlist and search what was listed, taken and skipped (known, filtered, duplicate) and why a walk stopped; translation and speech-to-text counters; the source counters.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videos": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("hyperbach/youtube-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"] }

# Run the Actor and wait for it to finish
run = client.actor("hyperbach/youtube-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ]
}' |
apify call hyperbach/youtube-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,hyperbach/youtube-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/1dIFscBp53cFAAefM/builds/vqzU7ReZA2vj9qo1Q/openapi.json
