YouTube Transcript Scraper avatar

YouTube Transcript Scraper

Pricing

$4.00 / 1,000 transcripts

Go to Apify Store
YouTube Transcript Scraper

YouTube Transcript Scraper

Get YouTube video and Shorts transcripts as plain text, timed segments or SRT, with title, channel, views and publish date. Pick languages in priority order. Videos without captions are never charged.

Pricing

$4.00 / 1,000 transcripts

Rating

0.0

(0)

Developer

Cairnware

Cairnware

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

21 hours ago

Last modified

Categories

Share

Get the transcript (captions / subtitles) of any YouTube video or Short as clean plain text, timestamped segments or an SRT subtitle file, together with the video's title, channel, duration, views, likes and publish date. One row per video, ready for JSON, CSV, Excel, Google Sheets, the Apify API or your LLM pipeline.

Built for runs you do not have to babysit:

  • You only pay for transcripts delivered. Videos without captions, unavailable or private videos, invalid links and failures are returned as free rows with a clear reason. No start fee.
  • Every input is accounted for. Each link you enter appears in the output with status ok, no_transcript or error. The run ends with a summary such as "48 of 50 videos have a transcript; 2 have no captions in the requested languages (not charged)".
  • Smart retries. When YouTube blocks or rate-limits a request, it is retried with exponential backoff on a fresh residential IP, switching between YouTube's app clients.
  • Language control. Give your preferred languages in order; human-made captions are preferred over auto-generated ones, and you can fall back to the video's original language.
  • Respects your budget. The Actor stops when your maximum cost per run is reached and lists the remaining videos as free, unscraped rows.

What data you get

FieldWhat it contains
textThe full transcript as one plain-text string (HTML entities decoded, line breaks joined).
segmentsEvery caption line with start and duration in seconds (on by default).
srtThe transcript as an SRT subtitle file (off by default).
language, languageName, isAutoGeneratedWhich caption track you got, e.g. en, "English (auto-generated)", true.
availableLanguagesEvery caption track the video has, so you can re-run in another language.
title, channelName, channelId, channelUrl, durationSeconds, viewCount, isLiveContentAlways included.
publishDate, likeCount, category, description, keywords, thumbnailUrlWith Full video details (on by default).
wordCount, segmentCount, warnings, errorMessage, scrapedAtRun bookkeeping.

Use cases

  • AI and LLM workflows: feed transcripts into summarisation, RAG, Q&A or content repurposing pipelines.
  • Content research and SEO: see what top videos in your niche actually say; turn videos into blog posts, show notes or quotes.
  • Market and brand research: analyse product reviews, interviews and talks at scale.
  • Subtitles: download SRT files for editing or translation workflows.
  • Education and accessibility: create study notes and searchable text from lectures.

How to use

  1. Paste one or more YouTube video or Shorts links (or video IDs), one per line.
  2. Optionally set your preferred languages (default en).
  3. Click Start. Download the results as JSON, CSV, Excel or HTML, or read them through the Apify API.

All common link formats work: youtube.com/watch?v=..., youtu.be/..., /shorts/..., /embed/..., /live/..., m.youtube.com, music.youtube.com, and extra parameters like &t=30s or &list=... are ignored. Playlist, channel and search links are not supported yet: add the video links themselves.

Input example

{
"videos": [
"https://www.youtube.com/watch?v=jNQXAC9IVRw",
"https://youtu.be/dQw4w9WgXcQ",
"https://www.youtube.com/shorts/VIDEO_ID"
],
"languages": ["en", "es"],
"fallbackToAnyLanguage": true,
"includeAutoGenerated": true,
"includeTimestamps": true,
"includeSrt": false,
"includeMetadata": true,
"proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}
FieldNotes
videosRequired. Up to 10,000 per run. Duplicates are removed.
languagesLanguage codes in order of preference (en, es, pt-BR, zh-Hans...). en matches every English variant; pt-BR prefers Brazilian Portuguese, then other Portuguese. Empty = each video's original language.
fallbackToAnyLanguageIf none of your languages exists, return the original-language transcript (with a warning). Off = no_transcript row, not charged.
includeAutoGeneratedUse YouTube's automatic captions when the uploader added none (most videos only have these).
includeTimestampsAdd the segments array.
includeSrtAdd the srt field.
includeMetadataAdd publish date, likes, category, description, tags and thumbnail.
maxRetriesRetries per request before a video is marked error (default 6).
maxConcurrencyVideos processed in parallel (default 5).

Output example

Real output for "Me at the zoo", collected on 2026-10-03 (description and segments shortened):

{
"input": "jNQXAC9IVRw",
"videoId": "jNQXAC9IVRw",
"url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
"status": "ok",
"errorMessage": null,
"title": "Me at the zoo",
"channelName": "jawed",
"channelId": "UC4QobU6STFB0P71PMvOGN5A",
"channelUrl": "https://www.youtube.com/channel/UC4QobU6STFB0P71PMvOGN5A",
"durationSeconds": 19,
"viewCount": 439481354,
"likeCount": 19982756,
"publishDate": "2005-04-23T20:31:52-07:00",
"category": "Film & Animation",
"description": "...",
"keywords": ["me at the zoo", "jawed karim", "first youtube video"],
"thumbnailUrl": "https://i.ytimg.com/vi/jNQXAC9IVRw/hqdefault.jpg",
"isLiveContent": false,
"language": "en",
"languageName": "English",
"isAutoGenerated": false,
"availableLanguages": [
{ "languageCode": "en", "languageName": "English", "isAutoGenerated": false },
{ "languageCode": "de", "languageName": "German", "isAutoGenerated": false }
],
"text": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say",
"wordCount": 39,
"segmentCount": 6,
"segments": [
{ "start": 1.2, "duration": 2.16, "text": "All right, so here we are, in front of the elephants" },
{ "start": 5.318, "duration": 2.656, "text": "the cool thing about these guys is that they have really..." },
{ "start": 7.974, "duration": 4.642, "text": "really really long trunks" }
],
"srt": null,
"warnings": [],
"scrapedAt": "2026-10-03T10:25:02Z"
}

A video that has no captions in your languages (with fallback off) looks like this and is not charged:

{
"videoId": "9bZkp7q19f0",
"status": "no_transcript",
"errorMessage": "No captions in en. Available: ko (auto). Turn on \"Fall back to the original language\" or add one of these languages.",
"title": "PSY - GANGNAM STYLE(강남스타일) M/V",
"availableLanguages": [{ "languageCode": "ko", "languageName": "Korean (auto-generated)", "isAutoGenerated": true }],
"text": null
}

The dataset has two views: Overview (one row per video) and Timestamped lines (one row per caption line, handy for spreadsheets).

Pricing

$4 per 1,000 transcripts ($0.004 per video). No start fee. You only pay for transcripts delivered.

  • Charged once per row with status: "ok".
  • no_transcript and error rows (no captions, unavailable, private, members-only, invalid link, failed after retries, skipped by your cost limit) are free.
  • Residential proxy traffic and the metadata lookup are included in the price.
  • Set a maximum cost per run in the run options; the Actor never goes over it.

Example: 500 videos a month ≈ $2.

FAQ

Why do some videos return no_transcript? The video has no caption track that matches your settings: the uploader added none and YouTube did not create automatic captions (common for music without speech, very short or very new videos), or captions exist only in other languages. availableLanguages lists what the video has. You are not charged.

What is the difference between auto-generated and human-made captions? Human-made captions were uploaded by the creator and are usually accurate and punctuated. Auto-generated captions come from YouTube's speech recognition: they cover most videos but have no punctuation and some recognition errors. isAutoGenerated tells you which one you got.

Can I get a translation into another language? Not yet. YouTube's machine translation of captions is heavily rate-limited, so the Actor returns the caption tracks that actually exist. Use availableLanguages to see them.

In what order are the rows? In the order videos finish, which can differ from your input order when videos run in parallel. Each row has the input you gave and the videoId, so you can sort or join on them.

Why is a residential proxy the default? YouTube blocks many datacenter IP ranges ("Sign in to confirm you're not a bot"). Residential IPs, with a new IP on every retry, keep the failure rate low. You can choose another proxy in Proxy configuration, but expect more retries and errors.

Can I schedule it or call it from code? Yes. Use Apify Schedules for recurring runs, and the API tab of this Actor for ready-made examples in Python, JavaScript and cURL. Integrations exist for Google Sheets, Make, n8n, Zapier and webhooks.

Limitations

  • Only videos and Shorts. Playlist, channel and search links are not expanded yet (planned).
  • Age-restricted, private and members-only videos cannot be read without signing in, so they return error (free). The Actor never signs in to YouTube.
  • Live streams return captions only after the stream has ended and YouTube has processed them.
  • Machine translation of captions is not offered (see FAQ).
  • The Actor uses the same public endpoints as YouTube's apps. If YouTube changes them, results may break until the Actor is updated. Please report it and we will fix it.
  • It does not download video or audio and does not collect comments or personal data. Transcripts are the creators' content: you are responsible for using them in line with YouTube's terms, copyright and the laws that apply to you.

Support

Questions, bugs or feature requests: open an issue on the Issues tab or email hello@cairnware.com. Please include the run ID if something went wrong. Support replies may be AI-assisted.

Made by Cairnware.