YouTube Transcript Scraper avatar

YouTube Transcript Scraper

Pricing

Pay per event

Go to Apify Store
YouTube Transcript Scraper

YouTube Transcript Scraper

Get YouTube transcripts as timed segments or plain text, plus title, channel, duration and views. Works on videos, Shorts, playlists and channels; picks your preferred caption language. No browser, no API key. Videos without captions are not charged.

Pricing

Pay per event

Rating

0.0

(0)

Developer

axly

axly

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

3 days ago

Last modified

Share

Turn YouTube videos into text. Give the actor video links, Shorts links, video IDs, playlists or whole channels and get one JSON row per video with the transcript as timed segments or plain text, the caption language that was used, and the video's title, channel, duration and view count. It reads YouTube the way the mobile app does, so it needs no browser and no API key. It runs through Apify's residential proxy by default, because YouTube asks datacenter IPs to sign in.

Built for AI and LLM pipelines (RAG over video libraries, summarisation, fine-tuning data), content teams repurposing videos into articles and posts, researchers building interview or lecture corpora, SEO tools mining spoken keywords, and accessibility work.

Three things set it apart from other transcript scrapers:

  • You know what happened to every video. Videos without captions, private, removed or age-restricted videos still get a row, with transcript: null and an error code such as no_captions or video_unavailable. Those rows are not charged as transcripts.
  • Language control. Pass a preference list (["en", "es"]): the actor picks the first match, prefers a human-made track over the auto-generated one, and tells you which track it used (language, isAutoGenerated) and which others exist (availableLanguages).
  • Playlists and channels in one input. A playlist or channel URL is expanded to its videos (newest first) with one request per 100 videos.

What you get

One row per video.

FieldDescription
videoId, urlThe video and its watch URL
title, channelName, channelId, channelUrlVideo and channel identity
durationSeconds, viewCount, isLive, thumbnailUrl, keywords, descriptionVideo metadata (switch off with includeMetadata: false)
transcriptArray of {start, duration, text} in seconds (outputFormat: segments, default) or one string (text); null when the video has no usable captions
transcriptTextThe joined text, only with outputFormat: both
language, languageName, isAutoGeneratedWhich caption track was used and whether YouTube generated it automatically
availableLanguagesEvery caption track of the video: [{code, name, kind}] (kind is manual or asr)
isTranslated, translationErrorResult of the optional translateTo request
segmentCount, wordCountSize of the transcript
error, errorReasonnull on success; otherwise no_captions, language_not_available, video_unavailable, private, age_restricted, live, unplayable, bot_detected, rate_limited or fetch_failed, plus YouTube's own message
sourceThe input line that produced the row (the playlist or channel URL for expanded videos)
scrapedAtISO-8601 UTC timestamp

Use cases

  • RAG and chat over video. Pull a channel's transcripts with timestamps and index them; answers can link to watch?v=ID&t=<start>.
  • Video to blog, newsletter or social posts. Fetch outputFormat: text for a playlist and hand the text to your writing model.
  • Research corpora. Interviews, lectures, earnings calls and political speeches as clean text with speaker-free timing, in the language you need.
  • SEO and topic research. Count what competitors actually say in their videos; keywords and description are included in every row.
  • Monitoring. Schedule a run on a channel URL; the checkpoint makes a re-run with the same input skip videos already delivered.

Input

ParameterTypeDefaultDescription
videoUrlsarray–Video URLs (watch?v=, youtu.be/, /shorts/, /embed/, /live/), bare video IDs, playlist URLs or PL… IDs, channel URLs (/@handle, /channel/UC…, /c/…, /user/…).
languagesarray["en"]Preferred caption languages in order. en also matches en-US and en-GB.
fallbackToAnyLanguagebooleantrueUse the video's default track when none of the preferred languages exists.
translateTostring–Ask YouTube to translate the chosen track (de, en, …). Best effort: YouTube throttles translations per IP, and the row falls back to the original language with translationError: rate_limited.
outputFormatenumsegmentssegments, text or both.
includeMetadatabooleantrueTitle, channel, duration, views, thumbnail, keywords, description.
maxVideosinteger100Cap on videos per run, including those expanded from playlists and channels.
concurrencyinteger5Parallel videos (1–10).
proxyConfigurationobjectApify residentialYouTube bot-checks datacenter IPs; keep residential unless you bring your own clean proxies.

Example input

{
"videoUrls": [
"https://www.youtube.com/watch?v=UF8uR6Z6KLc",
"https://www.youtube.com/shorts/-_aMgxyQo_8",
"https://www.youtube.com/playlist?list=PLbpi6ZahtOH6Blw3RGYpWkSByi_T7Rygb",
"https://www.youtube.com/@TED"
],
"languages": ["en"],
"outputFormat": "segments",
"maxVideos": 50
}

Example output row

{
"videoId": "UF8uR6Z6KLc",
"url": "https://www.youtube.com/watch?v=UF8uR6Z6KLc",
"title": "Steve Jobs' 2005 Stanford Commencement Address",
"channelName": "Stanford",
"channelId": "UC-EnprmCZ3OXyAoG7vjVNCA",
"durationSeconds": 904,
"viewCount": 49183529,
"language": "en",
"languageName": "English - English",
"isAutoGenerated": false,
"isTranslated": false,
"isLive": false,
"availableLanguages": [{"code": "ar", "name": "Arabic", "kind": "manual"}, {"code": "en", "name": "English (auto-generated)", "kind": "asr"}, {"code": "en", "name": "English - English", "kind": "manual"}],
"transcript": [
{"start": 7.47, "duration": 2.899, "text": "This program is brought to you by Stanford University."},
{"start": 10.47, "duration": 3.973, "text": "Please visit us at stanford.edu"}
],
"segmentCount": 244,
"wordCount": 2298,
"error": null,
"errorReason": null,
"source": "https://www.youtube.com/watch?v=UF8uR6Z6KLc",
"scrapedAt": "2026-10-05T09:12:44Z"
}

A video without captions looks like {"videoId": "aqz-KE-bpKQ", "title": "Big Buck Bunny 60fps 4K", "transcript": null, "error": "no_captions", ...}.

Scheduling, webhooks and integrations

Run it on a schedule with a channel URL to keep a transcript archive current; a re-run with the same input resumes from its checkpoint and skips videos it already delivered. Add a webhook on ACTOR.RUN.SUCCEEDED to push the dataset to your pipeline, or use the Apify integrations for Google Sheets, Make, Zapier, n8n and LangChain / LlamaIndex loaders.

Use with AI assistants (MCP)

Through the Apify MCP server any MCP-capable assistant (Claude, Cursor, custom agents) can call this actor as a tool: "get the transcript of this YouTube video and summarise it". Input and output are plain JSON, so no glue code is needed.

FAQ

Which videos have transcripts? Any video with captions: uploaded subtitles or YouTube's automatic speech recognition (available for most spoken-word videos in major languages a few hours after upload). Music videos, silent clips and videos whose owner disabled captions return error: no_captions.

Are auto-generated transcripts punctuated? No. YouTube's ASR tracks have no punctuation or speaker labels; manual tracks keep whatever the uploader wrote. isAutoGenerated tells you which one you got.

Can I translate? translateTo asks YouTube for its machine translation of the chosen track. YouTube rate-limits that path heavily, so treat it as a bonus: when it refuses, the row carries the original language and translationError: rate_limited. For reliable translation, feed the original text to your own translation model.

Does it work for Shorts, live streams and premieres? Shorts yes (same captions system). Live streams and premieres have no captions until the video is processed after the stream; they return error: live or no_captions.

How fast is it and are there limits? About 2 requests and 0.3–0.5 s per video; with the default concurrency of 5, a 1,000-video channel takes a few minutes. Age-restricted and private videos cannot be read without signing in and are reported as age_restricted / private.

Is this legal? The actor reads only publicly available caption data that YouTube serves to any viewer. You are responsible for how you use the text (copyright on the spoken content stays with its owner).