YouTube Transcript Scraper
Pricing
Pay per event
YouTube Transcript Scraper
Get YouTube transcripts as timed segments or plain text, plus title, channel, duration and views. Works on videos, Shorts, playlists and channels; picks your preferred caption language. No browser, no API key. Videos without captions are not charged.
Pricing
Pay per event
Rating
0.0
(0)
Developer
axly
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
3 days ago
Last modified
Categories
Share
Turn YouTube videos into text. Give the actor video links, Shorts links, video IDs, playlists or whole channels and get one JSON row per video with the transcript as timed segments or plain text, the caption language that was used, and the video's title, channel, duration and view count. It reads YouTube the way the mobile app does, so it needs no browser and no API key. It runs through Apify's residential proxy by default, because YouTube asks datacenter IPs to sign in.
Built for AI and LLM pipelines (RAG over video libraries, summarisation, fine-tuning data), content teams repurposing videos into articles and posts, researchers building interview or lecture corpora, SEO tools mining spoken keywords, and accessibility work.
Three things set it apart from other transcript scrapers:
- You know what happened to every video. Videos without captions, private, removed or age-restricted videos still get a row, with
transcript: nulland anerrorcode such asno_captionsorvideo_unavailable. Those rows are not charged as transcripts. - Language control. Pass a preference list (
["en", "es"]): the actor picks the first match, prefers a human-made track over the auto-generated one, and tells you which track it used (language,isAutoGenerated) and which others exist (availableLanguages). - Playlists and channels in one input. A playlist or channel URL is expanded to its videos (newest first) with one request per 100 videos.
What you get
One row per video.
| Field | Description |
|---|---|
videoId, url | The video and its watch URL |
title, channelName, channelId, channelUrl | Video and channel identity |
durationSeconds, viewCount, isLive, thumbnailUrl, keywords, description | Video metadata (switch off with includeMetadata: false) |
transcript | Array of {start, duration, text} in seconds (outputFormat: segments, default) or one string (text); null when the video has no usable captions |
transcriptText | The joined text, only with outputFormat: both |
language, languageName, isAutoGenerated | Which caption track was used and whether YouTube generated it automatically |
availableLanguages | Every caption track of the video: [{code, name, kind}] (kind is manual or asr) |
isTranslated, translationError | Result of the optional translateTo request |
segmentCount, wordCount | Size of the transcript |
error, errorReason | null on success; otherwise no_captions, language_not_available, video_unavailable, private, age_restricted, live, unplayable, bot_detected, rate_limited or fetch_failed, plus YouTube's own message |
source | The input line that produced the row (the playlist or channel URL for expanded videos) |
scrapedAt | ISO-8601 UTC timestamp |
Use cases
- RAG and chat over video. Pull a channel's transcripts with timestamps and index them; answers can link to
watch?v=ID&t=<start>. - Video to blog, newsletter or social posts. Fetch
outputFormat: textfor a playlist and hand the text to your writing model. - Research corpora. Interviews, lectures, earnings calls and political speeches as clean text with speaker-free timing, in the language you need.
- SEO and topic research. Count what competitors actually say in their videos;
keywordsanddescriptionare included in every row. - Monitoring. Schedule a run on a channel URL; the checkpoint makes a re-run with the same input skip videos already delivered.
Input
| Parameter | Type | Default | Description |
|---|---|---|---|
videoUrls | array | – | Video URLs (watch?v=, youtu.be/, /shorts/, /embed/, /live/), bare video IDs, playlist URLs or PL… IDs, channel URLs (/@handle, /channel/UC…, /c/…, /user/…). |
languages | array | ["en"] | Preferred caption languages in order. en also matches en-US and en-GB. |
fallbackToAnyLanguage | boolean | true | Use the video's default track when none of the preferred languages exists. |
translateTo | string | – | Ask YouTube to translate the chosen track (de, en, …). Best effort: YouTube throttles translations per IP, and the row falls back to the original language with translationError: rate_limited. |
outputFormat | enum | segments | segments, text or both. |
includeMetadata | boolean | true | Title, channel, duration, views, thumbnail, keywords, description. |
maxVideos | integer | 100 | Cap on videos per run, including those expanded from playlists and channels. |
concurrency | integer | 5 | Parallel videos (1–10). |
proxyConfiguration | object | Apify residential | YouTube bot-checks datacenter IPs; keep residential unless you bring your own clean proxies. |
Example input
{"videoUrls": ["https://www.youtube.com/watch?v=UF8uR6Z6KLc","https://www.youtube.com/shorts/-_aMgxyQo_8","https://www.youtube.com/playlist?list=PLbpi6ZahtOH6Blw3RGYpWkSByi_T7Rygb","https://www.youtube.com/@TED"],"languages": ["en"],"outputFormat": "segments","maxVideos": 50}
Example output row
{"videoId": "UF8uR6Z6KLc","url": "https://www.youtube.com/watch?v=UF8uR6Z6KLc","title": "Steve Jobs' 2005 Stanford Commencement Address","channelName": "Stanford","channelId": "UC-EnprmCZ3OXyAoG7vjVNCA","durationSeconds": 904,"viewCount": 49183529,"language": "en","languageName": "English - English","isAutoGenerated": false,"isTranslated": false,"isLive": false,"availableLanguages": [{"code": "ar", "name": "Arabic", "kind": "manual"}, {"code": "en", "name": "English (auto-generated)", "kind": "asr"}, {"code": "en", "name": "English - English", "kind": "manual"}],"transcript": [{"start": 7.47, "duration": 2.899, "text": "This program is brought to you by Stanford University."},{"start": 10.47, "duration": 3.973, "text": "Please visit us at stanford.edu"}],"segmentCount": 244,"wordCount": 2298,"error": null,"errorReason": null,"source": "https://www.youtube.com/watch?v=UF8uR6Z6KLc","scrapedAt": "2026-10-05T09:12:44Z"}
A video without captions looks like {"videoId": "aqz-KE-bpKQ", "title": "Big Buck Bunny 60fps 4K", "transcript": null, "error": "no_captions", ...}.
Scheduling, webhooks and integrations
Run it on a schedule with a channel URL to keep a transcript archive current; a re-run with the same input resumes from its checkpoint and skips videos it already delivered. Add a webhook on ACTOR.RUN.SUCCEEDED to push the dataset to your pipeline, or use the Apify integrations for Google Sheets, Make, Zapier, n8n and LangChain / LlamaIndex loaders.
Use with AI assistants (MCP)
Through the Apify MCP server any MCP-capable assistant (Claude, Cursor, custom agents) can call this actor as a tool: "get the transcript of this YouTube video and summarise it". Input and output are plain JSON, so no glue code is needed.
FAQ
Which videos have transcripts? Any video with captions: uploaded subtitles or YouTube's automatic speech recognition (available for most spoken-word videos in major languages a few hours after upload). Music videos, silent clips and videos whose owner disabled captions return error: no_captions.
Are auto-generated transcripts punctuated? No. YouTube's ASR tracks have no punctuation or speaker labels; manual tracks keep whatever the uploader wrote. isAutoGenerated tells you which one you got.
Can I translate? translateTo asks YouTube for its machine translation of the chosen track. YouTube rate-limits that path heavily, so treat it as a bonus: when it refuses, the row carries the original language and translationError: rate_limited. For reliable translation, feed the original text to your own translation model.
Does it work for Shorts, live streams and premieres? Shorts yes (same captions system). Live streams and premieres have no captions until the video is processed after the stream; they return error: live or no_captions.
How fast is it and are there limits? About 2 requests and 0.3–0.5 s per video; with the default concurrency of 5, a 1,000-video channel takes a few minutes. Age-restricted and private videos cannot be read without signing in and are reported as age_restricted / private.
Is this legal? The actor reads only publicly available caption data that YouTube serves to any viewer. You are responsible for how you use the text (copyright on the spoken content stays with its owner).