YouTube Transcript Scraper Pro (Captions + AI Fallback)
Pricing
from $0.70 / 1,000 transcript extracteds
YouTube Transcript Scraper Pro (Captions + AI Fallback)
Extract YouTube transcripts at scale without burning through your budget. It starts with free captions whenever they're available, then switches to AI only for videos that don't have them. You stay in control of costs, and the output — JSON, SRT, VTT, plain text, or LLM-ready format
Pricing
from $0.70 / 1,000 transcript extracteds
Rating
5.0
(1)
Developer
CodePoetry
Maintained by CommunityActor stats
12
Bookmarked
2.1K
Total users
548
Monthly active users
4 hours ago
Last modified
Categories
Share
YouTube Transcript Scraper — Captions + AI Speech-to-Text
Extract transcripts from any YouTube video — even when captions don't exist.
Give it a single video, a playlist, or a whole channel. You get each video's transcript in JSON, plain text, SRT, VTT, or an LLM-ready format. When a video has no captions, turn on AI transcription and the built-in speech-to-text model transcribes the audio — no API key, no extra setup. You pay per transcript, not per minute of server time.
Quick start
-
Click Try for free.
-
Paste one or more YouTube URLs into YouTube URLs:
- Single video:
youtube.com/watch?v=...oryoutu.be/... - Playlist:
youtube.com/playlist?list=... - Channel:
youtube.com/@channelname
Playlists and channels return 10 videos by default — raise Max videos per playlist/channel for more.
- Single video:
-
Pick your Output formats. Not sure? Plain Text is the words as one block.
-
AI transcription is off by default. Videos without native captions are reported with
NO_CAPTIONS_AVAILABLEunless you turn on Enable AI transcription. When you do, set Max AI minutes per run before processing an unknown playlist or channel. -
Click Start, then download results from the Dataset tab as JSON, CSV or Excel, or read them through the API.
How it works
- Expand — every URL becomes the videos it names. A channel URL is read from its
videos tab, newest first. A video named more than once — as a
youtu.belink and in a playlist, say — is transcribed and charged once. - Captions — for each video, the scraper looks for a caption track in your languages, in order: manual captions first (unless you chose auto-generated only), then auto-generated. The first track that downloads is used.
- Machine translation (optional, on by default) — with Machine-translate captions
into your language on, a video that has captions only in other languages has them
machine-translated into your first requested language, keeping every timestamp and charged
as a normal transcript. Items carry
is_machine_translated: trueandtranslated_from. - Original-language fallback (optional) — with Use captions in any available language on, a video with no track in your languages delivers its track in the language it was spoken in, charged as a normal transcript.
- AI transcription (optional) — the audio of a video with no captions at all is
transcribed. A video that has captions only in other languages, when translation and the
original-language fallback are both off, is returned as
LANGUAGE_NOT_FOUNDwith the languages it has.
A video that fails never stops the run: it becomes a dataset item with an error_code,
and the rest of the batch continues.
What you get
Each dataset item is one video.
Video metadata
| Field | Type | Description |
|---|---|---|
metadata.id | string | YouTube video ID |
metadata.title | string | Video title |
metadata.url | string | Canonical watch URL |
metadata.description | string | Full video description |
metadata.duration | number | Duration in seconds |
metadata.view_count | number | Total views |
metadata.like_count | number | Total likes |
metadata.channel | string | Channel display name |
metadata.channel_id | string | Channel ID (UC-prefixed) |
metadata.channel_url | string | Channel URL |
metadata.thumbnail | string | Highest-resolution thumbnail URL |
metadata.upload_date | string | Upload date (YYYYMMDD) |
metadata.tags | array | Creator-set tags |
metadata.categories | array | YouTube categories |
Transcript fields
A field appears only on the items it applies to — a format you did not request is absent, not empty.
| Field | Type | Description |
|---|---|---|
language | string | Language code of the transcript (e.g. en, zh-TW) |
is_auto_generated | boolean | true if YouTube auto-generated the captions |
is_ai_generated | boolean | true if the built-in AI model transcribed the audio |
is_language_fallback | boolean | true when the video's original-language track was delivered instead of your languages |
is_machine_translated | boolean | true when the transcript was machine-translated into your language from the video's original captions |
translated_from | string | Language code of the original captions a machine-translated transcript came from |
requested_languages | array | The languages you asked for, on language-fallback and machine-translated items |
transcript_json | array | Timestamped segments [{start, end, text}]. With wordLevel: true, each segment also has words: [{start, end, text}] — end is estimated for native captions, exact for AI transcriptions |
transcript_text | string | Plain text transcript |
transcript_llm | string | Text with [Music], (laughter) and similar tokens stripped — ready for AI pipelines |
transcript_srt | string | SRT subtitles |
transcript_vtt | string | WebVTT subtitles |
language_probability | number | The AI model's confidence in the language it detected (0–1). AI transcription only; absent when you forced the language |
language_was_forced | boolean | true when forceTranscriptionLanguage was set. AI transcription only |
ai_duration_charged_min | number | AI minutes charged for this video. AI transcription only |
ai_speech_duration_sec | number | Speech detected in the audio, in seconds. AI transcription only |
For a very long transcript, per-word timings are left out, and then the timestamped segments, so the item stays within the platform's size limit for one dataset item. The text formats are always kept.
Error items
| Field | Type | Description |
|---|---|---|
error_code | string | What went wrong — see the table below |
error | string | What happened and what to do about it |
retryable | string | Whether running the video again can help: yes, no or conditional |
url | string | The URL as you gave it, when no metadata could be read |
available_languages | array | The caption languages the video does have, on LANGUAGE_NOT_FOUND and CAPTION_FETCH_FAILED items |
Error items are never charged.
| Error code | Meaning | Retryable |
|---|---|---|
INVALID_URL | The URL is not a YouTube video, playlist or channel | no |
RUN_VIDEO_LIMIT_REACHED | The run already took 500 videos, the most one run transcribes | yes |
URL_EXPANSION_FAILED | A playlist or channel could not be listed | yes |
AGE_RESTRICTED | The video requires age verification | no |
PRIVATE_OR_UNAVAILABLE | The video is private, removed or blocked in the region | no |
LIVE_VIDEO | Live streams cannot be transcribed until they are archived | conditional |
BOT_DETECTION | YouTube asked for human verification | yes |
EXTRACTION_ERROR | The video's information could not be read | yes |
CAPTION_FETCH_FAILED | A caption track exists but could not be downloaded | yes |
LANGUAGE_NOT_FOUND | Captions exist, but not in your languages, and translation, the original-language fallback and AI are all off (or translation failed) | no |
NO_CAPTIONS_AVAILABLE | The video has no captions and AI transcription is off | no |
CAPTIONS_NOT_READY | The video's captions are listed but not generated yet — typical for a recent upload. Run it again later, or enable AI fallback | yes |
AI_FALLBACK_SKIPPED_TOO_LONG | The video is longer than 120 minutes, or than Skip AI for long videos | no |
AI_MINUTES_LIMIT_REACHED | The run's Max AI minutes budget would be exceeded | conditional |
AI_TRANSCRIPTION_FAILED | The audio could not be downloaded or transcribed | yes |
Output examples
Native captions
{"metadata": {"id": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up","channel": "Rick Astley","duration": 213,"upload_date": "20091025"},"language": "en","is_auto_generated": false,"is_ai_generated": false,"transcript_json": [{ "start": 18.5, "end": 21.0, "text": "We're no strangers to love" },{ "start": 21.0, "end": 24.5, "text": "You know the rules and so do I" }],"transcript_text": "We're no strangers to love You know the rules and so do I ..."}
AI transcription
{"metadata": { "id": "jNQXAC9IVRw", "title": "Me at the zoo", "duration": 19 },"language": "en","language_probability": 0.9996,"language_was_forced": false,"is_auto_generated": false,"is_ai_generated": true,"transcript_text": "Alright, so here we are, one of the elephants. ...","ai_duration_charged_min": 1,"ai_speech_duration_sec": 19}
Error
{"error_code": "LANGUAGE_NOT_FOUND","error": "Captions exist, but not in your requested language. Check available_languages and adjust Caption languages.","metadata": { "id": "dQw4w9WgXcQ", "title": "Rick Astley - Never Gonna Give You Up" },"available_languages": ["en", "es", "fr"],"retryable": "no"}
Pricing
Pay per event, with no charge for server time:
- per transcript delivered from captions;
- per minute of video transcribed by AI, billed on the video's full length, rounded up to the minute;
- a start fee per run.
The prices for your plan are on the Pricing page; higher Apify plans pay less per event. There is no separate proxy charge on your account.
Keeping AI cost predictable
- Max AI minutes per run caps the AI minutes a run may use (default 30, up to 600; 0
means the maximum, 600). A video that would exceed it is reported as
AI_MINUTES_LIMIT_REACHED. - Skip AI for long videos leaves out videos longer than the minutes you set. Videos over 120 minutes are always skipped, whatever you set here.
- Captions are always tried first. AI only runs on videos with no captions in any language;
a video captioned in other languages is machine-translated into your language (when
translation is on, the default) or returned as
LANGUAGE_NOT_FOUND, never AI-transcribed.
Advanced options
| Option | UI label | Default | What it does |
|---|---|---|---|
startUrls | YouTube URLs | — | Videos, playlists and channels |
maxResults | Max videos per playlist/channel | 10 | Videos taken from each playlist or channel; single videos ignore it. One run transcribes at most 500 videos; each one past that is returned as RUN_VIDEO_LIMIT_REACHED |
languages | Caption languages | ["en"] | Caption languages in order of preference |
subType | Caption source | "both" | "manual", "auto", or "both" (manual first) |
outputFormats | Output formats | json, text, llm | Which transcript formats each item carries |
wordLevel | Word-level timestamps | false | Per-word timings inside transcript_json |
enableAiFallback | Enable AI transcription | false | Transcribe videos that have no captions |
forceTranscriptionLanguage | AI transcription language | auto-detect | The language the AI transcribes in, instead of the one it detects in the audio |
maxAiMinutes | Max AI minutes per run | 30 | AI minute budget for the run, up to 600; 0 means the maximum (600) |
skipAiFallbackIfLongerThan | Skip AI for long videos (minutes) | 0 (off) | Skip AI for videos longer than this; videos over 120 minutes are skipped regardless |
useAnyAvailableCaptionLanguage | Use captions in any available language | false | Deliver the video's original-language captions when none match your languages |
Integration examples
Python — Apify client
from apify_client import ApifyClientclient = ApifyClient("YOUR_API_TOKEN")run = client.actor("codepoetry/youtube-transcript-ai-scraper").call(run_input={"startUrls": [{"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ"}],"languages": ["en"],"outputFormats": ["json", "llm"],})for item in client.dataset(run["defaultDatasetId"]).iterate_items():if "error_code" in item:print(item["error_code"], item["error"])continueprint(item["metadata"]["title"], item["transcript_llm"][:200])
JavaScript / Node.js
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });const run = await client.actor('codepoetry/youtube-transcript-ai-scraper').call({startUrls: [{ url: 'https://www.youtube.com/watch?v=dQw4w9WgXcQ' }],languages: ['en'],outputFormats: ['json', 'llm'],});const { items } = await client.dataset(run.defaultDatasetId).listItems();for (const item of items) {if (item.error_code) continue;console.log(item.metadata.title, item.transcript_llm.slice(0, 200));}
LangChain / RAG
from apify_client import ApifyClientfrom langchain.docstore.document import Documentclient = ApifyClient("YOUR_API_TOKEN")run = client.actor("codepoetry/youtube-transcript-ai-scraper").call(run_input={"startUrls": [{"url": "https://www.youtube.com/@channelname"}],"outputFormats": ["llm"],"maxResults": 50,})docs = [Document(page_content=item["transcript_llm"],metadata={"source": item["metadata"]["url"], "title": item["metadata"]["title"]},)for item in client.dataset(run["defaultDatasetId"]).iterate_items()if "transcript_llm" in item]
Frequently asked questions
What happens if a video has no captions?
With AI transcription off, the video is reported with NO_CAPTIONS_AVAILABLE. Turn on
Enable AI transcription and its audio is transcribed instead.
A video with captions only in other languages is never transcribed by AI: the model
transcribes the language the video is spoken in, not the one you asked for. Instead, with
Machine-translate captions into your language on (the default), its captions are
machine-translated into your language, timestamps kept. With translation off and no
transcript in your languages deliverable, it is reported with LANGUAGE_NOT_FOUND and the
languages it has in available_languages — add one of those, turn translation back on, or
turn on Use captions in any available language to receive its original-language track.
Does it work for playlists, channels and Shorts?
Yes. A playlist or channel is expanded into its videos (use Max videos per playlist/channel to cap it), and a Shorts URL is treated like any other video.
Which languages are supported?
Native captions: any language YouTube offers captions in. AI transcription: 99 languages.
Does it translate transcripts?
No. Captions are returned in the language the video has them in. An original-language fallback is in the language spoken, and an AI transcription is in the language the model hears (or the one you set in AI transcription language).
Changelog
2026-10-10
- A run stopped early — by you, or when its time limit is reached — now saves every result it has already collected, instead of returning an empty dataset.
- A video whose captions are listed but not generated yet (typical for a recent upload) is
now returned as
CAPTIONS_NOT_READYwith a retry hint, or transcribed with AI when AI fallback is on — instead of being reported asLANGUAGE_NOT_FOUND. - One unreadable caption track no longer fails the whole run; that video comes back as an error item and the rest of the run completes.
Notes
This is an unofficial scraper. It is not affiliated with, authorised by, sponsored by or endorsed by YouTube or Google, and those names are used only to say which public website it reads.