YouTube Transcript Scraper Pro (Captions + AI Fallback) avatar

YouTube Transcript Scraper Pro (Captions + AI Fallback)

Pricing

from $0.70 / 1,000 transcript extracteds

Go to Apify Store
YouTube Transcript Scraper Pro (Captions + AI Fallback)

YouTube Transcript Scraper Pro (Captions + AI Fallback)

Extract YouTube transcripts at scale without burning through your budget. It starts with free captions whenever they're available, then switches to AI only for videos that don't have them. You stay in control of costs, and the output — JSON, SRT, VTT, plain text, or LLM-ready format

Pricing

from $0.70 / 1,000 transcript extracteds

Rating

5.0

(1)

Developer

CodePoetry

CodePoetry

Maintained by Community

Actor stats

12

Bookmarked

2.1K

Total users

548

Monthly active users

4 hours ago

Last modified

Share

YouTube Transcript Scraper — Captions + AI Speech-to-Text

Extract transcripts from any YouTube video — even when captions don't exist.

Give it a single video, a playlist, or a whole channel. You get each video's transcript in JSON, plain text, SRT, VTT, or an LLM-ready format. When a video has no captions, turn on AI transcription and the built-in speech-to-text model transcribes the audio — no API key, no extra setup. You pay per transcript, not per minute of server time.


Quick start

  1. Click Try for free.

  2. Paste one or more YouTube URLs into YouTube URLs:

    • Single video: youtube.com/watch?v=... or youtu.be/...
    • Playlist: youtube.com/playlist?list=...
    • Channel: youtube.com/@channelname

    Playlists and channels return 10 videos by default — raise Max videos per playlist/channel for more.

  3. Pick your Output formats. Not sure? Plain Text is the words as one block.

  4. AI transcription is off by default. Videos without native captions are reported with NO_CAPTIONS_AVAILABLE unless you turn on Enable AI transcription. When you do, set Max AI minutes per run before processing an unknown playlist or channel.

  5. Click Start, then download results from the Dataset tab as JSON, CSV or Excel, or read them through the API.


How it works

  1. Expand — every URL becomes the videos it names. A channel URL is read from its videos tab, newest first. A video named more than once — as a youtu.be link and in a playlist, say — is transcribed and charged once.
  2. Captions — for each video, the scraper looks for a caption track in your languages, in order: manual captions first (unless you chose auto-generated only), then auto-generated. The first track that downloads is used.
  3. Machine translation (optional, on by default) — with Machine-translate captions into your language on, a video that has captions only in other languages has them machine-translated into your first requested language, keeping every timestamp and charged as a normal transcript. Items carry is_machine_translated: true and translated_from.
  4. Original-language fallback (optional) — with Use captions in any available language on, a video with no track in your languages delivers its track in the language it was spoken in, charged as a normal transcript.
  5. AI transcription (optional) — the audio of a video with no captions at all is transcribed. A video that has captions only in other languages, when translation and the original-language fallback are both off, is returned as LANGUAGE_NOT_FOUND with the languages it has.

A video that fails never stops the run: it becomes a dataset item with an error_code, and the rest of the batch continues.


What you get

Each dataset item is one video.

Video metadata

FieldTypeDescription
metadata.idstringYouTube video ID
metadata.titlestringVideo title
metadata.urlstringCanonical watch URL
metadata.descriptionstringFull video description
metadata.durationnumberDuration in seconds
metadata.view_countnumberTotal views
metadata.like_countnumberTotal likes
metadata.channelstringChannel display name
metadata.channel_idstringChannel ID (UC-prefixed)
metadata.channel_urlstringChannel URL
metadata.thumbnailstringHighest-resolution thumbnail URL
metadata.upload_datestringUpload date (YYYYMMDD)
metadata.tagsarrayCreator-set tags
metadata.categoriesarrayYouTube categories

Transcript fields

A field appears only on the items it applies to — a format you did not request is absent, not empty.

FieldTypeDescription
languagestringLanguage code of the transcript (e.g. en, zh-TW)
is_auto_generatedbooleantrue if YouTube auto-generated the captions
is_ai_generatedbooleantrue if the built-in AI model transcribed the audio
is_language_fallbackbooleantrue when the video's original-language track was delivered instead of your languages
is_machine_translatedbooleantrue when the transcript was machine-translated into your language from the video's original captions
translated_fromstringLanguage code of the original captions a machine-translated transcript came from
requested_languagesarrayThe languages you asked for, on language-fallback and machine-translated items
transcript_jsonarrayTimestamped segments [{start, end, text}]. With wordLevel: true, each segment also has words: [{start, end, text}] — end is estimated for native captions, exact for AI transcriptions
transcript_textstringPlain text transcript
transcript_llmstringText with [Music], (laughter) and similar tokens stripped — ready for AI pipelines
transcript_srtstringSRT subtitles
transcript_vttstringWebVTT subtitles
language_probabilitynumberThe AI model's confidence in the language it detected (0–1). AI transcription only; absent when you forced the language
language_was_forcedbooleantrue when forceTranscriptionLanguage was set. AI transcription only
ai_duration_charged_minnumberAI minutes charged for this video. AI transcription only
ai_speech_duration_secnumberSpeech detected in the audio, in seconds. AI transcription only

For a very long transcript, per-word timings are left out, and then the timestamped segments, so the item stays within the platform's size limit for one dataset item. The text formats are always kept.

Error items

FieldTypeDescription
error_codestringWhat went wrong — see the table below
errorstringWhat happened and what to do about it
retryablestringWhether running the video again can help: yes, no or conditional
urlstringThe URL as you gave it, when no metadata could be read
available_languagesarrayThe caption languages the video does have, on LANGUAGE_NOT_FOUND and CAPTION_FETCH_FAILED items

Error items are never charged.

Error codeMeaningRetryable
INVALID_URLThe URL is not a YouTube video, playlist or channelno
RUN_VIDEO_LIMIT_REACHEDThe run already took 500 videos, the most one run transcribesyes
URL_EXPANSION_FAILEDA playlist or channel could not be listedyes
AGE_RESTRICTEDThe video requires age verificationno
PRIVATE_OR_UNAVAILABLEThe video is private, removed or blocked in the regionno
LIVE_VIDEOLive streams cannot be transcribed until they are archivedconditional
BOT_DETECTIONYouTube asked for human verificationyes
EXTRACTION_ERRORThe video's information could not be readyes
CAPTION_FETCH_FAILEDA caption track exists but could not be downloadedyes
LANGUAGE_NOT_FOUNDCaptions exist, but not in your languages, and translation, the original-language fallback and AI are all off (or translation failed)no
NO_CAPTIONS_AVAILABLEThe video has no captions and AI transcription is offno
CAPTIONS_NOT_READYThe video's captions are listed but not generated yet — typical for a recent upload. Run it again later, or enable AI fallbackyes
AI_FALLBACK_SKIPPED_TOO_LONGThe video is longer than 120 minutes, or than Skip AI for long videosno
AI_MINUTES_LIMIT_REACHEDThe run's Max AI minutes budget would be exceededconditional
AI_TRANSCRIPTION_FAILEDThe audio could not be downloaded or transcribedyes

Output examples

Native captions

{
"metadata": {
"id": "dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up",
"channel": "Rick Astley",
"duration": 213,
"upload_date": "20091025"
},
"language": "en",
"is_auto_generated": false,
"is_ai_generated": false,
"transcript_json": [
{ "start": 18.5, "end": 21.0, "text": "We're no strangers to love" },
{ "start": 21.0, "end": 24.5, "text": "You know the rules and so do I" }
],
"transcript_text": "We're no strangers to love You know the rules and so do I ..."
}

AI transcription

{
"metadata": { "id": "jNQXAC9IVRw", "title": "Me at the zoo", "duration": 19 },
"language": "en",
"language_probability": 0.9996,
"language_was_forced": false,
"is_auto_generated": false,
"is_ai_generated": true,
"transcript_text": "Alright, so here we are, one of the elephants. ...",
"ai_duration_charged_min": 1,
"ai_speech_duration_sec": 19
}

Error

{
"error_code": "LANGUAGE_NOT_FOUND",
"error": "Captions exist, but not in your requested language. Check available_languages and adjust Caption languages.",
"metadata": { "id": "dQw4w9WgXcQ", "title": "Rick Astley - Never Gonna Give You Up" },
"available_languages": ["en", "es", "fr"],
"retryable": "no"
}

Pricing

Pay per event, with no charge for server time:

  • per transcript delivered from captions;
  • per minute of video transcribed by AI, billed on the video's full length, rounded up to the minute;
  • a start fee per run.

The prices for your plan are on the Pricing page; higher Apify plans pay less per event. There is no separate proxy charge on your account.

Keeping AI cost predictable

  • Max AI minutes per run caps the AI minutes a run may use (default 30, up to 600; 0 means the maximum, 600). A video that would exceed it is reported as AI_MINUTES_LIMIT_REACHED.
  • Skip AI for long videos leaves out videos longer than the minutes you set. Videos over 120 minutes are always skipped, whatever you set here.
  • Captions are always tried first. AI only runs on videos with no captions in any language; a video captioned in other languages is machine-translated into your language (when translation is on, the default) or returned as LANGUAGE_NOT_FOUND, never AI-transcribed.

Advanced options

OptionUI labelDefaultWhat it does
startUrlsYouTube URLs—Videos, playlists and channels
maxResultsMax videos per playlist/channel10Videos taken from each playlist or channel; single videos ignore it. One run transcribes at most 500 videos; each one past that is returned as RUN_VIDEO_LIMIT_REACHED
languagesCaption languages["en"]Caption languages in order of preference
subTypeCaption source"both""manual", "auto", or "both" (manual first)
outputFormatsOutput formatsjson, text, llmWhich transcript formats each item carries
wordLevelWord-level timestampsfalsePer-word timings inside transcript_json
enableAiFallbackEnable AI transcriptionfalseTranscribe videos that have no captions
forceTranscriptionLanguageAI transcription languageauto-detectThe language the AI transcribes in, instead of the one it detects in the audio
maxAiMinutesMax AI minutes per run30AI minute budget for the run, up to 600; 0 means the maximum (600)
skipAiFallbackIfLongerThanSkip AI for long videos (minutes)0 (off)Skip AI for videos longer than this; videos over 120 minutes are skipped regardless
useAnyAvailableCaptionLanguageUse captions in any available languagefalseDeliver the video's original-language captions when none match your languages

Integration examples

Python — Apify client

from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("codepoetry/youtube-transcript-ai-scraper").call(
run_input={
"startUrls": [{"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ"}],
"languages": ["en"],
"outputFormats": ["json", "llm"],
}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
if "error_code" in item:
print(item["error_code"], item["error"])
continue
print(item["metadata"]["title"], item["transcript_llm"][:200])

JavaScript / Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });
const run = await client.actor('codepoetry/youtube-transcript-ai-scraper').call({
startUrls: [{ url: 'https://www.youtube.com/watch?v=dQw4w9WgXcQ' }],
languages: ['en'],
outputFormats: ['json', 'llm'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) {
if (item.error_code) continue;
console.log(item.metadata.title, item.transcript_llm.slice(0, 200));
}

LangChain / RAG

from apify_client import ApifyClient
from langchain.docstore.document import Document
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("codepoetry/youtube-transcript-ai-scraper").call(
run_input={
"startUrls": [{"url": "https://www.youtube.com/@channelname"}],
"outputFormats": ["llm"],
"maxResults": 50,
}
)
docs = [
Document(
page_content=item["transcript_llm"],
metadata={"source": item["metadata"]["url"], "title": item["metadata"]["title"]},
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items()
if "transcript_llm" in item
]

Frequently asked questions

What happens if a video has no captions?

With AI transcription off, the video is reported with NO_CAPTIONS_AVAILABLE. Turn on Enable AI transcription and its audio is transcribed instead.

A video with captions only in other languages is never transcribed by AI: the model transcribes the language the video is spoken in, not the one you asked for. Instead, with Machine-translate captions into your language on (the default), its captions are machine-translated into your language, timestamps kept. With translation off and no transcript in your languages deliverable, it is reported with LANGUAGE_NOT_FOUND and the languages it has in available_languages — add one of those, turn translation back on, or turn on Use captions in any available language to receive its original-language track.

Does it work for playlists, channels and Shorts?

Yes. A playlist or channel is expanded into its videos (use Max videos per playlist/channel to cap it), and a Shorts URL is treated like any other video.

Which languages are supported?

Native captions: any language YouTube offers captions in. AI transcription: 99 languages.

Does it translate transcripts?

No. Captions are returned in the language the video has them in. An original-language fallback is in the language spoken, and an AI transcription is in the language the model hears (or the one you set in AI transcription language).


Changelog

2026-10-10

  • A run stopped early — by you, or when its time limit is reached — now saves every result it has already collected, instead of returning an empty dataset.
  • A video whose captions are listed but not generated yet (typical for a recent upload) is now returned as CAPTIONS_NOT_READY with a retry hint, or transcribed with AI when AI fallback is on — instead of being reported as LANGUAGE_NOT_FOUND.
  • One unreadable caption track no longer fails the whole run; that video comes back as an error item and the rest of the run completes.

Notes

This is an unofficial scraper. It is not affiliated with, authorised by, sponsored by or endorsed by YouTube or Google, and those names are used only to say which public website it reads.