YouTube Transcripts & Metadata Scraper
Pricing
from $3.00 / 1,000 transcript extracteds
YouTube Transcripts & Metadata Scraper
Extract reliable batch transcripts and metadata from YouTube videos, Shorts, channels and playlists. Multi-language, AI-ready output (text, markdown, SRT, VTT). Built for RAG, LLM and agent pipelines.
Pricing
from $3.00 / 1,000 transcript extracteds
Rating
0.0
(0)
Developer
Eric Z. Casaucao
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
15 hours ago
Last modified
Categories
Share
YouTube transcript API, captions and subtitles — extract clean, timestamped transcripts and metadata from YouTube videos, Shorts, channels and playlists, reliably and at scale. Output as text, markdown, SRT or VTT, ready for RAG, LLM and AI agent pipelines.
What does YouTube Transcript Scraper do?
It turns any public YouTube video into structured text: the full transcript segmented by timestamp, plus video metadata. It works on single videos, full URL batches, Shorts, entire playlists and the latest videos of a channel, in the languages you choose — and it can translate captions when a language is missing.
It does not transcribe audio: it extracts captions/subtitles that already exist on the video (auto-generated or uploaded). It's a reliable YouTube transcript API alternative that needs no Google API key and no OAuth.
Why scrape YouTube transcripts?
Video is where a huge amount of knowledge lives, but it isn't searchable, embeddable or queryable. Transcripts make it all three.
- Feed RAG and LLM pipelines. Drop transcripts straight into a vector store, a LangChain or
LlamaIndex loader, or an agent's context window. Each row is AI-ready with
segments,fullTextandmarkdown. - Repurpose video into text. Turn webinars, interviews and tutorials into blog posts, newsletters, subtitles and clip scripts.
- Research and analyse at scale. Pull hundreds of videos on a topic and run topic classification, entity extraction, sentiment or search over the text.
- Publish transcripts for SEO and accessibility. Make video content indexable and provide a text alternative for viewers who need it.
- Monitor channels and playlists. Run on a schedule to keep a keyword or knowledge base fresh.
Built for reliability and scale
- Batch thousands of URLs in one run, with configurable concurrency.
- Per-item
status/error— one bad video never breaks the run. - Retries with exponential backoff and Apify Proxy support to avoid IP blocks.
- Multiple output formats:
segments,text,markdown,SRT,VTT.
What data can it extract?
Every video produces one dataset row:
| Field | Type | Description |
|---|---|---|
videoId | string | YouTube video ID. |
title | string | Video title. |
channelName / channelId | string | Publishing channel. |
publishDate | string | Publication date. |
durationSec | number | Duration in seconds. |
viewCount | number | View count at run time. |
tags | array | Video keywords/tags. |
description | string | Video description. |
thumbnailUrl | string | Max-resolution thumbnail. |
language | string | Transcript language (ISO 639-1). |
isGenerated | boolean | Whether captions are auto-generated. |
translatedFrom | string | Source language if translated. |
segments | array | {start, end, text} timestamped segments. |
fullText | string | Full transcript as one string. |
markdown | string | Transcript as markdown with the video link. |
srt / vtt | string | Subtitle formats (when requested). |
status / error | string | Per-item outcome. |
How to scrape YouTube transcripts (step-by-step)
- Open the Actor in Apify Console or call it via API.
- Add Video URLs (watch/
youtu.be/Shorts), or add Channel URLs / Playlist URLs to expand them automatically. - Set Language priority (e.g.
["en", "es"]). - Choose the Output format (
segments,text,markdown,srt,vtt). - Start the run. Watch the log and results in the Output tab.
- Download the dataset as JSON, CSV or Excel, or read it via the API. Schedule the run to keep your data fresh.
How much does it cost to scrape YouTube?
You pay per event, not per compute time. Prices:
- $3 per 1,000 transcripts (
transcript-extracted) - $0.50 per 1,000 videos of metadata (
metadata-enriched)
Failed videos and videos without captions are not charged. On Apify's free plan you get monthly credits to test, and this Actor caps free-plan runs to the first few videos, so you can validate the output before paying.
Example: 1,000 videos with transcripts and metadata ≈ $3.50.
Input
See the Input tab for full configuration. Key fields:
| Field | Type | Default | Description |
|---|---|---|---|
videoUrls | array | — | Video/Shorts URLs (watch, youtu.be, shorts). |
channelUrls | array | — | Channel URLs, expanded to their latest videos. |
playlistUrls | array | — | Playlist URLs, expanded to their videos. |
languages | array | ["en"] | ISO 639-1 codes in priority order. |
includeAutoGenerated | boolean | true | Fall back to auto-generated captions. |
outputFormat | string | segments | segments, text, markdown, srt, vtt. |
includeMetadata | boolean | true | Collect video metadata. |
maxItems | integer | 0 | Max videos. 0 = unlimited (free plan capped). |
maxRetries | integer | 3 | Per-request retries. |
concurrency | integer | 5 | Parallel video workers. |
proxyConfiguration | object | Apify Proxy | Apify Proxy (auto) by default; set Residential at high volume if blocked. |
Output
One JSON object per video. Example:
{"videoId": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video)","channelName": "Rick Astley","publishDate": "2009-10-25","durationSec": 213,"viewCount": 1820610046,"language": "en","isGenerated": false,"segments": [{"start": 1.36, "end": 3.04, "text": "We're no strangers to love"}],"fullText": "We're no strangers to love ...","status": "ok","error": null}
Integrate (API / SDK / MCP)
from apify_client import ApifyClientclient = ApifyClient("<APIFY_TOKEN>")run = client.actor("youtube-transcript-scraper").call(run_input={"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"languages": ["en", "es"],"outputFormat": "srt",})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item["videoId"], item["status"])
This Actor is callable as a tool for AI agents through the Apify MCP server
(https://mcp.apify.com), so an agent can fetch a transcript on demand — no glue code.
FAQ
Can I download subtitles as SRT or VTT?
Yes. Set outputFormat to srt or vtt and each row includes subtitles ready to use.
Does it work with YouTube Shorts and playlists?
Yes. Shorts URLs are accepted and playlist/channel URLs are expanded into their videos automatically.
Can I get transcripts in another language?
Yes. Pass a language priority list; if the language isn't available, the transcript is
translated when possible (the translatedFrom field tells you the source language).
Does it transcribe videos without captions?
No. It extracts existing captions/subtitles; it does not perform speech-to-text.
How many videos can I process per run?
There is no fixed limit — batch as many URLs as you need, subject to maxItems and your plan.
Disclaimers & support
Use this Actor only for lawful purposes and content you are permitted to process. Transcripts may contain copyrighted material or personal data; keep source attribution and comply with YouTube's terms and your local laws. For bug reports or feature requests, open an issue on the Actor's page.