YouTube Transcripts & Metadata Scraper avatar

YouTube Transcripts & Metadata Scraper

Pricing

from $3.00 / 1,000 transcript extracteds

Go to Apify Store
YouTube Transcripts & Metadata Scraper

YouTube Transcripts & Metadata Scraper

Extract reliable batch transcripts and metadata from YouTube videos, Shorts, channels and playlists. Multi-language, AI-ready output (text, markdown, SRT, VTT). Built for RAG, LLM and agent pipelines.

Pricing

from $3.00 / 1,000 transcript extracteds

Rating

0.0

(0)

Developer

Eric Z. Casaucao

Eric Z. Casaucao

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

15 hours ago

Last modified

Share

YouTube transcript API, captions and subtitles — extract clean, timestamped transcripts and metadata from YouTube videos, Shorts, channels and playlists, reliably and at scale. Output as text, markdown, SRT or VTT, ready for RAG, LLM and AI agent pipelines.

What does YouTube Transcript Scraper do?

It turns any public YouTube video into structured text: the full transcript segmented by timestamp, plus video metadata. It works on single videos, full URL batches, Shorts, entire playlists and the latest videos of a channel, in the languages you choose — and it can translate captions when a language is missing.

It does not transcribe audio: it extracts captions/subtitles that already exist on the video (auto-generated or uploaded). It's a reliable YouTube transcript API alternative that needs no Google API key and no OAuth.

Why scrape YouTube transcripts?

Video is where a huge amount of knowledge lives, but it isn't searchable, embeddable or queryable. Transcripts make it all three.

  • Feed RAG and LLM pipelines. Drop transcripts straight into a vector store, a LangChain or LlamaIndex loader, or an agent's context window. Each row is AI-ready with segments, fullText and markdown.
  • Repurpose video into text. Turn webinars, interviews and tutorials into blog posts, newsletters, subtitles and clip scripts.
  • Research and analyse at scale. Pull hundreds of videos on a topic and run topic classification, entity extraction, sentiment or search over the text.
  • Publish transcripts for SEO and accessibility. Make video content indexable and provide a text alternative for viewers who need it.
  • Monitor channels and playlists. Run on a schedule to keep a keyword or knowledge base fresh.

Built for reliability and scale

  • Batch thousands of URLs in one run, with configurable concurrency.
  • Per-item status/error — one bad video never breaks the run.
  • Retries with exponential backoff and Apify Proxy support to avoid IP blocks.
  • Multiple output formats: segments, text, markdown, SRT, VTT.

What data can it extract?

Every video produces one dataset row:

FieldTypeDescription
videoIdstringYouTube video ID.
titlestringVideo title.
channelName / channelIdstringPublishing channel.
publishDatestringPublication date.
durationSecnumberDuration in seconds.
viewCountnumberView count at run time.
tagsarrayVideo keywords/tags.
descriptionstringVideo description.
thumbnailUrlstringMax-resolution thumbnail.
languagestringTranscript language (ISO 639-1).
isGeneratedbooleanWhether captions are auto-generated.
translatedFromstringSource language if translated.
segmentsarray{start, end, text} timestamped segments.
fullTextstringFull transcript as one string.
markdownstringTranscript as markdown with the video link.
srt / vttstringSubtitle formats (when requested).
status / errorstringPer-item outcome.

How to scrape YouTube transcripts (step-by-step)

  1. Open the Actor in Apify Console or call it via API.
  2. Add Video URLs (watch/youtu.be/Shorts), or add Channel URLs / Playlist URLs to expand them automatically.
  3. Set Language priority (e.g. ["en", "es"]).
  4. Choose the Output format (segments, text, markdown, srt, vtt).
  5. Start the run. Watch the log and results in the Output tab.
  6. Download the dataset as JSON, CSV or Excel, or read it via the API. Schedule the run to keep your data fresh.

How much does it cost to scrape YouTube?

You pay per event, not per compute time. Prices:

  • $3 per 1,000 transcripts (transcript-extracted)
  • $0.50 per 1,000 videos of metadata (metadata-enriched)

Failed videos and videos without captions are not charged. On Apify's free plan you get monthly credits to test, and this Actor caps free-plan runs to the first few videos, so you can validate the output before paying.

Example: 1,000 videos with transcripts and metadata ≈ $3.50.

Input

See the Input tab for full configuration. Key fields:

FieldTypeDefaultDescription
videoUrlsarray—Video/Shorts URLs (watch, youtu.be, shorts).
channelUrlsarray—Channel URLs, expanded to their latest videos.
playlistUrlsarray—Playlist URLs, expanded to their videos.
languagesarray["en"]ISO 639-1 codes in priority order.
includeAutoGeneratedbooleantrueFall back to auto-generated captions.
outputFormatstringsegmentssegments, text, markdown, srt, vtt.
includeMetadatabooleantrueCollect video metadata.
maxItemsinteger0Max videos. 0 = unlimited (free plan capped).
maxRetriesinteger3Per-request retries.
concurrencyinteger5Parallel video workers.
proxyConfigurationobjectApify ProxyApify Proxy (auto) by default; set Residential at high volume if blocked.

Output

One JSON object per video. Example:

{
"videoId": "dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video)",
"channelName": "Rick Astley",
"publishDate": "2009-10-25",
"durationSec": 213,
"viewCount": 1820610046,
"language": "en",
"isGenerated": false,
"segments": [{"start": 1.36, "end": 3.04, "text": "We're no strangers to love"}],
"fullText": "We're no strangers to love ...",
"status": "ok",
"error": null
}

Integrate (API / SDK / MCP)

from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("youtube-transcript-scraper").call(run_input={
"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"languages": ["en", "es"],
"outputFormat": "srt",
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item["videoId"], item["status"])

This Actor is callable as a tool for AI agents through the Apify MCP server (https://mcp.apify.com), so an agent can fetch a transcript on demand — no glue code.

FAQ

Can I download subtitles as SRT or VTT?

Yes. Set outputFormat to srt or vtt and each row includes subtitles ready to use.

Does it work with YouTube Shorts and playlists?

Yes. Shorts URLs are accepted and playlist/channel URLs are expanded into their videos automatically.

Can I get transcripts in another language?

Yes. Pass a language priority list; if the language isn't available, the transcript is translated when possible (the translatedFrom field tells you the source language).

Does it transcribe videos without captions?

No. It extracts existing captions/subtitles; it does not perform speech-to-text.

How many videos can I process per run?

There is no fixed limit — batch as many URLs as you need, subject to maxItems and your plan.

Disclaimers & support

Use this Actor only for lawful purposes and content you are permitted to process. Transcripts may contain copyrighted material or personal data; keep source attribution and comply with YouTube's terms and your local laws. For bug reports or feature requests, open an issue on the Actor's page.