Build a YouTube transcript corpus for RAG
Created by
Khadin Akbar
Collect transcript text and LLM-ready context from public YouTube videos to support retrieval, indexing, or grounded analysis.
YouTube Transcript and Subtitle Data Scraperkhadinakbar/youtube-transcript-extractor
Video title
Full transcript (plain text)
Word count
Estimated token count
+2 fieldsTextNumberBooleanListObject
Input
Search Query:customer support best practices
Max videos:5
Transcript language:en
Transcript format:all
Include video metadata:false
Apify proxy
Output fields
Video title
Full transcript (plain text)
Word count
Estimated token count
Language used
Auto-generated captions
Sign up on Apify01
Create your Apify account to access the YouTube Transcript and Subtitle Data Scraper.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
