Go to example tasks

Build a RAG Dataset from Y Combinator YouTube Videos

Pulls transcripts of the 20 newest Y Combinator videos as clean text plus timestamped segments you can chunk and cite, with title, publish date, views and word count for each. The run also saves one combined markdown file of every transcript. Use it to seed a startup-advice knowledge base, or paste your own channel link.

Try for free
YouTube Transcript Scraper – Subtitles to Text for LLM & RAG
YouTube Transcript Scraper – Subtitles to Text for LLM & RAGinovaflow/youtube-transcript-scraper
Video
Channel
Length
Lang
+7 fields
Text
Number
Boolean
List
Object

Input

YouTube URLs (videos, Shorts, playlists, channels)
url:https://www.youtube.com/@ycombinator
Max videos per playlist / channel / search:20
Include publish date, likes & category:true

Output fields

Video
Channel
Length
Lang
Auto-generated
Words
Published
Views
Status
URL
Transcript

How it works

Sign up on Apify01

Create your Apify account to access the YouTube Transcript Scraper – Subtitles to Text for LLM & RAG.

Start the run02

The Actor will start running based on the input automatically.

Receive the output03

Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.

Integrate into your workflow04

The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.

Image

Integrate Actor directly into your workflow

Choose from one of 100+ integration options we provide or integrate via API

Webhook

Webhook

n8n

n8n

Make

Make

Zapier

Zapier

Airbyte

Airbyte

Keboola

Keboola

IFTTT

IFTTT

Hubspot

Hubspot

GDrive

GDrive

Gmail

Gmail

Apify MCP

Apify MCP

GitHub

GitHub

Slack

Slack

LangChain

LangChain

LlamaIndex

LlamaIndex

Flowise

Flowise

Pinecone

Pinecone

OpenAI

OpenAI

Mastra

Mastra

Clay

Clay