YouTube All-in-One Downloader & Scraper
Pricing
from $90.00 / 1,000 video downloads
YouTube All-in-One Downloader & Scraper
Download YouTube videos, Shorts, playlists, and channels as MP4. Up to 10 concurrent downloads with no browser needed. Extract comments, captions, and rich metadata. Metadata-only mode for fast, cheap research. Quality selection with automatic fallback. From $0.10/video.
Pricing
from $90.00 / 1,000 video downloads
Rating
0.0
(0)
Developer
jy-labs
Maintained by CommunityActor stats
0
Bookmarked
115
Total users
14
Monthly active users
11 days ago
Last modified
Categories
Share
YouTube Data Pipeline for AI/ML Teams
YouTube transcripts, metadata, and downloads in one API call -- built for AI pipelines.
$0.005/transcript | $0.005/metadata | $0.10/video download | No monthly fee | No quota limits | LLM-ready output mode
Extract YouTube transcripts for RAG, embeddings, and fine-tuning. Get structured metadata for research and analytics. Download video files for content analysis. All through a fast, concurrent API with no browser required.
Quick Start
- Click "Try for free" at the top of this page.
- Paste one or more YouTube URLs into
startUrlsand pick aquality. - Click Start. Each video file lands in the Key-Value Store and one JSON row per video lands in the dataset.
Minimal input -- one video at 360p, everything else on defaults:
{"startUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"quality": "360p"}
What comes back (trimmed -- the complete row is in Rich Metadata Output):
{"sourceUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","downloadUrl": "https://api.apify.com/v2/key-value-stores/STORE_ID/records/dQw4w9WgXcQ_video","videoId": "dQw4w9WgXcQ","title": "Never Gonna Give You Up","channelName": "Rick Astley","duration": "3:33","quality": "360p","fileSize": "9.42 MB","viewCount": 1547892341,"uploadDate": "Oct 25, 2009"}
Only need the text? Set outputFormat: "llm_ready" and the actor skips the file download and returns a flat transcript row instead -- see LLM-Ready Output Mode.
Why AI/ML Teams Use This
- LLM-ready transcripts at $0.005/video -- Clean plaintext transcripts ready to feed into OpenAI, Anthropic, or any embedding model. No parsing needed.
- No YouTube API quota limits -- YouTube Data API caps you at 10,000 units/day and doesn't even provide transcripts. This actor has no daily quota.
- Structured JSON output -- Every field is typed and consistent. Drop results directly into vector DBs (Pinecone, Weaviate, Chroma) or data warehouses.
- Batch processing at scale -- Process playlists, channels, or search results. Up to 10 concurrent extractions. Feed entire YouTube channels into your training pipeline.
Cost Comparison
1,000 YouTube transcripts for your RAG pipeline:
| Approach | Cost | Effort |
|---|---|---|
| YouTube Data API | Can't get transcripts | 10K units/day quota limit |
youtube-transcript-api + your server | $5-15/month (server costs) | Setup, maintenance, IP bans |
| This Actor (LLM-ready mode) | $5.00 total | Zero infrastructure, zero maintenance |
Pricing Tiers
| Tier | Price | Best For |
|---|---|---|
Transcript / Metadata extraction (outputFormat: "llm_ready" or downloadVideo: false) | $0.005/video | RAG, embeddings, LLM training, research, analytics |
| Video download | $0.10/video | Archiving, content analysis, multimodal AI |
Video download that could not be delivered (age-restricted, DRM, or over maxFileSizeMb) | $0.005/video | Metadata row only, no file -- billed as metadata extraction, never as a video download |
Plus Apify platform fee (~$0.25-0.50/1,000 videos for compute). Proxy enabled by default: each video is tried on a cheap datacenter proxy first and only falls back to the RESIDENTIAL group (~$8/GB data transfer) when that attempt fails. Apify Free plan includes $5/month in platform credits.
LLM-Ready Output Mode
Set outputFormat: "llm_ready" to get transcripts optimized for AI/ML workflows. This mode automatically enables transcript extraction, disables video download, and charges $0.005/video (same as metadata extraction).
What you get:
{"videoId": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","channelName": "Rick Astley","channelUrl": "http://www.youtube.com/@RickAstleyYT","transcript": "We're no strangers to love You know the rules and so do I A full commitment's what I'm thinking of...","wordCount": 427,"language": "English","languageCode": "en","sourceUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","duration": "3:33","durationSeconds": 213,"viewCount": 1751798914,"uploadDate": "Oct 25, 2009","category": "Music","tags": [],"description": "The official video for 'Never Gonna Give You Up' by Rick Astley..."}
The transcript field contains the full transcript as clean plaintext -- ready to chunk and embed. The wordCount field gives you the token estimate for chunking strategies.
Batch download YouTube videos as MP4 (API)
Pass a list of URLs and the actor downloads each one as an MP4 into the Apify Key-Value Store, then writes one dataset row per video with a downloadUrl pointing at the stored file. Nothing is rendered in a browser -- the actor calls YouTube's InnerTube API directly, so a batch is limited by bandwidth rather than by page loads.
{"startUrls": ["https://www.youtube.com/watch?v=VIDEO_ID_1","https://youtu.be/VIDEO_ID_2","https://www.youtube.com/playlist?list=PLAYLIST_ID"],"quality": "720p","maxVideos": 100,"maxConcurrency": 8,"maxFileSizeMb": 200}
maxConcurrencyvideos are processed in parallel -- default 4, maximum 10. Raise it for throughput, lower it if a run hits its memory limit.maxVideoscaps the whole run at up to 1,000 videos, counted across every URL and search query together.maxFileSizeMbaborts a download mid-stream once the file passes the cap, so one oversized upload cannot drain the proxy budget. Per-quality caps apply on top of it (see Limitations).- Failures are retried with exponential backoff (
maxRequestRetries, default 3). Errors classified as permanent -- private, deleted, or unavailable videos -- are not retried at all. - To start runs from your own code, see Integration Examples for Python, Node.js, and cURL against the Apify API.
Download YouTube comments and captions
Comments, replies, and subtitles are extracted from the same metadata call, so you can collect them without downloading a single video file. Enabling any extraction flag turns metadata extraction on automatically.
{"startUrls": ["https://www.youtube.com/watch?v=VIDEO_ID"],"downloadVideo": false,"extractComments": true,"extractReplies": true,"maxComments": 200,"extractCaptions": true,"downloadCaptions": true,"captionFormat": "srt","captionLanguage": "en"}
- Comments come back as an array with author, author channel URL, text, like count, published time, and reply count.
extractReplies: truenests the reply threads inside each comment.maxCommentsdefaults to 100 and accepts up to 500 per video. - Captions are placed in the row as plain text under
captions. AddingdownloadCaptions: truealso writes an SRT or VTT file to the Key-Value Store and returnscaptionFileUrl. - Language follows
captionLanguage. If that track does not exist, the first available track is used instead.autoTranslateLanguageasks YouTube to translate the track, which is how you get English text out of a Korean or Japanese video. - A video with no caption track returns an empty
captionsarray. The row is still delivered, so a missing transcript never costs you the rest of the metadata.
YouTube metadata extraction without the official API quota
The official YouTube Data API hands every project 10,000 quota units a day, and it does not serve transcripts at all. This actor reads the InnerTube API instead, so there is no daily unit budget to exhaust and no API key to provision. Set downloadVideo: false and each video is processed as a metadata extraction with no file transfer.
{"startUrls": ["https://www.youtube.com/@ChannelHandle"],"downloadVideo": false,"maxVideos": 500,"extractChannelInfo": true,"extractRelatedVideos": true}
A metadata row carries the video id, title, description, channel name and URL, view and like counts, duration, upload date, tags, category, chapters, and thumbnail URL. extractChannelInfo adds subscriber count, banner URL, video count, and join date. extractRelatedVideos adds up to 20 suggested videos with title, channel, view count, and duration.
What replaces the quota, stated plainly: YouTube rate-limits by IP address. The actor routes the first attempt through a cheap datacenter proxy and escalates to the residential group when that attempt is refused, so the constraint becomes a data transfer cost rather than a hard daily ceiling. See Pricing Details for what that costs.
Download YouTube Shorts and playlists
Shorts, playlists, channels, and plain video links can be mixed in a single startUrls list. Playlists and channels are expanded into their video ids before processing, and maxVideos caps how many of them the run takes.
{"startUrls": ["https://www.youtube.com/shorts/SHORT_ID","https://www.youtube.com/playlist?list=PLAYLIST_ID","https://www.youtube.com/@ChannelHandle"],"quality": "highest","maxVideos": 50}
- Shorts download exactly like a normal video. A row is flagged
isShort: truewhen the URL you passed uses the/shorts/form. Shorts discovered by expanding a playlist or channel are addressed by theirwatch?v=URL, so they are not flagged. - Playlists add
playlistIndexandplaylistTitleto every row, so the original ordering survives into your dataset. - A URL carrying both
v=andlist=is treated as that single video, not as the whole playlist. Use the bareplaylist?list=form when you want the entire list. - Channels are accepted as
@handle,/channel/ID, and/c/Name. Pair them withmonitorModeandstateStoreNameto process only the videos that are new since the last scheduled run.
Integration Examples
Python + OpenAI Embeddings
from apify_client import ApifyClientimport openaiclient = ApifyClient("YOUR_API_TOKEN")# Extract transcripts for RAGrun = client.actor("jy-labs/youtube-all-in-one-downloader-scraper").call(run_input={"startUrls": ["https://www.youtube.com/watch?v=VIDEO_ID"],"outputFormat": "llm_ready","proxyConfiguration": {"useApifyProxy": True, "apifyProxyGroups": ["RESIDENTIAL"]},})for item in client.dataset(run["defaultDatasetId"]).iterate_items():transcript = item["transcript"]if transcript:# Feed to OpenAI embeddingsembedding = openai.embeddings.create(model="text-embedding-3-small",input=transcript)# Store in your vector DB (Pinecone, Weaviate, Chroma, etc.)print(f"Embedded: {item['title']} ({item['wordCount']} words)")
Python -- Batch Transcripts for Fine-Tuning
from apify_client import ApifyClientimport jsonclient = ApifyClient("YOUR_API_TOKEN")# Extract transcripts from an entire playlistrun = client.actor("jy-labs/youtube-all-in-one-downloader-scraper").call(run_input={"startUrls": ["https://www.youtube.com/playlist?list=YOUR_PLAYLIST_ID"],"outputFormat": "llm_ready","maxVideos": 100,"captionLanguage": "en","proxyConfiguration": {"useApifyProxy": True, "apifyProxyGroups": ["RESIDENTIAL"]},})# Save as JSONL for fine-tuningwith open("training_data.jsonl", "w") as f:for item in client.dataset(run["defaultDatasetId"]).iterate_items():if item.get("transcript"):f.write(json.dumps({"text": item["transcript"],"metadata": {"title": item["title"],"channel": item["channelName"],"video_id": item["videoId"],"word_count": item["wordCount"],}}) + "\n")
Node.js
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });const run = await client.actor('jy-labs/youtube-all-in-one-downloader-scraper').call({startUrls: ['https://www.youtube.com/watch?v=VIDEO_ID'],outputFormat: 'llm_ready',proxyConfiguration: { useApifyProxy: true, apifyProxyGroups: ['RESIDENTIAL'] },});const { items } = await client.dataset(run.defaultDatasetId).listItems();for (const item of items) {console.log(`${item.title}: ${item.wordCount} words`);console.log(`Transcript: ${item.transcript.substring(0, 100)}...`);// Feed item.transcript to your embedding pipeline...}
cURL (REST API)
# Start a runcurl "https://api.apify.com/v2/acts/jy-labs~youtube-all-in-one-downloader-scraper/runs?token=YOUR_API_TOKEN" \-X POST \-H "Content-Type: application/json" \-d '{"startUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"outputFormat": "llm_ready","proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }}'# Fetch results (after run completes)curl "https://api.apify.com/v2/datasets/DATASET_ID/items?token=YOUR_API_TOKEN"
vs YouTube Data API
| This Actor | YouTube Data API | |
|---|---|---|
| Daily quota | No limit | 10,000 units/day |
| Transcripts | Built-in | Not available |
| Comments | Built-in with replies | Available (costs quota) |
| Video download | Built-in | Not available |
| Auth required | No | Yes (OAuth/API key) |
| Output format | Structured JSON, LLM-ready mode | Nested JSON with pagination tokens |
| Setup time | ~1 minute | ~30 minutes (GCP project + OAuth) |
vs Competing Apify Actors (Streamers)
| This Actor | Streamers ($30/mo actors) | |
|---|---|---|
| Pricing model | Pay per video, no monthly fee | $30/month subscription + usage |
| Transcript extraction | $0.005/video | Not available as separate tier |
| Metadata only | $0.005/video | Not available (pay full price) |
| LLM-ready output | Built-in | Not available |
| Video download | $0.10/video | Requires $30/mo subscription |
| SRT/VTT file download | Built-in | Not available |
| Chapter extraction | Built-in | Not available |
| Channel info (subscribers) | Built-in | Not available |
| Comment replies | Built-in | Not available |
| Thumbnail download | Built-in | Not available |
| Webhook notification | Built-in | Not available |
| YouTube search | Built-in | Not available |
| Smart error classification | 5 categories + adaptive retry | Basic |
Features Overview
Data Extraction
- Transcripts/captions -- Full plaintext transcripts in 100+ languages. LLM-ready mode returns clean text optimized for chunking and embedding.
- Caption file download (SRT/VTT) -- Download subtitle files in industry-standard SRT or VTT format for video annotation or training data.
- Rich metadata -- Title, description, channel, views, likes, duration, upload date, tags, thumbnail URL, chapters, category, and more.
- Comments with replies -- Top comments with author, likes, timestamps, and full reply threads. Up to 500 comments per video.
- Channel info -- Subscriber count, description, banner URL, total video count, and join date.
- Related videos -- Up to 20 related video suggestions plus video category.
- Chapter markers -- Auto-extracted chapter titles and timestamps for content segmentation.
Processing
- Concurrent processing -- Up to 10 videos in parallel for fast batch extraction.
- Playlist and channel expansion -- Automatically discovers and processes all videos from playlists and channels.
- YouTube search -- Search by keywords with
searchQueriesto collect videos without manual URL hunting. - Quality selection -- Choose from highest, 1080p, 720p, 480p, 360p, or audio-only (M4A).
- Automatic quality fallback -- If requested resolution is unavailable, selects the nearest lower quality.
- Retry with exponential backoff -- Automatic retries with proxy rotation for reliable large-batch runs.
- Smart error classification -- Five error categories (fatal, retryable, rate_limited, geo_blocked, age_restricted) with adaptive retry logic.
Output and Integration
- LLM-ready mode -- Set
outputFormat: "llm_ready"for transcripts optimized for AI pipelines at $0.005/video. - Metadata-only mode -- Set
downloadVideo: falsefor fast, cheap metadata extraction at $0.005/video. - Configurable output -- Set
extractMetadata: falsefor lightweight output (sourceUrl + downloadUrl only). - Thumbnail download -- Highest-quality thumbnail image stored in Key-Value Store.
- Custom filename template -- Name files using placeholders:
{videoId},{title},{quality},{channelName},{date},{type}. - Webhook notification -- POST notification to any URL when processing completes.
- Full REST API -- Call programmatically from Python, Node.js, cURL, or any HTTP client.
Input Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
startUrls | string[] | (optional) | YouTube URLs to process. Supports videos, Shorts, playlists, and channels. At least one of startUrls or searchQueries is required. |
outputFormat | string | "default" | Output mode: "default" or "llm_ready". LLM-ready mode auto-enables transcripts, disables video download, and charges $0.005/video. |
quality | string | "highest" | Video quality: highest, 1080p, 720p, 480p, 360p, or audio_only |
maxFileSizeMb | integer | 200 | Max size per downloaded file in MB. Larger videos are skipped; with extractMetadata on (default) they are delivered as metadata-only ($0.005/video instead of $0.10), otherwise they count as failed. Keeps proxy data transfer costs predictable. |
downloadVideo | boolean | true | Download the video file. Set to false for metadata-only mode ($0.005/video). |
maxConcurrency | integer | 4 | Parallel processing (1--10). Higher values = faster but more memory. |
maxRequestRetries | integer | 3 | Retry attempts per video before marking as failed (0--10). |
includeFailedVideos | boolean | false | Include failed videos in output with error details for debugging. |
extractMetadata | boolean | true | Include full metadata. Auto-enabled when captions, comments, or downloadVideo: false is set. Set to false for lightweight output. |
extractCaptions | boolean | false | Extract captions/subtitles as full transcript text. |
captionLanguage | string | "en" | Preferred caption language code (e.g., en, ko, ja, es). Falls back to first available. |
extractComments | boolean | false | Extract top comments (author, text, likes, timestamp). |
maxComments | integer | 100 | Max comments per video (1--500). |
maxVideos | integer | 100 | Max videos from playlists/channels (1--500). |
proxyConfiguration | object | Apify Proxy, RESIDENTIAL | Proxy settings. Keep RESIDENTIAL selected -- the actor sends the first attempt for each video through a cheap datacenter proxy and escalates to the groups you pick here only when that attempt fails, so RESIDENTIAL is the fallback that keeps runs reliable. Disabling the proxy removes the fallback and will make most runs fail. The RESIDENTIAL fallback incurs data transfer costs (~$8/GB). |
searchQueries | string[] | (optional) | Search YouTube by keywords. Found videos are added to the processing queue. |
downloadThumbnail | boolean | false | Download highest-quality thumbnail to Key-Value Store. |
downloadCaptions | boolean | false | Download caption file to Key-Value Store. Use captionFormat for SRT or VTT. |
captionFormat | string | "srt" | Caption file format: srt (SubRip) or vtt (WebVTT). Only when downloadCaptions is true. |
extractReplies | boolean | false | Fetch full reply threads for each comment. Requires extractComments: true. |
extractChannelInfo | boolean | false | Extract channel details: subscriber count, description, banner, video count, join date. |
extractRelatedVideos | boolean | false | Extract up to 20 related videos and video category. |
filenameTemplate | string | "{videoId}_{type}" | Custom filename with placeholders: {videoId}, {title}, {quality}, {channelName}, {date}, {type}. |
webhookUrl | string | (optional) | URL to receive POST notification when the run completes. |
Output Schema
LLM-Ready Output (outputFormat: "llm_ready")
Optimized for AI pipelines. Auto-enables transcript extraction and disables video download. Returns a flat structure with clean transcript text, word count, and essential metadata only.
{"videoId": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","channelName": "Rick Astley","channelUrl": "http://www.youtube.com/@RickAstleyYT","transcript": "We're no strangers to love You know the rules and so do I A full commitment's what I'm thinking of...","wordCount": 427,"language": "English","languageCode": "en","sourceUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","duration": "3:33","durationSeconds": 213,"viewCount": 1751798914,"uploadDate": "Oct 25, 2009","category": "Music","tags": [],"description": "The official video for 'Never Gonna Give You Up' by Rick Astley..."}
Rich Metadata Output (default)
Full details including download URL, comments, channel info, and related videos when enabled.
{"sourceUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","downloadUrl": "https://api.apify.com/v2/key-value-stores/STORE_ID/records/dQw4w9WgXcQ_video","videoId": "dQw4w9WgXcQ","title": "Rick Astley - Never Gonna Give You Up","description": "The official video for 'Never Gonna Give You Up' by Rick Astley...","channelName": "Rick Astley","channelUrl": "https://www.youtube.com/channel/UCuAXFkgsw1L7xaCfnd5JJOw","viewCount": 1500000000,"likeCount": 16000000,"duration": "3:33","durationSeconds": 213,"uploadDate": "Oct 25, 2009","thumbnailUrl": "https://i.ytimg.com/vi/dQw4w9WgXcQ/maxresdefault.jpg","thumbnailDownloadUrl": "https://api.apify.com/v2/key-value-stores/STORE_ID/records/dQw4w9WgXcQ_thumbnail","quality": "720p","fileSize": "11.28 MB","category": "Music","chapters": [{ "title": "Intro", "startTime": "0:00", "startTimeSeconds": 0 }],"captions": [{"language": "English","languageCode": "en","text": "We're no strangers to love You know the rules and so do I..."}],"captionFileUrl": "https://api.apify.com/v2/key-value-stores/STORE_ID/records/dQw4w9WgXcQ_captions.srt","comments": [{"author": "YouTube User","authorChannelUrl": "https://www.youtube.com/channel/UC...","text": "This song is timeless!","likes": 42000,"publishedTime": "2 years ago","replyCount": 150,"replies": [{"author": "Another User","text": "Agreed, classic forever!","likes": 1200,"publishedTime": "1 year ago"}]}],"channelInfo": {"channelId": "@RickAstleyYT","subscriberCount": "4.47M subscribers","description": "Official YouTube channel of Rick Astley...","bannerUrl": "https://yt3.googleusercontent.com/...","videoCount": "402 videos","joinedDate": "Joined Feb 2, 2015"},"relatedVideos": [{"videoId": "yPYZpwSpKmA","title": "Rick Astley - Together Forever","channelName": "Rick Astley","viewCount": "198M views","duration": "3:24"}],"tags": ["rick astley", "never gonna give you up", "official video"],"isLive": false,"isShort": false,"playlistIndex": 1,"playlistTitle": "My Playlist"}
Note: playlistIndex and playlistTitle appear only for videos from playlists. isShort is true when the input URL uses the /shorts/ format.
Lightweight Output (extractMetadata: false)
{"sourceUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","downloadUrl": "https://api.apify.com/v2/key-value-stores/STORE_ID/records/dQw4w9WgXcQ_video"}
Metadata-Only Row (download requested but not delivered)
When downloadVideo is on but the file cannot be delivered -- age-restricted, private, DRM-protected, region-blocked, or larger than maxFileSizeMb -- the actor does not drop the video. With extractMetadata: true (the default) it returns the full metadata row with downloadUrl: null and a downloadSkippedReason marker, and bills it at the metadata extraction rate ($0.005) instead of the video download rate.
{"sourceUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","downloadUrl": null,"videoId": "dQw4w9WgXcQ","title": "Never Gonna Give You Up","channelName": "Rick Astley","viewCount": 1547892341,"duration": "3:33","quality": "none","fileSize": "","downloadSkippedReason": "download_failed"}
Every other metadata field (description, captions, comments, tags, chapters, ...) is populated exactly as in the rich output above. downloadSkippedReason is "download_failed" when every streaming client was blocked, and "size_limit_exceeded" when the file exceeded maxFileSizeMb. The field is absent from successful downloads and from metadata-only runs (downloadVideo: false), so downloadSkippedReason is a reliable filter for "metadata delivered, file not delivered". These rows are not marked status: "failed" -- that status is reserved for videos where nothing at all was delivered.
Failed Video Output (includeFailedVideos: true)
{"sourceUrl": "https://www.youtube.com/watch?v=INVALID_ID","downloadUrl": null,"error": "This video is unavailable","status": "failed"}
Pricing Details
This actor uses pay-per-event pricing. No monthly subscription. Because it uses the InnerTube API directly (no browser), it runs faster and cheaper than browser-based alternatives.
| Event | Price | When |
|---|---|---|
| Transcript / Metadata extraction | $0.005/video | outputFormat: "llm_ready" or downloadVideo: false |
| Video download | $0.10/video | Default (with video file) |
| Video download that could not be delivered | $0.005/video | Metadata delivered but no file (row carries downloadSkippedReason) -- billed as metadata extraction, never as a video download |
Estimated costs by volume:
| Usage | Actor Fee | Best For |
|---|---|---|
| 1,000 transcripts (LLM-ready) | ~$5.00 | RAG pipelines, embeddings |
| 10,000 transcripts (LLM-ready) | ~$50.00 | Large-scale training data |
| 1,000 metadata extractions | ~$5.00 | Research, analytics |
| 100 video downloads | ~$10.00 | Content analysis, archiving |
| 500 video downloads | ~$50.00 | Large-scale archiving |
Actor fees are listed above. Apify platform costs (compute time, proxy data transfer) are billed separately. Each video starts on a cheap datacenter proxy and falls back to RESIDENTIAL (~$8/GB) only when that attempt fails, so most runs transfer their bytes at the datacenter rate. Apify Free plan includes $5/month in platform credits.
Cost optimization tips:
- Use
outputFormat: "llm_ready"for transcript-only workloads -- same price as metadata mode ($0.005/video) but with optimized flat output for AI pipelines. - Use
audio_onlyquality for faster, cheaper video downloads when you only need the audio track. - Use
downloadVideo: falsewhen you only need metadata, captions, and comments. - Lower
maxFileSizeMbto cap proxy data transfer -- withextractMetadataon (the default), oversized videos are billed at the metadata rate ($0.005) instead of $0.10 and markeddownloadSkippedReason: "size_limit_exceeded"; with it off they count as failed and are not billed. - Keep the RESIDENTIAL proxy on. The actor already tries the cheap datacenter proxy first on every video; RESIDENTIAL is the fallback that catches the videos YouTube blocks there. Remove it and those videos fail instead of costing a few cents.
AI/ML Use Cases
RAG (Retrieval-Augmented Generation)
Extract transcripts from YouTube videos and store as embeddings in a vector database. When users ask questions, retrieve relevant transcript chunks and feed them to your LLM for grounded, accurate answers. Use outputFormat: "llm_ready" for clean transcript text at $0.005/video.
Training Data Collection
Build fine-tuning datasets from educational YouTube channels. Extract transcripts from entire playlists or channels, pair with metadata (title, channel, tags), and export as JSONL for model training. Use searchQueries to find domain-specific content automatically.
Content Monitoring and Competitive Intelligence
Track competitor YouTube channels with scheduled runs. Extract metadata (views, likes, comments) to monitor content performance over time. Use metadata-only mode at $0.005/video to minimize costs. Set up webhookUrl for automated alerts.
Sentiment Analysis
Extract comments (with replies) from YouTube videos for sentiment analysis. Use metadata-only mode with extractComments: true and extractReplies: true. Feed comment text into your NLP pipeline for brand monitoring or market research.
Video Understanding (Multimodal AI)
Download video files alongside transcripts and metadata for multimodal AI research. Combine transcript text with video frames for tasks like video summarization, scene classification, or visual question answering.
Knowledge Base Construction
Build searchable knowledge bases from educational YouTube content. Extract transcripts with chapter markers to create structured, segmented documents. Chapters provide natural topic boundaries for better document chunking.
Example Inputs
LLM-Ready Transcripts (Cheapest)
{"startUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"outputFormat": "llm_ready","captionLanguage": "en","proxyConfiguration": {"useApifyProxy": true,"apifyProxyGroups": ["RESIDENTIAL"]}}
Batch Transcripts from Playlist
{"startUrls": ["https://www.youtube.com/playlist?list=PLrAXtmErZgOeiKm4sgNOknGvNjby9efdf"],"outputFormat": "llm_ready","maxVideos": 100,"maxConcurrency": 6,"captionLanguage": "en","proxyConfiguration": {"useApifyProxy": true,"apifyProxyGroups": ["RESIDENTIAL"]}}
Research Dataset (Metadata + Comments)
{"startUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"downloadVideo": false,"extractComments": true,"extractReplies": true,"extractCaptions": true,"extractChannelInfo": true,"maxComments": 200,"proxyConfiguration": {"useApifyProxy": true,"apifyProxyGroups": ["RESIDENTIAL"]}}
Search and Extract
{"searchQueries": ["machine learning tutorial", "transformer architecture explained"],"outputFormat": "llm_ready","maxVideos": 20,"proxyConfiguration": {"useApifyProxy": true,"apifyProxyGroups": ["RESIDENTIAL"]}}
Full Extraction (All Features)
{"startUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"quality": "1080p","extractMetadata": true,"extractCaptions": true,"captionLanguage": "en","downloadCaptions": true,"captionFormat": "srt","downloadThumbnail": true,"extractComments": true,"maxComments": 100,"extractReplies": true,"extractChannelInfo": true,"extractRelatedVideos": true,"filenameTemplate": "{channelName}_{title}_{quality}","webhookUrl": "https://your-server.com/webhook/youtube","proxyConfiguration": {"useApifyProxy": true,"apifyProxyGroups": ["RESIDENTIAL"]}}
Supported URL Formats
| URL Type | Example |
|---|---|
| Standard video | https://www.youtube.com/watch?v=VIDEO_ID |
| Short URL | https://youtu.be/VIDEO_ID |
| Shorts | https://www.youtube.com/shorts/VIDEO_ID |
| Embed | https://www.youtube.com/embed/VIDEO_ID |
| Playlist | https://www.youtube.com/playlist?list=PLAYLIST_ID |
| Channel (handle) | https://www.youtube.com/@ChannelHandle |
| Channel (ID) | https://www.youtube.com/channel/CHANNEL_ID |
| Channel (custom URL) | https://www.youtube.com/c/ChannelName |
Mix and match URL types in a single run -- pass playlists, channels, and individual videos together.
Integrations -- Zapier, Make, n8n
Zapier
- Add a Webhooks by Zapier action (POST request)
- URL:
https://api.apify.com/v2/acts/jy-labs~youtube-all-in-one-downloader-scraper/runs?token=YOUR_TOKEN - Payload: JSON with your input (
startUrls,outputFormat, etc.) - Use
webhookUrlto trigger the next Zapier step when the run completes
Make (formerly Integromat)
- Add an HTTP > Make a request module
- POST to:
https://api.apify.com/v2/acts/jy-labs~youtube-all-in-one-downloader-scraper/runs?token=YOUR_TOKEN - Body: JSON with actor input
- Use Make's Apify module to poll for completion and fetch dataset items
n8n
- Add an HTTP Request node (POST) with actor input
- Use
webhookUrlpointing to an n8n webhook trigger for completion notification - Add a second HTTP Request node to GET dataset items
Apify Scheduler
Built-in scheduling -- set a cron expression on the Schedules tab to run the actor automatically. Combine with webhookUrl to send results to Slack, email, or any endpoint.
FAQ
What is LLM-ready mode?
Set outputFormat: "llm_ready" and the actor automatically enables transcript extraction, disables video download, and charges $0.005/video (same as metadata extraction). The output includes clean plaintext transcripts ready to chunk, embed, or feed into any LLM. No extra configuration needed.
Do I need a proxy?
Apify Proxy is enabled by default with the RESIDENTIAL group -- no configuration needed. The actor runs a two-tier chain: the first attempt on each video goes through a cheap datacenter proxy, and only if that attempt fails does it retry on the RESIDENTIAL group you configured. That keeps most of the data transfer at the datacenter rate while RESIDENTIAL catches the videos YouTube blocks on datacenter IPs. Turning the proxy off removes the fallback and will make most runs fail, not merely risk throttling. Note: Proxy usage incurs data transfer costs (~$8/GB for the RESIDENTIAL fallback); use maxFileSizeMb or downloadVideo: false to control that spend rather than disabling the proxy.
What quality options are available?
highest (best available), 1080p, 720p, 480p, 360p, and audio_only. Automatic fallback to nearest lower resolution if requested quality is unavailable.
What happens when the quality I asked for is not available?
The actor inspects the formats the video actually offers and takes the highest one at or below your request -- ask for 1080p on a video published at 720p and you get 720p. If every available format sits above your request, it falls back to the lowest one that exists rather than failing the video. The quality field on the output row always reports what was really delivered, so downgrades are filterable after the run.
Can I get transcripts in other languages?
Yes. Set captionLanguage to any ISO language code (e.g., "ko", "ja", "es", "de"). The actor extracts captions in 100+ languages. If the preferred language is unavailable, it falls back to the first available.
Can I download age-restricted or private videos?
Neither can be downloaded as a video file, but the two land in the output differently.
Age-restricted and DRM-protected videos usually still return their metadata while every streaming client refuses the file. With extractMetadata: true (the default) you get a row with downloadUrl: null, quality: "none", and downloadSkippedReason: "download_failed", billed at the metadata extraction rate ($0.005) rather than the video download rate. Set extractMetadata: false if you would rather have them counted as failures and not billed at all.
Private, deleted, and unavailable videos cannot be read at all -- YouTube returns no metadata to deliver, so they never produce a metadata row. They come out as status: "failed" rows, visible only with includeFailedVideos: true, and are not billed.
What is the maximum batch size?
Up to 1,000 videos per run (via maxVideos, default 100). For larger batches, schedule multiple runs or chain them via the Apify API.
How long are files stored?
Downloaded files are stored in Apify Key-Value Store. Free plan: 7 days retention. Paid plans: longer retention. Export files before they expire.
Can I get audio only?
Yes. Set quality: "audio_only" to extract the audio track as an M4A file (AAC audio in MP4 container). Supported by all major players. Useful for podcast archiving, music extraction, or audio-based NLP.
Can I get SRT/VTT subtitle files?
Yes. Set downloadCaptions: true and captionFormat: "srt" or "vtt". The file is stored in Key-Value Store and the output includes captionFileUrl.
Is this suitable for production use?
Yes. The actor includes retry logic with exponential backoff, proxy rotation, concurrent processing, structured error handling, and webhook notifications. It is designed for automated pipelines and scheduled runs.
Can I search YouTube and extract at the same time?
Yes. Use searchQueries to search by keywords. The actor fetches search results and processes found videos alongside any startUrls. Useful for building training datasets on specific topics.
What format are downloaded videos in?
Videos: MP4 (.mp4). Audio-only: M4A (.m4a, AAC in MP4 container). Both are universally supported.
How many videos are processed at the same time?
maxConcurrency decides it: default 4, minimum 1, maximum 10. A value outside that range is rejected and replaced with 4, with a warning in the run log. Lower it for long or high-resolution downloads if a run runs out of memory. Raise it toward 10 for transcript and metadata work, where each video costs almost no memory.
Can I extract metadata without downloading the video?
Yes -- set downloadVideo: false. No file is fetched or stored, the run finishes far faster, and each video is billed as a metadata extraction rather than a video download (see Pricing Details). Metadata extraction is enabled automatically in this mode, and also whenever you turn on captions, comments, channel info, related videos, caption files, or thumbnails.
Why do some rows have downloadUrl: null?
Because the metadata was readable but the file was not deliverable -- age-restricted, DRM-protected, region-blocked, or larger than maxFileSizeMb. Those rows carry a downloadSkippedReason marker and are billed at the metadata extraction rate instead of the video download rate. The full behaviour is described under Metadata-Only Row (download requested but not delivered) and Can I download age-restricted or private videos? above.
Can I run this on a schedule and get only the new videos?
Yes. Set monitorMode: true together with a stateStoreName, then schedule the actor. Processed video ids are persisted in that named Key-Value Store and skipped on the next run, so a channel or playlist you watch daily only costs you the videos that were actually added. Give each monitoring job its own store name.
Limitations
- Age-restricted videos require YouTube authentication, so the video file cannot be downloaded -- with
extractMetadataon (the default) they are delivered as metadata-only rows markeddownloadSkippedReason: "download_failed"and billed at $0.005 whenever the metadata itself is still readable - Private and unlisted videos accessible only to the uploader cannot be processed
- Live streams currently broadcasting cannot be downloaded (completed live streams work)
- DRM-protected content (YouTube Premium originals) cannot be downloaded -- same metadata-only fallback and $0.005 billing as age-restricted videos
- File size limits per quality tier: 360p=100MB, 480p=150MB, 720p=250MB, 1080p=400MB, highest=500MB, audio=50MB -- oversized videos are delivered as metadata-only rows marked
downloadSkippedReason: "size_limit_exceeded" - Very long videos (>2 hours) may require more memory; lower
maxConcurrencyfor these - YouTube rate limiting may affect large batches without proxy -- always use proxy for production
- Geographically restricted videos may fail depending on proxy location; when metadata still comes back, the download is downgraded to a metadata-only row
Technology
- youtubei.js -- YouTube InnerTube API client (no browser required)
- Apify SDK -- Actor framework with dataset, key-value store, and proxy management
- p-limit -- Concurrent download management
- TypeScript -- Type-safe ESM implementation
Changelog
See the Changelog tab for version history and updates.
Support
If you encounter issues or have feature requests, open an issue on the Apify Store page. For custom integrations or enterprise use cases, reach out through the Apify platform.