YouTube Music Video Scraper
Pricing
from $2.99 / 1,000 music videos
YouTube Music Video Scraper
Search and browse YouTube Music videos with first-party YouTube Music endpoints, pagination, deduplication, and optional player metadata.
Pricing
from $2.99 / 1,000 music videos
Rating
0.0
(0)
Developer
w3crawler
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
17 hours ago
Last modified
Categories
Share
What this Actor does
YouTube Music Video Scraper collects deduplicated public music-video results from YouTube search pages and public YouTube Music or YouTube playlist, browse, and channel pages. It records the public title, artist/channel context, album text when exposed, duration, displayed counts, thumbnails, canonical links, source rank, and extraction provenance.
The Actor requests ordinary public rendered HTML and parses embedded page data. Optional player enrichment reads the public watch page for a description, dates, keywords, caption-language summaries, playability status, and audio/video format counts. It never stores signed caption or media URLs and does not download media.
Typical uses include music-video discovery, artist or album research, public playlist inventory, chart-like snapshots, deduplicated candidate lists for editorial workflows, and point-in-time metadata exports.
Supported targets and boundaries
- Search target: public YouTube results for music-video phrases.
- Start targets: public YouTube or YouTube Music playlist, browse, and channel URLs.
- When both input arrays are supplied,
startUrlsare processed first. All sources share one run-widemaxItemscap. - The source order is the public page order. The Actor does not claim an official chart ranking or stable historical counts.
- The implementation uses public HTML and embedded
ytInitialData/ytInitialPlayerResponseonly. It does not call private YouTube Music endpoints, alternate clients, hidden APIs, login-only pages, or media endpoints.
Input
Provide at least one non-empty searchQueries or startUrls entry. The runtime trims and deduplicates both arrays before processing.
| Field | Type and default | Description |
|---|---|---|
searchQueries | string array; no default; max 100 | Public keyword phrases. Each trimmed phrase is 1–200 characters. Search sources run after all startUrls. |
startUrls | string array; no default; max 100 | HTTPS YouTube/YouTube Music playlist, browse, or channel URLs. Each start URL uses one public page and runs before search sources. |
maxItems | integer; 20; 1–100 | Maximum unique video rows across every source combined. Deduplication is by videoId. |
maxPages | integer; 3; 1–10 | Maximum public search-page requests per search query. Start URLs always use one page. |
maxRetries | integer; 1; 0–3 | Retries the same public request with bounded backoff and the same transport context. |
requestTimeoutSecs | integer; 60; 15–180 | Timeout for each public result-page or watch-page request. |
requestDelayMs | integer; 250; 0–5000 | Delay between successive public search pages. It does not add pagination to start URLs or player requests. |
countryCode | string; US | Two-letter country context, normalized to uppercase, such as US, GB, or IN. |
languageCode | string; en | Language context, normalized to lowercase, such as en, es, or pt-BR. |
includePlayerDetails | boolean; true | Fetch one public watch page per retained video and attach bounded metadata when available. |
enableProxyFallback | boolean; true | After a direct request fails, allow one request through the configured Apify Proxy. This is fallback behavior, not access-control bypass. |
includeDiagnostics | boolean; true | Retain page and optional-enrichment diagnostics. Each diagnostic has exactly url, error, errorCode, and scrapedAt. |
proxyConfiguration | object; none | Optional account-authorized Apify Proxy configuration or credential-free HTTP/SOCKS URLs. Credentials and proxy URLs are never written to output. |
The runtime rejects unknown input fields, invalid URLs, proxy credentials, unsupported proxy fields, unsafe proxy protocols, and values outside the ranges above.
Minimal search input
{"searchQueries": ["Adele official music video"],"maxItems": 10}
Known public playlist input
{"startUrls": ["https://www.youtube.com/playlist?list=PL2JtvykrieUxaLXkeuXRHb-pJB9XpVWtd"],"maxItems": 10,"includePlayerDetails": false}
Combined sources with bounded options
{"startUrls": ["https://music.youtube.com/playlist?list=PL2JtvykrieUxaLXkeuXRHb-pJB9XpVWtd"],"searchQueries": ["Adele live music video"],"maxItems": 20,"maxPages": 2,"maxRetries": 1,"requestTimeoutSecs": 60,"requestDelayMs": 250,"countryCode": "US","languageCode": "en","includePlayerDetails": true,"enableProxyFallback": true,"includeDiagnostics": true,"proxyConfiguration": {"useApifyProxy": true,"apifyProxyCountry": "US"}}
Run it in Apify Console
- Open the Actor and click Start.
- In Input, paste one of the JSON examples above or enter the fields in the form.
- Confirm the public source URLs,
maxItems, and whether optional player enrichment is needed. - Click Start and wait for the run to finish.
- Open the Dataset for result rows and the
OUTPUTkey in the run’s key-value store for counts and diagnostics; export the dataset as needed.
Output records
Rich normal record
This is a representative source-shaped record. Optional values are omitted when the public page does not expose them.
{"recordType": "youtube_music_video","videoId": "dQw4w9WgXcQ","title": "Adele - Rolling in the Deep","description": "Official music video","descriptionLength": 20,"accessibilityText": "Adele - Rolling in the Deep 4 minutes, 03 seconds","artists": [{"name": "Adele","id": "UCsRM0YB_dabtEPGPTKo-gcw","url": "https://www.youtube.com/channel/UCsRM0YB_dabtEPGPTKo-gcw"}],"artist": "Adele","album": "21","duration": "4:03","durationSeconds": 243,"views": 1200000000,"viewsText": "1.2B views","releaseYear": 2011,"publishedTimeText": "13 years ago","explicit": false,"badges": ["Official"],"thumbnails": [{"url": "https://i.ytimg.com/vi/dQw4w9WgXcQ/hq720.jpg","width": 720,"height": 404}],"thumbnailUrl": "https://i.ytimg.com/vi/dQw4w9WgXcQ/hq720.jpg","thumbnailCount": 1,"channelName": "Adele","channelId": "UCsRM0YB_dabtEPGPTKo-gcw","channelUrl": "https://www.youtube.com/channel/UCsRM0YB_dabtEPGPTKo-gcw","youtubeMusicUrl": "https://music.youtube.com/watch?v=dQw4w9WgXcQ","youtubeUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","source": "search","searchQuery": "Adele official music video","sourceUrl": "https://www.youtube.com/results?search_query=Adele+official+music+video&hl=en&gl=US&sp=EgIQAQ%3D%3D","rank": 1,"sourceRank": 1,"searchRank": 1,"page": 1,"resultSource": "ytInitialData.videoRenderer","sourceTransport": "public-rendered-html-direct","extractionMethod": "youtube-public-page-embedded-initial-data","metadataSource": "youtube-public-search-page","metadataEnriched": true,"isShort": false,"isLive": false,"isUpcoming": false,"continuationDetected": true,"continuationCount": 1,"initialItemsExposed": 20,"requestedMaxItems": 10,"languageCode": "en","countryCode": "US","playerDetailsRequested": true,"playerDetailsAvailable": true,"playerDetailsStatus": "available","playerTransport": "youtube-public-watch-page","player": {"title": "Adele - Rolling in the Deep","description": "Official music video","descriptionLength": 20,"durationSeconds": 243,"viewCount": 1200000000,"channelName": "Adele","channelId": "UCsRM0YB_dabtEPGPTKo-gcw","channelUrl": "https://www.youtube.com/channel/UCsRM0YB_dabtEPGPTKo-gcw","isLive": false,"keywords": ["Adele", "music"],"publishDate": "2011-01-24","uploadDate": "2011-01-24","category": "Music","tags": ["Adele", "Rolling in the Deep"],"captions": [{"languageCode": "en","name": "English","kind": "asr","isAutoGenerated": true,"isTranslatable": true}],"captionCount": 1,"playabilityStatus": "OK","hasVideo": true,"hasAudio": true,"formatCount": 34,"videoFormatCount": 20,"audioFormatCount": 14,"metadataSource": "youtube-public-watch-page","extractionMethod": "youtube-public-watch-page-embedded-player-response"},"scrapedAt": "2026-09-08T12:00:00.000Z"}
Fallback normal record and diagnostic
If the source result is usable but its optional watch page is unavailable, the normal row is retained with playerDetailsStatus: "unavailable"; the separate diagnostic explains the omission.
{"recordType": "youtube_music_video","videoId": "9bZkp7q19f0","title": "Public music-video result","youtubeMusicUrl": "https://music.youtube.com/watch?v=9bZkp7q19f0","youtubeUrl": "https://www.youtube.com/watch?v=9bZkp7q19f0","source": "search","sourceUrl": "https://www.youtube.com/results?search_query=music&hl=en&gl=US&sp=EgIQAQ%3D%3D","rank": 1,"sourceRank": 1,"searchRank": 1,"page": 1,"resultSource": "ytInitialData.videoRenderer","sourceTransport": "public-rendered-html-via-apify-proxy","extractionMethod": "youtube-public-page-embedded-initial-data","metadataSource": "youtube-public-search-page","metadataEnriched": true,"continuationDetected": false,"continuationCount": 0,"initialItemsExposed": 1,"requestedMaxItems": 1,"languageCode": "en","countryCode": "US","playerDetailsRequested": true,"playerDetailsAvailable": false,"playerDetailsStatus": "unavailable","scrapedAt": "2026-09-08T12:00:00.000Z"}
{"url": "https://www.youtube.com/watch?v=9bZkp7q19f0","error": "Public YouTube watch page did not expose matching player metadata","errorCode": "PLAYER_METADATA_UNAVAILABLE","scrapedAt": "2026-09-08T12:00:00.001Z"}
OUTPUT run summary
The Actor writes this JSON object under the OUTPUT key in the default key-value store.
{"runType": "youtube-public-music-video-run","status": "success","dataAvailable": true,"sourceCount": 2,"sourcesWithItems": 2,"maxItems": 20,"maxPages": 3,"pagesFetched": 2,"pageAttempts": 4,"itemCount": 2,"diagnosticCount": 0,"failedCount": 0,"blockedCount": 0,"playerDetailsRequested": 2,"playerDetailsAvailableCount": 2,"playerDetailsUnavailableCount": 0,"continuationDetected": true,"continuationCount": 1,"proxyRequested": false,"proxyConfigured": false,"includeDiagnostics": true,"languageCode": "en","countryCode": "US","maxRetries": 1,"requestDelayMs": 250,"requestTimeoutSecs": 60,"paginationMethod": "public-page-query","extractionMethod": "youtube-public-page-embedded-initial-data","durationMs": 4661,"finishedAt": "2026-09-08T12:00:04.661Z"}
Field reference and semantics
| Group | Fields | Semantics |
|---|---|---|
| Identity | recordType, videoId, title, description, descriptionLength, accessibilityText | Identity and text are copied or parsed from public result/watch metadata. Missing values are omitted. |
| Music context | artists[], artist, album, duration, durationSeconds | Artist/channel links and album text are present only when exposed by the renderer. durationSeconds is parsed from a public duration string. |
| Public stats | views, viewsText, releaseYear, publishedTimeText, explicit, badges | views is a parsed count; viewsText preserves the displayed text. These are point-in-time observations. |
| Channel and images | channelName, channelId, channelUrl, thumbnails[], thumbnailUrl, thumbnailCount | URLs are restricted to public YouTube/Google image hosts and are not signed media URLs. |
| Canonical links | youtubeUrl, youtubeMusicUrl | Both are derived from the validated 11-character videoId. |
| Source and rank | source, searchQuery, sourceUrl, rank, sourceRank, searchRank, page, resultSource | rank is run order after deduplication. sourceRank is order among unique rows from the source; searchRank is present for search sources. |
| Provenance | sourceTransport, extractionMethod, metadataSource, metadataEnriched, scrapedAt | Identifies the public page family, transport, parser, and emission time. |
| Flags and limits | isShort, isLive, isUpcoming, continuationDetected, continuationCount, initialItemsExposed, requestedMaxItems, languageCode, countryCode | Records flags and the bounded request context visible for that row. Continuation markers are observed, not followed. |
| Player status | playerDetailsRequested, playerDetailsAvailable, playerDetailsStatus, playerTransport, player | Optional public watch-page enrichment. player contains metadata, caption-language summaries, and format counts only. |
| Diagnostics | url, error, errorCode, scrapedAt | Diagnostic rows have exactly these four fields and never masquerade as normal results. |
Nested artists entries contain name, id, and url. Nested thumbnails entries contain url, width, and height. Nested player fields include title/description, dates, keywords/tags, captions language summaries, playability, and format counts. Caption and media URLs are deliberately omitted.
Pagination, deduplication, and omissions
- Search pagination uses the ordinary public
pagequery parameter up tomaxPagesfor each query. Start URLs use one public page. - The Actor reports embedded continuation markers in
continuationDetectedandcontinuationCountbut does not replay private continuation requests. - Rows are deduplicated by
videoId; the first occurrence in start-URL-then-search source order is retained.maxItemsis a run-wide cap. - There are no user-configurable filters or sorts beyond the built-in music-video search filter and the source page’s own order. The Actor does not fabricate ranking, artist, album, price, date, or availability values.
- Player enrichment is one bounded public watch-page request per retained video when
includePlayerDetailsis true. It does not expandmaxPages. - Missing public fields are omitted. An empty result page, blocked page, or failed public request increments the relevant summary count and emits a diagnostic when enabled. A usable result with unavailable player enrichment remains a normal row with the player fields omitted.
- Output never includes cookies, authorization headers, proxy credentials/URLs, request identities, fingerprint data, private client parameters, signed caption URLs, streaming URLs, playback URLs, or downloaded media.
Proxy behavior and cost
Direct public HTML is attempted first unless an explicit proxy configuration is supplied. With enableProxyFallback, a failed direct request may be retried once through the configured account-authorized Apify Proxy. A proxy is not a guarantee of access and is not used to bypass CAPTCHAs, authentication, or access controls. sourceTransport and OUTPUT.proxyConfigured state what happened without exposing proxy details.
Actor cost is driven by Apify compute time and memory, plus any proxy traffic used by your account; the Actor does not add a fixed per-result charge. Actual cost depends on the selected plan, run duration, page count, player enrichment, and proxy use. Check the current Apify pricing and the run’s resource usage for the authoritative amount.
Troubleshooting
The run returns diagnostics and no rows
Confirm that at least one URL is a public HTTPS YouTube/YouTube Music playlist, browse, or channel page, or try a broader public search phrase. Review OUTPUT.status, failedCount, blockedCount, and the diagnostic errorCode. A public page can change shape or be temporarily unavailable.
Search rows are present but player metadata is missing
Set includePlayerDetails to false when only search-page metadata is needed, or leave it enabled and treat playerDetailsStatus: "unavailable" as an honest fallback. The Actor does not substitute private endpoints or signed URLs.
Results are fewer than maxItems
maxItems is a maximum, not a promise. The public page may expose fewer music-video candidates, duplicates may be removed, a source may be blocked, or continuation markers may be visible without being followed. Use more distinct queries or public start URLs within the input limits.
Proxy fallback is not used
Fallback requires enableProxyFallback: true and a proxy configuration that can resolve in the current Apify account. Explicit proxy mode uses that configuration for the request. Credential-bearing proxy URLs are rejected.
API and support
The Actor’s Console API tab exposes the current run, dataset, and key-value store links. For programmatic access, use the Apify API documentation and read the default dataset plus the OUTPUT key from the run’s default key-value store. Use the Actor’s Console Issues or support channel to report a reproducible public URL, input, run ID, and diagnostic code; do not include secrets or proxy credentials.
Privacy, legal, and non-affiliation
Use this Actor only for public pages and in accordance with YouTube’s Terms of Service, applicable law, and any site-specific rules. You are responsible for choosing lawful targets, respecting rights and rate limits, and assessing whether an export contains personal information. The Actor does not log in, access private content, bypass CAPTCHAs, or evade access controls. It stores public metadata and bounded diagnostics in the run’s Apify dataset and key-value store; it does not intentionally store cookies, credentials, signed media URLs, or downloaded media.
This Actor is an independent community tool and is not affiliated with, endorsed by, or sponsored by YouTube, Google, or YouTube Music. Names and trademarks belong to their respective owners.
You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.