YouTube Music Video Scraper avatar

YouTube Music Video Scraper

Pricing

from $2.99 / 1,000 music videos

Go to Apify Store
YouTube Music Video Scraper

YouTube Music Video Scraper

Search and browse YouTube Music videos with first-party YouTube Music endpoints, pagination, deduplication, and optional player metadata.

Pricing

from $2.99 / 1,000 music videos

Rating

0.0

(0)

Developer

w3crawler

w3crawler

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

17 hours ago

Last modified

Categories

Share

What this Actor does

YouTube Music Video Scraper collects deduplicated public music-video results from YouTube search pages and public YouTube Music or YouTube playlist, browse, and channel pages. It records the public title, artist/channel context, album text when exposed, duration, displayed counts, thumbnails, canonical links, source rank, and extraction provenance.

The Actor requests ordinary public rendered HTML and parses embedded page data. Optional player enrichment reads the public watch page for a description, dates, keywords, caption-language summaries, playability status, and audio/video format counts. It never stores signed caption or media URLs and does not download media.

Typical uses include music-video discovery, artist or album research, public playlist inventory, chart-like snapshots, deduplicated candidate lists for editorial workflows, and point-in-time metadata exports.

Supported targets and boundaries

  • Search target: public YouTube results for music-video phrases.
  • Start targets: public YouTube or YouTube Music playlist, browse, and channel URLs.
  • When both input arrays are supplied, startUrls are processed first. All sources share one run-wide maxItems cap.
  • The source order is the public page order. The Actor does not claim an official chart ranking or stable historical counts.
  • The implementation uses public HTML and embedded ytInitialData / ytInitialPlayerResponse only. It does not call private YouTube Music endpoints, alternate clients, hidden APIs, login-only pages, or media endpoints.

Input

Provide at least one non-empty searchQueries or startUrls entry. The runtime trims and deduplicates both arrays before processing.

FieldType and defaultDescription
searchQueriesstring array; no default; max 100Public keyword phrases. Each trimmed phrase is 1–200 characters. Search sources run after all startUrls.
startUrlsstring array; no default; max 100HTTPS YouTube/YouTube Music playlist, browse, or channel URLs. Each start URL uses one public page and runs before search sources.
maxItemsinteger; 20; 1–100Maximum unique video rows across every source combined. Deduplication is by videoId.
maxPagesinteger; 3; 1–10Maximum public search-page requests per search query. Start URLs always use one page.
maxRetriesinteger; 1; 0–3Retries the same public request with bounded backoff and the same transport context.
requestTimeoutSecsinteger; 60; 15–180Timeout for each public result-page or watch-page request.
requestDelayMsinteger; 250; 0–5000Delay between successive public search pages. It does not add pagination to start URLs or player requests.
countryCodestring; USTwo-letter country context, normalized to uppercase, such as US, GB, or IN.
languageCodestring; enLanguage context, normalized to lowercase, such as en, es, or pt-BR.
includePlayerDetailsboolean; trueFetch one public watch page per retained video and attach bounded metadata when available.
enableProxyFallbackboolean; trueAfter a direct request fails, allow one request through the configured Apify Proxy. This is fallback behavior, not access-control bypass.
includeDiagnosticsboolean; trueRetain page and optional-enrichment diagnostics. Each diagnostic has exactly url, error, errorCode, and scrapedAt.
proxyConfigurationobject; noneOptional account-authorized Apify Proxy configuration or credential-free HTTP/SOCKS URLs. Credentials and proxy URLs are never written to output.

The runtime rejects unknown input fields, invalid URLs, proxy credentials, unsupported proxy fields, unsafe proxy protocols, and values outside the ranges above.

Minimal search input

{
"searchQueries": ["Adele official music video"],
"maxItems": 10
}

Known public playlist input

{
"startUrls": [
"https://www.youtube.com/playlist?list=PL2JtvykrieUxaLXkeuXRHb-pJB9XpVWtd"
],
"maxItems": 10,
"includePlayerDetails": false
}

Combined sources with bounded options

{
"startUrls": [
"https://music.youtube.com/playlist?list=PL2JtvykrieUxaLXkeuXRHb-pJB9XpVWtd"
],
"searchQueries": ["Adele live music video"],
"maxItems": 20,
"maxPages": 2,
"maxRetries": 1,
"requestTimeoutSecs": 60,
"requestDelayMs": 250,
"countryCode": "US",
"languageCode": "en",
"includePlayerDetails": true,
"enableProxyFallback": true,
"includeDiagnostics": true,
"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyCountry": "US"
}
}

Run it in Apify Console

  1. Open the Actor and click Start.
  2. In Input, paste one of the JSON examples above or enter the fields in the form.
  3. Confirm the public source URLs, maxItems, and whether optional player enrichment is needed.
  4. Click Start and wait for the run to finish.
  5. Open the Dataset for result rows and the OUTPUT key in the run’s key-value store for counts and diagnostics; export the dataset as needed.

Output records

Rich normal record

This is a representative source-shaped record. Optional values are omitted when the public page does not expose them.

{
"recordType": "youtube_music_video",
"videoId": "dQw4w9WgXcQ",
"title": "Adele - Rolling in the Deep",
"description": "Official music video",
"descriptionLength": 20,
"accessibilityText": "Adele - Rolling in the Deep 4 minutes, 03 seconds",
"artists": [
{
"name": "Adele",
"id": "UCsRM0YB_dabtEPGPTKo-gcw",
"url": "https://www.youtube.com/channel/UCsRM0YB_dabtEPGPTKo-gcw"
}
],
"artist": "Adele",
"album": "21",
"duration": "4:03",
"durationSeconds": 243,
"views": 1200000000,
"viewsText": "1.2B views",
"releaseYear": 2011,
"publishedTimeText": "13 years ago",
"explicit": false,
"badges": ["Official"],
"thumbnails": [
{
"url": "https://i.ytimg.com/vi/dQw4w9WgXcQ/hq720.jpg",
"width": 720,
"height": 404
}
],
"thumbnailUrl": "https://i.ytimg.com/vi/dQw4w9WgXcQ/hq720.jpg",
"thumbnailCount": 1,
"channelName": "Adele",
"channelId": "UCsRM0YB_dabtEPGPTKo-gcw",
"channelUrl": "https://www.youtube.com/channel/UCsRM0YB_dabtEPGPTKo-gcw",
"youtubeMusicUrl": "https://music.youtube.com/watch?v=dQw4w9WgXcQ",
"youtubeUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"source": "search",
"searchQuery": "Adele official music video",
"sourceUrl": "https://www.youtube.com/results?search_query=Adele+official+music+video&hl=en&gl=US&sp=EgIQAQ%3D%3D",
"rank": 1,
"sourceRank": 1,
"searchRank": 1,
"page": 1,
"resultSource": "ytInitialData.videoRenderer",
"sourceTransport": "public-rendered-html-direct",
"extractionMethod": "youtube-public-page-embedded-initial-data",
"metadataSource": "youtube-public-search-page",
"metadataEnriched": true,
"isShort": false,
"isLive": false,
"isUpcoming": false,
"continuationDetected": true,
"continuationCount": 1,
"initialItemsExposed": 20,
"requestedMaxItems": 10,
"languageCode": "en",
"countryCode": "US",
"playerDetailsRequested": true,
"playerDetailsAvailable": true,
"playerDetailsStatus": "available",
"playerTransport": "youtube-public-watch-page",
"player": {
"title": "Adele - Rolling in the Deep",
"description": "Official music video",
"descriptionLength": 20,
"durationSeconds": 243,
"viewCount": 1200000000,
"channelName": "Adele",
"channelId": "UCsRM0YB_dabtEPGPTKo-gcw",
"channelUrl": "https://www.youtube.com/channel/UCsRM0YB_dabtEPGPTKo-gcw",
"isLive": false,
"keywords": ["Adele", "music"],
"publishDate": "2011-01-24",
"uploadDate": "2011-01-24",
"category": "Music",
"tags": ["Adele", "Rolling in the Deep"],
"captions": [
{
"languageCode": "en",
"name": "English",
"kind": "asr",
"isAutoGenerated": true,
"isTranslatable": true
}
],
"captionCount": 1,
"playabilityStatus": "OK",
"hasVideo": true,
"hasAudio": true,
"formatCount": 34,
"videoFormatCount": 20,
"audioFormatCount": 14,
"metadataSource": "youtube-public-watch-page",
"extractionMethod": "youtube-public-watch-page-embedded-player-response"
},
"scrapedAt": "2026-09-08T12:00:00.000Z"
}

Fallback normal record and diagnostic

If the source result is usable but its optional watch page is unavailable, the normal row is retained with playerDetailsStatus: "unavailable"; the separate diagnostic explains the omission.

{
"recordType": "youtube_music_video",
"videoId": "9bZkp7q19f0",
"title": "Public music-video result",
"youtubeMusicUrl": "https://music.youtube.com/watch?v=9bZkp7q19f0",
"youtubeUrl": "https://www.youtube.com/watch?v=9bZkp7q19f0",
"source": "search",
"sourceUrl": "https://www.youtube.com/results?search_query=music&hl=en&gl=US&sp=EgIQAQ%3D%3D",
"rank": 1,
"sourceRank": 1,
"searchRank": 1,
"page": 1,
"resultSource": "ytInitialData.videoRenderer",
"sourceTransport": "public-rendered-html-via-apify-proxy",
"extractionMethod": "youtube-public-page-embedded-initial-data",
"metadataSource": "youtube-public-search-page",
"metadataEnriched": true,
"continuationDetected": false,
"continuationCount": 0,
"initialItemsExposed": 1,
"requestedMaxItems": 1,
"languageCode": "en",
"countryCode": "US",
"playerDetailsRequested": true,
"playerDetailsAvailable": false,
"playerDetailsStatus": "unavailable",
"scrapedAt": "2026-09-08T12:00:00.000Z"
}
{
"url": "https://www.youtube.com/watch?v=9bZkp7q19f0",
"error": "Public YouTube watch page did not expose matching player metadata",
"errorCode": "PLAYER_METADATA_UNAVAILABLE",
"scrapedAt": "2026-09-08T12:00:00.001Z"
}

OUTPUT run summary

The Actor writes this JSON object under the OUTPUT key in the default key-value store.

{
"runType": "youtube-public-music-video-run",
"status": "success",
"dataAvailable": true,
"sourceCount": 2,
"sourcesWithItems": 2,
"maxItems": 20,
"maxPages": 3,
"pagesFetched": 2,
"pageAttempts": 4,
"itemCount": 2,
"diagnosticCount": 0,
"failedCount": 0,
"blockedCount": 0,
"playerDetailsRequested": 2,
"playerDetailsAvailableCount": 2,
"playerDetailsUnavailableCount": 0,
"continuationDetected": true,
"continuationCount": 1,
"proxyRequested": false,
"proxyConfigured": false,
"includeDiagnostics": true,
"languageCode": "en",
"countryCode": "US",
"maxRetries": 1,
"requestDelayMs": 250,
"requestTimeoutSecs": 60,
"paginationMethod": "public-page-query",
"extractionMethod": "youtube-public-page-embedded-initial-data",
"durationMs": 4661,
"finishedAt": "2026-09-08T12:00:04.661Z"
}

Field reference and semantics

GroupFieldsSemantics
IdentityrecordType, videoId, title, description, descriptionLength, accessibilityTextIdentity and text are copied or parsed from public result/watch metadata. Missing values are omitted.
Music contextartists[], artist, album, duration, durationSecondsArtist/channel links and album text are present only when exposed by the renderer. durationSeconds is parsed from a public duration string.
Public statsviews, viewsText, releaseYear, publishedTimeText, explicit, badgesviews is a parsed count; viewsText preserves the displayed text. These are point-in-time observations.
Channel and imageschannelName, channelId, channelUrl, thumbnails[], thumbnailUrl, thumbnailCountURLs are restricted to public YouTube/Google image hosts and are not signed media URLs.
Canonical linksyoutubeUrl, youtubeMusicUrlBoth are derived from the validated 11-character videoId.
Source and ranksource, searchQuery, sourceUrl, rank, sourceRank, searchRank, page, resultSourcerank is run order after deduplication. sourceRank is order among unique rows from the source; searchRank is present for search sources.
ProvenancesourceTransport, extractionMethod, metadataSource, metadataEnriched, scrapedAtIdentifies the public page family, transport, parser, and emission time.
Flags and limitsisShort, isLive, isUpcoming, continuationDetected, continuationCount, initialItemsExposed, requestedMaxItems, languageCode, countryCodeRecords flags and the bounded request context visible for that row. Continuation markers are observed, not followed.
Player statusplayerDetailsRequested, playerDetailsAvailable, playerDetailsStatus, playerTransport, playerOptional public watch-page enrichment. player contains metadata, caption-language summaries, and format counts only.
Diagnosticsurl, error, errorCode, scrapedAtDiagnostic rows have exactly these four fields and never masquerade as normal results.

Nested artists entries contain name, id, and url. Nested thumbnails entries contain url, width, and height. Nested player fields include title/description, dates, keywords/tags, captions language summaries, playability, and format counts. Caption and media URLs are deliberately omitted.

Pagination, deduplication, and omissions

  • Search pagination uses the ordinary public page query parameter up to maxPages for each query. Start URLs use one public page.
  • The Actor reports embedded continuation markers in continuationDetected and continuationCount but does not replay private continuation requests.
  • Rows are deduplicated by videoId; the first occurrence in start-URL-then-search source order is retained. maxItems is a run-wide cap.
  • There are no user-configurable filters or sorts beyond the built-in music-video search filter and the source page’s own order. The Actor does not fabricate ranking, artist, album, price, date, or availability values.
  • Player enrichment is one bounded public watch-page request per retained video when includePlayerDetails is true. It does not expand maxPages.
  • Missing public fields are omitted. An empty result page, blocked page, or failed public request increments the relevant summary count and emits a diagnostic when enabled. A usable result with unavailable player enrichment remains a normal row with the player fields omitted.
  • Output never includes cookies, authorization headers, proxy credentials/URLs, request identities, fingerprint data, private client parameters, signed caption URLs, streaming URLs, playback URLs, or downloaded media.

Proxy behavior and cost

Direct public HTML is attempted first unless an explicit proxy configuration is supplied. With enableProxyFallback, a failed direct request may be retried once through the configured account-authorized Apify Proxy. A proxy is not a guarantee of access and is not used to bypass CAPTCHAs, authentication, or access controls. sourceTransport and OUTPUT.proxyConfigured state what happened without exposing proxy details.

Actor cost is driven by Apify compute time and memory, plus any proxy traffic used by your account; the Actor does not add a fixed per-result charge. Actual cost depends on the selected plan, run duration, page count, player enrichment, and proxy use. Check the current Apify pricing and the run’s resource usage for the authoritative amount.

Troubleshooting

The run returns diagnostics and no rows

Confirm that at least one URL is a public HTTPS YouTube/YouTube Music playlist, browse, or channel page, or try a broader public search phrase. Review OUTPUT.status, failedCount, blockedCount, and the diagnostic errorCode. A public page can change shape or be temporarily unavailable.

Search rows are present but player metadata is missing

Set includePlayerDetails to false when only search-page metadata is needed, or leave it enabled and treat playerDetailsStatus: "unavailable" as an honest fallback. The Actor does not substitute private endpoints or signed URLs.

Results are fewer than maxItems

maxItems is a maximum, not a promise. The public page may expose fewer music-video candidates, duplicates may be removed, a source may be blocked, or continuation markers may be visible without being followed. Use more distinct queries or public start URLs within the input limits.

Proxy fallback is not used

Fallback requires enableProxyFallback: true and a proxy configuration that can resolve in the current Apify account. Explicit proxy mode uses that configuration for the request. Credential-bearing proxy URLs are rejected.

API and support

The Actor’s Console API tab exposes the current run, dataset, and key-value store links. For programmatic access, use the Apify API documentation and read the default dataset plus the OUTPUT key from the run’s default key-value store. Use the Actor’s Console Issues or support channel to report a reproducible public URL, input, run ID, and diagnostic code; do not include secrets or proxy credentials.

Use this Actor only for public pages and in accordance with YouTube’s Terms of Service, applicable law, and any site-specific rules. You are responsible for choosing lawful targets, respecting rights and rate limits, and assessing whether an export contains personal information. The Actor does not log in, access private content, bypass CAPTCHAs, or evade access controls. It stores public metadata and bounded diagnostics in the run’s Apify dataset and key-value store; it does not intentionally store cookies, credentials, signed media URLs, or downloaded media.

This Actor is an independent community tool and is not affiliated with, endorsed by, or sponsored by YouTube, Google, or YouTube Music. Names and trademarks belong to their respective owners.

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.