YouTube Music Podcast Scraper avatar

YouTube Music Podcast Scraper

Pricing

from $2.99 / 1,000 podcasts

Go to Apify Store
YouTube Music Podcast Scraper

YouTube Music Podcast Scraper

Search current YouTube Music podcast shows and episodes through first-party youtubei endpoints, continuation pagination, enrichment, and deduplication.

Pricing

from $2.99 / 1,000 podcasts

Rating

0.0

(0)

Developer

w3crawler

w3crawler

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

18 hours ago

Last modified

Categories

Share

Search public YouTube pages for podcast episodes, shows/playlists, and creator channels. The Actor parses the public HTML delivered to an unauthenticated visitor and can optionally enrich episodes from public watch pages, shows from public playlist pages, and channels from public channel pages.

What this Actor does

Use this Actor to:

  • Discover public podcast episodes, show playlists, and creator channels from search terms.
  • Inspect a known public watch, playlist, channel, or handle URL.
  • Keep source-backed titles, descriptions, durations, publication text, views, identities, artwork, and public links together.
  • Add optional public watch-page episode metadata or bounded playlist/channel details.
  • Measure public access boundaries with small, explicit diagnostic rows.

The Actor emits a creator channel as type: "channel"; it does not relabel a channel as a podcast show. podcastMatch is a transparent title/description/query signal, not a hidden ranking or guarantee that the source describes a podcast.

Input

Provide at least one searchQueries entry or one public startUrls entry. Both may be supplied. Direct startUrls are processed first because they are deterministic detail targets; search queries run afterward until the global maxItems cap is reached. Search terms are normalized and deduplicated case-insensitively. Start URLs are canonicalized and duplicate canonical URLs are rejected.

Input fields

FieldType and limitsBehavior
searchQueriesstring array, up to 25 items; each 1–200 charactersOptional public podcast, show, or episode searches.
startUrlsstring array, up to 50 itemsOptional public YouTube watch, playlist, channel, @handle, /c/, or /user/ URLs. Credentials and non-YouTube hosts are rejected.
maxItemsinteger, 1–100; default 50Global cap on deduplicated normal episode/show/channel records.
maxPagesinteger, 1–5; default 1Maximum public search pages per query. Page URLs are bounded; continuation markers are reported but continuation tokens are not replayed.
maxShowEpisodesinteger, 1–50; default 10Maximum visible episodes retained inside each show detail object.
maxRetriesinteger, 0–3; default 1Bounded retries after a public page request fails.
requestTimeoutSecsinteger, 15–180; default 60Timeout for one public HTML request.
requestDelayMsinteger, 0–5000; default 250Delay between page requests.
countryCodetwo letters; default USCountry context for public pages.
languageCode2–3 letters with optional locale suffix; default enLanguage context for public pages.
includeShowDetailsboolean; default trueEnrich show results from public playlist pages and channel URL results when available.
includeEpisodeDetailsboolean; default trueAttempt optional public watch-page metadata for episode results.
includeDiagnosticsboolean; default trueEmit exact four-field diagnostics for access or optional-enrichment failures.
enableProxyFallbackboolean; default trueTry one configured Apify Proxy request after a direct public-page failure.
proxyConfigurationApify Proxy object; optionalAccount-authorized proxy configuration. Proxy URLs must be credential-free.

Runnable input examples

Public podcast search:

{
"searchQueries": ["technology podcast"],
"maxItems": 5,
"maxPages": 1,
"includeShowDetails": true,
"includeEpisodeDetails": true,
"includeDiagnostics": true,
"enableProxyFallback": true,
"proxyConfiguration": {
"useApifyProxy": false
}
}

Known public episode:

{
"startUrls": ["https://www.youtube.com/watch?v=7ARBJQn6QkM"],
"maxItems": 1,
"includeEpisodeDetails": true,
"includeShowDetails": false,
"enableProxyFallback": false
}

Known playlist plus search, with bounded enrichment:

{
"searchQueries": ["technology podcast", "AI interview podcast"],
"startUrls": [
"https://www.youtube.com/playlist?list=PLADd6sStSis77HKfbf4KCY6SvthfxeUgn"
],
"maxItems": 8,
"maxPages": 2,
"maxShowEpisodes": 5,
"maxRetries": 1,
"requestTimeoutSecs": 45,
"requestDelayMs": 250,
"countryCode": "US",
"languageCode": "en",
"includeShowDetails": true,
"includeEpisodeDetails": false,
"includeDiagnostics": true,
"enableProxyFallback": true,
"proxyConfiguration": {
"useApifyProxy": false
}
}

Run it in the Apify Console

  1. Open the Actor and select the Input tab.
  2. Paste one of the JSON objects above, or enter the fields in the form.
  3. If a proxy route is needed, configure the account-authorized Apify Proxy options and keep proxy URLs credential-free.
  4. Click Start and open the Dataset tab after the run finishes.
  5. Open the OUTPUT key-value record to review counts, enrichment availability, access state, limits, and duration.

Output

Normal episode record

Normal rows have recordType: "youtube_public_podcast_result" and type: "episode", "show", or "channel". This rich example is source-shaped; values not exposed by the public page are omitted.

{
"recordType": "youtube_public_podcast_result",
"type": "episode",
"searchQuery": "technology podcast",
"title": "The Big Technology Podcast",
"description": "A public conversation about technology and business.",
"descriptionLength": 52,
"subtitle": "Big Technology · 2 years ago · 120K views",
"showName": "The Big Technology Podcast",
"channelName": "Big Technology",
"channelId": "UCye1YedIypHffYb8k6Gp9wg",
"channelUrl": "https://www.youtube.com/channel/UCye1YedIypHffYb8k6Gp9wg",
"duration": "1:02:15",
"durationSeconds": 3735,
"publishedText": "2 years ago",
"views": 120000,
"viewsText": "120K views",
"videoId": "7ARBJQn6QkM",
"videoUrl": "https://www.youtube.com/watch?v=7ARBJQn6QkM",
"youtubeMusicUrl": "https://music.youtube.com/watch?v=7ARBJQn6QkM",
"thumbnails": [
{
"url": "https://i.ytimg.com/vi/7ARBJQn6QkM/hqdefault.jpg",
"width": 480,
"height": 360
}
],
"thumbnailUrl": "https://i.ytimg.com/vi/7ARBJQn6QkM/hqdefault.jpg",
"thumbnailCount": 1,
"badges": [],
"isLive": false,
"isUpcoming": false,
"podcastMatch": true,
"sourceRenderer": "public-video-renderer",
"episodeDetailsAvailable": true,
"episodeDetailsStatus": "available",
"episodeTransport": "youtube-public-watch-page",
"episode": {
"title": "The Big Technology Podcast",
"description": "A public conversation about technology and business.",
"descriptionLength": 52,
"channelName": "Big Technology",
"channelId": "UCye1YedIypHffYb8k6Gp9wg",
"channelUrl": "https://www.youtube.com/channel/UCye1YedIypHffYb8k6Gp9wg",
"lengthSeconds": 3735,
"publishDate": "2024-01-10",
"uploadDate": "2024-01-10",
"viewCount": 120000,
"category": "Science & Technology",
"keywords": ["technology", "podcast"],
"isLive": false,
"playerStatus": "OK",
"captionCount": 2,
"captions": [
{
"languageCode": "en",
"name": "English",
"kind": "asr"
}
],
"hasVideo": true,
"hasAudio": true,
"formatCount": 8,
"metadataSource": "youtube-public-watch-page",
"extractionMethod": "youtube-public-watch-page-embedded-player-data"
},
"searchUrl": "https://www.youtube.com/results?search_query=technology+podcast&hl=en&gl=US",
"source": "search",
"sourceUrl": "https://www.youtube.com/results?search_query=technology+podcast&hl=en&gl=US",
"rank": 1,
"page": 1,
"initialItemsExposed": 6,
"continuationDetected": false,
"continuationCount": 0,
"sourceTransport": "public-rendered-html-direct",
"extractionMethod": "youtube-public-page-embedded-initial-data",
"metadataSource": "youtube-public-watch-page",
"metadataEnriched": true,
"requestedMaxItems": 5,
"requestedMaxPages": 1,
"requestedMaxShowEpisodes": 10,
"languageCode": "en",
"countryCode": "US",
"scrapedAt": "2026-09-08T12:00:00.000Z"
}

Show and channel records

type: "show" may include a show object with public episode/view counts, creator identity, artwork, and at most maxShowEpisodes visible episodes. type: "channel" may include a channel object with public creator, subscriber/video text, description, artwork, and the canonical public channel URL. Detail objects are optional and are not synthesized when the page does not expose them.

Fallback and diagnostic record

With enableProxyFallback: true, a configured Apify Proxy route is tried after direct public-page failure. Optional enrichment failure does not discard a valid core result; its availability status becomes unavailable and a diagnostic is emitted when diagnostics are enabled. If a required page remains blocked or unparsable, the diagnostic has exactly these four fields:

{
"url": "https://www.youtube.com/watch?v=7ARBJQn6QkM",
"error": "Public YouTube page was blocked (HTTP 403).",
"errorCode": "EPISODE_ACCESS_BLOCKED",
"scrapedAt": "2026-09-08T12:00:00.000Z"
}

Diagnostics are not normal podcast records. The Actor never invents episode titles, show identities, dates, counts, captions, or media URLs.

OUTPUT summary

The OUTPUT key-value record reconciles normal and diagnostic counts and records optional-enrichment availability:

{
"runType": "youtube-public-podcast-run",
"status": "success",
"dataAvailable": true,
"sourceCount": 2,
"sourcesWithItems": 2,
"searchPagesFetched": 1,
"playlistPagesFetched": 1,
"watchPagesFetched": 0,
"pageAttempts": 4,
"itemCount": 3,
"diagnosticCount": 0,
"failedCount": 0,
"blockedCount": 0,
"continuationDetected": true,
"continuationCount": 1,
"episodeDetailsRequested": 2,
"episodeDetailsAvailableCount": 2,
"episodeDetailsUnavailableCount": 0,
"showDetailsRequested": 1,
"showDetailsAvailableCount": 1,
"showDetailsUnavailableCount": 0,
"channelDetailsRequested": 0,
"channelDetailsAvailableCount": 0,
"channelDetailsUnavailableCount": 0,
"proxyRequested": true,
"proxyConfigured": true,
"includeDiagnostics": true,
"includeShowDetails": true,
"includeEpisodeDetails": true,
"maxItems": 5,
"maxPages": 1,
"maxShowEpisodes": 10,
"maxRetries": 1,
"requestDelayMs": 250,
"requestTimeoutSecs": 60,
"languageCode": "en",
"countryCode": "US",
"paginationMethod": "bounded-public-search-pages-no-continuation-calls",
"extractionMethod": "youtube-public-page-embedded-initial-data",
"durationMs": 2450,
"finishedAt": "2026-09-08T12:00:00.000Z"
}

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

Field reference

All normal fields are copied from public source data or identify the request and extraction path. Optional fields are omitted when not exposed or not applicable.

Field groupFields and meaning
IdentityrecordType, type, title, videoId, episodeBrowseId, showBrowseId, showId, channelBrowseId, channelId identify the public result.
Public textdescription, descriptionLength, accessibilityText, subtitle, showName, channelName, duration, publishedText, viewsText, podcastFormat, badges preserve visible text and labels.
Public metricsdurationSeconds and views are parsed values; source text remains in duration, viewsText, and publishedText.
Public links and artworkchannelUrl, showUrl, videoUrl, youtubeMusicUrl, thumbnailUrl, thumbnailCount, and thumbnails contain normalized public URLs and dimensions.
ClassificationisLive, isUpcoming, and podcastMatch retain source flags and the transparent podcast signal.
Optional episode detailepisodeDetailsAvailable, episodeDetailsStatus, episodeTransport, and episode contain public watch-page metadata without caption URLs, signatures, streaming URLs, or playback URLs.
Optional show detailshowDetailsAvailable, showDetailsStatus, and show contain public playlist metadata and bounded nested episodes.
Optional channel detailchannelDetailsAvailable, channelDetailsStatus, and channel contain public channel metadata.
Provenancesource, sourceUrl, searchUrl, rank, page, initialItemsExposed, continuationDetected, continuationCount, sourceRenderer, sourceTransport, extractionMethod, metadataSource, and metadataEnriched explain where and how the row was obtained.
Effective settingsrequestedMaxItems, requestedMaxPages, requestedMaxShowEpisodes, languageCode, countryCode, and scrapedAt record run settings and UTC emission time.
Diagnosticsurl, error, errorCode, and scrapedAt are the only fields in a diagnostic row.

An episode object can contain title, description, descriptionLength, channelName, channelId, channelUrl, lengthSeconds, publishDate, uploadDate, viewCount, category, keywords, isLive, playerStatus, playerReason, captionCount, captions, hasVideo, hasAudio, formatCount, metadataSource, and extractionMethod. Caption items contain only languageCode, name, and kind.

A show object can contain id, title, description, episodeCount, episodeCountText, viewCount, viewCountText, updatedText, creatorName, creatorId, creatorUrl, subscriberCountText, videoCountText, thumbnailUrl, thumbnails, showUrl, episodes, and metadataSource. Each nested show episode can contain title, description, videoId, duration, durationSeconds, publishedText, views, viewsText, videoUrl, youtubeMusicUrl, thumbnailUrl, and thumbnails.

A channel object can contain id, title, description, creatorName, creatorId, creatorUrl, subscriberCountText, videoCountText, thumbnailUrl, thumbnails, channelUrl, and metadataSource.

Pagination, deduplication, and omissions

  • Direct startUrls run before searchQueries and all normal records count toward maxItems.
  • A search query fetches up to maxPages bounded public search pages. Continuation markers are counted, but undocumented continuation tokens are not replayed.
  • A playlist start URL emits a show record and then visible episode records until maxItems; a watch URL emits one episode record; a channel or handle URL emits one channel record.
  • Normal rows are deduplicated by videoId, showBrowseId/showId, or channelId. The same public result found by multiple sources is emitted once.
  • There are no hidden filters or sort options. Source ordering and rank are preserved; podcastMatch is a signal only.
  • includeShowDetails: false skips optional show/channel enrichment. includeEpisodeDetails: false skips optional watch-page enrichment. Missing source fields and disabled optional objects are omitted rather than fabricated.

Proxy, access boundaries, and cost

The default route is a direct public HTML request. With enableProxyFallback: true, a configured account-authorized Apify Proxy route is attempted after direct failure. Explicit proxy input uses the requested proxy configuration first. Transport and proxy state are reported in each row and in OUTPUT; proxy routing does not bypass access controls.

This Actor adds no separate API or licensing fee. Your Apify account may incur the normal compute and, when used, proxy usage charges. Request limits, retries, and pacing are exposed so runs can be sized deliberately.

Troubleshooting

The dataset contains only diagnostics

Check errorCode and OUTPUT. ACCESS_BLOCKED, SEARCH_PAGE_FAILED, or START_URL_FAILED indicates a public access or parsing boundary. Retry later or configure an account-authorized Apify Proxy route. A diagnostic is an honest availability result, not a substitute normal row.

Optional details are unavailable

The core search or start-URL record can still be valid. Review episodeDetailsStatus, showDetailsStatus, or channelDetailsStatus; optional failure is bounded and does not trigger private request methods.

A search contains unrelated results

YouTube search is broad. Use a more specific query and inspect podcastMatch, type, sourceRenderer, and the public title/description. The Actor does not claim that every search result is a podcast.

The run is slow or uses too many requests

Lower maxItems, maxPages, or maxShowEpisodes; set includeShowDetails or includeEpisodeDetails to false; increase requestDelayMs only when pacing is needed. Retries and proxy fallback multiply attempts, and pageAttempts reports the total.

API and support

For programmatic runs, use the Apify Actor API run documentation and pass the same JSON input object. The dataset and OUTPUT key-value record are available through the run’s default storage links.

For a reproducible issue, include the sanitized input, run ID, OUTPUT summary, diagnostic rows, and the public URL category involved. Do not include proxy credentials, cookies, tokens, or private account data.

Use only public pages and comply with YouTube’s terms, robots guidance, applicable law, and the rights of creators. Do not use this Actor to access private content, evade authentication, solve CAPTCHAs, or bypass access controls. The Actor does not use login state, CAPTCHA bypass, stealth browser automation, fingerprint spoofing, private request clients, or signed media extraction. This Actor is not affiliated with YouTube or Google. Results depend on what the public page exposes at run time.