Apple Podcasts Scraper avatar

Apple Podcasts Scraper

Pricing

from $1.26 / 1,000 results

Go to Apify Store
Apple Podcasts Scraper

Apple Podcasts Scraper

Podcast shows and their COMPLETE episode lists. Apple's own API silently caps episodes at 200 - a 558-episode show loses 358 - so this reads the show's RSS feed, whose URL Apple hands back in every result, and flags any row that came from the capped path instead.

Pricing

from $1.26 / 1,000 results

Rating

0.0

(0)

Developer

Ibnu Adzim

Ibnu Adzim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Categories

Share

Podcast shows and their complete episode lists — search Apple Podcasts or look up shows directly, then pull every episode with its audio URL, duration and publish date. HTTP-only, no API key, no login, no browser.

The thing that makes this different

Apple's own API silently caps episodes at 200 — and hands you the way past it in every result.

Measured on Talk Python To Me (id 979020229), which reports trackCount: 558:

limit sentepisodes returned
1010
100100
200200
300200 — byte-identical 394,700-byte body
500200 — byte-identical 394,700-byte body

No error, no echo of the limit you asked for, and the same bytes back for 300 as for 500. So 358 of 558 episodes are simply unreachable that way.

The escape hatch is feedUrl — a field the API volunteers in every result. Fetching that show's own RSS feed returned 558 <item> elements, exactly matching the count Apple itself reported. episodeSource therefore defaults to feed; the api path is offered for speed and flags itself as truncated (episodesTruncatedByApiCap) whenever the show has more episodes than it can return.

Three more upstream quirks it corrects

1. results[0] is the show, not an episode

A lookup with entity=podcastEpisode&limit=5 returns resultCount: 6:

[0] wrapperType "track" kind "podcast" <- the SHOW
[1] wrapperType "podcastEpisode" kind "podcast-episode"
... four more episodes

So the count is always episodes + 1, and anything reading results[0] as an episode publishes the show's metadata as a row — plausible title, no audio URL. Rows are selected by wrapperType, never by index.

2. The two episode sources emit different date formats in the same field

The RSS feed emits RFC-822 ("Wed, 19 Aug 2026 18:57:47 +0000"); the API emits ISO ("2026-08-19T18:57:47Z"). Since this actor makes it easy to mix them, both are normalised into publishedAt as ISO, with the original kept beside it as publishedAtRaw.

3. Feed tag counts are off by one

The Talk Python feed contains 559 <title> tags for 558 episodes — the extra is the channel's own, and the same is true of <pubDate> and <description>. Any document-wide count overshoots by exactly one, so parsing is scoped to <item> elements.

Output

One SEARCH_SUMMARY per run, one PODCAST per show, one EPISODE per episode, one ERROR per id that could not be resolved.

PODCAST carries the upstream object verbatim plus podcastId, podcastName, publisher, feedUrl, podcastUrl, artworkUrl, primaryGenre, genreList, episodeCountReported, episodesCollected, episodeSource, episodesTruncatedByApiCap and feedError.

EPISODE carries episodeTitle, description, publishedAt (ISO), publishedAtRaw, guid, audioUrl, audioType, audioLengthBytes, durationSeconds, durationRaw, episodeNumber, seasonNumber and episodeUrl.

Limits

  • Search caps at 100 shows: limit=200 and limit=300 both return 100.
  • Feeds are third-party hosts. A feed that is slow, moved or malformed degrades that one show's episode list — the show row is still complete and feedError says what happened.
  • A show id that does not exist returns HTTP 200 with resultCount: 0 — an honest zero, reported rather than inferred from a failure.
  • No WAF on Apple's side; the proxy is offered but off by default — podcast feeds are ordinary web servers run by small publishers, and there is nothing here to get past.