Apple Podcasts Scraper
Pricing
from $1.26 / 1,000 results
Apple Podcasts Scraper
Podcast shows and their COMPLETE episode lists. Apple's own API silently caps episodes at 200 - a 558-episode show loses 358 - so this reads the show's RSS feed, whose URL Apple hands back in every result, and flags any row that came from the capped path instead.
Pricing
from $1.26 / 1,000 results
Rating
0.0
(0)
Developer
Ibnu Adzim
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 days ago
Last modified
Categories
Share
Podcast shows and their complete episode lists — search Apple Podcasts or look up shows directly, then pull every episode with its audio URL, duration and publish date. HTTP-only, no API key, no login, no browser.
The thing that makes this different
Apple's own API silently caps episodes at 200 — and hands you the way past it in every result.
Measured on Talk Python To Me (id 979020229), which reports trackCount: 558:
limit sent | episodes returned |
|---|---|
| 10 | 10 |
| 100 | 100 |
| 200 | 200 |
| 300 | 200 — byte-identical 394,700-byte body |
| 500 | 200 — byte-identical 394,700-byte body |
No error, no echo of the limit you asked for, and the same bytes back for 300 as for 500. So 358 of 558 episodes are simply unreachable that way.
The escape hatch is feedUrl — a field the API volunteers in every result.
Fetching that show's own RSS feed returned 558 <item> elements, exactly
matching the count Apple itself reported. episodeSource therefore defaults
to feed; the api path is offered for speed and flags itself as
truncated (episodesTruncatedByApiCap) whenever the show has more episodes
than it can return.
Three more upstream quirks it corrects
1. results[0] is the show, not an episode
A lookup with entity=podcastEpisode&limit=5 returns resultCount: 6:
[0] wrapperType "track" kind "podcast" <- the SHOW[1] wrapperType "podcastEpisode" kind "podcast-episode"... four more episodes
So the count is always episodes + 1, and anything reading results[0] as an
episode publishes the show's metadata as a row — plausible title, no audio
URL. Rows are selected by wrapperType, never by index.
2. The two episode sources emit different date formats in the same field
The RSS feed emits RFC-822 ("Wed, 19 Aug 2026 18:57:47 +0000"); the API
emits ISO ("2026-08-19T18:57:47Z"). Since this actor makes it easy to mix
them, both are normalised into publishedAt as ISO, with the original kept
beside it as publishedAtRaw.
3. Feed tag counts are off by one
The Talk Python feed contains 559 <title> tags for 558 episodes —
the extra is the channel's own, and the same is true of <pubDate> and
<description>. Any document-wide count overshoots by exactly one, so parsing
is scoped to <item> elements.
Output
One SEARCH_SUMMARY per run, one PODCAST per show, one EPISODE per
episode, one ERROR per id that could not be resolved.
PODCAST carries the upstream object verbatim plus podcastId,
podcastName, publisher, feedUrl, podcastUrl, artworkUrl,
primaryGenre, genreList, episodeCountReported, episodesCollected,
episodeSource, episodesTruncatedByApiCap and feedError.
EPISODE carries episodeTitle, description, publishedAt (ISO),
publishedAtRaw, guid, audioUrl, audioType, audioLengthBytes,
durationSeconds, durationRaw, episodeNumber, seasonNumber and
episodeUrl.
Limits
- Search caps at 100 shows:
limit=200andlimit=300both return 100. - Feeds are third-party hosts. A feed that is slow, moved or malformed
degrades that one show's episode list — the show row is still complete and
feedErrorsays what happened. - A show id that does not exist returns HTTP 200 with
resultCount: 0— an honest zero, reported rather than inferred from a failure. - No WAF on Apple's side; the proxy is offered but off by default — podcast feeds are ordinary web servers run by small publishers, and there is nothing here to get past.