Podcast Directory & Episodes Scraper (Apple Podcasts + RSS) avatar

Podcast Directory & Episodes Scraper (Apple Podcasts + RSS)

Pricing

from $2.00 / 1,000 podcast returneds

Go to Apify Store
Podcast Directory & Episodes Scraper (Apple Podcasts + RSS)

Podcast Directory & Episodes Scraper (Apple Podcasts + RSS)

Find podcasts by keyword, ID or chart, and export show metadata plus full episode lists with direct audio links, ready for transcription pipelines, research, sponsorship prospecting or content monitoring. No keys required.

Pricing

from $2.00 / 1,000 podcast returneds

Rating

0.0

(0)

Developer

Paul Vasquez

Paul Vasquez

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Podcast Directory + Episodes

Discover podcasts through Apple's public directory, retrieve publisher RSS metadata, and export normalized episode records. This actor supports topic research, podcast outreach preparation, catalogue enrichment, and monitoring recent releases. It uses public endpoints without API keys. It downloads metadata and feeds, never the audio files themselves. Results reflect what Apple and publishers expose at run time; this is not a complete historical archive or a listening analytics service.

Quick start

The supplied INPUT.json searches for python programming and personal finance, returns at most ten unique podcasts, and includes at most ten episodes per show. Ratings are disabled and each HTTP attempt has a fifteen-second deadline. This small default is intended for the daily automated test without credentials. Healthy sources should finish within two minutes, although upstream timeouts and retries can increase runtime. Run the actor with Python 3.12 and the dependencies in requirements.txt:

python -m venv .venv
.venv/Scripts/python.exe -m pip install -r requirements.txt
apify run

For the reproducible local evidence run, execute powershell -File validation/run_live.ps1. The script creates isolated storage, copies INPUT.json, invokes the real SDK entry point, and records the complete output and elapsed time in validation/results.json. It does not deploy the actor. See VALIDATION.md for observed results and limitations.

Discovery inputs

mode accepts search, lookup, or charts, with search as the default. Search requires a nonempty queries array and requests up to 200 Apple matches for each term. Results retain query order, so the first query may fill the entire output cap. Duplicate Apple collection IDs across terms are emitted once. Search does not paginate beyond Apple's 200-result response.

Lookup requires podcastIds, containing numeric Apple IDs as strings or HTTPS podcasts.apple.com URLs. Charts retrieves the current overall top 100 for the selected country, then resolves entries through lookup. The optional genre filters genre IDs within that overall top 100; it does not claim to retrieve a separate genre-specific chart. Sparse genre matches may therefore produce fewer rows than requested.

country defaults to us and selects the Apple storefront. It is not the publisher's geographic location. maxPodcasts defaults to 50 and caps unique selected shows across all inputs. Inputs accept at most 100 terms or IDs. timeoutSecs defaults to 20 and bounds each network attempt. Responses with HTTP 429 or 5xx receive two retries with one- and two-second backoff; transport failures also receive two retries. Other HTTP errors are not retried.

Episode and enrichment controls

includeEpisodes defaults to true. maxEpisodesPerPodcast defaults to 50 and applies after deduplication and filtering. Episodes are sorted newest first. publishedAfter is an exclusive ISO date or timestamp; timestamps without an offset are interpreted as UTC. When this filter is present, episodes without a usable date are excluded.

RSS is fetched even when episodes are disabled, because it supplies language, description, website, and latest-release enrichment. The actor reads all entries present in that feed, but publishers may expose only a recent window. It does not follow archive pagination or recover deleted episodes. Entries require an audio enclosure; GUID, then audio URL, identifies an episode. If RSS fails or exposes no usable audio entries, Apple lookup supplies recent episodes, capped at 200. A feed failure still produces a free error row even when fallback succeeds.

includeRatings defaults to false. Enabling it requests each Apple show page and searches JSON-LD or embedded JSON for aggregateRating. Missing, blocked, or differently structured ratings remain null. Ratings are storefront-specific observations, not guaranteed worldwide totals. No rating-page failure prevents the podcast row from being returned.

Output and pricing

The dataset uses rowType to distinguish podcast, episode, error, and summary rows. Podcast records include identity, author, feed and Apple links, artwork, genres, episode count, language, explicit flag, latest release date, description, website, and optional rating/count. Episode records include podcast identity, GUID, title, plain-text description truncated to 2,000 characters, audio URL/type/bytes, duration in seconds, publication date, season, episode number/type, link, image, and source. Missing upstream values remain null. Apple's episode count may differ from the number currently exposed by RSS.

Each successful podcast row requests one podcast-returned event at $0.002. Each successful episode row requests one episode-returned event at $0.0002. Error and zero-result summary rows are free. SUMMARY in the key-value store records counts, eligible events, per-show processing times, and reached charge limits. Local SDK runs do not bill. Charging precedes persistence; a storage failure cannot automatically reverse an accepted charge. Configure both custom events and disable synthetic charges before publication.

Verification and boundaries

Run .venv/Scripts/python.exe -m unittest discover -s tests -v for mocked coverage and apify validate-schema .actor/input_schema.json for input validation. Fixtures cover search, lookup, RSS durations, filtering, retries, and charging. Public feeds can change, disappear, or reject requests; diagnostics preserve partial results. No proxy is configured by default. Docker execution, hosted storage, and actual platform billing require separate deployment validation. This package has not been pushed or published.

Example output

Recorded local validation output from validation/results.json, the first item in rows. Fields are omitted for brevity; retained values are unchanged. This is a historical example, not a current-source claim.

{
"rowType": "podcast",
"podcastId": "979020229",
"name": "Talk Python To Me",
"author": "Michael Kennedy",
"feedUrl": "https://talkpython.fm/episodes/rss",
"episodeCount": 563,
"language": "en-us",
"country": "us",
"rating": null,
"ratingCount": null
}

Use cases

  • A podcast advertising planner searches a topic and reviews recent episode titles to build a shortlist for manual sponsorship research.
  • A public relations agency finds relevant shows and follows publisher website links when preparing a guest-pitch research list.
  • A media monitoring analyst exports episodes after a chosen date and compares saved results to track releases in a topic area.
  • A podcast directory operator enriches known Apple IDs with RSS language, descriptions, and episode metadata for editorial review.

Pricing example

100 podcast rows and 1,000 episode rows cost (100 x $0.002) + (1,000 x $0.0002) = $0.40 in declared events. Rates come from .actor/pay_per_event.json. This calculation is an event subtotal, not a measured invoice; local validation does not bill.

Limitations

Directory presence and episode counts are not audience measurements. Search and chart bounds can omit relevant shows, while publisher feed windows limit historical coverage. Null ratings mean no usable value was extracted. Review feed errors alongside successful fallback episodes before judging freshness or completeness.