Apple Podcasts Scraper — RSS Feeds, iTunes Search & API avatar

Apple Podcasts Scraper — RSS Feeds, iTunes Search & API

Pricing

from $0.87 / 1,000 podcast/episodes

Go to Apify Store
Apple Podcasts Scraper — RSS Feeds, iTunes Search & API

Apple Podcasts Scraper — RSS Feeds, iTunes Search & API

Scrape Apple Podcasts via the official iTunes API — search shows, get metadata (genre, artwork, feed URL, episode count) and full episode lists from RSS. No auth, no proxy. Each record has parse_confidence. Pay per result.

Pricing

from $0.87 / 1,000 podcast/episodes

Rating

0.0

(0)

Developer

Vitalii Bondarev

Vitalii Bondarev

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

2

Monthly active users

19 hours ago

Last modified

Categories

Share

Apple Podcasts Scraper — iTunes Search, Episodes & RSS | from $0.50/1K

Built for podcast-app developers, content marketers, and lead-gen teams that need structured Apple Podcasts data — show metadata, episode lists from RSS, genre filters, and country selection — via the official iTunes API. No auth, no proxy, zero COGS.

Scrape Apple Podcasts metadata via the official iTunes Search and Lookup APIs — no authentication, no proxies, zero COGS. Optionally retrieve episodes directly from each podcast's RSS feed.

What you can scrape

  • Podcast metadata — title, author, genre, episode count, artwork, feed URL, iTunes URL, release date
  • Episodes (optional) — title, GUID, publish date, duration, description, audio file URL from the podcast's RSS feed

Use cases

  • Podcast directory and competitive research
  • Lead generation (find podcasts by topic, contact via feed)
  • Content monitoring — track episode counts, new releases
  • App development — build podcast apps with rich metadata

Input

FieldTypeDefaultDescription
searchTermsstring[]["technology"]Keywords to search Apple Podcasts
podcastIdsstring[]—Numeric iTunes podcast IDs for direct lookup
countrystring"us"Two-letter country code (us, gb, de, fr, jp…)
maxItemsinteger50Max total podcasts to return (0 = no limit)
includeEpisodesbooleanfalseAlso fetch episodes from RSS feed
maxEpisodesPerPodcastinteger20Max episodes per podcast (0 = all)
maxSearchResultsinteger50Max search results per term (iTunes cap: 200)

Finding a podcast ID

Look at the Apple Podcasts URL: podcasts.apple.com/us/podcast/name/id**470624027** — the number after id is the podcast ID.

Output schema

Every row in the dataset has these fields:

FieldTypeDescription
record_typestring"podcast" or "episode"
podcast_idstringNumeric iTunes podcast ID
titlestringPodcast or episode title
artiststringAuthor / creator name
genrestringPrimary genre (Technology, True Crime, etc.)
genresstring[]All genres
episode_countinteger|nullTotal episodes (podcast rows only)
ratingfloat|nullAverage user rating (where available)
rating_countinteger|nullNumber of ratings
countrystringInput country code
feed_urlstring|nullRSS feed URL
artwork_urlstring|null600×600 artwork image
release_datestring|nullISO 8601 UTC — latest episode release
itunes_urlstring|nullApple Podcasts browse URL
episode_guidstring|nullEpisode GUID (episode rows only)
pub_date_rawstring|nullEpisode publish date (RFC 2822)
durationstring|nullEpisode duration HH:MM:SS
episode_descstring|nullEpisode description (first 2000 chars)
enclosure_urlstring|nullEpisode audio file URL
parse_confidencefloat0.0–1.0 data quality score
warningsstring[]List of any missing/unexpected fields
scraped_atstringISO 8601 UTC scrape timestamp

Pricing

Pay per result — each podcast row or episode row = 1 billable event: from $0.50 per 1,000 records.

VolumeCost
100 podcasts~$0.05
1,000 podcasts~$0.50
1,000 podcasts + 20 episodes each = 21,000 rows~$10.50

FAQ

Do I need a proxy or API key? No. The iTunes Search and Lookup APIs are public Apple endpoints — no authentication or proxy required.

What output formats are available? JSON, CSV, and Excel — downloadable from the Apify dataset UI or via the REST API.

Can I schedule this to run automatically? Yes. Use Apify's scheduler to monitor new episodes or track podcast rankings on a daily schedule, with webhook delivery to your pipeline.

Why are rating and rating_count null for most podcasts? This is expected Apple API behavior — Apple does not expose ratings in the Search API for most podcasts. This is not a bug in the scraper.

Why this scraper beats the rest

  • Official APIs only — iTunes Search + Lookup APIs are Apple's own public endpoints. Zero DOM fragility, no CAPTCHA risk, no proxy required.
  • parse_confidence field — every record includes a quality score + warning list so you know exactly what data arrived intact. No competitor offers this.
  • Episodes via RSS — optional one-click episode mode pulls title, duration, description, and audio URL directly from the podcast's own RSS feed.
  • Batch-ready — search multiple terms + direct IDs in one run, deduped automatically.
  • Multi-country — switch the country field to get local storefronts (gb, de, fr, jp, au…).

Technical notes

  • iTunes Search API limit: 200 results per query.
  • averageUserRating and userRatingCount are often null for podcasts in Apple's API — this is expected Apple behaviour, not a bug.
  • RSS episode fetch adds one HTTP request per podcast. For large batches with includeEpisodes=true, runtime increases proportionally.
  • Not affiliated with Apple Inc.

Example output

{
"record_type": "podcast",
"podcast_id": "470624027",
"title": "TED Tech",
"artist": "TED Tech",
"genre": "Technology",
"genres": ["Technology", "Podcasts"],
"episode_count": 268,
"rating": null,
"rating_count": null,
"country": "us",
"feed_url": "https://feeds.acast.com/public/shows/...",
"artwork_url": "https://is1-ssl.mzstatic.com/.../600x600bb.jpg",
"release_date": "2026-05-29T04:00:00Z",
"itunes_url": "https://podcasts.apple.com/us/podcast/ted-tech/id470624027",
"episode_guid": null,
"pub_date_raw": null,
"duration": null,
"episode_desc": null,
"enclosure_url": null,
"parse_confidence": 1.0,
"warnings": [],
"scraped_at": "2026-05-31T12:00:00Z"
}

Use with AI agents (MCP)

This actor is available as an MCP tool for Claude, GPT-4, and other AI agents that support the Model Context Protocol:

https://mcp.apify.com/?tools=bovi/podcast-scraper

Search Apple Podcasts by keyword or look up shows by iTunes ID — ideal for AI podcast discovery assistants, content monitoring bots, and lead-gen pipelines targeting podcasters by niche.


vs. competitors

This actorTypical podcast scraper
Data sourceOfficial iTunes API + RSSHTML / unofficial APIs
Episode content (RSS)✓ (includeEpisodes toggle)Rarely
Country targeting✓Usually US-only
parse_confidence✓No
Proxy neededNoOften required
Pricefrom $0.50/1K$2–5/1K

Integrations

Built for podcast-app developers and content marketers sourcing show metadata and episode feeds by genre and country — the JSON/dataset output drops into the tools you already run, no glue code:

  • n8n / Make / Zapier — trigger a run or pipe every new dataset item into 500+ apps (Google Sheets, Airtable, Slack, HubSpot, your database) with no code: n8n, Make, Zapier.
  • Webhooks — fire your own endpoint the moment a run finishes, to push results straight into your pipeline (docs).
  • MCP server — expose this actor as a tool to Claude, Cursor, or any MCP client so an AI agent can pull this data mid-conversation (guide).
  • API & SDKs — fetch the dataset as JSON, CSV, or Excel through the Apify REST API or the Python / JS SDKs.

See all Apify integrations.

Usage statistics

This Actor creates a small, content-free summary at the end of each run. It is used only to monitor reliability and improve this Actor. A copy is saved as USAGE_STATS in your own Apify key-value store, so you can see the exact record created for your run.

Set disableUsageStats to true in the input to opt out. Nothing is sent then; your USAGE_STATS record only says that statistics were disabled.

Only these fields are recorded:

  • schema version, Actor name and build number;
  • UTC start and finish hour (not a precise timestamp);
  • run duration, number of results and time to the first result, each as a coarse range;
  • whether the result was empty, the end status, and an error type from a fixed list;
  • memory setting and counts of charged events;
  • names of the input fields you set, never their values;
  • the selected option for input fields that offer a fixed list of choices (for example a sort order).

We do not collect input text, search terms, URLs, domains, usernames, email addresses, names, proxy credentials, tokens, scraped records, output items, raw error messages, stack traces, or hashes of any of those values. Records are kept for no longer than 13 months, used only as aggregated operational statistics, and never sold or shared.

Additional fields (Phase 2)

This Actor also records your Apify user ID, whether Apify marks the account as paying, the size range of list inputs, the selected country when the input offers a fixed list of countries, and one category from a fixed Actor taxonomy. We use these fields only for aggregate reliability, repeat-use and cross-Actor analysis; reports suppress any cell with fewer than five distinct users.

The same disableUsageStats: true input flag turns these fields off too. The user ID is removed after 13 months; we do not export, sell, share, or attempt to re-identify this data.

Run-outcome signals (v2)

To learn whether a run did what it was asked to do, the record also holds a few more coarse ranges and yes/no flags. None of them contains content:

  • the result limit you asked for (a range, when the input has one) and what share of it was delivered;
  • results delivered per input item you listed (a range);
  • output quality as ranges: how fully the result fields were filled, the share of rows that look like errors, the share of duplicate rows, and how many different fields appeared. These are counted in memory while results are saved; no result content is kept;
  • how the run was started (console, API, schedule, webhook, another Actor);
  • how it ended: stopped by you, timed out, reached the requested limit, stopped by the charge limit, and how many times the platform moved the run;
  • if this Actor reports it: how many items to process worked or failed (ranges) and one failure reason from a fixed list;
  • a short code made from the names of the input fields you set, never their values.

Repeat-run fingerprint (v2)

When your Apify user ID is recorded (see above), the record also holds an 8-character one-way code made from your input (proxy settings left out) and this Actor's name. It only lets us see that the same account ran the same input again soon after an unsatisfying run; we never see the input itself. It is stored only in the database, never published, and reports use it in aggregate with the same five-user minimum. It is the one exception to the statement above that no hashes are collected, and disableUsageStats: true turns it off.