# Apple Podcasts Scraper (`scrapyx/apple-podcasts-scraper`) Actor

Podcast shows and their COMPLETE episode lists. Apple's own API silently caps episodes at 200 - a 558-episode show loses 358 - so this reads the show's RSS feed, whose URL Apple hands back in every result, and flags any row that came from the capped path instead.

- **URL**: https://apify.com/scrapyx/apple-podcasts-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Videos
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.26 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Apple Podcasts Scraper

Podcast shows and their **complete episode lists** — search Apple Podcasts or
look up shows directly, then pull every episode with its audio URL, duration
and publish date. HTTP-only, no API key, no login, no browser.

### The thing that makes this different

**Apple's own API silently caps episodes at 200 — and hands you the way past
it in every result.**

Measured on *Talk Python To Me* (id 979020229), which reports `trackCount: 558`:

| `limit` sent | episodes returned |
| ---: | ---: |
| 10 | 10 |
| 100 | 100 |
| 200 | 200 |
| **300** | **200** — byte-identical 394,700-byte body |
| **500** | **200** — byte-identical 394,700-byte body |

No error, no echo of the limit you asked for, and the *same bytes* back for
300 as for 500. So 358 of 558 episodes are simply unreachable that way.

The escape hatch is `feedUrl` — a field the API volunteers in every result.
Fetching that show's own RSS feed returned **558 `<item>` elements**, exactly
matching the count Apple itself reported. `episodeSource` therefore defaults
to `feed`; the `api` path is offered for speed and **flags itself** as
truncated (`episodesTruncatedByApiCap`) whenever the show has more episodes
than it can return.

### Three more upstream quirks it corrects

#### 1. `results[0]` is the show, not an episode

A lookup with `entity=podcastEpisode&limit=5` returns `resultCount: 6`:

```
[0]  wrapperType "track"           kind "podcast"           <- the SHOW
[1]  wrapperType "podcastEpisode"  kind "podcast-episode"
...  four more episodes
```

So the count is always episodes + 1, and anything reading `results[0]` as an
episode publishes the show's metadata as a row — plausible title, no audio
URL. Rows are selected by `wrapperType`, never by index.

#### 2. The two episode sources emit different date formats in the same field

The RSS feed emits RFC-822 (`"Wed, 19 Aug 2026 18:57:47 +0000"`); the API
emits ISO (`"2026-08-19T18:57:47Z"`). Since this actor makes it easy to mix
them, both are normalised into `publishedAt` as ISO, with the original kept
beside it as `publishedAtRaw`.

#### 3. Feed tag counts are off by one

The Talk Python feed contains **559** `<title>` tags for **558** episodes —
the extra is the channel's own, and the same is true of `<pubDate>` and
`<description>`. Any document-wide count overshoots by exactly one, so parsing
is scoped to `<item>` elements.

### Output

One `SEARCH_SUMMARY` per run, one `PODCAST` per show, one `EPISODE` per
episode, one `ERROR` per id that could not be resolved.

`PODCAST` carries the upstream object verbatim plus `podcastId`,
`podcastName`, `publisher`, `feedUrl`, `podcastUrl`, `artworkUrl`,
`primaryGenre`, `genreList`, `episodeCountReported`, `episodesCollected`,
`episodeSource`, `episodesTruncatedByApiCap` and `feedError`.

`EPISODE` carries `episodeTitle`, `description`, `publishedAt` (ISO),
`publishedAtRaw`, `guid`, `audioUrl`, `audioType`, `audioLengthBytes`,
`durationSeconds`, `durationRaw`, `episodeNumber`, `seasonNumber` and
`episodeUrl`.

### Limits

- **Search caps at 100** shows: `limit=200` and `limit=300` both return 100.
- **Feeds are third-party hosts.** A feed that is slow, moved or malformed
  degrades that one show's episode list — the show row is still complete and
  `feedError` says what happened.
- A show id that does not exist returns HTTP 200 with `resultCount: 0` — an
  honest zero, reported rather than inferred from a failure.
- No WAF on Apple's side; the proxy is offered but **off by default** — podcast
  feeds are ordinary web servers run by small publishers, and there is nothing
  here to get past.

# Actor input Schema

## `mode` (type: `string`):

search = find shows by keyword. podcasts = look up shows you name.

## `searchTerm` (type: `string`):

For mode='search'. Apple caps search at 100 shows — limit=200 and limit=300 both return 100.

## `podcastIds` (type: `array`):

For mode='podcasts'. A numeric Apple id like 979020229, or a podcasts.apple.com URL. An id that does not exist returns an honest zero rather than an error.

## `episodeSource` (type: `string`):

feed = the show's own RSS feed, whose URL Apple hands back in every result. This is the COMPLETE list. api = Apple's lookup endpoint, which silently caps at 200 episodes (limit=300 and limit=500 both returned exactly 200, byte-identical) — fast, but a 558-episode show loses 358 of them. Rows fetched this way flag themselves as truncated. none = skip episodes for a fast catalogue sweep.

## `country` (type: `string`):

Two-letter code used for the search and lookup.

## `maxPodcasts` (type: `integer`):

Set 0 for unlimited (still bounded by Apple's 100-show search cap).

## `maxEpisodesPerPodcast` (type: `integer`):

Set 0 for unlimited. With episodeSource='feed' this is the only bound — the feed carries every episode the show has published.

## `maxConcurrency` (type: `integer`):

Shows processed at once. Each show's feed is an independent third-party fetch.

## `minRequestInterval` (type: `integer`):

Politeness pacing shared across all workers. 0 uses the built-in default.

## `proxyConfiguration` (type: `object`):

OFF by default, deliberately. There is no WAF on Apple's side, and podcast feeds are ordinary web servers run by small publishers — there is nothing here to get past, and routing their traffic through a shared residential pool serves no purpose.

## Actor input object example

```json
{
  "mode": "search",
  "searchTerm": "python programming",
  "podcastIds": [
    "979020229",
    "https://podcasts.apple.com/us/podcast/talk-python-to-me/id979020229"
  ],
  "episodeSource": "feed",
  "country": "US",
  "maxPodcasts": 10,
  "maxEpisodesPerPodcast": 200,
  "maxConcurrency": 4,
  "minRequestInterval": 0,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/apple-podcasts-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/apple-podcasts-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call scrapyx/apple-podcasts-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/apple-podcasts-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/E69IgkxTIRN0fX5qZ/builds/oyhNaHGF0RdAX8N6u/openapi.json
