# Podcast Directory & Episodes Scraper (Apple Podcasts + RSS) (`everyotherfriday/podcast-directory`) Actor

Find podcasts by keyword, ID or chart, and export show metadata plus full episode lists with direct audio links, ready for transcription pipelines, research, sponsorship prospecting or content monitoring. No keys required.

- **URL**: https://apify.com/everyotherfriday/podcast-directory.md
- **Developed by:** [Paul Vasquez](https://apify.com/everyotherfriday) (community)
- **Categories:** Social media, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 podcast returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Podcast Directory + Episodes

Discover podcasts through Apple's public directory, retrieve publisher RSS metadata, and export normalized episode records. This actor supports topic research, podcast outreach preparation, catalogue enrichment, and monitoring recent releases. It uses public endpoints without API keys. It downloads metadata and feeds, never the audio files themselves. Results reflect what Apple and publishers expose at run time; this is not a complete historical archive or a listening analytics service.

### Quick start

The supplied INPUT.json searches for `python programming` and `personal finance`, returns at most ten unique podcasts, and includes at most ten episodes per show. Ratings are disabled and each HTTP attempt has a fifteen-second deadline. This small default is intended for the daily automated test without credentials. Healthy sources should finish within two minutes, although upstream timeouts and retries can increase runtime. Run the actor with Python 3.12 and the dependencies in requirements.txt:

```powershell
python -m venv .venv
.venv/Scripts/python.exe -m pip install -r requirements.txt
apify run
```

For the reproducible local evidence run, execute `powershell -File validation/run_live.ps1`. The script creates isolated storage, copies INPUT.json, invokes the real SDK entry point, and records the complete output and elapsed time in validation/results.json. It does not deploy the actor. See VALIDATION.md for observed results and limitations.

### Discovery inputs

`mode` accepts `search`, `lookup`, or `charts`, with search as the default. Search requires a nonempty `queries` array and requests up to 200 Apple matches for each term. Results retain query order, so the first query may fill the entire output cap. Duplicate Apple collection IDs across terms are emitted once. Search does not paginate beyond Apple's 200-result response.

Lookup requires `podcastIds`, containing numeric Apple IDs as strings or HTTPS podcasts.apple.com URLs. Charts retrieves the current overall top 100 for the selected country, then resolves entries through lookup. The optional `genre` filters genre IDs within that overall top 100; it does not claim to retrieve a separate genre-specific chart. Sparse genre matches may therefore produce fewer rows than requested.

`country` defaults to `us` and selects the Apple storefront. It is not the publisher's geographic location. `maxPodcasts` defaults to 50 and caps unique selected shows across all inputs. Inputs accept at most 100 terms or IDs. `timeoutSecs` defaults to 20 and bounds each network attempt. Responses with HTTP 429 or 5xx receive two retries with one- and two-second backoff; transport failures also receive two retries. Other HTTP errors are not retried.

### Episode and enrichment controls

`includeEpisodes` defaults to true. `maxEpisodesPerPodcast` defaults to 50 and applies after deduplication and filtering. Episodes are sorted newest first. `publishedAfter` is an exclusive ISO date or timestamp; timestamps without an offset are interpreted as UTC. When this filter is present, episodes without a usable date are excluded.

RSS is fetched even when episodes are disabled, because it supplies language, description, website, and latest-release enrichment. The actor reads all entries present in that feed, but publishers may expose only a recent window. It does not follow archive pagination or recover deleted episodes. Entries require an audio enclosure; GUID, then audio URL, identifies an episode. If RSS fails or exposes no usable audio entries, Apple lookup supplies recent episodes, capped at 200. A feed failure still produces a free error row even when fallback succeeds.

`includeRatings` defaults to false. Enabling it requests each Apple show page and searches JSON-LD or embedded JSON for aggregateRating. Missing, blocked, or differently structured ratings remain null. Ratings are storefront-specific observations, not guaranteed worldwide totals. No rating-page failure prevents the podcast row from being returned.

### Output and pricing

The dataset uses `rowType` to distinguish podcast, episode, error, and summary rows. Podcast records include identity, author, feed and Apple links, artwork, genres, episode count, language, explicit flag, latest release date, description, website, and optional rating/count. Episode records include podcast identity, GUID, title, plain-text description truncated to 2,000 characters, audio URL/type/bytes, duration in seconds, publication date, season, episode number/type, link, image, and source. Missing upstream values remain null. Apple's episode count may differ from the number currently exposed by RSS.

Each successful podcast row requests one `podcast-returned` event at $0.002. Each successful episode row requests one `episode-returned` event at $0.0002. Error and zero-result summary rows are free. SUMMARY in the key-value store records counts, eligible events, per-show processing times, and reached charge limits. Local SDK runs do not bill. Charging precedes persistence; a storage failure cannot automatically reverse an accepted charge. Configure both custom events and disable synthetic charges before publication.

### Verification and boundaries

Run `.venv/Scripts/python.exe -m unittest discover -s tests -v` for mocked coverage and `apify validate-schema .actor/input_schema.json` for input validation. Fixtures cover search, lookup, RSS durations, filtering, retries, and charging. Public feeds can change, disappear, or reject requests; diagnostics preserve partial results. No proxy is configured by default. Docker execution, hosted storage, and actual platform billing require separate deployment validation. This package has not been pushed or published.

### Example output

Recorded local validation output from [validation/results.json](validation/results.json), the first item in `rows`. Fields are omitted for brevity; retained values are unchanged. This is a historical example, not a current-source claim.

```json
{
    "rowType":  "podcast",
    "podcastId":  "979020229",
    "name":  "Talk Python To Me",
    "author":  "Michael Kennedy",
    "feedUrl":  "https://talkpython.fm/episodes/rss",
    "episodeCount":  563,
    "language":  "en-us",
    "country":  "us",
    "rating":  null,
    "ratingCount":  null
}
```

### Use cases

- A podcast advertising planner searches a topic and reviews recent episode titles to build a shortlist for manual sponsorship research.
- A public relations agency finds relevant shows and follows publisher website links when preparing a guest-pitch research list.
- A media monitoring analyst exports episodes after a chosen date and compares saved results to track releases in a topic area.
- A podcast directory operator enriches known Apple IDs with RSS language, descriptions, and episode metadata for editorial review.

### Pricing example

100 podcast rows and 1,000 episode rows cost (100 x $0.002) + (1,000 x $0.0002) = **$0.40** in declared events. Rates come from [the local event declaration](.actor/pay_per_event.json). This calculation is an event subtotal, not a measured invoice; local validation does not bill.

### Limitations

Directory presence and episode counts are not audience measurements. Search and chart bounds can omit relevant shows, while publisher feed windows limit historical coverage. Null ratings mean no usable value was extracted. Review feed errors alongside successful fallback episodes before judging freshness or completeness.

# Actor input Schema

## `mode` (type: `string`):

Discovery mode.

## `queries` (type: `array`):

Search terms; required in search mode.

## `podcastIds` (type: `array`):

Apple IDs or HTTPS Apple Podcasts URLs; required in lookup mode.

## `country` (type: `string`):

Two-letter Apple storefront code.

## `genre` (type: `string`):

Optional Apple genre ID; filters the overall charts top 100.

## `maxPodcasts` (type: `integer`):

Maximum unique podcasts across all inputs.

## `includeEpisodes` (type: `boolean`):

Return RSS episodes, with Apple lookup fallback.

## `maxEpisodesPerPodcast` (type: `integer`):

Maximum unique matching episodes per podcast.

## `publishedAfter` (type: `string`):

Exclusive ISO timestamp/date; excludes episodes without dates.

## `includeRatings` (type: `boolean`):

Best-effort Apple page aggregate ratings; null if absent.

## `timeoutSecs` (type: `integer`):

Deadline per HTTP attempt; two retries on 429/5xx.

## Actor input object example

```json
{
  "mode": "search",
  "queries": [
    "python programming",
    "personal finance"
  ],
  "country": "us",
  "maxPodcasts": 5,
  "includeEpisodes": true,
  "maxEpisodesPerPodcast": 10,
  "includeRatings": false,
  "timeoutSecs": 20
}
```

# Actor output Schema

## `rows` (type: `string`):

All output rows.

## `summary` (type: `string`):

Counts and timings.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxPodcasts": 5,
    "maxEpisodesPerPodcast": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("everyotherfriday/podcast-directory").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxPodcasts": 5,
    "maxEpisodesPerPodcast": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("everyotherfriday/podcast-directory").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxPodcasts": 5,
  "maxEpisodesPerPodcast": 10
}' |
apify call everyotherfriday/podcast-directory --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,everyotherfriday/podcast-directory"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XMVhtVBaq39Ncn4HA/builds/l3JFzmbuzyrKHsiLl/openapi.json
