# Podcast RSS Scraper & Transcriber (`celebrated-quadraphonic/podcast-rss-scraper`) Actor

Scrape any podcast RSS feed for episodes with metadata, durations and chapters. Get FREE publisher transcripts (Podcasting 2.0) or fall back to Whisper with your own key. Bulk-import subscriptions via OPML, filter by date/duration/title, and monitor shows for new episodes.

- **URL**: https://apify.com/celebrated-quadraphonic/podcast-rss-scraper.md
- **Developed by:** [XiaoZhi DataTools](https://apify.com/celebrated-quadraphonic) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 episode transcripts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Podcast RSS Scraper & Transcriber — Episodes, Free Transcripts, SRT to Text

Scrape **any podcast** from its RSS feed or discover shows via **Apple Podcasts search** —
titles, descriptions, dates, durations, audio URLs, chapters, **and transcripts**.
Pure HTTP: **no browser, no proxies, no CAPTCHAs**. Fast and cheap.

### Why this Actor

- **Free transcripts first.** Many shows publish official transcripts with every
  episode (Podcasting 2.0 `<podcast:transcript>`). This Actor fetches them at
  **zero cost** — no API key, no per-minute ASR charges. Plain text, SRT and VTT
  are preserved as published.
- **Whisper fallback.** Episodes without a publisher transcript can be transcribed
  with the OpenAI Whisper API using *your own* key — you pay OpenAI directly,
  nothing is marked up here.
- **Rich metadata.** Show author, categories, language, explicit flag, artwork;
  per-episode season/episode numbers, duration, GUID, chapters URL.
- **Filters.** Date range, min/max duration, and keyword matching on episode titles.
- **Monitor mode.** Run on a schedule with `newEpisodesOnly` and get just the
  fresh episodes since the last run — perfect for tracking shows.
- **OPML bulk import.** Point it at podcast-app subscription exports and
  scrape dozens of shows in one run.
- **Transcript formats & diagnostics.** Choose `text`/`srt`/`vtt` output per
  episode; `transcriptStatus` tells you exactly why an episode has no
  transcript.
- **Fast.** Concurrent fetching with no headless browser in the loop.

### Input

| Field | What it does |
|---|---|
| `feedUrls` | Direct RSS feed URLs, one per line |
| `opmlUrls` | OPML subscription lists (URLs or pasted OPML XML) to bulk-import feeds — monitor many shows at once |
| `searchTerms` | Keywords resolved to show feeds via the Apple Podcasts (iTunes) Search API — no key needed |
| `country` | iTunes storefront for search (`US`, `GB`, `DE`, `JP`…) |
| `maxResultsPerSearch` / `maxEpisodesPerFeed` | Caps per search term and per feed |
| `publishedAfter` / `publishedBefore` | Date window, `YYYY-MM-DD` |
| `minDurationSeconds` / `maxDurationSeconds` | Skip trailers, teasers, or over-long episodes |
| `titleContains` | Keyword filter on titles, `|`-separated (`interview|founder`) |
| `usePublisherTranscripts` | Fetch free official transcripts (default: on) |
| `transcriptFormats` | Which transcript formats to emit per episode: `text`, `srt`, `vtt` (default: text + srt) |
| `transcriptLanguage` | ISO-639-1 hint (e.g. `en`) passed to Whisper for faster, more accurate transcription |
| `transcribe` + `openaiApiKey` | Whisper fallback for episodes without a publisher transcript |
| `newEpisodesOnly` | Monitor mode — only episodes unseen by previous runs |
| `requestConcurrency` | Parallel fetches (default 5) |

```json
{
  "searchTerms": ["artificial intelligence"],
  "maxResultsPerSearch": 5,
  "maxEpisodesPerFeed": 20,
  "publishedAfter": "2025-01-01",
  "minDurationSeconds": 300,
  "usePublisherTranscripts": true,
  "transcribe": false
}
```

### Output

One dataset item per episode:

- `podcastTitle`, `podcastAuthor`, `podcastCategories`, `podcastLanguage`, `podcastExplicit`, `podcastFeedUrl`, `podcastLink`, `podcastImage`
- `episodeTitle`, `episodeDescription`, `publishedAt`, `durationSeconds`, `seasonNumber`, `episodeNumber`, `audioUrl`, `episodeUrl`, `episodeGuid`, `episodeImage`, `chaptersUrl`
- `transcriptUrl`, `transcriptType`, `transcriptSource` (`publisher` | `whisper`), `transcriptText`, `transcriptSrt`, `transcriptVtt`
- `transcriptStatus` — per-episode transcript diagnostics: `found`, `no_publisher_tag`, `download_failed`, `empty_transcript`, `format_fallback`, `whisper_failed`, `not_attempted`. The run log also prints a status breakdown so you can see transcript coverage at a glance.

### OPML bulk import

Podcast apps (Overcast, Pocket Casts, Apple Podcasts…) export subscriptions
as OPML. Paste OPML URLs — or raw OPML XML — into `opmlUrls` to scrape all
subscribed shows in one run. Combined with `newEpisodesOnly` on a schedule,
this becomes a full monitoring pipeline for dozens of shows.

### Cost

- Metadata-only and publisher-transcript runs cost almost nothing — a few HTTP
  requests per feed, **no proxy spend**.
- Whisper transcription downloads at most ~1 MB/minute of audio and caps uploads
  at 20 MB per episode (Whisper's limit is 25 MB). `maxTranscribeEpisodes`
  bounds paid transcriptions per run. You pay OpenAI directly for Whisper usage.

### Use cases

- Build podcast search / recommendation datasets
- Content research: mine transcripts for topics, quotes, and trends
- Monitor shows for new episodes on a schedule
- Feed RAG pipelines and AI agents with episode text
- Lead generation: find shows and hosts in any niche

### Run locally

```bash
pip install -r requirements.txt
apify run
```

### Changelog

- **0.6** — Fixed `publishedBefore` excluding the whole end day ("on or before
  this date" now really includes that day); publisher transcript downloads are
  throttled by `requestConcurrency` instead of all firing at once; Whisper
  fallback now transcribes concurrently; fixed invalid SRT timestamps like
  `00:00:05,1000` when JSON transcript segments round up to a full second;
  invalid `publishedAfter`/`publishedBefore` values now log a warning instead
  of being silently ignored; hardened HTTP client setup against malformed
  `NO_PROXY` environment values.
- **0.5** — Added output and dataset schemas for the Store listing.

# Actor input Schema

## `country` (type: `string`):

Two-letter iTunes Store country code used for search (e.g. US, GB, DE, JP). Affects which shows Apple returns.

## `feedUrls` (type: `array`):

Direct podcast RSS feed URLs, one per line. Example: https://feeds.npr.org/510318/podcast.xml

## `opmlUrls` (type: `array`):

OPML subscription lists to bulk-import podcast feeds: paste OPML URLs or raw OPML XML, one per line. Great for monitoring many shows at once.

## `maxAudioMinutes` (type: `integer`):

Only the first N minutes of each episode's audio are sent to Whisper (also capped by the 25 MB upload limit).

## `maxDurationSeconds` (type: `integer`):

Skip episodes longer than this. 0 = no maximum.

## `maxEpisodesPerFeed` (type: `integer`):

How many of each feed's newest episodes to scrape.

## `maxResultsPerSearch` (type: `integer`):

Maximum podcast shows resolved per search term (iTunes limit: 200).

## `maxTranscribeEpisodes` (type: `integer`):

Safety cap on how many episodes Whisper transcribes per run (publisher transcripts are always free and unlimited).

## `minDurationSeconds` (type: `integer`):

Skip episodes shorter than this. Useful to exclude trailers and teasers. 0 = no minimum.

## `newEpisodesOnly` (type: `boolean`):

Output only episodes not seen in previous runs (state is kept in the run's key-value store). Ideal for scheduled monitoring of shows.

## `openaiApiKey` (type: `string`):

Your OpenAI API key, used only when Whisper transcription is enabled. Stored securely as a secret.

## `proxyConfiguration` (type: `object`):

Optional. RSS feeds and the iTunes API have no anti-bot protection, so proxies are not needed.

## `publishedAfter` (type: `string`):

Only episodes published on or after this date. Format: YYYY-MM-DD (e.g. 2025-01-01). Leave empty for no lower bound.

## `publishedBefore` (type: `string`):

Only episodes published on or before this date. Format: YYYY-MM-DD. Leave empty for no upper bound.

## `requestConcurrency` (type: `integer`):

How many feeds / transcripts to fetch in parallel. Pure HTTP, no browser or proxy needed.

## `searchTerms` (type: `array`):

Keywords to discover podcasts via the Apple Podcasts (iTunes) Search API. Each term resolves to matching show feeds.

## `titleContains` (type: `string`):

Case-insensitive keyword filter on episode titles. Separate alternatives with |, e.g. interview|founder. Leave empty to keep all episodes.

## `transcribe` (type: `boolean`):

Transcribe episode audio via the OpenAI Whisper API for episodes that have no free publisher transcript. Requires your own OpenAI API key; you pay OpenAI directly.

## `usePublisherTranscripts` (type: `boolean`):

Fetch free transcripts published by the show itself via the Podcasting 2.0 <podcast:transcript> tag. No API key and no transcription cost. Falls back to Whisper only when enabled below.

## `transcriptFormats` (type: `array`):

Which transcript formats to include per episode, one per line: text, srt, vtt. Defaults to text + srt. If a found transcript cannot provide a requested format, plain text is kept as fallback (transcriptStatus = format\_fallback).

## `transcriptLanguage` (type: `string`):

ISO-639-1 language code (e.g. en, es, de) passed to Whisper as a hint for faster, more accurate transcription. Leave empty for auto-detect.

## `whisperModel` (type: `string`):

OpenAI speech-to-text model used for transcription.

## Actor input object example

```json
{
  "country": "US",
  "feedUrls": [
    "https://feeds.npr.org/510318/podcast.xml"
  ],
  "opmlUrls": [],
  "maxAudioMinutes": 30,
  "maxDurationSeconds": 0,
  "maxEpisodesPerFeed": 20,
  "maxResultsPerSearch": 5,
  "maxTranscribeEpisodes": 3,
  "minDurationSeconds": 0,
  "newEpisodesOnly": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "publishedAfter": "",
  "publishedBefore": "",
  "requestConcurrency": 5,
  "searchTerms": [
    "artificial intelligence"
  ],
  "titleContains": "",
  "transcribe": false,
  "usePublisherTranscripts": true,
  "transcriptFormats": [
    "text",
    "srt"
  ],
  "whisperModel": "whisper-1"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "country": "US",
    "feedUrls": [
        "https://feeds.npr.org/510318/podcast.xml"
    ],
    "opmlUrls": [],
    "searchTerms": [
        "artificial intelligence"
    ],
    "transcriptFormats": [
        "text",
        "srt"
    ],
    "transcriptLanguage": "",
    "whisperModel": "whisper-1"
};

// Run the Actor and wait for it to finish
const run = await client.actor("celebrated-quadraphonic/podcast-rss-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "country": "US",
    "feedUrls": ["https://feeds.npr.org/510318/podcast.xml"],
    "opmlUrls": [],
    "searchTerms": ["artificial intelligence"],
    "transcriptFormats": [
        "text",
        "srt",
    ],
    "transcriptLanguage": "",
    "whisperModel": "whisper-1",
}

# Run the Actor and wait for it to finish
run = client.actor("celebrated-quadraphonic/podcast-rss-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "country": "US",
  "feedUrls": [
    "https://feeds.npr.org/510318/podcast.xml"
  ],
  "opmlUrls": [],
  "searchTerms": [
    "artificial intelligence"
  ],
  "transcriptFormats": [
    "text",
    "srt"
  ],
  "transcriptLanguage": "",
  "whisperModel": "whisper-1"
}' |
apify call celebrated-quadraphonic/podcast-rss-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,celebrated-quadraphonic/podcast-rss-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/FbUiVVjVBnvxFM4BL/builds/RTqAjChybF5aXSeoS/openapi.json
