# RSS, Atom & Podcast Feed Scraper (`springlike_meadowland/rss-podcast-feed-scraper`) Actor

Turn public RSS 2.0, Atom 1.0 and podcast feeds into structured entries with dates, source text, categories and enclosure links. Add direct HTTPS feed URLs; inspect per-feed coverage in OUTPUT.

- **URL**: https://apify.com/springlike\_meadowland/rss-podcast-feed-scraper.md
- **Developed by:** [Akshay Aggarwal](https://apify.com/springlike_meadowland) (community)
- **Categories:** Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 feed entries

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## RSS, Atom & Podcast Feed Scraper

Turn public RSS 2.0 and Atom 1.0 feeds into a structured dataset for news monitoring, podcast research, and release tracking. Give it direct feed URLs; get one row per saved entry with dates, source text, categories, and podcast enclosure links when the feed provides them.

### Quick start

Paste this input in Apify Console:

```json
{
  "feedUrls": [
    "https://www.nasa.gov/news-release/feed/",
    "https://github.com/python/cpython/releases.atom"
  ],
  "maxItemsPerFeed": 2,
  "deduplicate": true
}
```

The default is 10 entries per feed, so a single feed is a small first run. `maxItemsPerFeed` accepts 1–500, and `feedUrls` accepts 1–50. `sinceDate` is optional: use an ISO 8601 date (`2026-09-01`, treated as UTC midnight) or timezone-aware datetime. When a date filter is set, entries with no usable publication or update date are skipped.

Here is an entry from the Python release Atom feed ([full example](examples/python-release-entry.json)):

```json
{
  "title": "v3.15.0rc2",
  "url": "https://github.com/python/cpython/releases/tag/v3.15.0rc2",
  "author": "hugovk",
  "updated_at": "2026-09-01T09:14:36Z",
  "content_if_feed_supplies_it": "<p>Python 3.15.0rc2</p>",
  "enclosures": []
}
```

Each full dataset row also has `feed_url`, source `id`/GUID, `published_at`, `summary`, `categories`, and `source_url` (the final feed URL after redirects). Missing source fields are `null`; the Actor does not make up dates or summaries. RSS `description` and Atom `summary` are copied from the feed. `content_if_feed_supplies_it` comes only from RSS `content:encoded` or Atom `content`. Podcast `<enclosure>` and Atom enclosure links become `{url,type,length}` objects; audio is not downloaded.

### Coverage and billing

The supported sources are public HTTPS RSS 2.0 and Atom 1.0 XML feeds, including podcast RSS feeds and public release feeds such as GitHub's Atom release feed. Enter the direct feed URL, not a website homepage or search query. This Actor does not search the web, open linked articles, extract paywalled full text, download audio, or transcribe it. It does not use AI. Feed entries with neither a usable entry URL nor an ID are skipped.

One **saved feed entry** is the pay-per-event billing unit (`feed-entry`) when pay-per-event pricing is enabled. Check the current price before running. Duplicate entries, skipped entries, errors, and the run summary are not result rows. The default dataset contains only saved entries. `OUTPUT` in the default key-value store records fetched, saved, skipped, error and limit counts by feed. If a feed fails or the run hits a charge limit, earlier rows may still be present; inspect `OUTPUT` before treating the dataset as complete. A run can have partial results because a source is malformed, unavailable, redirected unsafely, or larger than the 2 MiB per-feed response cap.

The Actor fetches at most 50 feeds, one request per feed, with a 90-second overall fetch budget and a 12-second timeout per request. It allows up to five redirects; every hop must be public HTTPS on port 443. Feeds are limited to 2 MiB each; XML DTDs and external entities are refused. Source timestamps are normalized to UTC when they contain a timezone; unparseable or timezone-free values become `null`. It does not crawl archive pages or guarantee all historical entries: the source's current feed window determines what is available.

### Fields

| Field | Meaning |
| --- | --- |
| `feed_url` | Input URL |
| `id` | RSS GUID or Atom ID, if supplied |
| `title`, `url`, `author` | Feed-supplied entry metadata |
| `published_at`, `updated_at` | Source dates normalized to UTC, or `null` |
| `summary` | Source description or summary; no AI summary |
| `content_if_feed_supplies_it` | Source feed content only, if present |
| `categories` | Source categories or terms |
| `enclosures` | Media link metadata, no media content |
| `source_url` | Final fetched feed URL after HTTPS redirects |

Deduplication uses the entry URL across feeds, or the GUID within a feed when no URL exists. A per-feed cap counts selected entries. For each feed, `OUTPUT` reports `fetched` and `saved` plus skip reasons (`date`, `duplicate`, `missing_identity`, `limit`).

# Actor input Schema

## `feedUrls` (type: `array`):

Add 1–50 direct RSS 2.0 or Atom 1.0 feed URLs. Start with one feed for a small sample.

## `maxItemsPerFeed` (type: `integer`):

Default 10 for a small sample; up to 500 per feed. Source feeds may contain fewer entries.

## `sinceDate` (type: `string`):

ISO 8601 date (UTC midnight) or timezone-aware datetime. Entries without a usable date are skipped when set.

## `deduplicate` (type: `boolean`):

Deduplicate by entry URL across feeds, or by GUID within a feed if no URL is supplied.

## Actor input object example

```json
{
  "feedUrls": [
    "https://www.nasa.gov/news-release/feed/"
  ],
  "maxItemsPerFeed": 10,
  "deduplicate": true
}
```

# Actor output Schema

## `results` (type: `string`):

Saved RSS, Atom and podcast entries.

## `summary` (type: `string`):

Fetched, saved, skipped and error counts by feed, including partial results.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "feedUrls": [
        "https://www.nasa.gov/news-release/feed/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("springlike_meadowland/rss-podcast-feed-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "feedUrls": ["https://www.nasa.gov/news-release/feed/"] }

# Run the Actor and wait for it to finish
run = client.actor("springlike_meadowland/rss-podcast-feed-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "feedUrls": [
    "https://www.nasa.gov/news-release/feed/"
  ]
}' |
apify call springlike_meadowland/rss-podcast-feed-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,springlike_meadowland/rss-podcast-feed-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/eSSyjQeD7ZKgfUGrF/builds/B8GmoTMutaATg1a1A/openapi.json
