# RSS Feed Reader Pro: multi-feed, dedupe, full text (`spongy_frame/rss-feed-reader-pro`) Actor

Read RSS 2.0, Atom, RSS 1.0 and JSON feeds into one clean dataset. Feed autodiscovery, cross-run dedupe, keyword filters, optional full-text extraction.

- **URL**: https://apify.com/spongy\_frame/rss-feed-reader-pro.md
- **Developed by:** [Spongy Frame Tools](https://apify.com/spongy_frame) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.75 / 1,000 feed items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does RSS Feed Reader Pro do?

RSS Feed Reader Pro reads any number of RSS 2.0, RSS 1.0 (RDF), Atom and JSON Feed 1.1 feeds and turns them into one clean, consistent dataset. Every item has the same snake\_case fields regardless of the source format: a stable `id`, the `link` and a tracking-free `canonical_link`, `published` and `updated` as ISO 8601 UTC timestamps, a plain-text `summary`, `content_text`, optional `content_html`, `categories`, `enclosures`, an `image_url`, and the feed it came from.

You can also give it a normal web page URL. The Actor looks for the `<link rel="alternate">` tags pages use to advertise their feed and reads the first one it can parse.

Two features make it useful for monitoring rather than one-off reads:

- **Cross-run de-duplication.** With `onlyNewItems` on, the Actor remembers which item ids it has already returned for each feed (in a named key-value store, up to 5,000 ids per feed) and only returns new ones next time. Run it on a schedule and you get a stream of new articles, not the same 20 posts every hour.
- **Optional full-text extraction.** With `fetchFullText` on, the Actor opens each item's link and extracts the main article body using a built-in readability-style heuristic (paragraph density scoring, navigation/footer/sidebar/comment removal). No external services are involved.

### Why use this one?

Most feed readers either parse one format cleanly or crawl pages. This Actor is built for the messy middle: a list of mixed feeds and a need for reliable, deduplicated output.

- It handles what real feeds ship: CDATA, HTML entities, escaped HTML inside `description`, relative links, `content:encoded`, `media:thumbnail`, Atom entries with several `<link>` elements and `hreflang` variants, JSON Feed attachments.
- Malformed XML is repaired where possible (unclosed tags in descriptions, bare ampersands, control characters) instead of failing the feed.
- Encoding is detected from the BOM, the XML declaration or the HTTP `Content-Type`, so Windows-1252 feeds do not come out as mojibake.
- One broken feed never fails the run. Failures are recorded per feed and the rest continue.
- The same article in several feeds is collapsed to one record (matched by link with `utm_*` parameters and `#fragment` removed).
- Author fields hold display names only. E-mail addresses in `<author>` are never included.

What it does not do: render JavaScript, log in to paywalled sites, or guarantee perfect extraction on every layout. Full-text extraction is a heuristic and works best on article-style pages.

### How to use it

1. **Add feed URLs.** Paste one or more feed URLs (or page URLs) into `feedUrls`.
2. **Choose your filters.** Set `maxItems`, keep `onlyNewItems` on for scheduled runs, and optionally add keyword filters or a `publishedAfter` date. Turn on `fetchFullText` if you need article bodies.
3. **Run and export.** Download the dataset as JSON, CSV or Excel, or connect it through the API or an integration. Schedule the run to keep it topped up with new items.

### Input

```json
{
  "feedUrls": ["https://blog.apify.com/rss/", "https://hnrss.org/frontpage"],
  "maxItems": 200,
  "maxItemsPerFeed": 100,
  "onlyNewItems": true,
  "publishedAfter": "2026-09-01T00:00:00Z",
  "fetchFullText": false,
  "includeContentHtml": true,
  "dedupeAcrossFeeds": true,
  "keywordsInclude": ["python", "scraping"],
  "keywordsExclude": ["sponsored"],
  "sortBy": "published_desc"
}
```

All fields except `feedUrls` are optional. Keyword filters are case-insensitive substring matches on the title and summary. Items filtered out, already seen, or dropped as duplicates are never charged.

### Output

One record per feed item:

```json
{
  "id": "post-1001",
  "title": "Shipping faster with & without CI",
  "link": "https://example.com/blog/shipping-faster?utm_source=rss",
  "canonical_link": "https://example.com/blog/shipping-faster",
  "published": "2024-09-10T06:30:00Z",
  "updated": null,
  "author_name": "Jane Doe",
  "summary": "We cut build times by 40%. Here's how & why.",
  "content_html": "<p>We cut build times by <strong>40%</strong>. Here's how &amp; why.</p>",
  "content_text": "We cut build times by 40%. Here's how & why.",
  "categories": ["Engineering", "CI/CD"],
  "enclosures": [{"url": "https://example.com/podcast/episode-1.mp3", "type": "audio/mpeg", "length": 12345678}],
  "image_url": "https://cdn.example.com/thumb-1001.jpg",
  "feed_url": "https://example.com/blog/rss.xml",
  "feed_title": "Example Engineering Blog",
  "feed_language": "en-gb",
  "full_text": null,
  "full_text_word_count": null,
  "source_url": "https://example.com/blog/rss.xml",
  "fetched_at": "2026-09-24T10:15:00Z"
}
```

`summary` is plain text capped at 1,000 characters. `full_text` and `full_text_word_count` are filled only when `fetchFullText` is on and extraction succeeded. A run summary with per-feed `failures` is saved as `SUMMARY` in the run's key-value store.

### How much does it cost?

This Actor uses pay-per-event pricing. You pay only for items that are actually written to the dataset.

| Event | Price | When it is charged |
| --- | --- | --- |
| `feed-item` | $0.001 | Once per item pushed to the dataset |
| `full-text-extraction` | $0.003 | Once per item where full text was extracted (at least 50 words), on top of `feed-item` |

Duplicates, filtered items and failed feeds cost nothing. Platform compute for a typical run is a fraction of a cent.

**Example 1: hourly monitoring of 10 feeds.** Roughly 30 new items per run with `onlyNewItems` on. 30 × $0.001 = $0.03 per run, about $0.72 per day or $22 per month at 24 runs a day.

**Example 2: 1,000 articles with full text.** 1,000 × $0.001 = $1.00 for the items, plus about 900 successful extractions × $0.003 = $2.70. Total around $3.70.

### Limits and fair use

- Feed responses and article pages are capped at 2 MB each. Each full-text fetch has a 20-second budget; slower pages are skipped (the item is still returned without `full_text`).
- Requests carry an identifying User-Agent, are retried up to 3 times with exponential backoff, and honour `Retry-After` on 429 and 503. Please keep schedules reasonable; polling more often than every 15 minutes rarely yields anything new.
- The seen-item memory holds the 5,000 most recent ids per feed. Extremely high-volume feeds may re-surface very old items after that window.
- Auto-discovery tries up to 5 advertised feeds per page and uses the first that parses.

### Legal note

The Actor reads public syndication feeds that publishers provide for exactly this purpose and, optionally, the public pages they link to. It does not log in, bypass paywalls or collect personal data: author e-mail addresses are stripped and only display names are kept. You are responsible for complying with each publisher's terms and applicable law when you reuse the content.

### FAQ

**Why did my run stop early?** Usually because it hit `maxItems` or the spending limit you set for the run; the Actor stops cleanly as soon as the platform reports the charge limit. The final log line shows items produced, items charged, failures and elapsed time.

**Why did a feed return zero items?** With `onlyNewItems` on, a feed that has not published anything since the last run correctly yields nothing. Set `onlyNewItems` to false to re-read everything, or look at `SUMMARY` in the key-value store for a per-feed error.

**Which feed formats are supported?** RSS 2.0 (and 0.9x), RSS 1.0 / RDF, Atom 1.0 and JSON Feed 1.1. Podcasts and media feeds work too; enclosures are listed with URL, MIME type and length.

**How is the item `id` chosen?** The feed's `guid` or Atom/JSON `id` when present, otherwise the item link, otherwise a SHA-1 of the title and published date. The same rule drives cross-run de-duplication.

**Can I reset the de-duplication memory?** Yes. Delete the named key-value store `rss-feed-reader-pro-state` in your Apify Console (or just the key for one feed, which is the SHA-1 of the feed URL), and the next run starts from scratch.

### Support

Found a feed that does not parse, or a page where full-text extraction picks the wrong block? Open a ticket on the Issues tab of this Actor with the URL and we will respond within 24 hours.

# Actor input Schema

## `feedUrls` (type: `array`):

List of feed URLs (RSS 2.0, RSS 1.0/RDF, Atom or JSON Feed 1.1). A regular HTML page URL also works: the Actor looks for <link rel="alternate" type="application/rss+xml|atom+xml|feed+json"> and reads the first feed it can parse.

## `maxItems` (type: `integer`):

Stop after this many items have been pushed to the dataset across all feeds. This is also the upper bound on feed-item charges for the run.

## `maxItemsPerFeed` (type: `integer`):

Keep at most this many items from each individual feed (after filtering). Feeds are read in document order, which for nearly all feeds is newest first.

## `onlyNewItems` (type: `boolean`):

Remember which item ids were already returned (per feed, in the named key-value store 'rss-feed-reader-pro-state', up to 5,000 ids per feed) and skip them on later runs. Skipped items are never charged. Turn off to get the full feed every run.

## `sortBy` (type: `string`):

Order of items in the dataset before maxItems is applied.

## `fetchFullText` (type: `boolean`):

Open each item's link and extract the main article text (readability-style heuristic: no external services). Adds one HTTP request per item and a 'full-text-extraction' charge for every item where at least 50 words were extracted. Pages larger than 2 MB are truncated; each item has a 20-second budget.

## `includeContentHtml` (type: `boolean`):

Include the raw HTML of the item's content/description in the 'content\_html' field. Turn off to get smaller records (plain-text 'content\_text' and 'summary' are always included).

## `publishedAfter` (type: `string`):

Only keep items published on or after this date/time (ISO 8601, e.g. 2026-09-01 or 2026-09-01T00:00:00Z). Items without a parseable date are kept.

## `keywordsInclude` (type: `array`):

Keep only items whose title or summary contains at least one of these words or phrases (case-insensitive substring match).

## `keywordsExclude` (type: `array`):

Drop items whose title or summary contains any of these words or phrases (case-insensitive substring match). Applied after the include list.

## `dedupeAcrossFeeds` (type: `boolean`):

When the same article appears in several feeds (matched by its link with utm\_\* parameters and #fragment removed), keep only the first occurrence.

## Actor input object example

```json
{
  "feedUrls": [
    "https://blog.apify.com/rss/",
    "https://hnrss.org/frontpage"
  ],
  "maxItems": 200,
  "maxItemsPerFeed": 100,
  "onlyNewItems": true,
  "sortBy": "published_desc",
  "fetchFullText": false,
  "includeContentHtml": true,
  "keywordsInclude": [],
  "keywordsExclude": [],
  "dedupeAcrossFeeds": true
}
```

# Actor output Schema

## `dataset` (type: `string`):

One record per feed item (title, link, published, summary, content, categories, enclosures, feed\_url, optional full\_text).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "feedUrls": [
        "https://blog.apify.com/rss/",
        "https://hnrss.org/frontpage"
    ],
    "maxItems": 200,
    "onlyNewItems": false
};

// Run the Actor and wait for it to finish
const run = await client.actor("spongy_frame/rss-feed-reader-pro").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "feedUrls": [
        "https://blog.apify.com/rss/",
        "https://hnrss.org/frontpage",
    ],
    "maxItems": 200,
    "onlyNewItems": False,
}

# Run the Actor and wait for it to finish
run = client.actor("spongy_frame/rss-feed-reader-pro").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "feedUrls": [
    "https://blog.apify.com/rss/",
    "https://hnrss.org/frontpage"
  ],
  "maxItems": 200,
  "onlyNewItems": false
}' |
apify call spongy_frame/rss-feed-reader-pro --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,spongy_frame/rss-feed-reader-pro"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KTWWFCVOcOhiPzuSV/builds/hHIiutC9FuwlRt8VQ/openapi.json
