RSS Feed Reader Pro: multi-feed, dedupe, full text avatar

RSS Feed Reader Pro: multi-feed, dedupe, full text

Pricing

from $0.75 / 1,000 feed items

Go to Apify Store
RSS Feed Reader Pro: multi-feed, dedupe, full text

RSS Feed Reader Pro: multi-feed, dedupe, full text

Read RSS 2.0, Atom, RSS 1.0 and JSON feeds into one clean dataset. Feed autodiscovery, cross-run dedupe, keyword filters, optional full-text extraction.

Pricing

from $0.75 / 1,000 feed items

Rating

0.0

(0)

Developer

Spongy Frame Tools

Spongy Frame Tools

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

19 hours ago

Last modified

Categories

Share

What does RSS Feed Reader Pro do?

RSS Feed Reader Pro reads any number of RSS 2.0, RSS 1.0 (RDF), Atom and JSON Feed 1.1 feeds and turns them into one clean, consistent dataset. Every item has the same snake_case fields regardless of the source format: a stable id, the link and a tracking-free canonical_link, published and updated as ISO 8601 UTC timestamps, a plain-text summary, content_text, optional content_html, categories, enclosures, an image_url, and the feed it came from.

You can also give it a normal web page URL. The Actor looks for the <link rel="alternate"> tags pages use to advertise their feed and reads the first one it can parse.

Two features make it useful for monitoring rather than one-off reads:

  • Cross-run de-duplication. With onlyNewItems on, the Actor remembers which item ids it has already returned for each feed (in a named key-value store, up to 5,000 ids per feed) and only returns new ones next time. Run it on a schedule and you get a stream of new articles, not the same 20 posts every hour.
  • Optional full-text extraction. With fetchFullText on, the Actor opens each item's link and extracts the main article body using a built-in readability-style heuristic (paragraph density scoring, navigation/footer/sidebar/comment removal). No external services are involved.

Why use this one?

Most feed readers either parse one format cleanly or crawl pages. This Actor is built for the messy middle: a list of mixed feeds and a need for reliable, deduplicated output.

  • It handles what real feeds ship: CDATA, HTML entities, escaped HTML inside description, relative links, content:encoded, media:thumbnail, Atom entries with several <link> elements and hreflang variants, JSON Feed attachments.
  • Malformed XML is repaired where possible (unclosed tags in descriptions, bare ampersands, control characters) instead of failing the feed.
  • Encoding is detected from the BOM, the XML declaration or the HTTP Content-Type, so Windows-1252 feeds do not come out as mojibake.
  • One broken feed never fails the run. Failures are recorded per feed and the rest continue.
  • The same article in several feeds is collapsed to one record (matched by link with utm_* parameters and #fragment removed).
  • Author fields hold display names only. E-mail addresses in <author> are never included.

What it does not do: render JavaScript, log in to paywalled sites, or guarantee perfect extraction on every layout. Full-text extraction is a heuristic and works best on article-style pages.

How to use it

  1. Add feed URLs. Paste one or more feed URLs (or page URLs) into feedUrls.
  2. Choose your filters. Set maxItems, keep onlyNewItems on for scheduled runs, and optionally add keyword filters or a publishedAfter date. Turn on fetchFullText if you need article bodies.
  3. Run and export. Download the dataset as JSON, CSV or Excel, or connect it through the API or an integration. Schedule the run to keep it topped up with new items.

Input

{
"feedUrls": ["https://blog.apify.com/rss/", "https://hnrss.org/frontpage"],
"maxItems": 200,
"maxItemsPerFeed": 100,
"onlyNewItems": true,
"publishedAfter": "2026-09-01T00:00:00Z",
"fetchFullText": false,
"includeContentHtml": true,
"dedupeAcrossFeeds": true,
"keywordsInclude": ["python", "scraping"],
"keywordsExclude": ["sponsored"],
"sortBy": "published_desc"
}

All fields except feedUrls are optional. Keyword filters are case-insensitive substring matches on the title and summary. Items filtered out, already seen, or dropped as duplicates are never charged.

Output

One record per feed item:

{
"id": "post-1001",
"title": "Shipping faster with & without CI",
"link": "https://example.com/blog/shipping-faster?utm_source=rss",
"canonical_link": "https://example.com/blog/shipping-faster",
"published": "2024-09-10T06:30:00Z",
"updated": null,
"author_name": "Jane Doe",
"summary": "We cut build times by 40%. Here's how & why.",
"content_html": "<p>We cut build times by <strong>40%</strong>. Here's how &amp; why.</p>",
"content_text": "We cut build times by 40%. Here's how & why.",
"categories": ["Engineering", "CI/CD"],
"enclosures": [{"url": "https://example.com/podcast/episode-1.mp3", "type": "audio/mpeg", "length": 12345678}],
"image_url": "https://cdn.example.com/thumb-1001.jpg",
"feed_url": "https://example.com/blog/rss.xml",
"feed_title": "Example Engineering Blog",
"feed_language": "en-gb",
"full_text": null,
"full_text_word_count": null,
"source_url": "https://example.com/blog/rss.xml",
"fetched_at": "2026-09-24T10:15:00Z"
}

summary is plain text capped at 1,000 characters. full_text and full_text_word_count are filled only when fetchFullText is on and extraction succeeded. A run summary with per-feed failures is saved as SUMMARY in the run's key-value store.

How much does it cost?

This Actor uses pay-per-event pricing. You pay only for items that are actually written to the dataset.

EventPriceWhen it is charged
feed-item$0.001Once per item pushed to the dataset
full-text-extraction$0.003Once per item where full text was extracted (at least 50 words), on top of feed-item

Duplicates, filtered items and failed feeds cost nothing. Platform compute for a typical run is a fraction of a cent.

Example 1: hourly monitoring of 10 feeds. Roughly 30 new items per run with onlyNewItems on. 30 × $0.001 = $0.03 per run, about $0.72 per day or $22 per month at 24 runs a day.

Example 2: 1,000 articles with full text. 1,000 × $0.001 = $1.00 for the items, plus about 900 successful extractions × $0.003 = $2.70. Total around $3.70.

Limits and fair use

  • Feed responses and article pages are capped at 2 MB each. Each full-text fetch has a 20-second budget; slower pages are skipped (the item is still returned without full_text).
  • Requests carry an identifying User-Agent, are retried up to 3 times with exponential backoff, and honour Retry-After on 429 and 503. Please keep schedules reasonable; polling more often than every 15 minutes rarely yields anything new.
  • The seen-item memory holds the 5,000 most recent ids per feed. Extremely high-volume feeds may re-surface very old items after that window.
  • Auto-discovery tries up to 5 advertised feeds per page and uses the first that parses.

The Actor reads public syndication feeds that publishers provide for exactly this purpose and, optionally, the public pages they link to. It does not log in, bypass paywalls or collect personal data: author e-mail addresses are stripped and only display names are kept. You are responsible for complying with each publisher's terms and applicable law when you reuse the content.

FAQ

Why did my run stop early? Usually because it hit maxItems or the spending limit you set for the run; the Actor stops cleanly as soon as the platform reports the charge limit. The final log line shows items produced, items charged, failures and elapsed time.

Why did a feed return zero items? With onlyNewItems on, a feed that has not published anything since the last run correctly yields nothing. Set onlyNewItems to false to re-read everything, or look at SUMMARY in the key-value store for a per-feed error.

Which feed formats are supported? RSS 2.0 (and 0.9x), RSS 1.0 / RDF, Atom 1.0 and JSON Feed 1.1. Podcasts and media feeds work too; enclosures are listed with URL, MIME type and length.

How is the item id chosen? The feed's guid or Atom/JSON id when present, otherwise the item link, otherwise a SHA-1 of the title and published date. The same rule drives cross-run de-duplication.

Can I reset the de-duplication memory? Yes. Delete the named key-value store rss-feed-reader-pro-state in your Apify Console (or just the key for one feed, which is the SHA-1 of the feed URL), and the next run starts from scratch.

Support

Found a feed that does not parse, or a page where full-text extraction picks the wrong block? Open a ticket on the Issues tab of this Actor with the URL and we will respond within 24 hours.