RSS Feed Reader – RSS, Atom & JSON Feed to JSON, Only New Items avatar

RSS Feed Reader – RSS, Atom & JSON Feed to JSON, Only New Items

Pricing

from $0.30 / 1,000 feed items

Go to Apify Store
RSS Feed Reader – RSS, Atom & JSON Feed to JSON, Only New Items

RSS Feed Reader – RSS, Atom & JSON Feed to JSON, Only New Items

Read any RSS 2.0, Atom, RDF or JSON Feed – or just a website URL, the feed is auto-discovered – and get clean JSON items: title, link, author, ISO dates, summary, categories, images, enclosures. OPML import, date filter, only-new-items mode for monitoring and optional full article text as Markdown.

Pricing

from $0.30 / 1,000 feed items

Rating

0.0

(0)

Developer

Cemal Atakli

Cemal Atakli

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Categories

Share

RSS Feed Reader – RSS, Atom & JSON Feed to JSON

Read any RSS 2.0, Atom, RSS 1.0 (RDF) or JSON Feed, or just paste a website URL and the feed is found for you. Every item comes back in one clean, normalized schema: title, link, author, ISO 8601 dates, plain-text summary, categories, image and enclosures (podcast audio, video). The Actor imports OPML subscription lists, filters by date, returns only new items on scheduled runs and can add the full article text as Markdown for LLM summaries and RAG.

There is no browser, so it is fast and costs $0.30 per 1,000 items.

What it does

  • Any feed format. RSS 0.9x/2.0, RSS 1.0 (RDF), Atom 0.3/1.0 and JSON Feed 1.0/1.1, including podcast (iTunes) and Media RSS tags. Broken XML is parsed leniently instead of failing.
  • Feed auto-discovery. Give theverge.com or https://example.com/blog and the Actor finds the feed:
    • <link rel="alternate"> tags first (comment feeds are skipped),
    • then feed links on the page,
    • then common paths: /feed, /rss, /feed.xml, /rss.xml, /atom.xml, /index.xml, /feed.json and others.
  • OPML import. Paste an OPML export from Feedly, Inoreader, NetNewsWire or Thunderbird, or give its URL. Folder names are kept as opmlCategory.
  • One normalized schema for all formats, so you never handle pubDate vs published vs date_published again. Dates are UTC ISO 8601.
  • Date filter. publishedAfter (2026-09-01 or 24 hours) or maxAgeDays.
  • Only-new mode for monitoring. Seen items are remembered in a named key-value store. Feeds that did not change (HTTP 304) cost nothing.
  • Full article text (optional). Downloads each article and adds the main content as clean Markdown (navigation, ads and footers removed), using the same extractor as our Website to Markdown Actor. It respects robots.txt and is charged only when the text was extracted.
  • Per-feed errors instead of crashes. Dead feeds, 404s and sites without a feed become feed-error rows (free), and the run continues.
  • Polite: honest User-Agent with a contact address, max 3 requests per host, retries with backoff on 429/5xx.

Use cases

  • News and brand monitoring: schedule hourly with onlyNew and send new items to Slack, email, Google Sheets or a webhook.
  • AI news digests: feed titles + full-text Markdown into an LLM for daily summaries.
  • RAG / knowledge bases: keep a vector database in sync with blogs and docs changelogs.
  • Content aggregation: build a niche news site, newsletter or dashboard from dozens of sources via OPML.
  • Podcast data: episode titles, dates and audio enclosure URLs (enclosures[].url, type, length).
  • Competitor tracking: blog posts, release notes and press releases of competitors, auto-discovered from their homepages.

Input example

{
"startUrls": [
"https://xkcd.com/atom.xml",
"https://www.jsonfeed.org/feed.json",
"https://blog.cloudflare.com",
"theverge.com"
],
"opmlUrl": "https://example.com/my-subscriptions.opml",
"maxItemsPerFeed": 20,
"maxAgeDays": 7,
"onlyNew": true,
"fetchFullText": false
}

The default input reads 3 feeds (Atom, JSON Feed and a site URL with discovery) in about 2 seconds.

FieldDefaultNotes
startUrls3 samplesFeed URLs and/or website URLs (feed auto-discovered)
opmlUrl / opmlText–OPML file URL or pasted OPML (or a plain list of feed URLs)
maxItemsPerFeed100Newest first, 0 = all
maxItems0Total cap across feeds, 0 = no limit
publishedAfter–2026-09-01 or relative 7 days
maxAgeDays0Alternative date filter
keepItemsWithoutDatefalseWhen a date filter is set
onlyNew / stateKeyfalse / –Monitoring memory
fetchFullTextfalseArticle text as Markdown, charged per success
includeContentHtmlfalseFull item HTML from the feed (content:encoded etc.)
includeAllDiscoveredFeedsfalseRead every feed a site lists, not just the main one
respectRobotstrueFor full-text fetches

Output example

One dataset row per item. The Feed items, Full article text, Images & enclosures and Feed errors views are in the Output tab.

{
"type": "item",
"feedUrl": "https://www.theverge.com/rss/index.xml",
"feedTitle": "The Verge",
"feedLink": "https://www.theverge.com",
"feedFormat": "Atom 1.0",
"guid": "https://www.theverge.com/?p=1003877",
"title": "Apple’s reportedly developing a smart home camera that doesn’t record video",
"link": "https://www.theverge.com/tech/1003877/apple-security-camera-no-video",
"author": "Stevie Bonifield",
"published": "2026-10-01T22:51:36Z",
"updated": "2026-10-01T22:51:36Z",
"summary": "Apple's rumored push into smart home tech could include a smart home security camera that only gives …",
"categories": ["Apple", "Apple Rumors", "Cameras", "Gadgets", "News", "Smart Home", "Tech"],
"enclosures": [],
"image": "https://platform.theverge.com/wp-content/uploads/sites/2/2025/03/STK071_APPLE_I.jpg?quality=90&strip=all&crop=0,0,100,100",
"commentsUrl": null,
"language": "en-US",
"fetchedAt": "2026-10-02T09:14:28Z"
}
  • With fetchFullText: fullTextStatus (ok / failed / skipped / needsBrowser / empty), fullTextMarkdown, fullTextWordCount, fullTextUrl, fullTextError. Only ok is charged.
  • With includeContentHtml: contentHtml (sanitized).
  • From OPML: opmlCategory (folder path, e.g. Tech / Python).
  • Errors: rows with "type": "feed-error", input, feedUrl, httpStatus, error. Not charged.
  • The key-value store record OUTPUT holds the run summary: per feed the discovered feed URL, format, items in feed / saved / filtered / already seen, and warnings.

Pricing

Pay per event. You pay only for items saved.

EventPrice
Feed item$0.0003 ($0.30 / 1,000)
Full article text extracted (optional)$0.0005 ($0.50 / 1,000)
Actor start$0.00005

Compared with other Store Actors (public Store prices, 2026-10-01):

ActorPrice per 1,000 items
RSS Feed Reader (this Actor)$0.30
automation-lab/rss-feed-reader$1.15
santamaria-automations$2.00
technicaldost$8.00

Set Maximum cost per run on the run options to cap spending. The Actor stops cleanly when the limit is reached, and in only-new mode the items it could not save are returned on the next run.

FAQ

I only have a website, not a feed URL. Paste the website or blog URL. The feed is discovered from <link rel="alternate">, page links or common paths. The OUTPUT record shows which feed was found and how (discoveredVia). Set includeAllDiscoveredFeeds to read every feed a site lists.

How does only-new mode work? Item IDs (guid, else link) are stored per feed in a named key-value store derived from your input (rss-feed-reader-…), or from stateKey if you set one. The first run returns the current items; later runs return only items that appeared since. Items skipped by maxItemsPerFeed or the date filter in a run are also marked as seen, so a monitor never "back-fills" old posts. ETag / Last-Modified are sent, so unchanged feeds return HTTP 304 and cost nothing. Only-new mode is cheapest on feeds that support ETag or Last-Modified (HTTP 304); feeds without them are downloaded again on every run.

Why are Google News feeds refused? news.google.com feeds contain Google redirect links rather than publisher URLs, and Google's terms do not allow this kind of automated reuse. Add the publishers' own feeds instead; you can paste their homepages.

Why are Reddit feeds refused? Reddit's Data API terms require Reddit's approval for commercial use of its content, so reddit.com, old.reddit.com, redd.it and any *.reddit.com feed is refused with a feed-error row and not charged.

Does the full-text option work on every site? It works on server-rendered news sites, blogs and docs. JavaScript-only pages and paywalls return needsBrowser / empty and are not charged. Pages disallowed by robots.txt are skipped.

Very large podcast feeds? Feeds of over 1 MB with hundreds of episodes are cut after the newest 3 × maxItemsPerFeed (min 100) items before parsing, which keeps runs fast and cheap. Set maxItemsPerFeed: 0 to read the whole archive.

Is it legal? RSS, Atom and JSON feeds are published for syndication. The Actor fetches only the feeds you supply, identifies itself and respects robots.txt for article pages. You are responsible for how you reuse the content (copyright, the site's terms).

Use with AI agents / Apify MCP

  • Apify MCP server: add gazidev/rss-feed-reader to your MCP client (Claude Desktop, Cursor, VS Code) via https://mcp.apify.com?actors=gazidev/rss-feed-reader. An agent can call it with {"startUrls":["techcrunch.com"],"maxAgeDays":1} and get today's posts as JSON.
  • API: POST https://api.apify.com/v2/acts/gazidev~rss-feed-reader/run-sync-get-dataset-items?token=... with the input JSON returns the items directly. This works well as a "news tool" for LangChain or LlamaIndex agents.
  • Scheduled monitoring: create a Task with onlyNew: true, schedule it, and connect the Slack, Google Sheets or webhook integration.

More from the website toolkit