# RSS Feed Reader – RSS, Atom & JSON Feed to JSON, Only New Items (`gazidev/rss-feed-reader`) Actor

Read any RSS 2.0, Atom, RDF or JSON Feed – or just a website URL, the feed is auto-discovered – and get clean JSON items: title, link, author, ISO dates, summary, categories, images, enclosures. OPML import, date filter, only-new-items mode for monitoring and optional full article text as Markdown.

- **URL**: https://apify.com/gazidev/rss-feed-reader.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.30 / 1,000 feed items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## RSS Feed Reader – RSS, Atom & JSON Feed to JSON

Read **any RSS 2.0, Atom, RSS 1.0 (RDF) or JSON Feed**, or just paste a **website URL** and the feed is found for you. Every item comes back in one clean, normalized schema: title, link, author, ISO 8601 dates, plain-text summary, categories, image and enclosures (podcast audio, video). The Actor imports **OPML** subscription lists, filters by date, returns **only new items** on scheduled runs and can add the **full article text as Markdown** for LLM summaries and RAG.

There is no browser, so it is fast and costs **$0.30 per 1,000 items**.

### What it does

- **Any feed format.** RSS 0.9x/2.0, RSS 1.0 (RDF), Atom 0.3/1.0 and JSON Feed 1.0/1.1, including podcast (iTunes) and Media RSS tags. Broken XML is parsed leniently instead of failing.
- **Feed auto-discovery.** Give `theverge.com` or `https://example.com/blog` and the Actor finds the feed:
  - `<link rel="alternate">` tags first (comment feeds are skipped),
  - then feed links on the page,
  - then common paths: `/feed`, `/rss`, `/feed.xml`, `/rss.xml`, `/atom.xml`, `/index.xml`, `/feed.json` and others.
- **OPML import.** Paste an OPML export from Feedly, Inoreader, NetNewsWire or Thunderbird, or give its URL. Folder names are kept as `opmlCategory`.
- **One normalized schema** for all formats, so you never handle `pubDate` vs `published` vs `date_published` again. Dates are UTC ISO 8601.
- **Date filter.** `publishedAfter` (`2026-09-01` or `24 hours`) or `maxAgeDays`.
- **Only-new mode** for monitoring. Seen items are remembered in a named key-value store. Feeds that did not change (HTTP 304) cost nothing.
- **Full article text (optional).** Downloads each article and adds the main content as clean Markdown (navigation, ads and footers removed), using the same extractor as our [Website to Markdown](https://apify.com/gazidev/website-to-markdown) Actor. It respects robots.txt and is charged only when the text was extracted.
- **Per-feed errors instead of crashes.** Dead feeds, 404s and sites without a feed become `feed-error` rows (free), and the run continues.
- Polite: honest User-Agent with a contact address, max 3 requests per host, retries with backoff on 429/5xx.

### Use cases

- **News and brand monitoring:** schedule hourly with `onlyNew` and send new items to Slack, email, Google Sheets or a webhook.
- **AI news digests:** feed titles + full-text Markdown into an LLM for daily summaries.
- **RAG / knowledge bases:** keep a vector database in sync with blogs and docs changelogs.
- **Content aggregation:** build a niche news site, newsletter or dashboard from dozens of sources via OPML.
- **Podcast data:** episode titles, dates and audio enclosure URLs (`enclosures[].url`, `type`, `length`).
- **Competitor tracking:** blog posts, release notes and press releases of competitors, auto-discovered from their homepages.

### Input example

```json
{
  "startUrls": [
    "https://xkcd.com/atom.xml",
    "https://www.jsonfeed.org/feed.json",
    "https://blog.cloudflare.com",
    "theverge.com"
  ],
  "opmlUrl": "https://example.com/my-subscriptions.opml",
  "maxItemsPerFeed": 20,
  "maxAgeDays": 7,
  "onlyNew": true,
  "fetchFullText": false
}
```

The default input reads 3 feeds (Atom, JSON Feed and a site URL with discovery) in about 2 seconds.

| Field | Default | Notes |
|---|---|---|
| `startUrls` | 3 samples | Feed URLs and/or website URLs (feed auto-discovered) |
| `opmlUrl` / `opmlText` | – | OPML file URL or pasted OPML (or a plain list of feed URLs) |
| `maxItemsPerFeed` | 100 | Newest first, 0 = all |
| `maxItems` | 0 | Total cap across feeds, 0 = no limit |
| `publishedAfter` | – | `2026-09-01` or relative `7 days` |
| `maxAgeDays` | 0 | Alternative date filter |
| `keepItemsWithoutDate` | false | When a date filter is set |
| `onlyNew` / `stateKey` | false / – | Monitoring memory |
| `fetchFullText` | false | Article text as Markdown, charged per success |
| `includeContentHtml` | false | Full item HTML from the feed (`content:encoded` etc.) |
| `includeAllDiscoveredFeeds` | false | Read every feed a site lists, not just the main one |
| `respectRobots` | true | For full-text fetches |

### Output example

One dataset row per item. The **Feed items**, **Full article text**, **Images & enclosures** and **Feed errors** views are in the Output tab.

```json
{
  "type": "item",
  "feedUrl": "https://www.theverge.com/rss/index.xml",
  "feedTitle": "The Verge",
  "feedLink": "https://www.theverge.com",
  "feedFormat": "Atom 1.0",
  "guid": "https://www.theverge.com/?p=1003877",
  "title": "Apple’s reportedly developing a smart home camera that doesn’t record video",
  "link": "https://www.theverge.com/tech/1003877/apple-security-camera-no-video",
  "author": "Stevie Bonifield",
  "published": "2026-10-01T22:51:36Z",
  "updated": "2026-10-01T22:51:36Z",
  "summary": "Apple's rumored push into smart home tech could include a smart home security camera that only gives …",
  "categories": ["Apple", "Apple Rumors", "Cameras", "Gadgets", "News", "Smart Home", "Tech"],
  "enclosures": [],
  "image": "https://platform.theverge.com/wp-content/uploads/sites/2/2025/03/STK071_APPLE_I.jpg?quality=90&strip=all&crop=0,0,100,100",
  "commentsUrl": null,
  "language": "en-US",
  "fetchedAt": "2026-10-02T09:14:28Z"
}
```

- With `fetchFullText`: `fullTextStatus` (`ok` / `failed` / `skipped` / `needsBrowser` / `empty`), `fullTextMarkdown`, `fullTextWordCount`, `fullTextUrl`, `fullTextError`. Only `ok` is charged.
- With `includeContentHtml`: `contentHtml` (sanitized).
- From OPML: `opmlCategory` (folder path, e.g. `Tech / Python`).
- Errors: rows with `"type": "feed-error"`, `input`, `feedUrl`, `httpStatus`, `error`. Not charged.
- The key-value store record `OUTPUT` holds the run summary: per feed the discovered feed URL, format, items in feed / saved / filtered / already seen, and warnings.

### Pricing

Pay per event. You pay only for items saved.

| Event | Price |
|---|---|
| Feed item | **$0.0003** ($0.30 / 1,000) |
| Full article text extracted (optional) | $0.0005 ($0.50 / 1,000) |
| Actor start | $0.00005 |

Compared with other Store Actors (public Store prices, 2026-10-01):

| Actor | Price per 1,000 items |
|---|---|
| **RSS Feed Reader (this Actor)** | **$0.30** |
| automation-lab/rss-feed-reader | $1.15 |
| santamaria-automations | $2.00 |
| technicaldost | $8.00 |

Set **Maximum cost per run** on the run options to cap spending. The Actor stops cleanly when the limit is reached, and in only-new mode the items it could not save are returned on the next run.

### FAQ

**I only have a website, not a feed URL.**
Paste the website or blog URL. The feed is discovered from `<link rel="alternate">`, page links or common paths. The `OUTPUT` record shows which feed was found and how (`discoveredVia`). Set `includeAllDiscoveredFeeds` to read every feed a site lists.

**How does only-new mode work?**
Item IDs (guid, else link) are stored per feed in a named key-value store derived from your input (`rss-feed-reader-…`), or from `stateKey` if you set one. The first run returns the current items; later runs return only items that appeared since. Items skipped by `maxItemsPerFeed` or the date filter in a run are also marked as seen, so a monitor never "back-fills" old posts. ETag / Last-Modified are sent, so unchanged feeds return HTTP 304 and cost nothing. Only-new mode is cheapest on feeds that support ETag or Last-Modified (HTTP 304); feeds without them are downloaded again on every run.

**Why are Google News feeds refused?**
`news.google.com` feeds contain Google redirect links rather than publisher URLs, and Google's terms do not allow this kind of automated reuse. Add the publishers' own feeds instead; you can paste their homepages.

**Why are Reddit feeds refused?**
Reddit's Data API terms require Reddit's approval for commercial use of its content, so `reddit.com`, `old.reddit.com`, `redd.it` and any `*.reddit.com` feed is refused with a `feed-error` row and not charged.

**Does the full-text option work on every site?**
It works on server-rendered news sites, blogs and docs. JavaScript-only pages and paywalls return `needsBrowser` / `empty` and are not charged. Pages disallowed by robots.txt are skipped.

**Very large podcast feeds?**
Feeds of over 1 MB with hundreds of episodes are cut after the newest `3 × maxItemsPerFeed` (min 100) items before parsing, which keeps runs fast and cheap. Set `maxItemsPerFeed: 0` to read the whole archive.

**Is it legal?**
RSS, Atom and JSON feeds are published for syndication. The Actor fetches only the feeds you supply, identifies itself and respects robots.txt for article pages. You are responsible for how you reuse the content (copyright, the site's terms).

### Use with AI agents / Apify MCP

- **Apify MCP server:** add `gazidev/rss-feed-reader` to your MCP client (Claude Desktop, Cursor, VS Code) via `https://mcp.apify.com?actors=gazidev/rss-feed-reader`. An agent can call it with `{"startUrls":["techcrunch.com"],"maxAgeDays":1}` and get today's posts as JSON.
- **API:** `POST https://api.apify.com/v2/acts/gazidev~rss-feed-reader/run-sync-get-dataset-items?token=...` with the input JSON returns the items directly. This works well as a "news tool" for LangChain or LlamaIndex agents.
- **Scheduled monitoring:** create a Task with `onlyNew: true`, schedule it, and connect the Slack, Google Sheets or webhook integration.

### More from the website toolkit

- [Website to Markdown](https://apify.com/gazidev/website-to-markdown): any URL, sitemap or whole site to clean Markdown with RAG chunks and llms.txt.
- [Sitemap Extractor](https://apify.com/gazidev/sitemap-url-extractor): every URL from sitemap.xml, with an HTTP status check and only-new mode.
- [SEO Audit Crawler](https://apify.com/gazidev/seo-audit-crawler): broken links, meta tags and redirects.
- [Tech Stack Detector](https://apify.com/gazidev/tech-stack-detector): bulk Wappalyzer alternative.
- [Domain Checker](https://apify.com/gazidev/domain-intel): WHOIS/RDAP, DNS, SSL and email security in bulk.
- [Website Contact Finder](https://apify.com/gazidev/website-contact-finder): emails, phones and social links from websites.

# Actor input Schema

## `startUrls` (type: `array`):

Any mix of RSS 2.0, Atom, RSS 1.0 (RDF) or JSON Feed URLs, and plain website or blog URLs (`theverge.com`, `https://example.com/blog`). For a website the feed is found automatically: `<link rel="alternate">` tags first, then feed links on the page, then common paths such as `/feed`, `/rss.xml`, `/atom.xml` and `/index.xml`. Google News feeds (news.google.com) are not supported.

## `opmlUrl` (type: `string`):

Optional. URL of an OPML subscription list exported from Feedly, Inoreader, NetNewsWire, Thunderbird and others. Every `<outline xmlUrl=…>` is read; folder names are saved as `opmlCategory`.

## `opmlText` (type: `string`):

Optional. Paste the content of an `.opml` file here instead of a URL (or a plain list of feed URLs, one per line).

## `maxItemsPerFeed` (type: `integer`):

Newest items first. 0 = every item in the feed. Most feeds hold the latest 10–50 items.

## `maxItems` (type: `integer`):

Stop after saving this many items across all feeds (0 = no limit). You pay only for items saved.

## `publishedAfter` (type: `string`):

Keep only items published on or after this date. Absolute (`2026-09-01`) or relative (`24 hours`, `7 days`).

## `maxAgeDays` (type: `integer`):

Alternative to the date above: keep only items from the last N days (0 = off). If both are set, the later date wins.

## `keepItemsWithoutDate` (type: `boolean`):

When a date filter is set, also keep items whose feed entry has no publication date (dropped by default).

## `onlyNew` (type: `boolean`):

Remembers the items already seen in each feed (in a named key-value store derived from this input) and outputs only items that appeared since the previous run. Schedule the Actor hourly or daily for alerts, Slack/email digests or a news pipeline. The first run returns the current items. Unchanged feeds (HTTP 304) cost nothing.

## `stateKey` (type: `string`):

Optional. By default the memory is tied to this exact list of feeds. Set a name to keep the same memory when you edit the list, or to run separate memories for the same feeds.

## `fetchFullText` (type: `boolean`):

Download each item's article page and add `fullTextMarkdown` (main content only: nav, ads, footers removed) and `fullTextWordCount`. Useful when the feed has only a short summary, e.g. for LLM summaries or RAG. Respects robots.txt. Charged as a separate event only when the text was extracted.

## `includeContentHtml` (type: `boolean`):

Add `contentHtml` with the item's full HTML from the feed (`content:encoded`, Atom `content`, JSON Feed `content_html`), sanitized. Off by default to keep rows small; the plain-text `summary` is always included.

## `includeAllDiscoveredFeeds` (type: `boolean`):

By default a website URL yields its main feed (comment feeds are skipped). Turn on to read every feed the site lists (e.g. per-category feeds, podcasts), up to 10.

## `respectRobots` (type: `boolean`):

Article pages disallowed by robots.txt are skipped (and not charged). Feeds themselves are published for syndication and are always read.

## `maxConcurrency` (type: `integer`):

How many feeds are read at the same time. At most 3 requests per host run in parallel.

## `requestTimeoutSecs` (type: `integer`):

Per-request timeout. Failed requests are retried twice with backoff.

## `proxyConfiguration` (type: `object`):

Optional. Feeds are rarely blocked; use a proxy only for sites that block datacenter IPs.

## Actor input object example

```json
{
  "startUrls": [
    "https://xkcd.com/atom.xml",
    "https://www.jsonfeed.org/feed.json",
    "https://blog.cloudflare.com"
  ],
  "maxItemsPerFeed": 10,
  "maxItems": 0,
  "maxAgeDays": 0,
  "keepItemsWithoutDate": false,
  "onlyNew": false,
  "fetchFullText": false,
  "includeContentHtml": false,
  "includeAllDiscoveredFeeds": false,
  "respectRobots": true,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

No description

## `fullText` (type: `string`):

No description

## `media` (type: `string`):

No description

## `errors` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://xkcd.com/atom.xml",
        "https://www.jsonfeed.org/feed.json",
        "https://blog.cloudflare.com"
    ],
    "maxItemsPerFeed": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/rss-feed-reader").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        "https://xkcd.com/atom.xml",
        "https://www.jsonfeed.org/feed.json",
        "https://blog.cloudflare.com",
    ],
    "maxItemsPerFeed": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("gazidev/rss-feed-reader").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://xkcd.com/atom.xml",
    "https://www.jsonfeed.org/feed.json",
    "https://blog.cloudflare.com"
  ],
  "maxItemsPerFeed": 10
}' |
apify call gazidev/rss-feed-reader --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/rss-feed-reader"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/mr47sVACevaGdImif/builds/JSGOSMEaXJhgCJGM9/openapi.json
