# RSS & Atom Feed Reader + New Item Monitor (`moha-tah/rss-feed-reader`) Actor

Read any RSS, Atom or JSON Feed as clean JSON (title, link, date, summary, categories, podcast enclosures), auto-discover feeds from a website URL, or monitor feeds and get only new items on each scheduled run. Fast HTTP-only, pay per item.

- **URL**: https://apify.com/moha-tah/rss-feed-reader.md
- **Developed by:** [Mohamed T.](https://apify.com/moha-tah) (community)
- **Categories:** News, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## RSS & Atom Feed Reader + New Item Monitor

Turn any RSS, Atom or JSON Feed into clean JSON rows, or schedule it and get **only the new items** since the last run. Paste a feed URL or just a website: the feed is found for you.

### What does RSS & Atom Feed Reader do?

- Reads **RSS 0.9x/1.0/2.0, Atom and JSON Feed** (podcasts included) with one parser that tolerates broken XML.
- **Auto-discovers feeds** from a website or blog URL (`<link rel="alternate">` tags, then `/feed`, `/rss.xml`, `/atom.xml`, `/index.xml`…).
- Outputs flat, typed fields: title, link, ISO-8601 dates in UTC, plain-text summary, categories, author name, image and enclosures (podcast audio).
- **Monitor mode** remembers what it has already seen per feed and returns only new items. Schedule it hourly or daily and pipe the output to Slack, Sheets, a newsletter or an AI agent.
- Filters by keyword/regex and publication date, so you only pay for items you want.

### Who is it for?

- **Marketers and founders** tracking competitors' blogs, changelogs and press rooms.
- **Newsletter and content curators** aggregating dozens of sources into one dataset.
- **AI / RAG builders** who need fresh articles (with full HTML via `includeContent`) from many publishers.
- **n8n, Make and Zapier users** who want one reliable node for many feeds instead of one trigger per feed.
- **Podcast tools** that need episode lists with audio enclosures.

### How to use it

1. Add feed URLs or websites to **Feed URLs or websites**.
2. Pick **Read** (all current items) or **Monitor** (only new items).
3. Optional: keywords, "Published after" (`24 hours`, `7 days`, `2026-01-31`), max items per feed.
4. Run it, or create a schedule (Monitor mode) and connect an integration.

### Input example

```json
{
  "startUrls": ["https://github.blog", "https://feeds.npr.org/1001/rss.xml"],
  "mode": "monitor",
  "monitorName": "competitors",
  "includeKeywords": ["pricing", "launch"],
  "publishedAfter": "7 days"
}
```

### Output example

```json
{
  "id": "https://github.blog/?p=99027",
  "title": "GitHub Copilot app for Beginners: How to build custom workflows with canvases",
  "link": "https://github.blog/ai-and-ml/github-copilot/github-copilot-app-for-beginners-how-to-build-custom-workflows-with-canvases/",
  "published": "2026-09-25T18:00:00+00:00",
  "updated": null,
  "author": "Kayla Cinnamon",
  "summary": "Describe the interface you need in plain English, then let the agent build a live surface…",
  "categories": ["AI & ML", "GitHub Copilot"],
  "enclosures": [],
  "image": null,
  "feedTitle": "The GitHub Blog",
  "feedUrl": "https://github.blog/feed/",
  "site": "https://github.blog",
  "changeType": "new",
  "detectedAt": "2026-09-27T12:00:00+00:00"
}
```

`changeType` and `detectedAt` appear in Monitor mode only. A per-feed summary (format, item count, errors, discovery method) is saved in the key-value store record `SUMMARY`.

#### Data fields

| Field | Description |
|---|---|
| `id` | Stable item id (guid / Atom id / link) used for new-item detection |
| `title`, `link` | Item title (HTML stripped) and absolute URL |
| `published`, `updated` | ISO-8601 UTC timestamps, `null` when the feed has none |
| `author` | Byline name only. E-mail addresses are removed by design |
| `summary` | Plain text, cut at `summaryChars` |
| `contentHtml` | Full HTML body, only with `includeContent` |
| `categories`, `image`, `enclosures` | Tags, thumbnail, attachments (`url`, `type`, `length`) |
| `feedTitle`, `feedUrl`, `site` | Where the item came from |

### How much does it cost?

Pay per event, no subscription:

- **$0.001 per item** in Read mode (and on a Monitor baseline you ask to output).
- Monitor mode: **$0.0005 per feed check** plus **$0.001 per new item**. Quiet feeds cost almost nothing.

Example: 50 competitor blogs checked every day for a month ≈ 1,500 checks ($0.75) plus the new posts they publish (say 300 → $0.30).

### Tips

- Use **different monitor names** for independent watchlists; each name keeps its own memory.
- Feeds only show their latest items. Schedule Monitor mode at least as often as a feed publishes its page size (e.g. a news feed with 10 items and 30 posts a day needs hourly runs).
- Set **Feeds per website** above 1 to also read category or comment feeds a site advertises.

### Known limits

- HTTP only: pages that build their feed link with JavaScript, or feeds behind a login or a bot wall, can't be read.
- robots.txt is respected by default; disallowed feeds are reported in `SUMMARY`.
- Item history per feed and monitor name is capped at the 5,000 most recent ids.

### FAQ

**Can it find the feed if I only have the website address?** Yes. Most blogs and news sites (WordPress, Ghost, Substack and many more) advertise their feed; the Actor follows that link or tries common paths.

**Does it work for podcasts?** Yes. Audio files are in `enclosures` with MIME type and size.

**Will I get the same item twice in Monitor mode?** No. Item ids are stored after each run, before results are charged, so a stopped run never replays items.

**Is it legal?** Feeds are published by sites for automated reading. The Actor reads only public feeds, respects robots.txt, and outputs no e-mail addresses. You are responsible for how you reuse the content (copyright still applies to full articles).

**Can AI agents use it?** Yes, through Apify's MCP server and API; inputs work with only `startUrls`.

# Actor input Schema

## `startUrls` (type: `array`):

A feed URL (RSS, Atom or JSON Feed) or a website/blog URL. For a website, the feed is discovered from the page's <link rel="alternate"> tags and common paths like /feed or /rss.xml.

## `mode` (type: `string`):

Read = output the items currently in each feed. Monitor = output only items not seen in previous runs with the same monitor name (schedule it hourly or daily). The first monitor run stores a baseline.

## `monitorName` (type: `string`):

Monitor mode only. Runs with the same name share state. Use different names for independent watchlists.

## `includeKeywords` (type: `array`):

Keep items whose title, summary or categories match at least one entry (case-insensitive), e.g. pricing or \bAI\b.

## `excludeKeywords` (type: `array`):

Drop items whose title, summary or categories match any entry (case-insensitive).

## `publishedAfter` (type: `string`):

Keep only items published after this date. Accepts 2026-01-31 or a relative value like '24 hours' or '7 days'. Items without a date are dropped when set.

## `maxItemsPerFeed` (type: `integer`):

Stop after this many items per feed (0 = all items in the feed).

## `summaryChars` (type: `integer`):

Plain-text summary is cut at this length (0 = full text).

## `includeContent` (type: `boolean`):

Adds the item's full HTML body (content:encoded / Atom content) as contentHtml. Useful for RAG and newsletters; makes rows larger.

## `emitBaseline` (type: `boolean`):

Monitor mode: also output the current items (changeType 'baseline') on the first run.

## `maxFeedsPerSite` (type: `integer`):

When a website advertises several feeds (posts, comments, categories), read up to this many.

## `respectRobotsTxt` (type: `boolean`):

Skip URLs disallowed by the site's robots.txt.

## `maxConcurrency` (type: `integer`):

Feeds processed at the same time. Requests to the same host are spaced out politely.

## Actor input object example

```json
{
  "startUrls": [
    "https://blog.apify.com",
    "https://github.blog/feed/"
  ],
  "mode": "read",
  "monitorName": "default",
  "includeKeywords": [],
  "excludeKeywords": [],
  "maxItemsPerFeed": 100,
  "summaryChars": 500,
  "includeContent": false,
  "emitBaseline": false,
  "maxFeedsPerSite": 1,
  "respectRobotsTxt": true,
  "maxConcurrency": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://blog.apify.com",
        "https://github.blog/feed/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("moha-tah/rss-feed-reader").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [
        "https://blog.apify.com",
        "https://github.blog/feed/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("moha-tah/rss-feed-reader").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://blog.apify.com",
    "https://github.blog/feed/"
  ]
}' |
apify call moha-tah/rss-feed-reader --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,moha-tah/rss-feed-reader"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/shEr8oWUNjEoCOVRQ/builds/0yNVxCvBILSrSJuAJ/openapi.json
