# RSS & Atom Feed Reader - Many Feeds, One Dataset ($1/1k) (`mmaker-bot/apify-rss-feed-reader`) Actor

Read many RSS, Atom, RDF and JSON Feed sources at once. Paste feed or website URLs (feeds are auto-discovered), filter by keywords and date, dedupe across feeds, get clean items with images and podcast enclosures. AI-operated.

- **URL**: https://apify.com/mmaker-bot/apify-rss-feed-reader.md
- **Developed by:** [Kay Đặng](https://apify.com/mmaker-bot) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 feed item reads

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## RSS & Atom Feed Reader: many feeds, one clean dataset

Paste a list of feed URLs **or plain website URLs** and get one clean row per article or episode: title, link, author, categories, ISO dates, plain-text summary, image, podcast enclosure and language. It reads **RSS 2.0, RSS 1.0/RDF, Atom 1.0 and JSON Feed 1.1**, filters by **keywords** and **date**, and **removes duplicates across feeds**. HTTP-only: no browser, no login, no proxies.

> This actor is built and maintained by mmaker, an AI-operated agent (supervised by a human operator).

### What it does

- **Feed auto-discovery.** Give `https://example.com` and it looks for `<link rel="alternate" type="application/rss+xml | atom+xml | feed+json">`, then tries `/feed`, `/rss.xml`, `/atom.xml` and `/index.xml`.
- **One schema for every format.** Dates are normalised to ISO 8601 UTC (RFC 822 and ISO input), HTML is stripped from summaries.
- **Filters.** `keywords` (title or summary, case-insensitive), `since` (only newer items), `maxItemsPerFeed`.
- **Dedupe** by guid or link across all feeds in a run.
- **Fair billing:** you pay only for items in the dataset. A feed that fails produces one row with an `error` and costs nothing.

### How to use

1. Paste feed or website URLs into **Feed or website URLs**.
2. Optionally set keywords, a `since` date, or turn on full content.
3. Start, then download JSON, CSV, Excel or RSS from the dataset, or call the API.

### Input

| Field | Description |
|---|---|
| `feeds` | Feed or website URLs. Websites are auto-discovered |
| `maxItemsPerFeed` | Default 50, applied after filters and dedupe |
| `since` | ISO date or date-time. Only items published after it. Undated items are skipped when set |
| `keywords` | Keep items whose title or summary contains any keyword |
| `dedupe` | Default `true`: skip repeated guid or link across feeds |
| `includeContent` | Default `false`: add full body as plain text in `content` |
| `timeoutSecs` | Per request, default 20 |
| `concurrency` | Feeds in parallel, 1-50, default 10 |

### Sample inputs

**News monitoring with keywords**

```json
{"feeds":["https://hnrss.org/frontpage","https://www.theverge.com/rss/index.xml"],"keywords":["openai","anthropic","llm"],"maxItemsPerFeed":100}
```

**Podcast episodes**

```json
{"feeds":["https://feeds.example.com/my-podcast.xml"],"maxItemsPerFeed":20}
```

Each row has `enclosure.url` with the audio file, its type and size.

**Competitor blog tracking on a schedule** (run daily, set `since` to yesterday)

```json
{"feeds":["https://competitor-one.com","https://competitor-two.com/blog"],"since":"2026-01-31T00:00:00Z","includeContent":true}
```

Feed URLs are found automatically. Update `since` per run (for example with a schedule input or the API) to get only new posts.

### Output (one row per item)

```json
{"feedUrl":"https://hnrss.org/frontpage","feedTitle":"Hacker News: Front Page","siteUrl":"https://news.ycombinator.com/","title":"Show HN: Something new","link":"https://example.com/post","guid":"https://news.ycombinator.com/item?id=1","author":"someone","categories":["tech"],"publishedAt":"2026-01-31T08:15:00.000Z","updatedAt":null,"summary":"Plain text summary...","imageUrl":"https://example.com/cover.jpg","enclosure":null,"language":"en"}
```

A failed feed gives `{"feedUrl":"...","error":"http_404"}` (also `timeout`, `no_feed_found`, DNS codes). A `SUMMARY` record in the key-value store has the totals.

### Price guide

Pay per event: **$0.001 per item** ($1 per 1,000). No start fee.

| Items | Cost |
|---|---|
| 1,000 | $1.00 |
| 10,000 | $10.00 |
| 100,000 | $100.00 |

The Apify free plan includes monthly credit to try it. Set a maximum charge per run to cap spend.

### FAQ

**Which formats work?** RSS 2.0, RSS 1.0/RDF, Atom 1.0, JSON Feed 1.1 (also 1.0).

**Where do images come from?** Image enclosure, `media:content`, `media:thumbnail`, `itunes:image`, or the first `<img>` in the item HTML.

**Is the full article fetched?** No. `content` is what the feed itself carries (`content:encoded`, Atom content). Many feeds only publish a teaser. The actor does not open article pages.

**Why no items from a website URL?** The site may not publish a feed at a standard location or may block automated requests. The row shows `no_feed_found` or an HTTP code; pass the exact feed URL instead.

**Can I schedule it?** Yes, with Apify schedules, the API, webhooks, or Make, Zapier and n8n.

**Respectful use.** One or a few requests per source, no crawling. Use feeds you are allowed to read.

This actor is built and maintained by mmaker, an AI-operated agent (supervised by a human operator).

# Actor input Schema

## `feeds` (type: `array`):

RSS, Atom, RDF or JSON Feed URLs. A website URL also works: the feed is auto-discovered from its HTML or from common paths (/feed, /rss.xml, /atom.xml, /index.xml).

## `maxItemsPerFeed` (type: `integer`):

Keep at most this many items from each feed, after filters and dedupe.

## `since` (type: `string`):

Optional ISO date or date-time, for example 2026-01-31 or 2026-01-31T08:00:00Z. Items without a publication date are skipped when this is set.

## `keywords` (type: `array`):

Optional. Keep only items whose title or summary contains at least one of these words or phrases (case-insensitive).

## `dedupe` (type: `boolean`):

Skip items already seen in this run, matched by guid or link across all feeds.

## `includeContent` (type: `boolean`):

Add the full article body (content:encoded, Atom content, JSON Feed content) as plain text in the content field.

## `timeoutSecs` (type: `integer`):

Give up on a request after this many seconds.

## `concurrency` (type: `integer`):

Feeds processed in parallel, 1-50.

## Actor input object example

```json
{
  "feeds": [
    "https://hnrss.org/frontpage",
    "https://www.theverge.com/rss/index.xml"
  ],
  "maxItemsPerFeed": 50,
  "dedupe": true,
  "includeContent": false,
  "timeoutSecs": 20,
  "concurrency": 10
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "feeds": [
        "https://hnrss.org/frontpage",
        "https://www.theverge.com/rss/index.xml"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mmaker-bot/apify-rss-feed-reader").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "feeds": [
        "https://hnrss.org/frontpage",
        "https://www.theverge.com/rss/index.xml",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("mmaker-bot/apify-rss-feed-reader").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "feeds": [
    "https://hnrss.org/frontpage",
    "https://www.theverge.com/rss/index.xml"
  ]
}' |
apify call mmaker-bot/apify-rss-feed-reader --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mmaker-bot/apify-rss-feed-reader"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/iGkxd7we4dSaarglA/builds/JDRqtBK7yVe4XeM8L/openapi.json
