# RSS & Atom Feed Reader – JSON, Only New Items (`far003/rss-atom-feed-reader-only-new`) Actor

Turn RSS, Atom and JSON feeds into clean JSON. Scheduled, it delivers only the items it has never delivered before. Keyword filters, feed discovery from any page.

- **URL**: https://apify.com/far003/rss-atom-feed-reader-only-new.md
- **Developed by:** [Francesco Antonio Russo](https://apify.com/far003) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 feed items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## RSS & Atom Feed Reader – JSON, Only New Items

Reads any list of feeds — **RSS 2.0, RSS 1.0 (RDF), Atom and JSON Feed** — and returns one clean JSON record per item: title, link, dates normalized to ISO 8601, author, plain-text summary, full content (HTML and text), categories, enclosures, image. Give it a normal web page and it reads the feed the page declares.

**Scheduled, it delivers only the items it has never delivered before.** It remembers every item it hands over (per feed, per monitor name) in a named key-value store, so the same article is never returned twice — across runs, across days, even when the feed reorders or drops it. Checking a feed that has nothing new costs nothing: no fixed fee per run, no fee per feed, and a feed that answers `304 Not Modified` is not even downloaded.

| | |
|---|---|
| **Formats** | RSS 2.0 / 0.9x, RSS 1.0 (RDF), Atom 1.0 / 0.3, JSON Feed 1.x; `dc:` (creator, date, subject), `content:encoded`, `media:` (YouTube, Vimeo: description, thumbnails, keywords), `itunes:` (podcasts: author, duration, episode, season, explicit) |
| **Only new** | persistent per-feed memory keyed by `guid`/`id`, else link, else a hash of title + date + summary; forgets an item 90 days after it leaves the feed |
| **Discovery** | an HTML page → its `<link rel="alternate" type="application/rss+xml|atom+xml|feed+json">` |
| **Filters** | include / exclude keywords (whole words or substrings, phrases allowed; searched in title, summary, full content, categories, author), max age, published-after date, max items per feed |
| **Dates** | RFC 822, RFC 3339 and the localized forms real feeds use (`mer, 07 ott 2026`, `Mi, 07 Okt 2026`, `CEST`, `GMT+01:00`), all normalized to ISO 8601 UTC; a date without a zone is read as UTC, never as the server's local time |
| **Polite** | conditional requests (`ETag` / `Last-Modified`), 1 request per second per host, identifies itself as `FeedReaderBot` |
| **Robust** | charset from BOM / header / XML declaration, HTML entities, malformed XML parsed leniently (with a warning), duplicate ids collapsed, oversized feeds cut at the last complete item, dates that cannot be read are reported, not guessed |

**Typical uses:** feed a Slack/Telegram/email alert with only the new posts · build a news or blog aggregator without duplicates · watch competitors' blogs, release notes, job pages, podcast feeds · collect articles for a RAG pipeline or an LLM summarizer · turn any site that has a feed into a JSON API.

### Quick start

1. Put one or more feed URLs (or site pages) in **Feed URLs**.
2. Run. The first run returns what the feed currently holds (newest first, up to **Max items per feed**).
3. Schedule it. Every later run returns only the items that appeared since — nothing when there is nothing new.

```json
{
  "feedUrls": ["https://blog.apify.com/rss/", "https://hnrss.org/frontpage"],
  "maxItemsPerFeed": 20
}
```

Prefer a silent start? Set `firstRunDelivers` to `false`: the first run remembers the current items as a baseline, returns nothing and costs nothing; from the second run on you get the additions only.

### Input

| Field | Default | Meaning |
|---|---|---|
| `feedUrls` | — | Feed URLs or web pages, one per line. `example.com/feed` is enough (https:// is assumed) |
| `onlyNew` | `true` | Remember delivered items and skip them in later runs. `false`: every run returns everything in the feed (`isNew` is then `null`) |
| `firstRunDelivers` | `true` | First run of a feed returns its current items. `false`: first run only builds the baseline, free |
| `stateKey` | `default` | Monitor name. Runs with the same name share one memory; use different names for independent schedules on the same feeds |
| `maxItemsPerFeed` | `100` | Newest first; the rest stay new for the next run. It is your cost cap per feed |
| `maxItemAgeDays` | `0` | Skip items published more than N days ago (0 = no limit). Undated items are kept |
| `publishedAfter` | — | ISO date; skip items published before it |
| `includeKeywords` | — | Deliver only items containing at least one of these (case-insensitive; searched in title, summary, full content, categories and author — also when `includeContent` is off) |
| `excludeKeywords` | — | Skip items containing any of these |
| `matchWholeWords` | `true` | `ai` matches "AI agents" but not "available"; a phrase such as `open source` matches as a phrase. Keywords in scripts without spaces (Chinese, Japanese, Thai…) always match as substrings. Off: plain substring |
| `includeContent` | `true` | Add `contentHtml` and `contentText`. Off keeps records small |
| `discoverFeeds` | `true` | When a URL is an HTML page, read the feed it declares |
| `timeoutSecs` | `20` | Per request |
| `maxFeedMegabytes` | `10` | A larger feed is cut at its last complete item (the SUMMARY tells you how many items were read and to raise this) |

Items skipped by a filter are remembered too, so loosening a filter later does not flood you with old posts. Items held back by `maxItemsPerFeed` or by your spending limit are **not** remembered: they come in the next run. An unreadable `publishedAfter` or a non-numeric number stops the run before anything is charged.

### Output

One dataset record per delivered item. Real example (Apify blog, shortened):

```json
{
  "itemId": "6a841aaa343bc200019df9e6",
  "title": "Martina Gelnerová: from travel agency to data analyst",
  "link": "https://blog.apify.com/martina-gelnerova-data-career-change/",
  "publishedAt": "2026-10-01T08:46:24.000Z",
  "updatedAt": null,
  "author": "Nathanael Durham",
  "authors": ["Nathanael Durham"],
  "summary": "Apify mentors a team at every Czechitas Digital Academy. Martina Gelnerová was one of those students, and now she mentors there herself. …",
  "contentText": "Martina Gelnerová spent twenty years at a travel agency, pricing packages, and she loved it. \"I was so satisfied with the job there,\" she says.\nThen COVID arriv …",
  "contentHtml": "<img src=\"https://storage.ghost.io/c/f2/6e/f26ec999-9a90-4aee-a0d4-9b3ca2bb668f/content/images/2026/10/travel-agency-to- …",
  "categories": ["Life at Apify"],
  "enclosures": [],
  "image": "https://storage.ghost.io/c/f2/6e/f26ec999-9a90-4aee-a0d4-9b3ca2bb668f/content/images/2026/10/travel-agency-to-data-analyst.png",
  "commentsUrl": null,
  "language": null,
  "durationSeconds": null,
  "episode": null,
  "season": null,
  "explicit": null,
  "feedUrl": "https://blog.apify.com/rss/",
  "feedTitle": "Apify Blog",
  "feedLink": "https://blog.apify.com/",
  "feedType": "rss2",
  "fetchedAt": "2026-10-07T11:48:02.894Z",
  "isNew": true,
  "firstSeenAt": "2026-10-07T11:48:02.894Z"
}
```

| Field | Notes |
|---|---|
| `itemId` | the key used for "only new": `guid` / Atom `id` / JSON Feed `id`, else the link, else `hash:…` |
| `publishedAt`, `updatedAt` | ISO 8601 UTC, from `pubDate`, `dc:date`, `published`, `updated`, `date_published`… `null` when absent or unreadable (the SUMMARY names the unreadable value) |
| `author`, `authors` | `dc:creator`, `author` ("email (Name)" → Name; a bare address is not a name), `itunes:author`, `media:credit` (Vimeo), Atom `author/name`, JSON Feed `authors`; falls back to the feed-level `itunes:author` / `dc:creator` |
| `summary` | plain text of `description` / `summary` / `media:description`; when the feed has none — or it says nothing ("…", or just the title again, as Substack does) — the first 500 characters of the content |
| `contentText`, `contentHtml` | `content:encoded`, Atom `content` (text, html or xhtml), `content_html` / `content_text`; `null` when the feed carries only a summary |
| `categories` | `category`, Atom `category` labels, `dc:subject` (RDF feeds), `itunes:keywords`, `media:keywords`, JSON Feed `tags` |
| `enclosures` | `{ url, type, length }` from `enclosure`, Atom `rel="enclosure"`, `media:content`, JSON Feed `attachments`. Dead Flash players (YouTube) are dropped; placeholder sizes some podcast hosts write (5242880) become `null` |
| `image` | `media:content` (image) / `media:thumbnail` / `itunes:image` / JSON Feed `image`, else the first real `<img>` in the content (tracking pixels, icons and feed-service beacons skipped) |
| `durationSeconds`, `episode`, `season`, `explicit` | podcasts and video: `itunes:duration` ("1:02:03" → 3723) or `media:content duration`, `itunes:episode`, `itunes:season`, `itunes:explicit` |
| `title` | the feed's title, HTML removed; when an item has none (Mastodon, micro.blog) the first line of its text, up to 120 characters |
| `feedUrl` | always the URL you gave, even when a page led to a discovered feed or a redirect was followed |

The **SUMMARY** record (key-value store) lists every feed with `status` (`ok`, `not_modified`, `error`), HTTP status, the discovered feed URL, counts (`itemsInFeed`, `itemsNew`, `itemsDelivered`, `itemsFiltered`, `itemsHeldBack`), warnings (malformed XML, unreadable dates, duplicate ids, held-back items) and the error text when a feed could not be read.

### Pricing

One event, **`feed-item`** ($0.001): one item delivered to the dataset. There is no fee per run and none per feed: a feed check with nothing new, a `304`, an error, an item skipped by a filter and the silent baseline run all cost $0. The platform usage of the run (compute, storage) is included in the event price, as with every pay-per-event Actor. If a run reaches your spending limit it stops delivering, leaves the rest as new for the next run, and says so in the SUMMARY.

Cost cap: `maxItemsPerFeed × number of feeds` per run.

**What it costs in practice**

| Use | Items | Cost |
|---|---|---|
| First run on 10 blogs with 20 items each | 200 | $0.20 |
| Hourly schedule on those 10 blogs, ~5 new posts a day in total | ~150 / month | **~$0.15 / month** |
| Daily digest of 50 news feeds, ~30 new items a day | ~900 / month | ~$0.90 / month |
| Reading a 1,000-item podcast archive once (`maxItemsPerFeed: 1000`) | 1,000 | $1.00 |

Readers that charge per feed or per run cost the same whether anything is new or not; a monitor that checks 10 feeds every hour at $0.002 a check is $0.48 a day. Here an empty check is free — that is the point of "only new".

### Why this one

- **Memory that survives.** The popular readers return the whole feed on every run and tell you to deduplicate downstream. This one remembers what it delivered, per feed and per monitor name, and never hands you the same item twice.
- **Every format, really.** RSS 2.0, RSS 1.0/RDF, Atom and JSON Feed, plus the namespaces that carry the useful fields (YouTube's `media:description`, podcast `itunes:` data, `dc:subject` categories). Give it a web page and it finds the feed.
- **Dates you can sort on.** Italian, German, French, Spanish, Portuguese and Dutch date strings, European zone names and zone-less dates all become ISO 8601 UTC.
- **Filters that mean what they say.** Whole-word keyword matching on the full content, even when you keep the output small.
- **Zero cost when nothing happened.** No start fee, no per-feed fee, conditional requests honoured.

### Use the output

```js
// Apify API, after a scheduled run (JavaScript)
const res = await fetch(`https://api.apify.com/v2/datasets/${datasetId}/items?token=${APIFY_TOKEN}`);
const newItems = await res.json(); // [{ title, link, publishedAt, summary, … }]
```

```python
## Python
from apify_client import ApifyClient
run = ApifyClient(token).actor("far003/rss-atom-feed-reader-only-new").call(run_input={"feedUrls": ["https://blog.apify.com/rss/"]})
items = ApifyClient(token).dataset(run["defaultDatasetId"]).list_items().items
```

n8n / Make / Zapier: use the Apify node or module with this Actor and read the run's dataset — each run already contains only the new items, so a "Split in batches → send message" flow needs no deduplication step.

### Limits and honest notes

- **No JavaScript, no login, no bot-protected hosts.** A feed served behind Cloudflare challenges or returning `403`/`419`/`429` to automated clients is reported as an error, not retried in disguise. Hacker News' own `/rss` is one of those; `hnrss.org` works.
- **Only new** needs a stable `guid`/`id` or link. A feed that changes its ids on every build (rare, but it happens) will re-deliver items; the SUMMARY shows `itemsInFeed` ≈ `itemsNew` every run when that is the case.
- A feed that changes an item's title or content after publication is **not** re-delivered: the id is the same. Use `onlyNew: false` if you need every current version.
- Memory per feed is capped at 20,000 ids; items that left the feed are forgotten after 90 days. Rename the monitor (`stateKey`) to start from scratch.
- Feeds larger than `maxFeedMegabytes` (default 10 MB) are cut at the last complete item: a 20 MB podcast archive gives its first ~1,500 episodes and a SUMMARY warning that says to raise the limit.
- A long first run: a feed with 3,000 items and `maxItemsPerFeed: 100` delivers 100 old items per run until the backlog is drained. Set `firstRunDelivers: false` (silent baseline) or `maxItemAgeDays` when you only want what is recent.
- Two runs with the same `stateKey` at the same moment are not locked against each other: schedule them apart.
- An HTML page that declares several feeds: the first is read; the others are listed in the SUMMARY warning.
- Private and loopback addresses are refused.

### Scheduling and integrations

Create a **Schedule** (e.g. every hour) with your feed list and `firstRunDelivers: false`; connect the run's dataset to a webhook, Make, Zapier, n8n or the Apify API (`GET …/datasets/{id}/items`). Each run's dataset holds only the new items — no deduplication needed on your side.

Use a different `stateKey` per schedule when two schedules read the same feeds (e.g. one for Slack, one for your archive), so each keeps its own memory.

# Actor input Schema

## `feedUrls` (type: `array`):

RSS 2.0, RSS 1.0, Atom or JSON Feed URLs. A normal web page works too: the feed it declares in its <head> is read instead.

## `onlyNew` (type: `boolean`):

Remembers every item it delivers (in a named key-value store) and skips it in later runs. Off: every run returns every item currently in the feed.

## `firstRunDelivers` (type: `boolean`):

On (default): the first run of a feed returns everything currently in it, later runs only the additions. Off: the first run silently remembers the current items as the baseline, costs nothing and returns nothing.

## `stateKey` (type: `string`):

Runs with the same name share one memory. Use different names for independent schedules on the same feeds.

## `maxItemsPerFeed` (type: `integer`):

Newest first. Items beyond this number stay "new" for the next run. This is your cost cap per feed (one item = one charged event).

## `maxItemAgeDays` (type: `integer`):

0 = no limit. Items published more than this many days ago are skipped (and never delivered later). Items without a readable date are kept.

## `publishedAfter` (type: `string`):

ISO date or date-time, e.g. 2026-10-01. Items published before it are skipped.

## `includeKeywords` (type: `array`):

Deliver only items whose title, summary, content, categories or author contain at least one of these (case-insensitive).

## `excludeKeywords` (type: `array`):

Skip items that contain any of these (case-insensitive). Applied after the include list.

## `matchWholeWords` (type: `boolean`):

On: "ai" matches "AI agents" but not "available"; a phrase matches as a phrase. Off: plain substring match.

## `includeContent` (type: `boolean`):

Adds contentHtml and contentText (the article body when the feed carries it). Off keeps records small; keyword filters still search the full content.

## `discoverFeeds` (type: `boolean`):

When a URL returns an HTML page, read the feed it declares (<link rel="alternate" type="application/rss+xml">). Off: an HTML page is reported as an error.

## `timeoutSecs` (type: `integer`):

3-120.

## `maxFeedMegabytes` (type: `integer`):

1-50. A larger feed is cut at its last complete item and a warning tells you to raise this.

## Actor input object example

```json
{
  "feedUrls": [
    "https://blog.apify.com/rss/",
    "https://hnrss.org/frontpage"
  ],
  "onlyNew": true,
  "firstRunDelivers": true,
  "stateKey": "default",
  "maxItemsPerFeed": 20,
  "maxItemAgeDays": 0,
  "matchWholeWords": true,
  "includeContent": true,
  "discoverFeeds": true,
  "timeoutSecs": 20,
  "maxFeedMegabytes": 10
}
```

# Actor output Schema

## `items` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "feedUrls": [
        "https://blog.apify.com/rss/",
        "https://hnrss.org/frontpage"
    ],
    "maxItemsPerFeed": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("far003/rss-atom-feed-reader-only-new").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "feedUrls": [
        "https://blog.apify.com/rss/",
        "https://hnrss.org/frontpage",
    ],
    "maxItemsPerFeed": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("far003/rss-atom-feed-reader-only-new").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "feedUrls": [
    "https://blog.apify.com/rss/",
    "https://hnrss.org/frontpage"
  ],
  "maxItemsPerFeed": 20
}' |
apify call far003/rss-atom-feed-reader-only-new --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,far003/rss-atom-feed-reader-only-new"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/6xmBFElYB3UCWfT4y/builds/IMDSOJlCXwHbbjnOY/openapi.json
