RSS & Atom Feed Monitor — New Items, Filters & Webhooks avatar

RSS & Atom Feed Monitor — New Items, Filters & Webhooks

Pricing

from $0.18 / 1,000 feed items

Go to Apify Store
RSS & Atom Feed Monitor — New Items, Filters & Webhooks

RSS & Atom Feed Monitor — New Items, Filters & Webhooks

Monitor RSS, Atom, RSS 1.0 and JSON feeds — news, blogs, YouTube channels, podcasts, GitHub releases — or an OPML list. Only new items after the first run, with keyword, author, category and age filters, cross-feed dedupe and a webhook. No API key.

Pricing

from $0.18 / 1,000 feed items

Rating

0.0

(0)

Developer

Insight Solutions

Insight Solutions

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 hours ago

Last modified

Share

Feeds in, items out — every item on the first run, only the new ones after that. Give it RSS, Atom, RSS 1.0 or JSON Feed addresses (or an OPML export from your feed reader) and it returns one flat row per item: title, link, publish time in UTC, author, summary, the full body as Markdown-ish text, categories, image, and — for podcasts and YouTube — the enclosure, duration, episode, video and channel ids.

Schedule it and it becomes a monitor: a named state store remembers what each feed has already shown, sends the feed's ETag back so an unchanged feed costs one tiny request, and each run emits only what is new. Filter by keyword, author, category or age; merge the same story arriving from several feeds; POST the new items to a webhook.

No API key. No login. No browser. $0.30 per 1,000 new items, and a run that finds nothing new costs nothing.

At a glance

Input — this is the Store prefill; paste it and run:

{ "feedUrls": ["https://feeds.bbci.co.uk/news/rss.xml", "https://news.ycombinator.com/rss",
"https://github.com/apify/crawlee/releases.atom", "https://www.theverge.com/rss/index.xml"],
"opmlUrl": "", "opml": "", "mode": "snapshot", "firstRunBehavior": "emit-all", "keywords": [],
"excludeKeywords": [], "authors": [], "categories": [], "maxItemsPerFeed": 25, "includeContent": true,
"dedupe": true, "webhookUrl": "", "stateStoreName": "rss-feed-monitor-state", "maxConcurrency": 5,
"maxRunSecs": 240, "proxyConfiguration": { "useApifyProxy": true } }

That is four requests and about seventy item rows (up to 25 from each feed; the GitHub and Verge feeds hold 10 each) plus four free feed rows. It uses snapshot so it returns items every time you try it; switch mode to monitor for a scheduled watch.

Output — one item row per new item. The fields you will use most are title, link, publishedAt, author, summary, content, categories and feedTitle (full list under Output reference). Each feed also gets a free feed summary row, and a feed that could not be read comes back as a free diagnostic row (ok: false, errorType, error) instead of a charge.

Price — $0.30 per 1,000 new items (+ $0.001 per run); feed summaries, diagnostics and empty runs free; no API key, no browser, limited permissions, works over the Apify MCP server and with x402.

From code — client.actor("insight.solutions/rss-feed-monitor").call(run_input={…}) with apify-client, or POST https://api.apify.com/v2/acts/insight.solutions~rss-feed-monitor/run-sync-get-dataset-items.


What you get

An item row from The Verge's Atom feed, as the Actor writes it (long text abridged where marked):

{
"ok": true,
"rowType": "item", // "item" | "feed" | "diagnostic"
"itemId": "https://www.theverge.com/?p=1002779",
"title": "The AI Tamagotchis are coming",
"link": "https://www.theverge.com/ai-artificial-intelligence/1002779/openai-dots-meta-muse-ai-agents-hardware-devices",
"externalUrl": null,
"publishedAt": "2026-09-30T18:07:38.000Z", // UTC, whatever zone the feed wrote
"updatedAt": "2026-09-30T18:07:38.000Z",
"author": "Hayden Field",
"authorUrl": null,
"summary": "While AI has made plenty of inroads on people's phones and computers, it's largely failed in dedicated devices. …",
"content": "Sam Altman onstage at OpenAI’s DevDay 2026. | Image: Hayden Field / The Verge\n\nWhile AI has made plenty of inroads … [abridged]",
"contentHtml": "<figure>\n\n<img alt=\"\" … [abridged]",
"contentText": "Sam Altman onstage at OpenAI’s DevDay 2026. … [abridged]",
"wordCount": 125,
"categories": ["AI", "Meta", "OpenAI", "Report", "Tech"],
"imageUrl": null,
"enclosureUrl": null, "enclosureType": null, "enclosureLength": null,
"durationSec": null, "episode": null, "season": null,
"videoId": null, "channelId": null,
"feedTitle": "The Verge",
"feedUrl": "https://www.theverge.com/rss/index.xml",
"feedKind": "atom", // "rss" | "atom" | "rdf" | "json"
"discoveredFrom": null,
"isNew": true, // monitor mode; null in snapshot mode
"firstSeenAt": "2026-09-30T19:00:00.000Z",
"alsoIn": [], // other feeds in this run that carried the same item
"matchedKeywords": [],
"rank": 1, // position in the feed as listed
"input": "https://www.theverge.com/rss/index.xml",
"source": "www.theverge.com",
"sourceUrl": "https://www.theverge.com/rss/index.xml",
"error": null, "errorType": null,
"scrapedAt": "2026-09-30T19:00:00.000Z"
// …plus the feed-row columns below, all null on an item row
}

A podcast episode fills the media columns:

{
"title": "LOW - Now Available",
"author": "Jack Rhysider",
"enclosureUrl": "https://www.podtrac.com/pts/redirect.mp3/dovetail.prxu.org/7057/4148b4aa-…/low_trailer.mp3",
"enclosureType": "audio/mpeg",
"enclosureLength": 3990156,
"durationSec": 219,
"imageUrl": "https://f.prxu.org/7057/4148b4aa-…/lowlogolg.jpg",
"feedTitle": "Darknet Diaries"
}

And every feed read gets one free feed row:

{
"rowType": "feed",
"feedTitle": "The Verge", "feedUrl": "https://www.theverge.com/rss/index.xml", "feedKind": "atom",
"homeUrl": "https://www.theverge.com", "language": "en-US",
"lastBuildDate": "2026-09-30T18:07:38.000Z",
"itemCount": 10, // items the feed holds
"newItems": 1, // item rows this run emitted from it
"filteredOut": 0, // new items your filters excluded (free)
"notModified": false, // true when the feed answered 304 to our ETag
"etag": "W/\"296a4a9251b3a1d67ca6f735c3d33d8f\"",
"isBaseline": true, // first time monitor mode saw this feed
"bytes": 34325
}

Quick start

Four feeds, everything they hold now — the prefill above ("mode": "snapshot").

Monitor 50 feeds hourly and get the new items on a webhook. Save this as a task and schedule it every hour:

{ "feedUrls": ["https://feeds.bbci.co.uk/news/rss.xml", "https://www.theguardian.com/world/rss", "…"],
"mode": "monitor", "firstRunBehavior": "baseline-only",
"keywords": ["election", "central bank"], "excludeKeywords": ["live blog"],
"webhookUrl": "https://hooks.example.com/your-secret-path",
"stateStoreName": "news-watch", "maxRunSecs": 600 }

baseline-only makes the first run record what each feed holds without emitting it, so you only ever pay for items published after you started watching.

Import your feed reader's subscriptions. Export OPML from your reader and paste the file into opml, or put its address in opmlUrl. Nested folders are fine; up to 500 feeds per run.

{ "opmlUrl": "https://example.com/my-subscriptions.opml", "mode": "monitor", "includeContent": false }

YouTube channels and podcasts. A channel's feed is https://www.youtube.com/feeds/videos.xml?channel_id=UC… (rows carry videoId, channelId and the thumbnail); a podcast's RSS address comes with its audio file, duration, episode and season.

{ "feedUrls": ["https://www.youtube.com/feeds/videos.xml?channel_id=UCsBjURrPoezykLs9EqgamOA",
"https://podcast.darknetdiaries.com/"],
"mode": "monitor", "maxItemsPerFeed": 20 }

Formats supported

FormatWhere you meet itWhat is read
RSS 2.0most news sites, WordPress, Medium, dev.to, Hacker News, podcaststitle, link, guid, pubDate, description, content:encoded, category (with domain), author, dc:creator, dc:subject, enclosure
RSS 1.0 / RDFSlashdot and older portals<item rdf:about> outside <channel>, dc:date, dc:creator, dc:subject; ISO-8859-1 and other declared charsets are decoded
AtomGitHub releases, YouTube, gov.uk, The Verge<entry>, <id>, <link rel="alternate">, <published>/<updated>, <summary>, `<content type="html
JSON Feed 1.0 / 1.1Daring Fireball and other JSON Feed publishersid, url, external_url, title, content_html, content_text, summary, image, date_published, date_modified, authors[], tags[], attachments[]
Extensionsmedia:content, media:thumbnail, media:group/media:description (YouTube), itunes:duration, itunes:episode, itunes:season, itunes:image, itunes:author, yt:videoId, yt:channelId

The format is decided by the document's root, never by the URL. A page address that is not a feed is checked for the feed it advertises in its <head> (<link rel="alternate" type="application/rss+xml|atom+xml|feed+json">) — one extra request — and that feed is read; the rows say discoveredFrom.

Input

FieldTypeDefaultWhat it does
feedUrlsarray[]Feed addresses, or pages that advertise one. feed:// links and bare host/path are accepted.
opmlUrlstring""An OPML file's address; every <outline xmlUrl> is added (duplicates once, 500 at most).
opmlstring""OPML text pasted directly.
modeselectmonitormonitor: only items not seen before (uses the state store). snapshot: everything the feeds hold now, no state.
firstRunBehaviorselectemit-allMonitor mode, first sight of a feed: emit-all emits its current items; baseline-only records them and emits nothing.
keywordsarray[]Keep items where any appears (case-insensitive) in title, summary, body or categories.
excludeKeywordsarray[]Drop items where any appears.
authorsarray[]Keep items whose byline contains any.
categoriesarray[]Keep items with a category/tag equal to any (case-insensitive).
sinceHoursinteger—Keep items published in the last N hours; undated items pass.
maxItemsPerFeedinteger100Items considered per feed per run, in the feed's own order. 0 = all.
includeContentbooleantrueOff: content, contentText and contentHtml are null (the summary stays).
dedupebooleantrueMerge the same item arriving from several feeds in one run into one row (alsoIn).
webhookUrlstring""One POST per run with the new items, only when there are any.
stateStoreNamestringrss-feed-monitor-stateThe named key-value store monitor mode remembers in. One per independent schedule.
maxConcurrencyinteger5Feeds read at once (1–20). Requests to one host always start ≥ 1 s apart.
maxRunSecsinteger240Run budget, 30–3600 s.
proxyConfigurationobjectdatacenterApify Proxy settings.

At least one feed address, after OPML expansion, is required; with none the run finishes FAILED before it makes a request.

Output reference

Every row has every column (null where it does not apply), so the dataset exports as one rectangular table. Three views are defined: Items, Feeds and Diagnostics.

  • Envelope (every row): ok, rowType (item | feed | diagnostic), input, error, errorType, scrapedAt, source (the feed's host), sourceUrl (the address fetched, after redirects).
  • item rows (charged): itemId, title, link, externalUrl, publishedAt, updatedAt, author, authorUrl, summary (≤ 2,000 characters), content (Markdown-ish: paragraphs, headings, lists, code, links as [text](url); ≤ 100,000 characters), contentHtml (raw, ≤ 50,000 characters), contentText (tags stripped), wordCount (whole body), categories[], imageUrl, enclosureUrl, enclosureType, enclosureLength, durationSec, episode, season, videoId, channelId, feedTitle, feedUrl, feedKind, discoveredFrom, isNew, firstSeenAt, alsoIn[], matchedKeywords[], rank.
  • feed rows (free, one per feed read): feedUrl, feedTitle, feedKind, homeUrl, description, language, generator, lastBuildDate, itemCount, newItems, filteredOut, notModified, etag, lastModified, fetchMs, bytes, isBaseline; the feed's own author (a podcast's itunes:author, an Atom feed author) and imageUrl (channel image or Atom icon).
  • diagnostic rows (free): errorType is one of invalid-input (an address or webhook URL that could not be used), not-a-feed (HTML with no advertised feed, or something that is not a feed), blocked (403, 429, a challenge page or an empty body, twice, from two proxy sessions), http (404, 5xx, unreachable), timeout, too-large (over 8 MB), parse-error (says what failed), opml (an OPML that could not be fetched or read, or the 500-feed cap), deadline (a feed not reached before maxRunSecs), budget (your maximum charge was reached), state-locked (another run holds the state store, or the store could not be read or written), webhook (the POST did not land) or upstream-format (a feed whose items carry nothing this Actor can read).

E-mail addresses are never emitted. RSS defines <author> as an e-mail address: jane@example.com (Jane Doe) becomes Jane Doe, and a bare address becomes null. webMaster, managingEditor, itunes:owner and itunes:email are never read. As a second guard, every text column — body included — has e-mail addresses removed before it is written. Social handles and profile links (medium.com/@author/…) are not addresses and are kept.

How monitor mode decides what is new

  • Item identity. An item is known by the feed's own id for it — RSS guid, Atom <id>, JSON Feed id, RSS 1.0 rdf:about — else by its canonical link (scheme folded to https, utm_* and fragment dropped), else by sha256(title + published).
  • The state store. For each feed, the named key-value store keeps the ids of the last 5,000 items it has shown (hashed, oldest dropped first), the feed's ETag and Last-Modified, and when it was last read. An item whose id is in that list is not emitted again; everything else is new. Items your filters excluded are recorded as seen too, so they do not return later.
  • Conditional requests. The stored ETag goes back as If-None-Match and Last-Modified as If-Modified-Since. A feed that answers 304 Not Modified gets a free feed row with notModified: true and no item is read at all. (Feeds that send neither header are simply re-read and diffed by id.)
  • Order of writes. A feed's record is updated only after its rows are in the dataset and billed. An item held back by your charge limit is left out of the record, so the next run delivers it rather than losing it.
  • One run at a time per store. A run takes a lock on its state store; a second run started on the same store while the first is working stops at once with a free state-locked row (FAILED, nothing charged). Give concurrent schedules different stateStoreNames.
  • Retention. A feed's record that no run has touched for 90 days is deleted.

What you are never charged for

  • feed rows, every diagnostic row, and the webhook.
  • Items the state store had already seen, items your filters excluded, and duplicates merged into another feed's row (alsoIn) — one item, one charge.
  • A run that returns no item. A run whose input was usable but produced no new item — nothing new since last time, every feed blocked or down, filters that excluded everything, a first run with baseline-only, the time budget — finishes SUCCEEDED with zero results, a status message that says so and points at the diagnostic rows, and bills nothing, start fee included. A run finishes FAILED only when there was nothing to read at all (no usable feed address after OPML expansion), when its state store is locked by another run, or when the Actor itself hit an error.

Pricing

EventFREEBRONZESILVERGOLD
Run started$0.001$0.001$0.001$0.001
Feed item$0.0003$0.0003$0.00024$0.00018

Worked example. Twenty feeds checked hourly that publish about sixty new items a day: 60 × $0.0003 + 24 × $0.001 = $0.042 a day. Hours with nothing new cost nothing — the start fee is billed only once an item has been delivered.

Set a maximum charge on the run (ACTOR_MAX_TOTAL_CHARGE_USD) and it delivers what the budget covers, files a free budget row, leaves the rest unseen for the next run and finishes SUCCEEDED.

Use it from an AI agent, or from code

One JSON object in, one flat array out. The Actor runs with limited permissions, uses pay-per-event pricing and never enters Standby, so it works over the Apify MCP server (mcp.apify.com) and with x402 agentic payments.

curl -X POST "https://api.apify.com/v2/acts/insight.solutions~rss-feed-monitor/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"feedUrls":["https://news.ycombinator.com/rss"],"mode":"snapshot","maxItemsPerFeed":10}'
# pip install apify-client
from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("insight.solutions/rss-feed-monitor").call(run_input={
"feedUrls": ["https://github.com/apify/crawlee/releases.atom"],
"mode": "monitor",
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
if row["rowType"] == "item":
print(row["publishedAt"], row["feedTitle"], row["title"], row["link"])

The webhook payload (one POST per run, only when items were delivered, at most 256 KB — items are dropped from the end to fit and droppedItems says how many):

{ "actorRunId": "…", "runAt": "2026-09-30T19:00:00.000Z", "mode": "monitor", "feedsRead": 4,
"newItems": 2, "droppedItems": 0,
"datasetUrl": "https://api.apify.com/v2/datasets/…/items?clean=true&format=json",
"items": [ { "itemId": "…", "title": "…", "link": "…", "publishedAt": "…", "author": "…",
"feedTitle": "…", "feedUrl": "…", "summary": "…(≤500 chars)", "categories": [],
"imageUrl": null, "alsoIn": [] } ] }

The URL is treated as a secret: it is never written into a row, and only its origin (https://hooks.example.com) ever appears in the log. Redirects are followed once and only on the same host; a failed POST is a free webhook row, and the run still succeeds.

FAQ

I only have the site's address, not its feed. Paste the site address. If its page advertises a feed with <link rel="alternate"> — WordPress sites and many blogs and newsrooms do — the feed is found and read (one extra request) and the rows carry discoveredFrom. If it advertises none you get a free not-a-feed row.

Why are my Reddit or Craigslist feeds blocked? Both refuse traffic from datacenter addresses: in testing Reddit's .rss answered 403 with a block page, and Craigslist's ?format=rss a 403 "Your request has been blocked". The Actor retries once from a fresh proxy session and then files a free blocked row; it does not try to get around a site's refusal. Residential proxies may get through — whether they should is the site's terms' call, and yours.

Where does monitor mode keep its memory, and who can see it? In a named key-value store (stateStoreName) in your own Apify account, created by the Actor on its first run and reused by every later run — including every run of a schedule — which is what limited permissions allow. Runs with different stateStoreNames have separate memories, so two schedules watching different things do not interfere. Delete the store to start over.

What happens the first time monitor mode sees a feed? With emit-all (the default) you get what the feed holds now — up to maxItemsPerFeed — and are charged for it; with baseline-only you get only the free feed row with isBaseline: true, and items appear from the next new one on.

Does it read the full article? It returns what the feed carries. Many feeds carry the whole body (content:encoded, Atom <content>, JSON content_html — WordPress, dev.to, GitHub releases); many carry a teaser only (BBC, NYT, Hacker News). It never fetches the article page itself.

What if two of my feeds carry the same story? With dedupe on it is one row, from the first feed in your list, with the others in alsoIn, and one charge.

Is robots.txt checked? No — a feed is published to be fetched by programs. Blocks are respected: a 403, 429 or challenge page is a free diagnostic, never worked around.

Limitations

  • Only what the feed publishes. A feed that holds the last 10 items cannot tell you about the 11th; run often enough that nothing scrolls off between runs (hourly suits most news feeds).
  • Teaser-only feeds give teaser-only content. The article page is never fetched.
  • The upstream format may change. Feeds are read with a no-dependency scanner, not a validating XML parser. Malformed-but-common feeds (missing XML declaration, CDATA everywhere, HTML in titles) are handled; a feed whose items carry nothing recognisable comes back as a free upstream-format row rather than wrong data.
  • Rows are written after every feed is read, so that alsoIn is complete. The read phase stops at 90% of maxRunSecs (at most 20 s short of it) to leave time for writing.
  • Size caps: feeds over 8 MB are refused (too-large); contentHtml is capped at 50,000 characters, content and contentText at 100,000, summary at 2,000; the state store remembers 5,000 items per feed; OPML adds at most 500 feeds per run.
  • Some sites refuse datacenter traffic (Reddit and Craigslist among them) — see the FAQ.
  • Dates: a feed that writes a date with no time zone is read as UTC.

Sources, terms and attribution

  • Everything is read from public feeds, exactly as their publishers serve them to any feed reader. No login, no cookie, no API key belonging to anyone.
  • Feed content belongs to its publishers. Monitoring, alerting, analysis and linking are what this is built for; republishing is your call and your responsibility. Bylines are returned as published, as attribution.
  • Not affiliated with any publisher named here; names are used only to describe which public feeds were used in testing.

Our other Actors

Every Insight Solutions Actor is pay-per-result with no browser, no login and no API key, and every one of them returns free diagnostic rows instead of billing for failures. Prices are per 1,000 results.

Video, audio & social

News, documents & the web

Business, finance & jobs

Apps & games