# RSS & Atom Feed Monitor — New Items, Filters & Webhooks (`insight.solutions/rss-feed-monitor`) Actor

Monitor RSS, Atom, RSS 1.0 and JSON feeds — news, blogs, YouTube channels, podcasts, GitHub releases — or an OPML list. Only new items after the first run, with keyword, author, category and age filters, cross-feed dedupe and a webhook. No API key.

- **URL**: https://apify.com/insight.solutions/rss-feed-monitor.md
- **Developed by:** [Insight Solutions](https://apify.com/insight.solutions) (community)
- **Categories:** News, Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.18 / 1,000 feed items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## RSS & Atom Feed Monitor — New Items, Filters & Webhooks

**Feeds in, items out — every item on the first run, only the new ones after that.** Give it RSS, Atom, RSS 1.0 or JSON Feed addresses (or an OPML export from your feed reader) and it returns one flat row per item: title, link, publish time in UTC, author, summary, the full body as Markdown-ish text, categories, image, and — for podcasts and YouTube — the enclosure, duration, episode, video and channel ids.

Schedule it and it becomes a monitor: a named state store remembers what each feed has already shown, sends the feed's ETag back so an unchanged feed costs one tiny request, and each run emits only what is new. Filter by keyword, author, category or age; merge the same story arriving from several feeds; POST the new items to a webhook.

No API key. No login. No browser. **$0.30 per 1,000 new items**, and a run that finds nothing new costs nothing.

### At a glance

**Input** — this is the Store prefill; paste it and run:

```json
{ "feedUrls": ["https://feeds.bbci.co.uk/news/rss.xml", "https://news.ycombinator.com/rss",
               "https://github.com/apify/crawlee/releases.atom", "https://www.theverge.com/rss/index.xml"],
  "opmlUrl": "", "opml": "", "mode": "snapshot", "firstRunBehavior": "emit-all", "keywords": [],
  "excludeKeywords": [], "authors": [], "categories": [], "maxItemsPerFeed": 25, "includeContent": true,
  "dedupe": true, "webhookUrl": "", "stateStoreName": "rss-feed-monitor-state", "maxConcurrency": 5,
  "maxRunSecs": 240, "proxyConfiguration": { "useApifyProxy": true } }
```

That is four requests and about seventy `item` rows (up to 25 from each feed; the GitHub and Verge feeds hold 10 each) plus four free `feed` rows. It uses `snapshot` so it returns items every time you try it; switch `mode` to `monitor` for a scheduled watch.

**Output** — one `item` row per new item. The fields you will use most are `title`, `link`, `publishedAt`, `author`, `summary`, `content`, `categories` and `feedTitle` (full list under *Output reference*). Each feed also gets a free `feed` summary row, and a feed that could not be read comes back as a free diagnostic row (`ok: false`, `errorType`, `error`) instead of a charge.

**Price** — $0.30 per 1,000 new items (+ $0.001 per run); feed summaries, diagnostics and empty runs free; no API key, no browser, limited permissions, works over the Apify MCP server and with x402.

**From code** — `client.actor("insight.solutions/rss-feed-monitor").call(run_input={…})` with `apify-client`, or `POST https://api.apify.com/v2/acts/insight.solutions~rss-feed-monitor/run-sync-get-dataset-items`.

***

### What you get

An `item` row from The Verge's Atom feed, as the Actor writes it (long text abridged where marked):

```jsonc
{
  "ok": true,
  "rowType": "item",                     // "item" | "feed" | "diagnostic"
  "itemId": "https://www.theverge.com/?p=1002779",
  "title": "The AI Tamagotchis are coming",
  "link": "https://www.theverge.com/ai-artificial-intelligence/1002779/openai-dots-meta-muse-ai-agents-hardware-devices",
  "externalUrl": null,
  "publishedAt": "2026-09-30T18:07:38.000Z",   // UTC, whatever zone the feed wrote
  "updatedAt": "2026-09-30T18:07:38.000Z",
  "author": "Hayden Field",
  "authorUrl": null,
  "summary": "While AI has made plenty of inroads on people's phones and computers, it's largely failed in dedicated devices. …",
  "content": "Sam Altman onstage at OpenAI’s DevDay 2026. | Image: Hayden Field / The Verge\n\nWhile AI has made plenty of inroads … [abridged]",
  "contentHtml": "<figure>\n\n<img alt=\"\" … [abridged]",
  "contentText": "Sam Altman onstage at OpenAI’s DevDay 2026. … [abridged]",
  "wordCount": 125,
  "categories": ["AI", "Meta", "OpenAI", "Report", "Tech"],
  "imageUrl": null,
  "enclosureUrl": null, "enclosureType": null, "enclosureLength": null,
  "durationSec": null, "episode": null, "season": null,
  "videoId": null, "channelId": null,
  "feedTitle": "The Verge",
  "feedUrl": "https://www.theverge.com/rss/index.xml",
  "feedKind": "atom",                    // "rss" | "atom" | "rdf" | "json"
  "discoveredFrom": null,
  "isNew": true,                         // monitor mode; null in snapshot mode
  "firstSeenAt": "2026-09-30T19:00:00.000Z",
  "alsoIn": [],                          // other feeds in this run that carried the same item
  "matchedKeywords": [],
  "rank": 1,                             // position in the feed as listed

  "input": "https://www.theverge.com/rss/index.xml",
  "source": "www.theverge.com",
  "sourceUrl": "https://www.theverge.com/rss/index.xml",
  "error": null, "errorType": null,
  "scrapedAt": "2026-09-30T19:00:00.000Z"
  // …plus the feed-row columns below, all null on an item row
}
```

A podcast episode fills the media columns:

```jsonc
{
  "title": "LOW - Now Available",
  "author": "Jack Rhysider",
  "enclosureUrl": "https://www.podtrac.com/pts/redirect.mp3/dovetail.prxu.org/7057/4148b4aa-…/low_trailer.mp3",
  "enclosureType": "audio/mpeg",
  "enclosureLength": 3990156,
  "durationSec": 219,
  "imageUrl": "https://f.prxu.org/7057/4148b4aa-…/lowlogolg.jpg",
  "feedTitle": "Darknet Diaries"
}
```

And every feed read gets one free `feed` row:

```jsonc
{
  "rowType": "feed",
  "feedTitle": "The Verge", "feedUrl": "https://www.theverge.com/rss/index.xml", "feedKind": "atom",
  "homeUrl": "https://www.theverge.com", "language": "en-US",
  "lastBuildDate": "2026-09-30T18:07:38.000Z",
  "itemCount": 10,        // items the feed holds
  "newItems": 1,          // item rows this run emitted from it
  "filteredOut": 0,       // new items your filters excluded (free)
  "notModified": false,   // true when the feed answered 304 to our ETag
  "etag": "W/\"296a4a9251b3a1d67ca6f735c3d33d8f\"",
  "isBaseline": true,     // first time monitor mode saw this feed
  "bytes": 34325
}
```

### Quick start

**Four feeds, everything they hold now** — the prefill above (`"mode": "snapshot"`).

**Monitor 50 feeds hourly and get the new items on a webhook.** Save this as a task and schedule it every hour:

```json
{ "feedUrls": ["https://feeds.bbci.co.uk/news/rss.xml", "https://www.theguardian.com/world/rss", "…"],
  "mode": "monitor", "firstRunBehavior": "baseline-only",
  "keywords": ["election", "central bank"], "excludeKeywords": ["live blog"],
  "webhookUrl": "https://hooks.example.com/your-secret-path",
  "stateStoreName": "news-watch", "maxRunSecs": 600 }
```

`baseline-only` makes the first run record what each feed holds without emitting it, so you only ever pay for items published after you started watching.

**Import your feed reader's subscriptions.** Export OPML from your reader and paste the file into `opml`, or put its address in `opmlUrl`. Nested folders are fine; up to 500 feeds per run.

```json
{ "opmlUrl": "https://example.com/my-subscriptions.opml", "mode": "monitor", "includeContent": false }
```

**YouTube channels and podcasts.** A channel's feed is `https://www.youtube.com/feeds/videos.xml?channel_id=UC…` (rows carry `videoId`, `channelId` and the thumbnail); a podcast's RSS address comes with its audio file, duration, episode and season.

```json
{ "feedUrls": ["https://www.youtube.com/feeds/videos.xml?channel_id=UCsBjURrPoezykLs9EqgamOA",
               "https://podcast.darknetdiaries.com/"],
  "mode": "monitor", "maxItemsPerFeed": 20 }
```

### Formats supported

| Format | Where you meet it | What is read |
|---|---|---|
| **RSS 2.0** | most news sites, WordPress, Medium, dev.to, Hacker News, podcasts | `title`, `link`, `guid`, `pubDate`, `description`, `content:encoded`, `category` (with `domain`), `author`, `dc:creator`, `dc:subject`, `enclosure` |
| **RSS 1.0 / RDF** | Slashdot and older portals | `<item rdf:about>` outside `<channel>`, `dc:date`, `dc:creator`, `dc:subject`; ISO-8859-1 and other declared charsets are decoded |
| **Atom** | GitHub releases, YouTube, gov.uk, The Verge | `<entry>`, `<id>`, `<link rel="alternate">`, `<published>`/`<updated>`, `<summary>`, `<content type="html|xhtml|text">`, `<author><name>`/`<uri>` (inherited from the feed when an entry names none), `<category term>`, `<link rel="enclosure">` |
| **JSON Feed 1.0 / 1.1** | Daring Fireball and other JSON Feed publishers | `id`, `url`, `external_url`, `title`, `content_html`, `content_text`, `summary`, `image`, `date_published`, `date_modified`, `authors[]`, `tags[]`, `attachments[]` |
| **Extensions** | | `media:content`, `media:thumbnail`, `media:group`/`media:description` (YouTube), `itunes:duration`, `itunes:episode`, `itunes:season`, `itunes:image`, `itunes:author`, `yt:videoId`, `yt:channelId` |

The format is decided by the document's root, never by the URL. A page address that is not a feed is checked for the feed it advertises in its `<head>` (`<link rel="alternate" type="application/rss+xml|atom+xml|feed+json">`) — one extra request — and that feed is read; the rows say `discoveredFrom`.

### Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `feedUrls` | array | `[]` | Feed addresses, or pages that advertise one. `feed://` links and bare `host/path` are accepted. |
| `opmlUrl` | string | `""` | An OPML file's address; every `<outline xmlUrl>` is added (duplicates once, 500 at most). |
| `opml` | string | `""` | OPML text pasted directly. |
| `mode` | select | `monitor` | `monitor`: only items not seen before (uses the state store). `snapshot`: everything the feeds hold now, no state. |
| `firstRunBehavior` | select | `emit-all` | Monitor mode, first sight of a feed: `emit-all` emits its current items; `baseline-only` records them and emits nothing. |
| `keywords` | array | `[]` | Keep items where any appears (case-insensitive) in title, summary, body or categories. |
| `excludeKeywords` | array | `[]` | Drop items where any appears. |
| `authors` | array | `[]` | Keep items whose byline contains any. |
| `categories` | array | `[]` | Keep items with a category/tag equal to any (case-insensitive). |
| `sinceHours` | integer | — | Keep items published in the last N hours; undated items pass. |
| `maxItemsPerFeed` | integer | `100` | Items considered per feed per run, in the feed's own order. `0` = all. |
| `includeContent` | boolean | `true` | Off: `content`, `contentText` and `contentHtml` are null (the summary stays). |
| `dedupe` | boolean | `true` | Merge the same item arriving from several feeds in one run into one row (`alsoIn`). |
| `webhookUrl` | string | `""` | One POST per run with the new items, only when there are any. |
| `stateStoreName` | string | `rss-feed-monitor-state` | The named key-value store monitor mode remembers in. One per independent schedule. |
| `maxConcurrency` | integer | `5` | Feeds read at once (1–20). Requests to one host always start ≥ 1 s apart. |
| `maxRunSecs` | integer | `240` | Run budget, 30–3600 s. |
| `proxyConfiguration` | object | datacenter | Apify Proxy settings. |

At least one feed address, after OPML expansion, is required; with none the run finishes FAILED before it makes a request.

### Output reference

Every row has every column (null where it does not apply), so the dataset exports as one rectangular table. Three views are defined: **Items**, **Feeds** and **Diagnostics**.

- **Envelope (every row):** `ok`, `rowType` (`item` | `feed` | `diagnostic`), `input`, `error`, `errorType`, `scrapedAt`, `source` (the feed's host), `sourceUrl` (the address fetched, after redirects).
- **`item` rows (charged):** `itemId`, `title`, `link`, `externalUrl`, `publishedAt`, `updatedAt`, `author`, `authorUrl`, `summary` (≤ 2,000 characters), `content` (Markdown-ish: paragraphs, headings, lists, code, links as `[text](url)`; ≤ 100,000 characters), `contentHtml` (raw, ≤ 50,000 characters), `contentText` (tags stripped), `wordCount` (whole body), `categories[]`, `imageUrl`, `enclosureUrl`, `enclosureType`, `enclosureLength`, `durationSec`, `episode`, `season`, `videoId`, `channelId`, `feedTitle`, `feedUrl`, `feedKind`, `discoveredFrom`, `isNew`, `firstSeenAt`, `alsoIn[]`, `matchedKeywords[]`, `rank`.
- **`feed` rows (free, one per feed read):** `feedUrl`, `feedTitle`, `feedKind`, `homeUrl`, `description`, `language`, `generator`, `lastBuildDate`, `itemCount`, `newItems`, `filteredOut`, `notModified`, `etag`, `lastModified`, `fetchMs`, `bytes`, `isBaseline`; the feed's own `author` (a podcast's `itunes:author`, an Atom feed author) and `imageUrl` (channel image or Atom icon).
- **`diagnostic` rows (free):** `errorType` is one of `invalid-input` (an address or webhook URL that could not be used), `not-a-feed` (HTML with no advertised feed, or something that is not a feed), `blocked` (403, 429, a challenge page or an empty body, twice, from two proxy sessions), `http` (404, 5xx, unreachable), `timeout`, `too-large` (over 8 MB), `parse-error` (says what failed), `opml` (an OPML that could not be fetched or read, or the 500-feed cap), `deadline` (a feed not reached before `maxRunSecs`), `budget` (your maximum charge was reached), `state-locked` (another run holds the state store, or the store could not be read or written), `webhook` (the POST did not land) or `upstream-format` (a feed whose items carry nothing this Actor can read).

**E-mail addresses are never emitted.** RSS defines `<author>` as an e-mail address: `jane@example.com (Jane Doe)` becomes `Jane Doe`, and a bare address becomes null. `webMaster`, `managingEditor`, `itunes:owner` and `itunes:email` are never read. As a second guard, every text column — body included — has e-mail addresses removed before it is written. Social handles and profile links (`medium.com/@author/…`) are not addresses and are kept.

### How monitor mode decides what is new

- **Item identity.** An item is known by the feed's own id for it — RSS `guid`, Atom `<id>`, JSON Feed `id`, RSS 1.0 `rdf:about` — else by its canonical link (scheme folded to https, `utm_*` and fragment dropped), else by `sha256(title + published)`.
- **The state store.** For each feed, the named key-value store keeps the ids of the last 5,000 items it has shown (hashed, oldest dropped first), the feed's `ETag` and `Last-Modified`, and when it was last read. An item whose id is in that list is not emitted again; everything else is new. Items your filters excluded are recorded as seen too, so they do not return later.
- **Conditional requests.** The stored `ETag` goes back as `If-None-Match` and `Last-Modified` as `If-Modified-Since`. A feed that answers **304 Not Modified** gets a free `feed` row with `notModified: true` and no item is read at all. (Feeds that send neither header are simply re-read and diffed by id.)
- **Order of writes.** A feed's record is updated only after its rows are in the dataset and billed. An item held back by your charge limit is left out of the record, so the next run delivers it rather than losing it.
- **One run at a time per store.** A run takes a lock on its state store; a second run started on the same store while the first is working stops at once with a free `state-locked` row (FAILED, nothing charged). Give concurrent schedules different `stateStoreName`s.
- **Retention.** A feed's record that no run has touched for 90 days is deleted.

### What you are never charged for

- `feed` rows, every diagnostic row, and the webhook.
- Items the state store had already seen, items your filters excluded, and duplicates merged into another feed's row (`alsoIn`) — one item, one charge.
- A run that returns no item. A run whose input was usable but produced no new item — nothing new since last time, every feed blocked or down, filters that excluded everything, a first run with `baseline-only`, the time budget — finishes **SUCCEEDED with zero results**, a status message that says so and points at the diagnostic rows, and bills nothing, start fee included. A run finishes **FAILED** only when there was nothing to read at all (no usable feed address after OPML expansion), when its state store is locked by another run, or when the Actor itself hit an error.

### Pricing

| Event | FREE | BRONZE | SILVER | GOLD |
|---|---|---|---|---|
| Run started | $0.001 | $0.001 | $0.001 | $0.001 |
| **Feed item** | **$0.0003** | $0.0003 | $0.00024 | $0.00018 |

**Worked example.** Twenty feeds checked hourly that publish about sixty new items a day: 60 × $0.0003 + 24 × $0.001 = **$0.042 a day**. Hours with nothing new cost nothing — the start fee is billed only once an item has been delivered.

Set a maximum charge on the run (`ACTOR_MAX_TOTAL_CHARGE_USD`) and it delivers what the budget covers, files a free `budget` row, leaves the rest unseen for the next run and finishes SUCCEEDED.

### Use it from an AI agent, or from code

One JSON object in, one flat array out. The Actor runs with **limited permissions**, uses **pay-per-event** pricing and never enters Standby, so it works over the Apify MCP server (`mcp.apify.com`) and with x402 agentic payments.

```bash
curl -X POST "https://api.apify.com/v2/acts/insight.solutions~rss-feed-monitor/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"feedUrls":["https://news.ycombinator.com/rss"],"mode":"snapshot","maxItemsPerFeed":10}'
```

```python
## pip install apify-client
from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("insight.solutions/rss-feed-monitor").call(run_input={
    "feedUrls": ["https://github.com/apify/crawlee/releases.atom"],
    "mode": "monitor",
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    if row["rowType"] == "item":
        print(row["publishedAt"], row["feedTitle"], row["title"], row["link"])
```

**The webhook payload** (one POST per run, only when items were delivered, at most 256 KB — items are dropped from the end to fit and `droppedItems` says how many):

```json
{ "actorRunId": "…", "runAt": "2026-09-30T19:00:00.000Z", "mode": "monitor", "feedsRead": 4,
  "newItems": 2, "droppedItems": 0,
  "datasetUrl": "https://api.apify.com/v2/datasets/…/items?clean=true&format=json",
  "items": [ { "itemId": "…", "title": "…", "link": "…", "publishedAt": "…", "author": "…",
               "feedTitle": "…", "feedUrl": "…", "summary": "…(≤500 chars)", "categories": [],
               "imageUrl": null, "alsoIn": [] } ] }
```

The URL is treated as a secret: it is never written into a row, and only its origin (`https://hooks.example.com`) ever appears in the log. Redirects are followed once and only on the same host; a failed POST is a free `webhook` row, and the run still succeeds.

### FAQ

**I only have the site's address, not its feed.** Paste the site address. If its page advertises a feed with `<link rel="alternate">` — WordPress sites and many blogs and newsrooms do — the feed is found and read (one extra request) and the rows carry `discoveredFrom`. If it advertises none you get a free `not-a-feed` row.

**Why are my Reddit or Craigslist feeds `blocked`?** Both refuse traffic from datacenter addresses: in testing Reddit's `.rss` answered 403 with a block page, and Craigslist's `?format=rss` a 403 "Your request has been blocked". The Actor retries once from a fresh proxy session and then files a free `blocked` row; it does not try to get around a site's refusal. Residential proxies may get through — whether they should is the site's terms' call, and yours.

**Where does monitor mode keep its memory, and who can see it?** In a named key-value store (`stateStoreName`) in your own Apify account, created by the Actor on its first run and reused by every later run — including every run of a schedule — which is what limited permissions allow. Runs with different `stateStoreName`s have separate memories, so two schedules watching different things do not interfere. Delete the store to start over.

**What happens the first time monitor mode sees a feed?** With `emit-all` (the default) you get what the feed holds now — up to `maxItemsPerFeed` — and are charged for it; with `baseline-only` you get only the free `feed` row with `isBaseline: true`, and items appear from the next new one on.

**Does it read the full article?** It returns what the feed carries. Many feeds carry the whole body (`content:encoded`, Atom `<content>`, JSON `content_html` — WordPress, dev.to, GitHub releases); many carry a teaser only (BBC, NYT, Hacker News). It never fetches the article page itself.

**What if two of my feeds carry the same story?** With `dedupe` on it is one row, from the first feed in your list, with the others in `alsoIn`, and one charge.

**Is robots.txt checked?** No — a feed is published to be fetched by programs. Blocks are respected: a 403, 429 or challenge page is a free diagnostic, never worked around.

### Limitations

- **Only what the feed publishes.** A feed that holds the last 10 items cannot tell you about the 11th; run often enough that nothing scrolls off between runs (hourly suits most news feeds).
- **Teaser-only feeds give teaser-only content.** The article page is never fetched.
- **The upstream format may change.** Feeds are read with a no-dependency scanner, not a validating XML parser. Malformed-but-common feeds (missing XML declaration, CDATA everywhere, HTML in titles) are handled; a feed whose items carry nothing recognisable comes back as a free `upstream-format` row rather than wrong data.
- **Rows are written after every feed is read**, so that `alsoIn` is complete. The read phase stops at 90% of `maxRunSecs` (at most 20 s short of it) to leave time for writing.
- **Size caps:** feeds over 8 MB are refused (`too-large`); `contentHtml` is capped at 50,000 characters, `content` and `contentText` at 100,000, `summary` at 2,000; the state store remembers 5,000 items per feed; OPML adds at most 500 feeds per run.
- **Some sites refuse datacenter traffic** (Reddit and Craigslist among them) — see the FAQ.
- **Dates:** a feed that writes a date with no time zone is read as UTC.

### Sources, terms and attribution

- Everything is read from **public** feeds, exactly as their publishers serve them to any feed reader. No login, no cookie, no API key belonging to anyone.
- Feed content belongs to its publishers. Monitoring, alerting, analysis and linking are what this is built for; republishing is your call and your responsibility. Bylines are returned as published, as attribution.
- Not affiliated with any publisher named here; names are used only to describe which public feeds were used in testing.

### Our other Actors

Every Insight Solutions Actor is pay-per-result with no browser, no login and no API key, and every one of them returns free diagnostic rows instead of billing for failures. Prices are per 1,000 results.

**Video, audio & social**

- [YouTube Transcript API](https://apify.com/insight.solutions/youtube-transcript-api) — captions as timed segments, text, SRT or VTT, with language fallback and translation.
- [YouTube Comments API](https://apify.com/insight.solutions/youtube-comments-api) — comments and replies with likes, pinned and hearted flags, newest or top sort.
- [YouTube Channel API](https://apify.com/insight.solutions/youtube-channel-api) — a channel's videos, Shorts and live streams, plus YouTube search.
- [Podcast Search, Episodes & Charts API](https://apify.com/insight.solutions/podcast-api) — Apple Podcasts search, charts and full episode feeds.
- [Bluesky Scraper](https://apify.com/insight.solutions/bluesky-scraper) — profiles, posts, followers and follows from the public AT Protocol API.
- [Telegram Channel Scraper](https://apify.com/insight.solutions/telegram-channel-scraper) — posts, views and channel stats from public Telegram channels.
- [Substack Scraper](https://apify.com/insight.solutions/substack-scraper) — posts with full free text, comments and publication profiles.
- [Hacker News API](https://apify.com/insight.solutions/hacker-news-api) — stories, comments, users, front page and a structured "Who is hiring?" parser from the official HN APIs.
- [Discourse Forum API](https://apify.com/insight.solutions/discourse-forum-api) — topics, posts and categories from any Discourse community via its own JSON endpoints, usernames only.

**News, documents & the web**

- [Google News Search, Topics & Real Article URLs](https://apify.com/insight.solutions/google-news-api) — news search and topic feeds with the publisher's real URL decoded.
- [Website to Markdown — Content Extractor for LLMs & RAG](https://apify.com/insight.solutions/website-content-extractor) — any site as clean Markdown, text and heading-aware chunks.
- [Internet Archive API](https://apify.com/insight.solutions/internet-archive-api) — archive.org search, item metadata, files and reviews.
- [Wayback Machine Toolkit](https://apify.com/insight.solutions/wayback-toolkit) — archived URL inventories, snapshots and text diffs between dates.
- [Website Technology Detector](https://apify.com/insight.solutions/website-tech-detector) — the tech stack behind any site, with the evidence for each detection.
- [Domain Intelligence API](https://apify.com/insight.solutions/domain-intelligence-api) — DNS, RDAP registration, TLS certificate and HTTP facts in one row per domain.
- [SEO Page Audit](https://apify.com/insight.solutions/seo-page-audit) — sitemap crawl with on-page checks, structured data and broken-link reports.
- [Keyword Suggestions API](https://apify.com/insight.solutions/keyword-suggestions-api) — Google, YouTube, Bing, Amazon and eBay autocomplete with alphabet and question expansions.
- [Website Contact Extractor](https://apify.com/insight.solutions/website-contact-extractor) — emails, phone numbers and social profiles from any list of websites.
- [Web Search Results API](https://apify.com/insight.solutions/web-search-api) — Bing and DuckDuckGo organic results with snippets, no key, no browser.
- [Company Enrichment API](https://apify.com/insight.solutions/company-enrichment-api) — a domain in, a company profile out: firmographics, contacts, tech stack, DNS and hiring signal.
- [Company Dossier API](https://apify.com/insight.solutions/company-dossier-api) — one company in, twelve sections out: profile, tech, contacts, DNS, open roles, news, SEC filings, federal awards, recalls, YC batch and apps.
- [Press Releases API](https://apify.com/insight.solutions/press-releases-api) — GlobeNewswire and PR Newswire releases plus any newsroom feed, by keyword, company, ticker or subject.
- [Federal Register API](https://apify.com/insight.solutions/federal-register-api) — rules, proposed rules, notices and the Public Inspection desk with dockets, comment deadlines and CFR references.
- [Academic Papers Search API](https://apify.com/insight.solutions/academic-papers-api) — OpenAlex, Crossref, arXiv and PubMed in one row per paper: abstract, citations, open-access PDF, authors and venue.
- [Website Change Monitor](https://apify.com/insight.solutions/website-change-monitor) — watch any pages, diff the text between runs, get change rows with added/removed lines, keyword alerts and a webhook.
- [Wikipedia & Wikidata API](https://apify.com/insight.solutions/wikipedia-api) — article text, search, daily pageviews and Wikidata entity facts, any language edition.

**Business, finance & jobs**

- [Congress & Insider Trades API](https://apify.com/insight.solutions/congress-insider-trades-api) — STOCK Act periodic transaction reports and SEC Form 4 insider trades in one schema.
- [Federal Contracts, Grants & Lobbying API](https://apify.com/insight.solutions/federal-contracts-grants-api) — SAM.gov opportunities, USAspending awards, Grants.gov notices and Senate lobbying filings in one schema.
- [SEC EDGAR API](https://apify.com/insight.solutions/sec-edgar-api) — filings, XBRL financials and full-text search by ticker or CIK.
- [Clinical Trials & FDA API](https://apify.com/insight.solutions/clinical-trials-fda-api) — ClinicalTrials.gov studies plus openFDA recalls, labels, approvals, 510(k)s and adverse-event reports.
- [Product & Vehicle Recalls API](https://apify.com/insight.solutions/product-recalls-api) — CPSC, NHTSA, FDA and USDA recalls, vehicle complaints and ratings, plus a VIN decoder.
- [Y Combinator Companies, Batches & Founders](https://apify.com/insight.solutions/yc-companies-directory) — the YC directory with founders and social links, filterable by batch, industry and hiring status.
- [Career Site Jobs API](https://apify.com/insight.solutions/ats-jobs-api) — jobs straight from Greenhouse, Lever, Ashby, Workable and 10+ other ATS career sites.
- [New Job Postings Monitor](https://apify.com/insight.solutions/job-postings-monitor) — new, closed and changed postings on the career sites you watch.
- [Hiring Signals API — Open Roles & Hiring Surge by Company](https://apify.com/insight.solutions/hiring-signals-api) — one row per company per run: open roles, what opened and closed, department and seniority breakdowns, and a hiring-surge flag.
- [Remote Jobs API](https://apify.com/insight.solutions/remote-jobs-api) — RemoteOK, Remotive, We Work Remotely, Himalayas, Jobicy and more in one schema, deduplicated.
- [Shopify Products API](https://apify.com/insight.solutions/shopify-products-api) — any Shopify store's catalogue, variants, prices and stock signals.
- [Shopify Store Monitor](https://apify.com/insight.solutions/shopify-store-monitor) — price drops, sales, restocks, sell-outs and new products on any Shopify store, one row per change.
- [Public Tenders API](https://apify.com/insight.solutions/public-tenders-api) — EU TED, UK Find a Tender and Contracts Finder notices by keyword, CPV code, country, stage and deadline.
- [Nonprofit & IRS 990 Lookup API](https://apify.com/insight.solutions/nonprofit-990-api) — search US nonprofits and get EIN, NTEE code and multi-year Form 990 financials.
- [OpenStreetMap Places API](https://apify.com/insight.solutions/osm-places-api) — businesses and points of interest by category and area from OpenStreetMap: name, address, coordinates, website, phone, opening hours.

**Apps & games**

- [App Store & Google Play Reviews API](https://apify.com/insight.solutions/app-reviews-api) — reviews from both stores with ratings, versions and developer replies.
- [App Store Top Charts & App Search API](https://apify.com/insight.solutions/app-charts-api) — Apple top charts by country and genre, plus app search and details.
- [App Store Keyword Rank Tracker](https://apify.com/insight.solutions/app-store-keyword-rank-tracker) — where any app ranks for any keyword on the App Store and Google Play, with rank changes and ASO suggestions.
- [Steam Reviews API](https://apify.com/insight.solutions/steam-reviews-api) — Steam reviews with playtime, helpfulness and game details.
- [Steam Game Data API](https://apify.com/insight.solutions/steam-store-stats-api) — prices, tags, review scores, live player counts and top charts.

# Actor input Schema

## `feedUrls` (type: `array`):

RSS 2.0, RSS 1.0/RDF, Atom or JSON Feed addresses — news sites, blogs, YouTube channel feeds (`https://www.youtube.com/feeds/videos.xml?channel_id=…`), podcast feeds, GitHub release feeds (`…/releases.atom`), government newsrooms. A page URL that is not a feed is checked for the feed it advertises (`<link rel="alternate">`, one extra request); if it advertises none you get a free `not-a-feed` row.

## `opmlUrl` (type: `string`):

The address of an OPML subscription list (what every feed reader exports). Every `<outline xmlUrl="…">`, at any nesting depth, is added to the feed URLs above (duplicates once). At most 500 feeds from OPML per run.

## `opml` (type: `string`):

Or paste the OPML file itself here. Parsed exactly like the OPML file URL.

## `mode` (type: `string`):

`monitor` remembers every item it has shown in a named state store, so each run emits only items it has not seen before — schedule it hourly and you get only what is new. `snapshot` emits everything the feeds hold right now and touches no state.

## `firstRunBehavior` (type: `string`):

What monitor mode does the first time it sees a feed. `emit-all` emits the items the feed currently holds (charged like any item). `baseline-only` records their ids and emits nothing but the free `feed` row, so only items published after today are ever charged.

## `keywords` (type: `array`):

Keep an item only if any of these appears (case-insensitive) in its title, summary, body text or categories. The row's `matchedKeywords` says which. Leave empty to keep everything.

## `excludeKeywords` (type: `array`):

Drop an item if any of these appears in the same text. Excluded items are free.

## `authors` (type: `array`):

Keep an item only if its byline contains any of these (case-insensitive). An item with no byline does not pass an author filter.

## `categories` (type: `array`):

Keep an item only if one of its categories or tags equals any of these (case-insensitive, whole tag — `AI` does not match `Ukraine`).

## `sinceHours` (type: `integer`):

Keep only items published within this many hours. Items with no date pass. Leave empty for no window.

## `maxItemsPerFeed` (type: `integer`):

How many of each feed's items to consider per run — the first ones as the feed lists them, which is newest first in almost every feed. `0` means all. The `feed` row still reports the feed's full item count.

## `includeContent` (type: `boolean`):

Return `content` (Markdown-ish text with links kept), `contentText` (plain text) and `contentHtml` (raw, capped at 50,000 characters). Off: those three are null and rows are much smaller; `summary` and `wordCount` stay.

## `dedupe` (type: `boolean`):

When two feeds in one run carry the same item (same link once `utm_*` and fragments are stripped, or the same globally unique guid), return it once — from the first feed in your list — with the others in `alsoIn`. One row, one charge.

## `webhookUrl` (type: `string`):

Optional. After a run that found at least one new item, one JSON POST goes here with the new items (fitted to 256 KB) and a link to the dataset. Free. The URL is treated as a secret: only its origin is ever logged.

## `stateStoreName` (type: `string`):

The named key-value store monitor mode keeps its memory in: per feed, the ids of the last 5,000 items shown, plus the ETag and Last-Modified sent back on the next run. Use a different name for each independent schedule. Letters, digits and hyphens.

## `maxConcurrency` (type: `integer`):

How many feeds are read at once. Whatever this says, requests to one host start at least one second apart.

## `maxRunSecs` (type: `integer`):

Wall-clock budget for the whole run. Feeds are read until 90% of it (at most 20 s short of it) is used; the rest is kept for writing rows and state. A feed not reached in time gets a free `deadline` row and its state is untouched.

## `proxyConfiguration` (type: `object`):

Every feed captured while building this Actor answered Apify's datacenter proxy, so that is the default and its cost is inside the per-item price. Some sites (Reddit, Craigslist) refuse datacenter traffic; they come back as free `blocked` rows. Sessions rotate automatically when an exit IP is refused.

## Actor input object example

```json
{
  "feedUrls": [
    "https://feeds.bbci.co.uk/news/rss.xml",
    "https://www.youtube.com/feeds/videos.xml?channel_id=UCsBjURrPoezykLs9EqgamOA",
    "https://podcast.darknetdiaries.com/"
  ],
  "opmlUrl": "",
  "opml": "",
  "mode": "snapshot",
  "firstRunBehavior": "emit-all",
  "keywords": [
    "artificial intelligence",
    "release"
  ],
  "excludeKeywords": [],
  "authors": [],
  "categories": [],
  "maxItemsPerFeed": 25,
  "includeContent": true,
  "dedupe": true,
  "webhookUrl": "",
  "stateStoreName": "rss-feed-monitor-state",
  "maxConcurrency": 5,
  "maxRunSecs": 240,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One row per new feed item — title, link, publish time, author, summary, body, categories, media and podcast fields — plus a free summary row per feed and a free diagnostic row for anything that could not be read. Delivered as JSON items in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "feedUrls": [
        "https://feeds.bbci.co.uk/news/rss.xml",
        "https://news.ycombinator.com/rss",
        "https://github.com/apify/crawlee/releases.atom",
        "https://www.theverge.com/rss/index.xml"
    ],
    "opmlUrl": "",
    "opml": "",
    "mode": "snapshot",
    "firstRunBehavior": "emit-all",
    "keywords": [],
    "excludeKeywords": [],
    "authors": [],
    "categories": [],
    "maxItemsPerFeed": 25,
    "includeContent": true,
    "dedupe": true,
    "webhookUrl": "",
    "stateStoreName": "rss-feed-monitor-state",
    "maxConcurrency": 5,
    "maxRunSecs": 240,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("insight.solutions/rss-feed-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "feedUrls": [
        "https://feeds.bbci.co.uk/news/rss.xml",
        "https://news.ycombinator.com/rss",
        "https://github.com/apify/crawlee/releases.atom",
        "https://www.theverge.com/rss/index.xml",
    ],
    "opmlUrl": "",
    "opml": "",
    "mode": "snapshot",
    "firstRunBehavior": "emit-all",
    "keywords": [],
    "excludeKeywords": [],
    "authors": [],
    "categories": [],
    "maxItemsPerFeed": 25,
    "includeContent": True,
    "dedupe": True,
    "webhookUrl": "",
    "stateStoreName": "rss-feed-monitor-state",
    "maxConcurrency": 5,
    "maxRunSecs": 240,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("insight.solutions/rss-feed-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "feedUrls": [
    "https://feeds.bbci.co.uk/news/rss.xml",
    "https://news.ycombinator.com/rss",
    "https://github.com/apify/crawlee/releases.atom",
    "https://www.theverge.com/rss/index.xml"
  ],
  "opmlUrl": "",
  "opml": "",
  "mode": "snapshot",
  "firstRunBehavior": "emit-all",
  "keywords": [],
  "excludeKeywords": [],
  "authors": [],
  "categories": [],
  "maxItemsPerFeed": 25,
  "includeContent": true,
  "dedupe": true,
  "webhookUrl": "",
  "stateStoreName": "rss-feed-monitor-state",
  "maxConcurrency": 5,
  "maxRunSecs": 240,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call insight.solutions/rss-feed-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,insight.solutions/rss-feed-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gmOi4u5xo2GR2AStj/builds/hYIijGzir4uqflnUE/openapi.json
