# RSS and Atom Feed to JSON, with Feed Auto Discovery (`pistachio_implementation/rss-atom-feed-to-json`) Actor

Turn any RSS, Atom or JSON Feed, or any website URL, into clean JSON items: title, link, date, authors, categories, summary, full text, image and enclosures. Finds the feed for you from a site's home page. For AI agents, RAG pipelines, news monitoring and alerts.

- **URL**: https://apify.com/pistachio\_implementation/rss-atom-feed-to-json.md
- **Developed by:** [Hay Equipos](https://apify.com/pistachio_implementation) (community)
- **Categories:** AI, Developer tools, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.50 / 1,000 feed item saveds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## RSS and Atom Feed to JSON, with Feed Auto Discovery

Give the actor any feed address, or just a website or blog address, and get every item back as clean, consistent JSON: title, link, publish date, authors, categories, summary, full text, image and attachments. It reads RSS 2.0, RSS 1.0 (RDF), Atom and JSON Feed, and returns all four in one schema, so your code or your AI agent never has to care which format a site uses.

Do not know the feed address? Paste the home page. The actor finds the feed the site declares in its own page header, and if there is none it tries the usual addresses such as /feed, /rss.xml and /atom.xml.

No login, no browser, no proxies. robots.txt is respected, including rules that name Apify, and each site gets at most one request per second.

### What you can use it for

- **AI agents and RAG pipelines:** pull the latest posts from any blog, newsroom or changelog as plain text, ready to summarize, embed or index.
- **News and content monitoring:** run it on a schedule with "Only items published after" and send only new items to Slack, email, a sheet or a webhook.
- **Competitor tracking:** follow competitors' blogs, product update pages and press rooms in one run.
- **Newsletters and digests:** Substack, Ghost, WordPress, Medium publications and most news sites publish feeds this actor reads.
- **Podcast and media feeds:** audio and video attachments come back in `enclosures` with type and size.

### Input

| Field | What it does | Default |
|---|---|---|
| Feed or website URLs | Feed addresses or ordinary website addresses, one per line | required |
| Find the feed when a website URL is given | Auto discovery on or off | on |
| Maximum items per feed | Items taken from each feed, in the publisher's order (newest first on almost every site) | 50 |
| Maximum items in total | Stop after this many items across all feeds | 1,000 |
| Only items published after | A date such as 2026-09-01; older items are skipped and not charged | none |
| Include item text | Plain text of each item | on |
| Maximum text length | Cut the text to this many characters (0 for no limit) | 20,000 |
| Include item HTML | The publisher's HTML as well | off |

Example input:

```json
{
  "urls": [
    "https://blog.cloudflare.com",
    "https://github.blog/feed/",
    "https://www.jsonfeed.org/feed.json"
  ],
  "maxItemsPerFeed": 20,
  "publishedAfter": "2026-09-01"
}
```

### Output

One row per feed item. The same item is never saved twice in a run.

```json
{
  "input": "https://blog.cloudflare.com",
  "feedUrl": "https://blog.cloudflare.com/rss/",
  "feedTitle": "The Cloudflare Blog",
  "feedLink": "https://blog.cloudflare.com",
  "feedFormat": "rss2",
  "title": "Example post title",
  "url": "https://blog.cloudflare.com/example-post/",
  "id": "abc123",
  "authors": ["Jane Roe"],
  "publishedAt": "2026-09-25T14:00:00.000Z",
  "updatedAt": null,
  "categories": ["AI", "Developers"],
  "summary": "The first lines of the post...",
  "hasFullContent": true,
  "wordCount": 1450,
  "imageUrl": "https://blog.cloudflare.com/content/images/example.png",
  "enclosures": [],
  "text": "The full post as plain text...",
  "scrapedAt": "2026-09-27T07:18:37.239Z"
}
```

- `feedFormat` is `rss2`, `rss1`, `atom` or `jsonfeed`.
- `hasFullContent` tells you whether the publisher put the whole article in the feed or only a summary. Many sites publish only a summary; the actor never visits the article pages themselves.
- Dates are ISO 8601 in UTC. Entities and character sets (UTF 8, ISO 8859 1, Windows 1252 and others) are decoded for you.
- Inputs that give nothing (no feed found, the site blocked the request, robots.txt disallows it) are listed in `RUN_SUMMARY` in the run's key value store and cost nothing.

### Pricing

Pay per event. No start fee, no subscription, no platform usage charged on top.

| Event | Price |
|---|---|
| Feed item saved | $0.0005 (50 cents per 1,000 items) |

Example: an agent that checks 10 blogs for today's posts pays only for the handful of new items it gets back, often less than one cent. Failed feeds, skipped old items and duplicates are free. Set a maximum charge per run in Apify and the actor stops cleanly when it is reached.

### Limits

- A feed carries only what the publisher puts in it, usually the latest 10 to 50 items. This actor reads feeds; it does not crawl a site's archive. Run it on a schedule to build a history.
- Sites that block cloud servers, or whose robots.txt disallows automated reading of the feed (GitHub's release feeds are one example), return a free error entry instead of items.
- Feeds larger than 10 MB are skipped.
- Discovery tries the feeds the page declares and at most 8 addresses in total per input.

### FAQ

**Does it work with Substack, WordPress, Ghost and Medium?** Yes. They all publish standard feeds. Paste the publication's home page and the actor finds the feed.

**Can I get only new items on a schedule?** Yes. Set "Only items published after" to the date of your last run, or keep the output of each run and compare by `id` or `url`.

**Why is `text` short for some items?** The publisher only puts a summary in the feed. `hasFullContent` is false in that case.

**Can an AI agent call it?** Yes. It is a single job with one required input, pay per event pricing and no start fee, which suits agents calling it through the Apify API or MCP server.

**Is this legal?** Feeds are published so that software can read them. The actor reads only feeds, respects robots.txt and rate limits, and returns what the publisher chose to publish. You are responsible for how you use the content, including copyright in the full text.

# Actor input Schema

## `urls` (type: `array`):

One per line. A feed address (https://github.blog/feed/) or any website or blog address (https://blog.cloudflare.com). For a website the actor finds its feed automatically.

## `discoverFeeds` (type: `boolean`):

Look for the feed in the page's own feed links, then at common addresses such as /feed and /rss.xml. Switch off to accept feed URLs only.

## `maxItemsPerFeed` (type: `integer`):

Newest first as the publisher orders them. Most feeds carry 10 to 50 items.

## `maxItems` (type: `integer`):

Stop after saving this many items across all feeds.

## `publishedAfter` (type: `string`):

Optional date such as 2026-09-01. Older items are skipped and not charged. Useful for scheduled runs that only want new items.

## `includeContent` (type: `boolean`):

Add the item body as plain text (full text when the feed carries it, otherwise the summary).

## `maxContentChars` (type: `integer`):

Cut the text field to this many characters. 0 means no limit.

## `includeHtml` (type: `boolean`):

Also add the item body as the publisher's HTML.

## Actor input object example

```json
{
  "urls": [
    "https://blog.cloudflare.com",
    "https://github.blog/feed/"
  ],
  "discoverFeeds": true,
  "maxItemsPerFeed": 50,
  "maxItems": 1000,
  "includeContent": true,
  "maxContentChars": 20000,
  "includeHtml": false
}
```

# Actor output Schema

## `results` (type: `string`):

All rows the run saved to the default dataset.

## `summary` (type: `string`):

The RUN\_SUMMARY record: counts and problems for the whole run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://blog.cloudflare.com",
        "https://github.blog/feed/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pistachio_implementation/rss-atom-feed-to-json").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://blog.cloudflare.com",
        "https://github.blog/feed/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("pistachio_implementation/rss-atom-feed-to-json").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://blog.cloudflare.com",
    "https://github.blog/feed/"
  ]
}' |
apify call pistachio_implementation/rss-atom-feed-to-json --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pistachio_implementation/rss-atom-feed-to-json"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/2D7TsFJTN1063OHYg/builds/fvPK0L32KeVoAdowQ/openapi.json
