# Website Page to RSS Feed (`automation-lab/website-page-to-rss-feed`) Actor

Generate RSS 2.0 XML snapshots from public article-list pages and refresh a stable public feed in a named store.

- **URL**: https://apify.com/automation-lab/website-page-to-rss-feed.md
- **Developed by:** [Automation Lab](https://apify.com/automation-lab) (community)
- **Categories:** Automation
- **Stats:** 2 total users, 1 monthly users, 50.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.96 / 1,000 item extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Page to RSS Feed

Convert public article-list pages into RSS 2.0 XML snapshots. This RSS feed generator reads semantic article cards on supplied pages, collects titles and canonical-looking links, and includes dates and summaries when the page supplies them. Each output row includes the entire XML document and the key of a stored XML file. It does **not** manufacture a continuous live feed or discover every article across a website.

### Who is this for?

Newsletter editors can capture posts from an article-list page. Researchers can subscribe to a page without manually copying every article link. Automation teams can schedule an Actor run and reuse a named feed store so a subscriber can retrieve the latest snapshot at a stable key.

### Why use this Actor?

Unlike a reader for existing RSS or Atom feeds, this Actor starts with HTML. Unlike a general web crawler, it creates an RSS document directly, with each article link as a stable RSS GUID. The optional named key-value store overwrites the same source's XML key on every scheduled run.

### What is extracted?

| Field | Meaning |
| --- | --- |
| `sourceUrl` | Final article-list page URL after redirects |
| `title` | HTML page title, used as feed title |
| `entries` | Article title, same-origin URL, optional RFC 822 date and summary |
| `rssXml` | Complete RSS 2.0 XML snapshot |
| `feedStoreId`, `feedKey`, `feedUrl` | Store ID, deterministic XML key, and optional public RSS link on named cloud stores |
| `itemCount`, `scrapedAt` | Number of included articles and snapshot timestamp |

### Getting started

1. Supply one or more public article-list pages, not individual article pages.
2. Set `maxItems` to the desired maximum **per page** (1–200).
3. Run the Actor and inspect the default dataset's `overview` view.
4. Download XML from the run key-value store using `feedStoreId` and `feedKey`, or read `rssXml` directly from the dataset.
5. For recurring refreshes, set `feedStoreName` and schedule identical Actor input. The same page resolves to the same XML key within that store. On Apify cloud, providing a store name automatically makes the entire named store readable by anyone with its ID; use a unique, dedicated store name and share `feedUrl` only when its content is suitable for publication. An existing private store containing records is rejected without changing access.

### Input example

```json
{"startUrls":[{"url":"https://blog.python.org/"}],"maxItems":4,"feedStoreName":"python-blog-rss"}
```

### Output example

A local run on `https://blog.python.org/` produced four entries; one had title `The Python documentation is now available in Persian`, link `https://blog.python.org/2026/09/the-python-documentation-is-now-available-in-persian`, and date `Wed, 23 Sep 2026 00:00:00 GMT`. The row also includes `rssXml`, `feedKey`, and `feedStoreId`.

### How much does it cost to generate RSS feeds from article pages?

Pay per event: a one-time `start` event and one `item` event per **non-empty RSS feed snapshot**, not per article. Article counts inside a feed do not add a separate event charge. The BRONZE tier currently charges $0.00005 to start plus $0.0016 per produced feed snapshot: the total for N feeds in one run is the one-time start price plus N times the per-feed price. For example, a two-feed run includes one start and two feed events. FREE is $0.00184/feed, SILVER $0.001248/feed, and GOLD/PLATINUM/DIAMOND $0.00096/feed; tier eligibility depends on account-wide Store spend. Check the live pricing panel before scheduling; estimates can change with platform policy. Failed or unsupported pages yield no feed event. Prices shown are estimates, not guaranteed invoices or payouts; refunds, fraud, disputes, taxes, corrections and contractual clawbacks can affect final platform settlement.

### Scheduling a stable feed

Apify Schedules can invoke this Actor repeatedly with the same input. Set `feedStoreName` to reuse a named key-value store. The `feedKey` is based on the final source URL; a changed redirect destination changes the key. On Apify cloud, the Actor grants **anyone with the store ID read access** to a named store and returns `feedUrl` for token-free subscription. Use a dedicated, unique store name: all records in this store are publicly readable by anyone with its ID. An existing private populated store is rejected before writing or changing access; a new or empty private store can be used, and subsequent scheduled refreshes reuse the public store. Never put private data in a named feed store. A run-scoped store is not made public.

### Integration workflows

Use the dataset in Make or Zapier to forward new snapshots to a downstream feed host or compare GUIDs between scheduled runs. Use `feedUrl` from a named cloud run directly in a feed reader, or an authenticated API client for a run-scoped snapshot. This Actor itself does not schedule runs, send alerts, deduplicate across runs, host a domain, or create a feed-reader subscription.

### Supported pages and limits

Supports anonymously accessible HTML article-list pages with `<article>`, schema.org Article, `post-card`, or `article-card` containers and links inside them. Links must be on the same origin as the page. It does not render JavaScript, pass login challenges, parse RSS/Atom as input, crawl pagination, scrape article bodies, or fall back to arbitrary navigation links. Missing source dates and summaries are omitted, not invented. One source page can contribute up to 200 articles; one run accepts up to ten pages.

### Failures and troubleshooting

A page with no usable semantic cards logs a warning and produces no row; if **all** sources are empty, the run fails. If your blog returns a JavaScript shell, use a server-rendered listing or another tool. If a URL redirects to a private network destination or a page is not HTML, the run fails rather than fetching it. Confirm that the URL represents a listing and that its HTML contains article containers before retrying.

### API usage

Use the normal Apify Actor run API with `automation-lab/website-page-to-rss-feed` and JSON input. For example:

```bash
curl -X POST 'https://api.apify.com/v2/acts/automation-lab~website-page-to-rss-feed/runs?token=YOUR_TOKEN' -H 'Content-Type: application/json' -d '{"startUrls":[{"url":"https://blog.python.org/"}],"maxItems":4}'
```

In JavaScript, call the Actor with the Apify client:

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('automation-lab/website-page-to-rss-feed').call({ startUrls: [{ url: 'https://blog.python.org/' }], maxItems: 4 });
console.log(await client.dataset(run.defaultDatasetId).listItems());
```

In Python, use the ApifyClient:

```python
from apify_client import ApifyClient
import os
client = ApifyClient(os.environ['APIFY_TOKEN'])
run = client.actor('automation-lab/website-page-to-rss-feed').call(run_input={'startUrls': [{'url': 'https://blog.python.org/'}], 'maxItems': 4})
print(client.dataset(run['defaultDatasetId']).list_items().items)
```

Retrieve the run's dataset or read the XML record in the key-value store.

### MCP use

For Claude Code, configure `claude mcp add --transport http apify "https://mcp.apify.com?tools=automation-lab/website-page-to-rss-feed"`. For Claude Desktop, Cursor, or VS Code supporting JSON MCP configuration, add `{ "mcpServers": { "apify": { "url": "https://mcp.apify.com?tools=automation-lab/website-page-to-rss-feed" } } }`. Authenticate with your Apify account in your MCP client. Example prompts: “Generate an RSS snapshot from the Python blog listing and show its XML key”; “Refresh the Cloudflare blog feed into my named store and tell me the number of articles.”

### Legality and responsible use

Only request public pages you are authorized to access. Respect site terms and load; schedule at a reasonable interval. RSS titles and summaries are source content. Check copyright and redistribution rules before republishing a full feed.

### FAQ

**Does it create a native RSS feed on the source site?** No. It writes a point-in-time XML file to Apify storage. A schedule can refresh that file.

**Why is the feed empty?** The listing likely lacks semantic article/card containers or same-origin article links. An empty source is not a valid proof of extraction.

**Can a feed reader use the named store?** Yes: a named cloud store returns `feedUrl`, a token-free XML URL. This opts the whole named store into anyone-with-ID read access. Run-scoped storage remains private.

### Related Actors

[RSS & Atom Feed Parser](https://apify.com/automation-lab/rss-atom-feed-parser) parses feeds that already exist. [Website Content Crawler](https://apify.com/automation-lab/website-content-crawler) crawls page text rather than producing RSS XML. Choose the reader for an existing feed and this Actor for a supported HTML article-list page.

# Changelog

This Actor's version history is a separate document: https://apify.com/automation-lab/website-page-to-rss-feed/changelog.md

# Actor input Schema

## `startUrls` (type: `array`):

Up to 10 public HTTP(S) pages containing article cards; not individual article URLs or sites requiring login.

## `maxItems` (type: `integer`):

Maximum distinct article entries to include in each RSS snapshot.

## `feedStoreName` (type: `string`):

Reuse this named Apify key-value store across scheduled runs. Leave blank for a run-scoped snapshot. On Apify cloud, providing a name automatically makes the entire named store readable by anyone with its ID. Use a unique, dedicated name; an existing private store containing records is rejected.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://blog.cloudflare.com/"
    }
  ],
  "maxItems": 10
}
```

# Actor output Schema

## `overview` (type: `string`):

Dataset rows containing XML snapshots and extracted article entries.

## `files` (type: `string`):

Keys in the run's default key-value store; named stores have separate IDs in dataset rows.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://blog.cloudflare.com/"
        }
    ],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("automation-lab/website-page-to-rss-feed").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://blog.cloudflare.com/" }],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("automation-lab/website-page-to-rss-feed").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://blog.cloudflare.com/"
    }
  ],
  "maxItems": 10
}' |
apify call automation-lab/website-page-to-rss-feed --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automation-lab/website-page-to-rss-feed"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/9GWSe1HAyKHergM9y/builds/n6DfzvwhRQy7YRT0x/openapi.json
