# RSS & Atom to Markdown — JSON + RAG Chunks (`ingenious_quip_bxq/rss-atom-to-markdown`) Actor

Parse RSS 2.0 and Atom feeds into structured JSON with Markdown body text for RAG and LLM pipelines. Optional feed discovery from a page, content hash, and heading-aware chunks. Failed feeds are reported, not fatal. 256 MB default.

- **URL**: https://apify.com/ingenious\_quip\_bxq/rss-atom-to-markdown.md
- **Developed by:** [新世紀書僮](https://apify.com/ingenious_quip_bxq) (community)
- **Categories:** AI, Developer tools, Open source
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 feed items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## RSS & Atom to Markdown — JSON + RAG Chunks

**Turn RSS 2.0 and Atom feeds into structured JSON with Markdown body text, ready for RAG and LLM pipelines.**
Paste feed URLs (or a page URL to discover feeds from). Each entry becomes one dataset row with title, link, published date, authors, categories, summary, `contentMarkdown`, word count and a content hash. Optional heading-aware RAG chunks. Broken feeds are reported, not fatal. Default memory: 256 MB. No browser, no AI keys.

### What you get

- 📡 **RSS 2.0, Atom, JSON Feed** — parsed with open-source `feedparser`
- 📝 **Markdown body** — HTML in `content:encoded` / Atom `content` / summary converted to Markdown
- 🔎 **Optional feed discovery** — scan a page for `<link rel="alternate" type="application/rss+xml|atom+xml">`, plus common `/feed` paths
- 🧩 **Optional RAG chunks** — heading-aware chunks with token estimate (same idea as our document Actors)
- 🧮 **Filters** — max items, published-after date, include/exclude link regex, per-feed cap
- 🧯 **Broken feeds don't break the run** — each feed/page gets a status line in `FEED_REPORT`
- 💾 **Light** — HTTP only, 256 MB default

### Measured results

Local smoke + private Apify cloud (2026-09-30 Asia/Taipei, build **0.1.1**, 256 MB). Platform `$` from settled `usageTotalUsd` (≥2 min after finish).

| Test | Result |
|---|---|
| NASA news RSS (local) | 10 items in **0.3 s**, peak **88 MB**; Markdown from `content:encoded` |
| xkcd Atom + Mozilla Blog Atom (local) | 6 items in **0.3 s**, peak **86 MB** |
| NASA homepage discovery + Python Insider (local) | 8 items from 2 feeds in **0.7 s**, peak **94 MB** |
| W3C news + RAG chunks (local) | 5 items + 6 chunks in **0.1 s**, peak **84 MB** |
| Cloud smoke `TS6Txm6iVMXBPvnPI` (NASA, 10 items) | **SUCCEEDED** ~4.4 s wall; `usageTotalUsd` ≈ **$0.00027** |
| Cloud bench `JLRaUu51k2gsAAfBM` (20 public feeds → **1000** items) | **SUCCEEDED** 120 s wall; peak RSS **153 MB**; CU 0.00835; settled **$0.006855** (≈ **$0.00000686 / item**) |
| Cloud chunks `PiKjKVJTMTTrus1n4` (300 items + 300 chunks) | **SUCCEEDED** 8.3 s; peak RSS **119 MB**; settled **$0.003275** (≈ **$0.0000109 / charged item**) |

### Use cases

- **Monitor blogs and government feeds** for RAG / knowledge bases
- **Normalize mixed RSS + Atom sources** into one JSON schema with Markdown
- **Chunk feed content** for embedding pipelines without a separate splitter
- **Discover a site's feed URL** from its homepage, then parse it in the same run

### How to use

1. Add feed URLs in **Feed URLs**, and/or page URLs in **Page URLs (discover feeds)**.
2. Optional: set **Max items**, a **Published on or after** date, or **Output** = RAG chunks.
3. Click **Start**. Items appear in the **Dataset**; `FEED_REPORT` and `OUTPUT` are in the **Key-value store**.

#### Input example

```json
{
  "feedUrls": [
    { "url": "https://www.nasa.gov/rss/dyn/breaking_news.rss" }
  ],
  "maxItems": 20,
  "outputFormat": "items"
}
```

#### Output example (one dataset item per feed entry)

```json
{
  "kind": "item",
  "status": "ok",
  "title": "Example headline",
  "link": "https://www.example.gov/news/123",
  "published": "2026-09-30T12:00:00+00:00",
  "authors": ["Press Office"],
  "categories": ["News"],
  "summary": "Short plain-text or Markdown summary…",
  "contentMarkdown": "# Example headline\n\nFull body as Markdown…",
  "contentSource": "content",
  "contentHash": "sha256…",
  "wordCount": 420,
  "feedUrl": "https://www.example.gov/feed.xml",
  "feedTitle": "Example Feed",
  "feedFormat": "rss"
}
```

#### Key-value store records

| Key | Content |
|---|---|
| `OUTPUT` | Run summary: items/chunks saved, feeds read/failed, duplicates, duration, peak memory |
| `FEED_REPORT` | One entry per feed or discovery page: `status` (`ok`, `not_found`, `error`, `skipped`), HTTP status, format, entry counts, error |

### Pricing

Pay per event (validated on private cloud benches; numbers in **Measured results** above):

| Event | Price |
|---|---|
| Feed item saved (primary) | **$0.0005** per entry (= $0.50 per 1,000 items) |
| Actor start | Apify default ($0.00005 per GB of run memory) |

RAG chunk rows are not billed as a separate PPE event (dataset storage still applies on the platform). Failed feeds and duplicate items are never charged. Measured platform cost ≈ **$0.000007 / item** at 1,000 items (items-only).

### Known limits

- Does **not** fetch the full article page behind each item link (feed body / summary only). Full-page extraction may come later as an optional paid event.
- Does not execute JavaScript; feeds must be plain HTTP(S) XML/JSON.
- Malformed feeds: `feedparser` is tolerant, but severely broken XML may yield zero entries (`bozo` noted in `FEED_REPORT`).
- JSON Feed support depends on `feedparser`; treat as best-effort until covered in measured tests.
- Very large feeds: use **Max items** / **Max items per feed** to cap cost and memory.

### FAQ

**Can I pass a homepage instead of a feed URL?** Yes — put it in **Page URLs**. The Actor looks for `<link rel="alternate">` feed links and, if none are found, tries common paths like `/feed` and `/atom.xml`.

**Are failed feeds charged?** No. Only successfully saved feed entries are charged.

### License & source code

This Actor is open source under the **GNU Affero General Public License v3.0 (AGPL-3.0)** — see `LICENSE`. The full source code is public: https://github.com/xbox002000/rss-atom-to-markdown

Third-party notices: `NOTICE`. Changelog: `CHANGELOG.md`.

# Changelog

This Actor's version history is a separate document: https://apify.com/ingenious\_quip\_bxq/rss-atom-to-markdown/changelog.md

# Actor input Schema

## `feedUrls` (type: `array`):

Direct RSS 2.0 or Atom feed URLs (also accepts JSON Feed). Paste a list or upload a text file of URLs.

## `pageUrls` (type: `array`):

Optional. HTML pages to scan for `<link rel="alternate" type="application/rss+xml|atom+xml">` (and common `/feed`, `/rss`, `/atom.xml` paths). Discovered feeds are fetched in the same run.

## `maxItems` (type: `integer`):

Stop after saving this many feed entries (0 = no limit). Also caps your cost: you pay only for items saved.

## `maxItemsPerFeed` (type: `integer`):

Optional cap per feed (0 = no cap), so one large feed cannot use up the whole Max items budget.

## `publishedAfter` (type: `string`):

Keep only entries whose published/updated date is on or after this date (YYYY-MM-DD). Entries without a date are kept when *Keep undated items* is on.

## `keepUndatedItems` (type: `boolean`):

When a date filter is set, keep entries that have no published/updated date.

## `includeUrlPatterns` (type: `array`):

Keep only entries whose link matches at least one of these regular expressions.

## `excludeUrlPatterns` (type: `array`):

Drop entries whose link matches any of these regular expressions.

## `outputFormat` (type: `string`):

`items`: one dataset row per feed entry (title, link, published, authors, summary, contentMarkdown, …). `chunks`: RAG-ready chunks (one row per chunk) plus a per-entry summary row. `items_and_chunks`: both.

## `preferFullContent` (type: `boolean`):

When the feed provides both a summary and full content (Atom `content`, RSS `content:encoded`), use the full body for Markdown. Off = use summary when present.

## `chunkSize` (type: `integer`):

Maximum chunk size (in the unit below). Chunks break at headings and paragraphs.

## `chunkOverlap` (type: `integer`):

How much trailing context from the previous chunk is repeated at the start of the next one (same unit as chunk size).

## `chunkUnit` (type: `string`):

`tokens` uses a fast approximation (≈4 characters per token); `characters` is exact.

## `tryCommonFeedPaths` (type: `boolean`):

When discovering from a page, also try `/feed`, `/rss`, `/atom.xml`, `/feed.xml`, `/index.xml` if no `<link rel="alternate">` is found.

## `requestTimeoutSecs` (type: `integer`):

Timeout for each HTTP request.

## `maxRetries` (type: `integer`):

Retries for timeouts, network errors, HTTP 429 and 5xx (exponential backoff).

## `maxConcurrency` (type: `integer`):

Parallel feed/page fetches.

## `userAgent` (type: `string`):

Optional custom User-Agent header.

## `proxyConfiguration` (type: `object`):

Optional. Use a proxy if a site blocks data-center IPs. Not needed for most public feeds.

## Actor input object example

```json
{
  "feedUrls": [
    {
      "url": "https://www.nasa.gov/rss/dyn/breaking_news.rss"
    }
  ],
  "maxItems": 50,
  "maxItemsPerFeed": 0,
  "keepUndatedItems": true,
  "outputFormat": "items",
  "preferFullContent": true,
  "chunkSize": 1000,
  "chunkOverlap": 150,
  "chunkUnit": "tokens",
  "tryCommonFeedPaths": true,
  "requestTimeoutSecs": 30,
  "maxRetries": 3,
  "maxConcurrency": 4,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

Default dataset items. Feed entries include: title, link, published, updated, authors, categories, summary, contentMarkdown, contentHash, wordCount, feedUrl, feedTitle, feedFormat, status. Chunk rows (when enabled) include chunkIndex, text, headingPath, tokenEstimate, parent link/title. Views: overview, chunks.

## `feedReport` (type: `string`):

Key-value store record FEED\_REPORT (JSON array): one entry per feed or discovery page with feedUrl/pageUrl, status (ok / error / skipped / not\_found), httpStatus, feedFormat, item counts, error message. Broken feeds are reported here instead of failing the run.

## `summary` (type: `string`):

Key-value store record OUTPUT (JSON): itemsOutput, chunksOutput, feedsRead, feedErrors, duplicatesSkipped, filteredOut, maxItemsReached, durationSecs, peakMemoryMb.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "feedUrls": [
        {
            "url": "https://www.nasa.gov/rss/dyn/breaking_news.rss"
        }
    ],
    "maxItems": 50,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("ingenious_quip_bxq/rss-atom-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "feedUrls": [{ "url": "https://www.nasa.gov/rss/dyn/breaking_news.rss" }],
    "maxItems": 50,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("ingenious_quip_bxq/rss-atom-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "feedUrls": [
    {
      "url": "https://www.nasa.gov/rss/dyn/breaking_news.rss"
    }
  ],
  "maxItems": 50,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call ingenious_quip_bxq/rss-atom-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ingenious_quip_bxq/rss-atom-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ufD8eD6O26WgcV8BB/builds/dtyQDi9wNYW4LJjTY/openapi.json
