RSS & Atom to Markdown — JSON + RAG Chunks avatar

RSS & Atom to Markdown — JSON + RAG Chunks

Pricing

from $0.50 / 1,000 feed items

Go to Apify Store
RSS & Atom to Markdown — JSON + RAG Chunks

RSS & Atom to Markdown — JSON + RAG Chunks

Parse RSS 2.0 and Atom feeds into structured JSON with Markdown body text for RAG and LLM pipelines. Optional feed discovery from a page, content hash, and heading-aware chunks. Failed feeds are reported, not fatal. 256 MB default.

Pricing

from $0.50 / 1,000 feed items

Rating

0.0

(0)

Developer

新世紀書僮

新世紀書僮

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Turn RSS 2.0 and Atom feeds into structured JSON with Markdown body text, ready for RAG and LLM pipelines. Paste feed URLs (or a page URL to discover feeds from). Each entry becomes one dataset row with title, link, published date, authors, categories, summary, contentMarkdown, word count and a content hash. Optional heading-aware RAG chunks. Broken feeds are reported, not fatal. Default memory: 256 MB. No browser, no AI keys.

What you get

  • 📡 RSS 2.0, Atom, JSON Feed — parsed with open-source feedparser
  • 📝 Markdown body — HTML in content:encoded / Atom content / summary converted to Markdown
  • 🔎 Optional feed discovery — scan a page for <link rel="alternate" type="application/rss+xml|atom+xml">, plus common /feed paths
  • 🧩 Optional RAG chunks — heading-aware chunks with token estimate (same idea as our document Actors)
  • 🧮 Filters — max items, published-after date, include/exclude link regex, per-feed cap
  • 🧯 Broken feeds don't break the run — each feed/page gets a status line in FEED_REPORT
  • 💾 Light — HTTP only, 256 MB default

Measured results

Local smoke + private Apify cloud (2026-09-30 Asia/Taipei, build 0.1.1, 256 MB). Platform $ from settled usageTotalUsd (≥2 min after finish).

TestResult
NASA news RSS (local)10 items in 0.3 s, peak 88 MB; Markdown from content:encoded
xkcd Atom + Mozilla Blog Atom (local)6 items in 0.3 s, peak 86 MB
NASA homepage discovery + Python Insider (local)8 items from 2 feeds in 0.7 s, peak 94 MB
W3C news + RAG chunks (local)5 items + 6 chunks in 0.1 s, peak 84 MB
Cloud smoke TS6Txm6iVMXBPvnPI (NASA, 10 items)SUCCEEDED ~4.4 s wall; usageTotalUsd ≈ $0.00027
Cloud bench JLRaUu51k2gsAAfBM (20 public feeds → 1000 items)SUCCEEDED 120 s wall; peak RSS 153 MB; CU 0.00835; settled $0.006855 (≈ $0.00000686 / item)
Cloud chunks PiKjKVJTMTTrus1n4 (300 items + 300 chunks)SUCCEEDED 8.3 s; peak RSS 119 MB; settled $0.003275 (≈ $0.0000109 / charged item)

Use cases

  • Monitor blogs and government feeds for RAG / knowledge bases
  • Normalize mixed RSS + Atom sources into one JSON schema with Markdown
  • Chunk feed content for embedding pipelines without a separate splitter
  • Discover a site's feed URL from its homepage, then parse it in the same run

How to use

  1. Add feed URLs in Feed URLs, and/or page URLs in Page URLs (discover feeds).
  2. Optional: set Max items, a Published on or after date, or Output = RAG chunks.
  3. Click Start. Items appear in the Dataset; FEED_REPORT and OUTPUT are in the Key-value store.

Input example

{
"feedUrls": [
{ "url": "https://www.nasa.gov/rss/dyn/breaking_news.rss" }
],
"maxItems": 20,
"outputFormat": "items"
}

Output example (one dataset item per feed entry)

{
"kind": "item",
"status": "ok",
"title": "Example headline",
"link": "https://www.example.gov/news/123",
"published": "2026-09-30T12:00:00+00:00",
"authors": ["Press Office"],
"categories": ["News"],
"summary": "Short plain-text or Markdown summary…",
"contentMarkdown": "# Example headline\n\nFull body as Markdown…",
"contentSource": "content",
"contentHash": "sha256…",
"wordCount": 420,
"feedUrl": "https://www.example.gov/feed.xml",
"feedTitle": "Example Feed",
"feedFormat": "rss"
}

Key-value store records

KeyContent
OUTPUTRun summary: items/chunks saved, feeds read/failed, duplicates, duration, peak memory
FEED_REPORTOne entry per feed or discovery page: status (ok, not_found, error, skipped), HTTP status, format, entry counts, error

Pricing

Pay per event (validated on private cloud benches; numbers in Measured results above):

EventPrice
Feed item saved (primary)$0.0005 per entry (= $0.50 per 1,000 items)
Actor startApify default ($0.00005 per GB of run memory)

RAG chunk rows are not billed as a separate PPE event (dataset storage still applies on the platform). Failed feeds and duplicate items are never charged. Measured platform cost ≈ $0.000007 / item at 1,000 items (items-only).

Known limits

  • Does not fetch the full article page behind each item link (feed body / summary only). Full-page extraction may come later as an optional paid event.
  • Does not execute JavaScript; feeds must be plain HTTP(S) XML/JSON.
  • Malformed feeds: feedparser is tolerant, but severely broken XML may yield zero entries (bozo noted in FEED_REPORT).
  • JSON Feed support depends on feedparser; treat as best-effort until covered in measured tests.
  • Very large feeds: use Max items / Max items per feed to cap cost and memory.

FAQ

Can I pass a homepage instead of a feed URL? Yes — put it in Page URLs. The Actor looks for <link rel="alternate"> feed links and, if none are found, tries common paths like /feed and /atom.xml.

Are failed feeds charged? No. Only successfully saved feed entries are charged.

License & source code

This Actor is open source under the GNU Affero General Public License v3.0 (AGPL-3.0) — see LICENSE. The full source code is public: https://github.com/xbox002000/rss-atom-to-markdown

Third-party notices: NOTICE. Changelog: CHANGELOG.md.