RSS, Atom & Podcast Feed Scraper avatar

RSS, Atom & Podcast Feed Scraper

Pricing

from $1.00 / 1,000 feed entries

Go to Apify Store
RSS, Atom & Podcast Feed Scraper

RSS, Atom & Podcast Feed Scraper

Turn public RSS 2.0, Atom 1.0 and podcast feeds into structured entries with dates, source text, categories and enclosure links. Add direct HTTPS feed URLs; inspect per-feed coverage in OUTPUT.

Pricing

from $1.00 / 1,000 feed entries

Rating

0.0

(0)

Developer

Akshay Aggarwal

Akshay Aggarwal

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Categories

Share

Turn public RSS 2.0 and Atom 1.0 feeds into a structured dataset for news monitoring, podcast research, and release tracking. Give it direct feed URLs; get one row per saved entry with dates, source text, categories, and podcast enclosure links when the feed provides them.

Quick start

Paste this input in Apify Console:

{
"feedUrls": [
"https://www.nasa.gov/news-release/feed/",
"https://github.com/python/cpython/releases.atom"
],
"maxItemsPerFeed": 2,
"deduplicate": true
}

The default is 10 entries per feed, so a single feed is a small first run. maxItemsPerFeed accepts 1–500, and feedUrls accepts 1–50. sinceDate is optional: use an ISO 8601 date (2026-09-01, treated as UTC midnight) or timezone-aware datetime. When a date filter is set, entries with no usable publication or update date are skipped.

Here is an entry from the Python release Atom feed (examples/python-release-entry.json):

{
"title": "v3.15.0rc2",
"url": "https://github.com/python/cpython/releases/tag/v3.15.0rc2",
"author": "hugovk",
"updated_at": "2026-09-01T09:14:36Z",
"content_if_feed_supplies_it": "<p>Python 3.15.0rc2</p>",
"enclosures": []
}

Each full dataset row also has feed_url, source id/GUID, published_at, summary, categories, and source_url (the final feed URL after redirects). Missing source fields are null; the Actor does not make up dates or summaries. RSS description and Atom summary are copied from the feed. content_if_feed_supplies_it comes only from RSS content:encoded or Atom content. Podcast <enclosure> and Atom enclosure links become {url,type,length} objects; audio is not downloaded.

Coverage and billing

The supported sources are public HTTPS RSS 2.0 and Atom 1.0 XML feeds, including podcast RSS feeds and public release feeds such as GitHub's Atom release feed. Enter the direct feed URL, not a website homepage or search query. This Actor does not search the web, open linked articles, extract paywalled full text, download audio, or transcribe it. It does not use AI. Feed entries with neither a usable entry URL nor an ID are skipped.

One saved feed entry is the pay-per-event billing unit (feed-entry) when pay-per-event pricing is enabled. Check the current price before running. Duplicate entries, skipped entries, errors, and the run summary are not result rows. The default dataset contains only saved entries. OUTPUT in the default key-value store records fetched, saved, skipped, error and limit counts by feed. If a feed fails or the run hits a charge limit, earlier rows may still be present; inspect OUTPUT before treating the dataset as complete. A run can have partial results because a source is malformed, unavailable, redirected unsafely, or larger than the 2 MiB per-feed response cap.

The Actor fetches at most 50 feeds, one request per feed, with a 90-second overall fetch budget and a 12-second timeout per request. It allows up to five redirects; every hop must be public HTTPS on port 443. Feeds are limited to 2 MiB each; XML DTDs and external entities are refused. Source timestamps are normalized to UTC when they contain a timezone; unparseable or timezone-free values become null. It does not crawl archive pages or guarantee all historical entries: the source's current feed window determines what is available.

Fields

FieldMeaning
feed_urlInput URL
idRSS GUID or Atom ID, if supplied
title, url, authorFeed-supplied entry metadata
published_at, updated_atSource dates normalized to UTC, or null
summarySource description or summary; no AI summary
content_if_feed_supplies_itSource feed content only, if present
categoriesSource categories or terms
enclosuresMedia link metadata, no media content
source_urlFinal fetched feed URL after HTTPS redirects

Deduplication uses the entry URL across feeds, or the GUID within a feed when no URL exists. A per-feed cap counts selected entries. For each feed, OUTPUT reports fetched and saved plus skip reasons (date, duplicate, missing_identity, limit).