Feed Discovery Lookup — Find RSS, Atom & JSON Feeds avatar

Feed Discovery Lookup — Find RSS, Atom & JSON Feeds

Pricing

from $2.50 / 1,000 successful lookups

Go to Apify Store
Feed Discovery Lookup — Find RSS, Atom & JSON Feeds

Feed Discovery Lookup — Find RSS, Atom & JSON Feeds

Find any website's RSS, Atom, or JSON feeds by URL or domain. Reads the site's own autodiscovery tags and probes conventional feed paths, then verifies each candidate is a real, parseable feed. Pay only when a working feed is found.

Pricing

from $2.50 / 1,000 successful lookups

Rating

0.0

(0)

Developer

Adrian Voss

Adrian Voss

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 hours ago

Last modified

Share

Find a website's RSS, Atom, or JSON feeds by URL or domain. This actor reads the page's own <link rel="alternate"> autodiscovery tags first, and falls back to probing conventional feed paths (/feed, /rss.xml, /atom.xml, and similar) when a site doesn't declare them — then confirms each candidate is a real feed by checking its content, not just its URL.

Features

  • Autodiscovery-first. Reads the homepage's declared <link rel="alternate"> feed tags (RSS, Atom, JSON Feed) — the same signal browsers and feed readers use.
  • Fallback path probing. If a site declares no feed, checks conventional paths like /feed, /feed.xml, /rss.xml, /rss, /atom.xml, and /index.xml.
  • Content-verified, not URL-guessed. Every candidate is fetched and checked for real RSS/Atom XML or JSON Feed structure before being reported as a feed — a dead or redirected URL doesn't count.
  • Feed format & title. Each result reports its format (rss, atom, or json) and the feed's own declared title where available.
  • Pay only for hits. A site with no discoverable feed costs nothing — see Pricing.

How to use Feed Discovery Lookup — Find RSS, Atom & JSON Feeds

  1. In the Apify Console. Open the actor page and click Start — the items field is already pre-filled with a working example. Results land in the run's dataset as soon as each item is found.
  2. Via the API. Call it directly with a POST request — no Console needed once you have an API token:
    curl "https://api.apify.com/v2/acts/accountable_eel~feed-discovery-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
    -X POST \
    -H "Content-Type: application/json" \
    -d '{"items":["nytimes.com"]}'
  3. On a schedule. Save this actor as an Apify Task with the input you want, then add a Schedule (hourly, daily, weekly) so it runs on its own — no server of your own required.

Input

{
"items": ["theverge.com"],
"maxConcurrency": 5
}

items is a list of website URLs or bare domains to check — one per line. A bare domain like theverge.com is resolved to https://theverge.com. maxConcurrency controls how many sites are checked in parallel, and proxyConfiguration defaults to Apify Proxy.

Output

One row per site, for example:

{
"query": "theverge.com",
"found": true,
"data": {
"domain": "theverge.com",
"feedCount": 1,
"feeds": [
{ "url": "https://www.theverge.com/rss/index.xml", "format": "rss", "title": "The Verge" }
]
},
"scrapedAt": "2026-08-21T10:15:00.000Z"
}

A row with "found": false means neither the site's own autodiscovery tags nor any of the conventional fallback paths resolved to a real, parseable feed — these rows are never charged.

Use cases

  • Build a reading-list or RSS aggregator by feeding in a list of sites you follow.
  • Check whether a company blog or newsroom publishes a feed before setting up a content monitor.
  • Audit a list of competitor or industry sites for which ones are feed-discoverable.
  • Feed discovered feed URLs into a separate RSS-polling pipeline.
  • Spot-check whether a site migration or redesign broke feed autodiscovery.

Pricing

$5 per 1,000 results, plus a $0.00005 start fee. Misses (found:false) are never charged.

Use it from Clay, n8n, Make, or an AI agent

This actor runs synchronously over plain HTTP — call it directly from a script, a workflow tool, or an AI agent, no Apify Console needed once you have an API token.

curl "https://api.apify.com/v2/acts/accountable_eel~feed-discovery-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
-X POST \
-H "Content-Type: application/json" \
-d '{"items":["nytimes.com"]}'

n8n. Add an HTTP Request node: Method POST, URL https://api.apify.com/v2/acts/accountable_eel~feed-discovery-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>, Body Content Type JSON, JSON Body {"items":["nytimes.com"]} (swap in an expression from an earlier node for a real value).

Clay. Add an "HTTP API" column: Method POST, URL https://api.apify.com/v2/acts/accountable_eel~feed-discovery-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>, Body {"items":["{{value}}"]}, mapping the row's value into the items array.

MCP. In Claude, Cursor, or any MCP client with the Apify MCP server, ask for "Feed Discovery Lookup | Apify" — the agent will find and run this actor.

FAQ

What counts as "not found"? A site with no <link rel="alternate"> feed tag and none of the fallback paths (/feed, /feed.xml, /rss.xml, /rss, /atom.xml, /index.xml) returning real feed content. A site can be perfectly reachable and still come back found: false if it simply doesn't publish a feed — this is expected, not an error.

How is a "real feed" confirmed? Each candidate URL is fetched and its content is checked for actual RSS/Atom XML markup or a JSON Feed version field — a URL that merely looks like a feed path but 404s, redirects to an HTML page, or returns something else isn't counted.

Does it check more than one feed per site? Yes — up to 3 confirmed feeds per site are reported (checking at most 6 candidate URLs), covering sites that publish separate feeds for different sections.

What if I submit a bare domain instead of a full URL? Both work. A bare domain like example.com is automatically resolved to https://example.com before the check runs.

Does this call any third-party feed-discovery API? No — it reads the target site's own HTML and feed endpoints directly, the same signals a browser or feed reader would use.

Can I run this against many sites at once? Yes. Use maxConcurrency to control request rate across a large list of domains.