Sitemap URL Extractor – All Pages for RAG & SEO
Pricing
Pay per event
Sitemap URL Extractor – All Pages for RAG & SEO
Get every URL a website publishes: XML sitemaps, sitemap indexes, gzip and news sitemaps, RSS/Atom feeds, plus a shallow crawl when there's no sitemap. lastmod, source and robots.txt status per URL. Feed it to your RAG pipeline, crawler or SEO audit. $0.40 per 1,000 URLs.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Joshua White
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Sitemap URL Extractor lists every page a website publishes, from its XML sitemaps, nested sitemap indexes, gzip and Google News sitemaps and RSS/Atom feeds, and falls back to a polite shallow crawl when there's no sitemap. One clean row per URL with its last-modified date, where it was found and whether robots.txt allows crawling it. Built for RAG pipelines, AI agents and SEO audits.
Try it in one click: the input is prefilled with two docs sites, 100 URLs each. Apify's free plan covers about 12,000 URLs a month.
How to extract all URLs from a website in 3 steps
- Add websites, one per line (or separated by commas):
example.com, a section likehttps://example.com/blog/, or a sitemap URL. - Optional: set Max URLs per site, include or exclude patterns (
/blog/*,*.pdf), or Modified since (7 days) to get only changed pages. - Click Start, then download CSV, Excel or JSON, or pass the list straight to a content crawler.
How much does it cost to extract sitemap URLs?
$0.40 per 1,000 URLs returned, after your filters and dedupe, so Apify's $5 monthly free credit covers about 12,000 URLs. Reading sitemaps is free, filtered-out URLs are free, sites that fail are free, and rows for entries that aren't web addresses are free. A 5,000-page site costs $2. Max URLs per site and the run's maximum cost both cap what you spend; the run stops cleanly when the cap is reached.
What the sitemap extractor handles
- Finds every sitemap:
Sitemap:lines in robots.txt,/sitemap.xml,/sitemap_index.xml,/wp-sitemap.xml,<link rel="sitemap">, and a section's ownsitemap.xml(e.g./docs/sitemap.xml). It follows nested sitemap indexes (loops are detected) and reads.xml.gzgzip sitemaps, Google News sitemaps and plain-text ones. - Reads RSS and Atom feeds linked from the site. Feeds often list the newest posts first, with a title and publication date.
- Falls back to a shallow crawl when a site has no sitemap or feed (or always, if you ask): it follows the site's own links a few clicks deep, opens only HTML pages and honours robots.txt.
- Streams huge sitemaps. 50,000-URL sitemaps and sites with thousands of sitemaps are read as they download, never loaded whole into memory, and reading stops as soon as your limit is reached.
- Doesn't fall over on messy sitemaps: malformed XML, unescaped
&, unclosed tags, HTML error pages served assitemap.xml, truncated gzip files, relative URLs and odd date formats are all handled. - Only a section: start from
https://docs.example.com/guideto get just that part of the site. If the address redirects (stripe.com/docs→docs.stripe.com), the section follows it. - Filters: include and exclude patterns (globs like
/blog/*or*.pdf, or regular expressions), modified since a date (only pages that changed, for refreshing an index), and a max number of URLs per site. - Clean output: fragments and
utm_tracking parameters are removed, and duplicates are dropped, includinghttp/https,www/no-wwwand trailing-slash variants. - robots.txt status per URL: each row says whether robots.txt lets crawlers (
User-agent: *) fetch that page, so your downstream crawler can skip what it shouldn't touch.
Input example
{"startUrls": ["docs.apify.com", "https://www.theverge.com", "https://docs.python.org/3/"],"maxUrlsPerSite": 5000,"includePatterns": [],"excludePatterns": ["/tag/*", "/author/*"],"modifiedSince": "30 days","includeFeeds": true,"crawlMode": "fallback"}
Output example
One row per URL. lastmod is ISO 8601 (UTC) when the site gives a date; title comes from feeds and news sitemaps.
[{"url": "https://techcrunch.com/2026/10/05/at-19-ghost-founder-raises-11-million-to-build-a-3499-computer-for-your-personal-ai/","site": "techcrunch.com","sourceType": "sitemap","source": "https://techcrunch.com/news-sitemap.xml","lastmod": "2026-10-05T18:07:07.000Z","changefreq": null,"priority": null,"title": "At 19, founder raises $11 million for Ghost, maker of a $3,499 computer for personal AI","allowedByRobots": true},{"url": "https://developer.mozilla.org/en-US/docs/Web/HTTP","site": "https://developer.mozilla.org/en-US/docs/Web/HTTP","sourceType": "sitemap","source": "https://developer.mozilla.org/sitemaps/en-us/sitemap.xml.gz","lastmod": "2026-07-28T00:00:00.000Z","changefreq": null,"priority": null,"title": null,"allowedByRobots": true},{"url": "https://docs.python.org/3/download.html","site": "https://docs.python.org/3/","sourceType": "crawl","source": "https://docs.python.org/3/","lastmod": null,"changefreq": null,"priority": null,"title": null,"allowedByRobots": true}]
| Field | Meaning |
|---|---|
url | The page's address (fragment and utm_ parameters removed). |
site | The website from your input that this URL belongs to. |
sourceType | sitemap, rss, atom or crawl: how the URL was found. |
source | The exact sitemap, feed or page it was found in. |
lastmod | Last-modified or publication date, if the site gives one. |
changefreq, priority | As declared in the sitemap, if present. |
title | Page title, from feeds and Google News sitemaps. |
allowedByRobots | Whether robots.txt allows generic crawlers to fetch the URL (null if unknown). |
errorCode, error | null on URL rows. See below. |
Entries that aren't web addresses
Every input entry is accounted for. An entry that isn't a domain or URL (say docs apify com) gets one free row with
errorCode: "INVALID_INPUT", error: "Not a web address (expected e.g. example.com or https://example.com)",
site set to the entry as you typed it and url set to null. The status message counts them too: Found 200 URLs
on 2 of 3 entries; 1 wasn't a web address. Common typos like htps:// and http// are fixed for you.
Run report
The run's key-value store (Storage → Key-value store) has an OUTPUT record with one entry per site: how many URLs, which sitemaps and feeds were read, how many URLs were out of scope, filtered out or duplicates, and why a site gave no URLs (no sitemap, blocked, robots.txt, nothing under the section you chose). A site that can't be read doesn't fail the run; it's reported there and in the log. The URL counts in the status message and OUTPUT are exactly the rows in the dataset, which are exactly the URLs you're charged for.
Good to know
- Polite by design. It identifies itself as
SitemapURLsBot, honours robots.txt (including Crawl-delay, up to 10 s), sends at most 3 requests at a time to a site, and backs off on HTTP 429 and 5xx. If a site's robots.txt can't be read because the server errors, the site is treated as off-limits, as the robots standard says. - Public pages only. It never logs in and never fills forms.
- JavaScript-only sites: sitemaps and feeds work regardless of how the site is built. The crawl fallback reads the HTML the server sends, so a site that builds its links only in the browser and has no sitemap gives few URLs.
- Very large sites: the default 512 MB of memory is enough for big sites: 150,000 gov.uk URLs were tested in 37 seconds (peak 316 MB). Give the run 512 MB or more for very large sites, and 1–2 GB if you set Max URLs per site to 0 (no limit) on a site well past that size.
- Speed: a typical site with a sitemap takes 1–5 seconds; a site with hundreds of sitemaps takes a minute or two.
Ready-made examples
Each one opens this actor with the input already filled in. Click Try to run it, or change the input to fit your own list.
- Extract all blog post URLs from a website
- New and updated pages from the last 7 days
- All product URLs of a Shopify store
- Documentation site URL list for RAG and LLMs
- Latest news articles from news sitemaps
- List all pages of a website (even without a sitemap)
More tools from oldjard
- Tech Stack Detector: what any list of websites is built with.
- Shopify Products Scraper & Price Monitor: catalogs and price changes from any Shopify store.
- Workday, Greenhouse, Lever & Ashby Jobs Scraper: every open job from company career sites.
- Bulk Website Screenshot & URL to PDF: screenshots and PDFs of any list of pages.
- AI Web Scraper (your own key): describe fields in English, get JSON.
- Website Change Monitor: a before/after diff by webhook, Slack or Discord when a page changes.
- Company Registry Lookup: UK Companies House, Spain, France, Finland and Norway in one schema.
- UK & EU Public Tenders: Find a Tender and TED notices in one table, with daily only-new alerts.
Use it from an AI agent or the API
- Minimal input:
{"startUrls": ["crawlee.dev"], "maxUrlsPerSite": 100}. SetmaxUrlsPerSiteto cap the work and the cost. - Cost: $0.0004 per URL output (after filters and dedupe). Sitemaps, feeds and pages read are free. 100 URLs = $0.04.
- Run time (our runs): 3–8 s for up to 800 URLs from 5 sites; 78 s for 20,000 URLs.
- Results: the default dataset, one row per URL; the per-site report is the
OUTPUTrecord in the key-value store. - Works over the Apify MCP server (
search-actors, thencall-actor) and is eligible for agentic payments (x402).
FAQ
How do I find a website's sitemap? You don't need to; it checks robots.txt, /sitemap.xml,
/sitemap_index.xml, /wp-sitemap.xml and <link rel="sitemap">.
Sitemap to CSV? Download the run's dataset as CSV (or Excel or JSON) from the Storage tab or the API.
Why did a site return no URLs? Open the OUTPUT record in Storage → Key-value store: each site has a note. Common reasons: the
section you chose has no pages in the sitemap (turn off Only the start URL's section), the site blocks automated
visitors, or robots.txt disallows it.
Can I get only new or changed pages? Yes: set Modified since to a date or a period like 7 days. Only URLs
whose sitemap or feed date is on or after it are returned. Schedule the run to keep a RAG index fresh.
Does it download page content? No, it lists URLs, which is fast and cheap. Pass the list to a content crawler to fetch the text.