Sitemap URL Extractor: All Page URLs from XML Sitemaps
Pricing
Pay per event
Sitemap URL Extractor: All Page URLs from XML Sitemaps
Extract every page URL from any website's sitemaps. Finds sitemaps in robots.txt, follows sitemap indexes and gzip files, and returns lastmod, priority, hreflang alternates, image counts and optional HTTP status. Filter by pattern or date.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Hay Equipos
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 hours ago
Last modified
Categories
Share
Get a complete list of a website's pages in seconds. Enter a domain or a sitemap link and the actor finds the sitemaps (robots.txt first, then the usual paths), follows every sitemap index, opens gzip files, and returns one row per URL with last modified date, change frequency, priority, hreflang alternates, image and video counts, and Google News fields. Turn on the status check to find broken or redirected pages that are still listed in the sitemap.
No browser and no crawling of the pages themselves, so it is fast and cheap: a sitemap with 50,000 URLs is read in a few seconds.
Who uses it
- SEO specialists auditing sitemaps, finding 404s and redirects, and checking hreflang coverage.
- Content teams and agencies listing every blog post or product page of a competitor, with dates.
- Developers and AI builders who need the list of pages to feed a crawler, a RAG pipeline or a change monitor.
- Migration projects that need a before and after list of URLs.
Input
| Field | What it does |
|---|---|
startUrls | Websites (example.com) or sitemap links (.xml, .xml.gz, .txt, sitemap index, robots.txt, RSS or Atom feed) |
includeUrlPatterns | Keep only URLs matching any of these regular expressions or plain words, for example /blog/ |
excludeUrlPatterns | Drop URLs matching any of these, for example /tag/ |
lastmodAfter | Keep only URLs changed on or after a date, for example 2026-09-01 |
sameHostOnly | Drop URLs on other hosts |
checkStatus | Add HTTP status, final URL and a redirected flag for every URL |
maxSitemapsPerSite, maxUrlsPerSite, maxUrls | Caps to control run time and cost |
Example input
{"startUrls": ["https://blog.cloudflare.com", "https://www.nasa.gov/sitemap.xml"],"includeUrlPatterns": [],"excludeUrlPatterns": ["/tag/"],"lastmodAfter": "2026-01-01","checkStatus": false,"maxUrlsPerSite": 50000}
Output
One row per unique URL:
{"url": "https://blog.cloudflare.com/fr-fr/cloudflare-turns-8","lastmod": "2026-07-15T16:05:05.825Z","changefreq": null,"priority": null,"alternates": [{ "hreflang": "en-us", "url": "https://blog.cloudflare.com/cloudflare-turns-8" },{ "hreflang": "de-de", "url": "https://blog.cloudflare.com/de-de/cloudflare-turns-8" }],"imageCount": 0,"videoCount": 0,"newsTitle": null,"newsPublishedAt": null,"sitemapUrl": "https://blog.cloudflare.com/sitemap-posts.xml","site": "https://blog.cloudflare.com","statusCode": 200,"finalUrl": "https://blog.cloudflare.com/fr-fr/cloudflare-turns-8/","redirected": true,"scrapedAt": "2026-09-27T06:14:24.933Z"}
statusCode, finalUrl and redirected appear only with the status check on. A RUN_SUMMARY record in the key value store shows, for each site, how the sitemap was found, how many sitemap files were read, any sitemap errors, and how many URLs were saved.
Export as JSON, CSV, Excel or HTML, or use the Apify API, webhooks, Make, Zapier or the Apify MCP server.
Pricing
Pay per event, no subscription and no platform usage charges on top:
- $0.25 per 1,000 URLs saved ($0.00025 per URL).
- $1.00 per 1,000 status checks, only when you turn the status check on.
- Filtered URLs, duplicates and sites without a sitemap cost nothing.
Limits
- The actor lists what the sitemaps say. Pages missing from the sitemap are not found, because pages are not crawled.
- Sites that block automated requests to their sitemap (some sites behind strict bot protection) return an error in
RUN_SUMMARY. lastmodis whatever the site writes; some sites set every URL to today's date.- With
lastmodAfterset, URLs without a lastmod are dropped, and sitemap files whose own lastmod is older than the date are skipped. - The status check sends a HEAD request (a GET if HEAD is refused) at a polite pace of about three requests a second per site, so it is much slower than listing.
FAQ
Where does it look for sitemaps? First every Sitemap: line in robots.txt, then /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /wp-sitemap.xml and /sitemap.txt. You can also give sitemap links directly.
Does it handle huge sites? Yes. Sitemap indexes with hundreds of files are followed breadth first. Use maxSitemapsPerSite and maxUrlsPerSite to cap cost.
Can I find broken links in my sitemap? Yes, turn on checkStatus and filter the results for statusCode 404 or redirected true.
Is it allowed? Sitemaps and robots.txt exist so that machines can read them. The actor does not log in and keeps a polite pace.