Sitemap URL Extractor: Auto-Discovery, Filters & Health Report
Pricing
$0.50 / 1,000 results
Sitemap URL Extractor: Auto-Discovery, Filters & Health Report
Get every URL from a website's sitemaps: just enter the domain. Handles sitemap indexes and .xml.gz, adds lastmod, images, hreflang and news data, filters by date or URL pattern, and reports sitemap problems.
Pricing
$0.50 / 1,000 results
Rating
0.0
(0)
Developer
CreativeFour LLC
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
What does the Sitemap URL Extractor do?
Enter just a domain and get every URL from the site's sitemaps, with each page's last-modified date, change frequency, and priority, plus any images, language versions (hreflang), and news data the sitemap declares.
It finds sitemaps automatically (from robots.txt, or the usual paths like /sitemap.xml), follows sitemap indexes to any depth, reads gzipped .xml.gz files, and writes a sitemap health report: duplicates, off-domain URLs, invalid dates, and broken sitemap files.
On Apify you can schedule it, call it from the API, or send the list straight into Google Sheets, a crawler, or an SEO tool.
Why use it?
- Get a site's full page list in seconds, without crawling it.
- Watch what changes. Use Only URLs modified on or after with a schedule to see what a competitor published or updated this week.
- Audit your own sitemaps. Catch duplicates, URLs pointing at another domain, broken child sitemaps, and missing or invalid
<lastmod>dates. - Feed other tools. Pass the URLs to a status checker, crawler, Lighthouse audit, or RAG pipeline.
- Check international SEO. See each page's hreflang alternates side by side.
How to use it
- Open the Input tab and type a domain (for example
example.com), or paste sitemap URLs. - Optional: add a modified since date, or URL patterns to include or skip.
- Click Start.
- Open the Output tab for the URL list, and the SUMMARY record for the health report.
Input
| Field | What it does |
|---|---|
| Websites | Domains. Sitemaps are discovered from robots.txt or the usual paths. |
| Sitemap URLs | Or give sitemap or index URLs directly (.xml or .xml.gz). |
| Only URLs modified on or after | Keeps URLs whose <lastmod> is on or after this date. |
| Only URLs matching / Skip URLs matching | Regular expressions or plain text, such as /blog/ or /tag/. |
| Max URLs / Max sitemap files | Safety limits. |
{"websites": ["example.com"],"modifiedSince": "2026-09-01","includePatterns": ["/blog/"]}
Output
One row per URL. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.
{"site": "example.com","url": "https://example.com/blog/new-post","lastmod": "2026-09-24T10:00:00Z","changefreq": "weekly","priority": 0.8,"images": ["https://example.com/img/hero.jpg"],"alternates": [{ "hreflang": "es", "href": "https://example.com/es/blog/new-post" }],"newsTitle": null,"newsPublishedAt": null,"videos": 0,"sitemap": "https://example.com/post-sitemap.xml.gz","offDomain": false,"lastmodValid": true}
SUMMARY (key-value store) gives, for each site: the number of sitemaps read, URLs found and saved, duplicates, off-domain URLs, invalid and missing lastmod dates, and every sitemap file that failed to load, with the reason.
Data fields
| Field | Description |
|---|---|
url, lastmod, changefreq, priority | Standard sitemap fields |
images, videos | Image URLs and video count, from image and video sitemap extensions |
alternates | hreflang language versions |
newsTitle, newsPublishedAt | News sitemap fields |
sitemap | Which sitemap file listed the URL |
offDomain, lastmodValid | Health flags |
How much does it cost to extract sitemap URLs?
You pay per URL saved. Filters (modified-since and URL patterns) are applied first, so you only pay for the URLs you keep. Set a maximum charge per run in the run options, and the Actor stops cleanly at that limit.
Tips
- Check the URLs next with our Bulk URL Status & Redirect Checker, or audit their speed with the Bulk Lighthouse & Core Web Vitals Audit.
- Some sites have no sitemap. The summary says so, and the Broken Link Checker can crawl them instead.
Use it from AI agents (MCP)
AI agents can find and run this Actor through the Apify MCP server.
- Claude, ChatGPT, or any MCP client: add
https://mcp.apify.com?tools=creativefour/sitemap-url-extractoras a custom connector, and sign in to Apify when prompted. - Claude Code, Cursor, VS Code, or Codex: run
apify mcp install claude-code(swap in your client's name), then ask your agent to "list every URL on example.com with creativefour/sitemap-url-extractor".
FAQ and support
What if a site has no sitemap? The summary reports "No readable sitemap found". Sitemaps are optional, and some sites don't publish one.
Does it visit every page? No. It reads only the sitemap files, which makes it fast and light on the site.
Found a bug or need a feature? Open an issue on the Issues tab. Custom versions are available on request.