Sitemap URL Extractor Pro
Pricing
from $0.30 / 1,000 url extracteds
Sitemap URL Extractor Pro
Extract every URL from any website's XML sitemaps: auto-discovery via robots.txt, nested indexes, .gz, images, hreflang, news, lastmod filters. Reliable, fast HTTP, $0.30 per 1,000 URLs.
Pricing
from $0.30 / 1,000 url extracteds
Rating
0.0
(0)
Developer
Sai
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 days ago
Last modified
Categories
Share
Get every URL a website publishes in its XML sitemaps as a clean table you can export to CSV, Excel, JSON or Google Sheets. Give it plain domains and it finds the sitemaps for you. Nested sitemap indexes, gzipped .xml.gz files, text sitemaps and RSS/Atom feeds are all handled.
Built for reliability. Sites where other extractors return nothing (sitemap only listed in robots.txt, wrong path given, gzip served without the right headers, sm: namespace prefixes, CDATA, BOMs, 429 rate limits) are exactly what this Actor was designed and tested for.
What you get
| Field | Example |
|---|---|
url | https://wordpress.org/news/2026/09/... |
lastmod | 2026-09-30T12:00:00+00:00 |
changefreq / priority | weekly / 0.8 |
site | https://wordpress.org |
sitemap | the sitemap file the URL came from |
images | image URLs from <image:image> |
alternates | hreflang alternates [{hreflang, href}] |
news / videos | Google News and video sitemap fields |
A per-site SUMMARY record (key-value store) lists how each site's sitemaps were discovered, every sitemap file processed, any sitemap that failed and why, and the counts of duplicates and filtered URLs. You never have to guess why a site returned nothing.
Features
- Auto-discovery: reads
Sitemap:lines inrobots.txt, then probes/sitemap.xml,/sitemap_index.xml,/wp-sitemap.xml,/sitemap.xml.gz,/sitemap.txtand more. If a sitemap URL you give 404s, it falls back to discovery automatically. - Any depth of nested sitemap indexes, processed in parallel (4 files per site, several sites at once).
- gzip detection by content, not just by file extension.
- Tolerant parser: namespaces and prefixes, CDATA, HTML entities, BOMs, text sitemaps, RSS and Atom.
- Retries with exponential backoff and
Retry-Aftersupport on 429/5xx and network errors. - Filters: include/exclude regex, "modified after" date (
lastmodAfter), same-host only, and per-site and total caps. - Deduplicated per site.
- No browser and no proxy needed. It's plain HTTP, so it's fast and cheap.
Input example
{"startUrls": ["wordpress.org", "https://www.nasa.gov/sitemap.xml", "apify.com"],"includePatterns": ["/news/", "/blog/"],"lastmodAfter": "2026-09-01","maxUrlsPerSite": 50000}
Use cases
- SEO audits: compare sitemap URLs with crawled or indexed pages, and find orphan or stale pages.
- Seed lists for scrapers and crawlers (feed URLs into Website Content Crawler, Screenshot or Change Monitor actors).
- Content inventories and migrations.
- Competitor monitoring: what did they publish this week? (
lastmodAfter) - Building RAG / LLM datasets from a site's canonical pages.
Pricing
Pay per event:
- $0.30 per 1,000 URLs ($0.0003 per unique URL written to the dataset)
- Discovery, robots.txt, sitemap downloads, duplicates, filtered-out URLs and errors are free
- Plus Apify's tiny actor start fee ($0.00005)
Use maxUrlsTotal to put a hard ceiling on a run's cost.
FAQ
A site returns 0 URLs. Why? Check the SUMMARY record. Usually the site has no sitemap at all (no_sitemap_found), or it blocks automated requests (sitemaps_unreachable with e.g. http_403). You can try a different userAgent.
Is there a limit? No hard limit. Sites with millions of URLs work, but set maxUrlsTotal if you want to cap cost.
Does it crawl pages? No. It only reads sitemaps, which is why it's fast and cheap. Pair it with a crawler if you need page content.
Related Actors
- Website Change Monitor: get alerted when pages change
- Website Screenshot Pro: bulk screenshots and PDFs
- Website Contact & Tech Enricher: emails, phones, socials and tech stack