Sitemap URL Extractor & XML Sitemap Scraper (Index, Gzip) avatar

Sitemap URL Extractor & XML Sitemap Scraper (Index, Gzip)

Pricing

from $0.20 / 1,000 url extracteds

Go to Apify Store
Sitemap URL Extractor & XML Sitemap Scraper (Index, Gzip)

Sitemap URL Extractor & XML Sitemap Scraper (Index, Gzip)

Extract every URL from a website's sitemap.xml: give a domain and it finds sitemaps via robots.txt, follows sitemap indexes, reads .gz, text and RSS sitemaps, and recovers broken XML. Returns lastmod, changefreq, priority, optional HTTP status. $0.20 per 1,000 URLs.

Pricing

from $0.20 / 1,000 url extracteds

Rating

0.0

(0)

Developer

Ventura WorkAlong

Ventura WorkAlong

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 hours ago

Last modified

Share

Sitemap URL Extractor & XML Sitemap Scraper

Sitemap URL Extractor gets every page URL listed in a website's sitemap. Give it a domain and it finds the sitemaps for you. It reads the Sitemap: lines in robots.txt, follows sitemap indexes to the end, opens .xml.gz files, and parses plain-text and RSS/Atom sitemaps. It also recovers URLs from broken XML. The output is a flat list of URLs with lastmod, changefreq, priority and the sitemap each came from, ready for a crawler, an SEO audit, a spreadsheet or an AI agent.

  • Reliable by design: one broken or missing child sitemap never fails the run. Every other sitemap is still read, and the failure is reported with its reason.
  • Handles real-world sitemaps: sitemap indexes (nested, with loop protection), gzip, text and RSS/Atom sitemaps, BOMs, CDATA, namespace prefixes, unescaped & and truncated files.
  • Finds sitemaps automatically: robots.txt first, then /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml and other common locations, including WordPress installs in a subfolder. HTML "soft 404" pages are never mistaken for sitemaps.
  • Filters: regex include/exclude and a lastmod date range ("pages changed since last week").
  • Optional HTTP status check per URL to find broken pages and redirects listed in your sitemap.
  • $0.20 per 1,000 URLs. Sites without a sitemap are free.

How to extract all URLs from a sitemap

  1. Add one or more websites to Websites or sitemap URLs. Use a domain (example.com), any page URL, a robots.txt URL or a direct sitemap URL (https://example.com/sitemap_index.xml).
  2. Optional: set Max URLs per site (0 = all), URL patterns or a last-modified date range.
  3. Click Start. Download the results as JSON, CSV or Excel, or fetch them through the Apify API.

Input example

{
"startUrls": ["https://crawlee.dev", "https://www.gov.uk/sitemap.xml"],
"maxUrlsPerSite": 0,
"includeUrlPatterns": ["/blog/"],
"lastmodFrom": "2026-09-01"
}
FieldWhat it doesDefault
startUrlsDomains, page URLs, robots.txt URLs or sitemap URLsrequired
maxUrlsPerSiteStop after this many unique URLs per input (0 = all)0
includeUrlPatterns / excludeUrlPatternsRegular expressions matched against each URLnone
lastmodFrom / lastmodToKeep URLs whose lastmod is in this range. URLs without a lastmod are dropped while a date filter is setnone
includeImagesAdd image URLs from image sitemapsfalse
includeAlternatesAdd hreflang alternates (multilingual sites)false
checkStatusHEAD request per URL, returning httpStatus and redirectTofalse
maxStatusChecksCap on status checks per run1000
maxSitemapsPerSiteSafety cap on sitemap files per site1000
respectRobotsTxtHonor robots.txt rules and Crawl-delaytrue
maxRequestsPerSecondPerHostPoliteness limit per site (0.2–5)1
maxConcurrencySites processed in parallel5

Output example

One dataset item per unique URL. This is real output from https://crawlee.dev (2026-10-07):

{
"type": "url",
"url": "https://crawlee.dev/blog",
"host": "crawlee.dev",
"path": "/blog",
"lastmod": null,
"changefreq": "weekly",
"priority": 0.5,
"sitemapUrl": "https://crawlee.dev/sitemap.xml",
"input": "https://crawlee.dev"
}
  • lastmod is normalized: dates stay YYYY-MM-DD, and datetimes become UTC ISO 8601 (2026-10-07T09:36:29Z).
  • With checkStatus on, each item also gets httpStatus (e.g. 200, 301, 404) and redirectTo.
  • An input with no readable sitemap gets one free row with "type": "error" that says why:
{ "type": "error", "input": "https://www.python.org", "url": null,
"error": "No sitemap found: robots.txt lists none and none of /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /wp-sitemap.xml, /sitemap.xml.gz, /sitemap.txt exist." }
  • The key-value store record SUMMARY has a per-site report: how sitemaps were discovered, every sitemap file read, each one that failed and why, duplicates removed and warnings. Example for MDN, whose sitemaps are gzipped:
{ "input": "https://developer.mozilla.org", "discovery": "robots.txt",
"sitemapsFound": ["https://developer.mozilla.org/sitemap.xml", "https://developer.mozilla.org/sitemaps/en-us/sitemap.xml.gz"],
"sitemapsProcessed": 2, "sitemapsFailed": [], "urlsFound": 301, "duplicates": 1, "urlsOutput": 300,
"stoppedEarly": "maxUrlsPerSite (300) reached" }

Pricing

Pay per event, with no platform usage charges on top:

EventPrice
URL extracted (url-extracted)$0.0002 per URL ($0.20 per 1,000)
URL status checked (url-status-checked, only with checkStatus)$0.0005 per URL

Examples: a 5,000-page site costs $1.00. A 200-page site with status checks costs $0.04 + $0.10 = $0.14. Inputs without a sitemap cost nothing. Each URL is charged once, even if several sitemaps list it. The run stops cleanly at your maximum total charge.

Use cases

  • SEO audits: list every indexable URL, find sitemap URLs that redirect or 404, and check lastmod freshness.
  • Crawling and scraping: feed the URL list into Website Content Crawler or your own scraper instead of crawling links.
  • Competitor and content monitoring: run on a schedule with lastmodFrom to see new or updated pages.
  • Site migrations: export the old site's URL inventory before a redirect project.
  • AI agents (MCP): "list all blog posts on example.com" in one call.

How it works (and how it stays polite)

  • Requests are identified as SitemapExtractorBot. robots.txt is honored by default, including Crawl-delay, and each site gets at most 1 request per second unless you raise it.
  • Transient errors (HTTP 429/5xx, timeouts) are retried with backoff, and Retry-After is honored. If a site keeps rate-limiting, the Actor stops asking and reports it.
  • Files are capped at 60 MB even after decompression, so gzip bombs are safe. Private and internal network addresses are refused.
  • Only public sitemap files are read. No login, no proxies, no personal data.

Limits

  • Only sitemaps the site publishes are read. The Actor doesn't crawl HTML links, so pages missing from the sitemap won't appear.
  • Sitemaps that need JavaScript or a login can't be read. If robots.txt blocks bots from the sitemap, that sitemap is skipped (you'll see the reason).
  • News and video sitemap extensions are read as normal URLs; their extra fields aren't returned.
  • The status check reports the first response (it doesn't follow redirects) and uses HEAD, falling back to GET when HEAD isn't allowed.

FAQ

Do I need the exact sitemap URL?

No. Give the domain and the Actor discovers the sitemaps. A direct sitemap URL also works and skips discovery.

Can it read sitemap index files with thousands of sitemaps?

Yes. It follows indexes breadth-first, skips loops and duplicates, and stops at maxSitemapsPerSite (default 1,000; up to 50,000).

Why is lastmod empty for some URLs?

That site's sitemap doesn't publish it. The Actor never invents dates.

Is this the same as Apify's Sitemap Extractor?

No. This is an independent Actor focused on discovery, malformed-file recovery, filters and a transparent per-site report.

Found a sitemap this Actor can't read? Open an issue with the URL. Fixing those is the point of this Actor.