Sitemap Scraper - All URLs from XML Sitemaps + Status Check
Pricing
from $0.16 / 1,000 urls
Sitemap Scraper - All URLs from XML Sitemaps + Status Check
Every URL from a website's XML sitemaps (robots.txt discovery, indexes, .gz, text sitemaps) with lastmod, images, hreflang and news tags. Optional HTTP status/redirect check. Monitoring of new URLs.
Pricing
from $0.16 / 1,000 urls
Rating
0.0
(0)
Developer
TinyRex
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
8 hours ago
Last modified
Categories
Share
Sitemap Scraper: extract all URLs from any XML sitemap (and find broken links)
Get every URL a website lists in its XML sitemaps as clean JSON or CSV — page URL, lastmod, change frequency, priority, images, videos, hreflang alternates and Google News tags. Paste a domain: sitemaps are discovered from robots.txt and common paths, indexes are followed, .xml.gz and text sitemaps are supported. Optionally check the HTTP status of every URL to catch broken pages and redirects, or schedule runs to return only new URLs. $0.20 per 1,000 URLs (status checks $0.50/1k); sites without a sitemap cost nothing.
What you get from the Sitemap Scraper
- Complete URL inventory — every
<loc>from sitemap indexes (nested up to 5 levels), gzip and plain-text sitemaps, with lastmod, changefreq, priority, image/video extensions, hreflang and news tags. - Broken link & redirect audit (optional) — status code, final URL after redirects, redirect chain, response time and content type for each URL, while respecting robots.txt.
- New-URL monitoring — turn on Only new URLs and each scheduled run returns only pages that were not in the sitemap before (new blog posts, products, job pages, competitor landings).
Sample output (JSON)
{"site": "crawlee.dev","url": "https://crawlee.dev/python/blog","lastmod": null,"changefreq": "weekly","priority": 0.5,"sitemapUrl": "https://crawlee.dev/python/sitemap.xml","imageCount": 0,"statusCode": 200,"finalStatusCode": 200,"finalUrl": "https://crawlee.dev/python/blog","redirectChain": [],"responseTimeMs": 581,"contentType": "text/html"}
Per-run reports: SITES (robots.txt found?, sitemaps read/failed, URLs saved, broken counts) and SUMMARY (totals). Download as CSV, Excel or JSON, or send to Sheets / Make / Zapier / n8n / webhooks / API.
How to configure the Sitemap Scraper
- Add websites, sitemap URLs or robots.txt URLs, one per line (
apify.com,https://crawlee.dev/sitemap.xml). - Optional: include/exclude URL patterns (plain text or
/regex/), last modified within N days, max URLs per site. - Optional: enable Check HTTP status of every URL for a broken-link / redirect audit.
- Run.
Example input
{"startUrls": ["https://apify.com", "https://crawlee.dev"],"maxUrlsPerSite": 300,"includeUrlPatterns": [],"checkStatus": true}
| Setting | Default | Notes |
|---|---|---|
| startUrls | (required) | Domains, sitemap URLs or robots.txt URLs |
| maxUrlsPerSite | 500 | Cap per site (raise for full inventories) |
| checkStatus | false | HEAD→GET status check + redirect chain |
| includeUrlPatterns / excludeUrlPatterns | [] | e.g. /blog/, /regex/ |
| lastModifiedWithinDays | 0 | 0 = no lastmod filter |
| onlyNewUrls | false | Diff against previous run for monitoring |
Common use cases
- SEO audits — all indexable URLs with lastmod, plus 4xx/5xx and unexpected redirects still listed in the sitemap.
- Crawl seeding — a complete URL list for a content or product scraper instead of link-crawling.
- Competitor monitoring — new product, blog or landing pages daily with Only new URLs.
- Migrations — export old and new sitemaps to build redirect maps.
- AI / RAG — the list of documentation pages to index.
Pricing (failed = free)
Pay per event:
- URL (
url): one charge per URL saved (~$0.20 / 1,000). - URL status check (
url-status-check): one charge per checked URL when the option is on (~$0.50 / 1,000). - Free: sites without a sitemap, failed sitemaps, filtered-out, duplicate and already-seen URLs.
Set a max cost per run; the Actor stops gracefully when it is reached. See the Pricing tab for Bronze/Silver/Gold discounts.
Limitations
- Only URLs the site lists in its sitemaps; it does not crawl links. Sites with no sitemap are reported in
SITES(and are free). - Datacenter IP blocks may need a proxy.
- Status checks use HEAD (fallback GET); some servers answer HEAD differently from browsers.
FAQ — Sitemap Scraper / XML sitemap URL extractor
How do I extract all URLs from a website's sitemap?
Paste the domain (or the sitemap URL). The Actor reads robots.txt Sitemap: lines, then common paths (/sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml, …), follows indexes, and writes every URL to the dataset.
Where is sitemap.xml if it is not at /sitemap.xml?
Many sites only list it in robots.txt, or use /sitemap_index.xml, /wp-sitemap.xml or a .gz file. Paste the bare domain and auto-discovery covers the usual locations; if the path is unusual, paste the sitemap URL directly.
How can I find broken links in a sitemap?
Turn on Check HTTP status of every URL. Each row gets statusCode, finalStatusCode, finalUrl and redirectChain so you can filter 4xx/5xx and redirect chains that should not be indexed.
Does it support gzipped and sitemap indexes?
Yes — nested indexes up to 5 levels, .xml.gz, plain-text sitemaps, plus image, video, hreflang and Google News extensions.
How many URLs can it handle?
Hundreds of thousands per site; files are streamed and capped by maxUrlsPerSite.
Is scraping sitemaps legal? Sitemaps and robots.txt are published so automated tools can discover pages. The Actor reads those files (and, if enabled, one polite request per URL while respecting robots.txt). It does not log in or bypass protection. Use results in line with each site's terms.
Does it work with AI agents / MCP?
Yes. Connect to https://mcp.apify.com?tools=tinyrex/sitemap-scraper and ask e.g. "List all blog URLs on crawlee.dev updated in the last 30 days" or "Find broken URLs in the sitemap of example.com".
📘 Step-by-step guide (Python, JS, curl, MCP): How to extract all URLs from a website's sitemap (and find broken links)
Related actors
- Tech Stack Detector — CMS, ecommerce, analytics and email provider at scale
- Contact Details Extractor — emails, phones and socials from domain lists
- RSS Feed Scraper — RSS/Atom/JSON feeds as clean JSON/CSV
- Shopify Products Scraper — products, prices, variants and stock
Built by TinyRex. Questions or feature requests? Open an issue on the Actor page.