Sitemap Extractor - All Website URLs + Broken Link Check
Pricing
from $0.30 / 1,000 urls
Sitemap Extractor - All Website URLs + Broken Link Check
Find every page of a website from its sitemaps (robots.txt, sitemap indexes, .gz). Returns URLs with last-modified dates, filters by pattern, and can check each URL for 404s and redirects.
Pricing
from $0.30 / 1,000 urls
Rating
0.0
(0)
Developer
Björn Ólafur
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
What does Sitemap Extractor do?
Sitemap Extractor finds every page of a website from its XML sitemaps. Just enter a domain: it reads robots.txt, checks the common sitemap locations, follows sitemap indexes and .gz files, and returns each URL with its last-modified date, change frequency and priority. Optionally it checks every URL for 404s and redirects, a quick broken-link audit of your sitemap.
Run it in Apify Console, via the API, on a schedule, or from your AI agent through the Apify MCP server. Connect it to Google Sheets, Make, Zapier or n8n with Apify integrations.
Why use Sitemap Extractor?
- 🗺️ Automatic discovery:
robots.txt→sitemap.xml,sitemap_index.xml,wp-sitemap.xml,.gzand text sitemaps. - 🔁 Follows nested sitemap indexes, with duplicates removed.
- 🎯 Filters: keep only
/blog/or/products/URLs, or drop/tag/pages. - 🩺 Broken-link check (optional): HTTP status and redirect target for every URL.
- 🖼️ Image counts from image sitemaps, and news/video sitemaps are supported.
- 💸 Cheap and fair: pay per URL, and errors are free.
Typical uses: SEO audits and site migrations (find 404s and redirects listed in your sitemap), getting a full URL list before scraping or crawling, monitoring competitors' new pages via lastmod, feeding page lists to RAG and AI agents, and content inventories.
How to extract all URLs from a website's sitemap
- Click Try for free.
- Enter websites (e.g.
apify.com) or direct sitemap links in Websites or sitemap URLs. - Optionally add include/exclude patterns and switch on Check each URL's status.
- Click Start. URLs stream into the Output tab.
- Download them as CSV, Excel or JSON, or use the API.
Input
| Option | What it does | Default |
|---|---|---|
startUrls | Websites or sitemap URLs | – |
includePatterns | Keep only URLs matching any pattern (regex or text) | – |
excludePatterns | Drop URLs matching any pattern | – |
maxUrlsPerSite | Stop after this many URLs per website | 100,000 |
checkStatus | Request each URL and report status and redirects | false |
maxConcurrency | Parallel requests | 10 |
{"startUrls": ["apify.com", "https://www.theguardian.com/sitemaps/news.xml"],"includePatterns": ["/blog/"],"checkStatus": true}
Output
One item per URL:
{"url": "https://apify.com/pricing","lastmod": "2026-09-01","changefreq": "weekly","priority": 0.8,"imageCount": 0,"site": "https://apify.com","sitemapUrl": "https://apify.com/sitemap/pages.xml","success": true,"statusCode": 200,"finalUrl": null,"redirected": false}
statusCode, finalUrl and redirected are included only when Check each URL's status is on. Websites without a sitemap are listed with success: false (free). A per-website summary is saved as SUMMARY in the key-value store. You can download the dataset in various formats such as JSON, HTML, CSV or Excel.
How much does it cost to extract a sitemap?
- $0.30 per 1,000 URLs ($0.0003 each)
- + $1 per 1,000 status checks when Check each URL's status is on
- + $0.002 per run
Example: a 5,000-page website costs $1.50, or $6.50 with status checks. Platform usage is included. Set a maximum cost per run to stay in budget; the Actor stops cleanly when it's reached.
Tips
- Use Max URLs per website for a quick sample of a huge site.
- Keep Parallel requests low (5–10) with status checks on small websites, to be polite.
- Some sites list sitemaps only in
robots.txt. That's the first place we check. - If a site blocks data-centre traffic, enable Proxy.
FAQ
The site has no sitemap? It's listed with an error and not charged. You'd need a crawler instead, such as Apify's Website Content Crawler.
Is it legal? Sitemaps are published so machines can read them. Keep request rates reasonable when checking statuses, and respect each site's terms.
Something not working? Open an issue in the Issues tab with the website and I'll fix it quickly.
Related tools
- Website Screenshot: full-page screenshots and PDFs of any website, cookie banners removed
- Document to Markdown: PDF, Word, Excel and PowerPoint files to clean Markdown for AI
- RSS Feed Reader: read and monitor any RSS/Atom feed, only new items, full article text