Sitemap Extractor - All Website URLs + Broken Link Check avatar

Sitemap Extractor - All Website URLs + Broken Link Check

Pricing

from $0.30 / 1,000 urls

Go to Apify Store
Sitemap Extractor - All Website URLs + Broken Link Check

Sitemap Extractor - All Website URLs + Broken Link Check

Find every page of a website from its sitemaps (robots.txt, sitemap indexes, .gz). Returns URLs with last-modified dates, filters by pattern, and can check each URL for 404s and redirects.

Pricing

from $0.30 / 1,000 urls

Rating

0.0

(0)

Developer

Björn Ólafur

Björn Ólafur

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What does Sitemap Extractor do?

Sitemap Extractor finds every page of a website from its XML sitemaps. Just enter a domain: it reads robots.txt, checks the common sitemap locations, follows sitemap indexes and .gz files, and returns each URL with its last-modified date, change frequency and priority. Optionally it checks every URL for 404s and redirects, a quick broken-link audit of your sitemap.

Run it in Apify Console, via the API, on a schedule, or from your AI agent through the Apify MCP server. Connect it to Google Sheets, Make, Zapier or n8n with Apify integrations.

Why use Sitemap Extractor?

  • 🗺️ Automatic discovery: robots.txt → sitemap.xml, sitemap_index.xml, wp-sitemap.xml, .gz and text sitemaps.
  • 🔁 Follows nested sitemap indexes, with duplicates removed.
  • 🎯 Filters: keep only /blog/ or /products/ URLs, or drop /tag/ pages.
  • 🩺 Broken-link check (optional): HTTP status and redirect target for every URL.
  • 🖼️ Image counts from image sitemaps, and news/video sitemaps are supported.
  • 💸 Cheap and fair: pay per URL, and errors are free.

Typical uses: SEO audits and site migrations (find 404s and redirects listed in your sitemap), getting a full URL list before scraping or crawling, monitoring competitors' new pages via lastmod, feeding page lists to RAG and AI agents, and content inventories.

How to extract all URLs from a website's sitemap

  1. Click Try for free.
  2. Enter websites (e.g. apify.com) or direct sitemap links in Websites or sitemap URLs.
  3. Optionally add include/exclude patterns and switch on Check each URL's status.
  4. Click Start. URLs stream into the Output tab.
  5. Download them as CSV, Excel or JSON, or use the API.

Input

OptionWhat it doesDefault
startUrlsWebsites or sitemap URLs–
includePatternsKeep only URLs matching any pattern (regex or text)–
excludePatternsDrop URLs matching any pattern–
maxUrlsPerSiteStop after this many URLs per website100,000
checkStatusRequest each URL and report status and redirectsfalse
maxConcurrencyParallel requests10
{
"startUrls": ["apify.com", "https://www.theguardian.com/sitemaps/news.xml"],
"includePatterns": ["/blog/"],
"checkStatus": true
}

Output

One item per URL:

{
"url": "https://apify.com/pricing",
"lastmod": "2026-09-01",
"changefreq": "weekly",
"priority": 0.8,
"imageCount": 0,
"site": "https://apify.com",
"sitemapUrl": "https://apify.com/sitemap/pages.xml",
"success": true,
"statusCode": 200,
"finalUrl": null,
"redirected": false
}

statusCode, finalUrl and redirected are included only when Check each URL's status is on. Websites without a sitemap are listed with success: false (free). A per-website summary is saved as SUMMARY in the key-value store. You can download the dataset in various formats such as JSON, HTML, CSV or Excel.

How much does it cost to extract a sitemap?

  • $0.30 per 1,000 URLs ($0.0003 each)
  • + $1 per 1,000 status checks when Check each URL's status is on
  • + $0.002 per run

Example: a 5,000-page website costs $1.50, or $6.50 with status checks. Platform usage is included. Set a maximum cost per run to stay in budget; the Actor stops cleanly when it's reached.

Tips

  • Use Max URLs per website for a quick sample of a huge site.
  • Keep Parallel requests low (5–10) with status checks on small websites, to be polite.
  • Some sites list sitemaps only in robots.txt. That's the first place we check.
  • If a site blocks data-centre traffic, enable Proxy.

FAQ

The site has no sitemap? It's listed with an error and not charged. You'd need a crawler instead, such as Apify's Website Content Crawler.

Is it legal? Sitemaps are published so machines can read them. Keep request rates reasonable when checking statuses, and respect each site's terms.

Something not working? Open an issue in the Issues tab with the website and I'll fix it quickly.