Sitemap Extractor & Bulk URL Status Checker
Pricing
$1.00 / 1,000 url processeds
Sitemap Extractor & Bulk URL Status Checker
Sitemap validation & URL status checker: auto-discover a site's sitemaps (robots.txt, index & .gz), extract every URL with lastmod, then check link status (HTTP status codes), redirect chains, noindex headers and response time. Or paste your own URL list.
Pricing
$1.00 / 1,000 url processeds
Rating
0.0
(0)
Developer
Forever Tools
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
34 minutes ago
Last modified
Categories
Share
Give it a domain and it finds every sitemap (robots.txt Sitemap: lines, common paths such as /sitemap.xml, sitemap indexes, gzipped sitemaps), extracts every URL with lastmod, changefreq and priority, and checks each URL's health: HTTP status, redirect chain, final URL, X-Robots-Tag header, content type and response time. You can also paste any list of URLs to check in bulk.
Extract all URLs from a website's sitemap
Enter a domain or homepage in Sites or sitemap URLs and the actor discovers the sitemaps itself. You can also pass a direct .xml or .xml.gz sitemap URL. Sitemap indexes are followed, so you get the full URL inventory in one dataset, each row noting which sitemap file it came from. Turn off Check HTTP status & redirects to extract URLs only, which is faster and priced the same.
Find broken links in a sitemap
With status checking on (the default), every URL is requested and rows are tagged with issue codes: unreachable (DNS, connection or timeout problems), http-404 / http-5xx style codes for error responses, redirects, redirect-chain (more than one hop) and noindex-header. A SUMMARY record in the key-value store counts each issue, so you can see at a glance how many 404s your sitemap contains.
Check redirect chains and noindex URLs in bulk
Sitemaps should list final, indexable URLs. Rows include redirects (hop count), redirectChain (each hop with its status code), finalUrl and xRobotsTag, so you can spot sitemap entries that redirect or carry a noindex header.
Use cases
- Weekly SEO health check: schedule the actor and alert on new 404s or redirect chains.
- Site-migration QA: confirm every old URL redirects once to the right place.
- Crawl-budget clean-up: find sitemap URLs that are noindex or redirecting.
- URL inventory: get a site's complete URL list to feed a crawler or LLM pipeline.
- Bulk URL check: paste any list into Extra URLs to check without needing a sitemap.
Input
{"sites": ["example.com"],"urls": ["https://example.com/some-extra-page"],"checkStatus": true,"maxUrls": 1000,"maxConcurrency": 10}
- sites: domains, homepages or direct sitemap URLs.
- urls: optional extra URLs to check directly.
- checkStatus (default on), maxUrls (default 10,000), maxConcurrency (default 10, max 50), and an optional userAgent.
Output
One row per URL. Example:
{"url": "http://example.com/old-page","sitemap": "https://example.com/sitemap.xml","lastmod": "2026-08-01","statusCode": 200,"finalUrl": "https://example.com/new-page","redirects": 2,"redirectChain": [{ "url": "http://example.com/old-page", "statusCode": 301 },{ "url": "https://example.com/old-page", "statusCode": 301 }],"contentType": "text/html; charset=utf-8","responseTimeMs": 231,"issues": ["redirects", "redirect-chain"]}
Other fields: changefreq, priority, xRobotsTag, error. If a site is invalid or no sitemap is found, you get an error row with site and error (invalid-site or no-sitemap-found).
Pricing
Pay per event: $0.001 per URL ($1 per 1,000 URLs), no subscription.
- 500-URL sitemap: about $0.50
- 10,000 URLs: about $10
- 100,000 URLs: about $100
Max URLs caps the run, and the actor stops when the charge limit is reached. Platform usage is billed by Apify per your plan.
Limitations
- It reports status, redirects and headers only; it does not read page content, so it cannot detect soft 404s.
- Only sitemaps are used to discover URLs; it is not a link crawler.
- Response times are measured from the actor's location and vary.
- Use it on sites you own or may check.
FAQ
How do I find all URLs in a sitemap?
Enter the domain in Sites and switch off status checking; you get one row per URL with lastmod, changefreq and priority.
How do I check a sitemap for 404 errors?
Leave status checking on and filter rows whose issues include http-404, or read the counts in SUMMARY.
Does it handle sitemap index files and .gz sitemaps?
Yes, nested indexes and gzipped sitemaps are followed.
What if no sitemap is found?
You get a row with error: "no-sitemap-found"; paste the URLs into Extra URLs to check instead.
Does it download the full pages?
No. Requests are plain GETs and only status and headers are read.
How much does it cost to check a 1,000-URL sitemap?
$1 at $0.001 per URL.
Built and maintained with AI assistance. Problems or requests: use the Issues tab.
Related tools
Other actors by the same developer (same flat pay-per-result pricing, no subscription):
- Apple App Store Reviews Scraper (Multi-Country)
- Article Extractor – Clean Text & Markdown for LLM/RAG
- Company Jobs Scraper: Workday, Greenhouse, Lever, Ashby
- Bulk Domain Checker — WHOIS/RDAP, DNS, SPF/DMARC, SSL Expiry
- Bulk PageSpeed Insights & Core Web Vitals Checker
- PDF to Text Extractor (Bulk, with Metadata)
- Website SEO Audit Crawler
- Website Tech Stack Detector (CMS, Framework, Analytics)
- Website Screenshot – Bulk Full Page PNG, JPEG & PDF
Integrations
Run it from the Apify API, a schedule, or no-code tools: the Apify apps for Zapier, Make and n8n can start any public actor ("Run Actor") and read its dataset. AI agents can call it through the Apify MCP server.