Sitemap Extractor & Bulk URL Status Checker avatar

Sitemap Extractor & Bulk URL Status Checker

Pricing

$1.00 / 1,000 url processeds

Go to Apify Store
Sitemap Extractor & Bulk URL Status Checker

Sitemap Extractor & Bulk URL Status Checker

Sitemap validation & URL status checker: auto-discover a site's sitemaps (robots.txt, index & .gz), extract every URL with lastmod, then check link status (HTTP status codes), redirect chains, noindex headers and response time. Or paste your own URL list.

Pricing

$1.00 / 1,000 url processeds

Rating

0.0

(0)

Developer

Forever Tools

Forever Tools

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

34 minutes ago

Last modified

Share

Give it a domain and it finds every sitemap (robots.txt Sitemap: lines, common paths such as /sitemap.xml, sitemap indexes, gzipped sitemaps), extracts every URL with lastmod, changefreq and priority, and checks each URL's health: HTTP status, redirect chain, final URL, X-Robots-Tag header, content type and response time. You can also paste any list of URLs to check in bulk.

Extract all URLs from a website's sitemap

Enter a domain or homepage in Sites or sitemap URLs and the actor discovers the sitemaps itself. You can also pass a direct .xml or .xml.gz sitemap URL. Sitemap indexes are followed, so you get the full URL inventory in one dataset, each row noting which sitemap file it came from. Turn off Check HTTP status & redirects to extract URLs only, which is faster and priced the same.

With status checking on (the default), every URL is requested and rows are tagged with issue codes: unreachable (DNS, connection or timeout problems), http-404 / http-5xx style codes for error responses, redirects, redirect-chain (more than one hop) and noindex-header. A SUMMARY record in the key-value store counts each issue, so you can see at a glance how many 404s your sitemap contains.

Check redirect chains and noindex URLs in bulk

Sitemaps should list final, indexable URLs. Rows include redirects (hop count), redirectChain (each hop with its status code), finalUrl and xRobotsTag, so you can spot sitemap entries that redirect or carry a noindex header.

Use cases

  • Weekly SEO health check: schedule the actor and alert on new 404s or redirect chains.
  • Site-migration QA: confirm every old URL redirects once to the right place.
  • Crawl-budget clean-up: find sitemap URLs that are noindex or redirecting.
  • URL inventory: get a site's complete URL list to feed a crawler or LLM pipeline.
  • Bulk URL check: paste any list into Extra URLs to check without needing a sitemap.

Input

{
"sites": ["example.com"],
"urls": ["https://example.com/some-extra-page"],
"checkStatus": true,
"maxUrls": 1000,
"maxConcurrency": 10
}
  • sites: domains, homepages or direct sitemap URLs.
  • urls: optional extra URLs to check directly.
  • checkStatus (default on), maxUrls (default 10,000), maxConcurrency (default 10, max 50), and an optional userAgent.

Output

One row per URL. Example:

{
"url": "http://example.com/old-page",
"sitemap": "https://example.com/sitemap.xml",
"lastmod": "2026-08-01",
"statusCode": 200,
"finalUrl": "https://example.com/new-page",
"redirects": 2,
"redirectChain": [
{ "url": "http://example.com/old-page", "statusCode": 301 },
{ "url": "https://example.com/old-page", "statusCode": 301 }
],
"contentType": "text/html; charset=utf-8",
"responseTimeMs": 231,
"issues": ["redirects", "redirect-chain"]
}

Other fields: changefreq, priority, xRobotsTag, error. If a site is invalid or no sitemap is found, you get an error row with site and error (invalid-site or no-sitemap-found).

Pricing

Pay per event: $0.001 per URL ($1 per 1,000 URLs), no subscription.

  • 500-URL sitemap: about $0.50
  • 10,000 URLs: about $10
  • 100,000 URLs: about $100

Max URLs caps the run, and the actor stops when the charge limit is reached. Platform usage is billed by Apify per your plan.

Limitations

  • It reports status, redirects and headers only; it does not read page content, so it cannot detect soft 404s.
  • Only sitemaps are used to discover URLs; it is not a link crawler.
  • Response times are measured from the actor's location and vary.
  • Use it on sites you own or may check.

FAQ

How do I find all URLs in a sitemap?

Enter the domain in Sites and switch off status checking; you get one row per URL with lastmod, changefreq and priority.

How do I check a sitemap for 404 errors?

Leave status checking on and filter rows whose issues include http-404, or read the counts in SUMMARY.

Does it handle sitemap index files and .gz sitemaps?

Yes, nested indexes and gzipped sitemaps are followed.

What if no sitemap is found?

You get a row with error: "no-sitemap-found"; paste the URLs into Extra URLs to check instead.

Does it download the full pages?

No. Requests are plain GETs and only status and headers are read.

How much does it cost to check a 1,000-URL sitemap?

$1 at $0.001 per URL.

Built and maintained with AI assistance. Problems or requests: use the Issues tab.

Other actors by the same developer (same flat pay-per-result pricing, no subscription):

Integrations

Run it from the Apify API, a schedule, or no-code tools: the Apify apps for Zapier, Make and n8n can start any public actor ("Run Actor") and read its dataset. AI agents can call it through the Apify MCP server.