Sitemap Extractor – All URLs from sitemap.xml + Status Check avatar

Sitemap Extractor – All URLs from sitemap.xml + Status Check

Pricing

from $0.15 / 1,000 url extracteds

Go to Apify Store
Sitemap Extractor – All URLs from sitemap.xml + Status Check

Sitemap Extractor – All URLs from sitemap.xml + Status Check

Extract every URL from any website's sitemap.xml: auto-discovery via robots.txt, nested sitemap indexes, .xml.gz, image/news/video/hreflang tags and lastmod. Filter by date or glob, get only new URLs, and optionally check HTTP status, redirects and canonical per URL.

Pricing

from $0.15 / 1,000 url extracteds

Rating

0.0

(0)

Developer

Cemal Atakli

Cemal Atakli

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

7 days ago

Last modified

Categories

Share

Enter a domain and get back every URL from its sitemap. The Actor finds the sitemap itself: it reads the robots.txt Sitemap: lines, then tries /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml and /sitemap.xml.gz. It follows nested sitemap indexes, unpacks gzip (.xml.gz) files and reads image, news, video and hreflang tags as well as lastmod, changefreq and priority. You can also turn on an HTTP status check for every URL to find 404s, redirect chains, non-canonical pages and noindex pages listed in your sitemap.

  • Built to succeed. It retries on 429/5xx, uses a regex fallback for broken XML, skips HTML "soft 404" pages served as sitemap.xml, detects index loops and records an error per sitemap instead of failing the run. A site with no sitemap is reported clearly, not crashed on.
  • Every sitemap format. Supports XML urlset and sitemapindex (up to 10 levels deep), .xml.gz (gzip payload or transfer encoding), plain-text sitemap.txt, and RSS/Atom feeds used as sitemaps. Google image (image:loc), news (title, publication, language, date), video (title, thumbnail, content/player URL, duration) and hreflang xhtml:link alternates are all extracted.
  • Bulk URL status checker. It sends HEAD requests and falls back to GET when a server rejects HEAD. For each URL you get status, finalUrl, redirectHops and redirectChain, contentType, responseMs and X-Robots-Tag, and optionally the canonical URL and meta robots tag. It runs 20 requests in parallel but never more than 5 per host.
  • Filters. Keep URLs modified since a date (2025-01-31 or 7 days), or filter by include/exclude globs (/blog/*, */products/*, *.pdf).
  • "Only new URLs" monitoring. The Actor remembers which URLs it has already returned for each site, so on a schedule each run returns only new pages: new products, articles or landing pages from a competitor.
  • Cheap. It uses plain HTTP with no browser and no proxy. $0.15 per 1,000 URLs, about 3× cheaper than the official Apify sitemap extractor.

What can I use it for?

  • SEO audits: list every indexable URL, then find 404s, redirect chains and pages that canonicalise elsewhere, all of which are sitemap errors that waste crawl budget
  • Competitor monitoring: schedule a daily run with Only new URLs to see every new product, blog post or landing page a competitor publishes
  • Crawler seeding: feed a complete URL list to Website Content Crawler, a RAG pipeline or your own scraper instead of discovering pages by link-following
  • Site migrations: export all old URLs with their status before a relaunch, then check that each one returns a 200 or a single 301 afterwards
  • News and content tracking: read Google News sitemaps (title, publication, date) from publishers in near real time
  • Image and video inventory: list every image and video URL that a site exposes to Google
  • International SEO: export the hreflang alternates of every page to audit language versions

Input

FieldDescription
startUrlsDomains (example.com), sitemap URLs (…/sitemap_index.xml, .xml.gz, .txt, feeds) or robots.txt URLs, in any mix
maxUrlsStop after this many URLs in total (default 1,000, 0 = everything)
maxUrlsPerSiteOptional per-site limit
checkStatusHTTP status, final URL, redirects, content type, response time and X-Robots-Tag per URL
extractCanonicalAlso read <link rel="canonical"> and meta robots (GET, first 64 KB only)
includeGlobs / excludeGlobsURL patterns to keep or drop
lastmodSinceOnly URLs with lastmod on or after a date (absolute or relative)
keepUrlsWithoutLastmodKeep URLs that have no lastmod when a date filter is set
onlyNew, stateKeyReturn only URLs not seen in previous runs (monitoring)
includeSiteSummaryRows, discoverCommonPaths, maxSitemapDepth, maxConcurrency, statusConcurrency, requestTimeoutSecs, proxyConfigurationAdvanced
{
"startUrls": ["allbirds.com", "https://www.nytimes.com/sitemaps/new/news.xml.gz"],
"maxUrls": 5000,
"checkStatus": true,
"excludeGlobs": ["*/collections/*"],
"lastmodSince": "30 days"
}

Output

Each URL is one row. The Output tab has three tables, Sitemap URLs, HTTP status check and Images, videos, news & hreflang, plus a per-site summary: sitemaps found, URL counts, newest and oldest lastmod, a status-code histogram and any failed sitemaps. Here is a real row (shortened) with the status check on:

{
"url": "https://www.allbirds.com/products/mens-wool-runners",
"site": "allbirds.com",
"sitemapUrl": "https://www.allbirds.com/sitemap_products_1.xml?from=1878193471557&to=7369944137808",
"lastmod": "2026-10-01T01:54:12-07:00",
"changefreq": "daily",
"priority": null,
"images": ["https://cdn.shopify.com/s/files/1/1104/4168/files/Allbirds_WL_RN_SF_PDP_Natural_Grey_LAT.png?v=1751143404"],
"imageCount": 1,
"alternates": [],
"videos": [],
"newsTitle": null,
"status": 200,
"finalUrl": "https://www.allbirds.com/products/mens-wool-runners",
"redirectHops": 0,
"redirectChain": [],
"contentType": "text/html",
"responseMs": 412,
"xRobotsTag": null,
"statusError": null
}

A news sitemap row adds newsTitle, newsPublicationDate, newsPublicationName and newsLanguage. A video sitemap row has videos: [{title, thumbnailUrl, contentUrl, playerUrl, durationSecs}]. See SAMPLE_OUTPUT.json for real results from 6 sites.

Pricing

Pay per event:

EventPrice
URL extracted (one row in the dataset)$0.00015 ($0.15 per 1,000 URLs)
URL status checked (only with checkStatus)$0.0004 ($0.40 per 1,000 URLs)
Actor start$0.00005
Duplicates, filtered-out URLs, sitemaps not found, summaryfree

How that compares with other sitemap and status tools in the Apify Store (listed prices, October 2026):

ActorListed price30-day run success (Store)
apify/sitemap-extractor$0.50 / 1,000 URLs69%
crawlerbros/sitemap-url-extractor$2 / 1,000 URLs63%
onescales sitemap extractor$30 / 1,000 URLs–
logiover URL status checker$2.50 / 1,000 URLs–
khadinakbar URL status checker$1 / 1,000 URLs–
This Actor$0.15 / 1,000 URLs, or $0.55 with status checknew

Set a maximum cost per run in the run options. The Actor saves as many URLs as fit and then stops cleanly. It never runs a status check that it cannot bill.

FAQ

The site has no sitemap. What happens? The summary says No sitemap found and you pay nothing for that site. This Actor does not crawl links. For sites without a sitemap, use a crawler instead.

Does it respect robots.txt? Sitemaps and robots.txt are published specifically for machines to read, and the Actor uses robots.txt only to discover sitemaps. The optional status check makes one lightweight HEAD (or GET) request per URL, with at most 5 parallel requests per host.

How big a sitemap can it handle? The protocol maximum of 50,000 URLs or 50 MB per file, gzipped or not, and thousands of child sitemaps per index. The XML is stream-parsed with lxml, so memory stays low. Set maxUrls: 0 to extract everything.

Why do some URLs have no lastmod? The site's sitemap does not provide one. Many Shopify and Next.js sitemaps omit it for some pages. With lastmodSince set these URLs are dropped unless you enable keepUrlsWithoutLastmod.

How does "Only new URLs" work? A short hash of every URL already returned is stored per site in a named key-value store (sitemap-url-extractor-state) in your account. The next run returns only URLs that are not in that store. The first run returns everything. Use stateKey to keep separate histories.

HEAD or GET? HEAD by default, because it is fastest. If a server answers HEAD with 400, 403, 404, 405, 406 or 501, the URL is retried with GET so you get the real status. extractCanonical always uses GET and reads at most 64 KB per page.

Can I get the results as CSV or Excel? Yes. Every Apify dataset can be exported to CSV, Excel, JSON or XML, or sent to Google Sheets with an integration.

Use with AI agents / Apify MCP

This Actor makes a good agent tool. The input is tiny (a domain), the output is a clean URL list, and 1,000 URLs cost $0.15. Connect it through the Apify MCP server (https://mcp.apify.com?actors=gazidev/sitemap-url-extractor) and Claude, ChatGPT, Cursor and other MCP clients can call it with prompts such as "list all blog posts example.com published in the last 30 days" or "find broken URLs in competitor.com's sitemap". Over the API: POST https://api.apify.com/v2/acts/gazidev~sitemap-url-extractor/run-sync-get-dataset-items with the input JSON.

Categories

SEO tools · Developer tools · Automation