Sitemap Extractor – All URLs from sitemap.xml + Status Check
Pricing
from $0.15 / 1,000 url extracteds
Sitemap Extractor – All URLs from sitemap.xml + Status Check
Extract every URL from any website's sitemap.xml: auto-discovery via robots.txt, nested sitemap indexes, .xml.gz, image/news/video/hreflang tags and lastmod. Filter by date or glob, get only new URLs, and optionally check HTTP status, redirects and canonical per URL.
Pricing
from $0.15 / 1,000 url extracteds
Rating
0.0
(0)
Developer
Cemal Atakli
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
7 days ago
Last modified
Categories
Share
Enter a domain and get back every URL from its sitemap. The Actor finds the sitemap itself: it reads the robots.txt Sitemap: lines, then tries /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml and /sitemap.xml.gz. It follows nested sitemap indexes, unpacks gzip (.xml.gz) files and reads image, news, video and hreflang tags as well as lastmod, changefreq and priority. You can also turn on an HTTP status check for every URL to find 404s, redirect chains, non-canonical pages and noindex pages listed in your sitemap.
- Built to succeed. It retries on 429/5xx, uses a regex fallback for broken XML, skips HTML "soft 404" pages served as
sitemap.xml, detects index loops and records an error per sitemap instead of failing the run. A site with no sitemap is reported clearly, not crashed on. - Every sitemap format. Supports XML
urlsetandsitemapindex(up to 10 levels deep),.xml.gz(gzip payload or transfer encoding), plain-textsitemap.txt, and RSS/Atom feeds used as sitemaps. Google image (image:loc), news (title, publication, language, date), video (title, thumbnail, content/player URL, duration) and hreflangxhtml:linkalternates are all extracted. - Bulk URL status checker. It sends HEAD requests and falls back to GET when a server rejects HEAD. For each URL you get
status,finalUrl,redirectHopsandredirectChain,contentType,responseMsandX-Robots-Tag, and optionally the canonical URL and meta robots tag. It runs 20 requests in parallel but never more than 5 per host. - Filters. Keep URLs modified since a date (
2025-01-31or7 days), or filter by include/exclude globs (/blog/*,*/products/*,*.pdf). - "Only new URLs" monitoring. The Actor remembers which URLs it has already returned for each site, so on a schedule each run returns only new pages: new products, articles or landing pages from a competitor.
- Cheap. It uses plain HTTP with no browser and no proxy. $0.15 per 1,000 URLs, about 3× cheaper than the official Apify sitemap extractor.
What can I use it for?
- SEO audits: list every indexable URL, then find 404s, redirect chains and pages that canonicalise elsewhere, all of which are sitemap errors that waste crawl budget
- Competitor monitoring: schedule a daily run with Only new URLs to see every new product, blog post or landing page a competitor publishes
- Crawler seeding: feed a complete URL list to Website Content Crawler, a RAG pipeline or your own scraper instead of discovering pages by link-following
- Site migrations: export all old URLs with their status before a relaunch, then check that each one returns a 200 or a single 301 afterwards
- News and content tracking: read Google News sitemaps (title, publication, date) from publishers in near real time
- Image and video inventory: list every image and video URL that a site exposes to Google
- International SEO: export the hreflang alternates of every page to audit language versions
Input
| Field | Description |
|---|---|
startUrls | Domains (example.com), sitemap URLs (…/sitemap_index.xml, .xml.gz, .txt, feeds) or robots.txt URLs, in any mix |
maxUrls | Stop after this many URLs in total (default 1,000, 0 = everything) |
maxUrlsPerSite | Optional per-site limit |
checkStatus | HTTP status, final URL, redirects, content type, response time and X-Robots-Tag per URL |
extractCanonical | Also read <link rel="canonical"> and meta robots (GET, first 64 KB only) |
includeGlobs / excludeGlobs | URL patterns to keep or drop |
lastmodSince | Only URLs with lastmod on or after a date (absolute or relative) |
keepUrlsWithoutLastmod | Keep URLs that have no lastmod when a date filter is set |
onlyNew, stateKey | Return only URLs not seen in previous runs (monitoring) |
includeSiteSummaryRows, discoverCommonPaths, maxSitemapDepth, maxConcurrency, statusConcurrency, requestTimeoutSecs, proxyConfiguration | Advanced |
{"startUrls": ["allbirds.com", "https://www.nytimes.com/sitemaps/new/news.xml.gz"],"maxUrls": 5000,"checkStatus": true,"excludeGlobs": ["*/collections/*"],"lastmodSince": "30 days"}
Output
Each URL is one row. The Output tab has three tables, Sitemap URLs, HTTP status check and Images, videos, news & hreflang, plus a per-site summary: sitemaps found, URL counts, newest and oldest lastmod, a status-code histogram and any failed sitemaps. Here is a real row (shortened) with the status check on:
{"url": "https://www.allbirds.com/products/mens-wool-runners","site": "allbirds.com","sitemapUrl": "https://www.allbirds.com/sitemap_products_1.xml?from=1878193471557&to=7369944137808","lastmod": "2026-10-01T01:54:12-07:00","changefreq": "daily","priority": null,"images": ["https://cdn.shopify.com/s/files/1/1104/4168/files/Allbirds_WL_RN_SF_PDP_Natural_Grey_LAT.png?v=1751143404"],"imageCount": 1,"alternates": [],"videos": [],"newsTitle": null,"status": 200,"finalUrl": "https://www.allbirds.com/products/mens-wool-runners","redirectHops": 0,"redirectChain": [],"contentType": "text/html","responseMs": 412,"xRobotsTag": null,"statusError": null}
A news sitemap row adds newsTitle, newsPublicationDate, newsPublicationName and newsLanguage. A video sitemap row has videos: [{title, thumbnailUrl, contentUrl, playerUrl, durationSecs}]. See SAMPLE_OUTPUT.json for real results from 6 sites.
Pricing
Pay per event:
| Event | Price |
|---|---|
| URL extracted (one row in the dataset) | $0.00015 ($0.15 per 1,000 URLs) |
URL status checked (only with checkStatus) | $0.0004 ($0.40 per 1,000 URLs) |
| Actor start | $0.00005 |
| Duplicates, filtered-out URLs, sitemaps not found, summary | free |
How that compares with other sitemap and status tools in the Apify Store (listed prices, October 2026):
| Actor | Listed price | 30-day run success (Store) |
|---|---|---|
| apify/sitemap-extractor | $0.50 / 1,000 URLs | 69% |
| crawlerbros/sitemap-url-extractor | $2 / 1,000 URLs | 63% |
| onescales sitemap extractor | $30 / 1,000 URLs | – |
| logiover URL status checker | $2.50 / 1,000 URLs | – |
| khadinakbar URL status checker | $1 / 1,000 URLs | – |
| This Actor | $0.15 / 1,000 URLs, or $0.55 with status check | new |
Set a maximum cost per run in the run options. The Actor saves as many URLs as fit and then stops cleanly. It never runs a status check that it cannot bill.
FAQ
The site has no sitemap. What happens? The summary says No sitemap found and you pay nothing for that site. This Actor does not crawl links. For sites without a sitemap, use a crawler instead.
Does it respect robots.txt? Sitemaps and robots.txt are published specifically for machines to read, and the Actor uses robots.txt only to discover sitemaps. The optional status check makes one lightweight HEAD (or GET) request per URL, with at most 5 parallel requests per host.
How big a sitemap can it handle?
The protocol maximum of 50,000 URLs or 50 MB per file, gzipped or not, and thousands of child sitemaps per index. The XML is stream-parsed with lxml, so memory stays low. Set maxUrls: 0 to extract everything.
Why do some URLs have no lastmod?
The site's sitemap does not provide one. Many Shopify and Next.js sitemaps omit it for some pages. With lastmodSince set these URLs are dropped unless you enable keepUrlsWithoutLastmod.
How does "Only new URLs" work?
A short hash of every URL already returned is stored per site in a named key-value store (sitemap-url-extractor-state) in your account. The next run returns only URLs that are not in that store. The first run returns everything. Use stateKey to keep separate histories.
HEAD or GET?
HEAD by default, because it is fastest. If a server answers HEAD with 400, 403, 404, 405, 406 or 501, the URL is retried with GET so you get the real status. extractCanonical always uses GET and reads at most 64 KB per page.
Can I get the results as CSV or Excel? Yes. Every Apify dataset can be exported to CSV, Excel, JSON or XML, or sent to Google Sheets with an integration.
Use with AI agents / Apify MCP
This Actor makes a good agent tool. The input is tiny (a domain), the output is a clean URL list, and 1,000 URLs cost $0.15. Connect it through the Apify MCP server (https://mcp.apify.com?actors=gazidev/sitemap-url-extractor) and Claude, ChatGPT, Cursor and other MCP clients can call it with prompts such as "list all blog posts example.com published in the last 30 days" or "find broken URLs in competitor.com's sitemap". Over the API: POST https://api.apify.com/v2/acts/gazidev~sitemap-url-extractor/run-sync-get-dataset-items with the input JSON.
Categories
SEO tools · Developer tools · Automation