Sitemap Extractor & Change Monitor - All URLs, New Pages
Pricing
from $0.30 / 1,000 url results
Sitemap Extractor & Change Monitor - All URLs, New Pages
Extract every URL from any website's sitemaps with lastmod, changefreq & priority - and monitor which pages were added, removed or updated.
Pricing
from $0.30 / 1,000 url results
Rating
0.0
(0)
Developer
Dave West
Maintained by CommunityActor stats
1
Bookmarked
2
Total users
1
Monthly active users
21 hours ago
Last modified
Categories
Share
Sitemap Extractor & Change Monitor - get all URLs from any website, track new and removed pages
Paste a website, get every URL it publishes - with lastmod, changefreq and priority - as JSON, CSV or Excel. Turn on monitor mode to get only the pages that were added, removed or updated since your last run.
What you get
- ✅ All URLs of a website from its sitemaps - the Actor finds the sitemaps for you (
robots.txt,/sitemap.xml,/sitemap_index.xml, WordPress/Yoast locations), follows nested sitemap indexes to any depth and reads gzipped.xml.gz, plain-text, RSS/Atom and Google News sitemaps. - ✅ A clean, de-duplicated URL list with last-modified date, change frequency, priority and the sitemap each URL came from - export it as CSV, Excel, JSON, or read it via API.
- ✅ Website change monitoring - schedule it daily and get a feed of new pages, removed pages and updated pages per site, with built-in protection against false alarms.
- ✅ Reliable on big and messy sites - streaming parser tested on a site with 6.4 million URLs, automatic retries, broken-XML recovery, and a per-sitemap report that explains every problem instead of failing the run.
- ✅ Predictable cost - pay per result, with a default cap of 100,000 results per site so a huge site never surprises you.
It only downloads robots.txt and the sitemap files - it does not crawl the pages themselves - so it is fast, cheap and gentle on the target site.
Quick start: extract all URLs from a website
- Enter one or more websites (
https://example.comor justexample.com) or sitemap URLs (https://example.com/sitemap.xml). - Click Start.
- Open the Output tab and download the URL list as CSV, Excel or JSON - or use the API.
That's it - no sitemap URL needed. If you paste a URL that looks like a sitemap but isn't one, the Actor falls back to discovering the site's real sitemaps.
Use cases
Get a full list of all pages on a website
Export every URL a site submits to search engines - for content inventories, site migrations (compare old and new URL lists so nothing gets lost), redirect mapping, or as a seed list for your own crawler or scraper.
SEO audit of sitemaps
Find out what a site really tells Google: sitemaps that return 404 or an HTML page instead of XML, broken gzip files, URLs blocked by robots.txt, stale or fake lastmod dates, http:// URLs in an https:// site, duplicates across sitemaps. The run summary lists every sitemap file with its HTTP status, URL count, size, retries and exact error.
Competitor content monitoring - get alerted to new pages
Schedule a daily run on competitor websites with monitor mode on. Each run outputs only what changed: the new product pages, landing pages, blog posts, docs or pricing pages they published - and the ones they quietly removed. Connect the run to Slack, email, a webhook, Make, Zapier or n8n to get the changes delivered.
News and publication tracking
Point it at publishers' news sitemaps (e.g. https://www.nytimes.com/sitemaps/new/news.xml.gz) and run it every few minutes or hours to collect new articles shortly after they are published. Google News publication_date is used as lastmod.
Feed AI, RAG and scraping pipelines with only what changed
Use the URL list - or just the added/updated URLs from monitor mode - as input for Website Content Crawler or your own scraper, so you only re-scrape pages that actually changed. Works well from AI agents and LLM tools through the Apify API and MCP server.
Indexing workflows
Send new and updated URLs to IndexNow, the Google Indexing API or your internal search index.
Sitemap monitoring: new, removed and updated pages
With Monitor mode on, the Actor saves each site's URL list (URL + lastmod, compressed) in a named key-value store in your account (sitemap-extractor-monitor-state by default) and compares against it on every run. Each output row gets a changeType:
changeType | Meaning |
|---|---|
initial | First run for this site - the baseline. Switch off Output all URLs on the first (baseline) run to save it without paying for rows. |
added | New page: in the sitemaps now, not in the previous run. |
removed | Page no longer listed in any sitemap. |
lastmodChanged | Page still listed, but its <lastmod> changed (the page was updated). previousLastmod holds the old value. |
unchanged | Only output when Monitor output is set to All URLs. |
Built-in safeguards, so you can trust the alerts:
- No false "removed" alerts. If any sitemap of a site fails (timeout, server error, broken gzip, HTML instead of XML) or a limit cuts the scan short, removals are not reported for that site and unconfirmed URLs stay in the state. A listed sitemap that cleanly returns 404/410 counts as genuinely gone.
- No fake "updated" floods. Some sites stamp every URL with the time the sitemap was generated. The Actor detects this per sitemap file and ignores those
lastmodchanges (reported aslastmodIgnored). - Nothing is lost. If a site has more changes than
maxResultsPerSitein one run, the rest are reported on the next run. If your Maximum cost per run is reached, the state is not updated, so the next run picks up everything. - Filters are part of the state key, so changing URL filters starts a fresh baseline instead of reporting thousands of fake changes. Use a different State store name to run independent monitors of the same site.
Input example
{"startUrls": [{ "url": "https://nodejs.org" },{ "url": "https://www.theguardian.com" },{ "url": "https://www.nytimes.com/sitemaps/new/news.xml.gz" }],"monitorMode": true,"outputOnFirstRun": false,"urlExcludePatterns": ["/tag/", "\\?page="]}
All options are optional except startUrls:
| Field | Default | Description |
|---|---|---|
startUrls | - | Websites or sitemap URLs (also accepts a link to a text file with URLs). |
monitorMode | false | Compare with the previous run and output changes. |
monitorOutput | changes | changes or all (all URLs with a changeType). |
outputOnFirstRun | true | Output the baseline on the first monitor run. |
monitorStoreName | sitemap-extractor-monitor-state | Named key-value store holding the state. |
urlIncludePatterns / urlExcludePatterns | - | Regular expressions to keep / drop page URLs (e.g. only /blog/). |
sitemapExcludePatterns | - | Skip child sitemaps by URL (e.g. old archive years). |
modifiedSince | - | Only URLs with lastmod on/after a date (2024-05-01) or period (7 days). |
maxResultsPerSite | 100000 | Cost cap: max result rows per site per run; 0 = no limit. |
maxUrlsPerSite | 0 (no limit) | Stop reading a site after N unique URLs (bounds time and memory; a truncated scan cannot detect removals). |
maxSitemapsPerSite | 10000 | Safety limit on sitemap files per site. |
respectRobotsTxt | true | Honour robots.txt rules and Crawl-delay. |
probeCommonLocations | true | Also try /sitemap.xml, /sitemap_index.xml and other common locations. |
maxConcurrency / maxRequestsPerHost | 10 / 2 | Total and per-host parallel downloads. |
minDelayBetweenRequestsMs | 100 | Minimum gap between requests to the same host. |
requestTimeoutSecs / maxRetries | 60 / 4 | Timeout and retries per sitemap file. |
proxyConfiguration | none | Apify Proxy or your own proxies, for sites that block cloud IPs. |
Output: sitemap URLs as CSV, Excel or JSON
One dataset row per URL:
{"url": "https://apify.com/store","lastmod": "2024-05-01T08:00:00.000Z","changefreq": "weekly","priority": 0.8,"sitemapUrl": "https://apify.com/sitemap/pages.xml","site": "https://apify.com","extractedAt": "2026-10-01T12:00:00.000Z"}
In monitor mode each row also has changeType and, where relevant, previousLastmod:
{"url": "https://www.nytimes.com/live/2026/10/01/us/midterms-elections","lastmod": "2026-10-01T14:45:03.000Z","changefreq": null,"priority": null,"sitemapUrl": "https://www.nytimes.com/sitemaps/new/news.xml.gz","site": "https://www.nytimes.com/sitemaps/new/news.xml.gz","changeType": "lastmodChanged","previousLastmod": "2026-10-01T14:22:15.000Z","extractedAt": "2026-10-01T14:49:54.000Z"}
lastmod is normalised to ISO-8601 UTC (common malformed formats are repaired; values that cannot be parsed are kept as they are). URLs are de-duplicated per site (fragments removed, host lower-cased). The dataset has two views: All URLs and Changes (monitor mode). Download as CSV, Excel, JSON, XML or HTML from the Output tab, or via the API.
Run summary (SUMMARY in the key-value store)
For each input site you also get a report - a ready-made sitemap health check:
{"input": "https://www.bbc.com/","site": "https://www.bbc.com","mode": "site","status": "complete","durationMs": 133298,"robotsTxt": { "url": "https://www.bbc.com/robots.txt", "httpStatus": 200, "found": true, "sitemapsListed": 38 },"sitemapsDiscovered": 197,"sitemapsFetched": 196,"sitemapsFailed": 0,"sitemapsNotFound": 1,"sitemapsBlockedByRobots": 0,"urlsFound": 6407165,"urlsOutput": 6407165,"duplicateUrls": 562171,"invalidUrls": 0,"filteredOutUrls": 0,"truncated": false,"errors": [],"warnings": ["https://www.bbc.com/tajik/sitemap.xml: notFound - HTTP 404"],"sitemaps": [{ "url": "https://www.bbc.com/afrique/sitemap.xml", "source": "robots", "depth": 0, "status": "ok", "httpStatus": 200, "type": "urlset", "urlCount": 100, "childSitemapCount": 0, "bytes": 16330, "gzipped": false, "attempts": 1, "durationMs": 74 }]}
In monitor mode the report also has a monitor block (added, removed, lastmodChanged, lastmodIgnored, changesCarriedOver, stateSaved, ...).
Site status is complete (every sitemap read), partial (some sitemaps failed or a limit was hit - you still get everything that could be read), failed (sitemaps exist but none could be read) or noSitemap (the site publishes no sitemap).
Pricing
Pay per event - you pay for results, not for compute time:
| Event | Price | When |
|---|---|---|
| URL result | $0.30 per 1,000 rows | Each row saved to the dataset (in monitor mode: each added / removed / updated URL). |
| Site check | $0.01 | Each site processed in a run. Sites with more than 25,000 sitemap URLs count one site check per started 25,000 URLs scanned. |
| Example | Cost |
|---|---|
| Extract all URLs of a 5,000-page website once | $0.01 + $1.50 = $1.51 |
| Monitor a 10,000-page site daily, ~100 changes per day (baseline saved without output) | $0.01 + $0.03 = |
| Monitor 50 competitor sites (20,000 URLs each) daily, ~1% change | $0.50 + $3.00 = ~$3.50/day |
| Monitor a 600,000-URL site daily, ~2,000 changes per day | $0.24 + $0.60 = ~$0.84/day |
No surprise bills. Each site is capped at 100,000 result rows per run by default (maxResultsPerSite, about $30 per site at list price). Raise it or set 0 to get every URL of a multi-million-URL site. The cap never breaks monitoring: the whole sitemap is still scanned, removals are still detected, and changes over the cap are reported on the next run. You can also set Maximum cost per run - the Actor stops cleanly when it is reached. Use modifiedSince or URL filters if you only need part of a site.
Why it is reliable
| Finds sitemaps you would miss | robots.txt Sitemap: lines plus common locations (/sitemap.xml, /sitemap_index.xml, WordPress, Yoast, .gz, .txt); nested indexes to any depth. |
| No URL limits in the parser | Streaming parser - memory does not depend on sitemap file size. Tested with 6.4 million URLs from one site. Many simple sitemap tools stop at a few thousand URLs or one level of index nesting. |
| Survives broken sitemaps | Error-tolerant XML parsing (unescaped &, unclosed tags, junk, BOM); detects HTML soft-404 pages served instead of XML; handles gzip with or without Content-Encoding. |
| Survives network trouble | Retries with exponential backoff for timeouts, resets, truncated downloads, 408/429/5xx; honours Retry-After; stalled downloads time out. |
| Never fails the whole run for one bad site | Each site gets its own status and error report; everything that could be read is delivered. Runs out of memory? It stops gracefully and keeps what it collected. |
| Polite by default | Respects robots.txt (RFC 9309) including Crawl-delay; 2 parallel requests and a 100 ms gap per host. |
Performance
| Site | Sitemap files | Unique URLs | Time | Result |
|---|---|---|---|---|
| bbc.com | 196 | 6,407,165 | 2 min 14 s | complete |
| apify.com | 17 | 670,618 | 6-9 s | complete |
| techcrunch.com (WordPress) | 2,062 | 408,699 | 5 min 33 s | complete |
| wordpress.org | 273 | 184,753 | 72 s | complete |
| github.blog | 26 | 9,556 | 5 s | complete |
| docs.apify.com | 7 | 4,053 | 4 s | complete |
| nodejs.org | 2 | 1,740 | 4 s | complete |
The small sites were measured on the Apify platform (512 MB-1 GB, peak memory under 60 MB). The large ones were benchmarked outside the platform with the results cap off and dataset writes disabled, so on the platform add time for saving millions of rows. Run time depends mostly on the number of sitemap files and the polite per-host limit, not on the number of URLs.
Memory: roughly 250 bytes per unique URL (monitor mode holds the previous and the current list), so 1 GB handles about 2-3 million URLs per run - use 2-4 GB for the very largest news sites.
FAQ
How do I get a list of all URLs (all pages) on a website? Enter the website's address and click Start. The Actor finds the site's sitemaps and returns every listed page URL. This covers the pages the site itself publishes for search engines; pages that are not in any sitemap are not included (use a link crawler such as Website Content Crawler for those).
How do I find the sitemap of a website?
You don't need to. The Actor checks robots.txt and the usual locations automatically. The run summary lists every sitemap it found, with its URL and status - useful if you just want to know where a site's sitemaps are.
Can I convert a sitemap.xml to CSV or Excel? Yes. Paste the sitemap URL (or the website), run it, and download the dataset as CSV or Excel from the Output tab. Sitemap indexes and gzipped sitemaps are expanded automatically.
How do I get alerted when a competitor publishes a new page?
Turn on Monitor mode, add the competitor's website, and create a daily schedule in Apify Console. Each run outputs only added, removed and lastmodChanged URLs. Add a Slack, email or webhook integration to the run to receive them automatically.
Does it crawl the pages or check their status codes?
No. It only downloads robots.txt and sitemap files - typically a handful of requests per site - so it is fast and inexpensive. It reports the status of every sitemap file, not of every page.
What if the site has no sitemap, or blocks requests?
Sites without sitemaps get status noSitemap and the run continues with the other sites. If a site blocks cloud IP ranges (HTTP 403), enable Apify Proxy in Proxy configuration; with a proxy, 403 responses are retried with a new IP.
Why do I see "removed" URLs on news sites? News sitemaps are rolling windows (usually the last 48 hours), so articles drop out as they age. "Removed" means "no longer listed in any sitemap". Monitor the site's full sitemaps if you need actual deletions.
Can I use it from my code or an AI agent? Yes - call it via the Apify API, the JavaScript or Python client, or from AI agents through the Apify MCP server. The output is a structured dataset, so it is easy to feed into scripts, spreadsheets, LLM workflows and RAG pipelines.
Does it respect robots.txt? Is it legal?
By default it skips sitemap files disallowed by robots.txt, honours Crawl-delay, and treats a host whose robots.txt returns server errors as disallowed (RFC 9309). Sitemaps are published so that machines can read them; still, make sure your use of the data complies with the site's terms and applicable law.
How do I reset monitoring for a site?
Delete the site's record from the sitemap-extractor-monitor-state key-value store (the key is shown in the summary as monitor.stateKey), or use a new State store name.
Support
Found a site that doesn't work? Open an issue on the Actor's Issues tab with the site URL and the run link - the per-sitemap report in SUMMARY usually shows exactly what went wrong.