Sitemap URL Extractor & Change Monitor avatar

Sitemap URL Extractor & Change Monitor

Pricing

from $0.12 / 1,000 results

Go to Apify Store
Sitemap URL Extractor & Change Monitor

Sitemap URL Extractor & Change Monitor

Sitemap URL extractor that reads robots.txt, sitemap indexes, .xml.gz and plain-text sitemaps and returns one row per URL with lastmod, changefreq, priority — plus a new/removed diff between runs.

Pricing

from $0.12 / 1,000 results

Rating

0.0

(0)

Developer

Murat Uzun

Murat Uzun

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What is Sitemap URL Extractor & Change Monitor?

Sitemap URL Extractor & Change Monitor is an Apify Actor that turns any website's XML sitemaps into a flat dataset — one row per URL, with lastmod, changefreq, priority and image/video counts. Give it a bare domain and it reads the Sitemap: lines of /robots.txt, then falls back to /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /sitemap.xml.gz and /sitemap.txt. Give it a sitemap index and it follows every child sitemap (up to 500) by itself. Gzipped .xml.gz sitemaps, namespace-prefixed XML, CDATA-wrapped <loc> values, plain-text sitemaps and RSS/Atom feeds used as sitemaps are all handled. Switch Mode to diff and the Actor snapshots each source in a named key-value store, so every later run labels each URL new, unchanged or removed — a ready-made "what pages did this site add this week" feed.

What data does Sitemap URL Extractor & Change Monitor extract?

Sitemap URL Extractor & Change Monitor extracts 11 fields per URL:

FieldTypeDescription
urlstringThe page URL from <loc>, entity-decoded and resolved to absolute
lastmodstring<lastmod> normalised to an ISO 8601 UTC timestamp (2026-09-12T16:30:03.000Z)
changefreqstring<changefreq> hint: always, hourly, daily, weekly, monthly, yearly, never
prioritynumber<priority> clamped to 0.0-1.0
imageCountinteger<image:image> entries in the URL block (Google image sitemap extension)
videoCountinteger<video:video> entries in the URL block (Google video sitemap extension)
sitemapUrlstringWhich child sitemap file actually contained this URL, e.g. https://apify.com/sitemap/pages.xml
sourcestringThe domain or sitemap URL you entered
changestringDiff mode only: new, unchanged or removed versus the previous run
errorstringWhy a source produced nothing, e.g. No sitemap found; null on successful rows
scrapedAtstringWhen the run read the sitemaps (ISO 8601 UTC)

How to use Sitemap URL Extractor & Change Monitor

  1. Paste domains or sitemap URLs into Sitemaps or domains. apify.com, https://apify.com and https://www.allbirds.com/sitemap.xml all work — discovery, index recursion and gzip are automatic.
  2. Set Max URLs per source to the number of rows you want per site (default 1,000, maximum 200,000). The cap counts across every child sitemap of an index.
  3. Optionally narrow the crawl with URL filter (regex) — /products/ for a Shopify catalogue, /blog/ for content, \.pdf$ for documents.
  4. Click Start, then export as JSON, CSV, Excel or HTML — or schedule the Actor in diff mode and wire the Changes view into a webhook.

Example input

{
"sources": ["https://apify.com", "https://www.allbirds.com/sitemap.xml"],
"maxUrlsPerSource": 1000,
"urlFilter": "",
"mode": "list",
"maxConcurrency": 5
}

Example output

{
"source": "https://www.allbirds.com/sitemap.xml",
"sitemapUrl": "https://www.allbirds.com/sitemap_products_1.xml?from=1878194389061&to=7369944137808",
"url": "https://www.allbirds.com/products/mens-wool-runners-natural-white",
"lastmod": "2026-09-12T16:30:03.000Z",
"changefreq": "daily",
"priority": null,
"imageCount": 1,
"videoCount": 0,
"change": null,
"error": null,
"scrapedAt": "2026-09-12T17:05:00.000Z"
}

Input parameters

ParameterTypeDefaultDescription
sourcesarray["https://apify.com"]Sitemap URLs or bare domains; robots.txt discovery for domains
maxUrlsPerSourceinteger1000URL cap per source across all child sitemaps (1-200,000)
urlFilterstring""Optional regex; only matching URLs are kept
modestringlistlist = every URL, diff = new/unchanged/removed since last run
diffStoreNamestringsitemap-monitorNamed key-value store holding the between-run URL snapshots
maxConcurrencyinteger5Sitemap files downloaded in parallel (1-20)

Pricing

Sitemap URL Extractor & Change Monitor uses pay-per-event pricing: $0.0002 per URL row, i.e. $0.20 per 1,000 URLs, plus a negligible actor-start fee, platform usage included. A sitemap is one small HTTP request per file — the 1,591-URL apify.com page sitemap is a single 132 KB download — so compute stays in the cents even for 200,000-URL catalogues. Set Maximum cost per run and the Actor trims the result set to what the budget covers instead of overspending.

Sitemap URL Extractor & Change Monitor vs. a full site crawl

Sitemap URL Extractor & Change Monitor gets you a site's URL inventory in seconds instead of hours. A crawler has to fetch and parse every page to find the next link, burning proxy bandwidth and tripping rate limits; a sitemap is the site's own published index, already complete, already dated. Use this Actor to build the URL list, then feed it to a page-level scraper — you crawl only what you meant to. Versus Screaming Frog or a curl | grep one-liner, you also get gzip and index recursion handled, lastmod normalised to ISO, and a persistent diff between runs that no one-shot tool gives you.

Using Sitemap URL Extractor & Change Monitor with AI agents and MCP

Sitemap URL Extractor & Change Monitor is pay-per-event with limited permissions — the two requirements for an Actor to be callable through the Apify MCP server at mcp.apify.com. An agent passes sources and gets back a structured URL inventory it can use to decide what to read next, without writing a crawler. The same run works from n8n, Make, Zapier and LangChain through Apify's integrations; in diff mode a scheduled run plus a webhook gives you "alert me when this competitor publishes a new page".

FAQ

How does diff mode remember the previous run? Each source's URL set is stored in the named key-value store from Diff snapshot store name (default sitemap-monitor) under a key derived from a SHA-1 of the source string. Keep the store name identical across scheduled runs so they compare against each other. The first run has no snapshot, so every URL is reported as new; the second run of an unchanged site reports every URL as unchanged. A source that fails to fetch never overwrites its snapshot, so one transient error cannot report a whole site as removed.

Are .xml.gz sitemaps supported? Yes. Some servers send Content-Encoding: gzip and some send raw gzip bytes as binary/octet-stream — the Actor sniffs the gzip magic number instead of trusting the file extension, so both work.

What are the limitations? Only what the site publishes: pages missing from the sitemap are invisible here, and lastmod is the site's claim, not a verified fact. An index is followed up to 500 child sitemaps and 4 levels deep. Sitemaps behind an aggressive CDN bot wall can return HTML or 403; that source then produces one row with the reason in error instead of failing the run.

Is this legal to run? Yes. robots.txt and sitemaps exist specifically to be read by automated clients, and no personal data is collected.

Can I export to CSV or Excel? Yes, from the Output tab or the API, with ready-made Overview and Changes (diff mode) views.

Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:

Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.

Website & domain intelligence

Content for AI, LLMs and RAG

Search, video and social

Leads, jobs and company data

Developer, app and research data

Support and feedback

Hit a sitemap flavour it parses wrong, or want another extension counted? Open an issue on the Issues tab.