Sitemap URL Extractor + Change Monitor (robots.txt, gzip) avatar

Sitemap URL Extractor + Change Monitor (robots.txt, gzip)

Pricing

from $0.20 / 1,000 url rows

Go to Apify Store
Sitemap URL Extractor + Change Monitor (robots.txt, gzip)

Sitemap URL Extractor + Change Monitor (robots.txt, gzip)

Every URL from a site's sitemaps: robots.txt discovery, sitemap index recursion, gzip, lastmod filter, image/video/news/hreflang fields. Or run it on a schedule and get only the URLs added, removed or changed since last time.

Pricing

from $0.20 / 1,000 url rows

Rating

0.0

(0)

Developer

Tenzin Phuntsok

Tenzin Phuntsok

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Get every URL a website lists in its sitemaps, or only the URLs that were added, removed or changed since your last run. Give it a domain and it finds the sitemaps through robots.txt and the usual locations, follows sitemap indexes, unpacks gzip, and returns a clean URL list with last-modified dates, images, videos, news and hreflang data. Filter by date or pattern, then feed the list to a scraper, a Google Sheet, or an alert.

What does Sitemap URL Extractor + Change Monitor do?

A sitemap is the file a website publishes so search engines can find all of its pages. This Actor reads those files and gives you the list.

  • Discovery: paste a domain (example.com), a robots.txt URL, or a sitemap URL. For a domain it reads robots.txt for Sitemap: lines, then tries /sitemap.xml, /sitemap_index.xml and six other common locations.
  • Every format: XML sitemaps, sitemap indexes (recursively), gzipped .xml.gz, plain-text sitemaps, RSS and Atom feeds. Gzip is detected from the bytes, not the file name.
  • Every field: lastmod (raw and normalised ISO), changefreq, priority, the sitemap the URL came from, and the image, video, Google News and hreflang extensions.
  • Filters: only URLs modified after a date or in the last 7d, include/exclude regular expressions, duplicate removal across sitemaps.
  • Change monitoring: run it on a schedule with mode Only changes and get just the URLs that were added, removed or had their lastmod changed since the previous run. The first run stores a baseline.
  • robots.txt summary: which AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot and more) the site blocks, and whether it publishes an llms.txt.
  • Optional status checks: request each URL (cut off after the headers) and record the HTTP status and final URL. Off by default, so a 150,000-URL site takes seconds, not hours.

Why use it?

Most sitemap tools stop at url and lastmod, fail on gzip or malformed XML, or check every page's status by default and time out on big sites. This one is built to finish:

  • A broken child sitemap, an HTML error page where a sitemap should be, or a 500 from one file does not fail the run. The good sitemaps are read, the bad ones are listed in the run summary with the reason, and the summary says whether the run was complete.
  • Retries with back-off on 429/5xx and network errors; realistic browser headers so fewer sites answer 403; optional proxy for the rest.
  • Sitemaps are streamed, so a 50 MB file does not need 50 MB of memory. Caps on URL count, sitemap count, nesting depth and file size keep a runaway index from running forever.
  • Change detection is careful: removals are never reported from an incomplete run, a sitemap that briefly returns half its URLs does not produce a flood of false "removed" rows, and a run that stops at your cost limit remembers what it did not report so it shows up next time.
  • Restart-safe: if the platform migrates the run, it continues without writing duplicate rows or charging twice.

Use cases: seed a scraper with fresh URLs only (lastmodAfter: 7d), watch a competitor for new product or blog pages, audit a site's sitemap health, build a URL inventory for SEO, check which AI bots a site allows.

How to extract sitemap URLs

  1. In Websites or sitemaps, add a domain such as example.com, or a sitemap URL if you already know it. One entry per site.
  2. Leave What to output on All URLs for a full list. Optionally set Only URLs modified after (2026-09-01 or 30d) and a pattern such as /blog/.
  3. Click Start. The Output tab lists the URLs; the Storage > Key-value store > SUMMARY record has the full report (each sitemap fetched, robots.txt findings, counts, warnings).
  4. To monitor changes, switch What to output to Only changes since the last run and create a Schedule (daily works well). The first run stores the baseline and outputs nothing; every later run outputs only the differences.

Input

FieldWhat it does
startUrlsDomains, robots.txt URLs or sitemap URLs.
modeall (default) = every URL. changes = only URLs added, removed or changed since the previous run with the same snapshot.
maxUrlsStop after this many URLs (default 500,000).
lastmodAfterKeep URLs whose lastmod is after a date (2026-09-01), a datetime, or a span back from now (7d, 36h, 2 weeks).
includeUrlsWithoutLastmodWith lastmodAfter: keep URLs that have no lastmod (default off).
urlIncludePatterns, urlExcludePatternsRegular expressions the URL must match / must not match.
dedupeUrlsOutput each URL once even if several sitemaps list it (default on).
includeExtensionsExtract image, video, news and hreflang data (default on).
snapshotKey, firstRunOutput, allowMassRemoval, resetSnapshotChange-detection settings, see below.
checkStatus, maxStatusChecksRequest each URL (cut off after the headers) and record status and finalUrl. Off by default; capped at 5,000 per run by default.
aiBotSummaryAdd the AI-crawler policy table and the llms.txt check to the summary (default on).
maxSitemaps, maxDepth, maxSitemapMegabytesCaps: sitemap files per run (2,000), index nesting (5), size per file (100 MB uncompressed).
maxConcurrency, requestTimeoutSecsParallel sitemap requests (5) and per-request timeout (30 s).
customUserAgent, proxyConfigurationIdentify your crawler, or route through Apify Proxy when a site blocks datacenter IPs.

Output

One dataset row per URL. You can download it as JSON, CSV, Excel or HTML, or pipe it into another Actor (for example a scraper's startUrls, or Dataset to Google Sheets Sync).

{
"url": "https://example.com/blog/how-to-read-a-sitemap",
"lastmod": "2026-09-20T08:15:00+00:00",
"lastmodIso": "2026-09-20T08:15:00.000Z",
"changefreq": "weekly",
"priority": 0.8,
"sitemap": "https://example.com/sitemap-posts.xml",
"sitemapIndex": "https://example.com/sitemap_index.xml",
"images": [{ "loc": "https://example.com/img/sitemap.png", "title": "Sitemap diagram", "caption": null }],
"videos": [],
"news": null,
"alternates": [{ "hreflang": "de", "href": "https://example.com/de/blog/sitemap-lesen" }]
}

In changes mode each row also has change (added, changed or removed) and previousLastmod. With status checks on, status and finalUrl are added.

FieldMeaning
urlThe page URL exactly as listed (fragment removed).
lastmod / lastmodIsoLast-modified value as written in the sitemap, and normalised to ISO 8601 (null if unparseable).
changefreq, prioritySitemap hints, when present.
sitemap, sitemapIndexWhich sitemap file listed the URL and, if it came via an index, which index.
images, videos, newsImage, video and Google News extension data.
alternateshreflang alternates (xhtml:link rel="alternate").
status, finalUrlHTTP status and redirect target from the optional status check.
change, previousLastmodChange type and the previous lastmod, in changes mode.

The SUMMARY record in the run's key-value store contains: complete and incompleteReasons, counts (found, accepted, written, duplicates, dropped by pattern/date), the list of every sitemap fetched with status, kind, URL count, bytes, time and error, the discovery path per site, the robots.txt findings with the AI-bot table, and all warnings.

Change monitoring in detail

  • The URL set of each run is saved as a snapshot in a key-value store named sitemap-snapshots in your account. The snapshot name is derived from the start URLs and patterns, so the same input always compares against its own previous run. Set snapshotKey to control this yourself.
  • A changed row means the lastmod value differs from the previous run (including from empty to a date). Two spellings of the same moment (2026-01-05 and 2026-01-05T00:00:00Z) count as unchanged.
  • Removals are suppressed when the run was incomplete (a sitemap failed with a server error or timeout, a cap or the cost limit was hit, the run was aborted or stopped before its timeout), when more than half of a 100+ URL baseline disappears in one run, and when lastmodAfter is a moving window like 7d (URLs leave the window without being removed from the site). In all cases the missing URLs stay in the snapshot and the summary says why (changes.removedSuppressedReason). A sitemap that answers 404 is treated as really gone. Turn on allowMassRemoval to report a genuine site restructure.
  • The first run stores the baseline and outputs nothing (firstRunOutput: all-as-added outputs every URL as added instead). A run that could not read any sitemap stores no baseline. resetSnapshot starts over.
  • Run one schedule per site at a time; two runs comparing against the same snapshot at the same moment can each report the same change.

Pricing

Pay per event: the platform's standard Actor start fee, a fraction of a cent per URL row written, a separate small fee per URL status check (only when you turn checks on), and a per-run fee for change detection (charged only when a comparison happens, not on the baseline run). Rows that are not written because your maximum cost per run was reached are never charged; the summary says the run stopped early and, in changes mode, those changes are reported on the next run instead. When the budget is tight, URL rows come first and status checks are cut, never the other way round. Exact prices are on this page's pricing box.

A full extraction of a 150,000-URL site takes a few seconds and one request per sitemap file; a daily change check on the same site costs the start fee, the change-detection fee, and only the rows that changed.

Tips

  • Fresh URLs for a scraper: lastmodAfter: "7d" plus a schedule. URLs without a lastmod are dropped by default (turn on includeUrlsWithoutLastmod if the site omits dates).
  • One section of a site: urlIncludePatterns: ["/blog/"]. Patterns are regular expressions; escape dots and question marks (\\?page=).
  • Site answers 403 or a Cloudflare page: set Proxy to Apify Proxy (residential if datacenter is blocked). Or point startUrls straight at the sitemap URL, which is often less protected than the homepage.
  • Huge sites: raise maxUrls and give the run 2 to 4 GB of memory for several million URLs. maxConcurrency 10 speeds up indexes with hundreds of child files.
  • Status checks: keep maxStatusChecks modest; each check is one request to the site. Combine with urlIncludePatterns to check only what matters.
  • Sitemap not found: the summary's discovery section lists every location tried and what came back. Many sites only publish the sitemap location inside robots.txt, or use a non-standard path; paste it into startUrls.

FAQ and disclaimer

Does it read page content? No. It reads only robots.txt, the first bytes of llms.txt (existence check), and sitemap files, which sites publish for automated consumption. The output is URL lists and sitemap metadata. Status checks send one request per URL, cut off as soon as the headers arrive, and store only the status code and final URL.

Is this legal? Sitemaps and robots.txt are public protocol files intended for machines. You choose the sites you target and are responsible for using the results in line with those sites' terms and applicable law. Reading a sitemap is not the same as scraping the pages it lists.

Why is complete false? One or more sitemaps could not be read, a cap (maxUrls, maxSitemaps, maxDepth, size, cost) stopped the run, or the run stopped fetching shortly before its timeout so that the rows it had could still be written (raise the run timeout to read everything). incompleteReasons and the per-sitemap list say exactly what happened. A start page that is not a sitemap, or a stale robots.txt entry replaced by a sitemap found at a well-known path, does not count as a failure.

Something else? Report it in the Issues tab of this Actor. Custom variants (structured-data validation, sitemap generation, other formats) can be built on request.