Sitemap URL Extractor + Change Monitor (robots.txt, gzip)
Pricing
from $0.20 / 1,000 url rows
Sitemap URL Extractor + Change Monitor (robots.txt, gzip)
Every URL from a site's sitemaps: robots.txt discovery, sitemap index recursion, gzip, lastmod filter, image/video/news/hreflang fields. Or run it on a schedule and get only the URLs added, removed or changed since last time.
Pricing
from $0.20 / 1,000 url rows
Rating
0.0
(0)
Developer
Tenzin Phuntsok
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Get every URL a website lists in its sitemaps, or only the URLs that were added, removed or changed since your last run. Give it a domain and it finds the sitemaps through robots.txt and the usual locations, follows sitemap indexes, unpacks gzip, and returns a clean URL list with last-modified dates, images, videos, news and hreflang data. Filter by date or pattern, then feed the list to a scraper, a Google Sheet, or an alert.
What does Sitemap URL Extractor + Change Monitor do?
A sitemap is the file a website publishes so search engines can find all of its pages. This Actor reads those files and gives you the list.
- Discovery: paste a domain (
example.com), arobots.txtURL, or a sitemap URL. For a domain it readsrobots.txtforSitemap:lines, then tries/sitemap.xml,/sitemap_index.xmland six other common locations. - Every format: XML sitemaps, sitemap indexes (recursively), gzipped
.xml.gz, plain-text sitemaps, RSS and Atom feeds. Gzip is detected from the bytes, not the file name. - Every field:
lastmod(raw and normalised ISO),changefreq,priority, the sitemap the URL came from, and the image, video, Google News and hreflang extensions. - Filters: only URLs modified after a date or in the last
7d, include/exclude regular expressions, duplicate removal across sitemaps. - Change monitoring: run it on a schedule with mode Only changes and get just the URLs that were added, removed or had their
lastmodchanged since the previous run. The first run stores a baseline. - robots.txt summary: which AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot and more) the site blocks, and whether it publishes an
llms.txt. - Optional status checks: request each URL (cut off after the headers) and record the HTTP status and final URL. Off by default, so a 150,000-URL site takes seconds, not hours.
Why use it?
Most sitemap tools stop at url and lastmod, fail on gzip or malformed XML, or check every page's status by default and time out on big sites. This one is built to finish:
- A broken child sitemap, an HTML error page where a sitemap should be, or a
500from one file does not fail the run. The good sitemaps are read, the bad ones are listed in the run summary with the reason, and the summary says whether the run was complete. - Retries with back-off on
429/5xxand network errors; realistic browser headers so fewer sites answer403; optional proxy for the rest. - Sitemaps are streamed, so a 50 MB file does not need 50 MB of memory. Caps on URL count, sitemap count, nesting depth and file size keep a runaway index from running forever.
- Change detection is careful: removals are never reported from an incomplete run, a sitemap that briefly returns half its URLs does not produce a flood of false "removed" rows, and a run that stops at your cost limit remembers what it did not report so it shows up next time.
- Restart-safe: if the platform migrates the run, it continues without writing duplicate rows or charging twice.
Use cases: seed a scraper with fresh URLs only (lastmodAfter: 7d), watch a competitor for new product or blog pages, audit a site's sitemap health, build a URL inventory for SEO, check which AI bots a site allows.
How to extract sitemap URLs
- In Websites or sitemaps, add a domain such as
example.com, or a sitemap URL if you already know it. One entry per site. - Leave What to output on All URLs for a full list. Optionally set Only URLs modified after (
2026-09-01or30d) and a pattern such as/blog/. - Click Start. The Output tab lists the URLs; the Storage > Key-value store > SUMMARY record has the full report (each sitemap fetched, robots.txt findings, counts, warnings).
- To monitor changes, switch What to output to Only changes since the last run and create a Schedule (daily works well). The first run stores the baseline and outputs nothing; every later run outputs only the differences.
Input
| Field | What it does |
|---|---|
startUrls | Domains, robots.txt URLs or sitemap URLs. |
mode | all (default) = every URL. changes = only URLs added, removed or changed since the previous run with the same snapshot. |
maxUrls | Stop after this many URLs (default 500,000). |
lastmodAfter | Keep URLs whose lastmod is after a date (2026-09-01), a datetime, or a span back from now (7d, 36h, 2 weeks). |
includeUrlsWithoutLastmod | With lastmodAfter: keep URLs that have no lastmod (default off). |
urlIncludePatterns, urlExcludePatterns | Regular expressions the URL must match / must not match. |
dedupeUrls | Output each URL once even if several sitemaps list it (default on). |
includeExtensions | Extract image, video, news and hreflang data (default on). |
snapshotKey, firstRunOutput, allowMassRemoval, resetSnapshot | Change-detection settings, see below. |
checkStatus, maxStatusChecks | Request each URL (cut off after the headers) and record status and finalUrl. Off by default; capped at 5,000 per run by default. |
aiBotSummary | Add the AI-crawler policy table and the llms.txt check to the summary (default on). |
maxSitemaps, maxDepth, maxSitemapMegabytes | Caps: sitemap files per run (2,000), index nesting (5), size per file (100 MB uncompressed). |
maxConcurrency, requestTimeoutSecs | Parallel sitemap requests (5) and per-request timeout (30 s). |
customUserAgent, proxyConfiguration | Identify your crawler, or route through Apify Proxy when a site blocks datacenter IPs. |
Output
One dataset row per URL. You can download it as JSON, CSV, Excel or HTML, or pipe it into another Actor (for example a scraper's startUrls, or Dataset to Google Sheets Sync).
{"url": "https://example.com/blog/how-to-read-a-sitemap","lastmod": "2026-09-20T08:15:00+00:00","lastmodIso": "2026-09-20T08:15:00.000Z","changefreq": "weekly","priority": 0.8,"sitemap": "https://example.com/sitemap-posts.xml","sitemapIndex": "https://example.com/sitemap_index.xml","images": [{ "loc": "https://example.com/img/sitemap.png", "title": "Sitemap diagram", "caption": null }],"videos": [],"news": null,"alternates": [{ "hreflang": "de", "href": "https://example.com/de/blog/sitemap-lesen" }]}
In changes mode each row also has change (added, changed or removed) and previousLastmod. With status checks on, status and finalUrl are added.
| Field | Meaning |
|---|---|
url | The page URL exactly as listed (fragment removed). |
lastmod / lastmodIso | Last-modified value as written in the sitemap, and normalised to ISO 8601 (null if unparseable). |
changefreq, priority | Sitemap hints, when present. |
sitemap, sitemapIndex | Which sitemap file listed the URL and, if it came via an index, which index. |
images, videos, news | Image, video and Google News extension data. |
alternates | hreflang alternates (xhtml:link rel="alternate"). |
status, finalUrl | HTTP status and redirect target from the optional status check. |
change, previousLastmod | Change type and the previous lastmod, in changes mode. |
The SUMMARY record in the run's key-value store contains: complete and incompleteReasons, counts (found, accepted, written, duplicates, dropped by pattern/date), the list of every sitemap fetched with status, kind, URL count, bytes, time and error, the discovery path per site, the robots.txt findings with the AI-bot table, and all warnings.
Change monitoring in detail
- The URL set of each run is saved as a snapshot in a key-value store named
sitemap-snapshotsin your account. The snapshot name is derived from the start URLs and patterns, so the same input always compares against its own previous run. SetsnapshotKeyto control this yourself. - A
changedrow means thelastmodvalue differs from the previous run (including from empty to a date). Two spellings of the same moment (2026-01-05and2026-01-05T00:00:00Z) count as unchanged. - Removals are suppressed when the run was incomplete (a sitemap failed with a server error or timeout, a cap or the cost limit was hit, the run was aborted or stopped before its timeout), when more than half of a 100+ URL baseline disappears in one run, and when
lastmodAfteris a moving window like7d(URLs leave the window without being removed from the site). In all cases the missing URLs stay in the snapshot and the summary says why (changes.removedSuppressedReason). A sitemap that answers 404 is treated as really gone. Turn onallowMassRemovalto report a genuine site restructure. - The first run stores the baseline and outputs nothing (
firstRunOutput: all-as-addedoutputs every URL asaddedinstead). A run that could not read any sitemap stores no baseline.resetSnapshotstarts over. - Run one schedule per site at a time; two runs comparing against the same snapshot at the same moment can each report the same change.
Pricing
Pay per event: the platform's standard Actor start fee, a fraction of a cent per URL row written, a separate small fee per URL status check (only when you turn checks on), and a per-run fee for change detection (charged only when a comparison happens, not on the baseline run). Rows that are not written because your maximum cost per run was reached are never charged; the summary says the run stopped early and, in changes mode, those changes are reported on the next run instead. When the budget is tight, URL rows come first and status checks are cut, never the other way round. Exact prices are on this page's pricing box.
A full extraction of a 150,000-URL site takes a few seconds and one request per sitemap file; a daily change check on the same site costs the start fee, the change-detection fee, and only the rows that changed.
Tips
- Fresh URLs for a scraper:
lastmodAfter: "7d"plus a schedule. URLs without alastmodare dropped by default (turn onincludeUrlsWithoutLastmodif the site omits dates). - One section of a site:
urlIncludePatterns: ["/blog/"]. Patterns are regular expressions; escape dots and question marks (\\?page=). - Site answers 403 or a Cloudflare page: set Proxy to Apify Proxy (residential if datacenter is blocked). Or point
startUrlsstraight at the sitemap URL, which is often less protected than the homepage. - Huge sites: raise
maxUrlsand give the run 2 to 4 GB of memory for several million URLs.maxConcurrency10 speeds up indexes with hundreds of child files. - Status checks: keep
maxStatusChecksmodest; each check is one request to the site. Combine withurlIncludePatternsto check only what matters. - Sitemap not found: the summary's
discoverysection lists every location tried and what came back. Many sites only publish the sitemap location inside robots.txt, or use a non-standard path; paste it intostartUrls.
FAQ and disclaimer
Does it read page content? No. It reads only robots.txt, the first bytes of llms.txt (existence check), and sitemap files, which sites publish for automated consumption. The output is URL lists and sitemap metadata. Status checks send one request per URL, cut off as soon as the headers arrive, and store only the status code and final URL.
Is this legal? Sitemaps and robots.txt are public protocol files intended for machines. You choose the sites you target and are responsible for using the results in line with those sites' terms and applicable law. Reading a sitemap is not the same as scraping the pages it lists.
Why is complete false? One or more sitemaps could not be read, a cap (maxUrls, maxSitemaps, maxDepth, size, cost) stopped the run, or the run stopped fetching shortly before its timeout so that the rows it had could still be written (raise the run timeout to read everything). incompleteReasons and the per-sitemap list say exactly what happened. A start page that is not a sitemap, or a stale robots.txt entry replaced by a sitemap found at a well-known path, does not count as a failure.
Something else? Report it in the Issues tab of this Actor. Custom variants (structured-data validation, sitemap generation, other formats) can be built on request.