Sitemap URL Extractor — SEO & Change Monitor
Pricing
from $0.50 / 1,000 sitemap urls
Sitemap URL Extractor — SEO & Change Monitor
Extract URLs from XML and gzipped sitemaps, nested indexes and robots.txt discovery. Export lastmod, priority, hreflang, image/video data, filter URLs, audit status/redirects and page indexability, fall back to crawling, and compare persistent snapshots for URL changes.
Pricing
from $0.50 / 1,000 sitemap urls
Rating
0.0
(0)
Developer
Rosario Vitale
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Sitemap URL Extractor API — Indexability & Change Monitor
Why use this Actor?
Extract URLs from XML and gzipped sitemaps, nested indexes and robots.txt discovery. Export lastmod, priority, hreflang, image/video data, filter URLs, audit status/redirects and page indexability, fall back to crawling, and compare persistent snapshots for URL changes.
Features
- Websites or sitemap URLs — Public site homepages or direct .xml/.xml.gz sitemap URLs.
- Maximum sitemap files — Maximum sitemap files. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- Maximum unique URLs — Maximum unique URLs. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- Include lastmod/changefreq/priority — Include lastmod/changefreq/priority. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- Extract image and video sitemap data — Extract image and video sitemap data. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- Extract hreflang alternates — Extract hreflang alternates. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- Include URL regex patterns — When non-empty, only URLs matching at least one regular expression are emitted.
- Exclude URL regex patterns — URLs matching any expression are skipped.
- Audit live HTTP status and redirects — Audit live HTTP status and redirects. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- Maximum URLs to status-check — Maximum URLs to status-check. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- Status-check concurrency — Status-check concurrency. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- Fallback crawl when no sitemap works — Crawl same-origin HTML links only when a site yields no sitemap URLs.
Use cases
- Seo inventory.
- Migration qa.
- Indexation monitoring.
- Url change detection.
Example input
{"sites": ["https://example.com"],"maxSitemaps": 200,"maxUrls": 25000,"includeMetadata": true,"includeMedia": true,"includeHreflang": true}
Pricing & cost control
Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.
FAQ
What is this Actor for?
It is designed for SEO inventory, migration QA, indexation monitoring.
Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.
How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.
Search keywords
sitemap url extractor, sitemap scraper, sitemap scraper python, sitemap web scraper, sitemap url scraper, website sitemap scraper, xml sitemap scraper, sitemap json web scraper, web scraper sitemap wizard, what is a sitemap and why is it important, how do i find the sitemap of a website, sitemap url extractor extension, sitemap url extractor tool, sitemap url extractor chrome extension
Turn a sitemap into a repeatable SEO inventory, health check, and change-monitoring feed.
This Actor accepts website homepages, direct XML sitemaps, sitemap indexes, and compressed .xml.gz files. It discovers sitemap locations from robots.txt plus common WordPress/generic paths, recursively follows indexes, deduplicates URLs, extracts sitemap metadata, images, videos and hreflang alternates, and can optionally audit live HTTP status/redirect chains.
The important difference from a basic sitemap extractor is change detection. A named persistent snapshot can be compared on every scheduled run so the Dataset reports URLs that were added, removed, redirected, changed status, or changed lastmod.
Why use it
- XML sitemap and sitemap-index recursion.
- Real
.xml.gzdecompression. robots.txt,/sitemap.xml,/sitemap_index.xml, and/wp-sitemap.xmldiscovery.- Image/video sitemap fields.
- hreflang alternate links.
- Include/exclude regular-expression filters.
- Optional live status, content type, redirect chain and final URL audit.
- Optional same-origin fallback crawl when no sitemap works.
- Persistent scheduled-run snapshots and URL diffs.
- Explicit limits, retries, timeouts, size caps, and concurrency controls.
- Free diagnostic, summary and change rows; billable units remain unique sitemap URL rows.
Input example
{"sites": ["https://example.com"],"maxUrls": 25000,"includeMetadata": true,"includeMedia": true,"includeHreflang": true,"excludeUrlPatterns": ["[?&]preview=", "/cart/"],"checkStatusCodes": true,"statusCheckLimit": 2000,"snapshotMode": "compare_update","snapshotId": "example-production"}
URL rows
A normal URL row can contain:
{"recordType": "url","url": "https://example.com/products/a","sitemapUrl": "https://example.com/sitemap-products.xml.gz","rootSite": "https://example.com/","lastmod": "2026-09-26","images": [{"loc": "https://cdn.example.com/a.jpg"}],"hreflang": [{"hreflang": "it", "href": "https://example.com/it/products/a"}],"status": 200,"finalUrl": "https://example.com/products/a","redirected": false,"responseTimeMs": 143}
Change rows
With snapshotMode=compare_update, the first run establishes a baseline. Later runs can emit:
addedremovedlastmod_changedstatus_changedredirect_changed
A summary row reports counts for every change type, sitemap errors, gzip files, status errors, redirects, media references and fallback pages.
Fallback crawl
Enable fallbackCrawl only when you want coverage for sites that expose no usable sitemap. The fallback remains same-origin and obeys a strict page limit. It is deliberately off by default because crawling HTML costs more than reading XML.
Status auditing
checkStatusCodes performs bounded concurrent requests and records:
- HTTP status
- final URL
- redirect chain
- content type
- response time
- a clean per-URL error when a request fails
Use statusCheckLimit to bound network work on very large sitemaps.
Snapshot storage
Snapshots live in a named Key-Value Store (default sitemap-intelligence-snapshots) so they persist across runs. Set a stable snapshotId when you schedule the same property repeatedly. maxSnapshotUrls bounds persistent snapshot size and the summary tells you when a snapshot was truncated.
Reliability and limits
Malformed sitemap files are isolated into diagnostic rows rather than aborting an entire run. Inputs are capped, responses have size limits, gzip is decompressed explicitly, retries are bounded, duplicate URLs are removed, and regular expressions are validated before network work begins.
Pricing
The primary value unit remains one unique sitemap URL emitted. Diagnostics, summaries, and change rows are not intentionally billed as sitemap URL events. Live status checks and fallback crawling consume additional platform resources, so use their limits according to the size of the site.
Responsible use
Use this Actor on public URLs that you are allowed to access. Respect site policies, rate limits, and applicable law. Status auditing and fallback crawling should be configured conservatively on sites you do not control.
Extended capabilities
- Discover robots.txt sitemaps, common paths, nested indexes, and .xml.gz files with image/video/hreflang metadata.
- Audit URL status, redirects, content type, and response time, with optional same-origin fallback crawling.
- Persist snapshots and optionally audit title, canonical, meta robots, X-Robots-Tag, and factual indexability signals.