Sitemap URL Extractor & New Page Monitor
Pricing
Pay per event
Sitemap URL Extractor & New Page Monitor
Extract every URL from a website's XML sitemaps (auto-discovered from robots.txt, sitemap indexes, .gz, hreflang, news and image sitemaps) or monitor sitemaps for new, modified and removed pages. Fast HTTP-only, pay per URL.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Mohamed T.
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 hours ago
Last modified
Categories
Share
Get every URL of a website in seconds, or get alerted when it publishes, updates or removes pages. Give it a domain: it finds the sitemaps (robots.txt, sitemap indexes, .xml.gz, text and RSS sitemaps), and returns each URL with lastmod, changefreq, priority, hreflang alternates, image and news data.
In monitor mode it remembers the previous run and outputs only what changed: new product pages, fresh blog posts, updated landing pages, deleted URLs. Schedule it daily and plug it into Slack, email or a Google Sheet.
What does Sitemap URL Extractor do?
- Auto-discovers sitemaps from
robots.txt, then common locations (/sitemap.xml,/sitemap_index.xml,/wp-sitemap.xml, …), or uses the sitemap URL you give it. - Follows sitemap indexes recursively and decompresses
.gzfiles. It also handles huge sitemaps and tolerates slightly broken XML. - Extracts everything the protocol offers:
loc,lastmod(normalized to ISO-8601),changefreq,priority,xhtml:linkhreflang alternates, image URLs, Google News title and publication date. - Filters by regex (include/exclude) and by modification date ("modified in the last 7 days").
- Monitors changes between runs:
added,modified(lastmod changed) andremovedURLs.
It's HTTP-only and polite (robots.txt respected, per-host delays, retries with backoff), so it's fast and cheap.
Why use it?
- SEO audits and migrations: list all indexable URLs, compare staging vs production, find pages missing hreflang.
- Competitive intelligence: know the day a competitor launches a product, a landing page or a pricing page.
- Content monitoring: track new articles from news sites, blogs or documentation.
- Feed your crawlers and AI/RAG pipelines: get a clean URL list before scraping, instead of crawling blindly.
- E-commerce: detect new and discontinued product pages across competitor catalogs.
How to use it
- Add one or more websites or sitemap URLs to Websites or sitemap URLs.
- Keep Mode = Extract to list URLs, or pick Monitor changes and give the watchlist a Monitor name.
- Optionally filter with Only URLs matching (e.g.
/blog/) or Modified after (7 days). - Click Start, then download the results as JSON, CSV or Excel, or use the API.
- For monitoring, create a Schedule (e.g. every morning) and add an integration (Slack, email, webhook, Google Sheets).
Input example
{"startUrls": ["https://apify.com", "https://www.bbc.co.uk/food/sitemap.xml"],"mode": "extract","includeUrlPatterns": ["/recipes/"],"lastmodAfter": "30 days","maxUrlsPerSite": 10000}
Output example
Extract mode, one row per URL:
{"url": "https://www.bbc.co.uk/food/recipes/deep-filled_lemon_50034","site": "https://www.bbc.co.uk","sitemapUrl": "https://www.bbc.co.uk/food/sitemap.xml","lastmod": "2014-01-29","changefreq": null,"priority": null,"alternates": []}
Monitor mode, one row per change:
{"url": "https://www.example.com/products/new-sneaker","site": "https://www.example.com","lastmod": "2026-09-27T08:12:00Z","changeType": "added","detectedAt": "2026-09-27T09:00:03+00:00"}
A per-site summary (sitemaps found, URL count, discovery method, errors) is saved in the key-value store record SUMMARY.
Data fields
| Field | Description |
|---|---|
url | Page URL from the sitemap |
site | Origin of the website |
sitemapUrl | Sitemap file where the URL was found |
lastmod, changefreq, priority | Values declared in the sitemap (lastmod normalized to ISO-8601) |
alternates | hreflang alternates [{hreflang, url}] |
images | Image URLs (optional) |
newsTitle, newsPublicationDate | Google News sitemap data |
changeType, previousLastmod, detectedAt | Monitor mode only: added, modified, removed (or baseline) |
How much does it cost?
Pay per event. Extract mode: a small price per URL returned. Monitor mode: a small fee per site checked plus a price per detected change. Unchanged URLs cost nothing. See the Pricing tab. Use Max URLs per site to cap test runs.
Tips
- Monitoring many sites? Use one Monitor name per watchlist; each site keeps its own state.
- The first monitor run only stores a baseline (enable Output all URLs on the first run if you want them).
- When a limit truncates the crawl, removed pages aren't reported, to avoid false alarms.
- Big news sites publish separate news sitemaps. Pass them directly for the freshest articles.
FAQ
The site has no sitemap. What happens?
The site is reported in SUMMARY with the reason, and you aren't charged for it. Sitemap-less sites need a crawler instead.
Why are some lastmod values missing or old?
The Actor returns what the website declares. Many sites omit lastmod or never update it.
Is it legal? Sitemaps are published by websites precisely so machines can read them. The Actor respects robots.txt and collects no personal data.
Found a sitemap it can't read? Open an issue on the Issues tab with the URL. Fixes usually ship within days.