Sitemap URL Extractor & New Page Monitor avatar

Sitemap URL Extractor & New Page Monitor

Pricing

Pay per event

Go to Apify Store
Sitemap URL Extractor & New Page Monitor

Sitemap URL Extractor & New Page Monitor

Extract every URL from a website's XML sitemaps (auto-discovered from robots.txt, sitemap indexes, .gz, hreflang, news and image sitemaps) or monitor sitemaps for new, modified and removed pages. Fast HTTP-only, pay per URL.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Mohamed T.

Mohamed T.

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 hours ago

Last modified

Share

Get every URL of a website in seconds, or get alerted when it publishes, updates or removes pages. Give it a domain: it finds the sitemaps (robots.txt, sitemap indexes, .xml.gz, text and RSS sitemaps), and returns each URL with lastmod, changefreq, priority, hreflang alternates, image and news data.

In monitor mode it remembers the previous run and outputs only what changed: new product pages, fresh blog posts, updated landing pages, deleted URLs. Schedule it daily and plug it into Slack, email or a Google Sheet.

What does Sitemap URL Extractor do?

  • Auto-discovers sitemaps from robots.txt, then common locations (/sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml, …), or uses the sitemap URL you give it.
  • Follows sitemap indexes recursively and decompresses .gz files. It also handles huge sitemaps and tolerates slightly broken XML.
  • Extracts everything the protocol offers: loc, lastmod (normalized to ISO-8601), changefreq, priority, xhtml:link hreflang alternates, image URLs, Google News title and publication date.
  • Filters by regex (include/exclude) and by modification date ("modified in the last 7 days").
  • Monitors changes between runs: added, modified (lastmod changed) and removed URLs.

It's HTTP-only and polite (robots.txt respected, per-host delays, retries with backoff), so it's fast and cheap.

Why use it?

  • SEO audits and migrations: list all indexable URLs, compare staging vs production, find pages missing hreflang.
  • Competitive intelligence: know the day a competitor launches a product, a landing page or a pricing page.
  • Content monitoring: track new articles from news sites, blogs or documentation.
  • Feed your crawlers and AI/RAG pipelines: get a clean URL list before scraping, instead of crawling blindly.
  • E-commerce: detect new and discontinued product pages across competitor catalogs.

How to use it

  1. Add one or more websites or sitemap URLs to Websites or sitemap URLs.
  2. Keep Mode = Extract to list URLs, or pick Monitor changes and give the watchlist a Monitor name.
  3. Optionally filter with Only URLs matching (e.g. /blog/) or Modified after (7 days).
  4. Click Start, then download the results as JSON, CSV or Excel, or use the API.
  5. For monitoring, create a Schedule (e.g. every morning) and add an integration (Slack, email, webhook, Google Sheets).

Input example

{
"startUrls": ["https://apify.com", "https://www.bbc.co.uk/food/sitemap.xml"],
"mode": "extract",
"includeUrlPatterns": ["/recipes/"],
"lastmodAfter": "30 days",
"maxUrlsPerSite": 10000
}

Output example

Extract mode, one row per URL:

{
"url": "https://www.bbc.co.uk/food/recipes/deep-filled_lemon_50034",
"site": "https://www.bbc.co.uk",
"sitemapUrl": "https://www.bbc.co.uk/food/sitemap.xml",
"lastmod": "2014-01-29",
"changefreq": null,
"priority": null,
"alternates": []
}

Monitor mode, one row per change:

{
"url": "https://www.example.com/products/new-sneaker",
"site": "https://www.example.com",
"lastmod": "2026-09-27T08:12:00Z",
"changeType": "added",
"detectedAt": "2026-09-27T09:00:03+00:00"
}

A per-site summary (sitemaps found, URL count, discovery method, errors) is saved in the key-value store record SUMMARY.

Data fields

FieldDescription
urlPage URL from the sitemap
siteOrigin of the website
sitemapUrlSitemap file where the URL was found
lastmod, changefreq, priorityValues declared in the sitemap (lastmod normalized to ISO-8601)
alternateshreflang alternates [{hreflang, url}]
imagesImage URLs (optional)
newsTitle, newsPublicationDateGoogle News sitemap data
changeType, previousLastmod, detectedAtMonitor mode only: added, modified, removed (or baseline)

How much does it cost?

Pay per event. Extract mode: a small price per URL returned. Monitor mode: a small fee per site checked plus a price per detected change. Unchanged URLs cost nothing. See the Pricing tab. Use Max URLs per site to cap test runs.

Tips

  • Monitoring many sites? Use one Monitor name per watchlist; each site keeps its own state.
  • The first monitor run only stores a baseline (enable Output all URLs on the first run if you want them).
  • When a limit truncates the crawl, removed pages aren't reported, to avoid false alarms.
  • Big news sites publish separate news sitemaps. Pass them directly for the freshest articles.

FAQ

The site has no sitemap. What happens? The site is reported in SUMMARY with the reason, and you aren't charged for it. Sitemap-less sites need a crawler instead.

Why are some lastmod values missing or old? The Actor returns what the website declares. Many sites omit lastmod or never update it.

Is it legal? Sitemaps are published by websites precisely so machines can read them. The Actor respects robots.txt and collects no personal data.

Found a sitemap it can't read? Open an issue on the Issues tab with the URL. Fixes usually ship within days.