Sitemap URL Extractor & Change Monitor
Pricing
from $0.12 / 1,000 results
Sitemap URL Extractor & Change Monitor
Sitemap URL extractor that reads robots.txt, sitemap indexes, .xml.gz and plain-text sitemaps and returns one row per URL with lastmod, changefreq, priority — plus a new/removed diff between runs.
Pricing
from $0.12 / 1,000 results
Rating
0.0
(0)
Developer
Murat Uzun
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
What is Sitemap URL Extractor & Change Monitor?
Sitemap URL Extractor & Change Monitor is an Apify Actor that turns any website's XML sitemaps into a flat dataset — one row per URL, with lastmod, changefreq, priority and image/video counts. Give it a bare domain and it reads the Sitemap: lines of /robots.txt, then falls back to /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /sitemap.xml.gz and /sitemap.txt. Give it a sitemap index and it follows every child sitemap (up to 500) by itself. Gzipped .xml.gz sitemaps, namespace-prefixed XML, CDATA-wrapped <loc> values, plain-text sitemaps and RSS/Atom feeds used as sitemaps are all handled. Switch Mode to diff and the Actor snapshots each source in a named key-value store, so every later run labels each URL new, unchanged or removed — a ready-made "what pages did this site add this week" feed.
What data does Sitemap URL Extractor & Change Monitor extract?
Sitemap URL Extractor & Change Monitor extracts 11 fields per URL:
| Field | Type | Description |
|---|---|---|
url | string | The page URL from <loc>, entity-decoded and resolved to absolute |
lastmod | string | <lastmod> normalised to an ISO 8601 UTC timestamp (2026-09-12T16:30:03.000Z) |
changefreq | string | <changefreq> hint: always, hourly, daily, weekly, monthly, yearly, never |
priority | number | <priority> clamped to 0.0-1.0 |
imageCount | integer | <image:image> entries in the URL block (Google image sitemap extension) |
videoCount | integer | <video:video> entries in the URL block (Google video sitemap extension) |
sitemapUrl | string | Which child sitemap file actually contained this URL, e.g. https://apify.com/sitemap/pages.xml |
source | string | The domain or sitemap URL you entered |
change | string | Diff mode only: new, unchanged or removed versus the previous run |
error | string | Why a source produced nothing, e.g. No sitemap found; null on successful rows |
scrapedAt | string | When the run read the sitemaps (ISO 8601 UTC) |
How to use Sitemap URL Extractor & Change Monitor
- Paste domains or sitemap URLs into Sitemaps or domains.
apify.com,https://apify.comandhttps://www.allbirds.com/sitemap.xmlall work — discovery, index recursion and gzip are automatic. - Set Max URLs per source to the number of rows you want per site (default 1,000, maximum 200,000). The cap counts across every child sitemap of an index.
- Optionally narrow the crawl with URL filter (regex) —
/products/for a Shopify catalogue,/blog/for content,\.pdf$for documents. - Click Start, then export as JSON, CSV, Excel or HTML — or schedule the Actor in
diffmode and wire the Changes view into a webhook.
Example input
{"sources": ["https://apify.com", "https://www.allbirds.com/sitemap.xml"],"maxUrlsPerSource": 1000,"urlFilter": "","mode": "list","maxConcurrency": 5}
Example output
{"source": "https://www.allbirds.com/sitemap.xml","sitemapUrl": "https://www.allbirds.com/sitemap_products_1.xml?from=1878194389061&to=7369944137808","url": "https://www.allbirds.com/products/mens-wool-runners-natural-white","lastmod": "2026-09-12T16:30:03.000Z","changefreq": "daily","priority": null,"imageCount": 1,"videoCount": 0,"change": null,"error": null,"scrapedAt": "2026-09-12T17:05:00.000Z"}
Input parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
sources | array | ["https://apify.com"] | Sitemap URLs or bare domains; robots.txt discovery for domains |
maxUrlsPerSource | integer | 1000 | URL cap per source across all child sitemaps (1-200,000) |
urlFilter | string | "" | Optional regex; only matching URLs are kept |
mode | string | list | list = every URL, diff = new/unchanged/removed since last run |
diffStoreName | string | sitemap-monitor | Named key-value store holding the between-run URL snapshots |
maxConcurrency | integer | 5 | Sitemap files downloaded in parallel (1-20) |
Pricing
Sitemap URL Extractor & Change Monitor uses pay-per-event pricing: $0.0002 per URL row, i.e. $0.20 per 1,000 URLs, plus a negligible actor-start fee, platform usage included. A sitemap is one small HTTP request per file — the 1,591-URL apify.com page sitemap is a single 132 KB download — so compute stays in the cents even for 200,000-URL catalogues. Set Maximum cost per run and the Actor trims the result set to what the budget covers instead of overspending.
Sitemap URL Extractor & Change Monitor vs. a full site crawl
Sitemap URL Extractor & Change Monitor gets you a site's URL inventory in seconds instead of hours. A crawler has to fetch and parse every page to find the next link, burning proxy bandwidth and tripping rate limits; a sitemap is the site's own published index, already complete, already dated. Use this Actor to build the URL list, then feed it to a page-level scraper — you crawl only what you meant to. Versus Screaming Frog or a curl | grep one-liner, you also get gzip and index recursion handled, lastmod normalised to ISO, and a persistent diff between runs that no one-shot tool gives you.
Using Sitemap URL Extractor & Change Monitor with AI agents and MCP
Sitemap URL Extractor & Change Monitor is pay-per-event with limited permissions — the two requirements for an Actor to be callable through the Apify MCP server at mcp.apify.com. An agent passes sources and gets back a structured URL inventory it can use to decide what to read next, without writing a crawler. The same run works from n8n, Make, Zapier and LangChain through Apify's integrations; in diff mode a scheduled run plus a webhook gives you "alert me when this competitor publishes a new page".
FAQ
How does diff mode remember the previous run? Each source's URL set is stored in the named key-value store from Diff snapshot store name (default sitemap-monitor) under a key derived from a SHA-1 of the source string. Keep the store name identical across scheduled runs so they compare against each other. The first run has no snapshot, so every URL is reported as new; the second run of an unchanged site reports every URL as unchanged. A source that fails to fetch never overwrites its snapshot, so one transient error cannot report a whole site as removed.
Are .xml.gz sitemaps supported? Yes. Some servers send Content-Encoding: gzip and some send raw gzip bytes as binary/octet-stream — the Actor sniffs the gzip magic number instead of trusting the file extension, so both work.
What are the limitations? Only what the site publishes: pages missing from the sitemap are invisible here, and lastmod is the site's claim, not a verified fact. An index is followed up to 500 child sitemaps and 4 levels deep. Sitemaps behind an aggressive CDN bot wall can return HTML or 403; that source then produces one row with the reason in error instead of failing the run.
Is this legal to run? Yes. robots.txt and sitemaps exist specifically to be read by automated clients, and no personal data is collected.
Can I export to CSV or Excel? Yes, from the Output tab or the API, with ready-made Overview and Changes (diff mode) views.
Related Actors
Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:
Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.
Website & domain intelligence
- Email Extractor — Website Contact & Social Finder — e-mails, phones and social profiles per domain
- Tech Stack Detector — Wappalyzer & BuiltWith Alternative — CMS, e-commerce, analytics, pixels and payments per domain
- Domain DNS & Email Security Checker — SPF, DKIM, DMARC, MX provider, registrar and domain age
- Domain Security Audit (TLS, HTTP headers, redirects, robots) — TLS expiry, security headers, redirect chain, robots and llms.txt
- Subdomain Finder (Certificate Transparency) — every subdomain seen in CT logs, with a live DNS check
- Bulk Core Web Vitals & PageSpeed Audit — Lighthouse scores, LCP, CLS, INP and top fixes per URL
- On-Page SEO Audit — title, meta, headings, links, images and schema issues per page
- Wayback Machine Snapshot & Page Change Tracker — how a page changed over time, or every archived snapshot
- Bulk Domain WHOIS & RDAP Lookup — registrar, dates, status and nameservers per domain
- Web Scraper — CSS Selector & Data Extractor — pull any CSS selector off any page, one row per URL
- Website Screenshot Generator — full-page or viewport PNG/JPEG screenshots of any URL
Content for AI, LLMs and RAG
- AI Web Search & Read: Google results as clean Markdown — a query turned into clean Markdown from the top search results
- Website to Markdown — Content Crawler for LLM & RAG — any site as clean Markdown per page, no browser
- Article & News Extractor (clean text, author, date, markdown) — clean article text, author, date and Markdown per URL
- Structured Data & JSON-LD Extractor (Schema.org, Open Graph) — Schema.org and Open Graph data from any page
- Google News Scraper (RSS search by keyword, topic, site) — news results by keyword, topic or site
- Press Release Monitor: PR Newswire, BusinessWire, GlobeNewswire — PR Newswire, Business Wire and GlobeNewswire releases
Search, video and social
- YouTube Shorts Scraper — Shorts from channels, hashtags and searches with view counts
- Pinterest Pins Scraper — latest pins of public Pinterest profiles and boards
- YouTube Transcript Scraper — captions and subtitles as text + timed segments, per video or channel
- Google Search Results Scraper — SERP API — organic SERP results per keyword and country
- YouTube Comments Scraper — Comments & Replies — comments and replies with likes, no API key
- YouTube Channel Latest Videos (RSS, no API key) — the latest 15 videos of any channel from RSS
- YouTube Channel Scraper (videos, shorts, live) — a channel's full video, shorts and stream list
- YouTube Search Results Scraper (videos, channels, no API key) — videos, channels and playlists per query
- YouTube Video Details Scraper (views, likes, description, tags) — views, likes, description, tags and chapters per video
- Apple Podcasts Lookup & Episodes Scraper — podcast metadata and episodes from iTunes and RSS
- Bluesky Post, Search & Profile Scraper — posts, profiles, followers and threads from the AT Protocol API
- Telegram Channel Posts Scraper — posts, views and media flags from any public channel
- Substack Publication & Posts Scraper — archive, authors and paywall status per publication
- Google Play Reviews Scraper — reviews, ratings, replies and app versions per app
- App Store Reviews Scraper — iOS reviews and ratings per app and country
- Google Trends Scraper — interest over time, by region, and related queries per keyword
- Google Ads Transparency Scraper — ads any advertiser runs on Google, with format and dates
- Keyword Suggestions Scraper (Google, YouTube, Amazon, Bing) — autocomplete keyword ideas from four search engines
- Bilibili Scraper (Videos, Search, Popular) — Chinese video platform: views, likes, coins, danmaku, uploader
- Mastodon Scraper (Hashtags, Accounts, Trending) — public fediverse posts by hashtag, account or trending
- Meetup Events Scraper (Search by Keyword & City) — upcoming events with RSVPs, fees, venues and groups
- Eventbrite Scraper (Events by Keyword & City) — events by keyword and city with venue, dates and organizer
Leads, jobs and company data
- Career Site Jobs API (Greenhouse, Lever, Ashby, Workday +1) — company domains in, their open jobs out, ATS detected automatically
- Workday Jobs Scraper — jobs with full descriptions from any Workday career site
- Google Maps Scraper — businesses with phone, website, address, rating and coordinates per search
- LinkedIn Jobs Scraper — job titles, companies, locations and full descriptions from LinkedIn job search
- Company 360: full company profile from a domain — one row per domain: contacts, tech, security, hiring and company facts
- Hiring Signals Scraper (Greenhouse, Lever, Ashby, Workable) — open jobs and hiring velocity from 10 public ATS boards
- Y Combinator Companies & Founders Scraper — YC startups by batch, industry and hiring status
- Wikidata Entity & Company Enrichment (facts, IDs, links) — HQ, founders, employees, revenue and social IDs per company
- Email Validator & Verifier — Bulk Email Check — syntax, MX, disposable, role and free-provider checks
- OpenStreetMap POI Extractor (Overpass API: shops, amenities) — shops and amenities by radius, bbox or area
- Stock, Crypto & FX Quotes — one row per symbol from Yahoo, Binance and ECB rates
- Remote Jobs Aggregator (RemoteOK, WWR, Hacker News) — one clean row per remote job, de-duplicated across feeds
- Greenhouse Jobs Scraper — jobs with descriptions from any Greenhouse job board
- Lever Jobs Scraper — jobs with descriptions from any Lever careers page
- Ashby Jobs Scraper — jobs, salaries and descriptions from any Ashby job board
- SmartRecruiters Jobs Scraper — jobs with descriptions from any SmartRecruiters company
- Seek Jobs Scraper (Australia & New Zealand) — Seek job ads with salary, work type and location
- Dice Jobs Scraper — US tech jobs from Dice with salary and remote flag
- AutoScout24 Scraper — European car listings with price, mileage and seller
- Rightmove Scraper — UK property for sale or rent with price and agent
- Wellfound Jobs Scraper (AngelList Startup Jobs) — startup jobs with salary and equity ranges, company size and stage
- Yandex Maps Scraper (Places, Ratings, Phones) — businesses in Russia, Türkiye and the CIS with phones, ratings, hours
- Craigslist Scraper (Listings, Prices, Locations) — listings in any area and category with price, date and coordinates
- JobStreet Scraper (Malaysia, Singapore, PH, ID + JobsDB) — JobStreet and JobsDB jobs in 6 Asian countries with parsed salaries
- InfoJobs Scraper (Spain Jobs, Salaries, Companies) — Spanish jobs with salary range, contract type and full description
- Redfin Scraper (Homes for Sale, Prices, Details) — US homes for sale or sold from any Redfin search, with price and details
- Kleinanzeigen Scraper (Ads, Prices, Locations) — German classifieds with price, VB flag, ZIP, city and seller type
Developer, app and research data
- npm, PyPI & Crates.io Package Health Checker — releases, downloads, deprecation and a health score
- GitHub Repository Health & Activity Report — stars, commits, contributors and risk flags per repo
- VS Code Marketplace Extension Scraper (installs, ratings) — installs, ratings and versions per extension
- Chrome Web Store Extension Scraper (installs, ratings) — users, rating, version and developer per extension
- Google Play Scraper — apps, ratings, installs, developer contact and reviews
- App Store (iOS) App Metadata, Ratings & Top Charts Lookup — ratings, price, version and charts per app
- CrossRef DOI & Citation Metadata Lookup — papers, authors, journals and citation counts
- FDA Recalls & Adverse Events Monitor (openFDA) — food, drug and device recalls from openFDA
- iCal / ICS Calendar Feed to Events Extractor — any public calendar feed as event rows
- Shopify Store Products Scraper — catalog, prices, variants and stock per store
- Hacker News Search & Front Page Scraper — stories, comments and points by query or front page
- GitHub Trending Repositories Scraper — trending repos and developers by language and period
- Stack Overflow & Stack Exchange Q&A Scraper — questions, answers and scores by query, tag or site
- Bulk Image Downloader — download image URLs to storage with size, dimensions and a ZIP
- Google Flights Scraper (Prices, Airlines, Stops) — flight prices, airlines, times, stops and CO2 by route and date
- Google Hotels Scraper (Prices, Ratings, Reviews) — hotel prices per night, stars, rating and reviews by city and dates
- AliExpress Scraper (Search Products & Prices) — AliExpress search results with USD price, discount and rank
- Lazada Scraper (Products, Prices, Sold, Ratings) — Lazada products in 6 countries with price, rating, units sold and seller
Support and feedback
Hit a sitemap flavour it parses wrong, or want another extension counted? Open an issue on the Issues tab.