Sitemap URL Extractor & Scraper - All Page URLs of a Site
Pricing
Pay per event
Sitemap URL Extractor & Scraper - All Page URLs of a Site
Sitemap scraper and URL extractor: give it a website or sitemap.xml and get every page URL with lastmod, changefreq and priority. Finds sitemaps via robots.txt, follows sitemap indexes, reads .xml.gz and text sitemaps, survives broken XML. Plain HTTP, no browser. $0.20 per 1,000 URLs.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Data Gleaner
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Sitemap URL Extractor & Scraper - Extract All URLs From Any Website
Sitemap extractor that gets every page URL from a website's sitemap.xml: give it a domain, get all URLs with lastmod, changefreq and priority as JSON or CSV. $0.20 per 1,000 URLs.
It finds the sitemaps through robots.txt and common paths, follows sitemap indexes, reads .xml.gz and plain-text sitemaps, and returns one row per URL with last-modified date, change frequency, priority, image URLs, hreflang alternates and the sitemap it came from. An optional status check (checkStatus) adds each page's HTTP status and redirect target, for SEO audits. Plain HTTP, no browser, no login, no proxy needed for most sites.
It is built to finish the run. If one site has no sitemap, one sitemap file is broken, a host is down or a server returns an error, the Actor logs it, skips it and carries on with the rest. Every run ends with a per-site summary in the SUMMARY key-value record.
What does Sitemap URL Extractor do?
It finds a site's sitemaps, reads them all, and returns the page URLs as JSON, CSV or Excel through the Apify dataset and API. Typical uses:
- SEO audits and site inventories. List every indexable URL of a site, with
lastmod, before a migration or a crawl. - Competitor research. See how big a competitor's site is and which sections it grows (filter with
includeUrlPatternssuch as/blog/or/products/). - Seed lists for scrapers. Feed the URLs into a crawler or a product / article scraper instead of crawling links.
- Change monitoring. Run daily and compare
lastmodto find pages that are new or updated. - LLM and RAG pipelines. Get the page list of a documentation site, then fetch and embed only the pages you want.
How sitemaps are found
robots.txtSitemap:lines.- If those give nothing:
/sitemap.xml,/sitemap_index.xml,/sitemap-index.xml,/wp-sitemap.xml(WordPress). - Sitemap index files are followed up to
maxCrawlingDepthlevels (default 6). Sitemap files of an index are fetched 4 at a time, so large sites finish faster;requestDelaySecondsstill spaces the starts. Gzip files (.xml.gz) are unpacked by content, plain-text sitemaps (one URL per line) are read, and broken or truncated XML is recovered as far as possible.
You can also pass a direct sitemap URL (ending in .xml, .xml.gz or .txt) instead of a site.
Input
| Field | Type | What it does |
|---|---|---|
websites | list of text | Website URLs or bare domains (https://aioseo.com, stripe.com). Only the site root is used. Leave it empty to run a 5-URL example on aioseo.com. |
startUrls | list | Same as websites in the request-list format, so input written for other sitemap Actors works unchanged. Used only when websites is empty. |
maxUrlsPerSite | number | Cap per website. Default 5,000 when left out (the Console form prefills 100). It is also your cost cap per site. |
checkStatus | boolean | Default false. Sends one HEAD request per URL (GET if HEAD is refused) and adds status, finalUrl and checkError, to find 404s, 5xx errors and redirects listed in the sitemap. Fills a missing lastmod from the page's Last-Modified header (lastmodSource = http-header). Slower, about 5 pages a second. No extra charge. |
includeUrlPatterns | list of text | Optional regular expressions (Python syntax). Keep a URL only if it matches at least one. An invalid expression fails the run before anything is charged. |
excludeUrlPatterns | list of text | Optional regular expressions. Drop a URL if it matches any. Invalid ones fail the run too. |
requestDelaySeconds | number | Pause between sitemap and robots.txt requests, default 0.3 s (randomised up to 1.5x). Status checks are not paced by it. |
maxCrawlingDepth | number | How many levels of sitemap indexes to follow. Default 6, range 1 to 20. 1 reads the top-level index's child sitemaps but not indexes nested inside it. |
maxRequestRetries | number | Retries per sitemap file on timeouts, connection errors, 429 and 5xx gateway errors (500, 502, 503, 504), with exponential backoff that honours Retry-After. Default 3, range 0 to 10. |
proxyConfiguration | object | Optional Apify Proxy. Default is no proxy. |
{"websites": ["https://aioseo.com", "stripe.com"],"maxUrlsPerSite": 5000,"includeUrlPatterns": ["/blog/"],"excludeUrlPatterns": ["\\.pdf$"]}
URLs are deduplicated per site, so a page listed in several sitemaps is returned once.
Output
One dataset item per page URL. Export as JSON, CSV, Excel or via the API.
| Field | Meaning |
|---|---|
url | The page URL. |
lastmod | Last modification date as ISO 8601 (2026-09-30 or 2026-09-30T10:00:00+00:00), or null if the sitemap has none or an invalid value. |
changefreq | always, hourly, daily, weekly, monthly, yearly, never, or null. |
priority | Number from 0 to 1, or null. |
sitemapUrl | The sitemap file this URL came from. |
lastmodSource | sitemap, http-header (only with checkStatus), or null when there is no lastmod. |
domain | Host name of the site. |
imageCount | Number of image:image entries listed for the page (0 when none). |
images | Image URLs (image:loc) listed for the page. |
alternates | hreflang language versions as {hreflang, href} objects. |
status | HTTP status of the page. Only with checkStatus. |
finalUrl | URL after redirects. Only with checkStatus. |
checkError | Network error if the status check failed. Only with checkStatus. |
website | The input value this row came from. |
scrapedAt | ISO 8601 UTC timestamp. |
{"url": "https://stripe.com/payments","lastmod": "2026-09-30","changefreq": "weekly","priority": 0.8,"sitemapUrl": "https://stripe.com/sitemap/partition-0.xml","domain": "stripe.com","lastmodSource": "sitemap","imageCount": 0,"images": [],"alternates": [],"website": "stripe.com","scrapedAt": "2026-10-08T09:30:00+00:00"}
The dataset has two views: Page URLs and SEO audit (status, redirects, lastmodSource, hreflang). Many sites publish only <loc>. Then lastmod, changefreq and priority are null; that is what the site publishes, not a failed extraction.
Pricing
Pay per event: $0.20 per 1,000 URLs ($0.0002 for each URL saved to the dataset). With checkStatus on, each URL is instead charged as a checked URL at $0.40 per 1,000 ($0.0004, one event per URL, never both), because every URL costs an extra HTTP request. You pay only for URLs you receive: nothing for sites without a sitemap, failed requests or retries, and the run stops by itself when your maximum total charge is reached. Platform usage is included.
Example: 20 sites at the default 5,000 URLs is up to 100,000 URLs, so up to $20. A site with 1,200 URLs costs $0.24.
Use it from the API
Run it synchronously and get the URLs back as JSON. Replace <YOUR_APIFY_TOKEN> with your token.
curl -X POST "https://api.apify.com/v2/acts/datagleaner~sitemap-extractor/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \-H "Content-Type: application/json" \-d '{"websites": ["stripe.com"], "maxUrlsPerSite": 1000, "includeUrlPatterns": ["/docs/"]}'
It also works from n8n, Make or Zapier through the Apify integrations.
Use with Python
Install the client with pip install apify-client and set APIFY_TOKEN. This run returns at most 10 URLs, so it costs at most $0.002.
import osfrom apify_client import ApifyClientclient = ApifyClient(os.environ["APIFY_TOKEN"])run = client.actor("datagleaner/sitemap-extractor").call(run_input={"websites": ["stripe.com"],"maxUrlsPerSite": 10,})for item in client.dataset(run.default_dataset_id).iterate_items():print(item["url"], item["lastmod"])
Use with JavaScript / Node.js
Install the client with npm install apify-client and set APIFY_TOKEN. Save as run.mjs and run node run.mjs.
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('datagleaner/sitemap-extractor').call({websites: ['stripe.com'],maxUrlsPerSite: 10,});const { items } = await client.dataset(run.defaultDatasetId).listItems();for (const item of items) console.log(item.url, item.lastmod);
Use it from n8n, Make, Zapier or an AI agent
Actor ID: datagleaner/sitemap-extractor
Minimal input:
{"websites": ["stripe.com"], "maxUrlsPerSite": 10}
Each tool below runs this Actor with your own Apify API token.
-
n8n: add the Apify node (
@apify/n8n-nodes-apify). On n8n Cloud you install it from the community node registry. Choose Run an Actor and get dataset, set Actor todatagleaner/sitemap-extractorand paste the input above. -
Make: use the Apify app's Run an Actor module, then Get Dataset Items to read the results. Watch Actor Runs can trigger a scenario when a run finishes.
-
Zapier: use the Apify action Run Actor, then the search Fetch dataset items. The trigger Finished Actor run starts a Zap when a run ends.
-
AI agents (MCP): connect to
https://mcp.apify.com?tools=datagleaner/sitemap-extractor. In Claude Code:claude mcp add --transport http apify "https://mcp.apify.com?tools=datagleaner/sitemap-extractor"Then run
/mcpto sign in to Apify in your browser. Then ask in plain words, for example: "Use datagleaner/sitemap-extractor to list every /blog/ URL on stripe.com, capped at 500, and tell me which ones were updated in the last 30 days." -
LangChain (Python):
# pip install langchain-apify, then set APIFY_TOKEN in your environmentimport jsonfrom langchain_apify import ApifyActorsTooltool = ApifyActorsTool("datagleaner/sitemap-extractor")result = tool.invoke({"run_input": json.loads('{"websites": ["stripe.com"], "maxUrlsPerSite": 10}')})
Limits
- Only what the site publishes. A page missing from every sitemap is not found. Sites without any sitemap return nothing; the log and
SUMMARYsay so. - Per-site cap.
maxUrlsPerSitestops a site early. Very large sites (millions of URLs) need a higher cap or patterns that narrow them. - Blocking. Sites behind bot protection may answer 403 to datacenter IPs. Enable
proxyConfigurationor raiserequestDelaySeconds. - Index nesting is followed up to
maxCrawlingDepthlevels (default 6) and 3,000 sitemap files per site, as a safety stop. - Status message. It reports totals, sites stopped at
maxUrlsPerSite, sites with no URLs, how many sitemap files were "empty or unreadable", and withcheckStatushow many pages were "not 2xx".
FAQ
How do I extract all URLs from a website? Enter the domain in websites. The Actor reads every sitemap the site publishes and returns each page URL as a dataset row. Pages left out of every sitemap are not found, because it does not crawl links.
How do I count a site's sitemap URLs? Run it with a maxUrlsPerSite above the site's size. The urls field for each site in the SUMMARY record, and the dataset's item count, give the total.
Does it need a browser or login? No, plain HTTP only.
What if a site's sitemap is broken? The Actor recovers the URLs it can read from damaged XML, skips files it cannot read and continues.
Are .xml.gz sitemaps supported? Yes, detected by content, so a gzip file with a wrong name or header still works.
Why fewer URLs than maxUrlsPerSite? The site lists fewer, your patterns filtered some out, or duplicates were removed.
Is it a sitemap scraper or a crawler? A sitemap scraper: it reads the sitemap files a site publishes and does not crawl links, so it is fast and cheap, but it only finds pages the site lists.
Can I export the sitemap URLs to CSV or Excel? Yes. Open the run's dataset and export as CSV, Excel, JSON or XML, or fetch it from the API.
Can I run it with no input? Yes. It runs a built-in example on aioseo.com, capped at 5 URLs, and finishes in well under a minute.
Responsible use
This Actor reads only sitemap files that sites publish for crawlers, and its default pacing is polite. Respect each site's terms and robots.txt rules when you use the URLs afterwards. URLs can contain personal data in rare cases (for example a person's name in a profile path): you are responsible for a lawful basis and your obligations under GDPR and any other law that applies. This is not legal advice.