Sitemap URL Extractor & Scraper - All Page URLs of a Site avatar

Sitemap URL Extractor & Scraper - All Page URLs of a Site

Pricing

Pay per event

Go to Apify Store
Sitemap URL Extractor & Scraper - All Page URLs of a Site

Sitemap URL Extractor & Scraper - All Page URLs of a Site

Sitemap scraper and URL extractor: give it a website or sitemap.xml and get every page URL with lastmod, changefreq and priority. Finds sitemaps via robots.txt, follows sitemap indexes, reads .xml.gz and text sitemaps, survives broken XML. Plain HTTP, no browser. $0.20 per 1,000 URLs.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Data Gleaner

Data Gleaner

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Sitemap URL Extractor & Scraper - Extract All URLs From Any Website

Sitemap extractor that gets every page URL from a website's sitemap.xml: give it a domain, get all URLs with lastmod, changefreq and priority as JSON or CSV. $0.20 per 1,000 URLs.

It finds the sitemaps through robots.txt and common paths, follows sitemap indexes, reads .xml.gz and plain-text sitemaps, and returns one row per URL with last-modified date, change frequency, priority, image URLs, hreflang alternates and the sitemap it came from. An optional status check (checkStatus) adds each page's HTTP status and redirect target, for SEO audits. Plain HTTP, no browser, no login, no proxy needed for most sites.

It is built to finish the run. If one site has no sitemap, one sitemap file is broken, a host is down or a server returns an error, the Actor logs it, skips it and carries on with the rest. Every run ends with a per-site summary in the SUMMARY key-value record.

What does Sitemap URL Extractor do?

It finds a site's sitemaps, reads them all, and returns the page URLs as JSON, CSV or Excel through the Apify dataset and API. Typical uses:

  • SEO audits and site inventories. List every indexable URL of a site, with lastmod, before a migration or a crawl.
  • Competitor research. See how big a competitor's site is and which sections it grows (filter with includeUrlPatterns such as /blog/ or /products/).
  • Seed lists for scrapers. Feed the URLs into a crawler or a product / article scraper instead of crawling links.
  • Change monitoring. Run daily and compare lastmod to find pages that are new or updated.
  • LLM and RAG pipelines. Get the page list of a documentation site, then fetch and embed only the pages you want.

How sitemaps are found

  1. robots.txt Sitemap: lines.
  2. If those give nothing: /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /wp-sitemap.xml (WordPress).
  3. Sitemap index files are followed up to maxCrawlingDepth levels (default 6). Sitemap files of an index are fetched 4 at a time, so large sites finish faster; requestDelaySeconds still spaces the starts. Gzip files (.xml.gz) are unpacked by content, plain-text sitemaps (one URL per line) are read, and broken or truncated XML is recovered as far as possible.

You can also pass a direct sitemap URL (ending in .xml, .xml.gz or .txt) instead of a site.

Input

FieldTypeWhat it does
websiteslist of textWebsite URLs or bare domains (https://aioseo.com, stripe.com). Only the site root is used. Leave it empty to run a 5-URL example on aioseo.com.
startUrlslistSame as websites in the request-list format, so input written for other sitemap Actors works unchanged. Used only when websites is empty.
maxUrlsPerSitenumberCap per website. Default 5,000 when left out (the Console form prefills 100). It is also your cost cap per site.
checkStatusbooleanDefault false. Sends one HEAD request per URL (GET if HEAD is refused) and adds status, finalUrl and checkError, to find 404s, 5xx errors and redirects listed in the sitemap. Fills a missing lastmod from the page's Last-Modified header (lastmodSource = http-header). Slower, about 5 pages a second. No extra charge.
includeUrlPatternslist of textOptional regular expressions (Python syntax). Keep a URL only if it matches at least one. An invalid expression fails the run before anything is charged.
excludeUrlPatternslist of textOptional regular expressions. Drop a URL if it matches any. Invalid ones fail the run too.
requestDelaySecondsnumberPause between sitemap and robots.txt requests, default 0.3 s (randomised up to 1.5x). Status checks are not paced by it.
maxCrawlingDepthnumberHow many levels of sitemap indexes to follow. Default 6, range 1 to 20. 1 reads the top-level index's child sitemaps but not indexes nested inside it.
maxRequestRetriesnumberRetries per sitemap file on timeouts, connection errors, 429 and 5xx gateway errors (500, 502, 503, 504), with exponential backoff that honours Retry-After. Default 3, range 0 to 10.
proxyConfigurationobjectOptional Apify Proxy. Default is no proxy.
{
"websites": ["https://aioseo.com", "stripe.com"],
"maxUrlsPerSite": 5000,
"includeUrlPatterns": ["/blog/"],
"excludeUrlPatterns": ["\\.pdf$"]
}

URLs are deduplicated per site, so a page listed in several sitemaps is returned once.

Output

One dataset item per page URL. Export as JSON, CSV, Excel or via the API.

FieldMeaning
urlThe page URL.
lastmodLast modification date as ISO 8601 (2026-09-30 or 2026-09-30T10:00:00+00:00), or null if the sitemap has none or an invalid value.
changefreqalways, hourly, daily, weekly, monthly, yearly, never, or null.
priorityNumber from 0 to 1, or null.
sitemapUrlThe sitemap file this URL came from.
lastmodSourcesitemap, http-header (only with checkStatus), or null when there is no lastmod.
domainHost name of the site.
imageCountNumber of image:image entries listed for the page (0 when none).
imagesImage URLs (image:loc) listed for the page.
alternateshreflang language versions as {hreflang, href} objects.
statusHTTP status of the page. Only with checkStatus.
finalUrlURL after redirects. Only with checkStatus.
checkErrorNetwork error if the status check failed. Only with checkStatus.
websiteThe input value this row came from.
scrapedAtISO 8601 UTC timestamp.
{
"url": "https://stripe.com/payments",
"lastmod": "2026-09-30",
"changefreq": "weekly",
"priority": 0.8,
"sitemapUrl": "https://stripe.com/sitemap/partition-0.xml",
"domain": "stripe.com",
"lastmodSource": "sitemap",
"imageCount": 0,
"images": [],
"alternates": [],
"website": "stripe.com",
"scrapedAt": "2026-10-08T09:30:00+00:00"
}

The dataset has two views: Page URLs and SEO audit (status, redirects, lastmodSource, hreflang). Many sites publish only <loc>. Then lastmod, changefreq and priority are null; that is what the site publishes, not a failed extraction.

Pricing

Pay per event: $0.20 per 1,000 URLs ($0.0002 for each URL saved to the dataset). With checkStatus on, each URL is instead charged as a checked URL at $0.40 per 1,000 ($0.0004, one event per URL, never both), because every URL costs an extra HTTP request. You pay only for URLs you receive: nothing for sites without a sitemap, failed requests or retries, and the run stops by itself when your maximum total charge is reached. Platform usage is included.

Example: 20 sites at the default 5,000 URLs is up to 100,000 URLs, so up to $20. A site with 1,200 URLs costs $0.24.

Use it from the API

Run it synchronously and get the URLs back as JSON. Replace <YOUR_APIFY_TOKEN> with your token.

curl -X POST "https://api.apify.com/v2/acts/datagleaner~sitemap-extractor/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
-H "Content-Type: application/json" \
-d '{"websites": ["stripe.com"], "maxUrlsPerSite": 1000, "includeUrlPatterns": ["/docs/"]}'

It also works from n8n, Make or Zapier through the Apify integrations.

Use with Python

Install the client with pip install apify-client and set APIFY_TOKEN. This run returns at most 10 URLs, so it costs at most $0.002.

import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("datagleaner/sitemap-extractor").call(run_input={
"websites": ["stripe.com"],
"maxUrlsPerSite": 10,
})
for item in client.dataset(run.default_dataset_id).iterate_items():
print(item["url"], item["lastmod"])

Use with JavaScript / Node.js

Install the client with npm install apify-client and set APIFY_TOKEN. Save as run.mjs and run node run.mjs.

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagleaner/sitemap-extractor').call({
websites: ['stripe.com'],
maxUrlsPerSite: 10,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) console.log(item.url, item.lastmod);

Use it from n8n, Make, Zapier or an AI agent

Actor ID: datagleaner/sitemap-extractor

Minimal input:

{"websites": ["stripe.com"], "maxUrlsPerSite": 10}

Each tool below runs this Actor with your own Apify API token.

  • n8n: add the Apify node (@apify/n8n-nodes-apify). On n8n Cloud you install it from the community node registry. Choose Run an Actor and get dataset, set Actor to datagleaner/sitemap-extractor and paste the input above.

  • Make: use the Apify app's Run an Actor module, then Get Dataset Items to read the results. Watch Actor Runs can trigger a scenario when a run finishes.

  • Zapier: use the Apify action Run Actor, then the search Fetch dataset items. The trigger Finished Actor run starts a Zap when a run ends.

  • AI agents (MCP): connect to https://mcp.apify.com?tools=datagleaner/sitemap-extractor. In Claude Code:

    claude mcp add --transport http apify "https://mcp.apify.com?tools=datagleaner/sitemap-extractor"
    

    Then run /mcp to sign in to Apify in your browser. Then ask in plain words, for example: "Use datagleaner/sitemap-extractor to list every /blog/ URL on stripe.com, capped at 500, and tell me which ones were updated in the last 30 days."

  • LangChain (Python):

# pip install langchain-apify, then set APIFY_TOKEN in your environment
import json
from langchain_apify import ApifyActorsTool
tool = ApifyActorsTool("datagleaner/sitemap-extractor")
result = tool.invoke({"run_input": json.loads('{"websites": ["stripe.com"], "maxUrlsPerSite": 10}')})

Limits

  • Only what the site publishes. A page missing from every sitemap is not found. Sites without any sitemap return nothing; the log and SUMMARY say so.
  • Per-site cap. maxUrlsPerSite stops a site early. Very large sites (millions of URLs) need a higher cap or patterns that narrow them.
  • Blocking. Sites behind bot protection may answer 403 to datacenter IPs. Enable proxyConfiguration or raise requestDelaySeconds.
  • Index nesting is followed up to maxCrawlingDepth levels (default 6) and 3,000 sitemap files per site, as a safety stop.
  • Status message. It reports totals, sites stopped at maxUrlsPerSite, sites with no URLs, how many sitemap files were "empty or unreadable", and with checkStatus how many pages were "not 2xx".

FAQ

How do I extract all URLs from a website? Enter the domain in websites. The Actor reads every sitemap the site publishes and returns each page URL as a dataset row. Pages left out of every sitemap are not found, because it does not crawl links.

How do I count a site's sitemap URLs? Run it with a maxUrlsPerSite above the site's size. The urls field for each site in the SUMMARY record, and the dataset's item count, give the total.

Does it need a browser or login? No, plain HTTP only.

What if a site's sitemap is broken? The Actor recovers the URLs it can read from damaged XML, skips files it cannot read and continues.

Are .xml.gz sitemaps supported? Yes, detected by content, so a gzip file with a wrong name or header still works.

Why fewer URLs than maxUrlsPerSite? The site lists fewer, your patterns filtered some out, or duplicates were removed.

Is it a sitemap scraper or a crawler? A sitemap scraper: it reads the sitemap files a site publishes and does not crawl links, so it is fast and cheap, but it only finds pages the site lists.

Can I export the sitemap URLs to CSV or Excel? Yes. Open the run's dataset and export as CSV, Excel, JSON or XML, or fetch it from the API.

Can I run it with no input? Yes. It runs a built-in example on aioseo.com, capped at 5 URLs, and finishes in well under a minute.

Responsible use

This Actor reads only sitemap files that sites publish for crawlers, and its default pacing is polite. Respect each site's terms and robots.txt rules when you use the URLs afterwards. URLs can contain personal data in rare cases (for example a person's name in a profile path): you are responsible for a lawful basis and your obligations under GDPR and any other law that applies. This is not legal advice.