Sitemap URL Extractor avatar

Sitemap URL Extractor

Pricing

$0.30 / 1,000 url extracteds

Go to Apify Store
Sitemap URL Extractor

Sitemap URL Extractor

Extract every URL from a website's sitemaps: indexes, .gz, plain-text and RSS/Atom, with lastmod filtering.

Pricing

$0.30 / 1,000 url extracteds

Rating

0.0

(0)

Developer

Ali Alsaidi

Ali Alsaidi

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Point it at a domain or a sitemap and get back every page URL the site publishes — clean, de-duplicated, and filtered how you want it.

What does Sitemap URL Extractor do?

  • Turns one domain into a full URL list. Give it example.com and it reads robots.txt, then tries nine common sitemap paths until it finds one.
  • Handles the sitemaps that break other extractors — indexes nested several levels deep, gzipped .xml.gz files, plain-text and RSS/Atom sitemaps, and XML with a malformed tag in the middle (it returns every URL it can read instead of failing the run).
  • Streams instead of loading. A 500,000-URL sitemap is parsed in chunks and memory stays flat. Measured: 63,000 URLs/second, 10 MB peak on a 4 MB sitemap.
  • Filters before it charges you. Regex include/exclude and a lastmodAfter date cut both the noise and the bill.
  • Does not check every URL by default. Most extractors fire an HTTP request per URL to report a status code, which turns a 3-second job into a 30-minute one. Here that is a switch (checkStatus), off unless you ask.

What data can you get?

FieldWhat it means
urlThe page URL, absolute and de-duplicated
lastmodLast-modified date the site publishes, or null if it publishes none
changefreqHow often the site claims the page changes
priorityThe site's own priority hint, 0.0–1.0
sitemapWhich sitemap file this URL came from
depthHow many index levels deep that sitemap was
statusHTTP status — only when checkStatus is on
alternateshreflang versions of the page — only when includeAlternates is on
imagesImage URLs listed for the page — only when includeImages is on

How to use it

  1. Create a free Apify account.
  2. Open the Actor and click Try for free.
  3. Put your domain or sitemap URL in Start URLs, and set Maximum URLs to cap the run.
  4. Click Start. A typical site finishes in seconds.
  5. Open the Dataset tab and download as JSON, CSV, Excel or XML, or pull it through the Apify API.

Input

FieldTypeDefaultWhat it does
startUrlsarray—Domains or sitemap URLs (required)
maxUrlsinteger1000Stop after this many URLs. A hard ceiling on run cost
checkStatusbooleanfalseAdd an HTTP status per URL (slower)
includePatternsarray[]Keep only URLs matching these regexes, e.g. /blog/
excludePatternsarray[]Drop URLs matching these, e.g. /tag/
lastmodAfterstring—Only URLs changed after a date, e.g. 2026-01-01
includeAlternatesbooleanfalseAdd hreflang versions of each URL
includeImagesbooleanfalseAdd image URLs listed for the page
maxDepthinteger5How deep to follow nested sitemap indexes
concurrencyinteger5Parallel sitemap fetches
requestTimeoutSecsinteger30Per-request timeout
maxRetriesinteger3Retries per request
respectRobotsTxtbooleantrueObey robots.txt
proxyConfigurationobject—Optional; sitemaps are public, so rarely needed
{
"startUrls": [{ "url": "https://apify.com" }],
"maxUrls": 5000,
"excludePatterns": ["/tag/"],
"lastmodAfter": "2026-01-01"
}

Output

One item per URL. This is a real item from a run against apify.com:

{
"url": "https://apify.com/ai-agents",
"lastmod": null,
"changefreq": null,
"priority": null,
"sitemap": "https://apify.com/sitemap/pages.xml",
"depth": 1
}

The nulls are honest: apify.com's sitemap publishes no lastmod, changefreq or priority, and this Actor reports what is there rather than inventing it. Many sites do publish all three.

Every run also writes a RUN_SUMMARY record to the key-value store with urlsReturned, urlsFound, sitemapsParsed, sitemapsFailed, requests, retries, blockedByRobotsTxt and a problems list. A partial result is never a mystery.

How much does it cost?

$0.30 per 1,000 URLs returned ($0.0003 each), and nothing else. 10,000 URLs = $3. Apify platform usage — compute, storage, transfer — is included in that price, not billed on top.

Filters are applied before charging, so includePatterns, excludePatterns and lastmodAfter reduce the bill as well as the noise. maxUrls is a hard ceiling. Re-running a 5,000-URL site monthly, filtered to pages changed since last time, usually costs a few cents.

Integrations

Connect it to webhooks, Make, Zapier, Google Sheets, Slack, Airbyte or GitHub from the Integrations tab — for example, run it weekly and append new URLs to a sheet.

Call it from your own code with the Apify API:

from apify_client import ApifyClient
client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("palgenius/sitemap-url-extractor").call(
run_input={"startUrls": [{"url": "https://apify.com"}], "maxUrls": 5000}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item["url"])

AI agents can call it through the Apify MCP server at https://mcp.apify.com, e.g. https://mcp.apify.com?tools=palgenius/sitemap-url-extractor.

Error items and limits

  • No sitemap found. The run fails with a status message naming the first problem, and RUN_SUMMARY.problems lists every path that was tried. You are not charged for URLs that do not exist.
  • A sitemap could not be read (timeout, 5xx, broken gzip): counted in sitemapsFailed and described in problems; the other sitemaps still return their URLs.
  • robots.txt disallows a sitemap path: skipped, and counted in blockedByRobotsTxt.
  • lastmod is whatever the site publishes. Many sites omit it and some never update it. URLs without a lastmod are dropped when you set lastmodAfter.
  • A sitemap can list URLs that no longer exist. Turn on checkStatus, or pipe the output into the Bulk URL Status Checker below.

FAQ

Is this legal? Yes. Sitemaps are files sites publish specifically so crawlers can read them. This Actor reads only public sitemap files, obeys robots.txt by default, never logs in, and collects no personal data.

What if I only know the domain? That is the normal case. It reads robots.txt and tries nine common sitemap locations.

Will a huge sitemap time out? It streams and parses in chunks, so size is not the limit — maxUrls and your run timeout are.

Can I get only pages changed since last month? Set lastmodAfter, on sites that publish lastmod.

Do I pay for URLs I filter out? No. Filtering happens before charging.

Can I check which of these URLs are broken? Turn on checkStatus, or use the status checker for redirect chains and response times.

You might also like

  • Bulk URL Status Checker — feed it this run's dataset ID to check every URL for broken links and redirect chains.
  • Tech Stack Detector — find the platform, CMS and frameworks behind any domain, with evidence.

Feedback

Found a sitemap it cannot read, or want a field added? Open an issue on the Actor's Issues tab. Fixes ship within days.