Sitemap URL Extractor
Pricing
$0.30 / 1,000 url extracteds
Sitemap URL Extractor
Extract every URL from a website's sitemaps: indexes, .gz, plain-text and RSS/Atom, with lastmod filtering.
Pricing
$0.30 / 1,000 url extracteds
Rating
0.0
(0)
Developer
Ali Alsaidi
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Point it at a domain or a sitemap and get back every page URL the site publishes — clean, de-duplicated, and filtered how you want it.
What does Sitemap URL Extractor do?
- Turns one domain into a full URL list. Give it
example.comand it readsrobots.txt, then tries nine common sitemap paths until it finds one. - Handles the sitemaps that break other extractors — indexes nested several levels deep,
gzipped
.xml.gzfiles, plain-text and RSS/Atom sitemaps, and XML with a malformed tag in the middle (it returns every URL it can read instead of failing the run). - Streams instead of loading. A 500,000-URL sitemap is parsed in chunks and memory stays flat. Measured: 63,000 URLs/second, 10 MB peak on a 4 MB sitemap.
- Filters before it charges you. Regex include/exclude and a
lastmodAfterdate cut both the noise and the bill. - Does not check every URL by default. Most extractors fire an HTTP request per URL to
report a status code, which turns a 3-second job into a 30-minute one. Here that is a
switch (
checkStatus), off unless you ask.
What data can you get?
| Field | What it means |
|---|---|
url | The page URL, absolute and de-duplicated |
lastmod | Last-modified date the site publishes, or null if it publishes none |
changefreq | How often the site claims the page changes |
priority | The site's own priority hint, 0.0–1.0 |
sitemap | Which sitemap file this URL came from |
depth | How many index levels deep that sitemap was |
status | HTTP status — only when checkStatus is on |
alternates | hreflang versions of the page — only when includeAlternates is on |
images | Image URLs listed for the page — only when includeImages is on |
How to use it
- Create a free Apify account.
- Open the Actor and click Try for free.
- Put your domain or sitemap URL in Start URLs, and set Maximum URLs to cap the run.
- Click Start. A typical site finishes in seconds.
- Open the Dataset tab and download as JSON, CSV, Excel or XML, or pull it through the Apify API.
Input
| Field | Type | Default | What it does |
|---|---|---|---|
startUrls | array | — | Domains or sitemap URLs (required) |
maxUrls | integer | 1000 | Stop after this many URLs. A hard ceiling on run cost |
checkStatus | boolean | false | Add an HTTP status per URL (slower) |
includePatterns | array | [] | Keep only URLs matching these regexes, e.g. /blog/ |
excludePatterns | array | [] | Drop URLs matching these, e.g. /tag/ |
lastmodAfter | string | — | Only URLs changed after a date, e.g. 2026-01-01 |
includeAlternates | boolean | false | Add hreflang versions of each URL |
includeImages | boolean | false | Add image URLs listed for the page |
maxDepth | integer | 5 | How deep to follow nested sitemap indexes |
concurrency | integer | 5 | Parallel sitemap fetches |
requestTimeoutSecs | integer | 30 | Per-request timeout |
maxRetries | integer | 3 | Retries per request |
respectRobotsTxt | boolean | true | Obey robots.txt |
proxyConfiguration | object | — | Optional; sitemaps are public, so rarely needed |
{"startUrls": [{ "url": "https://apify.com" }],"maxUrls": 5000,"excludePatterns": ["/tag/"],"lastmodAfter": "2026-01-01"}
Output
One item per URL. This is a real item from a run against apify.com:
{"url": "https://apify.com/ai-agents","lastmod": null,"changefreq": null,"priority": null,"sitemap": "https://apify.com/sitemap/pages.xml","depth": 1}
The nulls are honest: apify.com's sitemap publishes no lastmod, changefreq or priority,
and this Actor reports what is there rather than inventing it. Many sites do publish all three.
Every run also writes a RUN_SUMMARY record to the key-value store with urlsReturned,
urlsFound, sitemapsParsed, sitemapsFailed, requests, retries, blockedByRobotsTxt
and a problems list. A partial result is never a mystery.
How much does it cost?
$0.30 per 1,000 URLs returned ($0.0003 each), and nothing else. 10,000 URLs = $3. Apify platform usage — compute, storage, transfer — is included in that price, not billed on top.
Filters are applied before charging, so includePatterns, excludePatterns and
lastmodAfter reduce the bill as well as the noise. maxUrls is a hard ceiling. Re-running a
5,000-URL site monthly, filtered to pages changed since last time, usually costs a few cents.
Integrations
Connect it to webhooks, Make, Zapier, Google Sheets, Slack, Airbyte or GitHub from the Integrations tab — for example, run it weekly and append new URLs to a sheet.
Call it from your own code with the Apify API:
from apify_client import ApifyClientclient = ApifyClient("<YOUR_API_TOKEN>")run = client.actor("palgenius/sitemap-url-extractor").call(run_input={"startUrls": [{"url": "https://apify.com"}], "maxUrls": 5000})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item["url"])
AI agents can call it through the Apify MCP server at https://mcp.apify.com, e.g.
https://mcp.apify.com?tools=palgenius/sitemap-url-extractor.
Error items and limits
- No sitemap found. The run fails with a status message naming the first problem, and
RUN_SUMMARY.problemslists every path that was tried. You are not charged for URLs that do not exist. - A sitemap could not be read (timeout, 5xx, broken gzip): counted in
sitemapsFailedand described inproblems; the other sitemaps still return their URLs. robots.txtdisallows a sitemap path: skipped, and counted inblockedByRobotsTxt.lastmodis whatever the site publishes. Many sites omit it and some never update it. URLs without alastmodare dropped when you setlastmodAfter.- A sitemap can list URLs that no longer exist. Turn on
checkStatus, or pipe the output into the Bulk URL Status Checker below.
FAQ
Is this legal? Yes. Sitemaps are files sites publish specifically so crawlers can read
them. This Actor reads only public sitemap files, obeys robots.txt by default, never logs
in, and collects no personal data.
What if I only know the domain? That is the normal case. It reads robots.txt and tries
nine common sitemap locations.
Will a huge sitemap time out? It streams and parses in chunks, so size is not the limit —
maxUrls and your run timeout are.
Can I get only pages changed since last month? Set lastmodAfter, on sites that publish
lastmod.
Do I pay for URLs I filter out? No. Filtering happens before charging.
Can I check which of these URLs are broken? Turn on checkStatus, or use the status
checker for redirect chains and response times.
You might also like
- Bulk URL Status Checker — feed it this run's dataset ID to check every URL for broken links and redirect chains.
- Tech Stack Detector — find the platform, CMS and frameworks behind any domain, with evidence.
Feedback
Found a sitemap it cannot read, or want a field added? Open an issue on the Actor's Issues tab. Fixes ship within days.