Sitemap URL Extractor & Scraper — indexes, gzip, robots.txt
Pricing
from $0.40 / 1,000 url extracteds
Sitemap URL Extractor & Scraper — indexes, gzip, robots.txt
Sitemap scraper that extracts page URLs from public XML sitemaps. Discovers robots.txt, sitemap.xml and sitemap indexes, unpacks gzip, and filters by date or regex. $0.40 per 1,000 URLs.
Pricing
from $0.40 / 1,000 url extracteds
Rating
0.0
(0)
Developer
drop-in apis
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 hours ago
Last modified
Categories
Share
Sitemap URL extractor and scraper that reads a site's public sitemap — including sitemap indexes and gzip — and returns one row per page URL. It does not open those pages. $0.40 per 1,000 URLs.
Last updated: 2026-10-02.
Give it a domain or a sitemap URL. It looks up robots.txt, /sitemap.xml and /sitemap_index.xml, follows a sitemap index, and unpacks .xml.gz.
Use cases
- A URL list for another Actor. Feed the rows into a crawler that only needs the links, not a fresh crawl of the whole site.
- Change detection. Keep rows whose
lastmodis after a date you pass in. - A scoped export. Keep or drop URLs with a regular expression (
/blog/, a locale, a file type).
Input
| Field | Type | Description |
|---|---|---|
startUrls | array (required) | A domain (example.com or https://example.com) or a sitemap URL (.xml or .xml.gz). |
includeRegex | string (optional) | Case-insensitive regular expression. Only matching URLs are kept. |
excludeRegex | string (optional) | Case-insensitive regular expression. Matching URLs are dropped. |
lastmodSince | string (optional) | ISO date. URLs with an older lastmod are dropped. URLs with no lastmod are kept. |
maxUrls | integer | Stop after this many unique URLs. Default 1000. The Store's daily test uses 10. |
Example input:
{"startUrls": [{ "url": "https://www.sitemaps.org" }],"maxUrls": 10}
Output
One dataset row per unique page URL.
{"url": "https://www.sitemaps.org/","lastmod": "2016-11-21","changefreq": null,"priority": null,"sourceSitemap": "https://www.sitemaps.org/sitemap.xml","domain": "www.sitemaps.org"}
Pricing
Pay per event. The primary event is apify-default-dataset-item at $0.0004 per URL ($0.40 per 1,000). Apify charges it when a row is saved. There is no monthly fee.
Limits
- Public
httpandhttpssitemaps only. No login, and private hosts (localhost,10.x,192.168.x) are refused. - At most 50 sitemap files per run, and 20 MB per file after unpacking.
- At most 5 fetches in flight, with retries on HTTP 429 and 5xx.
Crawl-delayin robots.txt is honored up to 5 seconds. - A path
Disallowed forUser-agent: *is not fetched. - The run fails if every input produces zero URLs, so an empty success is not reported as a finished dataset.
FAQ
Does this crawl the pages listed in the sitemap?
No. It only downloads sitemap files and returns the URLs inside them.
What if the site publishes a sitemap index or a .gz file?
An index is followed, one child file at a time, until maxUrls is reached. A file that starts with the gzip magic bytes is unpacked. Both were checked against live sites: https://apify.com (index) and https://www.gnu.org (a child sitemap0.xml.gz).
What does a domain input fetch?
/robots.txt (every Sitemap: line), then /sitemap.xml and /sitemap_index.xml if those were not already listed. A start URL that already ends in .xml or .xml.gz is fetched directly.
Why did the run fail with "0 URL(s)"?
None of the inputs had a readable public sitemap. A missing /sitemap_index.xml on its own is normal and is not a failure. The failure means every candidate was missing or was not sitemap XML.