Sitemap URL Extractor & Scraper — indexes, gzip, robots.txt avatar

Sitemap URL Extractor & Scraper — indexes, gzip, robots.txt

Pricing

from $0.40 / 1,000 url extracteds

Go to Apify Store
Sitemap URL Extractor & Scraper — indexes, gzip, robots.txt

Sitemap URL Extractor & Scraper — indexes, gzip, robots.txt

Sitemap scraper that extracts page URLs from public XML sitemaps. Discovers robots.txt, sitemap.xml and sitemap indexes, unpacks gzip, and filters by date or regex. $0.40 per 1,000 URLs.

Pricing

from $0.40 / 1,000 url extracteds

Rating

0.0

(0)

Developer

drop-in apis

drop-in apis

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 hours ago

Last modified

Categories

Share

Sitemap URL extractor and scraper that reads a site's public sitemap — including sitemap indexes and gzip — and returns one row per page URL. It does not open those pages. $0.40 per 1,000 URLs.

Last updated: 2026-10-02.

Give it a domain or a sitemap URL. It looks up robots.txt, /sitemap.xml and /sitemap_index.xml, follows a sitemap index, and unpacks .xml.gz.

Use cases

  • A URL list for another Actor. Feed the rows into a crawler that only needs the links, not a fresh crawl of the whole site.
  • Change detection. Keep rows whose lastmod is after a date you pass in.
  • A scoped export. Keep or drop URLs with a regular expression (/blog/, a locale, a file type).

Input

FieldTypeDescription
startUrlsarray (required)A domain (example.com or https://example.com) or a sitemap URL (.xml or .xml.gz).
includeRegexstring (optional)Case-insensitive regular expression. Only matching URLs are kept.
excludeRegexstring (optional)Case-insensitive regular expression. Matching URLs are dropped.
lastmodSincestring (optional)ISO date. URLs with an older lastmod are dropped. URLs with no lastmod are kept.
maxUrlsintegerStop after this many unique URLs. Default 1000. The Store's daily test uses 10.

Example input:

{
"startUrls": [{ "url": "https://www.sitemaps.org" }],
"maxUrls": 10
}

Output

One dataset row per unique page URL.

{
"url": "https://www.sitemaps.org/",
"lastmod": "2016-11-21",
"changefreq": null,
"priority": null,
"sourceSitemap": "https://www.sitemaps.org/sitemap.xml",
"domain": "www.sitemaps.org"
}

Pricing

Pay per event. The primary event is apify-default-dataset-item at $0.0004 per URL ($0.40 per 1,000). Apify charges it when a row is saved. There is no monthly fee.

Limits

  • Public http and https sitemaps only. No login, and private hosts (localhost, 10.x, 192.168.x) are refused.
  • At most 50 sitemap files per run, and 20 MB per file after unpacking.
  • At most 5 fetches in flight, with retries on HTTP 429 and 5xx. Crawl-delay in robots.txt is honored up to 5 seconds.
  • A path Disallowed for User-agent: * is not fetched.
  • The run fails if every input produces zero URLs, so an empty success is not reported as a finished dataset.

FAQ

Does this crawl the pages listed in the sitemap?

No. It only downloads sitemap files and returns the URLs inside them.

What if the site publishes a sitemap index or a .gz file?

An index is followed, one child file at a time, until maxUrls is reached. A file that starts with the gzip magic bytes is unpacked. Both were checked against live sites: https://apify.com (index) and https://www.gnu.org (a child sitemap0.xml.gz).

What does a domain input fetch?

/robots.txt (every Sitemap: line), then /sitemap.xml and /sitemap_index.xml if those were not already listed. A start URL that already ends in .xml or .xml.gz is fetched directly.

Why did the run fail with "0 URL(s)"?

None of the inputs had a readable public sitemap. A missing /sitemap_index.xml on its own is normal and is not a failure. The failure means every candidate was missing or was not sitemap XML.