Sitemap URLs: Exact File, No Page Crawl avatar

Sitemap URLs: Exact File, No Page Crawl

Pricing

$0.50 / 1,000 sitemap urls

Go to Apify Store
Sitemap URLs: Exact File, No Page Crawl

Sitemap URLs: Exact File, No Page Crawl

Extract the exact sitemap file you name, or discover a domain's sitemaps. Deduplicated URL, lastmod and source rows with explicit partial/error reports. No page crawl or proxy; empty runs are free.

Pricing

$0.50 / 1,000 sitemap urls

Rating

0.0

(0)

Developer

James Power

James Power

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Sitemap URLs from the exact file you choose

Give the Actor a specific sitemap URL and it reads that file, not a different sitemap it finds at the domain root. Or give it a bare domain to discover sitemaps. The dataset contains one row per distinct URL, with lastmod and the sitemap file that supplied it.

Only sitemap files are requested — robots.txt, sitemap.xml, .xml.gz files and sitemap indexes. A page URL is never fetched. This avoids page rendering and keeps requests light, but a site's WAF can still refuse a sitemap file; the output names the failed file and status.

Input

FieldWhat it does
Domains or sitemap URLsOne per line. https://example.com is discovered through its robots.txt, falling back to /sitemap.xml and /sitemap_index.xml. https://example.com/sitemap.xml is read directly. https:// is added if you leave it out. Up to 32 lines.
Maximum URLsDefault 2,000 (maximum 200,000). The run stops there and reports partial.
Maximum sitemap files per domainDefault 50 (maximum 1,000), for sites that split their URLs over many files.

That is the whole input — there is nothing to configure about parsing, because the sitemap protocol decides it.

Output

The dataset has exactly three fields per row:

{ "url": "https://example.com/pricing", "lastmod": "2026-08-14", "source": "https://example.com/sitemap.xml" }
  • url — exactly as your sitemap declared it. It was never requested.
  • lastmod — the sitemap's <lastmod> for that URL, normalised to ISO 8601 (2026-08-14 or 2026-08-14T09:30:00+00:00), or null when it was absent or not a date. When the same URL appears in more than one file, the newest lastmod wins.
  • source — the sitemap file that first listed the URL (after any redirects), so you always know where a row came from.

The run's output record (OUTPUT in the key-value store) is the report:

FieldMeaning
statusok, partial or error — see below.
urls, datasetRowsUnique URLs found, and rows actually stored (they match unless the run's cost limit was reached).
charge.eventsBilled events, one per stored row.
totalsduplicates skipped, URLs with a lastmod, invalid or future lastmod values, <loc> values that are not valid URLs, off-site URLs, files read.
inputs[]Per line: how the sitemap was found, every file that was tried with its HTTP status, bytes, whether it was gzipped, what was skipped and why, limits hit.
limitsHit, warningsWhich cap stopped the run, and any input line that was ignored and why.

status is never optimistic. ok means every sitemap you asked for was read and nothing was skipped. partial means you got URLs but something was cut short — a file that failed, a cap, a cross-site sitemap that was not fetched, or <loc> values that are not valid URLs (up to two are quoted back in the report with the reason). error means the dataset is empty. A partial or error run is honest about the reason instead of quietly returning a short list.

Reading sitemaps well

  • Discovery: robots.txt Sitemap: lines first (RFC 9309), then the conventional /sitemap.xml and /sitemap_index.xml. If robots.txt names a dead sitemap, the conventional paths are still tried.
  • Indexes: nested sitemap indexes are followed (up to three levels), with every child URL validated before it is fetched.
  • Compression: .xml.gz files are expanded under a byte budget, and gzip content-encoding is handled by the HTTP client. Compression bombs are refused, not decompressed into memory.
  • Deduplication: one row per unique URL, across every file and every line of input. A duplicate is never stored twice, so it is never charged twice.
  • Hostile XML: entity declarations and external entity references are refused; a plain DOCTYPE is ignored without fetching it. A broken file leaves the rest of the run alone.
  • Honest about bad data: a <loc> that is not a valid URL — whitespace in the middle, a mailto: entry, a URL longer than 2048 characters, credentials in the URL — is left out and reported with an example, rather than written to the dataset as if it were usable. Real sitemaps do contain such entries; a well-behaved reader says so instead of silently emitting them.

Safety by construction

  • Only globally routable addresses are dialled, and that is enforced by the resolver that opens the socket (not by a check that happens before it), so a DNS answer that changes between the check and the connection cannot redirect the request anywhere private.
  • http:// and https:// only; no credentials in URLs; known non-HTTP ports are refused; private, loopback, link-local and cloud-metadata addresses are refused.
  • Every redirect hop is validated before it is followed; the client itself is not allowed to follow redirects on its own.
  • Child sitemaps must be on the same site as the sitemap that lists them. A sitemap hosted on a different domain is not fetched from a robots.txt or index declaration — add it as its own input line if you want it read, and it will then be read as a first-class target.
  • URLs inside a sitemap that point at another domain are returned as data and counted in the output record; they are never requested.
  • No proxy, browser, JavaScript, API keys, login or cookies. The Actor never crawls profiles or extracts contact fields, but sitemap URLs can contain profile paths or personal details in query strings. Review the returned URLs before sharing them.

Billing

US$0.0005 per URL (US$0.50 per 1,000), one event per stored dataset row. No start fee. A run that finds nothing, or fails, stores nothing and is charged nothing. If your cost limit for the run is reached, the crawl stops early, the output record says charge.limitReached: true and the status is partial.

Limits and what is not included

  • No full-page crawling and no JavaScript rendering — this Actor reads sitemaps, nothing else.
  • Image, video and news extension data inside a sitemap is ignored; each entry's <loc> and <lastmod> are read.
  • A run reads at most 50 MB per file, 256 MB in total, and by default 50 files per domain (input maximum 1,000). The default time limit is 15 minutes; the output reports when a bound is hit.

Local development

python -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/pip install pytest
.venv/bin/python -m pytest # offline tests against a local synthetic fixture server

The tests never touch the public internet, an Apify account or anyone's real website; they cover discovery, gzip, indexes, hostile XML, byte and count caps, cancellation, redirects, private-address and DNS-rebinding refusal, and the billing rules.