Sitemap URL Extractor — Robots.txt & Sitemap.xml Parser avatar

Sitemap URL Extractor — Robots.txt & Sitemap.xml Parser

Pricing

Pay per event

Go to Apify Store
Sitemap URL Extractor — Robots.txt & Sitemap.xml Parser

Sitemap URL Extractor — Robots.txt & Sitemap.xml Parser

Extracts the URLs a website lists in its XML sitemaps (sitemap indexes, gzip and text sitemaps, robots.txt discovery), up to your Max URLs cap, with filters and a clear per-site report.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Tarnlight

Tarnlight

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Sitemap URL Extractor

Reliably pulls every URL a site's sitemaps list (up to your maxUrls cap) out of its XML sitemaps -- and reports, per site, exactly what it found and what it couldn't reach -- instead of quietly returning a partial list. Honours robots.txt. Built by Tarnlight.

What it does

Give it a homepage/site URL or a direct sitemap URL -- the two behave differently, on purpose (see "How discovery works" below):

  • A homepage/site URL (e.g. https://example.com) triggers full discovery: robots.txt is read for Sitemap: lines, and well-known common paths are probed too.
  • A direct sitemap URL (e.g. https://example.com/sitemap.xml, or anything else that looks like a sitemap file/path) is fetched by itself -- only that file and, if it's a sitemap index, its child sitemaps. There is no site-wide discovery and nothing else is guessed, so you're only ever charged for the file you actually asked for.

For every site, it will:

  1. Look up /robots.txt for Disallow: rules -- every request this actor makes (a common-path guess, a declared sitemap, a child of a sitemap index, or a start URL you gave it directly) is checked against those rules first (RFC 9309; see "How discovery works" below). This always happens, for both kinds of start URL.
  2. For a homepage/site URL only: also read robots.txt's Sitemap: lines, and probe common sitemap paths (/sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /wp-sitemap.xml, /sitemap.xml.gz, /sitemap.txt) -- both happen regardless of what the other finds, skipping any robots.txt disallows.
  3. Recursively follow sitemap indexes (gzip and plain, cycle-safe, depth capped) down to the individual page sitemaps.
  4. Parse XML sitemaps (namespaced or not) and gzip-compressed sitemaps. Parse a plain-text sitemap (one absolute URL per line) only when the response actually looks like one -- see "Plain-text sitemaps" below; anything else (an HTML page, humans.txt, arbitrary text) is rejected as "not-a-sitemap" and never billed.
  5. Push one dataset row per URL, deduplicated within the run, with lastmod, changefreq, priority, hreflang alternates and image counts when present.
  6. Record a SUMMARY key-value-store entry with per-site sitemap counts, failures (with reasons), how many URLs matched your filters vs. were actually written vs. skipped once a cap was hit, every discovery attempt made (found it, 404, blocked, malformed, ...), and a plain-English note explaining anything worth flagging -- especially when a site yields 0 URLs.

If a page returns an HTML "error" page with an HTTP 200 status where a sitemap was expected, that's detected and reported as a failure for that sitemap instead of being mis-parsed as XML. A single unreachable or broken sitemap never fails the whole run -- it's logged in SUMMARY and the run carries on with everything else.

Plain-text sitemaps

A non-XML 200 response is only treated as a plain-text sitemap when it was actually served as one: the URL ends in .txt/.txt.gz, or the response's Content-Type is text/plain. Even then, only lines that parse as a bare absolute http(s):// URL (no embedded whitespace) on the sitemap file's own site are kept -- per the sitemaps.org rule that a sitemap may only list URLs on its own host; cross-host submission requires a search-console verification this actor has no way to check, so it's never assumed. If no line qualifies, the whole response is rejected as "not-a-sitemap" rather than billing you for every line of, say, a site's humans.txt.

Who it's for

  • SEO audits -- get the full, current list of URLs a site is telling search engines about.
  • Content inventories -- see every page a site publishes without crawling the whole thing link by link.
  • Migration checks -- diff sitemap URL lists before/after a replatform.
  • Feeding crawlers or RAG pipelines -- a clean, deduplicated seed list of URLs to fetch next, instead of hand-rolling sitemap parsing.

Input

FieldTypeDefaultNotes
startUrlsarray(required)Homepages/site URLs (full discovery) or direct sitemap URLs (fetched by themselves only -- see above).
maxUrlsinteger5000Stop after this many URLs are written (0 = unlimited -- can run up an unbounded bill on a large site). Each URL costs $0.001, so the default caps a run at about $5.
includePatternsarray of regex[]Keep only URLs matching at least one.
excludePatternsarray of regex[]Drop URLs matching any. An invalid regex here (or in includePatterns) ends the run cleanly with an "invalid input" status message instead of crashing mid-run.
modifiedSincedate(none)Keep URLs with lastmod on/after this date.
keepUrlsWithoutLastmodbooleanfalseOnly matters when modifiedSince is set. By default, URLs with no usable lastmod (missing, or a placeholder such as WordPress/Jetpack's "0000-00-00") are dropped and not charged, so a date-filtered run only returns (and bills) URLs known to match your date. Set to true to keep them as well. SUMMARY reports urlsWithUnparseableLastmod per site either way.
discoverFromRobotsbooleantrueHomepage/site URLs only: use robots.txt's Sitemap: lines for discovery. robots.txt is always fetched and its Disallow: rules always honoured regardless of this setting -- it only controls whether its declared sitemaps are added to the discovery list. Has no effect on a direct sitemap URL.
tryCommonPathsbooleantrueHomepage/site URLs only: probe well-known sitemap paths, regardless of what robots.txt declares. Has no effect on a direct sitemap URL.
includeAlternatesbooleanfalseInclude hreflang alternates in output rows.
includeImagesbooleanfalseInclude imageUrls in output rows (a count is always included).
maxConcurrencyinteger5Max simultaneous requests, across all sites. A single site's own crawl is further limited to 2 concurrent requests per host, so this mostly matters when you pass several startUrls at once.
requestTimeoutSecsinteger30Per-request timeout.

Example input:

{
"startUrls": [
{ "url": "https://example.com" },
{ "url": "https://example.com/custom-sitemap.xml" }
],
"maxUrls": 5000,
"excludePatterns": ["/tag/", "/page/[0-9]+/"],
"modifiedSince": "2025-01-01",
"includeImages": true
}

Output

One dataset row per URL. Real rows from a test run (docs.apify.com and gov.uk), showing typical variation in which fields a site fills in:

{
"url": "https://docs.apify.com/api",
"lastmod": null,
"changefreq": "weekly",
"priority": 0.5,
"sitemapUrl": "https://docs.apify.com/sitemap_base.xml",
"site": "https://docs.apify.com",
"imageCount": 0
}
{
"url": "https://www.gov.uk/government/news/whole-system-approach-to-tackling-violent-crime-is-working",
"lastmod": "2026-07-23T19:37:26+01:00",
"changefreq": null,
"priority": 0.29,
"sitemapUrl": "https://www.gov.uk/sitemaps/sitemap_1.xml",
"site": "https://www.gov.uk",
"imageCount": 0
}

alternates (list of {hreflang, href}) and imageUrls are only added when includeAlternates / includeImages are enabled.

The SUMMARY key-value-store record looks like:

{
"sites": [
{
"site": "https://example.com",
"sitemapsFound": 3,
"sitemapUrls": ["https://example.com/sitemap_index.xml", "..."],
"sitemapsFailed": [],
"urlsMatched": 812,
"urlsWritten": 812,
"urlsSkippedAfterCap": 0,
"urlsFilteredOut": 4,
"urlsWithUnparseableLastmod": 0,
"stoppedEarly": false,
"discoveryAttempts": [
{ "url": "https://example.com/robots.txt", "source": "robots.txt", "httpStatus": 200,
"outcome": "not-a-sitemap", "detail": "found 1 Sitemap: line(s)" },
{ "url": "https://example.com/sitemap_index.xml", "source": "robots.txt", "httpStatus": 200,
"outcome": "sitemap", "detail": "sitemap index with 2 child sitemap(s)" },
{ "url": "https://example.com/wp-sitemap.xml", "source": "common-path", "httpStatus": 404,
"outcome": "not-found", "detail": "HTTP 404" }
],
"note": null
}
],
"totalUrlsWritten": 812,
"stopReason": null
}

urlsMatched is how many of that site's sitemap entries passed your includePatterns/excludePatterns/modifiedSince filters. urlsWritten is how many of those were actually written to the dataset (and charged) -- always the honest number; it can be lower than urlsMatched only when urlsSkippedAfterCap is non-zero, meaning maxUrls or the run's own charge limit was reached before every matched URL could be written. totalUrlsWritten (top level) is the sum of every site's urlsWritten, and always equals the number of url-extracted charged events for the run.

Every URL the actor tried in order to find a sitemap for a site -- whether it worked or not -- is listed in that site's discoveryAttempts: source says how the URL was arrived at ("start-url", "robots.txt" or "common-path"), outcome says what happened ("sitemap" / "not-found" / "blocked" / "error" / "not-a-sitemap" / "disallowed-by-robots" -- robots.txt forbade fetching this URL, so it wasn't fetched /

"robots- unreachable"
-- robots.txt itself returned a 5xx or a network error after retries, so the whole site was treated as disallowed), and detail is a short human-readable reason. note is null when there is nothing to explain, and a plain-English sentence otherwise, e.g.:

{ "site": "https://www.python.org", "sitemapsFound": 0, "urlsMatched": 0, "note":
"No sitemap published: robots.txt has no Sitemap: line and none of 6 common paths exist." }
{ "site": "https://blocked-example.com", "sitemapsFound": 0, "urlsMatched": 0, "note":
"Access blocked (HTTP 403) when fetching robots.txt and/or sitemap paths -- the site may block automated requests." }

When any site ends with 0 matched URLs, the actor also logs a WARNING with that site's note, and the final run status message adds a short, correctly-categorized hint -- a blocked/unreachable site is never reported the same way as one with no sitemap (e.g. "1 site was blocked or unreachable; 1 site had no sitemap -- see SUMMARY") -- so it's never silent about it, and a bot-blocking site is never confused with a genuinely sitemap-less one.

How discovery works

For each start URL's origin, robots.txt is fetched first, always -- before anything else is requested on that host. It's parsed for TarnlightSitemapBot's rules, falling back to * when there's no bot-specific group (per RFC 9309's group-selection rule), and the result gates every later request to that origin:

  • robots.txt itself returns 4xx (e.g. 404): no restrictions apply (RFC 9309 S2.3.1.3) -- a missing robots.txt is normal, and discovery proceeds exactly as if it didn't exist.
  • robots.txt returns 5xx, or is unreachable over the network, after retries: RFC 9309 S2.3.1.4 requires assuming complete disallow -- this actor treats the entire site as off-limits and makes no further request to it at all (not even a common-path guess), and says so clearly in that site's note.
  • robots.txt fetches fine (2xx): its Sitemap: lines are read, and every URL this actor might fetch on that origin -- a Sitemap:-declared sitemap, a common-path guess, a child of a sitemap index, or the start URL you gave it directly (RFC 9309 applies to all of these alike) -- is checked against its Disallow: rules before being requested. A disallowed URL is never fetched; it's recorded in discoveryAttempts with outcome "disallowed-by-robots" and simply skipped.

For each start URL, if the URL itself looks like a sitemap file (ends in .xml, .xml.gz, .txt, or has "sitemap" in the path), it's fetched directly (once robots.txt allows it) -- and only that URL and its sitemap-index children are ever fetched for that start URL. robots.txt Sitemap: discovery and common-path probing do not run for it, regardless of discoverFromRobots/tryCommonPaths, so a direct sitemap URL never silently expands into (and bills for) the rest of the site. Otherwise the actor treats the start URL as a site root: it reads robots.txt for Sitemap: lines and probes the common paths above -- both happen unconditionally (subject to discoverFromRobots/tryCommonPaths), not only when the other turns up nothing, since some sites keep a stale legacy sitemap at a common path alongside a different one declared in robots.txt. Every sitemap found is fetched with retries (exponential backoff on 429 and 5xx responses, honouring Retry-After), decompressed if gzipped, and parsed. A sitemap index's children are followed recursively (each one checked against robots.txt first; depth capped at 5, and already-visited URLs are skipped, so a site that accidentally -- or deliberately -- references a sitemap that references itself can't cause a loop).

Every URL this process tries -- robots.txt itself, each Sitemap: line it names, and each common-path guess -- is recorded as a discoveryAttempts entry in SUMMARY, whether it succeeded or not. That's what makes a 0-URL result explainable instead of silent: 401/403, and 429 that's still failing after retries, are recorded as "blocked" (a bot-blocking site is a common real-world cause of an empty result, not just a missing sitemap); a 404 on a mere guess is "not-found" and never treated as a failure; a 200 response that isn't valid sitemap XML/text (an HTML error page, malformed XML, an empty body) is "not-a-sitemap"; a URL robots.txt disallows is "disallowed-by-robots"; robots.txt itself being unreachable is "robots-unreachable". A per-site note in plain English summarizes whichever of these applies.

Limits

  • This actor only reports what a site's sitemaps actually publish. It does not crawl pages or discover URLs that aren't listed in any sitemap.
  • A site with no robots.txt Sitemap: line, no sitemap at a common path, and no sitemap URL given directly will correctly report zero URLs for that site -- that's expected, not a bug (see the FAQ).
  • Per-file size is capped at 100 MB decompressed as a safety limit against runaway or hostile files -- for both gzip and plain (uncompressed) files.
  • robots.txt is checked for the start URL's own origin. A sitemap index that points at a different host (e.g. a www. vs. bare-domain split, or a separate sitemap subdomain) has its children checked against that same start-origin policy, not a fresh fetch of the other host's own robots.txt; a redirect target is likewise not re-checked. This covers the overwhelmingly common case (a site's own sitemap on its own host) but isn't a substitute for reviewing a multi-host site's own policies.
  • maxConcurrency bounds total in-flight requests across all startUrls, but a single site's own crawl is additionally capped at 2 concurrent requests per host -- so raising it mostly helps when you pass many start URLs at once, not a single large site.

FAQ

A site returned 0 URLs -- is that broken? Not necessarily -- check that site's note and discoveryAttempts in SUMMARY (also logged as a WARNING, and hinted at in the final run status). Two common, non-broken causes:

  • No sitemap published. Nothing links to a sitemap from robots.txt, none of the common paths exist, and you didn't pass a direct sitemap URL -- there is genuinely nothing for the actor to find. Pass the sitemap URL directly if you know it.
  • The site blocked the request. Some sites return HTTP 401/403 (or keep returning 429 after retries) to automated requests, including to robots.txt itself. SUMMARY reports this as "outcome": "blocked" with a note like "Access blocked (HTTP 403) ..." -- that's a site policy, not a bug in the actor, and it's reported distinctly from "no sitemap" so the two aren't confused.

Does it crawl the site itself? No. It only follows sitemap files (and sitemap indexes) -- it never requests ordinary pages on purpose. If a sitemap path redirects to an ordinary HTML page, or a site serves HTML at a URL a sitemap named, that's detected and discarded rather than parsed as XML.

What happens with a huge site? Use maxUrls to cap the run, and includePatterns/excludePatterns to narrow scope. Sitemap index recursion and per-file size are both bounded automatically.

Why do some rows have empty lastmod/changefreq/priority? Those fields are optional in the sitemap spec; the actor reports exactly what the site published and leaves the rest null rather than guessing.

Pricing

Pay per event, in USD: $0.005 per run start (scales with the run's memory, in GB) + $0.001 per URL written to the dataset (e.g. 1,000 URLs = $1.005). The default maxUrls of 5,000 caps a run at about $5.005; raise it (or set it to 0 for no cap) only once you know roughly how many URLs you expect, since a large uncapped site can run up a large bill. Any tax Apify applies at checkout is extra; the Pricing tab is authoritative.

Acceptable use and data protection

Use responsibly: only run this on sites you're entitled to access, and follow each site's terms. The actor identifies itself as TarnlightSitemapBot and honours robots.txt (including treating a robots.txt server error as a full disallow -- see "How discovery works"), so site owners can block it there.

Output can include personal data where a site's URLs contain it (e.g. author pages). You decide how output is used and need a lawful basis. Tarnlight does not access your runs unless you share run data with developers.