Sitemap URL Extractor — All Page URLs from a Website avatar

Sitemap URL Extractor — All Page URLs from a Website

Pricing

from $0.20 / 1,000 urls

Go to Apify Store
Sitemap URL Extractor — All Page URLs from a Website

Sitemap URL Extractor — All Page URLs from a Website

Input: domains or sitemap URLs. Output: one row per page URL with its lastmod, changefreq and priority, plus a report row per sitemap file read. Finds sitemaps via robots.txt, follows nested indexes, unpacks gzip. Filter by URL pattern or by date. Hard cap of 25,000 URLs per run.

Pricing

from $0.20 / 1,000 urls

Rating

0.0

(0)

Developer

Power On Labs

Power On Labs

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

17 days ago

Last modified

Share

Sitemap URL Extractor — every page URL from a website's sitemaps

Give it a domain, a homepage, or a direct link to a sitemap.xml, and get back the full list of the site's page URLs — one dataset row per URL, with the lastmod, changefreq and priority the sitemap declares.

It is built to finish the job on sites where a naive sitemap reader gives up:

  • Finds the sitemap for you. Reads robots.txt, then tries the common locations (/sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml, …).
  • Follows sitemap-index files, to any depth. A site that splits its sitemap into 40 partitions behind two levels of index still comes back whole.
  • Unpacks gzip. .xml.gz sitemaps, and sitemaps served with a gzip body, are decompressed automatically.
  • Never fails the whole run for one bad file. Every sitemap gets a row in the report saying whether it was read, and why not if it wasn't.
  • No browser. It is HTTP plus XML parsing, so it is fast and the run is cheap.

Output

With one row per URL (the default), each row is:

{
"url": "https://example.com/blog/how-sitemaps-work",
"lastmod": "2026-08-14",
"changefreq": "monthly",
"priority": "0.7",
"sourceSitemap": "https://example.com/sitemap-posts.xml",
"website": "https://example.com"
}

That is exactly the shape a crawler or a screenshot tool wants as input, so the dataset drops straight into the next step — for example Website Screenshot & PDF Generator or any crawler that takes a URL list.

Turn "One row per URL" off to get instead one row per website, with the URLs as an array and a per-sitemap breakdown.

The sitemap report

Unless you turn it off, the dataset ends with one row per sitemap file found:

websitesitemapUrldiscoveredViastatusurlCountchildCountnote
https://example.com/sitemap.xmlrobots.txtok (index)01212 new child sitemap(s)
https://example.com/sitemap-posts.xmlindex: /sitemap.xmlok1 843
https://example.com/sitemap-authors.xmlindex: /sitemap.xmlerror0HTTP 404

An index file holds no page URLs, so its count goes in childCount: urlCount only ever counts pages.

This is the row that tells you why a count is lower than you expected — a sitemap the index points to that 404s, a file that isn't XML, a timeout.

Options

OptionWhat it does
Websites or sitemap URLsOne per line. A bare domain, a homepage, or a direct sitemap link.
Max URLs per websiteStop after N URLs for one site (0 = up to the run limit).
Include only URLs matching / Exclude URLs matchingJavaScript regular expressions applied to each URL.
Only URLs changed on or afterA date. Keeps only URLs whose <lastmod> is on or after it — the rest, and URLs with no <lastmod>, are dropped.
One row per URLOn: a flat URL list. Off: one row per site with an array and a per-sitemap breakdown.
Report onlyReturn the sitemap report and the page count, no URLs — and nothing is charged.
Max sitemap files per websiteSafety limit on how many sitemap files to fetch for one site.
ProxyOptional, for sites that block datacenter traffic or serve a region-specific sitemap.

Pricing

Pay per URL returned. A website with no readable sitemap costs nothing, and a Report only run costs nothing either — it returns counts, not URLs.

A run returns at most 25,000 URLs, so a single run can never surprise you. When a run stops at that limit — or at the maximum cost you allow for one run — it finishes normally and adds a report row saying so, with the sitemap files still left to read: point a new run straight at those, or narrow the result with Include only URLs matching.

Common uses

  • Get all URLs from a website to seed a crawl, an archive, or a migration checklist.
  • Convert a sitemap.xml to a plain URL list (JSON or CSV via the dataset export).
  • Find the pages that changed since your last crawl, using the lastmod filter, and re-process only those.
  • Count the pages of a website, or audit which sitemaps it publishes, with Report only: one row per sitemap file plus the total, no URLs returned and nothing charged.
  • Feed a screenshot or PDF Actor one row per page.
  • Feed a tech stack audit. The URL list this Actor returns is already the input format of Tech Stack Audit, which opens each page in a real browser and reports its CMS, framework, CDN, analytics and tracking pixels with the evidence for every detection.

Limits

  • It reads XML sitemaps (<urlset> and <sitemapindex>). It does not parse HTML sitemap pages or discover URLs by crawling links — give it the sitemap.
  • lastmod is reported as the sitemap writes it; it is not independently verified against the page.
  • Sites that publish no sitemap return an empty result with a report row saying so. Some large sites (for example python.org) genuinely have none at a standard location.

Issues

Found a sitemap it can't read? Open an issue on the Actor's Issues tab with the URL and it will be looked at.