Sitemap URL Extractor - All URLs from sitemap.xml & robots.txt
Pricing
from $0.40 / 1,000 url extracteds
Sitemap URL Extractor - All URLs from sitemap.xml & robots.txt
Get all URLs of a website from its XML sitemap: finds sitemaps via robots.txt and common paths, walks sitemap indexes, reads .gz files and extracts every page URL with lastmod, priority, images, news and hreflang. HTTP-only. $0.40 per 1,000 URLs; duplicates and failed sites are free.
Pricing
from $0.40 / 1,000 url extracteds
Rating
0.0
(0)
Developer
Tenfold Fleet
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
What does Sitemap URL Extractor do?
Sitemap URL Extractor gets every URL from a website's XML sitemaps. Give it a domain, a homepage or a direct sitemap link and it returns one row per page with its lastmod date, change frequency, priority, images, Google News data and hreflang alternates.
It finds sitemaps the way search engines do: it reads the Sitemap: lines in robots.txt, then tries /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml and /sitemap-index.xml. It walks sitemap index files recursively (up to 6 levels), reads gzipped .xml.gz sitemaps, plain-text sitemaps and RSS/Atom feeds, and removes duplicate URLs.
It's a fast, low-cost sitemap scraper API: plain HTTP, no browser, $0.40 per 1,000 URLs. Sites without a sitemap, broken sitemaps and duplicates are free.
What you get that basic sitemap scrapers don't: a lastmod date filter that skips stale child sitemaps entirely, Google News, image and hreflang data, free error rows that tell you why a site returned nothing, and deduplication across inputs (www and non-www, domain plus sitemap URL).
Why use a sitemap extractor?
- ๐ SEO audits: list every indexable page of a site, spot stale
lastmoddates and missing hreflang alternates. - ๐ท๏ธ Crawl seeding: feed a complete URL list into a scraper (Website Content Crawler, Cheerio Scraper) instead of crawling links.
- ๐ Content monitoring: with Only URLs modified after, get just the pages a competitor or publisher added or changed since a date.
- ๐ Ecommerce research: pull all product or category URLs of a Shopify, WooCommerce or Magento store with an include pattern like
/products/. - ๐ฐ News tracking: read Google News sitemaps with headline and publication time.
- ๐ Migrations and redirects: export the old site's URLs before a relaunch.
Runs on the Apify platform: you get an API, scheduling, webhooks and integrations with Make, Zapier, Google Sheets and more.
What data can it extract?
| Field | Description |
|---|---|
url | Page URL (<loc>) |
lastmod | Last modification date |
changefreq | Change frequency hint |
priority | Priority 0.0-1.0 |
images | Image sitemap entries: loc, title |
news | Google News: title, publicationDate |
alternates | hreflang alternates: hreflang, href |
sourceSitemap | The sitemap file the URL came from |
site | Input website |
depth | Sitemap index depth (0 = top-level sitemap) |
How to extract all URLs from a sitemap
- Open Sitemap URL Extractor and click Try for free.
- Add websites (
example.com) or sitemap URLs (https://example.com/sitemap_index.xml). - Optionally set Max URLs per site, include/exclude regex patterns or a lastmod date.
- Click Start and download the URLs as JSON, CSV, Excel or HTML, or fetch them through the API.
How much does it cost to extract sitemap URLs?
Pay per result: $0.40 per 1,000 unique URLs ($0.0004 per URL). You are not charged for duplicate URLs, sitemap index entries, websites without a sitemap, 404 or timeout errors, or malformed XML. Set a maximum cost per run and the actor stops cleanly when it is reached. Platform usage is small because it uses plain HTTP requests.
Input
{"startUrls": ["apify.com", "https://wordpress.org/news/sitemap.xml"],"maxUrlsPerSite": 10000,"includePatterns": ["/blog/"],"excludePatterns": ["/tag/"],"lastmodAfter": "2026-01-01","includeExtensions": true}
Output
{"url": "https://www.nytimes.com/interactive/2026/09/29/us/quake-tracker-washington-seattle.html","lastmod": "2026-10-01T05:03:58Z","changefreq": null,"priority": null,"images": [{ "loc": "https://static01.nyt.com/images/2026/09/29/...-articleLarge-v80.jpg", "title": null }],"news": { "title": "Map: 4.2-Magnitude Earthquake Strikes Near Seattle", "publicationDate": "2026-09-29T21:55:55Z" },"alternates": [],"sourceSitemap": "https://www.nytimes.com/sitemaps/new/news.xml.gz","site": "nytimes.com","depth": 0}
Sites that fail produce a free row such as { "site": "example.com", "error": "No sitemap found: ...", "status": 404 }, so you always know why a site returned nothing. Every row has the same keys: URL rows have error and status set to null, error rows have url set to null, so CSV and spreadsheet exports keep fixed columns. Download results as JSON, CSV, Excel, XML or HTML.
Tips
- Huge sites (news publishers, marketplaces) can list millions of URLs: use Max URLs per site and patterns to keep runs small.
- With Only URLs modified after, child sitemaps whose own
lastmodis older are skipped entirely, which makes monitoring runs fast and cheap. - With strict filters on a huge site, the actor stops a site after about 100 + maxUrlsPerSite / 50 sitemap files that contain no matching URL and returns a free note row. Start from a specific sitemap URL (for example the blog or news sitemap) to go deeper.
- Limits per site: up to 1,000,000 URLs, 30 minutes, 2 GB of sitemap downloads and 60 MB per sitemap file. A site that times out or fails 5 times in a row is stopped. Each limit that cuts a site short adds a free note row.
- A large run (several sites with hundreds of thousands of URLs each) needs more memory: give it 2 GB or more. If deduplication fills the memory, the actor stops the site with a free note row instead of crashing.
- Gzip is detected from the file content, so
.xml.gzsitemaps work whether or not the server already decompressed them. - If a site blocks datacenter IPs, switch the proxy to a residential group.
FAQ
Does it crawl the website pages?
No. It reads only robots.txt and sitemap files, which is why it is fast and cheap. Use a crawler if the site has no sitemap.
What if robots.txt lists no sitemap?
It tries the common sitemap paths. If none exists you get a free error row.
Can I give it a single sitemap file?
Yes. Any URL ending in .xml, .xml.gz, .txt or containing "sitemap" is read directly, including sitemap index files.
Is it legal to extract sitemaps?
Sitemaps are public files that sites publish for crawlers. This actor only reads public data. Make sure your use of the URLs complies with the target site's terms.