Sitemap URL Extractor: Get All Website URLs
Pricing
from $0.50 / 1,000 results
Sitemap URL Extractor: Get All Website URLs
Get every URL from any website's sitemaps: robots.txt discovery, nested sitemap indexes, gzip, RSS/Atom and TXT sitemaps, hreflang alternates, lastmod filters.
Pricing
from $0.50 / 1,000 results
Rating
0.0
(0)
Developer
Digitální produkty pro život
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
17 hours ago
Last modified
Categories
Share
Sitemap URL Extractor
Get every URL a website publishes in its sitemaps, in seconds, as a clean table you can export to CSV, Excel or JSON, or pipe into your next crawler, SEO audit or AI agent.
Just enter a domain like example.com. The Actor finds the sitemaps for you (from robots.txt and the usual locations), follows nested sitemap indexes, unpacks .xml.gz files and returns one row per unique URL.
What you can use it for
- SEO audits: compare what is in the sitemap with what is indexed, find orphan pages, check
lastmodhygiene and hreflang coverage. - Content monitoring: use the date filter to list only pages added or updated since a given day. Schedule it daily to track a competitor's new products, blog posts or landing pages.
- Crawl planning: feed the exact list of URLs into another scraper instead of crawling the whole site link by link. This is faster and cheaper.
- AI / RAG pipelines: get the canonical list of pages to load into your knowledge base.
- Migrations and QA: export the full URL inventory before and after a site redesign.
Features
- Automatic sitemap discovery:
robots.txtSitemap:lines, then/sitemap.xml,/sitemap_index.xml,/wp-sitemap.xml,/sitemap.txtand more. - Nested sitemap index files, gzip (
.xml.gz), plain-text sitemaps and RSS / Atom feeds. - Returns
lastmod,changefreq,priority, hreflang alternates, image and video counts, and news titles. - Filters: regex include/exclude rules and a
lastmoddate range. - De-duplicates URLs across all sitemaps and websites in the run.
- Many websites in one run, processed in parallel with polite limits.
- Never fails the whole run because of one broken sitemap. Problems are listed in the
SUMMARYrecord. - Lightweight HTTP only, with no browser. That keeps it fast and inexpensive.
Input
| Field | Description |
|---|---|
startUrls | Websites (example.com) or direct sitemap/feed URLs. |
maxUrls | Stop after this many unique URLs (0 = no limit). |
includePatterns / excludePatterns | Regular expressions matched against each URL. |
lastmodFrom / lastmodTo | Keep only URLs modified within this date range. |
keepUrlsWithoutLastmod | Whether URLs without lastmod pass a date filter (default yes). |
includeAlternates | Add hreflang alternates to each row. |
maxSitemaps, concurrency | Safety limits. |
proxyConfiguration | Optional. Only for sites that block cloud servers. |
Example:
{"startUrls": ["https://crawlee.dev", "example-shop.com"],"includePatterns": ["/blog/"],"lastmodFrom": "2026-09-01","maxUrls": 5000}
Output
One row per unique URL:
{"url": "https://crawlee.dev/blog/scrapy-vs-crawlee","lastmod": "2026-09-14","changefreq": "weekly","priority": 0.5,"title": null,"imageCount": 0,"videoCount": 0,"domain": "crawlee.dev","sitemapUrl": "https://crawlee.dev/sitemap.xml","startUrl": "https://crawlee.dev/","alternates": [{ "hreflang": "de", "href": "https://crawlee.dev/de/blog/scrapy-vs-crawlee" }]}
A SUMMARY record in the key-value store reports how many sitemaps were found, processed and failed, how many URLs were filtered or de-duplicated, and why the run stopped.
Pricing
You pay only for the URLs you get. Each saved URL is one result. Use Maximum cost per run or maxUrls to stay within your budget: the Actor stops cleanly when the limit is reached.
Tips
- No results for a site? Some websites simply have no sitemap. The
SUMMARYrecord says so explicitly. - To track only new pages, schedule the Actor daily with
lastmodFromset to yesterday's date. - Huge sites (millions of URLs) work too. Raise
maxSitemapsand setmaxUrlsto control cost.
Responsible use
The Actor reads only sitemaps and feeds, which websites publish specifically for automated discovery. It does not log in, bypass protections or collect personal data. Please respect each website's terms when you reuse the URLs you obtain.
How to use it via API
You can run the Actor from the Apify Console, on a schedule, or from your own code. Get your API token in Apify Console → Settings → Integrations.
Python
from apify_client import ApifyClientclient = ApifyClient("<YOUR_API_TOKEN>")run = client.actor("digitalni.produkty.pro.zivot/sitemap-url-extractor").call(run_input={"startUrls": ["https://crawlee.dev"],"maxUrls": 1000})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
JavaScript / Node.js
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });const run = await client.actor('digitalni.produkty.pro.zivot/sitemap-url-extractor').call({"startUrls": ["https://crawlee.dev"],"maxUrls": 1000});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
Integrations and AI agents
- Export results as JSON, CSV, Excel, XML or HTML, or open them directly in Google Sheets.
- Connect to Make, Zapier, n8n, Slack, Google Drive, Airbyte or any webhook to get all URLs from the sitemaps of a website into your workflow automatically.
- Use it from AI agents and LLM apps (Claude, ChatGPT, Cursor, LangChain…) through the Apify MCP server: the agent can call this Actor as a tool.
- Schedule runs (hourly, daily, weekly) to keep data fresh without any code.
FAQ
How much does it cost? $0.50 per 1,000 URLs. A 500-page website costs about $0.25, a 20,000-page shop about $10. Apify's free plan includes monthly credits, so small sites are effectively free to try.
Do I need to know where the sitemap is? No. Enter just the domain. The Actor reads robots.txt, tries common sitemap locations and follows nested sitemap indexes and gzipped files.
Can I get only new or updated pages? Yes. Use the lastmod date filter and schedule the Actor (e.g. weekly) to get only pages changed since a given date.
Is it legal? Sitemaps are published by website owners precisely so that machines can read them. The Actor only downloads sitemap files, not page content.
More tools from the same developer
All tools with code examples: github.com/Phenixik/apify-actors
- Broken Link Checker: find 404s and dead links on your site
- PageSpeed & Lighthouse Audit: bulk Core Web Vitals and SEO scores
- Website Screenshot: full-page, mobile and PDF screenshots
- Domain WHOIS & DNS Lookup: expiry, DNS, SPF/DMARC and SSL for many domains
- PDF Text Extractor: text, Markdown and RAG chunks from PDFs
- AI Image Upscaler: 4x Real-ESRGAN upscaling, no GPU
- RSS Feed Reader & Monitor: only-new-items feed monitoring
- ATS Jobs: jobs from Greenhouse, Lever, Ashby & more
- App Store Reviews: iOS reviews from 50+ countries
Support
Found a sitemap that is not parsed correctly? Open an issue with the URL and we will fix it, usually within a day or two.