Sitemap URL Extractor: Get All Website URLs avatar

Sitemap URL Extractor: Get All Website URLs

Pricing

from $0.50 / 1,000 results

Go to Apify Store
Sitemap URL Extractor: Get All Website URLs

Sitemap URL Extractor: Get All Website URLs

Get every URL from any website's sitemaps: robots.txt discovery, nested sitemap indexes, gzip, RSS/Atom and TXT sitemaps, hreflang alternates, lastmod filters.

Pricing

from $0.50 / 1,000 results

Rating

0.0

(0)

Developer

Digitální produkty pro život

Digitální produkty pro život

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

17 hours ago

Last modified

Share

Sitemap URL Extractor

Get every URL a website publishes in its sitemaps, in seconds, as a clean table you can export to CSV, Excel or JSON, or pipe into your next crawler, SEO audit or AI agent.

Just enter a domain like example.com. The Actor finds the sitemaps for you (from robots.txt and the usual locations), follows nested sitemap indexes, unpacks .xml.gz files and returns one row per unique URL.

What you can use it for

  • SEO audits: compare what is in the sitemap with what is indexed, find orphan pages, check lastmod hygiene and hreflang coverage.
  • Content monitoring: use the date filter to list only pages added or updated since a given day. Schedule it daily to track a competitor's new products, blog posts or landing pages.
  • Crawl planning: feed the exact list of URLs into another scraper instead of crawling the whole site link by link. This is faster and cheaper.
  • AI / RAG pipelines: get the canonical list of pages to load into your knowledge base.
  • Migrations and QA: export the full URL inventory before and after a site redesign.

Features

  • Automatic sitemap discovery: robots.txt Sitemap: lines, then /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml, /sitemap.txt and more.
  • Nested sitemap index files, gzip (.xml.gz), plain-text sitemaps and RSS / Atom feeds.
  • Returns lastmod, changefreq, priority, hreflang alternates, image and video counts, and news titles.
  • Filters: regex include/exclude rules and a lastmod date range.
  • De-duplicates URLs across all sitemaps and websites in the run.
  • Many websites in one run, processed in parallel with polite limits.
  • Never fails the whole run because of one broken sitemap. Problems are listed in the SUMMARY record.
  • Lightweight HTTP only, with no browser. That keeps it fast and inexpensive.

Input

FieldDescription
startUrlsWebsites (example.com) or direct sitemap/feed URLs.
maxUrlsStop after this many unique URLs (0 = no limit).
includePatterns / excludePatternsRegular expressions matched against each URL.
lastmodFrom / lastmodToKeep only URLs modified within this date range.
keepUrlsWithoutLastmodWhether URLs without lastmod pass a date filter (default yes).
includeAlternatesAdd hreflang alternates to each row.
maxSitemaps, concurrencySafety limits.
proxyConfigurationOptional. Only for sites that block cloud servers.

Example:

{
"startUrls": ["https://crawlee.dev", "example-shop.com"],
"includePatterns": ["/blog/"],
"lastmodFrom": "2026-09-01",
"maxUrls": 5000
}

Output

One row per unique URL:

{
"url": "https://crawlee.dev/blog/scrapy-vs-crawlee",
"lastmod": "2026-09-14",
"changefreq": "weekly",
"priority": 0.5,
"title": null,
"imageCount": 0,
"videoCount": 0,
"domain": "crawlee.dev",
"sitemapUrl": "https://crawlee.dev/sitemap.xml",
"startUrl": "https://crawlee.dev/",
"alternates": [{ "hreflang": "de", "href": "https://crawlee.dev/de/blog/scrapy-vs-crawlee" }]
}

A SUMMARY record in the key-value store reports how many sitemaps were found, processed and failed, how many URLs were filtered or de-duplicated, and why the run stopped.

Pricing

You pay only for the URLs you get. Each saved URL is one result. Use Maximum cost per run or maxUrls to stay within your budget: the Actor stops cleanly when the limit is reached.

Tips

  • No results for a site? Some websites simply have no sitemap. The SUMMARY record says so explicitly.
  • To track only new pages, schedule the Actor daily with lastmodFrom set to yesterday's date.
  • Huge sites (millions of URLs) work too. Raise maxSitemaps and set maxUrls to control cost.

Responsible use

The Actor reads only sitemaps and feeds, which websites publish specifically for automated discovery. It does not log in, bypass protections or collect personal data. Please respect each website's terms when you reuse the URLs you obtain.

How to use it via API

You can run the Actor from the Apify Console, on a schedule, or from your own code. Get your API token in Apify Console → Settings → Integrations.

Python

from apify_client import ApifyClient
client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("digitalni.produkty.pro.zivot/sitemap-url-extractor").call(run_input={
"startUrls": [
"https://crawlee.dev"
],
"maxUrls": 1000
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)

JavaScript / Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });
const run = await client.actor('digitalni.produkty.pro.zivot/sitemap-url-extractor').call({
"startUrls": [
"https://crawlee.dev"
],
"maxUrls": 1000
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

Integrations and AI agents

  • Export results as JSON, CSV, Excel, XML or HTML, or open them directly in Google Sheets.
  • Connect to Make, Zapier, n8n, Slack, Google Drive, Airbyte or any webhook to get all URLs from the sitemaps of a website into your workflow automatically.
  • Use it from AI agents and LLM apps (Claude, ChatGPT, Cursor, LangChain…) through the Apify MCP server: the agent can call this Actor as a tool.
  • Schedule runs (hourly, daily, weekly) to keep data fresh without any code.

FAQ

How much does it cost? $0.50 per 1,000 URLs. A 500-page website costs about $0.25, a 20,000-page shop about $10. Apify's free plan includes monthly credits, so small sites are effectively free to try.

Do I need to know where the sitemap is? No. Enter just the domain. The Actor reads robots.txt, tries common sitemap locations and follows nested sitemap indexes and gzipped files.

Can I get only new or updated pages? Yes. Use the lastmod date filter and schedule the Actor (e.g. weekly) to get only pages changed since a given date.

Is it legal? Sitemaps are published by website owners precisely so that machines can read them. The Actor only downloads sitemap files, not page content.

More tools from the same developer

All tools with code examples: github.com/Phenixik/apify-actors

Support

Found a sitemap that is not parsed correctly? Open an issue with the URL and we will fix it, usually within a day or two.