# Sitemap URL Extractor: Get All Website URLs (`digitalni.produkty.pro.zivot/sitemap-url-extractor`) Actor

Get every URL from any website's sitemaps: robots.txt discovery, nested sitemap indexes, gzip, RSS/Atom and TXT sitemaps, hreflang alternates, lastmod filters.

- **URL**: https://apify.com/digitalni.produkty.pro.zivot/sitemap-url-extractor.md
- **Developed by:** [Digitální produkty pro život](https://apify.com/digitalni.produkty.pro.zivot) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor

Get **every URL a website publishes in its sitemaps**, in seconds, as a clean table you can export to CSV, Excel or JSON, or pipe into your next crawler, SEO audit or AI agent.

Just enter a domain like `example.com`. The Actor finds the sitemaps for you (from `robots.txt` and the usual locations), follows nested sitemap indexes, unpacks `.xml.gz` files and returns one row per unique URL.

### What you can use it for

- **SEO audits:** compare what is in the sitemap with what is indexed, find orphan pages, check `lastmod` hygiene and hreflang coverage.
- **Content monitoring:** use the date filter to list only pages **added or updated since a given day**. Schedule it daily to track a competitor's new products, blog posts or landing pages.
- **Crawl planning:** feed the exact list of URLs into another scraper instead of crawling the whole site link by link. This is faster and cheaper.
- **AI / RAG pipelines:** get the canonical list of pages to load into your knowledge base.
- **Migrations and QA:** export the full URL inventory before and after a site redesign.

### Features

- Automatic sitemap discovery: `robots.txt` `Sitemap:` lines, then `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml`, `/sitemap.txt` and more.
- Nested **sitemap index** files, **gzip** (`.xml.gz`), **plain-text** sitemaps and **RSS / Atom** feeds.
- Returns `lastmod`, `changefreq`, `priority`, **hreflang alternates**, image and video counts, and news titles.
- **Filters:** regex include/exclude rules and a `lastmod` date range.
- De-duplicates URLs across all sitemaps and websites in the run.
- Many websites in one run, processed in parallel with polite limits.
- Never fails the whole run because of one broken sitemap. Problems are listed in the `SUMMARY` record.
- Lightweight HTTP only, with no browser. That keeps it fast and inexpensive.

### Input

| Field | Description |
|---|---|
| `startUrls` | Websites (`example.com`) or direct sitemap/feed URLs. |
| `maxUrls` | Stop after this many unique URLs (0 = no limit). |
| `includePatterns` / `excludePatterns` | Regular expressions matched against each URL. |
| `lastmodFrom` / `lastmodTo` | Keep only URLs modified within this date range. |
| `keepUrlsWithoutLastmod` | Whether URLs without `lastmod` pass a date filter (default yes). |
| `includeAlternates` | Add hreflang alternates to each row. |
| `maxSitemaps`, `concurrency` | Safety limits. |
| `proxyConfiguration` | Optional. Only for sites that block cloud servers. |

Example:

```json
{
  "startUrls": ["https://crawlee.dev", "example-shop.com"],
  "includePatterns": ["/blog/"],
  "lastmodFrom": "2026-09-01",
  "maxUrls": 5000
}
```

### Output

One row per unique URL:

```json
{
  "url": "https://crawlee.dev/blog/scrapy-vs-crawlee",
  "lastmod": "2026-09-14",
  "changefreq": "weekly",
  "priority": 0.5,
  "title": null,
  "imageCount": 0,
  "videoCount": 0,
  "domain": "crawlee.dev",
  "sitemapUrl": "https://crawlee.dev/sitemap.xml",
  "startUrl": "https://crawlee.dev/",
  "alternates": [{ "hreflang": "de", "href": "https://crawlee.dev/de/blog/scrapy-vs-crawlee" }]
}
```

A `SUMMARY` record in the key-value store reports how many sitemaps were found, processed and failed, how many URLs were filtered or de-duplicated, and why the run stopped.

### Pricing

You pay only for the URLs you get. Each saved URL is one result. Use **Maximum cost per run** or `maxUrls` to stay within your budget: the Actor stops cleanly when the limit is reached.

### Tips

- No results for a site? Some websites simply have no sitemap. The `SUMMARY` record says so explicitly.
- To track only new pages, schedule the Actor daily with `lastmodFrom` set to yesterday's date.
- Huge sites (millions of URLs) work too. Raise `maxSitemaps` and set `maxUrls` to control cost.

### Responsible use

The Actor reads only sitemaps and feeds, which websites publish specifically for automated discovery. It does not log in, bypass protections or collect personal data. Please respect each website's terms when you reuse the URLs you obtain.

### How to use it via API

You can run the Actor from the Apify Console, on a schedule, or from your own code. Get your API token in Apify Console → Settings → Integrations.

**Python**

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("digitalni.produkty.pro.zivot/sitemap-url-extractor").call(run_input={
    "startUrls": [
        "https://crawlee.dev"
    ],
    "maxUrls": 1000
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

**JavaScript / Node.js**

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });
const run = await client.actor('digitalni.produkty.pro.zivot/sitemap-url-extractor').call({
    "startUrls": [
        "https://crawlee.dev"
    ],
    "maxUrls": 1000
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

### Integrations and AI agents

- Export results as **JSON, CSV, Excel, XML or HTML**, or open them directly in **Google Sheets**.
- Connect to **Make, Zapier, n8n, Slack, Google Drive, Airbyte** or any **webhook** to get all URLs from the sitemaps of a website into your workflow automatically.
- Use it from **AI agents and LLM apps** (Claude, ChatGPT, Cursor, LangChain…) through the **Apify MCP server**: the agent can call this Actor as a tool.
- **Schedule** runs (hourly, daily, weekly) to keep data fresh without any code.

### FAQ

**How much does it cost?**
$0.50 per 1,000 URLs. A 500-page website costs about $0.25, a 20,000-page shop about $10. Apify's free plan includes monthly credits, so small sites are effectively free to try.

**Do I need to know where the sitemap is?**
No. Enter just the domain. The Actor reads robots.txt, tries common sitemap locations and follows nested sitemap indexes and gzipped files.

**Can I get only new or updated pages?**
Yes. Use the lastmod date filter and schedule the Actor (e.g. weekly) to get only pages changed since a given date.

**Is it legal?**
Sitemaps are published by website owners precisely so that machines can read them. The Actor only downloads sitemap files, not page content.

### More tools from the same developer

All tools with code examples: [github.com/Phenixik/apify-actors](https://github.com/Phenixik/apify-actors)

- [Broken Link Checker](https://apify.com/digitalni.produkty.pro.zivot/broken-link-checker): find 404s and dead links on your site
- [PageSpeed & Lighthouse Audit](https://apify.com/digitalni.produkty.pro.zivot/pagespeed-audit): bulk Core Web Vitals and SEO scores
- [Website Screenshot](https://apify.com/digitalni.produkty.pro.zivot/website-screenshot): full-page, mobile and PDF screenshots
- [Domain WHOIS & DNS Lookup](https://apify.com/digitalni.produkty.pro.zivot/domain-whois-lookup): expiry, DNS, SPF/DMARC and SSL for many domains
- [PDF Text Extractor](https://apify.com/digitalni.produkty.pro.zivot/pdf-text-extractor): text, Markdown and RAG chunks from PDFs
- [AI Image Upscaler](https://apify.com/digitalni.produkty.pro.zivot/ai-image-upscaler): 4x Real-ESRGAN upscaling, no GPU
- [RSS Feed Reader & Monitor](https://apify.com/digitalni.produkty.pro.zivot/rss-feed-reader): only-new-items feed monitoring
- [ATS Jobs](https://apify.com/digitalni.produkty.pro.zivot/ats-jobs): jobs from Greenhouse, Lever, Ashby & more
- [App Store Reviews](https://apify.com/digitalni.produkty.pro.zivot/app-store-reviews): iOS reviews from 50+ countries

### Support

Found a sitemap that is not parsed correctly? Open an issue with the URL and we will fix it, usually within a day or two.

# Actor input Schema

## `startUrls` (type: `array`):

Website addresses (e.g. https://example.com or just example.com) or direct links to sitemap.xml, sitemap indexes, .xml.gz, .txt sitemaps or RSS/Atom feeds. For websites, sitemaps are discovered from robots.txt and common locations.

## `maxUrls` (type: `integer`):

Stop after saving this many unique URLs in total. 0 means no limit.

## `includePatterns` (type: `array`):

Keep only URLs that match at least one of these regular expressions, e.g. /blog/ or /products/ . Case-insensitive.

## `excludePatterns` (type: `array`):

Drop URLs that match any of these regular expressions, e.g. /tag/ or page= . Case-insensitive.

## `lastmodFrom` (type: `string`):

Keep only URLs whose <lastmod> is on or after this date (YYYY-MM-DD or full ISO timestamp). Great for finding new or updated pages.

## `lastmodTo` (type: `string`):

Keep only URLs whose <lastmod> is on or before this date (YYYY-MM-DD or full ISO timestamp).

## `keepUrlsWithoutLastmod` (type: `boolean`):

Many sitemaps omit <lastmod>. If a date filter is set, decide whether such URLs are kept.

## `includeAlternates` (type: `boolean`):

Add the list of language/region alternates (xhtml:link hreflang) declared for each URL.

## `maxSitemaps` (type: `integer`):

Safety limit on how many sitemap files (including nested ones) are downloaded.

## `concurrency` (type: `integer`):

How many sitemap files are downloaded at the same time. Keep it low to be gentle to websites.

## `proxyConfiguration` (type: `object`):

Optional. Only needed for websites that block cloud servers. Without a proxy the run is cheaper.

## Actor input object example

```json
{
  "startUrls": [
    "https://crawlee.dev"
  ],
  "maxUrls": 100,
  "includePatterns": [],
  "excludePatterns": [],
  "keepUrlsWithoutLastmod": true,
  "includeAlternates": true,
  "maxSitemaps": 1000,
  "concurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `urls` (type: `string`):

Unique URLs with lastmod, changefreq, priority, hreflang alternates and the sitemap they came from.

## `summary` (type: `string`):

Counts of discovered, processed and failed sitemaps, filtered and duplicate URLs, and the reason the run stopped.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://crawlee.dev"
    ],
    "maxUrls": 100
};

// Run the Actor and wait for it to finish
const run = await client.actor("digitalni.produkty.pro.zivot/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["https://crawlee.dev"],
    "maxUrls": 100,
}

# Run the Actor and wait for it to finish
run = client.actor("digitalni.produkty.pro.zivot/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://crawlee.dev"
  ],
  "maxUrls": 100
}' |
apify call digitalni.produkty.pro.zivot/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,digitalni.produkty.pro.zivot/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZrNRhM8L7FXRosqGP/builds/iDEDqf8ghuPR41OyY/openapi.json
