# Sitemap URL Extractor — robots.txt, Sitemap Index, gzip (`chorelet/sitemap-url-extractor`) Actor

Every URL of a website from its sitemaps: robots.txt discovery, sitemap indexes, gzip, lastmod, changefreq, priority, hreflang alternates, image/video/news extensions. Include/exclude patterns and date filter, CSV/JSON export and API.

- **URL**: https://apify.com/chorelet/sitemap-url-extractor.md
- **Developed by:** [Chorelet](https://apify.com/chorelet) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.21 / 1,000 urls

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor — robots.txt, Sitemap Index, gzip

Get every URL a website publishes in its sitemaps — with `lastmod`, `changefreq`, `priority`, hreflang alternates and image/video/news extensions — as JSON, CSV or Excel, or via API. Paste website URLs (the sitemap is discovered through `robots.txt` and common paths) or sitemap URLs directly; sitemap indexes and `.gz` files are followed automatically.

### Why this Actor

- Paste a website — the sitemap is found via robots.txt and nine common paths
- Sitemap indexes and gzip files followed automatically
- lastmod, changefreq, priority, hreflang alternates, image/video/news extensions
- Include/exclude regular expressions and a modified-after filter
- Checked every day by an automated run

### Sample output

One item of the dataset (long values shortened):

```json
{
  "site": "https://apify.com",
  "url": "https://apify.com/",
  "lastmod": null,
  "changefreq": null,
  "priority": null,
  "sitemapUrl": "https://apify.com/sitemap/pages.xml"
}
```

### What you get

| Field | Description |
|---|---|
| `site`, `sitemapUrl` | Where the URL came from |
| `url` | The page |
| `lastmod`, `changefreq`, `priority` | As declared in the sitemap (`lastmod` normalised to ISO 8601) |
| `alternates` | `hreflang` variants |
| `images`, `videos`, `news` | Sitemap extensions when present |

A per-site summary (how the sitemap was found, sitemaps read, URLs extracted, errors) is saved as `SUMMARY`.

### Input

- **Websites or sitemaps** — `example.com`, `https://example.com`, or `https://example.com/sitemap_index.xml`.
- **Max URLs per site**, **Max sitemaps per site**.
- **Include / Exclude URLs matching** — regular expressions, e.g. include `/blog/`, exclude `\.pdf$`.
- **Only URLs modified after** — date filter on `lastmod`.

### Limits and notes

- Discovery order: `Sitemap:` lines in robots.txt, then `/sitemap.xml`, `/sitemap_index.xml`, `/sitemap-index.xml`, `/sitemap/sitemap.xml`, `/wp-sitemap.xml`, `/sitemap1.xml`, `/sitemaps.xml`, `/sitemap.xml.gz`, `/sitemap.txt`.
- Sitemaps behind bot protection or login return an error for that site; everything else still completes.
- Public data only; the Actor stores nothing beyond the dataset of your run.

### Input example

```json
{
  "urls": [
    "https://blog.cloudflare.com"
  ],
  "maxUrlsPerSite": 10000,
  "maxSitemapsPerSite": 200
}
```

### How much does it cost?

Pay per url — no subscription, no minimum, no charge for platform usage.

| Volume | Price |
|---|---|
| 1,000 URLs | $0.30 |
| 10,000 URLs | $3.00 |
| 100,000 URLs | $30.00 |

The Apify **free plan includes $5 of usage every month** — about 16,666 URLs with this Actor, no card needed. Nothing else is charged: platform usage is included in the price, and Apify Bronze, Silver and Gold subscribers get 10%, 20% and 30% off these prices.

### Use it from code, n8n, Make, Zapier or an AI agent

Run the Actor and download the dataset in one call (JSON by default; add `&format=csv` or `xlsx`):

```bash
curl -X POST "https://api.apify.com/v2/acts/chorelet~sitemap-url-extractor/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["https://blog.cloudflare.com"], "maxUrlsPerSite": 10000, "maxSitemapsPerSite": 200}'
```

Python:

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("chorelet/sitemap-url-extractor").call(run_input={"urls": ["https://blog.cloudflare.com"], "maxUrlsPerSite": 10000, "maxSitemapsPerSite": 200})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

- **n8n, Make, Zapier** — use the Apify node/module: run the Actor, then "get dataset items".
- **Google Sheets, Slack, webhooks** — add an integration on the run's *Integrations* tab.
- **AI agents** — the Actor is available as a tool through the Apify MCP server; the dataset schema describes every field for the model.
- **Schedules** — run it hourly, daily or weekly from the *Schedules* tab.

### FAQ

**What if the site has no sitemap?**

The Actor reports it for that site and continues with the rest. Try a direct sitemap URL if you know one.

**Can I take only part of a site?**

Yes — `includePattern` and `excludePattern` are regular expressions over the URL, e.g. `/blog/` or `\.pdf$`.

**How big can a site be?**

Hundreds of thousands of URLs are fine; set `maxUrlsPerSite` and `maxSitemapsPerSite` to cap the run.

**Does it crawl pages?**

No — it reads only the sitemaps, which is why it is fast and cheap. Feed the URLs into a crawler if you need page content.

**What does a run cost?**

$0.30 per 1,000 URLs. The free plan's $5 a month covers about 16,000 URLs.

### Support

Questions, missing fields or a source that changed? Open an issue on the *Issues* tab or write to support@chorelet.app — problems are usually fixed within a day, and the Actor is checked every morning by an automated test run. If the Actor saved you time, a short review on its Store page helps other people find it.

# Actor input Schema

## `urls` (type: `array`):

Website URLs (the sitemap is found via robots.txt or common paths) or direct sitemap URLs, including sitemap indexes and .gz files.

## `maxUrlsPerSite` (type: `integer`):

Stop after this many URLs for one site.

## `maxSitemapsPerSite` (type: `integer`):

Sitemap indexes can list hundreds of child sitemaps; this caps how many are read.

## `includePattern` (type: `string`):

Regular expression (case-insensitive), e.g. `/blog/`. Empty = all.

## `excludePattern` (type: `string`):

Regular expression (case-insensitive), e.g. `\.pdf$|/tag/`.

## `modifiedAfter` (type: `string`):

ISO date, e.g. `2026-01-01`. Applies to sitemaps that carry lastmod.

## Actor input object example

```json
{
  "urls": [
    "https://blog.cloudflare.com"
  ],
  "maxUrlsPerSite": 10000,
  "maxSitemapsPerSite": 200
}
```

# Actor output Schema

## `urls` (type: `string`):

All extracted URLs — items of the default dataset. Use ?format=csv or xlsx on this URL for spreadsheets.

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://blog.cloudflare.com"
    ],
    "maxUrlsPerSite": 10000,
    "maxSitemapsPerSite": 200
};

// Run the Actor and wait for it to finish
const run = await client.actor("chorelet/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://blog.cloudflare.com"],
    "maxUrlsPerSite": 10000,
    "maxSitemapsPerSite": 200,
}

# Run the Actor and wait for it to finish
run = client.actor("chorelet/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://blog.cloudflare.com"
  ],
  "maxUrlsPerSite": 10000,
  "maxSitemapsPerSite": 200
}' |
apify call chorelet/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,chorelet/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Aacn1lY2SUAAdwhj3/builds/mdVJsdfZvdvDCALs2/openapi.json
