# Sitemap URL Extractor: Auto-Discovery, Filters & Health Report (`creativefour/sitemap-url-extractor`) Actor

Get every URL from a website's sitemaps: just enter the domain. Handles sitemap indexes and .xml.gz, adds lastmod, images, hreflang and news data, filters by date or URL pattern, and reports sitemap problems.

- **URL**: https://apify.com/creativefour/sitemap-url-extractor.md
- **Developed by:** [CreativeFour LLC](https://apify.com/creativefour) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does the Sitemap URL Extractor do?

Enter **just a domain** and get **every URL from the site's sitemaps**, with each page's **last-modified date, change frequency, and priority**, plus any **images, language versions (hreflang), and news data** the sitemap declares.

It **finds sitemaps automatically** (from robots.txt, or the usual paths like /sitemap.xml), follows **sitemap indexes to any depth**, reads **gzipped .xml.gz** files, and writes a **sitemap health report**: duplicates, off-domain URLs, invalid dates, and broken sitemap files.

On Apify you can schedule it, call it from the API, or send the list straight into Google Sheets, a crawler, or an SEO tool.

### Why use it?

- **Get a site's full page list in seconds**, without crawling it.
- **Watch what changes.** Use **Only URLs modified on or after** with a schedule to see what a competitor published or updated this week.
- **Audit your own sitemaps.** Catch duplicates, URLs pointing at another domain, broken child sitemaps, and missing or invalid `<lastmod>` dates.
- **Feed other tools.** Pass the URLs to a status checker, crawler, Lighthouse audit, or RAG pipeline.
- **Check international SEO.** See each page's hreflang alternates side by side.

### How to use it

1. Open the **Input** tab and type a domain (for example `example.com`), or paste sitemap URLs.
2. Optional: add a **modified since** date, or URL patterns to include or skip.
3. Click **Start**.
4. Open the **Output** tab for the URL list, and the **SUMMARY** record for the health report.

### Input

| Field | What it does |
|---|---|
| **Websites** | Domains. Sitemaps are discovered from robots.txt or the usual paths. |
| **Sitemap URLs** | Or give sitemap or index URLs directly (.xml or .xml.gz). |
| **Only URLs modified on or after** | Keeps URLs whose `<lastmod>` is on or after this date. |
| **Only URLs matching / Skip URLs matching** | Regular expressions or plain text, such as `/blog/` or `/tag/`. |
| **Max URLs / Max sitemap files** | Safety limits. |

```json
{
  "websites": ["example.com"],
  "modifiedSince": "2026-09-01",
  "includePatterns": ["/blog/"]
}
```

### Output

One row per URL. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

```json
{
  "site": "example.com",
  "url": "https://example.com/blog/new-post",
  "lastmod": "2026-09-24T10:00:00Z",
  "changefreq": "weekly",
  "priority": 0.8,
  "images": ["https://example.com/img/hero.jpg"],
  "alternates": [{ "hreflang": "es", "href": "https://example.com/es/blog/new-post" }],
  "newsTitle": null,
  "newsPublishedAt": null,
  "videos": 0,
  "sitemap": "https://example.com/post-sitemap.xml.gz",
  "offDomain": false,
  "lastmodValid": true
}
```

**SUMMARY** (key-value store) gives, for each site: the number of sitemaps read, URLs found and saved, duplicates, off-domain URLs, invalid and missing `lastmod` dates, and every sitemap file that failed to load, with the reason.

### Data fields

| Field | Description |
|---|---|
| `url`, `lastmod`, `changefreq`, `priority` | Standard sitemap fields |
| `images`, `videos` | Image URLs and video count, from image and video sitemap extensions |
| `alternates` | hreflang language versions |
| `newsTitle`, `newsPublishedAt` | News sitemap fields |
| `sitemap` | Which sitemap file listed the URL |
| `offDomain`, `lastmodValid` | Health flags |

### How much does it cost to extract sitemap URLs?

You pay per URL saved. Filters (modified-since and URL patterns) are applied first, so you only pay for the URLs you keep. Set a **maximum charge per run** in the run options, and the Actor stops cleanly at that limit.

### Tips

- **Check the URLs next** with our [Bulk URL Status & Redirect Checker](https://apify.com/creativefour/url-status-audit), or audit their speed with the [Bulk Lighthouse & Core Web Vitals Audit](https://apify.com/creativefour/lighthouse-audit).
- **Some sites have no sitemap.** The summary says so, and the [Broken Link Checker](https://apify.com/creativefour/broken-link-crawler) can crawl them instead.

### Use it from AI agents (MCP)

AI agents can find and run this Actor through the [Apify MCP server](https://docs.apify.com/integrations/mcp).

- **Claude, ChatGPT, or any MCP client:** add `https://mcp.apify.com?tools=creativefour/sitemap-url-extractor` as a custom connector, and sign in to Apify when prompted.
- **Claude Code, Cursor, VS Code, or Codex:** run `apify mcp install claude-code` (swap in your client's name), then ask your agent to "list every URL on example.com with creativefour/sitemap-url-extractor".

### FAQ and support

**What if a site has no sitemap?** The summary reports "No readable sitemap found". Sitemaps are optional, and some sites don't publish one.

**Does it visit every page?** No. It reads only the sitemap files, which makes it fast and light on the site.

**Found a bug or need a feature?** Open an issue on the **Issues** tab. Custom versions are available on request.

# Actor input Schema

## `websites` (type: `array`):

Just the domain is enough (example.com). Sitemaps are found through robots.txt, or the usual paths such as /sitemap.xml and /sitemap\_index.xml.

## `sitemapUrls` (type: `array`):

Or give sitemap or sitemap-index URLs directly (.xml or .xml.gz). Nested indexes are followed.

## `modifiedSince` (type: `string`):

YYYY-MM-DD. Uses each URL's <lastmod>. Great for watching what changed on a competitor's site. URLs without <lastmod> are skipped when this is set.

## `includePatterns` (type: `array`):

Regular expressions or plain text, for example /blog/ or /products/. Leave empty to keep everything.

## `excludePatterns` (type: `array`):

Regular expressions or plain text, for example /tag/ or ?page=

## `maxUrls` (type: `integer`):

Stop after saving this many URLs in total.

## `maxSitemaps` (type: `integer`):

Safety limit on how many sitemap files to read per site.

## Actor input object example

```json
{
  "websites": [
    "apify.com"
  ],
  "maxUrls": 200,
  "maxSitemaps": 500
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "apify.com"
    ],
    "maxUrls": 200
};

// Run the Actor and wait for it to finish
const run = await client.actor("creativefour/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": ["apify.com"],
    "maxUrls": 200,
}

# Run the Actor and wait for it to finish
run = client.actor("creativefour/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "apify.com"
  ],
  "maxUrls": 200
}' |
apify call creativefour/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,creativefour/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/hM2mq1XRbSMQeVkLJ/builds/nK6rwyDDstF6n2xf6/openapi.json
