# Sitemap Extractor - All Website URLs + Broken Link Check (`adaptive_arbor_rtz/sitemap-extractor`) Actor

Find every page of a website from its sitemaps (robots.txt, sitemap indexes, .gz). Returns URLs with last-modified dates, filters by pattern, and can check each URL for 404s and redirects.

- **URL**: https://apify.com/adaptive\_arbor\_rtz/sitemap-extractor.md
- **Developed by:** [Björn Ólafur](https://apify.com/adaptive_arbor_rtz) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.30 / 1,000 urls

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Sitemap Extractor do?

**Sitemap Extractor** finds **every page of a website** from its XML sitemaps. Just enter a domain: it reads **robots.txt**, checks the **common sitemap locations**, follows **sitemap indexes** and **.gz files**, and returns **each URL with its last-modified date, change frequency and priority**. Optionally it **checks every URL for 404s and redirects**, a quick broken-link audit of your sitemap.

Run it in Apify Console, via the API, on a schedule, or from your AI agent through the Apify MCP server. Connect it to Google Sheets, Make, Zapier or n8n with Apify integrations.

### Why use Sitemap Extractor?

- 🗺️ **Automatic discovery**: `robots.txt` → `sitemap.xml`, `sitemap_index.xml`, `wp-sitemap.xml`, `.gz` and text sitemaps.
- 🔁 **Follows nested sitemap indexes**, with duplicates removed.
- 🎯 **Filters**: keep only `/blog/` or `/products/` URLs, or drop `/tag/` pages.
- 🩺 **Broken-link check** (optional): HTTP status and redirect target for every URL.
- 🖼️ **Image counts** from image sitemaps, and news/video sitemaps are supported.
- 💸 **Cheap and fair**: pay per URL, and errors are free.

**Typical uses:** SEO audits and site migrations (find 404s and redirects listed in your sitemap), getting a full URL list before scraping or crawling, monitoring competitors' new pages via `lastmod`, feeding page lists to RAG and AI agents, and content inventories.

### How to extract all URLs from a website's sitemap

1. Click **Try for free**.
2. Enter websites (e.g. `apify.com`) or direct sitemap links in **Websites or sitemap URLs**.
3. Optionally add **include/exclude patterns** and switch on **Check each URL's status**.
4. Click **Start**. URLs stream into the **Output** tab.
5. Download them as CSV, Excel or JSON, or use the API.

### Input

| Option            | What it does                                        | Default |
| ----------------- | --------------------------------------------------- | ------- |
| `startUrls`       | Websites or sitemap URLs                            | –       |
| `includePatterns` | Keep only URLs matching any pattern (regex or text) | –       |
| `excludePatterns` | Drop URLs matching any pattern                      | –       |
| `maxUrlsPerSite`  | Stop after this many URLs per website               | 100,000 |
| `checkStatus`     | Request each URL and report status and redirects    | `false` |
| `maxConcurrency`  | Parallel requests                                   | 10      |

```json
{
    "startUrls": ["apify.com", "https://www.theguardian.com/sitemaps/news.xml"],
    "includePatterns": ["/blog/"],
    "checkStatus": true
}
```

### Output

One item per URL:

```json
{
    "url": "https://apify.com/pricing",
    "lastmod": "2026-09-01",
    "changefreq": "weekly",
    "priority": 0.8,
    "imageCount": 0,
    "site": "https://apify.com",
    "sitemapUrl": "https://apify.com/sitemap/pages.xml",
    "success": true,
    "statusCode": 200,
    "finalUrl": null,
    "redirected": false
}
```

`statusCode`, `finalUrl` and `redirected` are included only when **Check each URL's status** is on. Websites without a sitemap are listed with `success: false` (free). A per-website summary is saved as `SUMMARY` in the key-value store. You can download the dataset in various formats such as JSON, HTML, CSV or Excel.

### How much does it cost to extract a sitemap?

- **$0.30 per 1,000 URLs** ($0.0003 each)
- **+ $1 per 1,000 status checks** when **Check each URL's status** is on
- **+ $0.002 per run**

Example: a 5,000-page website costs **$1.50**, or **$6.50** with status checks. Platform usage is included. Set a **maximum cost per run** to stay in budget; the Actor stops cleanly when it's reached.

### Tips

- Use **Max URLs per website** for a quick sample of a huge site.
- Keep **Parallel requests** low (5–10) with status checks on small websites, to be polite.
- Some sites list sitemaps only in `robots.txt`. That's the first place we check.
- If a site blocks data-centre traffic, enable **Proxy**.

### FAQ

**The site has no sitemap?** It's listed with an error and not charged. You'd need a crawler instead, such as Apify's Website Content Crawler.

**Is it legal?** Sitemaps are published so machines can read them. Keep request rates reasonable when checking statuses, and respect each site's terms.

**Something not working?** Open an issue in the **Issues** tab with the website and I'll fix it quickly.

### Related tools

- [Website Screenshot](https://apify.com/adaptive_arbor_rtz/website-screenshot): full-page screenshots and PDFs of any website, cookie banners removed
- [Document to Markdown](https://apify.com/adaptive_arbor_rtz/document-to-markdown): PDF, Word, Excel and PowerPoint files to clean Markdown for AI
- [RSS Feed Reader](https://apify.com/adaptive_arbor_rtz/rss-feed-reader): read and monitor any RSS/Atom feed, only new items, full article text

# Actor input Schema

## `startUrls` (type: `array`):

Website addresses (e.g. apify.com) or direct sitemap links (e.g. https://example.com/sitemap.xml), one per line. For websites, sitemaps are found automatically via robots.txt and common locations; sitemap indexes and .gz files are followed.

## `includePatterns` (type: `array`):

Optional. Keep only URLs that match at least one of these patterns (regular expressions or plain text), e.g. '/blog/' or '/products/'.

## `excludePatterns` (type: `array`):

Optional. Drop URLs that match any of these patterns, e.g. '/tag/' or '?page='.

## `maxUrlsPerSite` (type: `integer`):

Stop after this many URLs per website.

## `checkStatus` (type: `boolean`):

Also request every URL and report its HTTP status code and redirect target. Finds 404s and redirects listed in your sitemap. Charged per checked URL.

## `maxConcurrency` (type: `integer`):

How many requests to make at the same time. Keep it low for small websites.

## `proxyConfiguration` (type: `object`):

Optional. Use a proxy if a website blocks data-centre traffic.

## Actor input object example

```json
{
  "startUrls": [
    "https://playwright.dev"
  ],
  "maxUrlsPerSite": 100000,
  "checkStatus": false,
  "maxConcurrency": 10,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://playwright.dev"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("adaptive_arbor_rtz/sitemap-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": ["https://playwright.dev"] }

# Run the Actor and wait for it to finish
run = client.actor("adaptive_arbor_rtz/sitemap-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://playwright.dev"
  ]
}' |
apify call adaptive_arbor_rtz/sitemap-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,adaptive_arbor_rtz/sitemap-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/rnXZ75F64CKqtAeqw/builds/pgPK1KWKqR65ytd4Z/openapi.json
