# Sitemap URL Extractor & Scraper - All Page URLs of a Site (`datagleaner/sitemap-extractor`) Actor

Sitemap scraper and URL extractor: give it a website or sitemap.xml and get every page URL with lastmod, changefreq and priority. Finds sitemaps via robots.txt, follows sitemap indexes, reads .xml.gz and text sitemaps, survives broken XML. Plain HTTP, no browser. $0.20 per 1,000 URLs.

- **URL**: https://apify.com/datagleaner/sitemap-extractor.md
- **Developed by:** [Data Gleaner](https://apify.com/datagleaner) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor & Scraper - Extract All URLs From Any Website

Sitemap extractor that gets every page URL from a website's sitemap.xml: give it a domain, get all URLs with lastmod, changefreq and priority as JSON or CSV. $0.20 per 1,000 URLs.

It finds the sitemaps through `robots.txt` and common paths, follows sitemap indexes, reads `.xml.gz` and plain-text sitemaps, and returns one row per URL with last-modified date, change frequency, priority, image URLs, hreflang alternates and the sitemap it came from. An optional status check (`checkStatus`) adds each page's HTTP status and redirect target, for SEO audits. Plain HTTP, no browser, no login, no proxy needed for most sites.

It is built to **finish the run**. If one site has no sitemap, one sitemap file is broken, a host is down or a server returns an error, the Actor logs it, skips it and carries on with the rest. Every run ends with a per-site summary in the `SUMMARY` key-value record.

### What does Sitemap URL Extractor do?

It finds a site's sitemaps, reads them all, and returns the page URLs as JSON, CSV or Excel through the Apify dataset and API. Typical uses:

- **SEO audits and site inventories.** List every indexable URL of a site, with `lastmod`, before a migration or a crawl.
- **Competitor research.** See how big a competitor's site is and which sections it grows (filter with `includeUrlPatterns` such as `/blog/` or `/products/`).
- **Seed lists for scrapers.** Feed the URLs into a crawler or a product / article scraper instead of crawling links.
- **Change monitoring.** Run daily and compare `lastmod` to find pages that are new or updated.
- **LLM and RAG pipelines.** Get the page list of a documentation site, then fetch and embed only the pages you want.

### How sitemaps are found

1. `robots.txt` `Sitemap:` lines.
2. If those give nothing: `/sitemap.xml`, `/sitemap_index.xml`, `/sitemap-index.xml`, `/wp-sitemap.xml` (WordPress).
3. Sitemap index files are followed up to `maxCrawlingDepth` levels (default 6). Sitemap files of an index are fetched 4 at a time, so large sites finish faster; `requestDelaySeconds` still spaces the starts. Gzip files (`.xml.gz`) are unpacked by content, plain-text sitemaps (one URL per line) are read, and broken or truncated XML is recovered as far as possible.

You can also pass a direct sitemap URL (ending in `.xml`, `.xml.gz` or `.txt`) instead of a site.

### Input

| Field | Type | What it does |
|---|---|---|
| `websites` | list of text | Website URLs or bare domains (`https://aioseo.com`, `stripe.com`). Only the site root is used. Leave it empty to run a 5-URL example on aioseo.com. |
| `startUrls` | list | Same as `websites` in the request-list format, so input written for other sitemap Actors works unchanged. Used only when `websites` is empty. |
| `maxUrlsPerSite` | number | Cap per website. Default 5,000 when left out (the Console form prefills 100). It is also your cost cap per site. |
| `checkStatus` | boolean | Default false. Sends one HEAD request per URL (GET if HEAD is refused) and adds `status`, `finalUrl` and `checkError`, to find 404s, 5xx errors and redirects listed in the sitemap. Fills a missing `lastmod` from the page's `Last-Modified` header (`lastmodSource` = `http-header`). Slower, about 5 pages a second. No extra charge. |
| `includeUrlPatterns` | list of text | Optional regular expressions (Python syntax). Keep a URL only if it matches at least one. An invalid expression fails the run before anything is charged. |
| `excludeUrlPatterns` | list of text | Optional regular expressions. Drop a URL if it matches any. Invalid ones fail the run too. |
| `requestDelaySeconds` | number | Pause between sitemap and `robots.txt` requests, default 0.3 s (randomised up to 1.5x). Status checks are not paced by it. |
| `maxCrawlingDepth` | number | How many levels of sitemap indexes to follow. Default 6, range 1 to 20. 1 reads the top-level index's child sitemaps but not indexes nested inside it. |
| `maxRequestRetries` | number | Retries per sitemap file on timeouts, connection errors, 429 and 5xx gateway errors (500, 502, 503, 504), with exponential backoff that honours `Retry-After`. Default 3, range 0 to 10. |
| `proxyConfiguration` | object | Optional Apify Proxy. Default is no proxy. |

```json
{
  "websites": ["https://aioseo.com", "stripe.com"],
  "maxUrlsPerSite": 5000,
  "includeUrlPatterns": ["/blog/"],
  "excludeUrlPatterns": ["\\.pdf$"]
}
```

URLs are deduplicated per site, so a page listed in several sitemaps is returned once.

### Output

One dataset item per page URL. Export as JSON, CSV, Excel or via the API.

| Field | Meaning |
|---|---|
| `url` | The page URL. |
| `lastmod` | Last modification date as ISO 8601 (`2026-09-30` or `2026-09-30T10:00:00+00:00`), or `null` if the sitemap has none or an invalid value. |
| `changefreq` | `always`, `hourly`, `daily`, `weekly`, `monthly`, `yearly`, `never`, or `null`. |
| `priority` | Number from 0 to 1, or `null`. |
| `sitemapUrl` | The sitemap file this URL came from. |
| `lastmodSource` | `sitemap`, `http-header` (only with `checkStatus`), or `null` when there is no `lastmod`. |
| `domain` | Host name of the site. |
| `imageCount` | Number of `image:image` entries listed for the page (0 when none). |
| `images` | Image URLs (`image:loc`) listed for the page. |
| `alternates` | hreflang language versions as `{hreflang, href}` objects. |
| `status` | HTTP status of the page. Only with `checkStatus`. |
| `finalUrl` | URL after redirects. Only with `checkStatus`. |
| `checkError` | Network error if the status check failed. Only with `checkStatus`. |
| `website` | The input value this row came from. |
| `scrapedAt` | ISO 8601 UTC timestamp. |

```json
{
  "url": "https://stripe.com/payments",
  "lastmod": "2026-09-30",
  "changefreq": "weekly",
  "priority": 0.8,
  "sitemapUrl": "https://stripe.com/sitemap/partition-0.xml",
  "domain": "stripe.com",
  "lastmodSource": "sitemap",
  "imageCount": 0,
  "images": [],
  "alternates": [],
  "website": "stripe.com",
  "scrapedAt": "2026-10-08T09:30:00+00:00"
}
```

The dataset has two views: **Page URLs** and **SEO audit** (status, redirects, `lastmodSource`, hreflang). Many sites publish only `<loc>`. Then `lastmod`, `changefreq` and `priority` are `null`; that is what the site publishes, not a failed extraction.

### Pricing

Pay per event: **$0.20 per 1,000 URLs** ($0.0002 for each URL saved to the dataset). With `checkStatus` on, each URL is instead charged as a **checked URL** at **$0.40 per 1,000** ($0.0004, one event per URL, never both), because every URL costs an extra HTTP request. You pay only for URLs you receive: nothing for sites without a sitemap, failed requests or retries, and the run stops by itself when your **maximum total charge** is reached. Platform usage is included.

Example: 20 sites at the default 5,000 URLs is up to 100,000 URLs, so up to $20. A site with 1,200 URLs costs $0.24.

### Use it from the API

Run it synchronously and get the URLs back as JSON. Replace `<YOUR_APIFY_TOKEN>` with your token.

```bash
curl -X POST "https://api.apify.com/v2/acts/datagleaner~sitemap-extractor/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"websites": ["stripe.com"], "maxUrlsPerSite": 1000, "includeUrlPatterns": ["/docs/"]}'
```

It also works from n8n, Make or Zapier through the Apify integrations.

### Use with Python

Install the client with `pip install apify-client` and set `APIFY_TOKEN`. This run returns at most 10 URLs, so it costs at most $0.002.

```python
import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("datagleaner/sitemap-extractor").call(run_input={
    "websites": ["stripe.com"],
    "maxUrlsPerSite": 10,
})
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item["url"], item["lastmod"])
```

### Use with JavaScript / Node.js

Install the client with `npm install apify-client` and set `APIFY_TOKEN`. Save as `run.mjs` and run `node run.mjs`.

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagleaner/sitemap-extractor').call({
    websites: ['stripe.com'],
    maxUrlsPerSite: 10,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) console.log(item.url, item.lastmod);
```

### Use it from n8n, Make, Zapier or an AI agent

Actor ID: `datagleaner/sitemap-extractor`

Minimal input:

```
{"websites": ["stripe.com"], "maxUrlsPerSite": 10}
```

Each tool below runs this Actor with your own Apify API token.

- **n8n:** add the **Apify** node (`@apify/n8n-nodes-apify`). On n8n Cloud you install it from the community node registry. Choose **Run an Actor and get dataset**, set Actor to `datagleaner/sitemap-extractor` and paste the input above.
- **Make:** use the Apify app's **Run an Actor** module, then **Get Dataset Items** to read the results. **Watch Actor Runs** can trigger a scenario when a run finishes.
- **Zapier:** use the Apify action **Run Actor**, then the search **Fetch dataset items**. The trigger **Finished Actor run** starts a Zap when a run ends.
- **AI agents (MCP):** connect to `https://mcp.apify.com?tools=datagleaner/sitemap-extractor`. In Claude Code:

  ```
  claude mcp add --transport http apify "https://mcp.apify.com?tools=datagleaner/sitemap-extractor"
  ```

  Then run `/mcp` to sign in to Apify in your browser. Then ask in plain words, for example: "Use datagleaner/sitemap-extractor to list every /blog/ URL on stripe.com, capped at 500, and tell me which ones were updated in the last 30 days."
- **LangChain (Python):**

```python
## pip install langchain-apify, then set APIFY_TOKEN in your environment
import json
from langchain_apify import ApifyActorsTool
tool = ApifyActorsTool("datagleaner/sitemap-extractor")
result = tool.invoke({"run_input": json.loads('{"websites": ["stripe.com"], "maxUrlsPerSite": 10}')})
```

### Limits

- **Only what the site publishes.** A page missing from every sitemap is not found. Sites without any sitemap return nothing; the log and `SUMMARY` say so.
- **Per-site cap.** `maxUrlsPerSite` stops a site early. Very large sites (millions of URLs) need a higher cap or patterns that narrow them.
- **Blocking.** Sites behind bot protection may answer 403 to datacenter IPs. Enable `proxyConfiguration` or raise `requestDelaySeconds`.
- Index nesting is followed up to `maxCrawlingDepth` levels (default 6) and 3,000 sitemap files per site, as a safety stop.
- **Status message.** It reports totals, sites stopped at `maxUrlsPerSite`, sites with no URLs, how many sitemap files were "empty or unreadable", and with `checkStatus` how many pages were "not 2xx".

### FAQ

**How do I extract all URLs from a website?** Enter the domain in `websites`. The Actor reads every sitemap the site publishes and returns each page URL as a dataset row. Pages left out of every sitemap are not found, because it does not crawl links.

**How do I count a site's sitemap URLs?** Run it with a `maxUrlsPerSite` above the site's size. The `urls` field for each site in the `SUMMARY` record, and the dataset's item count, give the total.

**Does it need a browser or login?** No, plain HTTP only.

**What if a site's sitemap is broken?** The Actor recovers the URLs it can read from damaged XML, skips files it cannot read and continues.

**Are `.xml.gz` sitemaps supported?** Yes, detected by content, so a gzip file with a wrong name or header still works.

**Why fewer URLs than `maxUrlsPerSite`?** The site lists fewer, your patterns filtered some out, or duplicates were removed.

**Is it a sitemap scraper or a crawler?** A sitemap scraper: it reads the sitemap files a site publishes and does not crawl links, so it is fast and cheap, but it only finds pages the site lists.

**Can I export the sitemap URLs to CSV or Excel?** Yes. Open the run's dataset and export as CSV, Excel, JSON or XML, or fetch it from the API.

**Can I run it with no input?** Yes. It runs a built-in example on aioseo.com, capped at 5 URLs, and finishes in well under a minute.

### Responsible use

This Actor reads only sitemap files that sites publish for crawlers, and its default pacing is polite. Respect each site's terms and `robots.txt` rules when you use the URLs afterwards. URLs can contain personal data in rare cases (for example a person's name in a profile path): **you** are responsible for a lawful basis and your obligations under GDPR and any other law that applies. This is not legal advice.

# Actor input Schema

## `websites` (type: `array`):

Website URLs or bare domains, for example https://aioseo.com or stripe.com. Only the site root is used; any path is ignored. A direct sitemap URL (ending in .xml or .xml.gz) is also accepted.

## `maxUrlsPerSite` (type: `integer`):

Upper bound on page URLs returned for each website. Each returned URL is billed, so this is also your cost cap per site. The status message names every site that had more.

## `checkStatus` (type: `boolean`):

Sends one HEAD request per URL (GET if HEAD is refused) and adds status, finalUrl (after redirects) and checkError, so you can find 404s, 5xx errors and redirects listed in the sitemap. When the sitemap has no lastmod, it is filled from the page's Last-Modified header (lastmodSource: http-header). Slower: about 5 pages a second.

## `includeUrlPatterns` (type: `array`):

Optional regular expressions (Python syntax). A URL is kept only if it matches at least one. Example: /blog/

## `excludeUrlPatterns` (type: `array`):

Optional regular expressions. A URL matching any of them is dropped. Example: .pdf$

## `requestDelaySeconds` (type: `number`):

Base pause between requests to the same site; the actual pause is randomised up to 1.5x. Raise it if a site throttles you.

## `maxCrawlingDepth` (type: `integer`):

How many levels of sitemap indexes to follow. 1 = follow only the top-level index (its child sitemaps are read, but indexes nested inside it are not).

## `maxRequestRetries` (type: `integer`):

Retries per sitemap file on timeouts, 429 and 5xx errors (exponential backoff).

## `proxyConfiguration` (type: `object`):

Optional. Most sites work without a proxy; enable Apify Proxy if a site blocks datacenter IPs.

## `startUrls` (type: `array`):

Same as Websites, in the request-list format other sitemap Actors use, so their input works here unchanged. Used only when Websites is empty.

## Actor input object example

```json
{
  "websites": [
    "https://aioseo.com"
  ],
  "maxUrlsPerSite": 100,
  "checkStatus": false,
  "requestDelaySeconds": 0.3,
  "maxCrawlingDepth": 6,
  "maxRequestRetries": 3,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "https://aioseo.com"
    ],
    "maxUrlsPerSite": 100,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagleaner/sitemap-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": ["https://aioseo.com"],
    "maxUrlsPerSite": 100,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("datagleaner/sitemap-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "https://aioseo.com"
  ],
  "maxUrlsPerSite": 100,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call datagleaner/sitemap-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagleaner/sitemap-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bR213JK0T1PwsfuFp/builds/m7Occ8r2diwfeATbO/openapi.json
