# Sitemap URL Extractor: All Page URLs from XML Sitemaps (`pistachio_implementation/sitemap-url-extractor`) Actor

Extract every page URL from any website's sitemaps. Finds sitemaps in robots.txt, follows sitemap indexes and gzip files, and returns lastmod, priority, hreflang alternates, image counts and optional HTTP status. Filter by pattern or date.

- **URL**: https://apify.com/pistachio\_implementation/sitemap-url-extractor.md
- **Developed by:** [Hay Equipos](https://apify.com/pistachio_implementation) (community)
- **Categories:** SEO tools, Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor: all page URLs from XML sitemaps

Get a complete list of a website's pages in seconds. Enter a domain or a sitemap link and the actor finds the sitemaps (robots.txt first, then the usual paths), follows every sitemap index, opens gzip files, and returns one row per URL with last modified date, change frequency, priority, hreflang alternates, image and video counts, and Google News fields. Turn on the status check to find broken or redirected pages that are still listed in the sitemap.

No browser and no crawling of the pages themselves, so it is fast and cheap: a sitemap with 50,000 URLs is read in a few seconds.

### Who uses it

- **SEO specialists** auditing sitemaps, finding 404s and redirects, and checking hreflang coverage.
- **Content teams and agencies** listing every blog post or product page of a competitor, with dates.
- **Developers and AI builders** who need the list of pages to feed a crawler, a RAG pipeline or a change monitor.
- **Migration projects** that need a before and after list of URLs.

### Input

| Field | What it does |
|---|---|
| `startUrls` | Websites (`example.com`) or sitemap links (`.xml`, `.xml.gz`, `.txt`, sitemap index, robots.txt, RSS or Atom feed) |
| `includeUrlPatterns` | Keep only URLs matching any of these regular expressions or plain words, for example `/blog/` |
| `excludeUrlPatterns` | Drop URLs matching any of these, for example `/tag/` |
| `lastmodAfter` | Keep only URLs changed on or after a date, for example `2026-09-01` |
| `sameHostOnly` | Drop URLs on other hosts |
| `checkStatus` | Add HTTP status, final URL and a redirected flag for every URL |
| `maxSitemapsPerSite`, `maxUrlsPerSite`, `maxUrls` | Caps to control run time and cost |

#### Example input

```json
{
    "startUrls": ["https://blog.cloudflare.com", "https://www.nasa.gov/sitemap.xml"],
    "includeUrlPatterns": [],
    "excludeUrlPatterns": ["/tag/"],
    "lastmodAfter": "2026-01-01",
    "checkStatus": false,
    "maxUrlsPerSite": 50000
}
```

### Output

One row per unique URL:

```json
{
    "url": "https://blog.cloudflare.com/fr-fr/cloudflare-turns-8",
    "lastmod": "2026-07-15T16:05:05.825Z",
    "changefreq": null,
    "priority": null,
    "alternates": [
        { "hreflang": "en-us", "url": "https://blog.cloudflare.com/cloudflare-turns-8" },
        { "hreflang": "de-de", "url": "https://blog.cloudflare.com/de-de/cloudflare-turns-8" }
    ],
    "imageCount": 0,
    "videoCount": 0,
    "newsTitle": null,
    "newsPublishedAt": null,
    "sitemapUrl": "https://blog.cloudflare.com/sitemap-posts.xml",
    "site": "https://blog.cloudflare.com",
    "statusCode": 200,
    "finalUrl": "https://blog.cloudflare.com/fr-fr/cloudflare-turns-8/",
    "redirected": true,
    "scrapedAt": "2026-09-27T06:14:24.933Z"
}
```

`statusCode`, `finalUrl` and `redirected` appear only with the status check on. A `RUN_SUMMARY` record in the key value store shows, for each site, how the sitemap was found, how many sitemap files were read, any sitemap errors, and how many URLs were saved.

Export as JSON, CSV, Excel or HTML, or use the Apify API, webhooks, Make, Zapier or the Apify MCP server.

### Pricing

Pay per event, no subscription and no platform usage charges on top:

- **$0.25 per 1,000 URLs** saved ($0.00025 per URL).
- **$1.00 per 1,000 status checks**, only when you turn the status check on.
- Filtered URLs, duplicates and sites without a sitemap cost nothing.

### Limits

- The actor lists what the sitemaps say. Pages missing from the sitemap are not found, because pages are not crawled.
- Sites that block automated requests to their sitemap (some sites behind strict bot protection) return an error in `RUN_SUMMARY`.
- `lastmod` is whatever the site writes; some sites set every URL to today's date.
- With `lastmodAfter` set, URLs without a lastmod are dropped, and sitemap files whose own lastmod is older than the date are skipped.
- The status check sends a HEAD request (a GET if HEAD is refused) at a polite pace of about three requests a second per site, so it is much slower than listing.

### FAQ

**Where does it look for sitemaps?** First every `Sitemap:` line in robots.txt, then `/sitemap.xml`, `/sitemap_index.xml`, `/sitemap-index.xml`, `/wp-sitemap.xml` and `/sitemap.txt`. You can also give sitemap links directly.

**Does it handle huge sites?** Yes. Sitemap indexes with hundreds of files are followed breadth first. Use `maxSitemapsPerSite` and `maxUrlsPerSite` to cap cost.

**Can I find broken links in my sitemap?** Yes, turn on `checkStatus` and filter the results for `statusCode` 404 or `redirected` true.

**Is it allowed?** Sitemaps and robots.txt exist so that machines can read them. The actor does not log in and keeps a polite pace.

# Actor input Schema

## `startUrls` (type: `array`):

One per line. A website (example.com) makes the actor look in robots.txt and then the usual sitemap paths. A sitemap URL (https://example.com/sitemap.xml, .xml.gz, .txt, or a sitemap index) is read directly.

## `includeUrlPatterns` (type: `array`):

Regular expressions or plain text, for example /blog/ or .html$. Leave empty to keep every URL.

## `excludeUrlPatterns` (type: `array`):

Regular expressions or plain text, for example /tag/ or ?page=.

## `lastmodAfter` (type: `string`):

Keep only URLs whose lastmod is on or after this date, for example 2026-09-01. URLs without lastmod are dropped when this is set.

## `sameHostOnly` (type: `boolean`):

Drop URLs on other hosts than the website you entered.

## `checkStatus` (type: `boolean`):

Sends a HEAD request to each saved URL and adds status code, final URL and a redirected flag. Finds broken and redirected pages in your sitemap. Charged as an extra event.

## `maxSitemapsPerSite` (type: `integer`):

Large sites split their sitemap into many files. The actor stops reading new files for a site after this many.

## `maxUrlsPerSite` (type: `integer`):

Stop saving URLs for a site after this many.

## `maxUrls` (type: `integer`):

The run stops after saving this many URLs.

## Actor input object example

```json
{
  "startUrls": [
    "https://apify.com"
  ],
  "includeUrlPatterns": [],
  "excludeUrlPatterns": [],
  "sameHostOnly": false,
  "checkStatus": false,
  "maxSitemapsPerSite": 500,
  "maxUrlsPerSite": 50000,
  "maxUrls": 200000
}
```

# Actor output Schema

## `results` (type: `string`):

All rows the run saved to the default dataset.

## `summary` (type: `string`):

The RUN\_SUMMARY record: counts and problems for the whole run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://apify.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pistachio_implementation/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": ["https://apify.com"] }

# Run the Actor and wait for it to finish
run = client.actor("pistachio_implementation/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://apify.com"
  ]
}' |
apify call pistachio_implementation/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pistachio_implementation/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/FjUvgVQLTKxcMLUXe/builds/Kcj9S8kdcsCJdl5Ec/openapi.json
