# Sitemap URL Extractor Pro (`pinkish_gallop/sitemap-url-extractor-pro`) Actor

Extract every URL from any website's XML sitemaps: auto-discovery via robots.txt, nested indexes, .gz, images, hreflang, news, lastmod filters. Reliable, fast HTTP, $0.30 per 1,000 URLs.

- **URL**: https://apify.com/pinkish_gallop/sitemap-url-extractor-pro.md
- **Developed by:** [Sai](https://apify.com/pinkish_gallop) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.30 / 1,000 url extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor Pro

Get **every URL a website publishes in its XML sitemaps** as a clean table you can export to CSV, Excel, JSON or Google Sheets. Give it plain domains and it finds the sitemaps for you. Nested sitemap indexes, gzipped `.xml.gz` files, text sitemaps and RSS/Atom feeds are all handled.

Built for reliability. Sites where other extractors return nothing (sitemap only listed in robots.txt, wrong path given, gzip served without the right headers, `sm:` namespace prefixes, CDATA, BOMs, 429 rate limits) are exactly what this Actor was designed and tested for.

### What you get

| Field | Example |
|---|---|
| `url` | `https://wordpress.org/news/2026/09/...` |
| `lastmod` | `2026-09-30T12:00:00+00:00` |
| `changefreq` / `priority` | `weekly` / `0.8` |
| `site` | `https://wordpress.org` |
| `sitemap` | the sitemap file the URL came from |
| `images` | image URLs from `<image:image>` |
| `alternates` | hreflang alternates `[{hreflang, href}]` |
| `news` / `videos` | Google News and video sitemap fields |

A per-site **SUMMARY** record (key-value store) lists how each site's sitemaps were discovered, every sitemap file processed, any sitemap that failed and why, and the counts of duplicates and filtered URLs. You never have to guess why a site returned nothing.

### Features

- **Auto-discovery**: reads `Sitemap:` lines in `robots.txt`, then probes `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml`, `/sitemap.xml.gz`, `/sitemap.txt` and more. If a sitemap URL you give 404s, it falls back to discovery automatically.
- **Any depth of nested sitemap indexes**, processed in parallel (4 files per site, several sites at once).
- **gzip detection by content**, not just by file extension.
- **Tolerant parser**: namespaces and prefixes, CDATA, HTML entities, BOMs, text sitemaps, RSS and Atom.
- **Retries** with exponential backoff and `Retry-After` support on 429/5xx and network errors.
- **Filters**: include/exclude regex, "modified after" date (`lastmodAfter`), same-host only, and per-site and total caps.
- **Deduplicated** per site.
- No browser and no proxy needed. It's plain HTTP, so it's fast and cheap.

### Input example

```json
{
  "startUrls": ["wordpress.org", "https://www.nasa.gov/sitemap.xml", "apify.com"],
  "includePatterns": ["/news/", "/blog/"],
  "lastmodAfter": "2026-09-01",
  "maxUrlsPerSite": 50000
}
```

### Use cases

- SEO audits: compare sitemap URLs with crawled or indexed pages, and find orphan or stale pages.
- Seed lists for scrapers and crawlers (feed URLs into Website Content Crawler, Screenshot or Change Monitor actors).
- Content inventories and migrations.
- Competitor monitoring: what did they publish this week? (`lastmodAfter`)
- Building RAG / LLM datasets from a site's canonical pages.

### Pricing

Pay per event:

- **$0.30 per 1,000 URLs** ($0.0003 per unique URL written to the dataset)
- Discovery, robots.txt, sitemap downloads, duplicates, filtered-out URLs and errors are **free**
- Plus Apify's tiny actor start fee ($0.00005)

Use `maxUrlsTotal` to put a hard ceiling on a run's cost.

### FAQ

**A site returns 0 URLs. Why?** Check the `SUMMARY` record. Usually the site has no sitemap at all (`no_sitemap_found`), or it blocks automated requests (`sitemaps_unreachable` with e.g. `http_403`). You can try a different `userAgent`.

**Is there a limit?** No hard limit. Sites with millions of URLs work, but set `maxUrlsTotal` if you want to cap cost.

**Does it crawl pages?** No. It only reads sitemaps, which is why it's fast and cheap. Pair it with a crawler if you need page content.

### Related Actors

- [Website Change Monitor](https://apify.com/pinkish_gallop/website-change-monitor): get alerted when pages change
- [Website Screenshot Pro](https://apify.com/pinkish_gallop/website-screenshot-pro): bulk screenshots and PDFs
- [Website Contact & Tech Enricher](https://apify.com/pinkish_gallop/website-contact-enricher): emails, phones, socials and tech stack

# Actor input Schema

## `startUrls` (type: `array`):

Domains (example.com), any page URL, or direct sitemap URLs (sitemap.xml, sitemap_index.xml, .xml.gz, .txt, RSS). For plain domains the sitemaps are discovered automatically via robots.txt and common locations.

## `maxUrlsPerSite` (type: `integer`):

Stop after this many URLs per website (0 = no limit).

## `maxUrlsTotal` (type: `integer`):

Stop the whole run after this many URLs (0 = no limit). Useful to cap cost.

## `includePatterns` (type: `array`):

Only keep URLs matching at least one pattern, e.g. /blog/ or /products/. Case-insensitive.

## `excludePatterns` (type: `array`):

Drop URLs matching any of these patterns, e.g. /tag/ or ?page=

## `lastmodAfter` (type: `string`):

Keep only URLs whose <lastmod> is on/after this date (URLs without lastmod are dropped when set). Great for 'what was published this week'.

## `sameHostOnly` (type: `boolean`):

Drop URLs pointing to other hosts (www. is ignored).

## `discoverFromRobotsTxt` (type: `boolean`):

Read Sitemap: lines from /robots.txt for plain domains.

## `tryCommonLocations` (type: `boolean`):

If robots.txt lists none, probe /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml, /sitemap.xml.gz and others.

## `includeExtras` (type: `boolean`):

Add image URLs, hreflang alternates, Google News and video fields when the sitemap has them.

## `maxSitemapsPerSite` (type: `integer`):

Safety cap on nested sitemap files fetched per site.

## `maxConcurrency` (type: `integer`):

How many websites to process at the same time (each site also fetches up to 4 sitemap files in parallel).

## `timeoutSecs` (type: `integer`):

Per-request timeout. Failed requests are retried 3x with backoff.

## `userAgent` (type: `string`):

Override the User-Agent header if a site blocks the default one.

## Actor input object example

```json
{
  "startUrls": [
    "apify.com",
    "https://www.python.org/sitemap.xml"
  ],
  "maxUrlsPerSite": 0,
  "maxUrlsTotal": 0,
  "includePatterns": [],
  "excludePatterns": [],
  "sameHostOnly": false,
  "discoverFromRobotsTxt": true,
  "tryCommonLocations": true,
  "includeExtras": true,
  "maxSitemapsPerSite": 2000,
  "maxConcurrency": 5,
  "timeoutSecs": 30
}
```

# Actor output Schema

## `urls` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "apify.com",
        "https://www.python.org/sitemap.xml"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pinkish_gallop/sitemap-url-extractor-pro").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [
        "apify.com",
        "https://www.python.org/sitemap.xml",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("pinkish_gallop/sitemap-url-extractor-pro").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "apify.com",
    "https://www.python.org/sitemap.xml"
  ]
}' |
apify call pinkish_gallop/sitemap-url-extractor-pro --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pinkish_gallop/sitemap-url-extractor-pro"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XzeUGaZclE8fWJDfJ/builds/UpvnhbY9h6iYQqbS7/openapi.json
