# Sitemap URL Extractor & XML Sitemap Scraper (Index, Gzip) (`ventura_workalong/sitemap-url-extractor`) Actor

Extract every URL from a website's sitemap.xml: give a domain and it finds sitemaps via robots.txt, follows sitemap indexes, reads .gz, text and RSS sitemaps, and recovers broken XML. Returns lastmod, changefreq, priority, optional HTTP status. $0.20 per 1,000 URLs.

- **URL**: https://apify.com/ventura_workalong/sitemap-url-extractor.md
- **Developed by:** [Ventura WorkAlong](https://apify.com/ventura_workalong) (community)
- **Categories:** Developer tools, SEO tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.20 / 1,000 url extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor & XML Sitemap Scraper

**Sitemap URL Extractor** gets every page URL listed in a website's sitemap. Give it a domain and it finds the sitemaps for you. It reads the `Sitemap:` lines in robots.txt, follows sitemap indexes to the end, opens `.xml.gz` files, and parses plain-text and RSS/Atom sitemaps. It also recovers URLs from broken XML. The output is a flat list of URLs with `lastmod`, `changefreq`, `priority` and the sitemap each came from, ready for a crawler, an SEO audit, a spreadsheet or an AI agent.

- **Reliable by design:** one broken or missing child sitemap never fails the run. Every other sitemap is still read, and the failure is reported with its reason.
- **Handles real-world sitemaps:** sitemap indexes (nested, with loop protection), gzip, text and RSS/Atom sitemaps, BOMs, CDATA, namespace prefixes, unescaped `&` and truncated files.
- **Finds sitemaps automatically:** robots.txt first, then `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml` and other common locations, including WordPress installs in a subfolder. HTML "soft 404" pages are never mistaken for sitemaps.
- **Filters:** regex include/exclude and a `lastmod` date range ("pages changed since last week").
- **Optional HTTP status check** per URL to find broken pages and redirects listed in your sitemap.
- **$0.20 per 1,000 URLs.** Sites without a sitemap are free.

### How to extract all URLs from a sitemap

1. Add one or more websites to **Websites or sitemap URLs**. Use a domain (`example.com`), any page URL, a `robots.txt` URL or a direct sitemap URL (`https://example.com/sitemap_index.xml`).
2. Optional: set **Max URLs per site** (0 = all), URL patterns or a last-modified date range.
3. Click **Start**. Download the results as JSON, CSV or Excel, or fetch them through the Apify API.

### Input example

```json
{
  "startUrls": ["https://crawlee.dev", "https://www.gov.uk/sitemap.xml"],
  "maxUrlsPerSite": 0,
  "includeUrlPatterns": ["/blog/"],
  "lastmodFrom": "2026-09-01"
}
```

| Field | What it does | Default |
|---|---|---|
| `startUrls` | Domains, page URLs, robots.txt URLs or sitemap URLs | required |
| `maxUrlsPerSite` | Stop after this many unique URLs per input (0 = all) | 0 |
| `includeUrlPatterns` / `excludeUrlPatterns` | Regular expressions matched against each URL | none |
| `lastmodFrom` / `lastmodTo` | Keep URLs whose `lastmod` is in this range. URLs without a lastmod are dropped while a date filter is set | none |
| `includeImages` | Add image URLs from image sitemaps | false |
| `includeAlternates` | Add hreflang alternates (multilingual sites) | false |
| `checkStatus` | HEAD request per URL, returning `httpStatus` and `redirectTo` | false |
| `maxStatusChecks` | Cap on status checks per run | 1000 |
| `maxSitemapsPerSite` | Safety cap on sitemap files per site | 1000 |
| `respectRobotsTxt` | Honor robots.txt rules and Crawl-delay | true |
| `maxRequestsPerSecondPerHost` | Politeness limit per site (0.2–5) | 1 |
| `maxConcurrency` | Sites processed in parallel | 5 |

### Output example

One dataset item per unique URL. This is real output from `https://crawlee.dev` (2026-10-07):

```json
{
  "type": "url",
  "url": "https://crawlee.dev/blog",
  "host": "crawlee.dev",
  "path": "/blog",
  "lastmod": null,
  "changefreq": "weekly",
  "priority": 0.5,
  "sitemapUrl": "https://crawlee.dev/sitemap.xml",
  "input": "https://crawlee.dev"
}
```

- `lastmod` is normalized: dates stay `YYYY-MM-DD`, and datetimes become UTC ISO 8601 (`2026-10-07T09:36:29Z`).
- With `checkStatus` on, each item also gets `httpStatus` (e.g. `200`, `301`, `404`) and `redirectTo`.
- An input with no readable sitemap gets one free row with `"type": "error"` that says why:

```json
{ "type": "error", "input": "https://www.python.org", "url": null,
  "error": "No sitemap found: robots.txt lists none and none of /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /wp-sitemap.xml, /sitemap.xml.gz, /sitemap.txt exist." }
```

- The key-value store record **SUMMARY** has a per-site report: how sitemaps were discovered, every sitemap file read, each one that failed and why, duplicates removed and warnings. Example for MDN, whose sitemaps are gzipped:

```json
{ "input": "https://developer.mozilla.org", "discovery": "robots.txt",
  "sitemapsFound": ["https://developer.mozilla.org/sitemap.xml", "https://developer.mozilla.org/sitemaps/en-us/sitemap.xml.gz"],
  "sitemapsProcessed": 2, "sitemapsFailed": [], "urlsFound": 301, "duplicates": 1, "urlsOutput": 300,
  "stoppedEarly": "maxUrlsPerSite (300) reached" }
```

### Pricing

Pay per event, with no platform usage charges on top:

| Event | Price |
|---|---|
| URL extracted (`url-extracted`) | $0.0002 per URL ($0.20 per 1,000) |
| URL status checked (`url-status-checked`, only with `checkStatus`) | $0.0005 per URL |

Examples: a 5,000-page site costs $1.00. A 200-page site with status checks costs $0.04 + $0.10 = $0.14. Inputs without a sitemap cost nothing. Each URL is charged once, even if several sitemaps list it. The run stops cleanly at your maximum total charge.

### Use cases

- **SEO audits:** list every indexable URL, find sitemap URLs that redirect or 404, and check `lastmod` freshness.
- **Crawling and scraping:** feed the URL list into Website Content Crawler or your own scraper instead of crawling links.
- **Competitor and content monitoring:** run on a schedule with `lastmodFrom` to see new or updated pages.
- **Site migrations:** export the old site's URL inventory before a redirect project.
- **AI agents (MCP):** "list all blog posts on example.com" in one call.

### How it works (and how it stays polite)

- Requests are identified as `SitemapExtractorBot`. robots.txt is honored by default, including `Crawl-delay`, and each site gets at most 1 request per second unless you raise it.
- Transient errors (HTTP 429/5xx, timeouts) are retried with backoff, and `Retry-After` is honored. If a site keeps rate-limiting, the Actor stops asking and reports it.
- Files are capped at 60 MB even after decompression, so gzip bombs are safe. Private and internal network addresses are refused.
- Only public sitemap files are read. No login, no proxies, no personal data.

### Limits

- Only sitemaps the site publishes are read. The Actor doesn't crawl HTML links, so pages missing from the sitemap won't appear.
- Sitemaps that need JavaScript or a login can't be read. If robots.txt blocks bots from the sitemap, that sitemap is skipped (you'll see the reason).
- News and video sitemap extensions are read as normal URLs; their extra fields aren't returned.
- The status check reports the first response (it doesn't follow redirects) and uses HEAD, falling back to GET when HEAD isn't allowed.

### FAQ

#### Do I need the exact sitemap URL?

No. Give the domain and the Actor discovers the sitemaps. A direct sitemap URL also works and skips discovery.

#### Can it read sitemap index files with thousands of sitemaps?

Yes. It follows indexes breadth-first, skips loops and duplicates, and stops at `maxSitemapsPerSite` (default 1,000; up to 50,000).

#### Why is `lastmod` empty for some URLs?

That site's sitemap doesn't publish it. The Actor never invents dates.

#### Is this the same as Apify's Sitemap Extractor?

No. This is an independent Actor focused on discovery, malformed-file recovery, filters and a transparent per-site report.

### Related Actors

- [PageSpeed Insights Bulk Checker](https://apify.com/ventura_workalong/pagespeed-insights-bulk): audit Core Web Vitals for every page in a sitemap.
- [Tech Stack Detector](https://apify.com/ventura_workalong/tech-stack-detector): find the CMS, frameworks and analytics behind a site.
- [WHOIS Domain Lookup with DNS, SSL & Security Headers](https://apify.com/ventura_workalong/domain-lookup-bundle).
- [PDF to Markdown & Tables](https://apify.com/ventura_workalong/doc-to-markdown-tables): turn documents found in a sitemap into LLM-ready text.

Found a sitemap this Actor can't read? Open an issue with the URL. Fixing those is the point of this Actor.

# Actor input Schema

## `startUrls` (type: `array`):

Domains (`example.com`), page URLs, `robots.txt` URLs or direct sitemap URLs (`https://example.com/sitemap_index.xml`, `.xml.gz`, `.txt`). For a domain, sitemaps are discovered from robots.txt, then from common locations like /sitemap.xml.

## `maxUrlsPerSite` (type: `integer`):

Stop after this many unique URLs per input. 0 = no limit (all URLs). The prefilled 100 is for a quick try.

## `includeUrlPatterns` (type: `array`):

Keep only URLs matching at least one of these regular expressions (case-insensitive), e.g. `/blog/` or `/products/.+`.

## `excludeUrlPatterns` (type: `array`):

Drop URLs matching any of these regular expressions, e.g. `/tag/` or `\?page=`.

## `lastmodFrom` (type: `string`):

Only URLs whose sitemap `lastmod` is on or after this date (YYYY-MM-DD or ISO 8601). URLs without a lastmod are excluded when a date filter is set.

## `lastmodTo` (type: `string`):

Only URLs whose sitemap `lastmod` is on or before this date. URLs without a lastmod are excluded when a date filter is set.

## `includeImages` (type: `boolean`):

Add the `images` list from image sitemap extensions (`image:loc`).

## `includeAlternates` (type: `boolean`):

Add the `alternates` list (`xhtml:link rel=alternate hreflang`) for multilingual sites.

## `checkStatus` (type: `boolean`):

Send one HEAD request per URL and return `httpStatus` (redirects are reported, not followed, with `redirectTo`). Slower: requests to one site are rate-limited. Charged per checked URL.

## `maxStatusChecks` (type: `integer`):

Upper limit on HTTP status checks in this run (only with status checking on).

## `maxSitemapsPerSite` (type: `integer`):

Safety cap on sitemap files read per input (sitemap indexes can list thousands).

## `respectRobotsTxt` (type: `boolean`):

Skip sitemaps and pages that robots.txt disallows for this bot, and honor Crawl-delay.

## `maxRequestsPerSecondPerHost` (type: `number`):

Politeness limit for each host (0.2 to 5). Robots.txt Crawl-delay overrides it when slower.

## `maxConcurrency` (type: `integer`):

How many inputs are processed at the same time.

## Actor input object example

```json
{
  "startUrls": [
    "https://crawlee.dev"
  ],
  "maxUrlsPerSite": 100,
  "includeUrlPatterns": [],
  "excludeUrlPatterns": [],
  "includeImages": false,
  "includeAlternates": false,
  "checkStatus": false,
  "maxStatusChecks": 1000,
  "maxSitemapsPerSite": 1000,
  "respectRobotsTxt": true,
  "maxRequestsPerSecondPerHost": 1,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

One item per unique page URL found in the sitemaps (plus a free error row for any input without a readable sitemap).

## `summary` (type: `string`):

Per-input report: how sitemaps were discovered, sitemaps read and failed (with reasons), URL counts, filters and warnings.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://crawlee.dev"
    ],
    "maxUrlsPerSite": 100
};

// Run the Actor and wait for it to finish
const run = await client.actor("ventura_workalong/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["https://crawlee.dev"],
    "maxUrlsPerSite": 100,
}

# Run the Actor and wait for it to finish
run = client.actor("ventura_workalong/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://crawlee.dev"
  ],
  "maxUrlsPerSite": 100
}' |
apify call ventura_workalong/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ventura_workalong/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Ngvvwa4ZOxolqmgc3/builds/NwA9ikEhXkmjvJKrB/openapi.json
