# Sitemap URL Extractor — SEO & Change Monitor (`zenomastro/sitemap-url-extractor-pro`) Actor

Extract URLs from XML and gzipped sitemaps, nested indexes and robots.txt discovery. Export lastmod, priority, hreflang, image/video data, filter URLs, audit status/redirects and page indexability, fall back to crawling, and compare persistent snapshots for URL changes.

- **URL**: https://apify.com/zenomastro/sitemap-url-extractor-pro.md
- **Developed by:** [Rosario Vitale](https://apify.com/zenomastro) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 sitemap urls

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor API — Indexability & Change Monitor

### Why use this Actor?

Extract URLs from XML and gzipped sitemaps, nested indexes and robots.txt discovery. Export lastmod, priority, hreflang, image/video data, filter URLs, audit status/redirects and page indexability, fall back to crawling, and compare persistent snapshots for URL changes.

### Features

- **Websites or sitemap URLs** — Public site homepages or direct .xml/.xml.gz sitemap URLs.
- **Maximum sitemap files** — Maximum sitemap files. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Maximum unique URLs** — Maximum unique URLs. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Include lastmod/changefreq/priority** — Include lastmod/changefreq/priority. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Extract image and video sitemap data** — Extract image and video sitemap data. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Extract hreflang alternates** — Extract hreflang alternates. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Include URL regex patterns** — When non-empty, only URLs matching at least one regular expression are emitted.
- **Exclude URL regex patterns** — URLs matching any expression are skipped.
- **Audit live HTTP status and redirects** — Audit live HTTP status and redirects. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Maximum URLs to status-check** — Maximum URLs to status-check. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Status-check concurrency** — Status-check concurrency. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Fallback crawl when no sitemap works** — Crawl same-origin HTML links only when a site yields no sitemap URLs.

### Use cases

- Seo inventory.
- Migration qa.
- Indexation monitoring.
- Url change detection.

### Example input

```json
{
  "sites": [
    "https://example.com"
  ],
  "maxSitemaps": 200,
  "maxUrls": 25000,
  "includeMetadata": true,
  "includeMedia": true,
  "includeHreflang": true
}
```

### Pricing & cost control

Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.

### FAQ

**What is this Actor for?**\
It is designed for SEO inventory, migration QA, indexation monitoring.

**Can I run it on a schedule?**\
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

**How do I control cost and run size?**\
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

### Search keywords

sitemap url extractor, sitemap scraper, sitemap scraper python, sitemap web scraper, sitemap url scraper, website sitemap scraper, xml sitemap scraper, sitemap json web scraper, web scraper sitemap wizard, what is a sitemap and why is it important, how do i find the sitemap of a website, sitemap url extractor extension, sitemap url extractor tool, sitemap url extractor chrome extension

Turn a sitemap into a repeatable SEO inventory, health check, and change-monitoring feed.

This Actor accepts website homepages, direct XML sitemaps, sitemap indexes, and compressed `.xml.gz` files. It discovers sitemap locations from `robots.txt` plus common WordPress/generic paths, recursively follows indexes, deduplicates URLs, extracts sitemap metadata, images, videos and hreflang alternates, and can optionally audit live HTTP status/redirect chains.

The important difference from a basic sitemap extractor is **change detection**. A named persistent snapshot can be compared on every scheduled run so the Dataset reports URLs that were added, removed, redirected, changed status, or changed `lastmod`.

### Why use it

- XML sitemap and sitemap-index recursion.
- Real `.xml.gz` decompression.
- `robots.txt`, `/sitemap.xml`, `/sitemap_index.xml`, and `/wp-sitemap.xml` discovery.
- Image/video sitemap fields.
- hreflang alternate links.
- Include/exclude regular-expression filters.
- Optional live status, content type, redirect chain and final URL audit.
- Optional same-origin fallback crawl when no sitemap works.
- Persistent scheduled-run snapshots and URL diffs.
- Explicit limits, retries, timeouts, size caps, and concurrency controls.
- Free diagnostic, summary and change rows; billable units remain unique sitemap URL rows.

### Input example

```json
{
  "sites": ["https://example.com"],
  "maxUrls": 25000,
  "includeMetadata": true,
  "includeMedia": true,
  "includeHreflang": true,
  "excludeUrlPatterns": ["[?&]preview=", "/cart/"],
  "checkStatusCodes": true,
  "statusCheckLimit": 2000,
  "snapshotMode": "compare_update",
  "snapshotId": "example-production"
}
```

### URL rows

A normal URL row can contain:

```json
{
  "recordType": "url",
  "url": "https://example.com/products/a",
  "sitemapUrl": "https://example.com/sitemap-products.xml.gz",
  "rootSite": "https://example.com/",
  "lastmod": "2026-09-26",
  "images": [{"loc": "https://cdn.example.com/a.jpg"}],
  "hreflang": [{"hreflang": "it", "href": "https://example.com/it/products/a"}],
  "status": 200,
  "finalUrl": "https://example.com/products/a",
  "redirected": false,
  "responseTimeMs": 143
}
```

### Change rows

With `snapshotMode=compare_update`, the first run establishes a baseline. Later runs can emit:

- `added`
- `removed`
- `lastmod_changed`
- `status_changed`
- `redirect_changed`

A summary row reports counts for every change type, sitemap errors, gzip files, status errors, redirects, media references and fallback pages.

### Fallback crawl

Enable `fallbackCrawl` only when you want coverage for sites that expose no usable sitemap. The fallback remains same-origin and obeys a strict page limit. It is deliberately off by default because crawling HTML costs more than reading XML.

### Status auditing

`checkStatusCodes` performs bounded concurrent requests and records:

- HTTP status
- final URL
- redirect chain
- content type
- response time
- a clean per-URL error when a request fails

Use `statusCheckLimit` to bound network work on very large sitemaps.

### Snapshot storage

Snapshots live in a named Key-Value Store (default `sitemap-intelligence-snapshots`) so they persist across runs. Set a stable `snapshotId` when you schedule the same property repeatedly. `maxSnapshotUrls` bounds persistent snapshot size and the summary tells you when a snapshot was truncated.

### Reliability and limits

Malformed sitemap files are isolated into diagnostic rows rather than aborting an entire run. Inputs are capped, responses have size limits, gzip is decompressed explicitly, retries are bounded, duplicate URLs are removed, and regular expressions are validated before network work begins.

### Pricing

The primary value unit remains one unique sitemap URL emitted. Diagnostics, summaries, and change rows are not intentionally billed as sitemap URL events. Live status checks and fallback crawling consume additional platform resources, so use their limits according to the size of the site.

### Responsible use

Use this Actor on public URLs that you are allowed to access. Respect site policies, rate limits, and applicable law. Status auditing and fallback crawling should be configured conservatively on sites you do not control.

### Extended capabilities

- Discover robots.txt sitemaps, common paths, nested indexes, and .xml.gz files with image/video/hreflang metadata.
- Audit URL status, redirects, content type, and response time, with optional same-origin fallback crawling.
- Persist snapshots and optionally audit title, canonical, meta robots, X-Robots-Tag, and factual indexability signals.

# Actor input Schema

## `sites` (type: `array`):

Public site homepages or direct .xml/.xml.gz sitemap URLs.

## `maxSitemaps` (type: `integer`):

Maximum sitemap files. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `maxUrls` (type: `integer`):

Maximum unique URLs. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `includeMetadata` (type: `boolean`):

Include lastmod/changefreq/priority. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `includeMedia` (type: `boolean`):

Extract image and video sitemap data. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `includeHreflang` (type: `boolean`):

Extract hreflang alternates. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `includeUrlPatterns` (type: `array`):

When non-empty, only URLs matching at least one regular expression are emitted.

## `excludeUrlPatterns` (type: `array`):

URLs matching any expression are skipped.

## `checkStatusCodes` (type: `boolean`):

Audit live HTTP status and redirects. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `statusCheckLimit` (type: `integer`):

Maximum URLs to status-check. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `statusConcurrency` (type: `integer`):

Status-check concurrency. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `fallbackCrawl` (type: `boolean`):

Crawl same-origin HTML links only when a site yields no sitemap URLs.

## `fallbackCrawlLimit` (type: `integer`):

Fallback crawl page limit. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `snapshotMode` (type: `string`):

Change detection. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `snapshotId` (type: `string`):

Use a stable ID when scheduling repeated audits. Otherwise one is derived from the input.

## `snapshotStoreName` (type: `string`):

Persistent snapshot store. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `maxSnapshotUrls` (type: `integer`):

Maximum URLs kept in a snapshot. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `emitChanges` (type: `boolean`):

Emit URL change rows. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `requestTimeoutSecs` (type: `integer`):

Request timeout. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `retries` (type: `integer`):

Request retries. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `maxSitemapSizeMb` (type: `integer`):

Maximum sitemap response size (MB). Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `auditIndexability` (type: `boolean`):

Fetch a bounded subset of sitemap URLs and record title, canonical, meta robots, X-Robots-Tag and a factual indexability state.

## `indexabilityAuditLimit` (type: `integer`):

Hard cap for HTML indexability audits.

## `indexabilityConcurrency` (type: `integer`):

Parallel page fetches used only for optional indexability auditing.

## Actor input object example

```json
{
  "sites": [
    "https://example.com"
  ],
  "maxSitemaps": 200,
  "maxUrls": 25000,
  "includeMetadata": true,
  "includeMedia": true,
  "includeHreflang": true,
  "includeUrlPatterns": [],
  "excludeUrlPatterns": [],
  "checkStatusCodes": false,
  "statusCheckLimit": 2000,
  "statusConcurrency": 12,
  "fallbackCrawl": false,
  "fallbackCrawlLimit": 100,
  "snapshotMode": "compare_update",
  "snapshotStoreName": "sitemap-intelligence-snapshots",
  "maxSnapshotUrls": 25000,
  "emitChanges": true,
  "requestTimeoutSecs": 25,
  "retries": 2,
  "maxSitemapSizeMb": 25,
  "auditIndexability": false,
  "indexabilityAuditLimit": 500,
  "indexabilityConcurrency": 8
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("zenomastro/sitemap-url-extractor-pro").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("zenomastro/sitemap-url-extractor-pro").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call zenomastro/sitemap-url-extractor-pro --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zenomastro/sitemap-url-extractor-pro"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/67G4t1WDT3dwitWK0/builds/f3qEfhcczVmMhQWbd/openapi.json
