# Sitemap Extractor Pro: URLs, lastmod, images, diff (`spongy_frame/sitemap-extractor-pro`) Actor

Extract every URL from XML, gzipped, index and text sitemaps with lastmod, priority, images, news and hreflang; filter by glob or date and diff against the previous run.

- **URL**: https://apify.com/spongy\_frame/sitemap-extractor-pro.md
- **Developed by:** [Spongy Frame Tools](https://apify.com/spongy_frame) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.15 / 1,000 sitemap urls

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Sitemap Extractor Pro do?

Sitemap Extractor Pro reads XML sitemaps and returns every URL they list as clean, structured data. Give it a sitemap, a sitemap index, a gzipped sitemap, a robots.txt, or just a homepage, and it will:

- Follow sitemap indexes recursively (nested indexes up to a configurable depth, with cycle detection).
- Decompress gzipped sitemaps, detected by content rather than by the `.gz` extension.
- Read plain-text sitemaps (one URL per line).
- Discover sitemaps from a homepage by checking `/robots.txt` `Sitemap:` lines and the common paths `/sitemap.xml`, `/sitemap_index.xml` and `/sitemap-index.xml`.
- Extract `lastmod`, `changefreq`, `priority`, image URLs, video count, Google News title and publication date, and `hreflang` alternates from each `<url>` entry.
- Filter URLs by glob pattern and by last-modified date range.
- Remove duplicates across sitemaps (the first occurrence wins).
- Optionally compare the URL set with your previous run and report which URLs were added or removed.

It never visits the pages themselves. It only downloads sitemap files, so runs are fast and cheap even for sites with hundreds of thousands of URLs.

### Why use this one?

Many sitemap tools either stop at the first index level, choke on gzipped or malformed files, or return a flat list of URLs with no metadata. This Actor is built for the messy reality of real sitemaps:

- Streaming XML parsing keeps memory flat on very large files, with a hard 50 MB per-file limit so a single giant sitemap cannot crash the run.
- Malformed XML is parsed in recovery mode; entries that can be salvaged are kept and the file is listed in the run summary as partially parsed.
- A bad URL never fails the run. Fetch errors, 404s, unknown formats and oversized files are recorded under `failures` in the `RUN_SUMMARY` key-value record and the crawl continues.
- Pricing is per URL, so you pay for what you get, and `maxItems` caps the spend up front.

What it does not do: it does not crawl pages, does not check whether URLs return 200, and does not render JavaScript. If a site has no sitemap at all, the result is empty.

### How to use it

1. Enter one or more start URLs. A sitemap URL is best, but a homepage or robots.txt URL works too.
2. Optionally set `maxItems`, glob or date filters, and turn on `diffAgainstPreviousRun` if you want change detection.
3. Run the Actor and download the dataset as JSON, CSV or Excel, or read it through the API.

### Input

```json
{
  "startUrls": ["https://www.apify.com/sitemap.xml"],
  "maxItems": 500,
  "outputFields": "full",
  "includeGlobs": ["https://www.apify.com/*"],
  "excludeGlobs": ["*.pdf"],
  "modifiedAfter": "2026-01-01",
  "maxDepth": 5,
  "followSitemapsFromRobots": true,
  "respectRobotsDisallow": false,
  "diffAgainstPreviousRun": false
}
```

Only `startUrls` is required. Globs use shell-style wildcards (`*`, `?`, `[abc]`) matched against the whole URL. Date filters accept ISO 8601 dates or datetimes; URLs without a `lastmod` are dropped when a date filter is active because their modification date is unknown.

### Output

With `outputFields: "full"` every dataset item looks like this:

```json
{
  "url": "https://example.com/news/story",
  "lastmod": "2026-09-20T08:30:00+00:00",
  "changefreq": "daily",
  "priority": 0.8,
  "sitemap_url": "https://example.com/sitemap-news.xml",
  "source_url": "https://example.com/sitemap_index.xml",
  "images": ["https://cdn.example.com/img/1.jpg"],
  "videos_count": 1,
  "news": {"title": "Big story", "publication_date": "2026-09-20T08:00:00+00:00"},
  "alternates": [{"hreflang": "de", "href": "https://example.com/de/news/story"}],
  "depth": 1
}
```

With `outputFields: "urlsOnly"` items contain just `url` and `sitemap_url`. Dates are normalised to ISO 8601 with a timezone; missing values are `null`. `depth` is 0 for URLs found directly in a start sitemap and increases by one for each index level.

When `diffAgainstPreviousRun` is on, one extra item with `record_type: "diff_report"` is appended containing `added`, `removed`, `added_count`, `removed_count`, `unchanged_count` and `previous_run_found`. A summary of every run (counts, failures, warnings, elapsed time) is saved as `RUN_SUMMARY` in the default key-value store.

### How much does it cost?

The Actor uses pay-per-event pricing. There is no subscription and no charge for failed fetches.

| Event | When it is charged | Price |
| --- | --- | --- |
| `sitemap-url` | Once per URL written to the dataset | $0.0002 |
| `diff-report` | Once per diff report, only when a previous run existed to compare against | $0.005 |

Examples:

- A 2,000-URL site with `maxItems: 1000` produces 1,000 items: 1,000 × $0.0002 = **$0.20**.
- Weekly change monitoring of a 5,000-URL site with `diffAgainstPreviousRun: true`: 5,000 × $0.0002 + $0.005 = **$1.005** per run. The very first run stores the baseline and costs $1.00.

Platform compute is included in the event prices. If your account's spending limit is lower than the cost of `maxItems`, the Actor lowers `maxItems` to fit and says so in the log.

### Limits and fair use

- Sitemap files larger than 50 MB (after decompression) are skipped and reported. Split such sitemaps or point the Actor at the child sitemaps directly.
- `maxDepth` defaults to 5 nested index levels, which covers every real-world site we have seen.
- Requests use a descriptive User-Agent, a 30-second timeout, three retries with exponential backoff, and honour `Retry-After` on 429 and 503 responses. Please do not hammer small sites with repeated large runs.
- Duplicate URLs are removed per run. Cross-run deduplication is available through the diff feature.

### Legal note

The Actor only reads publicly served sitemap files, which site owners publish specifically so that they can be read by automated tools. It does not visit pages, does not collect personal data, and outputs only URLs and the metadata declared in the sitemap. You are responsible for how you use the resulting URL lists.

### FAQ

**Why did my run stop early?** Either `maxItems` was reached, or your Apify pay-per-event spending limit for this run was hit. Both cases are logged and recorded in `RUN_SUMMARY` under `limit_reason`. Raise the limit and run again; the diff feature will pick up where the previous set ended.

**The dataset is empty. What happened?** Check `RUN_SUMMARY.failures`. Common causes: the site has no sitemap at the usual locations, the sitemap URL returned an HTML page (some sites redirect unknown paths to the homepage), or a date filter removed every URL because the sitemap has no `lastmod` values.

**Can I get only new URLs since last time?** Yes. Turn on `diffAgainstPreviousRun` and use the same `startUrls` each time. The `diff_report` record lists `added` and `removed` URLs. Note that if a run is truncated by `maxItems`, the comparison is against a partial set, which the report flags with `current_run_truncated`.

**Does it respect robots.txt Disallow rules?** It reads robots.txt to find sitemaps. With `respectRobotsDisallow` enabled it adds a `robots_disallowed` flag to each item, but nothing is filtered, because the Actor never requests the pages themselves.

**Can I filter by path or file type?** Use `includeGlobs` and `excludeGlobs`, for example `https://example.com/blog/*` or `*.pdf`. Patterns match the full URL.

### Support

Found a bug or missing a feature? Open a ticket in the Issues tab of this Actor. We respond within 24 hours on working days.

# Actor input Schema

## `startUrls` (type: `array`):

Sitemap URLs to read. Each can be a sitemap.xml, a sitemap index, a gzipped sitemap (.xml.gz), a robots.txt URL, or a plain homepage (the Actor then looks for sitemaps in /robots.txt and at /sitemap.xml, /sitemap\_index.xml, /sitemap-index.xml).

## `maxItems` (type: `integer`):

Stop after this many URLs have been written to the dataset. Each URL is one charged 'sitemap-url' event, so this caps your spend.

## `outputFields` (type: `string`):

'full' writes url, lastmod, changefreq, priority, images, news, hreflang alternates and more. 'urlsOnly' writes just url and sitemap\_url (smaller dataset, same price per URL).

## `includeGlobs` (type: `array`):

Keep only URLs matching at least one of these glob patterns (fnmatch style, matched against the whole URL), e.g. https://example.com/blog/\*. Leave empty to keep everything.

## `excludeGlobs` (type: `array`):

Drop URLs matching any of these glob patterns, e.g. \*.pdf or */tag/*.

## `modifiedAfter` (type: `string`):

Keep only URLs whose <lastmod> is on or after this ISO 8601 date or datetime (e.g. 2026-01-01 or 2026-01-01T00:00:00Z). URLs without a lastmod are dropped when this is set.

## `modifiedBefore` (type: `string`):

Keep only URLs whose <lastmod> is on or before this ISO 8601 date or datetime. URLs without a lastmod are dropped when this is set.

## `maxDepth` (type: `integer`):

How many levels of nested sitemap indexes to follow. The start URL is depth 0.

## `followSitemapsFromRobots` (type: `boolean`):

When a start URL is a robots.txt or a homepage, follow the sitemaps it lists.

## `respectRobotsDisallow` (type: `boolean`):

Adds a robots\_disallowed field to each full item indicating whether the URL matches a Disallow rule for User-agent: \*. Nothing is filtered; the Actor only reads sitemaps, it never visits the pages.

## `diffAgainstPreviousRun` (type: `boolean`):

Remember the URL set in a named key-value store (sitemap-extractor-pro-state) and add one diff\_report record listing URLs added and removed since the last run with the same start URLs. The first run stores a baseline for free; later diff reports are charged as 'diff-report'.

## Actor input object example

```json
{
  "startUrls": [
    "https://www.apify.com/sitemap.xml"
  ],
  "maxItems": 500,
  "outputFields": "full",
  "maxDepth": 5,
  "followSitemapsFromRobots": true,
  "respectRobotsDisallow": false,
  "diffAgainstPreviousRun": false
}
```

# Actor output Schema

## `dataset` (type: `string`):

One record per URL (url, lastmod, changefreq, priority, sitemap\_url, images, alternates) plus optional diff\_report records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://www.apify.com/sitemap.xml"
    ],
    "maxItems": 500
};

// Run the Actor and wait for it to finish
const run = await client.actor("spongy_frame/sitemap-extractor-pro").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["https://www.apify.com/sitemap.xml"],
    "maxItems": 500,
}

# Run the Actor and wait for it to finish
run = client.actor("spongy_frame/sitemap-extractor-pro").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://www.apify.com/sitemap.xml"
  ],
  "maxItems": 500
}' |
apify call spongy_frame/sitemap-extractor-pro --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,spongy_frame/sitemap-extractor-pro"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/OGuuTgX00eJcdxzmQ/builds/IMn6W6DBb6iFZoujh/openapi.json
