# Sitemap URL Extractor + Change Monitor (robots.txt, gzip) (`osel_house/sitemap-extractor-monitor`) Actor

Every URL from a site's sitemaps: robots.txt discovery, sitemap index recursion, gzip, lastmod filter, image/video/news/hreflang fields. Or run it on a schedule and get only the URLs added, removed or changed since last time.

- **URL**: https://apify.com/osel\_house/sitemap-extractor-monitor.md
- **Developed by:** [Tenzin Phuntsok](https://apify.com/osel_house) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.20 / 1,000 url rows

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

Get **every URL a website lists in its sitemaps**, or only the URLs that were **added, removed or changed since your last run**. Give it a domain and it finds the sitemaps through robots.txt and the usual locations, follows sitemap indexes, unpacks gzip, and returns a clean URL list with last-modified dates, images, videos, news and hreflang data. Filter by date or pattern, then feed the list to a scraper, a Google Sheet, or an alert.

### What does Sitemap URL Extractor + Change Monitor do?

A [sitemap](https://www.sitemaps.org/protocol.html) is the file a website publishes so search engines can find all of its pages. This Actor reads those files and gives you the list.

- **Discovery**: paste a domain (`example.com`), a `robots.txt` URL, or a sitemap URL. For a domain it reads `robots.txt` for `Sitemap:` lines, then tries `/sitemap.xml`, `/sitemap_index.xml` and six other common locations.
- **Every format**: XML sitemaps, sitemap indexes (recursively), gzipped `.xml.gz`, plain-text sitemaps, RSS and Atom feeds. Gzip is detected from the bytes, not the file name.
- **Every field**: `lastmod` (raw and normalised ISO), `changefreq`, `priority`, the sitemap the URL came from, and the image, video, Google News and hreflang extensions.
- **Filters**: only URLs modified after a date or in the last `7d`, include/exclude regular expressions, duplicate removal across sitemaps.
- **Change monitoring**: run it on a schedule with mode **Only changes** and get just the URLs that were added, removed or had their `lastmod` changed since the previous run. The first run stores a baseline.
- **robots.txt summary**: which AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot and more) the site blocks, and whether it publishes an `llms.txt`.
- **Optional status checks**: request each URL (cut off after the headers) and record the HTTP status and final URL. Off by default, so a 150,000-URL site takes seconds, not hours.

### Why use it?

Most sitemap tools stop at `url` and `lastmod`, fail on gzip or malformed XML, or check every page's status by default and time out on big sites. This one is built to finish:

- A broken child sitemap, an HTML error page where a sitemap should be, or a `500` from one file does not fail the run. The good sitemaps are read, the bad ones are listed in the run summary with the reason, and the summary says whether the run was complete.
- Retries with back-off on `429`/`5xx` and network errors; realistic browser headers so fewer sites answer `403`; optional proxy for the rest.
- Sitemaps are streamed, so a 50 MB file does not need 50 MB of memory. Caps on URL count, sitemap count, nesting depth and file size keep a runaway index from running forever.
- Change detection is careful: removals are never reported from an incomplete run, a sitemap that briefly returns half its URLs does not produce a flood of false "removed" rows, and a run that stops at your cost limit remembers what it did not report so it shows up next time.
- Restart-safe: if the platform migrates the run, it continues without writing duplicate rows or charging twice.

**Use cases**: seed a scraper with fresh URLs only (`lastmodAfter: 7d`), watch a competitor for new product or blog pages, audit a site's sitemap health, build a URL inventory for SEO, check which AI bots a site allows.

### How to extract sitemap URLs

1. In **Websites or sitemaps**, add a domain such as `example.com`, or a sitemap URL if you already know it. One entry per site.
2. Leave **What to output** on **All URLs** for a full list. Optionally set **Only URLs modified after** (`2026-09-01` or `30d`) and a pattern such as `/blog/`.
3. Click **Start**. The **Output** tab lists the URLs; the **Storage > Key-value store > SUMMARY** record has the full report (each sitemap fetched, robots.txt findings, counts, warnings).
4. To monitor changes, switch **What to output** to **Only changes since the last run** and create a **Schedule** (daily works well). The first run stores the baseline and outputs nothing; every later run outputs only the differences.

### Input

| Field                                                                | What it does                                                                                                                         |
| -------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| `startUrls`                                                          | Domains, robots.txt URLs or sitemap URLs.                                                                                            |
| `mode`                                                               | `all` (default) = every URL. `changes` = only URLs added, removed or changed since the previous run with the same snapshot.          |
| `maxUrls`                                                            | Stop after this many URLs (default 500,000).                                                                                         |
| `lastmodAfter`                                                       | Keep URLs whose `lastmod` is after a date (`2026-09-01`), a datetime, or a span back from now (`7d`, `36h`, `2 weeks`).              |
| `includeUrlsWithoutLastmod`                                          | With `lastmodAfter`: keep URLs that have no `lastmod` (default off).                                                                 |
| `urlIncludePatterns`, `urlExcludePatterns`                           | Regular expressions the URL must match / must not match.                                                                             |
| `dedupeUrls`                                                         | Output each URL once even if several sitemaps list it (default on).                                                                  |
| `includeExtensions`                                                  | Extract image, video, news and hreflang data (default on).                                                                           |
| `snapshotKey`, `firstRunOutput`, `allowMassRemoval`, `resetSnapshot` | Change-detection settings, see below.                                                                                                |
| `checkStatus`, `maxStatusChecks`                                     | Request each URL (cut off after the headers) and record `status` and `finalUrl`. Off by default; capped at 5,000 per run by default. |
| `aiBotSummary`                                                       | Add the AI-crawler policy table and the `llms.txt` check to the summary (default on).                                                |
| `maxSitemaps`, `maxDepth`, `maxSitemapMegabytes`                     | Caps: sitemap files per run (2,000), index nesting (5), size per file (100 MB uncompressed).                                         |
| `maxConcurrency`, `requestTimeoutSecs`                               | Parallel sitemap requests (5) and per-request timeout (30 s).                                                                        |
| `customUserAgent`, `proxyConfiguration`                              | Identify your crawler, or route through Apify Proxy when a site blocks datacenter IPs.                                               |

### Output

One dataset row per URL. You can download it as JSON, CSV, Excel or HTML, or pipe it into another Actor (for example a scraper's `startUrls`, or [Dataset to Google Sheets Sync](https://apify.com/osel_house/dataset-to-google-sheets)).

```json
{
    "url": "https://example.com/blog/how-to-read-a-sitemap",
    "lastmod": "2026-09-20T08:15:00+00:00",
    "lastmodIso": "2026-09-20T08:15:00.000Z",
    "changefreq": "weekly",
    "priority": 0.8,
    "sitemap": "https://example.com/sitemap-posts.xml",
    "sitemapIndex": "https://example.com/sitemap_index.xml",
    "images": [{ "loc": "https://example.com/img/sitemap.png", "title": "Sitemap diagram", "caption": null }],
    "videos": [],
    "news": null,
    "alternates": [{ "hreflang": "de", "href": "https://example.com/de/blog/sitemap-lesen" }]
}
```

In **changes** mode each row also has `change` (`added`, `changed` or `removed`) and `previousLastmod`. With status checks on, `status` and `finalUrl` are added.

| Field                       | Meaning                                                                                          |
| --------------------------- | ------------------------------------------------------------------------------------------------ |
| `url`                       | The page URL exactly as listed (fragment removed).                                               |
| `lastmod` / `lastmodIso`    | Last-modified value as written in the sitemap, and normalised to ISO 8601 (null if unparseable). |
| `changefreq`, `priority`    | Sitemap hints, when present.                                                                     |
| `sitemap`, `sitemapIndex`   | Which sitemap file listed the URL and, if it came via an index, which index.                     |
| `images`, `videos`, `news`  | Image, video and Google News extension data.                                                     |
| `alternates`                | `hreflang` alternates (`xhtml:link rel="alternate"`).                                            |
| `status`, `finalUrl`        | HTTP status and redirect target from the optional status check.                                  |
| `change`, `previousLastmod` | Change type and the previous `lastmod`, in changes mode.                                         |

The **SUMMARY** record in the run's key-value store contains: `complete` and `incompleteReasons`, counts (found, accepted, written, duplicates, dropped by pattern/date), the list of every sitemap fetched with status, kind, URL count, bytes, time and error, the discovery path per site, the robots.txt findings with the AI-bot table, and all warnings.

### Change monitoring in detail

- The URL set of each run is saved as a **snapshot** in a key-value store named `sitemap-snapshots` in your account. The snapshot name is derived from the start URLs and patterns, so the same input always compares against its own previous run. Set `snapshotKey` to control this yourself.
- A `changed` row means the `lastmod` value differs from the previous run (including from empty to a date). Two spellings of the same moment (`2026-01-05` and `2026-01-05T00:00:00Z`) count as unchanged.
- **Removals are suppressed** when the run was incomplete (a sitemap failed with a server error or timeout, a cap or the cost limit was hit, the run was aborted or stopped before its timeout), when more than half of a 100+ URL baseline disappears in one run, and when `lastmodAfter` is a moving window like `7d` (URLs leave the window without being removed from the site). In all cases the missing URLs stay in the snapshot and the summary says why (`changes.removedSuppressedReason`). A sitemap that answers 404 is treated as really gone. Turn on `allowMassRemoval` to report a genuine site restructure.
- The first run stores the baseline and outputs nothing (`firstRunOutput: all-as-added` outputs every URL as `added` instead). A run that could not read any sitemap stores no baseline. `resetSnapshot` starts over.
- Run one schedule per site at a time; two runs comparing against the same snapshot at the same moment can each report the same change.

### Pricing

Pay per event: the platform's standard Actor start fee, a fraction of a cent per URL row written, a separate small fee per URL status check (only when you turn checks on), and a per-run fee for change detection (charged only when a comparison happens, not on the baseline run). Rows that are not written because your maximum cost per run was reached are never charged; the summary says the run stopped early and, in changes mode, those changes are reported on the next run instead. When the budget is tight, URL rows come first and status checks are cut, never the other way round. Exact prices are on this page's pricing box.

A full extraction of a 150,000-URL site takes a few seconds and one request per sitemap file; a daily change check on the same site costs the start fee, the change-detection fee, and only the rows that changed.

### Tips

- **Fresh URLs for a scraper**: `lastmodAfter: "7d"` plus a schedule. URLs without a `lastmod` are dropped by default (turn on `includeUrlsWithoutLastmod` if the site omits dates).
- **One section of a site**: `urlIncludePatterns: ["/blog/"]`. Patterns are regular expressions; escape dots and question marks (`\\?page=`).
- **Site answers 403 or a Cloudflare page**: set **Proxy** to Apify Proxy (residential if datacenter is blocked). Or point `startUrls` straight at the sitemap URL, which is often less protected than the homepage.
- **Huge sites**: raise `maxUrls` and give the run 2 to 4 GB of memory for several million URLs. `maxConcurrency` 10 speeds up indexes with hundreds of child files.
- **Status checks**: keep `maxStatusChecks` modest; each check is one request to the site. Combine with `urlIncludePatterns` to check only what matters.
- **Sitemap not found**: the summary's `discovery` section lists every location tried and what came back. Many sites only publish the sitemap location inside robots.txt, or use a non-standard path; paste it into `startUrls`.

### FAQ and disclaimer

**Does it read page content?** No. It reads only robots.txt, the first bytes of llms.txt (existence check), and sitemap files, which sites publish for automated consumption. The output is URL lists and sitemap metadata. Status checks send one request per URL, cut off as soon as the headers arrive, and store only the status code and final URL.

**Is this legal?** Sitemaps and robots.txt are public protocol files intended for machines. You choose the sites you target and are responsible for using the results in line with those sites' terms and applicable law. Reading a sitemap is not the same as scraping the pages it lists.

**Why is `complete` false?** One or more sitemaps could not be read, a cap (`maxUrls`, `maxSitemaps`, `maxDepth`, size, cost) stopped the run, or the run stopped fetching shortly before its timeout so that the rows it had could still be written (raise the run timeout to read everything). `incompleteReasons` and the per-sitemap list say exactly what happened. A start page that is not a sitemap, or a stale robots.txt entry replaced by a sitemap found at a well-known path, does not count as a failure.

**Something else?** Report it in the **Issues** tab of this Actor. Custom variants (structured-data validation, sitemap generation, other formats) can be built on request.

# Actor input Schema

## `startUrls` (type: `array`):

A domain or homepage (sitemaps are found through robots.txt and the usual locations), a robots.txt URL, or a sitemap URL (.xml, .xml.gz, sitemap index, .txt, RSS/Atom). One entry per site.

## `mode` (type: `string`):

All URLs = every URL in the sitemaps. Only changes = compare with the previous run for the same sites and output only URLs that were added, removed or had their lastmod changed (the first run stores a baseline).

## `maxUrls` (type: `integer`):

Stop after this many URLs (after filters). The run says so when the cap is hit.

## `lastmodAfter` (type: `string`):

Keep only URLs whose <lastmod> is after this point. A date (2026-09-01), a datetime (2026-09-01T00:00:00Z) or a span back from now: 7d, 36h, 2 weeks. Great for scheduled runs that feed a scraper only fresh pages.

## `includeUrlsWithoutLastmod` (type: `boolean`):

Only matters with "Only URLs modified after". Off: URLs without a <lastmod> are dropped (counted in the summary). On: they are kept, since they might be new.

## `urlIncludePatterns` (type: `array`):

Keep only URLs matching at least one of these regular expressions, e.g. /blog/ or .html$. Leave empty to keep all.

## `urlExcludePatterns` (type: `array`):

Drop URLs matching any of these regular expressions, e.g. ?page= or /tag/.

## `dedupeUrls` (type: `boolean`):

The same URL listed in several sitemaps is output once (first occurrence wins). Duplicates are counted in the summary.

## `includeExtensions` (type: `boolean`):

Extract sitemap extensions into the images, videos, news and alternates fields. Turn off for a leaner dataset.

## `snapshotKey` (type: `string`):

Name of the saved URL set to compare against. Leave empty to derive it from the start URLs and patterns (recommended). Set it when you want two different inputs to share one baseline, or to keep several baselines for one site.

## `firstRunOutput` (type: `string`):

What the first run (no snapshot yet) outputs. Baseline only = store the URL set, output nothing. All as added = also output every URL with change = added.

## `allowMassRemoval` (type: `boolean`):

Report removals even when more than half of the baseline disappeared in one run (a site restructure). Off by default so a broken sitemap never floods you with false removals.

## `resetSnapshot` (type: `boolean`):

Delete the saved baseline first; this run stores a new one. Turn it off again afterwards.

## `checkStatus` (type: `boolean`):

Request each URL (the transfer is cut off as soon as the headers arrive) and record the HTTP status and final URL after redirects. Off by default: it is one request per URL and slows large sites down.

## `maxStatusChecks` (type: `integer`):

Safety cap; rows beyond it get status null.

## `aiBotSummary` (type: `boolean`):

Report in the run summary whether robots.txt blocks GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot and other AI crawlers, and whether /llms.txt exists.

## `maxSitemaps` (type: `integer`):

Cap on sitemap files fetched per run (indexes can list thousands).

## `maxDepth` (type: `integer`):

How many levels of sitemap indexes to follow (0 = do not follow indexes).

## `maxSitemapMegabytes` (type: `integer`):

A sitemap larger than this is read up to the cap and marked truncated. The protocol limit is 50 MB.

## `maxConcurrency` (type: `integer`):

How many sitemap files to fetch at the same time.

## `requestTimeoutSecs` (type: `integer`):

Per request. Retries with backoff happen on 429/5xx and network errors.

## `customUserAgent` (type: `string`):

Leave empty for a real browser header set. Set it to identify your crawler (some sites whitelist known agents).

## `proxyConfiguration` (type: `object`):

Use Apify Proxy when a site blocks datacenter requests (403 on robots.txt or sitemap). Off by default; most sitemaps do not need it.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.sitemaps.org/"
    }
  ],
  "mode": "all",
  "maxUrls": 500000,
  "includeUrlsWithoutLastmod": false,
  "dedupeUrls": true,
  "includeExtensions": true,
  "firstRunOutput": "baseline-only",
  "allowMassRemoval": false,
  "resetSnapshot": false,
  "checkStatus": false,
  "maxStatusChecks": 5000,
  "aiBotSummary": true,
  "maxSitemaps": 2000,
  "maxDepth": 5,
  "maxSitemapMegabytes": 100,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `urls` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.sitemaps.org/"
        }
    ],
    "mode": "all",
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("osel_house/sitemap-extractor-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://www.sitemaps.org/" }],
    "mode": "all",
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("osel_house/sitemap-extractor-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.sitemaps.org/"
    }
  ],
  "mode": "all",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call osel_house/sitemap-extractor-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,osel_house/sitemap-extractor-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/UPwyKCEgpaPgsrf9I/builds/judefNpEinkYWXT4Z/openapi.json
