# Sitemap Extractor Plus (`gmanner/sitemap-extractor-plus`) Actor

Lists every URL from a site's robots.txt and nested or gzipped sitemaps with lastmod, a changed-since filter, per-section counts and optional HTTP status checks.

- **URL**: https://apify.com/gmanner/sitemap-extractor-plus.md
- **Developed by:** [Hwangjun Choi](https://apify.com/gmanner) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap Extractor Plus

Every URL a site publishes in its sitemaps, as a clean dataset: URL, last-modified date, change frequency, priority, section, and the sitemap it came from. Point it at a homepage and it reads `robots.txt`, follows sitemap indexes to any depth, decompresses `.gz` files, and dedupes the result. Optional: keep only URLs changed since a date, and check the HTTP status of each URL.

### What you can do with it

- **Content inventory** - list all pages of a site (yours or a competitor's) with per-section counts, ready for a spreadsheet or a crawler.
- **Change monitoring** - run weekly with `changedSince: 7d` to get only pages added or updated since the last run. Child sitemaps older than the cutoff are not even downloaded, so incremental runs on large sites take seconds.
- **Migration and SEO QA** - enable status checks to find sitemap URLs that 404, redirect, or point to the wrong content type.
- **Seed lists for scrapers** - feed the dataset into any crawler instead of discovering links page by page.

### Input

| Field | Meaning |
|---|---|
| `startUrls` | Homepages (robots.txt discovery, fallback to `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml` and similar) or direct sitemap URLs (`.xml`, `.xml.gz`, `.txt`, indexes). |
| `changedSince` | `2026-10-01`, an ISO datetime, or `24h` / `7d` / `2w` / `1m`. URLs without `lastmod` are dropped unless `keepUrlsWithoutLastmod` is on. |
| `includeUrlPattern` / `excludeUrlPattern` | Case-insensitive regular expressions applied to each URL. |
| `checkStatus` | HEAD request per URL (GET fallback), 5 in parallel, 200 ms apart per host. |
| `maxUrls`, `maxSitemaps`, `maxDepth` | Hard caps; the run cost cannot exceed `maxUrls`. |

Minimal input:

```json
{ "startUrls": [{ "url": "https://www.gov.uk" }], "changedSince": "7d", "maxUrls": 5000 }
```

### Output

One row per unique URL:

```json
{
  "url": "https://www.gov.uk/hmrc-internal-manuals/double-taxation-relief/dt4850pp",
  "lastmod": "2026-09-30T10:05:18.000Z",
  "lastmodRaw": "2026-09-30T11:05:18+01:00",
  "changefreq": null,
  "priority": 0.5,
  "site": "www.gov.uk",
  "section": "/hmrc-internal-manuals",
  "sitemapUrl": "https://www.gov.uk/sitemaps/sitemap_2.xml",
  "sitemapRoot": "https://www.gov.uk/sitemap.xml",
  "sitemapKind": "urlset",
  "sitemapDepth": 1,
  "discoveredVia": "input",
  "alternatesCount": 0,
  "alternates": null,
  "imageCount": 0,
  "videoCount": 0,
  "newsTitle": null,
  "newsPublicationDate": null
}
```

With `checkStatus` each row also gets `status`, `redirected`, `finalUrl`, `contentType` and `statusCheckedAt`. `alternates` holds `hreflang` links, `newsTitle` / `newsPublicationDate` come from Google News sitemaps, and RSS/Atom feeds listed as sitemaps are accepted too.

The key-value store record `SUMMARY` has the run statistics, per-section counts and one entry per sitemap file (kind, HTTP status, URL count, error), so a failing or empty sitemap is visible without reading logs.

### Pricing

Pay per event, no subscription:

| Event | Price |
|---|---|
| Run start | $0.003 |
| URL listed | $0.0004 |
| URL status checked (optional) | $0.0003 |

Examples: 2,000 URLs cost $0.003 + 2,000 x $0.0004 = **$0.80**; the same run with status checks costs $1.40. A weekly `changedSince: 7d` run that finds 150 changed pages costs $0.06. Set `maxUrls` to cap the spend.

### Limits

- Only URLs the site publishes in sitemaps are returned; pages missing from the sitemap are not discovered.
- Some sites answer sitemap requests with 403 or an HTML challenge from their CDN (seen on nytimes.com and npmjs.com). Those sitemaps are reported in `SUMMARY` with their status and skipped; the run continues with the others.
- `lastmod` is whatever the site declares; many sites omit it or set it to the crawl time. Rows keep the raw value in `lastmodRaw` so you can judge it.
- Sitemap files are parsed in memory; the spec's 50 MB uncompressed limit per file is fine, multi-hundred-MB files are not.
- Status checks are deliberately slow (5 parallel, 200 ms per host) to stay polite; 10,000 checks take about 7 minutes.

### Local test

`npm test` runs the Actor against a local synthetic site (robots discovery, gzipped index, nested, text sitemap, `changedSince`, status checks) and against gov.uk, bbc.com and apify.com, then asserts the dataset fields.

# Actor input Schema

## `startUrls` (type: `array`):

Site homepages (the Actor reads robots.txt and falls back to /sitemap.xml and friends) or direct sitemap URLs (.xml, .xml.gz, .txt, sitemap indexes).

## `changedSince` (type: `string`):

Keep only URLs whose <lastmod> is at or after this moment. Accepts an ISO date (2026-10-01), an ISO datetime, or a relative period: 24h, 7d, 2w, 1m. Child sitemaps whose own lastmod is older are skipped entirely, which makes incremental runs fast. Leave empty to list everything.

## `keepUrlsWithoutLastmod` (type: `boolean`):

When 'Only URLs changed since' is set, URLs that have no <lastmod> are dropped by default. Enable this to keep them.

## `includeUrlPattern` (type: `string`):

Only list URLs matching this case-insensitive regular expression, e.g. ^https://www.example.com/blog/

## `excludeUrlPattern` (type: `string`):

Skip URLs matching this case-insensitive regular expression, e.g. .(jpg|png|pdf)$

## `checkStatus` (type: `boolean`):

Send a HEAD request (GET fallback) to every listed URL, 5 at a time, and add status, finalUrl, redirected and contentType. Charged as a separate 'status-checked' event.

## `maxUrls` (type: `integer`):

Stop after this many unique URLs have been listed. The run cost is bounded by this number.

## `maxSitemaps` (type: `integer`):

Stop fetching sitemap files after this many (indexes and leaves combined).

## `maxDepth` (type: `integer`):

How many levels of sitemap indexes to follow. 0 reads only the sitemaps named in robots.txt or the input.

## `discoverFromRobots` (type: `boolean`):

For site URLs, read /robots.txt 'Sitemap:' lines first. If disabled or nothing is listed, common paths such as /sitemap.xml and /sitemap_index.xml are tried.

## `skipUnchangedSitemaps` (type: `boolean`):

When filtering by date, do not download child sitemaps whose <lastmod> in the index is older than the cutoff.

## `requestTimeoutSecs` (type: `integer`):

Per-request timeout for sitemap downloads and status checks.

## `userAgent` (type: `string`):

Sent with every request. The default is a plain desktop browser string.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.gov.uk/sitemap.xml"
    }
  ],
  "changedSince": "7d",
  "keepUrlsWithoutLastmod": false,
  "includeUrlPattern": "/blog/",
  "excludeUrlPattern": "\\?page=",
  "checkStatus": false,
  "maxUrls": 10000,
  "maxSitemaps": 500,
  "maxDepth": 5,
  "discoverFromRobots": true,
  "skipUnchangedSitemaps": true,
  "requestTimeoutSecs": 60,
  "userAgent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/129.0.0.0 Safari/537.36"
}
```

# Actor output Schema

## `results` (type: `string`):

One row per unique URL with lastmod, changefreq, priority, section, source sitemap and optional HTTP status; a final SUMMARY row lists per-sitemap status.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.gov.uk/sitemap.xml"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("gmanner/sitemap-extractor-plus").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://www.gov.uk/sitemap.xml" }] }

# Run the Actor and wait for it to finish
run = client.actor("gmanner/sitemap-extractor-plus").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.gov.uk/sitemap.xml"
    }
  ]
}' |
apify call gmanner/sitemap-extractor-plus --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gmanner/sitemap-extractor-plus"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Q8sm93ZZDKBqzVkbn/builds/SthTGggdJJ5yBoWh1/openapi.json
