# Sitemap URLs: Exact File, No Page Crawl (`jmspwr/sitemap-url-extractor`) Actor

Extract the exact sitemap file you name, or discover a domain's sitemaps. Deduplicated URL, lastmod and source rows with explicit partial/error reports. No page crawl or proxy; empty runs are free.

- **URL**: https://apify.com/jmspwr/sitemap-url-extractor.md
- **Developed by:** [James Power](https://apify.com/jmspwr) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.50 / 1,000 sitemap urls

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URLs from the exact file you choose

Give the Actor a **specific sitemap URL** and it reads that file, not a different sitemap it finds
at the domain root. Or give it a bare domain to discover sitemaps. The dataset contains one row
per distinct URL, with `lastmod` and the sitemap file that supplied it.

Only sitemap files are requested — `robots.txt`, `sitemap.xml`, `.xml.gz` files and sitemap
indexes. **A page URL is never fetched.** This avoids page rendering and keeps requests light,
but a site's WAF can still refuse a sitemap file; the output names the failed file and status.

### Input

| Field | What it does |
|---|---|
| **Domains or sitemap URLs** | One per line. `https://example.com` is discovered through its robots.txt, falling back to `/sitemap.xml` and `/sitemap_index.xml`. `https://example.com/sitemap.xml` is read directly. `https://` is added if you leave it out. Up to 32 lines. |
| **Maximum URLs** | Default 2,000 (maximum 200,000). The run stops there and reports `partial`. |
| **Maximum sitemap files per domain** | Default 50 (maximum 1,000), for sites that split their URLs over many files. |

That is the whole input — there is nothing to configure about parsing, because the sitemap protocol
decides it.

### Output

The dataset has exactly three fields per row:

```json
{ "url": "https://example.com/pricing", "lastmod": "2026-08-14", "source": "https://example.com/sitemap.xml" }
```

- `url` — exactly as your sitemap declared it. It was never requested.
- `lastmod` — the sitemap's `<lastmod>` for that URL, normalised to ISO 8601 (`2026-08-14` or
  `2026-08-14T09:30:00+00:00`), or `null` when it was absent or not a date. When the same URL
  appears in more than one file, the newest `lastmod` wins.
- `source` — the sitemap file that first listed the URL (after any redirects), so you always know
  where a row came from.

The run's **output record** (`OUTPUT` in the key-value store) is the report:

| Field | Meaning |
|---|---|
| `status` | `ok`, `partial` or `error` — see below. |
| `urls`, `datasetRows` | Unique URLs found, and rows actually stored (they match unless the run's cost limit was reached). |
| `charge.events` | Billed events, one per stored row. |
| `totals` | duplicates skipped, URLs with a lastmod, invalid or future lastmod values, `<loc>` values that are not valid URLs, off-site URLs, files read. |
| `inputs[]` | Per line: how the sitemap was found, every file that was tried with its HTTP status, bytes, whether it was gzipped, what was skipped and why, limits hit. |
| `limitsHit`, `warnings` | Which cap stopped the run, and any input line that was ignored and why. |

**`status` is never optimistic.** `ok` means every sitemap you asked for was read and nothing was
skipped. `partial` means you got URLs but something was cut short — a file that failed, a cap, a
cross-site sitemap that was not fetched, or `<loc>` values that are not valid URLs (up to two are
quoted back in the report with the reason). `error` means the
dataset is empty. A `partial` or `error` run is honest about the reason instead of quietly
returning a short list.

### Reading sitemaps well

- **Discovery**: `robots.txt` `Sitemap:` lines first (RFC 9309), then the conventional
  `/sitemap.xml` and `/sitemap_index.xml`. If robots.txt names a dead sitemap, the conventional
  paths are still tried.
- **Indexes**: nested sitemap indexes are followed (up to three levels), with every child URL
  validated before it is fetched.
- **Compression**: `.xml.gz` files are expanded under a byte budget, and gzip content-encoding is
  handled by the HTTP client. Compression bombs are refused, not decompressed into memory.
- **Deduplication**: one row per unique URL, across every file and every line of input. A duplicate
  is never stored twice, so it is never charged twice.
- **Hostile XML**: entity declarations and external entity references are refused; a plain
  DOCTYPE is ignored without fetching it. A broken file leaves the rest of the run alone.
- **Honest about bad data**: a `<loc>` that is not a valid URL — whitespace in the middle, a
  `mailto:` entry, a URL longer than 2048 characters, credentials in the URL — is left out and
  reported with an example, rather than written to the dataset as if it were usable. Real sitemaps
  do contain such entries; a well-behaved reader says so instead of silently emitting them.

### Safety by construction

- Only **globally routable** addresses are dialled, and that is enforced by the resolver that opens
  the socket (not by a check that happens before it), so a DNS answer that changes between the
  check and the connection cannot redirect the request anywhere private.
- `http://` and `https://` only; no credentials in URLs; known non-HTTP ports are refused; private,
  loopback, link-local and cloud-metadata addresses are refused.
- Every redirect hop is validated before it is followed; the client itself is not allowed to follow
  redirects on its own.
- **Child sitemaps must be on the same site** as the sitemap that lists them. A sitemap hosted on a
  different domain is *not* fetched from a `robots.txt` or index declaration — add it as its own
  input line if you want it read, and it will then be read as a first-class target.
- URLs inside a sitemap that point at another domain are returned as data and counted in the output
  record; they are never requested.
- No proxy, browser, JavaScript, API keys, login or cookies. The Actor never crawls profiles or
  extracts contact fields, but sitemap URLs can contain profile paths or personal details in
  query strings. Review the returned URLs before sharing them.

### Billing

US$0.0005 per URL (US$0.50 per 1,000), one event per stored dataset row. No start fee. A run
that finds nothing, or fails, stores nothing and is charged nothing. If your cost limit for the
run is reached, the crawl stops early, the output
record says `charge.limitReached: true` and the status is `partial`.

### Limits and what is not included

- No full-page crawling and no JavaScript rendering — this Actor reads sitemaps, nothing else.
- Image, video and news extension data inside a sitemap is ignored; each entry's `<loc>` and
  `<lastmod>` are read.
- A run reads at most 50 MB per file, 256 MB in total, and by default 50 files per domain (input
  maximum 1,000). The default time limit is 15 minutes; the output reports when a bound is hit.

### Local development

```
python -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/pip install pytest
.venv/bin/python -m pytest        # offline tests against a local synthetic fixture server
```

The tests never touch the public internet, an Apify account or anyone's real website; they cover
discovery, gzip, indexes, hostile XML, byte and count caps, cancellation, redirects, private-address
and DNS-rebinding refusal, and the billing rules.

# Actor input Schema

## `sitemaps` (type: `array`):

One domain or sitemap URL per line. A domain (https://example.com) is discovered through its robots.txt, falling back to /sitemap.xml; a path (https://example.com/sitemap.xml) is read directly. "https://" is added if you leave it out. Up to 32 lines are used.

## `maxUrls` (type: `integer`):

Stop after this many unique URLs across the whole run. The output record then says status "partial" and limitsHit \["maxUrls"]. Nothing beyond the limit is stored, so nothing beyond it is charged.

## `maxFiles` (type: `integer`):

Stop after this many sitemap files per domain. Large sites split their URLs across many files; raise this only if you need every one of them.

## Actor input object example

```json
{
  "sitemaps": [
    "https://onescales.com/sitemap_blogs_1.xml"
  ],
  "maxUrls": 2000,
  "maxFiles": 50
}
```

# Actor output Schema

## `urls` (type: `string`):

One row per unique URL: url, lastmod (ISO 8601 or null) and source (the sitemap file it was listed in). Rows are stored only for URLs that were actually read, and each stored row is one billed event.

## `report` (type: `string`):

status (ok, partial or error), counts, every sitemap file that was tried with its HTTP status or error, what was skipped and why, limits hit, and how many events were charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sitemaps": [
        "https://onescales.com/sitemap_blogs_1.xml"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("jmspwr/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "sitemaps": ["https://onescales.com/sitemap_blogs_1.xml"] }

# Run the Actor and wait for it to finish
run = client.actor("jmspwr/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sitemaps": [
    "https://onescales.com/sitemap_blogs_1.xml"
  ]
}' |
apify call jmspwr/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jmspwr/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QfVSqDWVmJzf9xXrh/builds/T91sWcbutgX1Pn6I7/openapi.json
