# Sitemap URL Extractor (`palgenius/sitemap-url-extractor`) Actor

Extract every URL from a website's sitemaps: indexes, .gz, plain-text and RSS/Atom, with lastmod filtering.

- **URL**: https://apify.com/palgenius/sitemap-url-extractor.md
- **Developed by:** [Ali Alsaidi](https://apify.com/palgenius) (community)
- **Categories:** Developer tools, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.30 / 1,000 url extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor

Point it at a domain or a sitemap and get back every page URL the site publishes — clean,
de-duplicated, and filtered how you want it.

### What does Sitemap URL Extractor do?

- **Turns one domain into a full URL list.** Give it `example.com` and it reads `robots.txt`,
  then tries nine common sitemap paths until it finds one.
- **Handles the sitemaps that break other extractors** — indexes nested several levels deep,
  gzipped `.xml.gz` files, plain-text and RSS/Atom sitemaps, and XML with a malformed tag in
  the middle (it returns every URL it can read instead of failing the run).
- **Streams instead of loading.** A 500,000-URL sitemap is parsed in chunks and memory stays
  flat. Measured: 63,000 URLs/second, 10 MB peak on a 4 MB sitemap.
- **Filters before it charges you.** Regex include/exclude and a `lastmodAfter` date cut both
  the noise and the bill.
- **Does not check every URL by default.** Most extractors fire an HTTP request per URL to
  report a status code, which turns a 3-second job into a 30-minute one. Here that is a
  switch (`checkStatus`), off unless you ask.

### What data can you get?

| Field | What it means |
|---|---|
| `url` | The page URL, absolute and de-duplicated |
| `lastmod` | Last-modified date the site publishes, or `null` if it publishes none |
| `changefreq` | How often the site claims the page changes |
| `priority` | The site's own priority hint, 0.0–1.0 |
| `sitemap` | Which sitemap file this URL came from |
| `depth` | How many index levels deep that sitemap was |
| `status` | HTTP status — only when `checkStatus` is on |
| `alternates` | `hreflang` versions of the page — only when `includeAlternates` is on |
| `images` | Image URLs listed for the page — only when `includeImages` is on |

### How to use it

1. Create a free Apify account.
2. Open the Actor and click **Try for free**.
3. Put your domain or sitemap URL in **Start URLs**, and set **Maximum URLs** to cap the run.
4. Click **Start**. A typical site finishes in seconds.
5. Open the **Dataset** tab and download as JSON, CSV, Excel or XML, or pull it through the
   Apify API.

### Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `startUrls` | array | — | Domains or sitemap URLs (required) |
| `maxUrls` | integer | 1000 | Stop after this many URLs. A hard ceiling on run cost |
| `checkStatus` | boolean | false | Add an HTTP `status` per URL (slower) |
| `includePatterns` | array | `[]` | Keep only URLs matching these regexes, e.g. `/blog/` |
| `excludePatterns` | array | `[]` | Drop URLs matching these, e.g. `/tag/` |
| `lastmodAfter` | string | — | Only URLs changed after a date, e.g. `2026-01-01` |
| `includeAlternates` | boolean | false | Add `hreflang` versions of each URL |
| `includeImages` | boolean | false | Add image URLs listed for the page |
| `maxDepth` | integer | 5 | How deep to follow nested sitemap indexes |
| `concurrency` | integer | 5 | Parallel sitemap fetches |
| `requestTimeoutSecs` | integer | 30 | Per-request timeout |
| `maxRetries` | integer | 3 | Retries per request |
| `respectRobotsTxt` | boolean | true | Obey `robots.txt` |
| `proxyConfiguration` | object | — | Optional; sitemaps are public, so rarely needed |

```json
{
  "startUrls": [{ "url": "https://apify.com" }],
  "maxUrls": 5000,
  "excludePatterns": ["/tag/"],
  "lastmodAfter": "2026-01-01"
}
```

### Output

One item per URL. This is a real item from a run against `apify.com`:

```json
{
  "url": "https://apify.com/ai-agents",
  "lastmod": null,
  "changefreq": null,
  "priority": null,
  "sitemap": "https://apify.com/sitemap/pages.xml",
  "depth": 1
}
```

The nulls are honest: apify.com's sitemap publishes no `lastmod`, `changefreq` or `priority`,
and this Actor reports what is there rather than inventing it. Many sites do publish all three.

Every run also writes a `RUN_SUMMARY` record to the key-value store with `urlsReturned`,
`urlsFound`, `sitemapsParsed`, `sitemapsFailed`, `requests`, `retries`, `blockedByRobotsTxt`
and a `problems` list. A partial result is never a mystery.

### How much does it cost?

**$0.30 per 1,000 URLs returned** ($0.0003 each), and nothing else. **10,000 URLs = $3.**
Apify platform usage — compute, storage, transfer — is included in that price, not billed
on top.

Filters are applied *before* charging, so `includePatterns`, `excludePatterns` and
`lastmodAfter` reduce the bill as well as the noise. `maxUrls` is a hard ceiling. Re-running a
5,000-URL site monthly, filtered to pages changed since last time, usually costs a few cents.

### Integrations

Connect it to webhooks, Make, Zapier, Google Sheets, Slack, Airbyte or GitHub from the
**Integrations** tab — for example, run it weekly and append new URLs to a sheet.

Call it from your own code with the Apify API:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("palgenius/sitemap-url-extractor").call(
    run_input={"startUrls": [{"url": "https://apify.com"}], "maxUrls": 5000}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["url"])
```

AI agents can call it through the **Apify MCP server** at `https://mcp.apify.com`, e.g.
`https://mcp.apify.com?tools=palgenius/sitemap-url-extractor`.

### Error items and limits

- **No sitemap found.** The run fails with a status message naming the first problem, and
  `RUN_SUMMARY.problems` lists every path that was tried. You are not charged for URLs that
  do not exist.
- **A sitemap could not be read** (timeout, 5xx, broken gzip): counted in `sitemapsFailed`
  and described in `problems`; the other sitemaps still return their URLs.
- **`robots.txt` disallows a sitemap path:** skipped, and counted in `blockedByRobotsTxt`.
- `lastmod` is whatever the site publishes. Many sites omit it and some never update it.
  URLs without a `lastmod` are dropped when you set `lastmodAfter`.
- A sitemap can list URLs that no longer exist. Turn on `checkStatus`, or pipe the output
  into the Bulk URL Status Checker below.

### FAQ

**Is this legal?** Yes. Sitemaps are files sites publish specifically so crawlers can read
them. This Actor reads only public sitemap files, obeys `robots.txt` by default, never logs
in, and collects no personal data.

**What if I only know the domain?** That is the normal case. It reads `robots.txt` and tries
nine common sitemap locations.

**Will a huge sitemap time out?** It streams and parses in chunks, so size is not the limit —
`maxUrls` and your run timeout are.

**Can I get only pages changed since last month?** Set `lastmodAfter`, on sites that publish
`lastmod`.

**Do I pay for URLs I filter out?** No. Filtering happens before charging.

**Can I check which of these URLs are broken?** Turn on `checkStatus`, or use the status
checker for redirect chains and response times.

### You might also like

- [Bulk URL Status Checker](https://apify.com/palgenius/bulk-url-status-checker) — feed it
  this run's dataset ID to check every URL for broken links and redirect chains.
- [Tech Stack Detector](https://apify.com/palgenius/tech-stack-detector) — find the platform,
  CMS and frameworks behind any domain, with evidence.

### Feedback

Found a sitemap it cannot read, or want a field added? Open an issue on the Actor's **Issues**
tab. Fixes ship within days.

# Actor input Schema

## `startUrls` (type: `array`):

A domain (example.com) or a direct sitemap URL (example.com/sitemap.xml). For a domain, the Actor reads robots.txt first and then tries the usual sitemap paths. Sitemap indexes, .gz, plain-text and RSS/Atom sitemaps all work.

## `maxUrls` (type: `integer`):

Stop after this many URLs. Every URL returned is one charged result.

## `checkStatus` (type: `boolean`):

Off by default, because one request per URL is what makes sitemap extractors slow and flaky on big sites. Turn it on to add a `status` field.

## `includePatterns` (type: `array`):

Regular expressions. A URL is kept if it matches at least one, e.g. /blog/ or .html$

## `excludePatterns` (type: `array`):

Regular expressions. A URL is dropped if it matches any of them, e.g. /tag/ or ?

## `lastmodAfter` (type: `string`):

ISO date, e.g. 2026-01-01. Uses the sitemap's <lastmod>; URLs without one are dropped. Handy for incremental crawls.

## `includeAlternates` (type: `boolean`):

Add the xhtml:link language versions of each URL (e.g. the Arabic version).

## `includeImages` (type: `boolean`):

Add image locations from image:image entries, where the sitemap has them.

## `maxDepth` (type: `integer`):

How deep to follow sitemap indexes that point at other indexes.

## `concurrency` (type: `integer`):

How many sitemap files to fetch at once. Keep it polite.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for a single sitemap file before giving up and retrying.

## `maxRetries` (type: `integer`):

Retries with backoff on timeouts and 429/5xx responses.

## `respectRobotsTxt` (type: `boolean`):

Skip anything robots.txt disallows. Leave this on unless you own the site.

## `proxyConfiguration` (type: `object`):

Optional. Sitemaps are public files, so most sites need no proxy.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://apify.com"
    }
  ],
  "maxUrls": 100,
  "checkStatus": false,
  "includePatterns": [],
  "excludePatterns": [],
  "includeAlternates": false,
  "includeImages": false,
  "maxDepth": 5,
  "concurrency": 5,
  "requestTimeoutSecs": 30,
  "maxRetries": 3,
  "respectRobotsTxt": true
}
```

# Actor output Schema

## `urls` (type: `string`):

One item per URL, with lastmod, changefreq, priority and the sitemap it came from.

## `runSummary` (type: `string`):

Counts for this run: URLs returned and found, sitemaps parsed and failed, requests, retries, robots.txt blocks, and any problems.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://apify.com"
        }
    ],
    "maxUrls": 100
};

// Run the Actor and wait for it to finish
const run = await client.actor("palgenius/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://apify.com" }],
    "maxUrls": 100,
}

# Run the Actor and wait for it to finish
run = client.actor("palgenius/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://apify.com"
    }
  ],
  "maxUrls": 100
}' |
apify call palgenius/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,palgenius/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WTTcmuBqAfPguX3TW/builds/0Z44y8iP6PDw1501p/openapi.json
