# Best Damn Sitemap URL Extractor (`josh99smith/sitemap-url-extractor`) Actor

Extract every URL from a website's XML sitemaps, including sitemap indexes, gzip files and robots.txt discovery, with lastmod, priority and hreflang. Give it a domain, get the whole map back, at a tiny price per URL.

- **URL**: https://apify.com/josh99smith/sitemap-url-extractor.md
- **Developed by:** [Joshua Smith](https://apify.com/josh99smith) (community)
- **Categories:** SEO tools, Developer tools, Open source
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.20 / 1,000 url extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

![sitemap-url-extractor banner](https://raw.githubusercontent.com/josh99smith/apify-actor-assets/main/banners/sitemap-url-extractor.png?v=bd1)

**Sitemap URL extractor**: extract all URLs from a sitemap, with the metadata that comes with them (last modification date, change frequency, priority, `hreflang` alternates and image counts). Paste website URLs or sitemap URLs, and the Actor finds the sitemaps (robots.txt, `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml`, ...), follows nested sitemap indexes, unpacks `.xml.gz` files and returns one clean record per page.

It is built for **SEO specialists, developers and data teams** who need a complete, structured list of a site's pages without crawling it. You pay a small flat price per URL, and sites whose sitemap cannot be found or loaded are reported **free of charge**.

### Features

- Extract all URLs from an XML sitemap or sitemap index
- Find a website's sitemap automatically via robots.txt and common paths
- Parse gzip-compressed `.xml.gz` sitemaps and plain-text sitemaps
- Get `lastmod`, `changefreq`, `priority` and `hreflang` alternates for every URL
- Filter sitemap URLs by glob, regex or substring patterns
- Export a full list of website URLs to CSV, Excel or JSON
- Audit sitemap structure: list every sitemap file with its URL count and depth

### What can you do with Best Damn Sitemap URL Extractor?

- **SEO audits**: compare the sitemap against your crawl or Google Search Console coverage, find stale `lastmod` dates, missing `hreflang` alternates or pages that should not be indexed.
- **Seed other scrapers and crawlers**: get the full URL list of a site in seconds and feed it to a content scraper, screenshot Actor or your own pipeline instead of discovering pages link by link.
- **Monitor competitors**: schedule a weekly run and diff the results to see which pages, products or articles a competitor added or changed.
- **Build AI / RAG corpora**: collect every documentation, blog or help-centre URL of a site and pass the list to a text extractor for embedding.
- **Content inventories and migrations**: export the whole site structure to CSV or Excel before a redesign or platform migration.
- **Replace Zapier/Make sitemap steps**: run on a schedule and push new URLs to Google Sheets, Airtable, Slack or a webhook with Apify integrations.

### How it works

For a website URL the Actor reads `robots.txt` and follows every `Sitemap:` directive. If there is none, it probes `/sitemap.xml`, `/sitemap_index.xml`, `/sitemap-index.xml`, `/wp-sitemap.xml` and `/sitemap.txt`. For a sitemap URL it starts there directly. Sitemap index files are followed recursively (up to the configured depth), gzip-compressed sitemaps are decompressed, and every `<url>` entry is parsed including the image and `xhtml:link` extensions. Requests are limited to 5 concurrent connections per host.

The Actor reads sitemap files only; it does not crawl the pages themselves. Sites without a sitemap, or sitemaps hidden behind a login or bot protection, cannot be extracted and are reported as failures.

### How to use it

1. Open the Actor and paste your URLs into **Website or sitemap URLs**, one per line. Website URLs trigger discovery; direct sitemap URLs are read as-is.
2. Optionally set **Max URLs per site** (default 5,000) to cap the cost of very large sites, and add **Include / Exclude URL patterns** to keep only the sections you need (for example `**/blog/**` or `/products/`).
3. Click **Start**. URLs appear in the **Output** tab while the run is in progress.
4. Download the dataset as JSON, CSV, Excel or XML, or pass it to another Actor or integration.

```json
{
    "urls": ["https://www.apify.com", "https://blog.apify.com/sitemap.xml"],
    "maxUrlsPerSite": 5000,
    "includePatterns": [],
    "excludePatterns": ["*.pdf"],
    "maxSitemapDepth": 5,
    "maxSitemapsPerSite": 500,
    "outputSitemapsOnly": false
}
```

### Output

![Sample output of sitemap-url-extractor](https://raw.githubusercontent.com/josh99smith/apify-actor-assets/main/previews/sitemap-url-extractor.png)

One record per URL:

```json
{
    "site": "https://www.apify.com",
    "sitemapUrl": "https://apify.com/sitemap.xml",
    "url": "https://apify.com/store",
    "lastmod": "2026-09-17T06:12:41.000Z",
    "changefreq": "daily",
    "priority": 0.8,
    "alternates": [{ "hreflang": "en", "href": "https://apify.com/store" }],
    "imageCount": 0,
    "fetchedAt": "2026-09-18T20:24:11.000Z"
}
```

Sitemaps that cannot be found or loaded are still recorded, so nothing silently disappears:

```json
{ "site": "https://this-domain-does-not-exist.example", "success": false, "errorType": "dns", "error": "getaddrinfo ENOTFOUND ...", "fetchedAt": "..." }
```

With **List sitemap files only** switched on, each record describes one sitemap file instead (`sitemapUrl`, `kind`, `urlCount`, `childSitemapCount`, `depth`, `lastmod`).

### Output fields

| Field | Description |
| --- | --- |
| `site` | Origin of the website the URL belongs to (as you supplied it). |
| `sitemapUrl` | The sitemap file the URL was found in. |
| `url` | The page URL (absolute). |
| `lastmod` | Last modification date from the sitemap, normalised to ISO 8601 when possible; `null` if absent. |
| `changefreq` | `always`, `hourly`, `daily`, `weekly`, `monthly`, `yearly`, `never` or `null`. |
| `priority` | Sitemap priority between 0.0 and 1.0, or `null`. |
| `alternates[]` | `hreflang` / `href` pairs from `xhtml:link rel="alternate"` entries. |
| `imageCount` | Number of `image:image` entries attached to the URL. |
| `fetchedAt` | When the sitemap file was read. |
| `errorType` | For failures only: `invalid-url`, `not-found`, `http-error`, `blocked`, `dns`, `timeout`, `network` or `other`. |

### Use it from the API, Python, JavaScript or an AI agent

Run the Actor and get the dataset back in one HTTP call:

```bash
curl -X POST "https://api.apify.com/v2/acts/josh99smith~sitemap-url-extractor/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["https://blog.apify.com/sitemap.xml"], "maxUrlsPerSite": 1000}'
```

Python, with the [apify-client](https://docs.apify.com/api/client/python) package:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("josh99smith/sitemap-url-extractor").call(
    run_input={"urls": ["https://www.apify.com"], "maxUrlsPerSite": 1000, "includePatterns": ["**/blog/**"]}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item.get("url"), item.get("lastmod"))
```

JavaScript or TypeScript, with the [apify-client](https://docs.apify.com/api/client/js) package:

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });
const run = await client.actor('josh99smith/sitemap-url-extractor').call({
    urls: ['https://www.apify.com'],
    maxUrlsPerSite: 1000,
    excludePatterns: ['*.pdf'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((item) => item.url));
```

#### Use it from Claude, Cursor, ChatGPT or any MCP client

The Actor is exposed as a tool by the [Apify MCP server](https://mcp.apify.com), so an AI agent can call it by name. Add this to your MCP client configuration (Claude Desktop, Claude Code, Cursor, VS Code, Windsurf and others):

```json
{
    "mcpServers": {
        "apify": {
            "url": "https://mcp.apify.com?tools=josh99smith/sitemap-url-extractor",
            "headers": { "Authorization": "Bearer <YOUR_API_TOKEN>" }
        }
    }
}
```

Then ask, for example: *"List every URL under /blog/ from the sitemap of https://blog.apify.com with josh99smith/sitemap-url-extractor."* The agent fills in the input, runs the Actor and reads the dataset back; you pay the same per-result price as in the Console.

The Actor can also be scheduled, or connected to Zapier, Make, n8n and Google Sheets in the **Integrations** tab.

### Pricing: how much does it cost to extract sitemap URLs?

You pay a **flat price per extracted URL** (shown next to the Start button); 5,000 URLs cost about $1. Nothing is charged for Actor start-up, for failed sites, or for sitemap files that could not be loaded. The Actor stops automatically when it reaches the maximum cost you set for a run, so a huge site never produces a surprise bill, and **Max URLs per site** caps each site individually.

**How it compares (September 2026).** Apify's own sitemap extractor charges $0.0005 per URL and shows a 16 percent failed-run rate in its public stats; other options charge $0.002 plus a start fee or $0.03 per URL. This Actor is $0.0002 per URL (a 5,000-URL site costs $1), handles sitemap indexes, gzip files and robots.txt discovery, and never bills a site where no sitemap could be found.

### Tips

- **Huge publishers**: sites like news archives expose thousands of monthly sitemap files. `maxSitemapsPerSite` (default 500) caps how many are fetched; raise it together with the run memory (1 GB or more) to walk the whole tree, and use `outputSitemapsOnly` first to see the tree's size cheaply.

- **Large sites**: news and ecommerce sites can list millions of URLs. Combine **Max URLs per site** with **Include URL patterns** to fetch only the section you care about, or run **List sitemap files only** first to see how the sitemap is structured.

- **Patterns**: globs (`**/blog/**`, `*.pdf`), regular expressions in slashes (`/\/products\/\d+$/`) and plain substrings (`/docs/`) are all accepted, case-insensitively.

- **Blocked sites**: a few CDNs refuse cloud IP addresses. Enable **Proxy configuration > Apify Proxy** in the Advanced section (proxy traffic is billed by Apify separately).

- **Change tracking**: use the **Schedule** tab to run weekly and compare `lastmod` values between runs.

### FAQ

#### Why are some pages of the site missing from the output?

Only URLs present in the site's sitemaps are returned; the Actor does not crawl pages. If the site does not maintain a sitemap, use a crawler such as Website Content Crawler instead.

#### Does it work with WordPress, Shopify, Wix, Webflow and Next.js sites?

Yes. All of them publish standard XML sitemaps (WordPress at `/wp-sitemap.xml` or via Yoast/RankMath sitemap indexes), which the Actor discovers automatically.

#### Which sitemap formats are supported?

XML `urlset` and `sitemapindex` files (including the image, video, news and `xhtml:link` extensions), gzip-compressed `.xml.gz` files and plain-text sitemaps. RSS/Atom feeds are not sitemaps and are reported as `not-found`.

#### What are the limits on URLs, depth and file size?

**Max URLs per site** goes up to 200,000 per run (default 5,000) and **Max sitemap index depth** up to 20 levels (default 5). A single sitemap file may be up to 64 MB uncompressed, which covers the 50 MB limit of the sitemap protocol. Up to 10 sites are processed in parallel with at most 5 concurrent requests per host, and each file request times out after at most 120 seconds.

#### Is it legal to extract URLs from a sitemap?

Sitemaps are published specifically so that automated clients can read them. The Actor sends a handful of requests per site at a polite rate and stores only the URLs and metadata the site publishes. You are responsible for using the results in compliance with the laws that apply to you.

#### Will the output fields change between runs?

No. Output fields are stable: existing fields are never renamed or removed without a major version bump announced in the changelog, and new fields are only ever added. You can build integrations on the schema without checking it after every run.

### Integrate Best Damn Sitemap URL Extractor and automate your workflow

Best Damn Sitemap URL Extractor plugs into the tools you already use through [Apify integrations](https://docs.apify.com/platform/integrations), so results can flow on without anyone downloading a file. Ready-made connectors include:

- [Make](https://docs.apify.com/platform/integrations/make)
- [Zapier](https://docs.apify.com/platform/integrations/zapier)
- [n8n](https://docs.apify.com/platform/integrations/n8n)
- [Slack](https://docs.apify.com/platform/integrations/slack)
- [Airbyte](https://docs.apify.com/platform/integrations/airbyte)
- [GitHub](https://docs.apify.com/platform/integrations/github)
- [Google Drive](https://docs.apify.com/platform/integrations/drive)
- and [many more](https://docs.apify.com/platform/integrations).

You can also attach [webhooks](https://docs.apify.com/platform/integrations/webhooks) to trigger your own endpoint whenever a run succeeds, fails or times out. For example, feed newly published URLs into your crawler or indexing workflow, or notify Slack when a competitor adds pages.

### Related Actors by the same developer

- [Best Damn Tech Stack Detector](https://apify.com/josh99smith/tech-stack-detector): find out what a website is built with.
- [Best Damn Website Screenshot API](https://apify.com/josh99smith/website-screenshot-api): full-page screenshots and PDFs of any URL.
- [Best Damn Google Autocomplete Scraper](https://apify.com/josh99smith/google-autocomplete-scraper): keyword suggestions from Google search.
- [Best Damn App Reviews Scraper](https://apify.com/josh99smith/app-reviews-scraper): app reviews from both stores.
- [Best Damn PageSpeed Insights Audit](https://apify.com/josh99smith/pagespeed-insights-audit): Core Web Vitals via Google's API.
- [Best Damn Remote Jobs Aggregator](https://apify.com/josh99smith/remote-jobs-aggregator): remote job listings in one dataset.
- [Best Damn PDF Text Extractor](https://apify.com/josh99smith/pdf-text-extractor): text and metadata from PDF URLs.
- [Best Damn RSS to JSON Converter](https://apify.com/josh99smith/rss-feed-to-json): RSS, Atom and JSON feeds as JSON items.
- [Best Damn YouTube Comments Scraper](https://apify.com/josh99smith/youtube-comments-scraper): comments and replies from YouTube videos and channels.
- [Best Damn YouTube Scraper](https://apify.com/josh99smith/youtube-scraper): videos, channels, playlists and search results with statistics.

### Support and feedback

Found a sitemap that is not parsed correctly? Open a ticket in the **Issues** tab of this Actor with the sitemap URL and we will look into it.

This Actor is open source under the MIT licence.

The full source code is on GitHub: [josh99smith/sitemap-url-extractor](https://github.com/josh99smith/sitemap-url-extractor). Stars and pull requests are welcome.

# Changelog

This Actor's version history is a separate document: https://apify.com/josh99smith/sitemap-url-extractor/changelog.md

# Actor input Schema

## `urls` (type: `array`):

One URL per line. A website URL (`https://example.com`) triggers sitemap discovery via robots.txt and common locations such as `/sitemap.xml`; a sitemap URL (`https://example.com/sitemap_index.xml` or `.xml.gz`) is read directly. Duplicates are removed automatically.

## `maxUrlsPerSite` (type: `integer`):

Stop reading the sitemaps of a site once this many page URLs have been extracted. Caps the cost of very large sites.

## `includePatterns` (type: `array`):

Optional. Keep only URLs matching at least one pattern. Use a glob (`**/blog/**`, `*.pdf`), a regular expression wrapped in slashes (`/\/products\/\d+/`) or a plain substring (`/docs/`). Leave empty to keep every URL.

## `excludePatterns` (type: `array`):

Optional. Drop URLs matching any pattern (same syntax as include patterns). Applied after the include patterns.

## `maxSitemapDepth` (type: `integer`):

How many levels of nested sitemap index files to follow. 0 reads only the sitemap files found directly; 1 also reads the sitemaps they reference, and so on.

## `maxSitemapsPerSite` (type: `integer`):

Stop discovering more sitemap files after this many were fetched. Large publishers expose thousands of monthly sitemaps; raise this (and the run memory) if you need the whole tree.

## `outputSitemapsOnly` (type: `boolean`):

Output one record per discovered sitemap file (with its type and URL count) instead of the page URLs themselves. Useful for auditing the sitemap structure of a site cheaply.

## `timeoutSecs` (type: `integer`):

Give up on a single sitemap file after this many seconds. Large compressed sitemaps on slow servers may need more.

## `proxyConfiguration` (type: `object`):

Optional. Route requests through Apify Proxy if a site blocks cloud IP addresses. Proxy usage is billed by Apify separately from the per-URL price of this Actor.

## Actor input object example

```json
{
  "urls": [
    "https://www.apify.com",
    "https://blog.apify.com/sitemap.xml"
  ],
  "maxUrlsPerSite": 5000,
  "includePatterns": [],
  "excludePatterns": [],
  "maxSitemapDepth": 5,
  "maxSitemapsPerSite": 500,
  "outputSitemapsOnly": false,
  "timeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One record per page URL found in the sitemaps (or per sitemap file in sitemaps-only mode), plus free failure records for sitemaps that could not be loaded.

## `summary` (type: `string`):

Counts of requested sites, sitemap files found, URLs extracted and billed, and failures.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.apify.com",
        "https://blog.apify.com/sitemap.xml"
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("josh99smith/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://www.apify.com",
        "https://blog.apify.com/sitemap.xml",
    ],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("josh99smith/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.apify.com",
    "https://blog.apify.com/sitemap.xml"
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call josh99smith/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,josh99smith/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/DdpfO4s9oTtFiZafi/builds/BOrReajsbpwcW20VF/openapi.json
