# Sitemap Extractor – All URLs from sitemap.xml + Status Check (`gazidev/sitemap-url-extractor`) Actor

Extract every URL from any website's sitemap.xml: auto-discovery via robots.txt, nested sitemap indexes, .xml.gz, image/news/video/hreflang tags and lastmod. Filter by date or glob, get only new URLs, and optionally check HTTP status, redirects and canonical per URL.

- **URL**: https://apify.com/gazidev/sitemap-url-extractor.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.15 / 1,000 url extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap Extractor – All URLs from sitemap.xml + Status Check

Enter a domain and get back **every URL from its sitemap**. The Actor finds the sitemap itself: it reads the **robots.txt `Sitemap:` lines**, then tries `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml` and `/sitemap.xml.gz`. It follows **nested sitemap indexes**, unpacks **gzip (.xml.gz)** files and reads **image, news, video and hreflang** tags as well as **lastmod**, changefreq and priority. You can also turn on an **HTTP status check** for every URL to find 404s, redirect chains, non-canonical pages and noindex pages listed in your sitemap.

- **Built to succeed.** It retries on 429/5xx, uses a regex fallback for broken XML, skips HTML "soft 404" pages served as `sitemap.xml`, detects index loops and records an error per sitemap instead of failing the run. A site with no sitemap is reported clearly, not crashed on.
- **Every sitemap format.** Supports XML `urlset` and `sitemapindex` (up to 10 levels deep), `.xml.gz` (gzip payload or transfer encoding), plain-text `sitemap.txt`, and RSS/Atom feeds used as sitemaps. Google **image** (`image:loc`), **news** (title, publication, language, date), **video** (title, thumbnail, content/player URL, duration) and **hreflang** `xhtml:link` alternates are all extracted.
- **Bulk URL status checker.** It sends HEAD requests and falls back to GET when a server rejects HEAD. For each URL you get `status`, `finalUrl`, `redirectHops` and `redirectChain`, `contentType`, `responseMs` and `X-Robots-Tag`, and optionally the **canonical URL** and **meta robots** tag. It runs 20 requests in parallel but never more than 5 per host.
- **Filters.** Keep URLs modified since a date (`2025-01-31` or `7 days`), or filter by include/exclude globs (`/blog/*`, `*/products/*`, `*.pdf`).
- **"Only new URLs" monitoring.** The Actor remembers which URLs it has already returned for each site, so on a schedule each run returns only **new pages**: new products, articles or landing pages from a competitor.
- **Cheap.** It uses plain HTTP with no browser and no proxy. **$0.15 per 1,000 URLs**, about 3× cheaper than the official Apify sitemap extractor.

### What can I use it for?

- **SEO audits:** list every indexable URL, then find 404s, redirect chains and pages that canonicalise elsewhere, all of which are sitemap errors that waste crawl budget
- **Competitor monitoring:** schedule a daily run with *Only new URLs* to see every new product, blog post or landing page a competitor publishes
- **Crawler seeding:** feed a complete URL list to [Website Content Crawler](https://apify.com/apify/website-content-crawler), a RAG pipeline or your own scraper instead of discovering pages by link-following
- **Site migrations:** export all old URLs with their status before a relaunch, then check that each one returns a 200 or a single 301 afterwards
- **News and content tracking:** read Google News sitemaps (title, publication, date) from publishers in near real time
- **Image and video inventory:** list every image and video URL that a site exposes to Google
- **International SEO:** export the hreflang alternates of every page to audit language versions

### Input

| Field | Description |
|---|---|
| `startUrls` | Domains (`example.com`), sitemap URLs (`…/sitemap_index.xml`, `.xml.gz`, `.txt`, feeds) or `robots.txt` URLs, in any mix |
| `maxUrls` | Stop after this many URLs in total (default 1,000, 0 = everything) |
| `maxUrlsPerSite` | Optional per-site limit |
| `checkStatus` | HTTP status, final URL, redirects, content type, response time and X-Robots-Tag per URL |
| `extractCanonical` | Also read `<link rel="canonical">` and meta robots (GET, first 64 KB only) |
| `includeGlobs` / `excludeGlobs` | URL patterns to keep or drop |
| `lastmodSince` | Only URLs with lastmod on or after a date (absolute or relative) |
| `keepUrlsWithoutLastmod` | Keep URLs that have no lastmod when a date filter is set |
| `onlyNew`, `stateKey` | Return only URLs not seen in previous runs (monitoring) |
| `includeSiteSummaryRows`, `discoverCommonPaths`, `maxSitemapDepth`, `maxConcurrency`, `statusConcurrency`, `requestTimeoutSecs`, `proxyConfiguration` | Advanced |

```json
{
  "startUrls": ["allbirds.com", "https://www.nytimes.com/sitemaps/new/news.xml.gz"],
  "maxUrls": 5000,
  "checkStatus": true,
  "excludeGlobs": ["*/collections/*"],
  "lastmodSince": "30 days"
}
```

### Output

Each URL is one row. The Output tab has three tables, **Sitemap URLs**, **HTTP status check** and **Images, videos, news & hreflang**, plus a **per-site summary**: sitemaps found, URL counts, newest and oldest lastmod, a status-code histogram and any failed sitemaps. Here is a real row (shortened) with the status check on:

```json
{
  "url": "https://www.allbirds.com/products/mens-wool-runners",
  "site": "allbirds.com",
  "sitemapUrl": "https://www.allbirds.com/sitemap_products_1.xml?from=1878193471557&to=7369944137808",
  "lastmod": "2026-10-01T01:54:12-07:00",
  "changefreq": "daily",
  "priority": null,
  "images": ["https://cdn.shopify.com/s/files/1/1104/4168/files/Allbirds_WL_RN_SF_PDP_Natural_Grey_LAT.png?v=1751143404"],
  "imageCount": 1,
  "alternates": [],
  "videos": [],
  "newsTitle": null,
  "status": 200,
  "finalUrl": "https://www.allbirds.com/products/mens-wool-runners",
  "redirectHops": 0,
  "redirectChain": [],
  "contentType": "text/html",
  "responseMs": 412,
  "xRobotsTag": null,
  "statusError": null
}
```

A news sitemap row adds `newsTitle`, `newsPublicationDate`, `newsPublicationName` and `newsLanguage`. A video sitemap row has `videos: [{title, thumbnailUrl, contentUrl, playerUrl, durationSecs}]`. See [SAMPLE_OUTPUT.json](SAMPLE_OUTPUT.json) for real results from 6 sites.

### Pricing

Pay per event:

| Event | Price |
|---|---|
| URL extracted (one row in the dataset) | **$0.00015** ($0.15 per 1,000 URLs) |
| URL status checked (only with `checkStatus`) | **$0.0004** ($0.40 per 1,000 URLs) |
| Actor start | $0.00005 |
| Duplicates, filtered-out URLs, sitemaps not found, summary | free |

How that compares with other sitemap and status tools in the Apify Store (listed prices, October 2026):

| Actor | Listed price | 30-day run success (Store) |
|---|---|---|
| apify/sitemap-extractor | $0.50 / 1,000 URLs | 69% |
| crawlerbros/sitemap-url-extractor | $2 / 1,000 URLs | 63% |
| onescales sitemap extractor | $30 / 1,000 URLs | – |
| logiover URL status checker | $2.50 / 1,000 URLs | – |
| khadinakbar URL status checker | $1 / 1,000 URLs | – |
| **This Actor** | **$0.15 / 1,000 URLs, or $0.55 with status check** | new |

Set a **maximum cost per run** in the run options. The Actor saves as many URLs as fit and then stops cleanly. It never runs a status check that it cannot bill.

### FAQ

**The site has no sitemap. What happens?**
The summary says *No sitemap found* and you pay nothing for that site. This Actor does not crawl links. For sites without a sitemap, use a crawler instead.

**Does it respect robots.txt?**
Sitemaps and robots.txt are published specifically for machines to read, and the Actor uses robots.txt only to discover sitemaps. The optional status check makes one lightweight HEAD (or GET) request per URL, with at most 5 parallel requests per host.

**How big a sitemap can it handle?**
The protocol maximum of 50,000 URLs or 50 MB per file, gzipped or not, and thousands of child sitemaps per index. The XML is stream-parsed with lxml, so memory stays low. Set `maxUrls: 0` to extract everything.

**Why do some URLs have no `lastmod`?**
The site's sitemap does not provide one. Many Shopify and Next.js sitemaps omit it for some pages. With `lastmodSince` set these URLs are dropped unless you enable `keepUrlsWithoutLastmod`.

**How does "Only new URLs" work?**
A short hash of every URL already returned is stored per site in a named key-value store (`sitemap-url-extractor-state`) in your account. The next run returns only URLs that are not in that store. The first run returns everything. Use `stateKey` to keep separate histories.

**HEAD or GET?**
HEAD by default, because it is fastest. If a server answers HEAD with 400, 403, 404, 405, 406 or 501, the URL is retried with GET so you get the real status. `extractCanonical` always uses GET and reads at most 64 KB per page.

**Can I get the results as CSV or Excel?**
Yes. Every Apify dataset can be exported to CSV, Excel, JSON or XML, or sent to Google Sheets with an integration.

### Use with AI agents / Apify MCP

This Actor makes a good agent tool. The input is tiny (a domain), the output is a clean URL list, and 1,000 URLs cost $0.15. Connect it through the [Apify MCP server](https://mcp.apify.com) (`https://mcp.apify.com?actors=gazidev/sitemap-url-extractor`) and Claude, ChatGPT, Cursor and other MCP clients can call it with prompts such as *"list all blog posts example.com published in the last 30 days"* or *"find broken URLs in competitor.com's sitemap"*. Over the API: `POST https://api.apify.com/v2/acts/gazidev~sitemap-url-extractor/run-sync-get-dataset-items` with the input JSON.

### Categories

SEO tools · Developer tools · Automation

# Actor input Schema

## `startUrls` (type: `array`):

Any mix of: a domain (`example.com` – the sitemap is found automatically via robots.txt `Sitemap:` lines, then `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml`, `/sitemap.xml.gz`…), a direct sitemap URL (`https://example.com/sitemap_index.xml`, `.xml.gz`, `.txt`, RSS/Atom feed) or a `robots.txt` URL. Sitemap indexes are followed recursively.

## `maxUrls` (type: `integer`):

Stop after saving this many URLs in total (0 = no limit, extract everything). You pay only for URLs saved.

## `maxUrlsPerSite` (type: `integer`):

Limit per website when you give several sites (0 = no per-site limit).

## `checkStatus` (type: `boolean`):

Request each URL (HEAD, falling back to GET when the server refuses HEAD) and add `status`, `finalUrl`, `redirectHops`, `redirectChain`, `contentType`, `responseMs` and `X-Robots-Tag`. Finds 404s, redirects and dead pages listed in your sitemap. Charged per URL checked.

## `extractCanonical` (type: `boolean`):

Uses GET instead of HEAD and reads the first 64 KB of each HTML page to add `canonical`, `canonicalMatchesUrl` (self-canonical?) and `metaRobots` (noindex…). Needs `Check HTTP status` on.

## `statusConcurrency` (type: `integer`):

Parallel status requests overall. At most 5 requests per host run at the same time, so target servers are not hammered.

## `includeGlobs` (type: `array`):

Keep only URLs matching at least one pattern. `*` matches anything. Patterns without a scheme match anywhere in the URL, e.g. `/blog/*`, `*/products/*`, `https://example.com/de/*`.

## `excludeGlobs` (type: `array`):

Drop URLs matching any of these patterns, e.g. `*/tag/*`, `*.pdf`, `*?page=*`.

## `lastmodSince` (type: `string`):

Keep only URLs whose `<lastmod>` (or news publication date / feed date) is on or after this date. Absolute (`2025-01-31`) or relative (`7 days`, `24 hours`, `3 months`).

## `keepUrlsWithoutLastmod` (type: `boolean`):

When `Only URLs modified since` is set, also keep URLs that have no `<lastmod>` at all (by default they are dropped).

## `onlyNew` (type: `boolean`):

Remembers the URLs already returned for each website (in a named key-value store) and outputs only URLs not seen in previous runs. Schedule it daily to monitor competitors' new pages, products or articles. The first run returns everything.

## `stateKey` (type: `string`):

Optional. Use different names to keep separate 'seen URLs' memories for different tasks watching the same site.

## `includeSiteSummaryRows` (type: `boolean`):

The per-site summary (sitemaps found, URL counts, newest/oldest lastmod, status-code histogram, failed sitemaps) is always saved to the key-value store record `SUMMARY`. Turn this on to also append it to the dataset as rows with `type: "site-summary"` (free).

## `discoverCommonPaths` (type: `boolean`):

If robots.txt lists no (working) sitemap, try `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml`, `/sitemap-index.xml`, `/sitemap.xml.gz` and `/sitemap.txt`.

## `maxSitemapDepth` (type: `integer`):

How deep nested sitemap indexes are followed (index → index → sitemap). Loops are detected automatically.

## `maxConcurrency` (type: `integer`):

How many websites are processed at the same time (each reads up to 3 sitemaps in parallel).

## `requestTimeoutSecs` (type: `integer`):

Per-request timeout. Failed requests are retried.

## `proxyConfiguration` (type: `object`):

Optional. Sitemaps are public and rarely blocked; only use a proxy for sites that block datacenter IPs.

## Actor input object example

```json
{
  "startUrls": [
    "https://crawlee.dev"
  ],
  "maxUrls": 50,
  "maxUrlsPerSite": 0,
  "checkStatus": false,
  "extractCanonical": false,
  "statusConcurrency": 20,
  "keepUrlsWithoutLastmod": false,
  "onlyNew": false,
  "includeSiteSummaryRows": false,
  "discoverCommonPaths": true,
  "maxSitemapDepth": 5,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `urls` (type: `string`):

No description

## `status` (type: `string`):

No description

## `media` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://crawlee.dev"
    ],
    "maxUrls": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["https://crawlee.dev"],
    "maxUrls": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("gazidev/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://crawlee.dev"
  ],
  "maxUrls": 50
}' |
apify call gazidev/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/AP8eFMasrzd7nWF9X/builds/9cdBqRMrcSYniZUaR/openapi.json
