# Sitemap URL Extractor - All URLs from sitemap.xml & robots.txt (`tenfoldfleet/sitemap-url-extractor`) Actor

Get all URLs of a website from its XML sitemap: finds sitemaps via robots.txt and common paths, walks sitemap indexes, reads .gz files and extracts every page URL with lastmod, priority, images, news and hreflang. HTTP-only. $0.40 per 1,000 URLs; duplicates and failed sites are free.

- **URL**: https://apify.com/tenfoldfleet/sitemap-url-extractor.md
- **Developed by:** [Tenfold Fleet](https://apify.com/tenfoldfleet) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.40 / 1,000 url extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Sitemap URL Extractor do?

**Sitemap URL Extractor** gets **every URL from a website's XML sitemaps**. Give it a domain, a homepage or a direct sitemap link and it returns one row per page with its **lastmod date, change frequency, priority, images, Google News data and hreflang alternates**.

It finds sitemaps the way search engines do: it reads the `Sitemap:` lines in **robots.txt**, then tries `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml` and `/sitemap-index.xml`. It walks **sitemap index files** recursively (up to 6 levels), reads **gzipped `.xml.gz` sitemaps**, plain-text sitemaps and RSS/Atom feeds, and removes duplicate URLs.

It's a fast, low-cost **sitemap scraper API**: plain HTTP, no browser, **$0.40 per 1,000 URLs**. Sites without a sitemap, broken sitemaps and duplicates are free.

What you get that basic sitemap scrapers don't: a **lastmod date filter** that skips stale child sitemaps entirely, **Google News, image and hreflang** data, **free error rows** that tell you why a site returned nothing, and **deduplication across inputs** (www and non-www, domain plus sitemap URL).

### Why use a sitemap extractor?

- 🔎 **SEO audits**: list every indexable page of a site, spot stale `lastmod` dates and missing hreflang alternates.
- 🕷️ **Crawl seeding**: feed a complete URL list into a scraper (Website Content Crawler, Cheerio Scraper) instead of crawling links.
- 🆕 **Content monitoring**: with **Only URLs modified after**, get just the pages a competitor or publisher added or changed since a date.
- 🛒 **Ecommerce research**: pull all product or category URLs of a Shopify, WooCommerce or Magento store with an include pattern like `/products/`.
- 📰 **News tracking**: read Google News sitemaps with headline and publication time.
- 🔁 **Migrations and redirects**: export the old site's URLs before a relaunch.

Runs on the Apify platform: you get an **API**, scheduling, webhooks and integrations with Make, Zapier, Google Sheets and more.

### What data can it extract?

| Field | Description |
|---|---|
| `url` | Page URL (`<loc>`) |
| `lastmod` | Last modification date |
| `changefreq` | Change frequency hint |
| `priority` | Priority 0.0-1.0 |
| `images` | Image sitemap entries: `loc`, `title` |
| `news` | Google News: `title`, `publicationDate` |
| `alternates` | hreflang alternates: `hreflang`, `href` |
| `sourceSitemap` | The sitemap file the URL came from |
| `site` | Input website |
| `depth` | Sitemap index depth (0 = top-level sitemap) |

### How to extract all URLs from a sitemap

1. Open **Sitemap URL Extractor** and click **Try for free**.
2. Add websites (`example.com`) or sitemap URLs (`https://example.com/sitemap_index.xml`).
3. Optionally set **Max URLs per site**, include/exclude regex patterns or a **lastmod** date.
4. Click **Start** and download the URLs as **JSON, CSV, Excel or HTML**, or fetch them through the API.

### How much does it cost to extract sitemap URLs?

Pay per result: **$0.40 per 1,000 unique URLs** ($0.0004 per URL). You are **not charged** for duplicate URLs, sitemap index entries, websites without a sitemap, 404 or timeout errors, or malformed XML. Set a maximum cost per run and the actor stops cleanly when it is reached. Platform usage is small because it uses plain HTTP requests.

### Input

```json
{
    "startUrls": ["apify.com", "https://wordpress.org/news/sitemap.xml"],
    "maxUrlsPerSite": 10000,
    "includePatterns": ["/blog/"],
    "excludePatterns": ["/tag/"],
    "lastmodAfter": "2026-01-01",
    "includeExtensions": true
}
```

### Output

```json
{
    "url": "https://www.nytimes.com/interactive/2026/09/29/us/quake-tracker-washington-seattle.html",
    "lastmod": "2026-10-01T05:03:58Z",
    "changefreq": null,
    "priority": null,
    "images": [{ "loc": "https://static01.nyt.com/images/2026/09/29/...-articleLarge-v80.jpg", "title": null }],
    "news": { "title": "Map: 4.2-Magnitude Earthquake Strikes Near Seattle", "publicationDate": "2026-09-29T21:55:55Z" },
    "alternates": [],
    "sourceSitemap": "https://www.nytimes.com/sitemaps/new/news.xml.gz",
    "site": "nytimes.com",
    "depth": 0
}
```

Sites that fail produce a free row such as `{ "site": "example.com", "error": "No sitemap found: ...", "status": 404 }`, so you always know why a site returned nothing. Every row has the same keys: URL rows have `error` and `status` set to `null`, error rows have `url` set to `null`, so CSV and spreadsheet exports keep fixed columns. Download results as JSON, CSV, Excel, XML or HTML.

### Tips

- Huge sites (news publishers, marketplaces) can list millions of URLs: use **Max URLs per site** and patterns to keep runs small.
- With **Only URLs modified after**, child sitemaps whose own `lastmod` is older are skipped entirely, which makes monitoring runs fast and cheap.
- With strict filters on a huge site, the actor stops a site after about 100 + maxUrlsPerSite / 50 sitemap files that contain no matching URL and returns a free note row. Start from a specific sitemap URL (for example the blog or news sitemap) to go deeper.
- Limits per site: up to 1,000,000 URLs, 30 minutes, 2 GB of sitemap downloads and 60 MB per sitemap file. A site that times out or fails 5 times in a row is stopped. Each limit that cuts a site short adds a free note row.
- A large run (several sites with hundreds of thousands of URLs each) needs more memory: give it 2 GB or more. If deduplication fills the memory, the actor stops the site with a free note row instead of crashing.
- Gzip is detected from the file content, so `.xml.gz` sitemaps work whether or not the server already decompressed them.
- If a site blocks datacenter IPs, switch the proxy to a residential group.

### FAQ

#### Does it crawl the website pages?

No. It reads only robots.txt and sitemap files, which is why it is fast and cheap. Use a crawler if the site has no sitemap.

#### What if robots.txt lists no sitemap?

It tries the common sitemap paths. If none exists you get a free error row.

#### Can I give it a single sitemap file?

Yes. Any URL ending in `.xml`, `.xml.gz`, `.txt` or containing "sitemap" is read directly, including sitemap index files.

#### Is it legal to extract sitemaps?

Sitemaps are public files that sites publish for crawlers. This actor only reads public data. Make sure your use of the URLs complies with the target site's terms.

### More tools from Tenfold Fleet

| Actor | Price |
|---|---|
| [ATS Jobs Scraper - Greenhouse, Lever, Ashby & 5 More](https://apify.com/tenfoldfleet/ats-jobs-scraper) | $2 per 1,000 (job posting) |
| [Website Contact Scraper - Emails, Phones & Socials](https://apify.com/tenfoldfleet/company-contact-finder) | $8 per 1,000 (website with contacts) |
| [Google Ads Transparency Center Scraper - Competitor Ads by Site](https://apify.com/tenfoldfleet/google-ads-transparency-scraper) | $1 per 1,000 (ad scraped) |
| [Google News Scraper - Articles, Real URLs & Full Text](https://apify.com/tenfoldfleet/google-news-scraper) | $2 per 1,000 (article) |
| [Keyword Search Volume, CPC & Difficulty Checker (Bulk)](https://apify.com/tenfoldfleet/keyword-search-volume) | $5 per 1,000 (keyword with data) |
| [Remote Jobs Aggregator - RemoteOK, WWR, Himalayas & 4 more](https://apify.com/tenfoldfleet/remote-jobs-aggregator) | $1.2 per 1,000 (unique remote job) |
| [SEO Audit Tool - Technical On-Page SEO Checker & Score](https://apify.com/tenfoldfleet/seo-audit-tool) | $20 per 1,000 (page audited) |
| [Shopify Products Scraper - Prices, Variants & Stock](https://apify.com/tenfoldfleet/shopify-products-scraper) | $0.8 per 1,000 (product) |
| [Shopify App & Theme Detector - Store Analyzer & Contacts](https://apify.com/tenfoldfleet/shopify-store-analyzer) | $8 per 1,000 (shopify store analyzed) |
| [Technology Detector - Wappalyzer & BuiltWith Alternative](https://apify.com/tenfoldfleet/tech-stack-detector) | $10 per 1,000 (website analyzed) |
| [URL to Markdown - Web Page to LLM-Ready Text for AI Agents](https://apify.com/tenfoldfleet/url-to-markdown) | $1 per 1,000 (page converted) |
| [Workday Jobs Scraper - myworkdayjobs.com API with Salary](https://apify.com/tenfoldfleet/workday-jobs-scraper) | $1.5 per 1,000 (job posting) |
| [YouTube Comments Scraper - Replies, Likes & Sort by Newest](https://apify.com/tenfoldfleet/youtube-comments-scraper) | $0.6 per 1,000 (comment scraped) |
| [YouTube Transcript Scraper - Captions & Subtitles API](https://apify.com/tenfoldfleet/youtube-transcript-scraper) | $3 per 1,000 (transcript extracted) |

# Actor input Schema

## `startUrls` (type: `array`):

Domains (example.com), homepages or direct sitemap URLs (https://example.com/sitemap_index.xml, .xml.gz or .txt). For a domain, sitemaps are discovered from robots.txt, then /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml and /sitemap-index.xml.

## `maxUrlsPerSite` (type: `integer`):

Stop collecting a site after this many unique URLs (up to 1,000,000). The cap is shared by all inputs of the same site. You only pay for URLs actually returned.

## `includePatterns` (type: `array`):

Only return URLs matching at least one of these regular expressions (case-insensitive), e.g. /blog/ or /products/.

## `excludePatterns` (type: `array`):

Drop URLs matching any of these regular expressions (case-insensitive), e.g. /tag/ or page=\d+.

## `lastmodAfter` (type: `string`):

Return only URLs whose <lastmod> is on or after this date (YYYY-MM-DD or ISO). URLs without a lastmod are skipped when this is set. Child sitemaps older than the date are not downloaded.

## `includeExtensions` (type: `boolean`):

Also return image sitemap entries (loc, title), Google News data (title, publication date) and hreflang alternates for each URL.

## `maxConcurrency` (type: `integer`):

How many websites to process in parallel.

## `proxyConfiguration` (type: `object`):

Apify datacenter proxy is used by default. Turn it off to connect directly.

## Actor input object example

```json
{
  "startUrls": [
    "apify.com",
    "https://wordpress.org/news/sitemap.xml"
  ],
  "maxUrlsPerSite": 500,
  "includePatterns": [],
  "excludePatterns": [],
  "includeExtensions": true,
  "maxConcurrency": 10,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "apify.com",
        "https://wordpress.org/news/sitemap.xml"
    ],
    "maxUrlsPerSite": 500,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("tenfoldfleet/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        "apify.com",
        "https://wordpress.org/news/sitemap.xml",
    ],
    "maxUrlsPerSite": 500,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("tenfoldfleet/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "apify.com",
    "https://wordpress.org/news/sitemap.xml"
  ],
  "maxUrlsPerSite": 500,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call tenfoldfleet/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tenfoldfleet/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XA7Va8lNUE6pyicpr/builds/usibIYsgv3xByfdyv/openapi.json
