# Sitemap URL Extractor – All Pages for RAG & SEO (`oldjard/sitemap-url-extractor`) Actor

Get every URL a website publishes: XML sitemaps, sitemap indexes, gzip and news sitemaps, RSS/Atom feeds, plus a shallow crawl when there's no sitemap. lastmod, source and robots.txt status per URL. Feed it to your RAG pipeline, crawler or SEO audit. $0.40 per 1,000 URLs.

- **URL**: https://apify.com/oldjard/sitemap-url-extractor.md
- **Developed by:** [Joshua White](https://apify.com/oldjard) (community)
- **Categories:** Developer tools, AI, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor – All Pages for RAG & SEO

**Sitemap URL Extractor** lists **every page a website publishes**, from its XML sitemaps, nested sitemap indexes,
gzip and Google News sitemaps and RSS/Atom feeds, and falls back to a polite shallow crawl when there's no sitemap.
One clean row per URL with its last-modified date, where it was found and whether robots.txt allows crawling it.
Built for **RAG pipelines, AI agents and SEO audits**.

**Try it in one click:** the input is prefilled with two docs sites, 100 URLs each. Apify's free plan covers about
12,000 URLs a month.

### How to extract all URLs from a website in 3 steps

1. Add websites, one per line (or separated by commas): `example.com`, a section like `https://example.com/blog/`,
   or a sitemap URL.
2. Optional: set **Max URLs per site**, include or exclude patterns (`/blog/*`, `*.pdf`), or **Modified since**
   (`7 days`) to get only changed pages.
3. Click **Start**, then download CSV, Excel or JSON, or pass the list straight to a content crawler.

### How much does it cost to extract sitemap URLs?

**$0.40 per 1,000 URLs** returned, after your filters and dedupe, so Apify's $5 monthly free credit covers about 12,000 URLs. Reading sitemaps is free, filtered-out URLs are
free, sites that fail are free, and rows for entries that aren't web addresses are free. A 5,000-page site costs $2. **Max URLs per site** and the run's maximum cost
both cap what you spend; the run stops cleanly when the cap is reached.

### What the sitemap extractor handles

- **Finds every sitemap:** `Sitemap:` lines in robots.txt, `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml`,
  `<link rel="sitemap">`, and a section's own `sitemap.xml` (e.g. `/docs/sitemap.xml`). It follows **nested sitemap
  indexes** (loops are detected) and reads **`.xml.gz` gzip sitemaps**, Google News sitemaps and plain-text ones.
- **Reads RSS and Atom feeds** linked from the site. Feeds often list the newest posts first, with a title and
  publication date.
- **Falls back to a shallow crawl** when a site has no sitemap or feed (or always, if you ask): it follows the site's
  own links a few clicks deep, opens only HTML pages and honours robots.txt.
- **Streams huge sitemaps.** 50,000-URL sitemaps and sites with thousands of sitemaps are read as they download, never
  loaded whole into memory, and reading stops as soon as your limit is reached.
- **Doesn't fall over on messy sitemaps:** malformed XML, unescaped `&`, unclosed tags, HTML error pages served as
  `sitemap.xml`, truncated gzip files, relative URLs and odd date formats are all handled.
- **Only a section:** start from `https://docs.example.com/guide` to get just that part of the site. If the address
  redirects (`stripe.com/docs` → `docs.stripe.com`), the section follows it.
- **Filters:** include and exclude patterns (globs like `/blog/*` or `*.pdf`, or regular expressions), **modified
  since** a date (only pages that changed, for refreshing an index), and a max number of URLs per site.
- **Clean output:** fragments and `utm_` tracking parameters are removed, and duplicates are dropped, including
  `http`/`https`, `www`/no-`www` and trailing-slash variants.
- **robots.txt status per URL:** each row says whether robots.txt lets crawlers (`User-agent: *`) fetch that page, so
  your downstream crawler can skip what it shouldn't touch.

### Input example

```json
{
    "startUrls": ["docs.apify.com", "https://www.theverge.com", "https://docs.python.org/3/"],
    "maxUrlsPerSite": 5000,
    "includePatterns": [],
    "excludePatterns": ["/tag/*", "/author/*"],
    "modifiedSince": "30 days",
    "includeFeeds": true,
    "crawlMode": "fallback"
}
```

### Output example

One row per URL. `lastmod` is ISO 8601 (UTC) when the site gives a date; `title` comes from feeds and news sitemaps.

```json
[
    {
        "url": "https://techcrunch.com/2026/10/05/at-19-ghost-founder-raises-11-million-to-build-a-3499-computer-for-your-personal-ai/",
        "site": "techcrunch.com",
        "sourceType": "sitemap",
        "source": "https://techcrunch.com/news-sitemap.xml",
        "lastmod": "2026-10-05T18:07:07.000Z",
        "changefreq": null,
        "priority": null,
        "title": "At 19, founder raises $11 million for Ghost, maker of a $3,499 computer for personal AI",
        "allowedByRobots": true
    },
    {
        "url": "https://developer.mozilla.org/en-US/docs/Web/HTTP",
        "site": "https://developer.mozilla.org/en-US/docs/Web/HTTP",
        "sourceType": "sitemap",
        "source": "https://developer.mozilla.org/sitemaps/en-us/sitemap.xml.gz",
        "lastmod": "2026-07-28T00:00:00.000Z",
        "changefreq": null,
        "priority": null,
        "title": null,
        "allowedByRobots": true
    },
    {
        "url": "https://docs.python.org/3/download.html",
        "site": "https://docs.python.org/3/",
        "sourceType": "crawl",
        "source": "https://docs.python.org/3/",
        "lastmod": null,
        "changefreq": null,
        "priority": null,
        "title": null,
        "allowedByRobots": true
    }
]
```

| Field | Meaning |
|---|---|
| `url` | The page's address (fragment and `utm_` parameters removed). |
| `site` | The website from your input that this URL belongs to. |
| `sourceType` | `sitemap`, `rss`, `atom` or `crawl`: how the URL was found. |
| `source` | The exact sitemap, feed or page it was found in. |
| `lastmod` | Last-modified or publication date, if the site gives one. |
| `changefreq`, `priority` | As declared in the sitemap, if present. |
| `title` | Page title, from feeds and Google News sitemaps. |
| `allowedByRobots` | Whether robots.txt allows generic crawlers to fetch the URL (`null` if unknown). |
| `errorCode`, `error` | `null` on URL rows. See below. |

#### Entries that aren't web addresses

Every input entry is accounted for. An entry that isn't a domain or URL (say `docs apify com`) gets one free row with
`errorCode: "INVALID_INPUT"`, `error: "Not a web address (expected e.g. example.com or https://example.com)"`,
`site` set to the entry as you typed it and `url` set to `null`. The status message counts them too: *Found 200 URLs
on 2 of 3 entries; 1 wasn't a web address.* Common typos like `htps://` and `http//` are fixed for you.

#### Run report

The run's key-value store (Storage → Key-value store) has an **OUTPUT** record with one entry per site: how many URLs, which sitemaps and feeds
were read, how many URLs were out of scope, filtered out or duplicates, and **why a site gave no URLs** (no sitemap,
blocked, robots.txt, nothing under the section you chose). A site that can't be read doesn't fail the run; it's
reported there and in the log. The URL counts in the status message and OUTPUT are exactly the rows in the dataset,
which are exactly the URLs you're charged for.

### Good to know

- **Polite by design.** It identifies itself as `SitemapURLsBot`, honours robots.txt (including Crawl-delay, up to
  10 s), sends at most 3 requests at a time to a site, and backs off on HTTP 429 and 5xx. If a site's robots.txt can't
  be read because the server errors, the site is treated as off-limits, as the robots standard says.
- **Public pages only.** It never logs in and never fills forms.
- **JavaScript-only sites:** sitemaps and feeds work regardless of how the site is built. The crawl fallback reads
  the HTML the server sends, so a site that builds its links only in the browser and has no sitemap gives few URLs.
- **Very large sites:** the default 512 MB of memory is enough for big sites: 150,000 gov.uk URLs were tested in 37
  seconds (peak 316 MB). Give the run 512 MB or more for very large sites, and 1–2 GB if you set **Max URLs per
  site** to 0 (no limit) on a site well past that size.
- **Speed:** a typical site with a sitemap takes 1–5 seconds; a site with hundreds of sitemaps takes a minute or two.

### Ready-made examples

Each one opens this actor with the input already filled in. Click **Try** to run it, or change the input to fit your own list.

- [Extract all blog post URLs from a website](https://apify.com/oldjard/sitemap-url-extractor/examples/extract-blog-post-urls)
- [New and updated pages from the last 7 days](https://apify.com/oldjard/sitemap-url-extractor/examples/new-pages-last-7-days)
- [All product URLs of a Shopify store](https://apify.com/oldjard/sitemap-url-extractor/examples/shopify-store-product-urls)
- [Documentation site URL list for RAG and LLMs](https://apify.com/oldjard/sitemap-url-extractor/examples/docs-site-urls-for-rag)
- [Latest news articles from news sitemaps](https://apify.com/oldjard/sitemap-url-extractor/examples/latest-news-articles-sitemap)
- [List all pages of a website (even without a sitemap)](https://apify.com/oldjard/sitemap-url-extractor/examples/website-pages-without-sitemap)

### More tools from oldjard

- [Tech Stack Detector](https://apify.com/oldjard/tech-stack-detector): what any list of websites is built with.
- [Shopify Products Scraper & Price Monitor](https://apify.com/oldjard/shopify-products-price-monitor): catalogs and price changes from any Shopify store.
- [Workday, Greenhouse, Lever & Ashby Jobs Scraper](https://apify.com/oldjard/ats-career-site-jobs): every open job from company career sites.
- [Bulk Website Screenshot & URL to PDF](https://apify.com/oldjard/screenshot-pdf): screenshots and PDFs of any list of pages.
- [AI Web Scraper (your own key)](https://apify.com/oldjard/ai-web-scraper): describe fields in English, get JSON.
- [Website Change Monitor](https://apify.com/oldjard/website-change-monitor): a before/after diff by webhook, Slack or Discord when a page changes.
- [Company Registry Lookup](https://apify.com/oldjard/company-registry-lookup): UK Companies House, Spain, France, Finland and Norway in one schema.
- [UK & EU Public Tenders](https://apify.com/oldjard/uk-eu-public-tenders): Find a Tender and TED notices in one table, with daily only-new alerts.

### Use it from an AI agent or the API

- **Minimal input:** `{"startUrls": ["crawlee.dev"], "maxUrlsPerSite": 100}`. Set `maxUrlsPerSite` to cap the work
  and the cost.
- **Cost:** $0.0004 per URL output (after filters and dedupe). Sitemaps, feeds and pages read are free. 100 URLs =
  $0.04.
- **Run time (our runs):** 3–8 s for up to 800 URLs from 5 sites; 78 s for 20,000 URLs.
- **Results:** the default dataset, one row per URL; the per-site report is the `OUTPUT` record in the key-value
  store.
- Works over the Apify MCP server (`search-actors`, then `call-actor`) and is eligible for agentic payments (x402).

### FAQ

**How do I find a website's sitemap?** You don't need to; it checks robots.txt, `/sitemap.xml`,
`/sitemap_index.xml`, `/wp-sitemap.xml` and `<link rel="sitemap">`.

**Sitemap to CSV?** Download the run's dataset as CSV (or Excel or JSON) from the Storage tab or the API.

**Why did a site return no URLs?** Open the OUTPUT record in Storage → Key-value store: each site has a `note`. Common reasons: the
section you chose has no pages in the sitemap (turn off *Only the start URL's section*), the site blocks automated
visitors, or robots.txt disallows it.

**Can I get only new or changed pages?** Yes: set **Modified since** to a date or a period like `7 days`. Only URLs
whose sitemap or feed date is on or after it are returned. Schedule the run to keep a RAG index fresh.

**Does it download page content?** No, it lists URLs, which is fast and cheap. Pass the list to a content crawler to
fetch the text.

# Actor input Schema

## `startUrls` (type: `array`):

Websites, one per line or comma-separated. example.com lists the whole site. A URL with a path (https://docs.example.com/guide) lists only that section; turn off limitToStartPath to change this. A sitemap or feed URL (https://example.com/sitemap_news.xml) is read directly. Duplicates and blank lines are removed; an entry that isn't a web address gets a free error row.

## `maxUrlsPerSite` (type: `integer`):

Stop a site after this many URLs (after filters and dedupe). 0 means no limit. Each URL output is one url-found charge, so this also caps the cost per site.

## `includePatterns` (type: `array`):

Keep only URLs matching at least one pattern. Globs with \* match the path and query, e.g. /blog/*, *.pdf, /docs/*/api/*. A glob starting with https:// matches the whole URL. Wrap a pattern in slashes for a regular expression on the whole URL, e.g. //v\[0-9]+//i. Empty = keep everything.

## `excludePatterns` (type: `array`):

Drop URLs matching any of these patterns (same syntax as above), e.g. /tag/*, /author/*, *?page=*.

## `modifiedSince` (type: `string`):

Only URLs whose sitemap or feed date (lastmod, publication date) is on or after this date, for refreshing a RAG index with just what changed. Accepts a date (2026-09-01) or a period (7 days, 2 weeks). URLs without any date are skipped when this is set.

## `includeFeeds` (type: `boolean`):

Read the RSS or Atom feeds the site links from its start page. Feeds often list the newest posts before the sitemap does, and give a title and publication date.

## `crawlMode` (type: `string`):

Only when there's no sitemap or feed (default): follow the site's own links a few levels deep. Always: also crawl sites that have a sitemap, to catch pages missing from it. Never: sitemaps and feeds only. The crawl stays on the site, honours robots.txt and fetches only HTML pages.

## `maxCrawlPages` (type: `integer`):

How many pages the link crawl may open per site. Every link on those pages is reported, so 50 pages usually find hundreds of URLs.

## `maxCrawlDepth` (type: `integer`):

How many clicks away from the start page the crawl goes. 1 = only links on the start page.

## `sameDomainOnly` (type: `boolean`):

Drop URLs on other domains that a sitemap or feed lists (subdomains like blog.example.com are kept).

## `limitToStartPath` (type: `boolean`):

For a start URL with a path (like https://example.com/docs), keep only URLs under that path.

## `maxSitemapsPerSite` (type: `integer`):

Safety cap on how many sitemap and feed files to read per site. Very large sites (news, marketplaces) can have thousands.

## `maxConcurrency` (type: `integer`):

How many sites to work on at the same time. Each site is visited politely: at most 3 requests at once, honouring any robots.txt Crawl-delay.

## `requestTimeoutSecs` (type: `integer`):

Give up on a request whose server hasn't started answering within this time.

## Actor input object example

```json
{
  "startUrls": [
    "docs.apify.com",
    "https://www.example.com/blog/",
    "https://example.com/sitemap.xml"
  ],
  "maxUrlsPerSite": 100,
  "includeFeeds": true,
  "crawlMode": "fallback",
  "maxCrawlPages": 50,
  "maxCrawlDepth": 2,
  "sameDomainOnly": true,
  "limitToStartPath": true,
  "maxSitemapsPerSite": 1000,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 30
}
```

# Actor output Schema

## `urls` (type: `string`):

Dataset, one row per discovered page: url, site, sourceType (sitemap/rss/atom/crawl), source, lastmod, changefreq, priority, title (feeds), allowedByRobots. Entries that are not web addresses get a free row with errorCode and error.

## `summary` (type: `string`):

Per-site report (JSON): sitemaps and feeds read, pages crawled, URLs found, filtered and deduplicated, robots.txt status, and why a site returned no URLs.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://docs.apify.com/academy",
        "crawlee.dev"
    ],
    "maxUrlsPerSite": 100
};

// Run the Actor and wait for it to finish
const run = await client.actor("oldjard/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        "https://docs.apify.com/academy",
        "crawlee.dev",
    ],
    "maxUrlsPerSite": 100,
}

# Run the Actor and wait for it to finish
run = client.actor("oldjard/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://docs.apify.com/academy",
    "crawlee.dev"
  ],
  "maxUrlsPerSite": 100
}' |
apify call oldjard/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,oldjard/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/tbYmWXo7KVGEOWM6S/builds/APHeJQYG1TsercjU4/openapi.json
