# Sitemap URL Extractor — Robots.txt & Sitemap.xml Parser (`tarnlight/sitemap-url-extractor`) Actor

Extracts the URLs a website lists in its XML sitemaps (sitemap indexes, gzip and text sitemaps, robots.txt discovery), up to your Max URLs cap, with filters and a clear per-site report.

- **URL**: https://apify.com/tarnlight/sitemap-url-extractor.md
- **Developed by:** [Tarnlight](https://apify.com/tarnlight) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor

Reliably pulls every URL a site's sitemaps list (up to your `maxUrls` cap)
out of its XML sitemaps -- and reports, per site, exactly what it found and
what it couldn't reach -- instead of quietly returning a partial list.
Honours `robots.txt`. Built by **Tarnlight**.

### What it does

Give it a homepage/site URL **or** a direct sitemap URL -- the two behave
differently, on purpose (see "How discovery works" below):

- **A homepage/site URL** (e.g. `https://example.com`) triggers full
  discovery: `robots.txt` is read for `Sitemap:` lines, and well-known
  common paths are probed too.
- **A direct sitemap URL** (e.g. `https://example.com/sitemap.xml`, or
  anything else that looks like a sitemap file/path) is fetched **by
  itself** -- only that file and, if it's a sitemap index, its child
  sitemaps. There is no site-wide discovery and nothing else is guessed,
  so you're only ever charged for the file you actually asked for.

For every site, it will:

1. Look up `/robots.txt` for `Disallow:` rules -- every request this actor
   makes (a common-path guess, a declared sitemap, a child of a sitemap
   index, or a start URL you gave it directly) is checked against those
   rules first (RFC 9309; see "How discovery works" below). This always
   happens, for both kinds of start URL.
2. For a homepage/site URL only: also read `robots.txt`'s `Sitemap:` lines,
   and probe common sitemap paths (`/sitemap.xml`, `/sitemap_index.xml`,
   `/sitemap-index.xml`, `/wp-sitemap.xml`, `/sitemap.xml.gz`,
   `/sitemap.txt`) -- both happen regardless of what the other finds,
   skipping any robots.txt disallows.
3. Recursively follow sitemap indexes (gzip and plain, cycle-safe, depth
   capped) down to the individual page sitemaps.
4. Parse XML sitemaps (namespaced or not) and gzip-compressed sitemaps.
   Parse a plain-text sitemap (one absolute URL per line) only when the
   response actually looks like one -- see "Plain-text sitemaps" below;
   anything else (an HTML page, `humans.txt`, arbitrary text) is rejected
   as `"not-a-sitemap"` and never billed.
5. Push one dataset row per URL, deduplicated within the run, with
   `lastmod`, `changefreq`, `priority`, hreflang alternates and image
   counts when present.
6. Record a `SUMMARY` key-value-store entry with per-site sitemap counts,
   failures (with reasons), how many URLs matched your filters vs. were
   actually written vs. skipped once a cap was hit, every discovery attempt
   made (found it, 404, blocked, malformed, ...), and a plain-English
   `note` explaining anything worth flagging -- especially when a site
   yields 0 URLs.

If a page returns an HTML "error" page with an HTTP 200 status where a
sitemap was expected, that's detected and reported as a failure for that
sitemap instead of being mis-parsed as XML. A single unreachable or broken
sitemap never fails the whole run -- it's logged in `SUMMARY` and the run
carries on with everything else.

#### Plain-text sitemaps

A non-XML 200 response is only treated as a plain-text sitemap when it was
actually served as one: the URL ends in `.txt`/`.txt.gz`, or the response's
`Content-Type` is `text/plain`. Even then, only lines that parse as a bare
absolute `http(s)://` URL (no embedded whitespace) **on the sitemap file's
own site** are kept -- per the sitemaps.org rule that a sitemap may only
list URLs on its own host; cross-host submission requires a search-console
verification this actor has no way to check, so it's never assumed. If no
line qualifies, the whole response is rejected as `"not-a-sitemap"` rather
than billing you for every line of, say, a site's `humans.txt`.

### Who it's for

- **SEO audits** -- get the full, current list of URLs a site is
  telling search engines about.
- **Content inventories** -- see every page a site publishes without
  crawling the whole thing link by link.
- **Migration checks** -- diff sitemap URL lists before/after a
  replatform.
- **Feeding crawlers or RAG pipelines** -- a clean, deduplicated seed
  list of URLs to fetch next, instead of hand-rolling sitemap parsing.

### Input

| Field | Type | Default | Notes |
|---|---|---|---|
| `startUrls` | array | *(required)* | Homepages/site URLs (full discovery) or direct sitemap URLs (fetched by themselves only -- see above). |
| `maxUrls` | integer | `5000` | Stop after this many URLs are **written** (0 = unlimited -- can run up an unbounded bill on a large site). Each URL costs $0.001, so the default caps a run at about $5. |
| `includePatterns` | array of regex | `[]` | Keep only URLs matching at least one. |
| `excludePatterns` | array of regex | `[]` | Drop URLs matching any. An invalid regex here (or in `includePatterns`) ends the run cleanly with an "invalid input" status message instead of crashing mid-run. |
| `modifiedSince` | date | *(none)* | Keep URLs with `lastmod` on/after this date. |
| `keepUrlsWithoutLastmod` | boolean | `false` | Only matters when `modifiedSince` is set. By default, URLs with no usable `lastmod` (missing, or a placeholder such as WordPress/Jetpack's `"0000-00-00"`) are dropped and **not charged**, so a date-filtered run only returns (and bills) URLs known to match your date. Set to `true` to keep them as well. `SUMMARY` reports `urlsWithUnparseableLastmod` per site either way. |
| `discoverFromRobots` | boolean | `true` | Homepage/site URLs only: use `robots.txt`'s `Sitemap:` lines for discovery. `robots.txt` is always fetched and its `Disallow:` rules always honoured regardless of this setting -- it only controls whether its declared sitemaps are added to the discovery list. Has no effect on a direct sitemap URL. |
| `tryCommonPaths` | boolean | `true` | Homepage/site URLs only: probe well-known sitemap paths, regardless of what `robots.txt` declares. Has no effect on a direct sitemap URL. |
| `includeAlternates` | boolean | `false` | Include hreflang `alternates` in output rows. |
| `includeImages` | boolean | `false` | Include `imageUrls` in output rows (a count is always included). |
| `maxConcurrency` | integer | `5` | Max simultaneous requests, across all sites. A single site's own crawl is further limited to 2 concurrent requests per host, so this mostly matters when you pass several `startUrls` at once. |
| `requestTimeoutSecs` | integer | `30` | Per-request timeout. |

Example input:

```json
{
  "startUrls": [
    { "url": "https://example.com" },
    { "url": "https://example.com/custom-sitemap.xml" }
  ],
  "maxUrls": 5000,
  "excludePatterns": ["/tag/", "/page/[0-9]+/"],
  "modifiedSince": "2025-01-01",
  "includeImages": true
}
```

### Output

One dataset row per URL. Real rows from a test run (`docs.apify.com` and
`gov.uk`), showing typical variation in which fields a site fills in:

```json
{
  "url": "https://docs.apify.com/api",
  "lastmod": null,
  "changefreq": "weekly",
  "priority": 0.5,
  "sitemapUrl": "https://docs.apify.com/sitemap_base.xml",
  "site": "https://docs.apify.com",
  "imageCount": 0
}
```

```json
{
  "url": "https://www.gov.uk/government/news/whole-system-approach-to-tackling-violent-crime-is-working",
  "lastmod": "2026-07-23T19:37:26+01:00",
  "changefreq": null,
  "priority": 0.29,
  "sitemapUrl": "https://www.gov.uk/sitemaps/sitemap_1.xml",
  "site": "https://www.gov.uk",
  "imageCount": 0
}
```

`alternates` (list of `{hreflang, href}`) and `imageUrls` are only added
when `includeAlternates` / `includeImages` are enabled.

The `SUMMARY` key-value-store record looks like:

```json
{
  "sites": [
    {
      "site": "https://example.com",
      "sitemapsFound": 3,
      "sitemapUrls": ["https://example.com/sitemap_index.xml", "..."],
      "sitemapsFailed": [],
      "urlsMatched": 812,
      "urlsWritten": 812,
      "urlsSkippedAfterCap": 0,
      "urlsFilteredOut": 4,
      "urlsWithUnparseableLastmod": 0,
      "stoppedEarly": false,
      "discoveryAttempts": [
        { "url": "https://example.com/robots.txt", "source": "robots.txt", "httpStatus": 200,
          "outcome": "not-a-sitemap", "detail": "found 1 Sitemap: line(s)" },
        { "url": "https://example.com/sitemap_index.xml", "source": "robots.txt", "httpStatus": 200,
          "outcome": "sitemap", "detail": "sitemap index with 2 child sitemap(s)" },
        { "url": "https://example.com/wp-sitemap.xml", "source": "common-path", "httpStatus": 404,
          "outcome": "not-found", "detail": "HTTP 404" }
      ],
      "note": null
    }
  ],
  "totalUrlsWritten": 812,
  "stopReason": null
}
```

`urlsMatched` is how many of that site's sitemap entries passed your
`includePatterns`/`excludePatterns`/`modifiedSince` filters. `urlsWritten`
is how many of those were **actually written to the dataset** (and
charged) -- always the honest number; it can be lower than `urlsMatched`
only when `urlsSkippedAfterCap` is non-zero, meaning `maxUrls` or the run's
own charge limit was reached before every matched URL could be written.
`totalUrlsWritten` (top level) is the sum of every site's `urlsWritten`,
and always equals the number of `url-extracted` charged events for the run.

Every URL the actor tried in order to find a sitemap for a site --
whether it worked or not -- is listed in that site's `discoveryAttempts`:
`source` says how the URL was arrived at (`"start-url"`, `"robots.txt"` or
`"common-path"`), `outcome` says what happened (`"sitemap"` / `"not-found"`
/ `"blocked"` / `"error"` / `"not-a-sitemap"` / `"disallowed-by-robots"` --
robots.txt forbade fetching this URL, so it wasn't fetched / `"robots-
unreachable"` -- robots.txt itself returned a 5xx or a network error after
retries, so the whole site was treated as disallowed), and `detail` is a
short human-readable reason. `note` is `null` when there is nothing to
explain, and a plain-English sentence otherwise, e.g.:

```json
{ "site": "https://www.python.org", "sitemapsFound": 0, "urlsMatched": 0, "note":
  "No sitemap published: robots.txt has no Sitemap: line and none of 6 common paths exist." }
```

```json
{ "site": "https://blocked-example.com", "sitemapsFound": 0, "urlsMatched": 0, "note":
  "Access blocked (HTTP 403) when fetching robots.txt and/or sitemap paths -- the site may block automated requests." }
```

When any site ends with 0 matched URLs, the actor also logs a `WARNING`
with that site's `note`, and the final run status message adds a short,
correctly-categorized hint -- a blocked/unreachable site is never reported
the same way as one with no sitemap (e.g. "1 site was blocked or
unreachable; 1 site had no sitemap -- see SUMMARY") -- so it's never silent
about it, and a bot-blocking site is never confused with a genuinely
sitemap-less one.

### How discovery works

For each start URL's origin, `robots.txt` is fetched first, always --
before anything else is requested on that host. It's parsed for
`TarnlightSitemapBot`'s rules, falling back to `*` when there's no
bot-specific group (per RFC 9309's group-selection rule), and the result
gates every later request to that origin:

- **robots.txt itself returns 4xx (e.g. 404):** no restrictions apply
  (RFC 9309 S2.3.1.3) -- a missing robots.txt is normal, and discovery
  proceeds exactly as if it didn't exist.
- **robots.txt returns 5xx, or is unreachable over the network, after
  retries:** RFC 9309 S2.3.1.4 requires assuming complete disallow -- this
  actor treats the *entire site* as off-limits and makes no further
  request to it at all (not even a common-path guess), and says so
  clearly in that site's `note`.
- **robots.txt fetches fine (2xx):** its `Sitemap:` lines are read, and
  every URL this actor might fetch on that origin -- a `Sitemap:`-declared
  sitemap, a common-path guess, a child of a sitemap index, or the start
  URL you gave it directly (RFC 9309 applies to all of these alike) -- is
  checked against its `Disallow:` rules before being requested. A
  disallowed URL is never fetched; it's recorded in `discoveryAttempts`
  with outcome `"disallowed-by-robots"` and simply skipped.

For each start URL, if the URL itself looks like a sitemap file (ends in
`.xml`, `.xml.gz`, `.txt`, or has "sitemap" in the path), it's fetched
directly (once robots.txt allows it) -- **and only that URL and its
sitemap-index children are ever fetched for that start URL.** `robots.txt`
Sitemap: discovery and common-path probing do not run for it, regardless
of `discoverFromRobots`/`tryCommonPaths`, so a direct sitemap URL never
silently expands into (and bills for) the rest of the site. Otherwise the
actor treats the start URL as a site root: it reads `robots.txt` for
`Sitemap:` lines **and** probes the common paths above -- both happen
unconditionally (subject to `discoverFromRobots`/`tryCommonPaths`), not
only when the other turns up nothing, since some sites keep a stale
legacy sitemap at a common path alongside a different one declared in
`robots.txt`. Every sitemap found is fetched with retries
(exponential backoff on 429 and 5xx responses, honouring `Retry-After`),
decompressed if gzipped, and parsed. A sitemap index's children are
followed recursively (each one checked against robots.txt first; depth
capped at 5, and already-visited URLs are skipped, so a site that
accidentally -- or deliberately -- references a sitemap that references
itself can't cause a loop).

Every URL this process tries -- robots.txt itself, each `Sitemap:` line it
names, and each common-path guess -- is recorded as a `discoveryAttempts`
entry in `SUMMARY`, whether it succeeded or not. That's what makes a 0-URL
result explainable instead of silent: 401/403, and 429 that's still
failing after retries, are recorded as `"blocked"` (a bot-blocking site is
a common real-world cause of an empty result, not just a missing
sitemap); a 404 on a mere guess is `"not-found"` and never treated as a
failure; a 200 response that isn't valid sitemap XML/text (an HTML error
page, malformed XML, an empty body) is `"not-a-sitemap"`; a URL robots.txt
disallows is `"disallowed-by-robots"`; robots.txt itself being unreachable
is `"robots-unreachable"`. A per-site `note` in plain English summarizes
whichever of these applies.

### Limits

- This actor only reports what a site's sitemaps actually publish. It
  does not crawl pages or discover URLs that aren't listed in any
  sitemap.
- A site with no `robots.txt` `Sitemap:` line, no sitemap at a common
  path, and no sitemap URL given directly will correctly report zero
  URLs for that site -- that's expected, not a bug (see the FAQ).
- Per-file size is capped at 100 MB decompressed as a safety limit against
  runaway or hostile files -- for both gzip and plain (uncompressed) files.
- `robots.txt` is checked for the *start URL's own origin*. A sitemap
  index that points at a different host (e.g. a `www.` vs. bare-domain
  split, or a separate sitemap subdomain) has its children checked against
  that same start-origin policy, not a fresh fetch of the other host's own
  `robots.txt`; a redirect target is likewise not re-checked. This covers
  the overwhelmingly common case (a site's own sitemap on its own host)
  but isn't a substitute for reviewing a multi-host site's own policies.
- `maxConcurrency` bounds total in-flight requests across all `startUrls`,
  but a single site's own crawl is additionally capped at 2 concurrent
  requests per host -- so raising it mostly helps when you pass many start
  URLs at once, not a single large site.

### FAQ

**A site returned 0 URLs -- is that broken?**
Not necessarily -- check that site's `note` and `discoveryAttempts` in
`SUMMARY` (also logged as a `WARNING`, and hinted at in the final run
status). Two common, non-broken causes:

- **No sitemap published.** Nothing links to a sitemap from `robots.txt`,
  none of the common paths exist, and you didn't pass a direct sitemap
  URL -- there is genuinely nothing for the actor to find. Pass the
  sitemap URL directly if you know it.
- **The site blocked the request.** Some sites return HTTP 401/403 (or
  keep returning 429 after retries) to automated requests, including to
  `robots.txt` itself. `SUMMARY` reports this as `"outcome": "blocked"`
  with a `note` like *"Access blocked (HTTP 403) ..."* -- that's a site
  policy, not a bug in the actor, and it's reported distinctly from "no
  sitemap" so the two aren't confused.

**Does it crawl the site itself?**
No. It only follows sitemap files (and sitemap indexes) -- it never
requests ordinary pages on purpose. If a sitemap path redirects to an
ordinary HTML page, or a site serves HTML at a URL a sitemap named, that's
detected and discarded rather than parsed as XML.

**What happens with a huge site?**
Use `maxUrls` to cap the run, and `includePatterns`/`excludePatterns` to
narrow scope. Sitemap index recursion and per-file size are both bounded
automatically.

**Why do some rows have empty `lastmod`/`changefreq`/`priority`?**
Those fields are optional in the sitemap spec; the actor reports exactly
what the site published and leaves the rest `null` rather than guessing.

### Pricing

Pay per event, in USD: $0.005 per run start (scales with the run's memory,
in GB) + $0.001 per URL written to the dataset (e.g. 1,000 URLs = $1.005).
The default `maxUrls` of 5,000 caps a run at about $5.005; raise it (or set
it to `0` for no cap) only once you know roughly how many URLs you expect,
since a large uncapped site can run up a large bill. Any tax Apify applies
at checkout is extra; the Pricing tab is authoritative.

### Acceptable use and data protection

Use responsibly: only run this on sites you're entitled to access, and
follow each site's terms. The actor identifies itself as
`TarnlightSitemapBot` and honours `robots.txt` (including treating a
robots.txt server error as a full disallow -- see "How discovery works"),
so site owners can block it there.

Output can include personal data where a site's URLs contain it (e.g.
author pages). You decide how output is used and need a lawful basis.
Tarnlight does not access your runs unless you share run data with
developers.

# Actor input Schema

## `startUrls` (type: `array`):

Site homepages (e.g. https://example.com) or direct sitemap URLs (e.g. https://example.com/sitemap.xml). A homepage/site URL triggers robots.txt Sitemap: discovery and common-path probing. A direct sitemap URL is fetched by itself, and ONLY that file (plus its sitemap-index children, if any) -- it never expands into a whole-site discovery, so you're never charged for URLs you didn't ask for.

## `maxUrls` (type: `integer`):

Stop once this many URLs have been written to the dataset across all start URLs (0 = unlimited -- can run up an unbounded bill on a large site; set a cap unless you really want everything). Each URL costs $0.001, so the default of 5,000 caps a run at about $5 (plus the $0.005 start fee).

## `includePatterns` (type: `array`):

Only keep extracted URLs matching at least one of these regular expressions. Leave empty to keep all.

## `excludePatterns` (type: `array`):

Drop extracted URLs matching any of these regular expressions.

## `modifiedSince` (type: `string`):

ISO date (e.g. 2025-01-01). Only keep URLs whose <lastmod> is on or after this date. See 'Keep URLs without lastmod' for URLs missing a lastmod value.

## `keepUrlsWithoutLastmod` (type: `boolean`):

Only matters when 'Modified since' is set. Off (default): URLs with no usable <lastmod> (missing, or a placeholder like 0000-00-00) are dropped and not charged, so you only pay for URLs known to match your date. On: keep them too (their date is unknown).

## `discoverFromRobots` (type: `boolean`):

For a homepage/site-root start URL: use robots.txt's 'Sitemap:' lines for discovery (in addition to any common-path probing). Has no effect on a direct sitemap URL, which is always fetched by itself. robots.txt is always fetched and its Disallow rules are always honoured regardless of this setting; it only controls whether declared sitemaps are added to the discovery list.

## `tryCommonPaths` (type: `boolean`):

For a homepage/site-root start URL: also probe well-known paths (/sitemap.xml, /sitemap\_index.xml, /sitemap-index.xml, /wp-sitemap.xml, /sitemap.xml.gz, /sitemap.txt) at its origin, regardless of what robots.txt declares (some sites keep a stale legacy sitemap at a common path alongside a different declared one). Has no effect on a direct sitemap URL, which is always fetched by itself.

## `includeAlternates` (type: `boolean`):

Include \<xhtml:link rel="alternate" hreflang=...> entries for each URL in the output.

## `includeImages` (type: `boolean`):

Include image count (and image URLs) from <image:image> entries in the output.

## `maxConcurrency` (type: `integer`):

Maximum number of sitemap/robots.txt requests in flight at once, across all sites.

## `requestTimeoutSecs` (type: `integer`):

Timeout for each HTTP request to a robots.txt or sitemap file.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://example.com"
    }
  ],
  "maxUrls": 5000,
  "includePatterns": [],
  "excludePatterns": [],
  "keepUrlsWithoutLastmod": false,
  "discoverFromRobots": true,
  "tryCommonPaths": true,
  "includeAlternates": false,
  "includeImages": false,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 30
}
```

# Actor output Schema

## `urls` (type: `string`):

One item per URL found in the site's sitemaps, with lastmod, changefreq, priority and source sitemap.

## `summary` (type: `string`):

SUMMARY record: sitemaps found and failed, every discovery attempt with HTTP status, and a plain-English note per site.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://example.com"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tarnlight/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://example.com" }] }

# Run the Actor and wait for it to finish
run = client.actor("tarnlight/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://example.com"
    }
  ]
}' |
apify call tarnlight/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tarnlight/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KjfWsj0HGbGRUEepD/builds/tHHIerGSCXxmR7OMn/openapi.json
