# Sitemap URL Extractor — PDF/DOCX Tags + robots.txt (`ingenious_quip_bxq/sitemap-url-discovery`) Actor

Discover every URL in XML sitemaps via robots.txt and indexes; tag PDF, DOCX, XLSX and more for SEO audits and RAG ingest. Auto-finds nested indexes and .xml.gz; exports DOC\_TO\_MARKDOWN\_INPUT for the PDF/DOCX Actor.

- **URL**: https://apify.com/ingenious\_quip\_bxq/sitemap-url-discovery.md
- **Developed by:** [新世紀書僮](https://apify.com/ingenious_quip_bxq) (community)
- **Categories:** SEO tools, Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.20 / 1,000 url discovereds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap URL Extractor — Find PDF & Document Links

**Get every URL a website publishes in its XML sitemaps — with lastmod, source sitemap and a file-type tag, so PDFs, Word, Excel and PowerPoint files stand out.**
Enter a homepage and the Actor finds the sitemaps for you (robots.txt `Sitemap:` lines plus `/sitemap.xml` and `/sitemap_index.xml`), follows nested sitemap indexes, reads gzipped `.xml.gz` and plain-text sitemaps, removes duplicates and saves one row per URL. Use it for SEO audits, content inventories, change monitoring, or to collect a site's PDF and DOCX documents for RAG and LLM pipelines.

### What you get

- 🔎 **Automatic sitemap discovery** — robots.txt `Sitemap:` lines + common paths; or paste sitemap URLs directly
- 🌳 **Sitemap index recursion** — nested indexes up to a depth you choose, loops and repeats skipped
- 🗜️ **All common formats** — XML `urlset` and `sitemapindex`, `.xml.gz` (gzip file or gzip transfer encoding), plain-text sitemaps, RSS/Atom feeds
- 🏷️ **File-type tag per URL** — `pdf`, `docx`, `doc`, `xlsx`, `xls`, `pptx`, `ppt`, `csv`, `html`, `image`, `other` (from the extension; optional HTTP HEAD check confirms the real Content-Type)
- 📄 **Ready-made input for PDF/DOCX conversion** — all PDF and DOCX URLs are saved as `DOC_TO_MARKDOWN_INPUT`, formatted as input for the [PDF & DOCX to Markdown](https://apify.com/ingenious_quip_bxq/pdf-docx-to-markdown) Actor
- 🧯 **Broken sitemaps don't break the run** — every sitemap gets its own status line (`ok`, `not_found`, `error`, `skipped`) with the reason in `SITEMAP_REPORT`; one bad sitemap or site never fails the whole run
- 🐢 **Polite and robust** — low per-host concurrency, timeouts, retries with backoff for timeouts / 429 / 5xx (honours `Retry-After`), size caps, tolerant parsing of malformed XML
- 🧮 **Filters** — file types, lastmod date, include/exclude regex, same domain only, max URLs (total and per site)
- 💾 **Light** — 256 MB memory by default; HTTP only, no browser

### Measured results

Tested by us on 2026-09-29 with small limits (local runs plus 3 runs on the Apify platform):

| Test | Sitemap setup | Result |
|---|---|---|
| fec.gov (US government) | robots.txt lists 3 sitemaps incl. a PDF-only sitemap | 69 PDF URLs tagged `pdf`; HEAD check on all 69: HTTP 200 `application/pdf` (69/69) |
| gov.uk | sitemap index → 5.6 MB child sitemap | **20,000 URLs in 11 s on the Apify platform at 256 MB** (peak memory 107 MB) |
| nist.gov | Drupal sitemap index with 56 child sitemaps | index followed, URLs saved up to the limit |
| developer.mozilla.org | index of 10 `.xml.gz` sitemaps | gzip sitemaps read |
| yoast.com | Yoast SEO index with 21 child sitemaps; `/sitemap.xml` is a second index to the same files | children read once (duplicate sitemaps and URLs skipped) |
| Non-existent host in the same run | – | recorded as `error` (DNS failure, not retried) in `SITEMAP_REPORT`; the run still succeeded with the other sites' URLs |
| Local test server | raw gzip file, plain-text sitemap, malformed XML, missing (404) child, index loop, 2-level nesting | all read; malformed XML recovered with the tolerant scan; 404 child reported; loop not followed twice |

These are small tests on a handful of sites, not a benchmark: sites with unusual setups can still fail — when they do, the reason is in `SITEMAP_REPORT`.

### Use cases

- **Find all PDFs on a website** — set *Only these file types* to `pdf` (and `docx`) to list a site's documents: reports, forms, policies, manuals.
- **Documents for RAG / LLMs** — feed the PDF and DOCX list straight into a PDF-to-Markdown converter (see *Chaining* below).
- **SEO audits and content inventories** — every indexed URL with `lastmod`, `changefreq`, `priority` and the sitemap it came from.
- **Change monitoring** — schedule the Actor with *Changed on or after* to get only pages updated recently.
- **Crawl seeding** — use the URL list as start URLs for a scraper or crawler instead of crawling links.

### How to use

1. Add one or more websites (`https://www.example.gov`) or sitemap URLs in **Websites or sitemap URLs**.
2. Optional: set **Max URLs**, **Only these file types**, a **lastmod** date or URL patterns.
3. Click **Start**. URLs appear in the **Dataset**; the per-sitemap report and the document list are in the **Key-value store**.

#### Input example

```json
{
  "startUrls": [
    { "url": "https://www.fec.gov" },
    { "url": "https://www.example.com/sitemap_index.xml" }
  ],
  "maxUrls": 5000,
  "fileTypes": ["pdf", "docx"],
  "checkContentType": "off"
}
```

#### Output example (one dataset item per URL)

```json
{
  "url": "https://www.fec.gov/resources/cms-content/documents/policy-guidance/fecfrm1.pdf",
  "fileType": "pdf",
  "extension": "pdf",
  "lastmod": "2023-08-17T10:00:00+00:00",
  "changefreq": null,
  "priority": null,
  "sitemapUrl": "https://www.fec.gov/resources/cms-content/documents/sitemap_pdf.xml",
  "inputUrl": "https://www.fec.gov",
  "host": "www.fec.gov"
}
```

With **Confirm file type with HTTP HEAD** enabled, items also get `httpStatus`, `contentType`, `contentLength`, `finalUrl`, `contentTypeFileType` and `checkError`.

#### Key-value store records

| Key | Content |
|---|---|
| `OUTPUT` | Run summary: URLs saved, counts per file type, sitemaps read / failed / skipped, duplicates, per-input status |
| `SITEMAP_REPORT` | One entry per robots.txt and sitemap: `status` (`ok`, `not_found`, `error`, `skipped`), HTTP status, type, gzip, URLs found/new, child sitemaps, attempts, error message |
| `DOC_TO_MARKDOWN_INPUT` | `{"urls": [{"url": "...pdf"}, ...]}` — every PDF and DOCX URL saved in this run (URLs that failed the HEAD check are left out) |

### Chaining: sitemap → PDF/DOCX → Markdown

`DOC_TO_MARKDOWN_INPUT` uses the exact input format of the [PDF & DOCX to Markdown](https://apify.com/ingenious_quip_bxq/pdf-docx-to-markdown) Actor (`urls` list). Example with the Apify Python client:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")

## 1) list a site's PDFs and DOCX files
run = client.actor("ingenious_quip_bxq/sitemap-url-discovery").call(run_input={
    "startUrls": [{"url": "https://www.fec.gov"}],
    "fileTypes": ["pdf", "docx"],
    "maxUrls": 20,
})
doc_input = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("DOC_TO_MARKDOWN_INPUT")["value"]

## 2) convert them to Markdown / RAG chunks
run2 = client.actor("ingenious_quip_bxq/pdf-docx-to-markdown").call(run_input={**doc_input, "outputFormat": "chunks"})
for item in client.dataset(run2["defaultDatasetId"]).iterate_items():
    print(item.get("fileName"), item.get("pageStart"), item.get("text", "")[:80])
```

No code? In Apify Console, add an **Integration → Actor** on this Actor's run and pass the key-value store record as input, or copy the `DOC_TO_MARKDOWN_INPUT` record into the other Actor's JSON input. Mind the second Actor's pricing (per document and per page) before sending hundreds of files.

### Pricing

Pay per event:

| Event | Price |
|---|---|
| URL saved to the dataset (primary) | **$0.0002** per URL (= $0.20 per 1,000 URLs) |
| Optional HEAD content-type check (off by default) | $0.0003 per URL checked |
| Actor start | Apify's default synthetic start event |

You pay only for unique URLs saved after filters. Duplicates, filtered URLs, sitemaps that fail to load and the per-sitemap report are free. Set **Max URLs** (or a spending limit on the run) to cap cost — the Actor stops at the limit.

### Known limits

- Only URLs listed in sitemaps (or feeds) are found — this Actor does not crawl links. Sites without a sitemap return no URLs (reported as `no_sitemap_found`).
- File type comes from the URL extension unless the HEAD check is on; download handlers such as `/download?id=123` are tagged `html` without it.
- Image and video URLs inside `<image:image>` / `<video:video>` sitemap extensions are not extracted as separate rows.
- Sites that block data-center IPs may need a proxy (Proxy setting).
- Sitemaps are read up to 100 MB compressed / 200 MB uncompressed each.

### FAQ

**Does it respect robots.txt?** It reads robots.txt to find sitemaps. Sitemaps exist to be read by bots, so it fetches them; it doesn't fetch pages unless you enable the HEAD check.

**Why do I see `not_found` entries?** `/sitemap.xml` and `/sitemap_index.xml` are probed on every site; when a site doesn't have them, that is recorded as `not_found`, not as an error.

**Is the output compatible with other tools?** Yes — download the dataset as CSV, JSON, Excel or XML, or read it via the API.

### License & source code

This Actor is open source under the **GNU Affero General Public License v3.0 (AGPL-3.0)**. The full source code is public: https://github.com/xbox002000/sitemap-url-discovery

Built with the Apify Python SDK, httpx and defusedxml.

See the [changelog](CHANGELOG.md) for version history.

# Changelog

This Actor's version history is a separate document: https://apify.com/ingenious\_quip\_bxq/sitemap-url-discovery/changelog.md

# Actor input Schema

## `startUrls` (type: `array`):

A homepage like `https://www.example.com` (sitemaps are discovered from robots.txt and /sitemap.xml, /sitemap\_index.xml) or a direct sitemap URL (`.xml`, `.xml.gz`, `.txt`, sitemap index).

## `maxUrls` (type: `integer`):

Stop after saving this many unique URLs (0 = no limit). Also caps your cost: you pay only for URLs saved.

## `maxUrlsPerInput` (type: `integer`):

Optional cap per input website/sitemap (0 = no cap), so one large site cannot use up the whole Max URLs budget.

## `fileTypes` (type: `array`):

Keep only URLs with these file-type tags (by extension). Leave empty to keep everything. Example: only `pdf` and `docx` to list a site's documents.

## `lastmodFrom` (type: `string`):

Keep only URLs whose sitemap `lastmod` is on or after this date (YYYY-MM-DD). Useful for monitoring new or updated pages.

## `includeUrlsWithoutLastmod` (type: `boolean`):

When a date filter is set, keep URLs that have no `lastmod` in the sitemap.

## `includeUrlPatterns` (type: `array`):

Keep only URLs matching at least one of these regular expressions, e.g. `/blog/`.

## `excludeUrlPatterns` (type: `array`):

Drop URLs matching any of these regular expressions, e.g. `/tag/`.

## `sameDomainOnly` (type: `boolean`):

Drop URLs whose host differs from the input website (www. ignored). Off by default because many sites list files on a CDN host.

## `checkContentType` (type: `string`):

Send a HEAD request to confirm HTTP status and Content-Type. `documents` checks only non-HTML, non-image URLs (e.g. files and URLs without a clear extension); `all` checks every URL. Costs an extra event per URL checked and makes runs slower.

## `useRobotsTxt` (type: `boolean`):

Discover sitemaps from `Sitemap:` lines in robots.txt.

## `tryCommonPaths` (type: `boolean`):

Also try /sitemap.xml and /sitemap\_index.xml. A missing path is reported as `not_found`, not as an error.

## `maxDepth` (type: `integer`):

How many levels of nested sitemap indexes to follow.

## `maxSitemaps` (type: `integer`):

Maximum number of sitemap files to read per input website (0 = no limit).

## `maxConcurrency` (type: `integer`):

Parallel requests overall.

## `maxConcurrencyPerHost` (type: `integer`):

Parallel requests to any single host (kept low to be polite).

## `requestTimeoutSecs` (type: `integer`):

Timeout for each HTTP request.

## `maxRetries` (type: `integer`):

Retries for timeouts, network errors, HTTP 429 and 5xx (exponential backoff, honours Retry-After up to 30 s).

## `userAgent` (type: `string`):

Optional custom User-Agent header.

## `proxyConfiguration` (type: `object`):

Optional. Use a proxy if a site blocks data-center IPs. Not needed for most sites.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.apify.com"
    }
  ],
  "maxUrls": 100,
  "maxUrlsPerInput": 0,
  "includeUrlsWithoutLastmod": true,
  "sameDomainOnly": false,
  "checkContentType": "off",
  "useRobotsTxt": true,
  "tryCommonPaths": true,
  "maxDepth": 5,
  "maxSitemaps": 500,
  "maxConcurrency": 4,
  "maxConcurrencyPerHost": 2,
  "requestTimeoutSecs": 30,
  "maxRetries": 3,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `urls` (type: `string`):

Default dataset items, one per unique URL: url, fileType (pdf, docx, xlsx, pptx, csv, html, image, other…), extension, lastmod, changefreq, priority, sitemapUrl (source sitemap), inputUrl, host. With the content-type check enabled also httpStatus, contentType, contentTypeFileType, contentLength, finalUrl, checkError. Views: overview (URLs), contentType (Content-type check).

## `sitemapReport` (type: `string`):

Key-value store record SITEMAP\_REPORT (JSON array): one entry per robots.txt / sitemap fetched or skipped with sitemapUrl, inputUrl, discoveredVia, parentSitemap, depth, status (ok / error / skipped), httpStatus, error and URL counts. Broken sitemaps are reported here instead of failing the run.

## `docToMarkdownInput` (type: `string`):

Key-value store record DOC\_TO\_MARKDOWN\_INPUT (JSON): {"urls": \[{"url": …}]} listing every PDF/DOCX URL found, ready to pass as input to the PDF & DOCX to Markdown Actor.

## `summary` (type: `string`):

Key-value store record OUTPUT (JSON): urlsOutput, fileTypes counts, documentUrlsForDocToMarkdown, sitemapsRead, sitemapErrors, sitemapsSkipped, duplicatesSkipped, filteredOut, maxUrlsReached, per-input status, durationSecs, peakMemoryMb.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.apify.com"
        }
    ],
    "maxUrls": 100,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("ingenious_quip_bxq/sitemap-url-discovery").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://www.apify.com" }],
    "maxUrls": 100,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("ingenious_quip_bxq/sitemap-url-discovery").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.apify.com"
    }
  ],
  "maxUrls": 100,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call ingenious_quip_bxq/sitemap-url-discovery --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ingenious_quip_bxq/sitemap-url-discovery"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/m37uyunL8QifiWEgD/builds/aav4teVcM8KNiuiQ4/openapi.json
