# Sitemap & robots.txt URL Discovery (no crawling) (`everyotherfriday/url-discovery`) Actor

Feed it site roots, get back URLs from their sitemaps plus each site's disallow rules and crawl-delay. Streams sitemaps, follows nested indexes and handles gzip without crawling page bodies. For seeding crawlers, SEO audits and migration checks.

- **URL**: https://apify.com/everyotherfriday/url-discovery.md
- **Developed by:** [Paul Vasquez](https://apify.com/everyotherfriday) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.30 / 1,000 url discovereds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap, robots.txt & URL Discovery

Turn a list of websites into a list of their URLs without crawling a single page. The actor reads each site's `robots.txt`, follows every sitemap it declares (including nested sitemap indexes and gzip files), falls back to the common sitemap locations when robots.txt is silent, and writes one dataset row per unique URL. Optionally it HEAD-checks every URL so you can see which ones are alive.

Typical uses: seeding a crawler with a complete URL list, SEO audits (which URLs a site advertises, when they changed), content inventories, migration checks, monitoring competitors' new pages, and finding a site's disallow rules and crawl-delay before you scrape it.

### What it does per site

1. Fetches `/robots.txt` and parses `Sitemap:` lines, user-agent groups, `Disallow`/`Allow` rules and `Crawl-delay`.
2. Queues every declared sitemap. If robots.txt declares none, it tries `/sitemap.xml`, `/sitemap_index.xml`, `/sitemap-index.xml`, `/wp-sitemap.xml` and `/sitemap.xml.gz`.
3. Streams each sitemap with an incremental XML parser (memory stays flat even for 50 MB files), detects gzip automatically, follows sitemap indexes up to 20 levels deep and 1,000 sitemap requests per site, and rejects HTML or DTD-bearing responses safely.
4. Deduplicates URLs, applies your optional regex filter, stops at `maxUrlsPerSite`, and optionally HEAD-checks each URL with 10 concurrent requests.
5. Writes the rows to the dataset and a `SUMMARY-<host>` record to the key-value store.

If you pass a sitemap URL directly (anything with a path, e.g. `https://example.com/news-sitemap.xml`) it is fetched as a sitemap for that origin without consulting robots.txt fallbacks.

### Input

| Field | Type | Default | Notes |
|---|---|---|---|
| `startUrls` | array of URLs | required | Site roots or explicit sitemap URLs, 1 to 100. Grouped by origin. |
| `includeRobots` | boolean | true | Parse robots.txt for sitemaps and rules. |
| `followSitemapIndexes` | boolean | true | Recurse into sitemap index files. |
| `maxUrlsPerSite` | integer | 5000 | Hard cap per origin, up to 1,000,000. |
| `includeLastmod` | boolean | true | Copy `lastmod` from the sitemap. |
| `includeStatusCheck` | boolean | false | HEAD every URL and record the HTTP status. |
| `filterPattern` | string | none | Python regular expression searched against each full URL. |
| `timeoutSecs` | integer | 30 | Per request, 1 to 300. |
| `proxyConfiguration` | object | none | Apify Proxy or custom proxy URLs. |

### Output

Dataset row per URL:

```json
{
  "site": "https://wordpress.org/",
  "url": "https://wordpress.org/news/2026/09/example/",
  "source": "sitemap-index",
  "sitemapUrl": "https://wordpress.org/news/sitemap-1.xml",
  "lastmod": "2026-09-24T10:12:00+00:00",
  "changefreq": null,
  "priority": null,
  "status": 200
}
```

`source` is one of `robots` (sitemap declared in robots.txt), `sitemap-index` (reached through an index), `fallback` (found at a common path) or `sitemap` (URL you supplied directly). `status` is present only when `includeStatusCheck` is on.

Key-value store record `SUMMARY-<host>` per site:

```json
{
  "site": "https://wordpress.org/",
  "robotsFound": true,
  "sitemapsFound": 19,
  "urlCount": 300,
  "disallowCount": 4,
  "crawlDelay": null,
  "warnings": ["maxUrlsPerSite reached; discovery stopped"],
  "robotsGroups": [{"userAgents": ["*"], "disallow": ["/wp-admin/"], "allow": [], "crawlDelay": null}],
  "elapsedSeconds": 7.9
}
```

Warnings list every sitemap that returned a non-200 status, was not XML, failed to parse, or timed out. A site with no sitemap at all produces a summary with `urlCount: 0` and the fallback 404 warnings, and costs nothing.

### Pricing

Pay per event: one `url-discovered` event ($0.0003) per unique URL row written. Robots-only lookups, failed sitemaps and sites without sitemaps are free. Discovering 100,000 URLs costs $30 plus nothing else; there is no per-run or per-site fee.

### Limits and behaviour

- Sitemaps larger than 256 MiB (after gzip expansion) are rejected with a warning.
- Sitemap index recursion stops at depth 20 or 1,000 sitemap fetches per site.
- Only `http` and `https` URLs are accepted; URLs with embedded credentials are rejected.
- Robots.txt bodies over 2 MiB or served as HTML are ignored with a warning.
- The actor never fetches page bodies, so it does not consume site bandwidth beyond the sitemap files themselves. It honours nothing in robots.txt except reading it; use the summary's disallow rules to configure your own crawler.

### Running locally

```powershell
python -m venv .venv
.venv\Scripts\python.exe -m pip install -r requirements.txt
$env:APIFY_LOCAL_STORAGE_DIR = "$PWD\storage\manual"
New-Item -ItemType Directory -Force "$env:APIFY_LOCAL_STORAGE_DIR\key_value_stores\default"
Copy-Item INPUT.json "$env:APIFY_LOCAL_STORAGE_DIR\key_value_stores\default\INPUT.json"
.venv\Scripts\python.exe -m src
.venv\Scripts\python.exe -m unittest discover -s tests -v
```

`validation/run_live.ps1` runs `INPUT.json` in fresh local storage and writes `validation/results.json`. See `VALIDATION.md` for the latest real-site results.

### Example output

One real dataset row from [storage/live-20260926-040335/datasets/default/000000001.json](storage/live-20260926-040335/datasets/default/000000001.json), trimmed by omitting fields without changing retained values:

```json
{
  "site": "https://wordpress.org/",
  "url": "https://wordpress.org/",
  "source": "robots",
  "sitemapUrl": "https://wordpress.org/news-sitemap.xml",
  "lastmod": null
}
```

This row came from the saved 900-row validation run. No HEAD check was requested, so status is absent. A null lastmod means the sitemap supplied no value for this row. The earlier Output examples illustrate the schema.

### Use cases

- SEO teams can inventory sitemap-advertised URLs before selecting pages for a technical audit.
- Website migration teams can compare exported URL lists with planned redirects and replacement pages.
- Content operations teams can compare dated exports to identify newly advertised articles for review.
- Data engineering teams can seed a downstream crawler and use the robots summary to configure its access rules.

**Pricing example:** 10,000 unique matching URL rows written x $0.0003 per `url-discovered` event = **$3.00 in event fees**, using `.actor/pay_per_event.json`. Local runs do not bill.

### Limitations

Discovery covers URLs advertised through the sitemaps reached within the configured bounds; it is not a complete inventory of every reachable page. A returned URL does not prove that its page is available or indexable. Inspect per-site warnings, especially when a cap stops discovery.

# Actor input Schema

## `startUrls` (type: `array`):

Site roots or explicit sitemap URLs, grouped by origin.

## `includeRobots` (type: `boolean`):

Read robots.txt sitemap declarations and user-agent rule groups.

## `followSitemapIndexes` (type: `boolean`):

Follow nested indexes up to depth 20 and 1000 sitemap requests per site.

## `includeLastmod` (type: `boolean`):

Include sitemap lastmod values; false emits null.

## `includeStatusCheck` (type: `boolean`):

HEAD each URL with concurrency 10; omit status when disabled.

## `maxUrlsPerSite` (type: `integer`):

Maximum unique matching URLs per origin.

## `filterPattern` (type: `string`):

Optional Python regular expression searched against each full URL.

## `timeoutSecs` (type: `integer`):

Timeout for each network operation in seconds, not the entire site.

## `proxyConfiguration` (type: `object`):

Optional Apify or custom proxy configuration; omitted means direct connections.

## Actor input object example

```json
{
  "startUrls": [
    "https://python.org/"
  ],
  "includeRobots": true,
  "followSitemapIndexes": true,
  "includeLastmod": true,
  "includeStatusCheck": false,
  "maxUrlsPerSite": 5000,
  "timeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `urls` (type: `string`):

All URL rows with source, sitemap, lastmod, changefreq, priority and optional HTTP status.

## `urlsCsv` (type: `string`):

Same rows as CSV.

## `summaries` (type: `string`):

Key-value store keys SUMMARY-<host> with robots, sitemap and URL counts, crawl-delay and warnings.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://python.org/"
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("everyotherfriday/url-discovery").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["https://python.org/"],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("everyotherfriday/url-discovery").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://python.org/"
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call everyotherfriday/url-discovery --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,everyotherfriday/url-discovery"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/MLo5dsFq77JSgTf4O/builds/GpdtiJksTpYzZQiP1/openapi.json
