Sitemap & robots.txt URL Discovery (no crawling) avatar

Sitemap & robots.txt URL Discovery (no crawling)

Pricing

$0.30 / 1,000 url discovereds

Go to Apify Store
Sitemap & robots.txt URL Discovery (no crawling)

Sitemap & robots.txt URL Discovery (no crawling)

Feed it site roots, get back URLs from their sitemaps plus each site's disallow rules and crawl-delay. Streams sitemaps, follows nested indexes and handles gzip without crawling page bodies. For seeding crawlers, SEO audits and migration checks.

Pricing

$0.30 / 1,000 url discovereds

Rating

0.0

(0)

Developer

Paul Vasquez

Paul Vasquez

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Sitemap, robots.txt & URL Discovery

Turn a list of websites into a list of their URLs without crawling a single page. The actor reads each site's robots.txt, follows every sitemap it declares (including nested sitemap indexes and gzip files), falls back to the common sitemap locations when robots.txt is silent, and writes one dataset row per unique URL. Optionally it HEAD-checks every URL so you can see which ones are alive.

Typical uses: seeding a crawler with a complete URL list, SEO audits (which URLs a site advertises, when they changed), content inventories, migration checks, monitoring competitors' new pages, and finding a site's disallow rules and crawl-delay before you scrape it.

What it does per site

  1. Fetches /robots.txt and parses Sitemap: lines, user-agent groups, Disallow/Allow rules and Crawl-delay.
  2. Queues every declared sitemap. If robots.txt declares none, it tries /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /wp-sitemap.xml and /sitemap.xml.gz.
  3. Streams each sitemap with an incremental XML parser (memory stays flat even for 50 MB files), detects gzip automatically, follows sitemap indexes up to 20 levels deep and 1,000 sitemap requests per site, and rejects HTML or DTD-bearing responses safely.
  4. Deduplicates URLs, applies your optional regex filter, stops at maxUrlsPerSite, and optionally HEAD-checks each URL with 10 concurrent requests.
  5. Writes the rows to the dataset and a SUMMARY-<host> record to the key-value store.

If you pass a sitemap URL directly (anything with a path, e.g. https://example.com/news-sitemap.xml) it is fetched as a sitemap for that origin without consulting robots.txt fallbacks.

Input

FieldTypeDefaultNotes
startUrlsarray of URLsrequiredSite roots or explicit sitemap URLs, 1 to 100. Grouped by origin.
includeRobotsbooleantrueParse robots.txt for sitemaps and rules.
followSitemapIndexesbooleantrueRecurse into sitemap index files.
maxUrlsPerSiteinteger5000Hard cap per origin, up to 1,000,000.
includeLastmodbooleantrueCopy lastmod from the sitemap.
includeStatusCheckbooleanfalseHEAD every URL and record the HTTP status.
filterPatternstringnonePython regular expression searched against each full URL.
timeoutSecsinteger30Per request, 1 to 300.
proxyConfigurationobjectnoneApify Proxy or custom proxy URLs.

Output

Dataset row per URL:

{
"site": "https://wordpress.org/",
"url": "https://wordpress.org/news/2026/09/example/",
"source": "sitemap-index",
"sitemapUrl": "https://wordpress.org/news/sitemap-1.xml",
"lastmod": "2026-09-24T10:12:00+00:00",
"changefreq": null,
"priority": null,
"status": 200
}

source is one of robots (sitemap declared in robots.txt), sitemap-index (reached through an index), fallback (found at a common path) or sitemap (URL you supplied directly). status is present only when includeStatusCheck is on.

Key-value store record SUMMARY-<host> per site:

{
"site": "https://wordpress.org/",
"robotsFound": true,
"sitemapsFound": 19,
"urlCount": 300,
"disallowCount": 4,
"crawlDelay": null,
"warnings": ["maxUrlsPerSite reached; discovery stopped"],
"robotsGroups": [{"userAgents": ["*"], "disallow": ["/wp-admin/"], "allow": [], "crawlDelay": null}],
"elapsedSeconds": 7.9
}

Warnings list every sitemap that returned a non-200 status, was not XML, failed to parse, or timed out. A site with no sitemap at all produces a summary with urlCount: 0 and the fallback 404 warnings, and costs nothing.

Pricing

Pay per event: one url-discovered event ($0.0003) per unique URL row written. Robots-only lookups, failed sitemaps and sites without sitemaps are free. Discovering 100,000 URLs costs $30 plus nothing else; there is no per-run or per-site fee.

Limits and behaviour

  • Sitemaps larger than 256 MiB (after gzip expansion) are rejected with a warning.
  • Sitemap index recursion stops at depth 20 or 1,000 sitemap fetches per site.
  • Only http and https URLs are accepted; URLs with embedded credentials are rejected.
  • Robots.txt bodies over 2 MiB or served as HTML are ignored with a warning.
  • The actor never fetches page bodies, so it does not consume site bandwidth beyond the sitemap files themselves. It honours nothing in robots.txt except reading it; use the summary's disallow rules to configure your own crawler.

Running locally

python -m venv .venv
.venv\Scripts\python.exe -m pip install -r requirements.txt
$env:APIFY_LOCAL_STORAGE_DIR = "$PWD\storage\manual"
New-Item -ItemType Directory -Force "$env:APIFY_LOCAL_STORAGE_DIR\key_value_stores\default"
Copy-Item INPUT.json "$env:APIFY_LOCAL_STORAGE_DIR\key_value_stores\default\INPUT.json"
.venv\Scripts\python.exe -m src
.venv\Scripts\python.exe -m unittest discover -s tests -v

validation/run_live.ps1 runs INPUT.json in fresh local storage and writes validation/results.json. See VALIDATION.md for the latest real-site results.

Example output

One real dataset row from storage/live-20260926-040335/datasets/default/000000001.json, trimmed by omitting fields without changing retained values:

{
"site": "https://wordpress.org/",
"url": "https://wordpress.org/",
"source": "robots",
"sitemapUrl": "https://wordpress.org/news-sitemap.xml",
"lastmod": null
}

This row came from the saved 900-row validation run. No HEAD check was requested, so status is absent. A null lastmod means the sitemap supplied no value for this row. The earlier Output examples illustrate the schema.

Use cases

  • SEO teams can inventory sitemap-advertised URLs before selecting pages for a technical audit.
  • Website migration teams can compare exported URL lists with planned redirects and replacement pages.
  • Content operations teams can compare dated exports to identify newly advertised articles for review.
  • Data engineering teams can seed a downstream crawler and use the robots summary to configure its access rules.

Pricing example: 10,000 unique matching URL rows written x $0.0003 per url-discovered event = $3.00 in event fees, using .actor/pay_per_event.json. Local runs do not bill.

Limitations

Discovery covers URLs advertised through the sitemaps reached within the configured bounds; it is not a complete inventory of every reachable page. A returned URL does not prove that its page is available or indexable. Inspect per-site warnings, especially when a cap stops discovery.