Sitemap & robots.txt URL Discovery (no crawling)
Pricing
$0.30 / 1,000 url discovereds
Sitemap & robots.txt URL Discovery (no crawling)
Feed it site roots, get back URLs from their sitemaps plus each site's disallow rules and crawl-delay. Streams sitemaps, follows nested indexes and handles gzip without crawling page bodies. For seeding crawlers, SEO audits and migration checks.
Pricing
$0.30 / 1,000 url discovereds
Rating
0.0
(0)
Developer
Paul Vasquez
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Sitemap, robots.txt & URL Discovery
Turn a list of websites into a list of their URLs without crawling a single page. The actor reads each site's robots.txt, follows every sitemap it declares (including nested sitemap indexes and gzip files), falls back to the common sitemap locations when robots.txt is silent, and writes one dataset row per unique URL. Optionally it HEAD-checks every URL so you can see which ones are alive.
Typical uses: seeding a crawler with a complete URL list, SEO audits (which URLs a site advertises, when they changed), content inventories, migration checks, monitoring competitors' new pages, and finding a site's disallow rules and crawl-delay before you scrape it.
What it does per site
- Fetches
/robots.txtand parsesSitemap:lines, user-agent groups,Disallow/Allowrules andCrawl-delay. - Queues every declared sitemap. If robots.txt declares none, it tries
/sitemap.xml,/sitemap_index.xml,/sitemap-index.xml,/wp-sitemap.xmland/sitemap.xml.gz. - Streams each sitemap with an incremental XML parser (memory stays flat even for 50 MB files), detects gzip automatically, follows sitemap indexes up to 20 levels deep and 1,000 sitemap requests per site, and rejects HTML or DTD-bearing responses safely.
- Deduplicates URLs, applies your optional regex filter, stops at
maxUrlsPerSite, and optionally HEAD-checks each URL with 10 concurrent requests. - Writes the rows to the dataset and a
SUMMARY-<host>record to the key-value store.
If you pass a sitemap URL directly (anything with a path, e.g. https://example.com/news-sitemap.xml) it is fetched as a sitemap for that origin without consulting robots.txt fallbacks.
Input
| Field | Type | Default | Notes |
|---|---|---|---|
startUrls | array of URLs | required | Site roots or explicit sitemap URLs, 1 to 100. Grouped by origin. |
includeRobots | boolean | true | Parse robots.txt for sitemaps and rules. |
followSitemapIndexes | boolean | true | Recurse into sitemap index files. |
maxUrlsPerSite | integer | 5000 | Hard cap per origin, up to 1,000,000. |
includeLastmod | boolean | true | Copy lastmod from the sitemap. |
includeStatusCheck | boolean | false | HEAD every URL and record the HTTP status. |
filterPattern | string | none | Python regular expression searched against each full URL. |
timeoutSecs | integer | 30 | Per request, 1 to 300. |
proxyConfiguration | object | none | Apify Proxy or custom proxy URLs. |
Output
Dataset row per URL:
{"site": "https://wordpress.org/","url": "https://wordpress.org/news/2026/09/example/","source": "sitemap-index","sitemapUrl": "https://wordpress.org/news/sitemap-1.xml","lastmod": "2026-09-24T10:12:00+00:00","changefreq": null,"priority": null,"status": 200}
source is one of robots (sitemap declared in robots.txt), sitemap-index (reached through an index), fallback (found at a common path) or sitemap (URL you supplied directly). status is present only when includeStatusCheck is on.
Key-value store record SUMMARY-<host> per site:
{"site": "https://wordpress.org/","robotsFound": true,"sitemapsFound": 19,"urlCount": 300,"disallowCount": 4,"crawlDelay": null,"warnings": ["maxUrlsPerSite reached; discovery stopped"],"robotsGroups": [{"userAgents": ["*"], "disallow": ["/wp-admin/"], "allow": [], "crawlDelay": null}],"elapsedSeconds": 7.9}
Warnings list every sitemap that returned a non-200 status, was not XML, failed to parse, or timed out. A site with no sitemap at all produces a summary with urlCount: 0 and the fallback 404 warnings, and costs nothing.
Pricing
Pay per event: one url-discovered event ($0.0003) per unique URL row written. Robots-only lookups, failed sitemaps and sites without sitemaps are free. Discovering 100,000 URLs costs $30 plus nothing else; there is no per-run or per-site fee.
Limits and behaviour
- Sitemaps larger than 256 MiB (after gzip expansion) are rejected with a warning.
- Sitemap index recursion stops at depth 20 or 1,000 sitemap fetches per site.
- Only
httpandhttpsURLs are accepted; URLs with embedded credentials are rejected. - Robots.txt bodies over 2 MiB or served as HTML are ignored with a warning.
- The actor never fetches page bodies, so it does not consume site bandwidth beyond the sitemap files themselves. It honours nothing in robots.txt except reading it; use the summary's disallow rules to configure your own crawler.
Running locally
python -m venv .venv.venv\Scripts\python.exe -m pip install -r requirements.txt$env:APIFY_LOCAL_STORAGE_DIR = "$PWD\storage\manual"New-Item -ItemType Directory -Force "$env:APIFY_LOCAL_STORAGE_DIR\key_value_stores\default"Copy-Item INPUT.json "$env:APIFY_LOCAL_STORAGE_DIR\key_value_stores\default\INPUT.json".venv\Scripts\python.exe -m src.venv\Scripts\python.exe -m unittest discover -s tests -v
validation/run_live.ps1 runs INPUT.json in fresh local storage and writes validation/results.json. See VALIDATION.md for the latest real-site results.
Example output
One real dataset row from storage/live-20260926-040335/datasets/default/000000001.json, trimmed by omitting fields without changing retained values:
{"site": "https://wordpress.org/","url": "https://wordpress.org/","source": "robots","sitemapUrl": "https://wordpress.org/news-sitemap.xml","lastmod": null}
This row came from the saved 900-row validation run. No HEAD check was requested, so status is absent. A null lastmod means the sitemap supplied no value for this row. The earlier Output examples illustrate the schema.
Use cases
- SEO teams can inventory sitemap-advertised URLs before selecting pages for a technical audit.
- Website migration teams can compare exported URL lists with planned redirects and replacement pages.
- Content operations teams can compare dated exports to identify newly advertised articles for review.
- Data engineering teams can seed a downstream crawler and use the robots summary to configure its access rules.
Pricing example: 10,000 unique matching URL rows written x $0.0003 per url-discovered event = $3.00 in event fees, using .actor/pay_per_event.json. Local runs do not bill.
Limitations
Discovery covers URLs advertised through the sitemaps reached within the configured bounds; it is not a complete inventory of every reachable page. A returned URL does not prove that its page is available or indexable. Inspect per-site warnings, especially when a cap stops discovery.