Robots.txt Auditor | Crawl Rules & Sitemap Discovery
Pricing
$5.00 / 1,000 completed checks
Robots.txt Auditor | Crawl Rules & Sitemap Discovery
Read robots.txt for supplied websites and export user-agent groups, allow/disallow rules, sitemap declarations, extension directives and parsing warnings. Identify missing files and compare content hashes in your crawl configuration workflows.
Pricing
$5.00 / 1,000 completed checks
Rating
0.0
(0)
Developer
Austin Aryain
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Robots.txt Auditor | Crawl Rules & Sitemap Discovery
Read robots.txt for supplied websites and export user-agent groups, allow/disallow rules, sitemap declarations, extension directives and parsing warnings. Identify missing files and compare content hashes in your crawl configuration workflows.
How it works
Supply a website URL. The Actor replaces the path with /robots.txt on that origin, follows up to four HTTP redirects, and parses the returned UTF-8 text. It groups consecutive user-agent declarations with their allow and disallow rules, collects Sitemap directives, and preserves unknown directives such as Crawl-delay or Content-Signal separately. The content hash lets your own workflow detect edits between runs. Missing files returning 404 or 410 are useful completed results and are charged.
Quick start
- Enter one or more public URLs in the Input tab, beginning with the supplied example.
- Set a maximum run charge. A completed check costs $0.005; checking 10 sources once costs $0.05.
- Start the Actor and inspect the dataset. Download JSON, CSV or Excel, or consume results through the Apify API.
- Inspect the OUTPUT run summary as well as the dataset: failed or unprocessed inputs appear there. Save the input as a task if you want to schedule future runs.
Pricing
$0.005 per completed check ($5 per 1,000), with platform usage included. There are no separate Actor-start or dataset-item fees. Empty and unchanged successful checks are charged. The maximum charge is checked before each source request and again before output. Failed network or format checks are free; see the specific HTTP-response cases below. Billing is per completed source check, not per nested array item, extracted URL, change or schema block.
Limits and interpretation
The robots file must fit within 500 KB and 10,000 directives. This is a directive inventory, not a crawler permission decision or a guarantee of search-engine behavior. It does not evaluate a target path against allow/disallow patterns, interpret extensions, visit sitemap links, or crawl the website. HTTP 401, 403, 429, 5xx, HTML masquerading as robots text, and network failures are free errors. There is no retained cross-run state.
A run accepts 1-50 unique input URLs and requests them sequentially. Each check has an 18-second network deadline; new checks stop after 160 seconds. Use a 240-second run timeout and 512 MB memory. If the time or charge limit stops a batch, OUTPUT lists uncheckedUrls for a later run. No response exceeding the configured byte limit is accepted, and a complete record must fit within 6 MB. The Actor permits only public HTTP(S) destinations on standard ports, pins a validated DNS address per request, and refuses redirects into private networks or from HTTPS to HTTP.
The Actor uses direct HTTP requests, without a browser, residential proxy, login, CAPTCHA solving or access-control bypass. Rate limits and blocks may prevent checks. Avoid secret-bearing URLs. Results describe the source and network observed at check time.
Integrations and support
Connect the dataset and OUTPUT summary to your own n8n, Make, Zapier or API workflow. This Actor produces data; it does not automatically send email, Slack messages or webhooks to third parties. No external account credentials are needed for the supplied public examples. Report reproducible issues in the Actor Issues tab, including a non-sensitive input and run link. This is an independent utility and is not endorsed by the websites, standards bodies or services it reads.
Input example
{"urls": ["https://www.wikipedia.org/"]}
See the Input tab for all supported fields. Results are available through the dataset API and can be downloaded as JSON, CSV or Excel.
Output fields
| Field | Meaning |
|---|---|
| inputUrl | Normalized supplied URL. |
| checkedAt | Check time in ISO format. |
| robotsUrl | Final robots.txt URL. |
| httpStatus | Observed HTTP status. |
| exists | False for HTTP 404 or 410. |
| groups | User-agent groups and their ordered allow/disallow rules. |
| sitemaps | Unique declared sitemap values; not downloaded or validated. |
| extensions | Other directives such as crawl-delay with observed group context. |
| warnings | Line numbers and structural parsing warnings. |
| contentHash | SHA-256 of decoded robots text; null for absent files. |
Output example
Example from a public source check; live values vary. Long items, changes, groups and blocks arrays are shortened to two entries here for readability; the actual record contains the complete arrays within the documented limits.
{"inputUrl": "https://www.wikipedia.org/","checkedAt": "2026-09-07T21:05:12.189Z","robotsUrl": "https://en.wikipedia.org/robots.txt","httpStatus": 200,"exists": true,"groups": [{"userAgents": ["MJ12bot"],"rules": [{"directive": "disallow","path": "/"}]},{"userAgents": ["Mediapartners-Google*"],"rules": [{"directive": "disallow","path": "/"}]}],"sitemaps": ["https://en.wikipedia.org/w/rest.php/site/v1/sitemap/0"],"extensions": [{"directive": "crawl-delay","value": "5","userAgents": ["SemrushBot"]}],"warnings": [],"contentHash": "62ede7ec5bff7cd9bdbd97ee37273e0162db5b1930db11ee0665b1fa1254fcc3"}