Robots.txt Policy Evidence API avatar

Robots.txt Policy Evidence API

Pricing

from $1.50 / 1,000 policy observations

Go to Apify Store
Robots.txt Policy Evidence API

Robots.txt Policy Evidence API

Extract public robots.txt directives, policy hashes and declared sitemap URLs. Compare policy changes and inspect explicitly labelled Python parser evaluations.

Pricing

from $1.50 / 1,000 policy observations

Rating

0.0

(0)

Developer

Salih Can Kurnaz

Salih Can Kurnaz

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Extract public robots.txt directives, policy hashes and declared sitemap URLs. Compare policy changes and inspect explicitly labelled Python parser evaluations.

Quick start

{
"urls": [
"https://www.python.org/downloads/"
],
"userAgents": [
"*",
"Googlebot",
"GPTBot"
]
}

Download the default dataset as JSON, CSV or Excel. Read OUTPUT before relying on completeness; source availability can change. Use sources you are authorized to access.

Documented behavior and limits

Reads the origin robots.txt and returns recognized directives with line numbers, unmodified declared sitemap addresses, a policy-text hash and target-path evaluations explicitly labelled PYTHON_ROBOTFILEPARSER. A 404/410 policy produces MISSING and a billable missing-policy observation; 403, 429, server errors or HTML responses are source errors with permission unknown. This Actor does not fetch the target page or sitemap addresses. Evaluation uses Python RobotFileParser, not a Googlebot/GPTBot implementation or complete RFC 9309 conformance validator: wildcard/end-anchor, group specificity and percent-encoding behavior can differ from actual crawlers. Treat the raw policy as evidence and the canFetch value as interpreter output, not access authorization, indexing or AI-training permission. Only syntactic directives are returned; sitemap addresses are not validated or fetched.

Changes and snapshots

Each result includes itemKey, contentHash and observedAt. Pass SNAPSHOT.previousHashes in the next input. NO_BASELINE means no prior map; FIRST_SEEN means absent from a supplied map; UNCHANGED/CHANGED compare hashes. FIRST_SEEN does not prove newly created content. Snapshot contains only emitted rows. Missing rows are not deletions. No hidden persistence or scheduled tasks. For header observations, redirect trace is excluded from the hash; it remains in output. Other observations hash their normalized fields, excluding pagination source URL and time.

Pricing

$0.0015 per emitted policy observation ($1.5 per 1,000), plus $0.005 Actor-start at the default 256 MB. Higher memory can multiply startup events under Apify rules. Event-only pricing includes platform usage; no paid upstream API, proxy or browser. Failed sources incur no result event, but the start event can apply. An intentional missing-policy observation is billable for the robots product. Check Store pricing before running. Budget truncation sets PARTIAL_BUDGET and complete=false.

Network and data

Public HTTP(S), ports 80/443 only; credentials and private IPs rejected. DNS addresses are validated and pinned while TLS hostname verification stays enabled. Up to five redirects, 2 MiB decoded responses and approximately 95 seconds network budget. Site/content/spec fetches respect robots.txt; missing 404/410 policy allows access, unavailable/denied policy is an error. The robots product reads the policy file itself and never requests the evaluated page. No login, cookies, proxy, access-block workaround, external reference fetching or active security tests. All failed sources are reported; successful sources can continue. A run with no rows and source errors fails; valid empty collections/specs can succeed. Sources, redirects and their robots files are the only network targets. Inputs and output stay in your Apify run storage under platform retention; no separate customer analytics endpoint. Do not provide secrets.

Integration and support

POST https://api.apify.com/v2/actors/soilair~robots-txt-policy-evidence-api/runs with your Apify authorization header and input JSON. Poll the returned run and retrieve its defaultDatasetId and defaultKeyValueStoreId for OUTPUT/SNAPSHOT. Configure integrations/schedules in Console if needed; the Actor creates neither. Developed with AI assistance, fixtures and public-source/cloud checks. Examples do not establish arbitrary-source compatibility. Report a run ID and expected behavior through the issue tab without credentials.

Source specification: https://www.rfc-editor.org/rfc/rfc9309.html