Bulk AI Crawler Policy Checker avatar

Bulk AI Crawler Policy Checker

Pricing

from $20.00 / 1,000 completed bulk crawler policy checks

Go to Apify Store
Bulk AI Crawler Policy Checker

Bulk AI Crawler Policy Checker

Check many domains' robots.txt AI crawler policy in one run.

Pricing

from $20.00 / 1,000 completed bulk crawler policy checks

Rating

0.0

(0)

Developer

Sentinel Signal

Sentinel Signal

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Check many domains' robots.txt AI crawler policy in a single run — the bulk-native sibling of AI Crawler Access Checker.

What it checks

  • robots.txt rules for named AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot, and others), for every domain submitted
  • Whether each domain's robots.txt itself is reachable and parseable

This reports declared policy only, not traffic. It does not measure or infer whether any crawler has crawled any of the submitted domains — only what each domain's robots.txt currently permits or disallows.

Pricing

This is a paid, pay-per-event Actor: $0.02 per successfully checked domain.

One crawler-access-check-bulk event is charged per domain, only after a real result exists for that domain (success or a useful partial). There is no charge for a domain rejected by local input validation, blocked by the public-network policy, or unreachable.

Input

  • domains (required): 1 to 100 public domains to check, e.g. example.com.

Example input:

{
"domains": ["example.com", "example.org"]
}

Output

The default dataset receives one result envelope per domain; RUN_SUMMARY in the default key-value store records processing, delivery, billing, and dependency totals across the whole run.

Representative abbreviated output (one dataset row):

{
"status": "success",
"target": {"identifier": "example.com"},
"result": {
"domain": "example.com",
"crawlers": {"GPTBot": "blocked", "ClaudeBot": "allowed", "PerplexityBot": "allowed"},
"robotsUrl": "https://example.com/robots.txt"
}
}

restricted means API-IFY rejected that domain locally before calling Verify. unreachable means the domain could not be reached. dependency_unavailable means Verify was temporarily unavailable. Duplicate domains within one run are collapsed to a single check with a warning.

How it works

This Actor is a thin wrapper around Sentinel Verify's POST /v1/utilities/inspect-ai-crawler-policy endpoint, called once per unique domain with bounded concurrency. Each domain is checked locally against API-IFY's public-network policy before being sent to Verify.

Limitations

  • Public HTTPS domains only; private, loopback, link-local, metadata, reserved, and CGNAT targets are rejected.
  • Up to 100 domains per run.
  • Reports declared robots.txt policy, not observed crawler activity.
  • Check availability depends on Sentinel Verify.

Privacy and security

Each target domain is sent to Sentinel's live Utility API (/v1/utilities/inspect-ai-crawler-policy) for analysis after local public-network validation. Results are stored in the customer's Apify run dataset and summary store. VERIFY_API_KEY is an operator-managed Actor secret — a scoped utilities:read Intelligence API key sent as a bearer token on every request, the same credential mechanism used by any other Utility API consumer; VERIFY_BASE_URL is an operator setting, not an Actor input.

Support

For a reproducible support request, provide the Apify run ID, result itemId, and status. Contact Sentinel Signal Systems through the support link on the Actor page.