AI Crawler Checker - robots.txt Checker for GPTBot & AI Bots avatar

AI Crawler Checker - robots.txt Checker for GPTBot & AI Bots

Pricing

Pay per event

Go to Apify Store
AI Crawler Checker - robots.txt Checker for GPTBot & AI Bots

AI Crawler Checker - robots.txt Checker for GPTBot & AI Bots

Which AI bots can read a site? Bulk-check up to 10,000 domains for 26 AI crawlers (GPTBot, ClaudeBot, PerplexityBot...) in robots.txt: AI search score, fix snippet, Content Signals. $2/1k sites.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Yukai Lin

Yukai Lin

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

What does AI Crawler Access Checker do?

It tells you which AI crawlers are allowed to read a website according to its robots.txt, scores it, gives you a ready-to-paste robots.txt fix, and can track changes week to week. It also reads Cloudflare Content Signals (Content-Signal: search=yes, ai-train=no), and checks for an llms.txt file and noai directives. Check one site or up to 10,000 domains per run for $2 per 1,000 sites.

Why it matters: if AI search crawlers such as OAI-SearchBot, Claude-SearchBot or PerplexityBot are blocked, the site is unlikely to be cited in ChatGPT, Claude or Perplexity answers. Blocking training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) does not have that effect.

Free daily benchmark: see how 1,005 of the world's most visited websites treat these 29 crawlers, by category and for the top 100, in the AI Crawler Index. It is updated every day with the same checks as this Actor, and you can download the full table as CSV.

What it checks

  • πŸ€– 29 AI crawlers from OpenAI, Anthropic, Perplexity, Google, Apple, Meta, Amazon, DuckDuckGo, Mistral, Common Crawl, ByteDance and Cohere, each labelled as training, search, user (fetches made when a person asks an assistant) or ads (ad review, shown for information and not scored). Crawler list last checked against each vendor's documentation: 2026-10-01 (see the list below)
  • 🏷️ Content Signals: Cloudflare's Content-Signal line in robots.txt (search, ai-input, ai-train = yes/no) is returned as contentSignals, mentioned in the verdict, and tracked between runs
  • πŸ“œ robots.txt rules evaluated like search engines do: the most specific user-agent group, longest matching rule, * and $ wildcards
  • 🚦 Per-bot status: allowed, partial (home page allowed but a tested path is blocked, or the bot's own group has Disallow rules), blocked, or unknown
  • πŸ”’ Honest "unknown": if robots.txt answers 401, 403, 418, 429 or a bot challenge page, the result is policy: "unknown" with scores null, not "open", and the site is not charged
  • πŸ“Š Scores and policy class: aiSearchScore (AI search and assistant crawlers allowed, 0-100), aiAccessScore (all AI crawlers), policy: open, search-only, blocks-search, restrictive, no-robots or unknown
  • πŸ› οΈ Fix snippets: recommendedRobotsSnippet re-allows the blocked AI search and assistant crawlers (keeping your other rules); suggestedPolicy is a complete template that blocks training crawlers and allows AI search
  • πŸ” Change tracking: give the run a monitor name and each site is compared with its previous result (comparison.changedBots, policy, llms.txt); optionally output only changed sites and get a Slack, Discord or JSON webhook
  • πŸ›£οΈ Any paths you choose (e.g. /blog/, /products/), not only the home page
  • πŸ“„ llms.txt and llms-full.txt: present or not (HTML "not found" pages are not counted)
  • 🚫 noai / noimageai in the home page's meta robots or X-Robots-Tag header
  • πŸ—ΊοΈ Sitemaps listed in robots.txt, and a plain-language verdict for every site

Which tool should I use?

AI Crawler Access Checker (this Actor)GEO Readiness Audit (by TidyTools)
Question"Do these 5,000 domains allow AI crawlers in robots.txt, and what changed since last week?""Why is my site not cited by AI answers, and what should I fix first?"
Depthrobots.txt policy only: custom paths, matched rules, fix snippet, change trackingLive firewall test (catches CDN "block AI bots" rules and challenge pages), content without JavaScript, structured data, score, HTML report
Best forResearchers, data teams, agencies watching many domainsSite owners and agencies auditing one site or a few competitors
Price$0.002 per siteabout $0.018 per site

For a firewall diagnosis (robots.txt allows a bot but the CDN blocks it), use GEO Readiness Audit.

How much does it cost?

EventPrice
Checked website$2.00 / 1,000 websites
Re-checked website, unchanged (monitoring)$0.50 / 1,000 websites

No start fee. Unreachable websites and sites whose robots.txt cannot be read are not charged. Scores, fix snippets and change tracking are included. With a monitor name, a site whose result did not change since the last run is charged at the re-check price ($0.0005); new and changed sites cost $0.002. Example: watching 1,000 domains weekly where 2% change costs about $0.53 per week. With Only output changed sites or Only output sites with a problem, skipped sites are still checked and charged. Higher Apify plans get volume discounts.

For comparison, other AI crawler checkers in the Apify Store charge $0.005 to $0.02 per site, some with a start fee, and some limit a run to 50 or 100 sites (checked September 2026).

Control your cost

  • Charged: each website whose robots.txt was read ($0.002), or $0.0005 for an unchanged site in a monitor.
  • Free: invalid input lines, duplicate lines (merged into one site), unreachable websites, and robots.txt that could not be read (401, 403, 418, 429, bot challenge). Every row has charged: true or false, and the error text of free rows ends with "(not charged)".
  • At the start, the log and status message show the plan: number of websites Γ— price = the most the run can cost, compared with your maximum charge per run (set in the run options).
  • When that maximum is reached, the run stops checking new sites. The status message says how many were not checked, and SUMMARY.notProcessed lists them (count and up to 100 inputs) so you can run them again. A monitor webhook sent from such a run carries incomplete: true.
  • If Apify restarts the run (server migration or Resurrect), items already finished are skipped and not charged again (SUMMARY.resumedSkipped).
  • Time limit per website (Advanced settings, default 120 seconds): a site whose robots.txt, llms.txt and home page take longer in total gets a row with errorType: "timeout" and is not charged.

How to use it

  1. Paste websites, one per line (example.com or full URLs). A line with several domains separated by commas or spaces is split. A line that is not a website gets its own error row (errorType: "invalid_input", not charged) and the other sites are still checked. A URL with a path (e.g. nytimes.com/section/world) also tests that path.
  2. Optional: add paths to test.
  3. Optional, for monitoring: set a monitor name, schedule the Actor (e.g. weekly), and add a webhook URL (Slack and Discord URLs get a formatted message).
  4. Click Start and export the results as JSON, CSV or Excel. The Fix snippets view lists the robots.txt lines to paste for each site.

Input example

{
"websites": ["nytimes.com", "python.org", "stripe.com, lowes.com"],
"paths": ["/", "/search"],
"monitorName": "clients-weekly",
"outputOnlyChanges": false,
"webhookUrl": "https://hooks.slack.com/services/..."
}

Output example (real result, September 2026, shortened)

{
"site": "https://nytimes.com",
"verdict": "Some AI search/assistant crawlers are blocked (OAI-SearchBot, Claude-SearchBot, PerplexityBot, meta-webindexer, DuckAssistBot, ChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher): the site may not appear in those AI answers.",
"policy": "blocks-search",
"aiSearchScore": 47,
"aiAccessScore": 35,
"blockedBots": ["GPTBot", "OAI-SearchBot", "ChatGPT-User", "ClaudeBot", "..."],
"partialBots": ["Google-GeminiNotebook", "Google-Agent", "Applebot", "Amazonbot", "..."],
"contentSignals": null,
"recommendedRobotsSnippet": "# Allow AI search and assistant crawlers (they decide whether the site can appear in AI answers)\n# Also remove \"Disallow: /\" from the existing group(s) for: OAI-SearchBot, ChatGPT-User, ...\nUser-agent: OAI-SearchBot\nUser-agent: ChatGPT-User\n...\nAllow: /\nDisallow: /ads/\n...",
"suggestedPolicy": "# AI model training crawlers: blocked (does not affect AI search answers)\nUser-agent: GPTBot\n...\nDisallow: /\n\n# AI search and assistant crawlers: allowed\nUser-agent: OAI-SearchBot\n...",
"robotsTxt": { "status": 200, "found": true, "state": "found", "reason": null, "sitemaps": ["https://www.nytimes.com/sitemaps/new/news.xml.gz", "..."] },
"llmsTxt": { "found": false, "url": "https://nytimes.com/llms.txt", "status": 404 },
"noAiDirective": false,
"comparison": {
"status": "changed",
"changedBots": [{ "bot": "MistralAI-User", "before": "allowed", "after": "partial" }],
"llmsTxtChanged": false,
"policyChanged": null,
"previousCheckedAt": "2026-09-29T02:18:46.163Z"
},
"bots": [
{ "bot": "GPTBot", "company": "OpenAI", "purpose": "training", "status": "blocked", "allowedAll": false, "matchedGroup": "GPTBot",
"paths": [{ "path": "/", "allowed": false, "rule": "Disallow: /" }] }
],
"charged": true
}

A site whose robots.txt refuses the request (lowes.com answered HTTP 403) is reported as unknown and not charged:

{ "site": "https://lowes.com", "policy": "unknown", "aiSearchScore": null, "robotsTxt": { "status": 403, "state": "unavailable", "reason": "robots.txt answered HTTP 403" }, "charged": false }

The SUMMARY record in the key-value store has status (SUCCESS, PARTIAL_RESULTS, FAILED, NO_RESULTS or LIMIT_REACHED), counts (including invalidInputs and duplicateInputs), changedSites, the webhook result, costPlan and notProcessed. Unreachable sites get success: false and an errorType.

Every row carries input, your original line, so results map back to your spreadsheet. When several lines point to the same site (e.g. example.com and EXAMPLE.com), the site is checked and charged once and inputs lists all of them. testedPaths lists the paths checked for every bot. A line that is not a website:

{ "input": "acme", "site": null, "success": false, "errorType": "invalid_input", "error": "\"acme\" is not a website: not a URL or domain (not charged)", "charged": false }

A line of words that are not domains (e.g. Acme Inc) stays one line and gives one error row; a line of domains separated by spaces (a.com b.com) is split.

A site that publishes Content Signals (real result, www.cloudflare.com, September 2026):

{
"site": "https://www.cloudflare.com",
"verdict": "Open to all listed AI crawlers. Content-Signal: search allowed, AI answers (ai-input) allowed, AI training (ai-train) allowed.",
"policy": "open",
"contentSignals": { "search": "yes", "aiInput": "yes", "aiTrain": "yes", "userAgent": "*", "raw": "ai-train=yes, search=yes, ai-input=yes" },
"comparison": { "status": "changed", "contentSignalChanged": { "before": null, "after": "ai-train=yes, search=yes, ai-input=yes" } }
}

contentSignals is null when robots.txt has no Content-Signal line. A missing signal means "no preference" (neither allowed nor refused). When different user-agent groups declare different signals, all of them are listed in contentSignalGroups.

Policy classes

policyMeaning
openNo listed AI crawler is blocked at the home page
search-onlyOnly training crawlers are blocked; AI search and assistants are allowed
blocks-searchSome AI search or assistant crawlers are blocked
restrictiveAll AI search and assistant crawlers are blocked (or robots.txt returns a server error)
no-robotsNo robots.txt: everything is allowed
unknownrobots.txt could not be read (401, 403, 418, 429 or a bot challenge)

Crawlers checked (last checked 2026-10-01)

CompanyTrainingAI searchUser-triggeredAds review
OpenAIGPTBotOAI-SearchBotChatGPT-User*OAI-AdsBot
AnthropicClaudeBotClaude-SearchBotClaude-User
PerplexityPerplexityBotPerplexity-User*
GoogleGoogle-ExtendedGoogle-CloudVertexBot (Vertex AI Agents, crawls requested by the site owner)Google-GeminiNotebook*, Google-Agent*
AppleApplebot-ExtendedApplebot (follows Googlebot's rules when not named)
MetaMeta-ExternalAgentmeta-webindexermeta-externalfetcher*meta-externalads
AmazonAmazonbotAmzn-SearchBotAmzn-User*
DuckDuckGoDuckAssistBot
MistralMistralAI-TrainingMistralAI-IndexMistralAI-User
Common CrawlCCBot
ByteDance, CohereBytespider, cohere-ai (no vendor documentation; widely listed in robots.txt files)

Every token was checked against the company's own crawler documentation. * The vendor says this user-triggered fetcher generally does not follow robots.txt; such bots carry ignoresRobots: true in bots, and robots.txt rules for them express your wish rather than a guarantee.

Ads review crawlers (purpose: "ads") check ad landing pages or improve advertising products. They are listed so you can see your rules for them, but they count in no score and no policy class. Amazonbot is labelled training since 2026-10-01: Amazon's page says it "may be used to train Amazon AI models", while Amzn-SearchBot is the crawler for search experiences such as Alexa.

Use with AI agents (MCP)

Connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/ai-crawler-access-checker) to Claude, Cursor or any MCP client, then ask e.g. "Which of these 50 domains block ChatGPT search in robots.txt, and what should they paste to fix it?"

{ "websites": ["example.com", "example.org"] }

Failed items are not charged and carry an errorType.

Use it from code and integrations

Run it from your own code with the Apify API. This call waits for the run and returns the results as JSON (replace YOUR_TOKEN with your Apify API token):

curl -X POST "https://api.apify.com/v2/acts/tidytools~ai-crawler-access-checker/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"websites":["https://www.nytimes.com","python.org","docs.anthropic.com"]}'

The synchronous endpoint waits up to 5 minutes. For bigger runs, start the run with POST https://api.apify.com/v2/acts/tidytools~ai-crawler-access-checker/runs and read the dataset when it finishes, or use the apify-client package for JavaScript or Python.

Schedules and integrations: run it daily or weekly with Apify Schedules, get a webhook when a run finishes, or send the results to Zapier, Make, n8n, Google Sheets, Slack and other apps with Apify integrations. Results can be exported as JSON, CSV, Excel or XML.

FAQ

Will blocking GPTBot remove my site from ChatGPT answers? No. GPTBot collects training data. ChatGPT search uses OAI-SearchBot, and ChatGPT-User fetches pages when a person asks. You can block GPTBot and still allow the other two; suggestedPolicy does exactly that.

What is Content-Signal? An extension Cloudflare introduced for robots.txt: a line such as Content-Signal: search=yes, ai-input=no, ai-train=no inside a user-agent group states how content may be used after it is fetched: search (search index with links and snippets), ai-input (feeding AI answers, e.g. RAG or grounding) and ai-train (model training). It does not block crawling by itself; it states the site's preference. This Actor returns the signals for all crawlers (*) in contentSignals.

My site has no llms.txt. Should it? It is an emerging proposal, not yet used by the major AI search engines. If you want one, llms.txt Generator (by TidyTools) builds llms.txt and llms-full.txt from your sitemap.

What does "partial" mean? The home page is allowed, but one of your tested paths is blocked, or the bot has its own group with Disallow rules.

I changed the paths to test (or the crawler list was updated). Will every site show as changed? No. Each monitor snapshot stores the tested paths and the crawler list version. When the paths changed, only bots that became blocked or unblocked are reported; when the crawler list changed, new crawlers and the policy class are not compared that time. comparison.note says so, and a site with no other change is charged the re-check price.

robots.txt allows AI search. Can the site still block AI crawlers? Yes, a CDN or firewall rule can block them regardless of robots.txt. The verdict mentions this; GEO Readiness Audit (by TidyTools) runs a live firewall test.

Use cases

  • Your own site: confirm AI search crawlers can read it (and training crawlers are handled the way you want), and copy the fix snippet if not
  • SEO/GEO agencies: audit client sites in bulk and get a weekly alert when a client's robots.txt starts blocking AI search
  • Research: measure how many sites in a list block AI crawlers, and how that changes over time (the AI Crawler Index does this daily for 1,005 popular sites)

Tips

  • Advanced settings: add a Proxy to read robots.txt and llms.txt through Apify Proxy (billed to your Apify account) for sites that refuse data-center requests. Plain HTTP requests from controls how the home page is read for the noai check.
  • The webhook can be signed: set a webhook signing secret and verify the X-Signature-256: sha256=<hex> header (HMAC-SHA256 of the raw body).

Limitations

  • robots.txt is a request, not an enforcement: this Actor reports what the site asks crawlers to do.
  • Firewalls or bot protection (e.g. blocking by IP or user agent at the CDN) are not detected here; use GEO Readiness Audit for that.
  • The crawler list covers the major AI companies (checked 2026-10-01); others may exist.
  • Adding crawlers to the list changes scores: a site that names only a few AI crawlers in robots.txt is judged on all of them (29 since 2026-10-01; ad-review crawlers are not scored).

Support

Open an issue in the Issues tab with the website. Issues are checked regularly.