AI Crawler Access Checker: robots.txt & llms.txt avatar

AI Crawler Access Checker: robots.txt & llms.txt

Pricing

from $2.00 / 1,000 domain checkeds

Go to Apify Store
AI Crawler Access Checker: robots.txt & llms.txt

AI Crawler Access Checker: robots.txt & llms.txt

Check which AI crawlers each domain allows in robots.txt: GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot and 20 more, each with the exact rule that matched. Adds a one-line AI visibility verdict (open to AI search, blocks training …), llms.txt and llms-full.txt checks and sitemap URLs.

Pricing

from $2.00 / 1,000 domain checkeds

Rating

0.0

(0)

Developer

Offera Studio

Offera Studio

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 hours ago

Last modified

Share

What does AI Crawler Access Checker do?

AI Crawler Access Checker tells you, for any list of domains (up to 5,000 per run), which AI crawlers the site lets in and whether it publishes an llms.txt file. For every domain you get:

  • πŸ€– 26 AI crawlers checked against robots.txt: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, CCBot, meta-externalagent, Amazonbot, DuckAssistBot, Mistral AI's crawlers, Bytespider, cohere-ai, Diffbot and more
  • βœ… for each one: allowed, partial or blocked, the group and rule that matched (with its line number) and a plain-English reason
  • 🧭 a one-line AI visibility verdict: open to all AI crawlers, open to AI search, blocks AI training, blocks some AI search crawlers, blocks all AI crawlers …
  • πŸ“„ /llms.txt and /llms-full.txt: found or not, H1 title, summary, link sections, size and format issues
  • πŸ—ΊοΈ sitemap URLs from robots.txt (or /sitemap.xml when none are listed)
  • πŸ”Ž Googlebot and Bingbot access for context

Paste your domains, click Start, and export the results to CSV, Excel or JSON, or use them through the Apify API, Google Sheets, Make, Zapier or n8n.

Who is this AI crawler checker for?

  • SEO and GEO (generative engine optimisation) teams: check that ChatGPT search, Perplexity and Claude can actually read your pages, and that a CDN or plugin didn't block them.
  • Publishers and content owners: verify that AI training crawlers are blocked while AI search crawlers stay allowed, across all your sites and subdomains.
  • Agencies: audit clients' and prospects' AI visibility and llms.txt in bulk.
  • Researchers and journalists: measure how many sites in a sector block GPTBot, ClaudeBot or CCBot, and track it over time with a schedule.
  • Developers of AI tools: check before crawling whether a site allows your agent (add your own user agent token).

Training, search and user-triggered crawlers

AI companies now run several crawlers with different jobs, and a site can allow one and block another:

PurposeWhat it meansExamples
AI trainingCollects content to train or improve AI modelsGPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, meta-externalagent, Amazonbot, MistralAI-Training
AI searchIndexes pages so AI assistants can find and cite them in answersOAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot, meta-webindexer, Amzn-SearchBot, MistralAI-Index
User-triggeredFetches a page when a user asks an assistant about itChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher, Amzn-User, MistralAI-User
OtherAI-related crawling of another kindDiffbot, Google-CloudVertexBot

Every token was checked against the vendor's own documentation on 30 September 2026, and each result links to it (docs). Bytespider, anthropic-ai and cohere-ai are not documented by their vendors today but are so common in robots.txt files that they are reported too (documented: false); they don't change the verdict. Some vendors say their user-triggered fetchers may not follow robots.txt; that is shown in note.

How robots.txt is read

The Actor follows RFC 9309, the Robots Exclusion Protocol standard that Google, OpenAI, Anthropic and others follow:

  • User-agent lines are matched case-insensitively; several groups for the same crawler are combined; a crawler without its own group follows User-agent: *; with no matching group, nothing is restricted.
  • The longest matching rule wins, and Allow wins a tie. * matches anything and $ anchors the end, so Disallow: /*.pdf$ works as crawlers read it.
  • A missing robots.txt (404) means no restrictions. Only the first 500 KiB are read, as the standard allows.

What the three results mean:

ResultMeaning
blockedThe home page is disallowed and no Allow rule opens anything.
partialOnly some paths are open (for example Disallow: / with Allow: /blog/), or the site wrote rules for this crawler that close some paths.
allowedEverything else. General User-agent: * rules that close a few paths for every crawler (like /admin/) don't count as a restriction on AI crawlers; they are listed in disallowedPaths.

How to check AI crawler access

  1. Click Try for free and sign in to Apify.
  2. Paste domains into Domains, one per line (up to 5,000 per run). www.example.com and example.com are checked separately because they can have different robots.txt files.
  3. Optional: add your own crawler tokens under Extra user agents to check, or tick Include the robots.txt text.
  4. Click Start. The Overview tab shows one row per domain; AI crawlers shows one row per crawler with the rule that matched; llms.txt and Sitemaps have their own tabs.

Input example

{
"domains": ["example.com", "docs.example.com", "https://www.example.org/blog"],
"checkLlmsTxt": true,
"additionalUserAgents": ["YourBot"],
"includeRobotsTxt": false
}

Output example

One item per domain (shortened, made-up data):

{
"domain": "news.example",
"robotsTxtUrl": "https://news.example/robots.txt",
"robotsTxtStatus": "found",
"aiVisibility": "blocks-training-only",
"aiVisibilityLabel": "Open to AI search, blocks AI training",
"aiVisibilitySummary": "AI training: 8 of 8 blocked (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and 4 more); AI search: 0 of 7 blocked; user-triggered fetchers: 0 of 6 blocked.",
"trainingBlockedCount": 8,
"searchBlockedCount": 0,
"blockedAgents": ["GPTBot", "ClaudeBot", "Google-Extended", "Applebot-Extended", "CCBot", "meta-externalagent", "Amazonbot", "MistralAI-Training", "Bytespider", "anthropic-ai", "cohere-ai"],
"partialAgents": [],
"blocksAllCrawlers": false,
"googlebotAccess": "allowed",
"bingbotAccess": "allowed",
"agents": [
{
"agent": "GPTBot",
"vendor": "OpenAI",
"purpose": "training",
"documented": true,
"access": "blocked",
"group": "User-agent: GPTBot",
"rule": "Disallow: /",
"ruleLine": 16,
"reason": "Blocked by \"Disallow: /\" (line 16) in the \"User-agent: GPTBot\" group.",
"docs": "https://platform.openai.com/docs/bots"
},
{
"agent": "OAI-SearchBot",
"vendor": "OpenAI",
"purpose": "search",
"access": "allowed",
"group": "User-agent: OAI-SearchBot",
"rule": "Allow: /",
"reason": "Allowed: the \"User-agent: OAI-SearchBot\" group closes no paths.",
"note": "OpenAI says sites that block it are not shown in ChatGPT search answers."
}
],
"hasLlmsTxt": true,
"hasLlmsFullTxt": false,
"llmsTxt": {
"url": "https://news.example/llms.txt",
"found": true,
"valid": true,
"title": "News Example",
"summary": "Independent local news since 1998.",
"sectionsCount": 3,
"linksCount": 24,
"hasOptionalSection": true,
"sizeBytes": 3120,
"issues": []
},
"sitemaps": ["https://news.example/sitemap.xml", "https://news.example/sitemap-news.xml"],
"sitemapSource": "robots.txt",
"error": null
}

Domains that can't be checked get a row with an error code (invalid-domain, dns-not-found, connection-failed, timeout, tls-error, robots-txt-forbidden, robots-txt-rate-limited, robots-txt-server-error) and cost nothing.

What is llms.txt?

llms.txt is a proposed standard: a Markdown file at /llms.txt that gives AI assistants a short, curated map of a site. It must start with an H1 title; it should have a blockquote summary, and H2 sections with lists of links (- [name](url): notes). An ## Optional section holds links that can be skipped. /llms-full.txt is a common companion with the full content in one file. The Actor checks:

  • presence (an HTML "not found" page served at the address doesn't count),
  • the H1 title, summary, number of sections and links, and an Optional section,
  • size, and plain-English issues such as a missing title or sections without links.

How much does it cost?

This Actor uses pay per event:

EventPrice
Domain checked$0.002 per domain
Domain that doesn't resolve, can't be reached or refuses the requestfree
  • 1,000 domains cost $2; 5,000 domains cost $10.
  • Checking llms.txt and llms-full.txt is included.
  • Apify also charges a tiny standard start fee per run (about $0.000025 at the default 512 MB).
  • Apify's free plan includes $5 of monthly usage, enough for about 2,500 domains a month.
  • Set Maximum cost per run in the run options and the Actor stops when it is reached.

Limitations

  • robots.txt is a request, not a lock. It shows what a site asks crawlers to do. Some crawlers may ignore it, and some vendors say their user-triggered fetchers may not follow it (see note). Sites can also block AI crawlers at the firewall or CDN, which robots.txt can't show.
  • Google-Extended is not about Google Search. It controls Gemini training and grounding. Google Search, including its AI features, uses Googlebot and Search's own controls such as nosnippet, which robots.txt tokens for AI crawlers don't change.
  • One hostname per row. example.com and www.example.com can have different rules; enter the ones you care about.
  • Politeness: the Actor makes at most four small requests per domain (robots.txt, llms.txt, llms-full.txt and, if robots.txt lists no sitemap, /sitemap.xml), half a second apart, and only fetches files that the site's robots.txt allows for crawlers. llms-full.txt is read up to 256 KB; its size comes from the server's Content-Length.
  • A 401, 403 or 429 on robots.txt usually means the server blocks automated requests; such domains get a free error row instead of a guess.
  • Not legal advice: whether AI companies may use content is a legal question that robots.txt alone doesn't answer.

FAQ

Does blocking GPTBot keep my site out of ChatGPT?

Not entirely. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and ChatGPT fetches pages that users ask about with ChatGPT-User. The verdict shows each group separately, so you can see, for example, "Open to AI search, blocks AI training".

Why is a crawler "partial" when I blocked it completely?

Check reason and rule: usually another rule opens some paths again, for example Allow: /blog/ next to Disallow: /, or the crawler matches a group with only some paths disallowed. The line number points to the rule in your robots.txt.

Why does a domain return robots-txt-forbidden?

The server answered 401 or 403 to our request for robots.txt, which usually means a firewall blocks automated requests. Its rules can't be read, so the row is free.

Can I check my own crawler?

Yes. Add its robots.txt token under Extra user agents to check. It appears in agents with purpose custom.

Can I run it on a schedule?

Yes. Save a task with your domains and add a schedule in Apify Console, then compare runs to see when a site changes its rules.

More tools from the same developer

All pay-per-result, no proxy or login needed, built and maintained by the same developer:

Website audits

Company data and compliance

Market signals

Feedback

A crawler missing from the list, or a result that looks wrong? Open an issue on the Issues tab with the domain. New AI crawlers are added when their vendors document them.