AI Crawler Access Checker: robots.txt & llms.txt
Pricing
from $2.00 / 1,000 domain checkeds
AI Crawler Access Checker: robots.txt & llms.txt
Check which AI crawlers each domain allows in robots.txt: GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot and 20 more, each with the exact rule that matched. Adds a one-line AI visibility verdict (open to AI search, blocks training β¦), llms.txt and llms-full.txt checks and sitemap URLs.
Pricing
from $2.00 / 1,000 domain checkeds
Rating
0.0
(0)
Developer
Offera Studio
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
9 hours ago
Last modified
Categories
Share
What does AI Crawler Access Checker do?
AI Crawler Access Checker tells you, for any list of domains (up to 5,000 per run), which AI crawlers the site lets in and whether it publishes an llms.txt file. For every domain you get:
- π€ 26 AI crawlers checked against robots.txt: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, CCBot, meta-externalagent, Amazonbot, DuckAssistBot, Mistral AI's crawlers, Bytespider, cohere-ai, Diffbot and more
- β for each one: allowed, partial or blocked, the group and rule that matched (with its line number) and a plain-English reason
- π§ a one-line AI visibility verdict: open to all AI crawlers, open to AI search, blocks AI training, blocks some AI search crawlers, blocks all AI crawlers β¦
- π /llms.txt and /llms-full.txt: found or not, H1 title, summary, link sections, size and format issues
- πΊοΈ sitemap URLs from robots.txt (or
/sitemap.xmlwhen none are listed) - π Googlebot and Bingbot access for context
Paste your domains, click Start, and export the results to CSV, Excel or JSON, or use them through the Apify API, Google Sheets, Make, Zapier or n8n.
Who is this AI crawler checker for?
- SEO and GEO (generative engine optimisation) teams: check that ChatGPT search, Perplexity and Claude can actually read your pages, and that a CDN or plugin didn't block them.
- Publishers and content owners: verify that AI training crawlers are blocked while AI search crawlers stay allowed, across all your sites and subdomains.
- Agencies: audit clients' and prospects' AI visibility and llms.txt in bulk.
- Researchers and journalists: measure how many sites in a sector block GPTBot, ClaudeBot or CCBot, and track it over time with a schedule.
- Developers of AI tools: check before crawling whether a site allows your agent (add your own user agent token).
Training, search and user-triggered crawlers
AI companies now run several crawlers with different jobs, and a site can allow one and block another:
| Purpose | What it means | Examples |
|---|---|---|
| AI training | Collects content to train or improve AI models | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, meta-externalagent, Amazonbot, MistralAI-Training |
| AI search | Indexes pages so AI assistants can find and cite them in answers | OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot, meta-webindexer, Amzn-SearchBot, MistralAI-Index |
| User-triggered | Fetches a page when a user asks an assistant about it | ChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher, Amzn-User, MistralAI-User |
| Other | AI-related crawling of another kind | Diffbot, Google-CloudVertexBot |
Every token was checked against the vendor's own documentation on 30 September 2026, and each result links to it (docs). Bytespider, anthropic-ai and cohere-ai are not documented by their vendors today but are so common in robots.txt files that they are reported too (documented: false); they don't change the verdict. Some vendors say their user-triggered fetchers may not follow robots.txt; that is shown in note.
How robots.txt is read
The Actor follows RFC 9309, the Robots Exclusion Protocol standard that Google, OpenAI, Anthropic and others follow:
- User-agent lines are matched case-insensitively; several groups for the same crawler are combined; a crawler without its own group follows
User-agent: *; with no matching group, nothing is restricted. - The longest matching rule wins, and Allow wins a tie.
*matches anything and$anchors the end, soDisallow: /*.pdf$works as crawlers read it. - A missing robots.txt (404) means no restrictions. Only the first 500 KiB are read, as the standard allows.
What the three results mean:
| Result | Meaning |
|---|---|
| blocked | The home page is disallowed and no Allow rule opens anything. |
| partial | Only some paths are open (for example Disallow: / with Allow: /blog/), or the site wrote rules for this crawler that close some paths. |
| allowed | Everything else. General User-agent: * rules that close a few paths for every crawler (like /admin/) don't count as a restriction on AI crawlers; they are listed in disallowedPaths. |
How to check AI crawler access
- Click Try for free and sign in to Apify.
- Paste domains into Domains, one per line (up to 5,000 per run).
www.example.comandexample.comare checked separately because they can have different robots.txt files. - Optional: add your own crawler tokens under Extra user agents to check, or tick Include the robots.txt text.
- Click Start. The Overview tab shows one row per domain; AI crawlers shows one row per crawler with the rule that matched; llms.txt and Sitemaps have their own tabs.
Input example
{"domains": ["example.com", "docs.example.com", "https://www.example.org/blog"],"checkLlmsTxt": true,"additionalUserAgents": ["YourBot"],"includeRobotsTxt": false}
Output example
One item per domain (shortened, made-up data):
{"domain": "news.example","robotsTxtUrl": "https://news.example/robots.txt","robotsTxtStatus": "found","aiVisibility": "blocks-training-only","aiVisibilityLabel": "Open to AI search, blocks AI training","aiVisibilitySummary": "AI training: 8 of 8 blocked (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and 4 more); AI search: 0 of 7 blocked; user-triggered fetchers: 0 of 6 blocked.","trainingBlockedCount": 8,"searchBlockedCount": 0,"blockedAgents": ["GPTBot", "ClaudeBot", "Google-Extended", "Applebot-Extended", "CCBot", "meta-externalagent", "Amazonbot", "MistralAI-Training", "Bytespider", "anthropic-ai", "cohere-ai"],"partialAgents": [],"blocksAllCrawlers": false,"googlebotAccess": "allowed","bingbotAccess": "allowed","agents": [{"agent": "GPTBot","vendor": "OpenAI","purpose": "training","documented": true,"access": "blocked","group": "User-agent: GPTBot","rule": "Disallow: /","ruleLine": 16,"reason": "Blocked by \"Disallow: /\" (line 16) in the \"User-agent: GPTBot\" group.","docs": "https://platform.openai.com/docs/bots"},{"agent": "OAI-SearchBot","vendor": "OpenAI","purpose": "search","access": "allowed","group": "User-agent: OAI-SearchBot","rule": "Allow: /","reason": "Allowed: the \"User-agent: OAI-SearchBot\" group closes no paths.","note": "OpenAI says sites that block it are not shown in ChatGPT search answers."}],"hasLlmsTxt": true,"hasLlmsFullTxt": false,"llmsTxt": {"url": "https://news.example/llms.txt","found": true,"valid": true,"title": "News Example","summary": "Independent local news since 1998.","sectionsCount": 3,"linksCount": 24,"hasOptionalSection": true,"sizeBytes": 3120,"issues": []},"sitemaps": ["https://news.example/sitemap.xml", "https://news.example/sitemap-news.xml"],"sitemapSource": "robots.txt","error": null}
Domains that can't be checked get a row with an error code (invalid-domain, dns-not-found, connection-failed, timeout, tls-error, robots-txt-forbidden, robots-txt-rate-limited, robots-txt-server-error) and cost nothing.
What is llms.txt?
llms.txt is a proposed standard: a Markdown file at /llms.txt that gives AI assistants a short, curated map of a site. It must start with an H1 title; it should have a blockquote summary, and H2 sections with lists of links (- [name](url): notes). An ## Optional section holds links that can be skipped. /llms-full.txt is a common companion with the full content in one file. The Actor checks:
- presence (an HTML "not found" page served at the address doesn't count),
- the H1 title, summary, number of sections and links, and an Optional section,
- size, and plain-English
issuessuch as a missing title or sections without links.
How much does it cost?
This Actor uses pay per event:
| Event | Price |
|---|---|
| Domain checked | $0.002 per domain |
| Domain that doesn't resolve, can't be reached or refuses the request | free |
- 1,000 domains cost $2; 5,000 domains cost $10.
- Checking llms.txt and llms-full.txt is included.
- Apify also charges a tiny standard start fee per run (about $0.000025 at the default 512 MB).
- Apify's free plan includes $5 of monthly usage, enough for about 2,500 domains a month.
- Set Maximum cost per run in the run options and the Actor stops when it is reached.
Limitations
- robots.txt is a request, not a lock. It shows what a site asks crawlers to do. Some crawlers may ignore it, and some vendors say their user-triggered fetchers may not follow it (see
note). Sites can also block AI crawlers at the firewall or CDN, which robots.txt can't show. - Google-Extended is not about Google Search. It controls Gemini training and grounding. Google Search, including its AI features, uses Googlebot and Search's own controls such as
nosnippet, which robots.txt tokens for AI crawlers don't change. - One hostname per row.
example.comandwww.example.comcan have different rules; enter the ones you care about. - Politeness: the Actor makes at most four small requests per domain (robots.txt, llms.txt, llms-full.txt and, if robots.txt lists no sitemap, /sitemap.xml), half a second apart, and only fetches files that the site's robots.txt allows for crawlers. llms-full.txt is read up to 256 KB; its size comes from the server's Content-Length.
- A 401, 403 or 429 on robots.txt usually means the server blocks automated requests; such domains get a free error row instead of a guess.
- Not legal advice: whether AI companies may use content is a legal question that robots.txt alone doesn't answer.
FAQ
Does blocking GPTBot keep my site out of ChatGPT?
Not entirely. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and ChatGPT fetches pages that users ask about with ChatGPT-User. The verdict shows each group separately, so you can see, for example, "Open to AI search, blocks AI training".
Why is a crawler "partial" when I blocked it completely?
Check reason and rule: usually another rule opens some paths again, for example Allow: /blog/ next to Disallow: /, or the crawler matches a group with only some paths disallowed. The line number points to the rule in your robots.txt.
Why does a domain return robots-txt-forbidden?
The server answered 401 or 403 to our request for robots.txt, which usually means a firewall blocks automated requests. Its rules can't be read, so the row is free.
Can I check my own crawler?
Yes. Add its robots.txt token under Extra user agents to check. It appears in agents with purpose custom.
Can I run it on a schedule?
Yes. Save a task with your domains and add a schedule in Apify Console, then compare runs to see when a site changes its rules.
More tools from the same developer
All pay-per-result, no proxy or login needed, built and maintained by the same developer:
Website audits
- Website Accessibility Checker: WCAG 2.2 & EAA: accessibility issues with fixes, SEO basics and security headers.
- Cookie & Tracker Audit: GDPR Consent Checker: cookies and tracking tags that load before consent.
- Website Change Monitor: Diffs, Prices & Alerts: get a row only when a page changes, with a clean diff.
Company data and compliance
- Company Contact Finder: Emails, Phones & Socials: contact details published on company websites.
- UK New Companies Feed: Companies House Daily: newly incorporated UK companies with sector filters.
- EU VAT Number Validator: Bulk VIES Checker: bulk VAT checks with name, address and consultation number.
- LEI Corporate Tree: GLEIF Parents & Subsidiaries: LEI lookup with parents, subsidiaries and a KYC summary.
Market signals
- US WARN Layoff Notices: 12 States Daily Feed: layoff and plant closure notices from official state sources.
- US Product Recalls Monitor: FDA & CPSC Feed: FDA and CPSC recalls in one feed, with severity.
Feedback
A crawler missing from the list, or a result that looks wrong? Open an issue on the Issues tab with the domain. New AI crawlers are added when their vendors document them.