AI Crawler Access Checker (robots.txt & llms.txt) avatar

AI Crawler Access Checker (robots.txt & llms.txt)

Pricing

from $2.00 / 1,000 sites

Go to Apify Store
AI Crawler Access Checker (robots.txt & llms.txt)

AI Crawler Access Checker (robots.txt & llms.txt)

See which of 33 AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot...) may crawl any website, per path, with an RFC 9309 robots.txt matcher. Also checks llms.txt, llms-full.txt, ai.txt, noai meta tags and TDMRep.

Pricing

from $2.00 / 1,000 sites

Rating

0.0

(0)

Developer

Deepak Ganesh

Deepak Ganesh

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

AI Crawler Access Checker (robots.txt & llms.txt)

Find out which AI crawlers can crawl any website. For 33 AI and search bots (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, Google-Extended, PerplexityBot, CCBot, Bytespider, Applebot-Extended, Meta-ExternalAgent, Amazonbot and more) you get an allowed / partial / blocked verdict for each path you test, plus the exact group and robots.txt line behind it.

  • ✅ Correct robots.txt parsing (RFC 9309). A purpose-built parser follows Google's documented behaviour: the most specific user-agent group wins, groups for the same bot are merged, the longest rule wins and Allow wins ties, * and $ wildcards work, percent-encoding is normalized, and status codes are handled per spec (404 = all allowed, 5xx/429/timeout = disallow all).
  • 🧭 Tests any path, not just the homepage. Add /blog/ or /products/ and see per-path verdicts. If an input URL has a path, that path is tested too.
  • 🏷️ Classifies each bot by purpose. Every bot is tagged with its owner and purpose: training, AI search, user-triggered or classic search engine. The summary tells you at a glance whether a site blocks all AI training while staying visible in AI search.
  • 📄 Checks AI files and signals in one pass. You also get llms.txt (title, summary, link and section counts, spec check), llms-full.txt (exact size via a Range request, without downloading it), ai.txt, noai/noimageai in <meta name="robots"> and X-Robots-Tag, the W3C TDMRep opt-out (header, meta or /.well-known/tdmrep.json), Cloudflare Content-Signal lines, RSL License: URLs, sitemaps and crawl-delays.
  • 🔀 Follows redirects like a crawler. example.com → www.example.com is followed, and the host that crawlers actually see is the one checked.
  • ⚡ Fast and cheap. HTTP only, with no browser, and checks hundreds of sites per minute. Failed sites are free.

What makes it different

Other AI-crawler checkers on the Store look at a fixed list of 6–30 bots and usually only test /. This Actor:

This ActorTypical alternatives
Robots matcherFull RFC 9309 implementation: longest match, Allow on ties, */$, percent-encoding, group merging. Covered by Google's documented test casesOften a plain prefix or "Disallow: /" check
PathsAny number of paths per site, each with its own deciding ruleHomepage only
Custom botsAdd any user-agent tokenFixed list
5xx / 429 / timeoutsTreated as disallow all (Google behaviour) and flaggedOften reported as "allowed"
www/canonical redirectsChecks the host crawlers actually reachOften checks the bare domain
Extra AI opt-out signalsnoai/noimageai, TDMRep, ai.txt, Content-Signal, RSL LicenseRarely
llms-full.txt sizeExact size via HTTP Range, without downloading the fileNot checked or fully downloaded

Use cases

  • SEO / GEO agencies. Audit clients to confirm they are visible to ChatGPT search, Perplexity and Claude, and not accidentally blocking OAI-SearchBot when they only wanted to block GPTBot.
  • Publishers and site owners. Verify that AI training opt-outs (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot...) actually work on every section of the site.
  • AI companies and researchers. Check consent signals (robots.txt, TDMRep, noai, Content-Signal, RSL) across many domains before collecting data.
  • Monitoring. Schedule the Actor to catch robots.txt changes, such as a new block on your search bot or a 5xx robots.txt that silently stops all crawling.

Input

FieldTypeDefaultDescription
urlsarray of strings(required)Domains or URLs, e.g. nytimes.com or https://example.com/blog/post. Duplicates and URLs on the same host are merged into one result.
testPathsarray of strings["/"]Paths tested for every bot on every site. The input URL's own path is always added. Up to 25 paths per site.
userAgentsarray of strings[]Extra robots.txt tokens to check, such as MyCompanyBot. A full user-agent string works too, because the product token is extracted.
maxConcurrencyinteger10Websites checked in parallel.
requestTimeoutSecsinteger20Timeout for each request. Transient errors are retried.
{
"urls": ["nytimes.com", "https://www.anthropic.com", "docs.apify.com/platform"],
"testPaths": ["/", "/blog/"],
"userAgents": ["MyCompanyBot"]
}

Output

You get one dataset item per website. The dataset has these views: Overview, Per-crawler verdicts (one row per site × bot) and llms.txt & AI signals.

Example, trimmed from a real run (3 of the 33 bots shown):

{
"inputUrl": "https://www.nytimes.com/section/technology",
"url": "https://www.nytimes.com/",
"domain": "nytimes.com",
"success": true,
"robotsTxtUrl": "https://www.nytimes.com/robots.txt",
"robotsTxtStatus": 200,
"robotsTxtFound": true,
"robotsTxtOutcome": "parsed",
"robotsTxtSizeBytes": 8547,
"testedPaths": ["/", "/section/technology"],
"bots": [
{
"name": "GPTBot", "owner": "OpenAI", "purpose": "training", "userAgentToken": "GPTBot",
"allowed": false, "access": "blocked", "matchedGroup": "GPTBot", "explicitlyListed": true,
"matchedRule": "Disallow: / (line 233)",
"paths": [
{ "path": "/", "allowed": false, "matchedRule": "Disallow: / (line 233)" },
{ "path": "/section/technology", "allowed": false, "matchedRule": "Disallow: / (line 233)" }
],
"crawlDelay": null, "contentSignal": null, "note": null
},
{
"name": "OAI-SearchBot", "owner": "OpenAI", "purpose": "search", "userAgentToken": "OAI-SearchBot",
"allowed": false, "access": "blocked", "matchedGroup": "OAI-SearchBot", "explicitlyListed": true,
"matchedRule": "Disallow: / (line 266)", "note": "ChatGPT search results"
},
{
"name": "Googlebot", "owner": "Google", "purpose": "search-engine", "userAgentToken": "Googlebot",
"allowed": true, "access": "allowed", "matchedGroup": "Googlebot", "explicitlyListed": true,
"matchedRule": "No matching rule (allowed by default)", "note": "Google Search incl. AI Overviews"
}
],
"summary": {
"allowedCount": 9,
"partialCount": 0,
"blockedCount": 24,
"blocksAllAiTraining": false,
"blocksAnyAiTraining": true,
"blocksAiSearch": true,
"blocksAiUserFetchers": true,
"blockedBots": ["GPTBot", "OAI-SearchBot", "ChatGPT-User", "ClaudeBot", "Claude-SearchBot", "..."],
"partiallyBlockedBots": [],
"allowedBots": ["Googlebot", "Bingbot", "Applebot", "Amazonbot", "cohere-training-data-crawler", "MistralAI-User", "AI2Bot", "Webzio-Extended", "PanguBot"],
"byPurpose": {
"training": { "allowed": 5, "partial": 0, "blocked": 12 },
"search": { "allowed": 0, "partial": 0, "blocked": 5 },
"user-triggered": { "allowed": 1, "partial": 0, "blocked": 7 },
"search-engine": { "allowed": 3, "partial": 0, "blocked": 0 }
},
"verdict": "Blocks 12 of 17 AI training bots; 5 of 5 AI search bots blocked; 7 of 8 user-triggered AI fetchers blocked."
},
"contentSignals": [],
"llmsTxt": { "found": false, "url": "https://www.nytimes.com/llms.txt", "status": 404, "sizeBytes": null, "title": null, "linkCount": 0 },
"llmsFullTxt": { "found": false, "url": "https://www.nytimes.com/llms-full.txt", "status": 404, "sizeBytes": null, "sizeIsLowerBound": false },
"aiTxt": { "found": false, "url": "https://www.nytimes.com/ai.txt", "status": 404, "sizeBytes": null, "blocksAll": false },
"metaRobotsAi": { "noai": false, "noimageai": false, "noindex": false, "nofollow": false, "directives": [] },
"tdmRep": { "reserved": null, "source": null, "policyUrl": null, "wellKnownFound": false },
"homepageStatus": 403,
"sitemaps": ["https://www.nytimes.com/sitemaps/new/news.xml.gz", "https://www.nytimes.com/sitemaps/new/sitemap.xml.gz"],
"rslLicenses": [],
"crawlDelay": {},
"issues": ["Homepage returned HTTP 403; meta robots tags may be missing from the response."],
"checkedAt": "2026-10-07T16:30:26.278Z"
}

When a site has an llms.txt, the llmsTxt field looks like this (from docs.apify.com): { "found": true, "sizeBytes": 96450, "title": "Apify Documentation", "linkCount": 460, "sectionCount": 17, "startsWithH1": true }.

Key fields

  • access can be allowed (all tested paths), partial (some paths) or blocked (no tested path).
  • matchedGroup is the bot's own group, *, or null when no group applies.
  • robotsTxtOutcome is parsed, not-found-allow-all (4xx) or unreachable-disallow-all (5xx, 429 or a network error).
  • issues lists problems found: soft-404 HTML robots.txt or llms.txt, rules before any User-agent, unparseable lines, files over 500 KiB, redirects, noai set without matching robots.txt blocks, and more.
  • Failed sites (invalid input, domains that don't resolve, sites that can't be reached) have success: false and an error.

Pricing

This Actor uses pay-per-event pricing:

EventPrice
Site checked (one website, all bots and paths)$0.002 ($2 per 1,000 sites)
Actor start$0.00005

Failed items are free. You are not charged for invalid inputs, domains that don't resolve or sites that can't be reached. Testing extra paths or extra bots costs nothing more. Set a maximum cost per run and the Actor stops cleanly when it is reached.

FAQ

Why is a bot "blocked" when robots.txt never mentions it? The * group applies to it. Check matchedGroup (*) and matchedRule.

robots.txt returns 503. Why is everything blocked? RFC 9309 and Google treat an unreachable robots.txt (5xx, 429 or a timeout) as "disallow all" until it recovers. The Actor flags this in issues, because it often means bot protection or a misconfiguration is silently stopping crawlers.

robots.txt returns 403. Why is everything allowed? Per the spec, 4xx means "no robots.txt", so everything is allowed. A 401/403 is usually a firewall, so it is flagged in issues.

Does Google-Extended block Google Search or AI Overviews? No. Google-Extended and Applebot-Extended are control tokens for AI training. They are not separate crawlers, and blocking them doesn't affect Googlebot or Applebot. That's why they are classified as training.

Does the Actor obey robots.txt itself? It only fetches robots.txt, the homepage and the well-known AI files (/llms.txt, /llms-full.txt, /ai.txt, /.well-known/tdmrep.json), about 6 small requests per site. It never crawls content.

Can I check my own crawler? Yes. Add its token to userAgents.

This Actor is not affiliated with, endorsed by or sponsored by OpenAI, Anthropic, Google, Microsoft, Perplexity, Apple, Meta, Amazon, ByteDance, Common Crawl, Cloudflare or any other company whose crawler is listed. Bot names are used only to identify them.

Changelog

  • 0.1 (2026-10): First release. Covers 33 crawlers, an RFC 9309 matcher with per-path tests, llms.txt / llms-full.txt / ai.txt checks, noai meta and X-Robots-Tag, TDMRep, Content-Signal and RSL License.