AI Crawler Access Audit: robots.txt, llms.txt, AI Opt-Out avatar

AI Crawler Access Audit: robots.txt, llms.txt, AI Opt-Out

Pricing

from $1.60 / 1,000 domain auditeds

Go to Apify Store
AI Crawler Access Audit: robots.txt, llms.txt, AI Opt-Out

AI Crawler Access Audit: robots.txt, llms.txt, AI Opt-Out

Check which AI crawlers each website allows or blocks: GPTBot, ClaudeBot, PerplexityBot, Google-Extended and 19 more, with the deciding robots.txt line. Plus llms.txt and AI training opt-out signals (TDMRep, noai). Bulk domains, change alerts, $2 per 1,000 domains. No login.

Pricing

from $1.60 / 1,000 domain auditeds

Rating

0.0

(0)

Developer

Don Mangu

Don Mangu

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 hours ago

Last modified

Share

AI Crawler Access Audit checks which AI crawlers each website lets in: for GPTBot, ClaudeBot, PerplexityBot, Google-Extended and 19 more AI and search crawlers it reads the site's robots.txt and returns allowed, partial or blocked, with the line that decides it. It also checks llms.txt and the signals sites use to opt out of AI training. You pay $2 per 1,000 domains; unreachable sites, duplicates and unchanged domains in monitor mode are free.

What does AI Crawler Access Audit do?

It checks, for each website you list, which AI crawlers the site lets in. For every crawler it reads the site's robots.txt the way the crawler should (the most specific group for that crawler, then the general group) and returns allowed, partial or blocked, with the exact line that decides it. It also looks for llms.txt and llms-full.txt, and for the signals sites use to opt out of AI training: a TDM reservation (the machine-readable opt-out of the EU text and data mining rules, from /.well-known/tdmrep.json, a header or a meta tag) and noai or noimageai in the robots meta tag or X-Robots-Tag header.

  • SEO and GEO agencies audit client sites in bulk: is the site visible to ChatGPT search, Perplexity and Claude, or blocked by an old rule nobody remembers?
  • Publishers and legal teams check that their AI training opt-out is in place on every domain they own, in both robots.txt and the TDM reservation.
  • Researchers and data teams measure how many sites in a list block AI training, by crawler and by company.
  • AI agents call it through the Apify API or Apify's MCP server.

Try it now. The form opens with three well-known sites. Click Start; it takes a few seconds and costs less than a cent.

AI crawlers checked

CompanyAI trainingAI searchFetch for a user
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBot, anthropic-aiClaude-SearchBotClaude-User
PerplexityPerplexityBotPerplexity-User
GoogleGoogle-Extended
AppleApplebot-ExtendedApplebot
Metameta-externalagentmeta-externalfetcher
AmazonAmazonbot
DuckDuckGoDuckAssistBot
MistralMistralAI-User
OthersCCBot (Common Crawl), Bytespider (ByteDance), cohere-ai, Diffbot

Googlebot and Bingbot are included as a reference for classic search. Add any other user-agent token under Extra crawlers.

How to audit robots.txt for AI crawlers, step by step

  1. Paste your domains, website addresses or work emails into Domains, one per line (up to 5,000).
  2. Optional: pick a policy under Suggest robots.txt lines for this policy to get the lines each site would need to add.
  3. Click Start. A few hundred domains take a few minutes.
  4. Open the Overview view for the verdict, AI training and AI search access and the blocked crawlers per domain, or export CSV, Excel or JSON.

Input

FieldWhat it does
DomainsDomains, website addresses or work emails, up to 5,000 per run
Suggest robots.txt lines for this policyNone, search only, block all, or allow all
Check llms.txt and llms-full.txtOn by default
Check AI training opt-out signalsTDMRep file, header and meta tag, and noai directives. On by default
Only domains that changed, Monitor nameMonitor mode
Extra crawlersMore user-agent tokens to check
Sites at a time1 to 20, default 5

Example:

{
"domains": ["example.com", "https://www.example.org/blog", "press@example.net"],
"targetPolicy": "search-only"
}

Verdicts

VerdictMeaning
open to AIEvery AI training and AI search crawler may read the site
AI search onlyTraining crawlers are blocked, search and user-request crawlers are allowed
AI training onlyThe reverse, which is rare and usually a mistake
closed to AIEvery AI crawler is blocked
mixedSome crawlers of a kind are blocked and others are not
unknownThe site refused our request for robots.txt (HTTP 401, 403 or 429), so its rules could not be read. Not charged

A crawler is partial when it may read the home page but some paths are kept from it, such as /admin or /search. For the verdict, partial counts as allowed, since most sites keep a few paths out of every crawler.

Suggested robots.txt lines

Pick a target policy (allow AI search and block AI training, block all AI crawlers, or allow all) and each row gets the robots.txt lines that site would need to add to reach it. Classic search crawlers are never touched.

llms.txt and AI training opt-out checks

Besides robots.txt, each domain gets:

  • llms.txt and llms-full.txt: whether the site has them, and the title and link count of llms.txt. A home page served at /llms.txt does not count.
  • TDM reservation: the machine-readable opt-out of the EU text and data mining rules, read from /.well-known/tdmrep.json, the tdm-reservation header or the meta tag (tdmReservation, tdmSource, tdmPolicy).
  • noai and noimageai in the robots meta tag or the X-Robots-Tag header (noaiDirective, noimageaiDirective).

Turn these checks off with Check llms.txt and llms-full.txt and Check AI training opt-out signals if you only need the robots.txt table.

Input example

This is the input the form is filled in with when you open the Actor, as JSON. Paste it into the JSON tab of the input, or send it as the run input through the API. Fields you leave out keep their defaults.

{
"domains": [
"nytimes.com",
"stripe.com",
"wikipedia.org"
]
}

How much does it cost?

$2 per 1,000 domains audited ($0.002 per domain). Each domain gets at most 5 small requests (robots.txt, llms.txt, llms-full.txt, the TDMRep file and the home page). Apify adds its small standard fee per run start. Apify's free plan gives $5 of credit a month, which covers about 2,500 domains.

Example: you audit 10,000 client domains once and then watch them weekly. The first run costs $20. If 150 domains change in a week, that week costs $0.30.

Change alerts

Turn on Only domains that changed, give the list a Monitor name, and schedule the Actor daily or weekly. Each run returns only the domains whose crawler rules, llms.txt or opt-out signals changed since the last run, with a changes list such as GPTBot allowed -> blocked or llms.txt added. Unchanged domains are free, so a quiet week costs only the run start.

Output

One row per domain. The Overview view shows the main columns; the full row has one entry per crawler.

{
"domain": "news.example",
"status": "ok",
"verdict": "AI search only",
"aiTrainingAccess": "blocked",
"aiSearchAccess": "allowed",
"blockedBots": ["GPTBot", "ClaudeBot", "anthropic-ai", "Google-Extended", "Applebot-Extended", "CCBot", "Bytespider", "meta-externalagent", "cohere-ai", "Diffbot"],
"partialBots": ["Googlebot", "Bingbot"],
"robotsTxtStatus": "found",
"sitemaps": ["https://news.example/sitemap.xml"],
"llmsTxtFound": true,
"llmsTxtTitle": "Example Docs",
"llmsTxtLinks": 3,
"llmsFullTxtFound": false,
"tdmReservation": true,
"tdmSource": "tdmrep.json",
"noaiDirective": true,
"suggestedRobotsTxtLines": "",
"bots": [
{ "bot": "GPTBot", "company": "OpenAI", "purpose": "training", "access": "blocked", "matchedGroup": "specific", "rule": "Disallow: /" },
{ "bot": "OAI-SearchBot", "company": "OpenAI", "purpose": "search", "access": "allowed", "matchedGroup": "specific", "rule": "Allow: /" }
],
"changes": null,
"checkedAt": "2026-09-27T10:00:00.000Z",
"charged": true
}

robotsTxtStatus is found, not_found (no file, which means every crawler may read the site), forbidden (the site refused our request; verdict unknown, not charged), server_error (the standard treats this as a full block, and so does the audit) or unreachable (not charged). A STATS record in the key-value store counts domains audited, charged, unreachable and unchanged, and the verdicts.

Use it from AI agents (MCP)

For AI agents: pass a list of domains; get one JSON row per domain with the access of 23 AI and search crawlers (allowed, partial or blocked, with the deciding robots.txt rule), a verdict, llms.txt presence, and AI training opt-out signals (TDMRep, noai). Set onlyChanges for change alerts.

The Actor works through the Apify API and Apify's MCP server, so Claude, Cursor and other agents can call it as a tool with a list of domains and read the rows.

FAQ

Does it tell me whether a crawler really obeys robots.txt? No. It reports what the site asks each crawler to do. It does not send requests as any AI crawler.

Does it respect robots.txt itself? Yes. It reads robots.txt, and reads llms.txt, the TDMRep file and the home page only where the site's robots.txt allows the token DonMangu-AICrawlerAudit. A site that blocks it still gets its crawler table from robots.txt.

Is a TDM reservation the same as blocking crawlers? No. robots.txt controls access; a TDM reservation states that the site reserves its text and data mining rights. Sites that want to opt out of AI training often use both, which is why the audit returns both.

How current is the crawler list? botsListVersion in each row shows the date of the list. For a crawler that is not on it yet, add its token under Extra crawlers.

Can it check sites behind a login? No. It reads only public files and the public home page.

The crawler names are trademarks of their owners. This Actor is not affiliated with or endorsed by any of them.