AI Crawler Access Audit: robots.txt, llms.txt, AI Opt-Out
Pricing
from $1.60 / 1,000 domain auditeds
AI Crawler Access Audit: robots.txt, llms.txt, AI Opt-Out
Check which AI crawlers each website allows or blocks: GPTBot, ClaudeBot, PerplexityBot, Google-Extended and 19 more, with the deciding robots.txt line. Plus llms.txt and AI training opt-out signals (TDMRep, noai). Bulk domains, change alerts, $2 per 1,000 domains. No login.
Pricing
from $1.60 / 1,000 domain auditeds
Rating
0.0
(0)
Developer
Don Mangu
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 hours ago
Last modified
Categories
Share
AI Crawler Access Audit checks which AI crawlers each website lets in: for GPTBot, ClaudeBot, PerplexityBot, Google-Extended and 19 more AI and search crawlers it reads the site's robots.txt and returns allowed, partial or blocked, with the line that decides it. It also checks llms.txt and the signals sites use to opt out of AI training. You pay $2 per 1,000 domains; unreachable sites, duplicates and unchanged domains in monitor mode are free.
What does AI Crawler Access Audit do?
It checks, for each website you list, which AI crawlers the site lets in. For every crawler it reads the site's robots.txt the way the crawler should (the most specific group for that crawler, then the general group) and returns allowed, partial or blocked, with the exact line that decides it. It also looks for llms.txt and llms-full.txt, and for the signals sites use to opt out of AI training: a TDM reservation (the machine-readable opt-out of the EU text and data mining rules, from /.well-known/tdmrep.json, a header or a meta tag) and noai or noimageai in the robots meta tag or X-Robots-Tag header.
- SEO and GEO agencies audit client sites in bulk: is the site visible to ChatGPT search, Perplexity and Claude, or blocked by an old rule nobody remembers?
- Publishers and legal teams check that their AI training opt-out is in place on every domain they own, in both robots.txt and the TDM reservation.
- Researchers and data teams measure how many sites in a list block AI training, by crawler and by company.
- AI agents call it through the Apify API or Apify's MCP server.
Try it now. The form opens with three well-known sites. Click Start; it takes a few seconds and costs less than a cent.
AI crawlers checked
| Company | AI training | AI search | Fetch for a user |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot, anthropic-ai | Claude-SearchBot | Claude-User |
| Perplexity | PerplexityBot | Perplexity-User | |
| Google-Extended | |||
| Apple | Applebot-Extended | Applebot | |
| Meta | meta-externalagent | meta-externalfetcher | |
| Amazon | Amazonbot | ||
| DuckDuckGo | DuckAssistBot | ||
| Mistral | MistralAI-User | ||
| Others | CCBot (Common Crawl), Bytespider (ByteDance), cohere-ai, Diffbot |
Googlebot and Bingbot are included as a reference for classic search. Add any other user-agent token under Extra crawlers.
How to audit robots.txt for AI crawlers, step by step
- Paste your domains, website addresses or work emails into Domains, one per line (up to 5,000).
- Optional: pick a policy under Suggest robots.txt lines for this policy to get the lines each site would need to add.
- Click Start. A few hundred domains take a few minutes.
- Open the Overview view for the verdict, AI training and AI search access and the blocked crawlers per domain, or export CSV, Excel or JSON.
Input
| Field | What it does |
|---|---|
| Domains | Domains, website addresses or work emails, up to 5,000 per run |
| Suggest robots.txt lines for this policy | None, search only, block all, or allow all |
| Check llms.txt and llms-full.txt | On by default |
| Check AI training opt-out signals | TDMRep file, header and meta tag, and noai directives. On by default |
| Only domains that changed, Monitor name | Monitor mode |
| Extra crawlers | More user-agent tokens to check |
| Sites at a time | 1 to 20, default 5 |
Example:
{"domains": ["example.com", "https://www.example.org/blog", "press@example.net"],"targetPolicy": "search-only"}
Verdicts
| Verdict | Meaning |
|---|---|
| open to AI | Every AI training and AI search crawler may read the site |
| AI search only | Training crawlers are blocked, search and user-request crawlers are allowed |
| AI training only | The reverse, which is rare and usually a mistake |
| closed to AI | Every AI crawler is blocked |
| mixed | Some crawlers of a kind are blocked and others are not |
| unknown | The site refused our request for robots.txt (HTTP 401, 403 or 429), so its rules could not be read. Not charged |
A crawler is partial when it may read the home page but some paths are kept from it, such as /admin or /search. For the verdict, partial counts as allowed, since most sites keep a few paths out of every crawler.
Suggested robots.txt lines
Pick a target policy (allow AI search and block AI training, block all AI crawlers, or allow all) and each row gets the robots.txt lines that site would need to add to reach it. Classic search crawlers are never touched.
llms.txt and AI training opt-out checks
Besides robots.txt, each domain gets:
- llms.txt and llms-full.txt: whether the site has them, and the title and link count of llms.txt. A home page served at
/llms.txtdoes not count. - TDM reservation: the machine-readable opt-out of the EU text and data mining rules, read from
/.well-known/tdmrep.json, thetdm-reservationheader or the meta tag (tdmReservation,tdmSource,tdmPolicy). - noai and noimageai in the robots meta tag or the
X-Robots-Tagheader (noaiDirective,noimageaiDirective).
Turn these checks off with Check llms.txt and llms-full.txt and Check AI training opt-out signals if you only need the robots.txt table.
Input example
This is the input the form is filled in with when you open the Actor, as JSON. Paste it into the JSON tab of the input, or send it as the run input through the API. Fields you leave out keep their defaults.
{"domains": ["nytimes.com","stripe.com","wikipedia.org"]}
How much does it cost?
$2 per 1,000 domains audited ($0.002 per domain). Each domain gets at most 5 small requests (robots.txt, llms.txt, llms-full.txt, the TDMRep file and the home page). Apify adds its small standard fee per run start. Apify's free plan gives $5 of credit a month, which covers about 2,500 domains.
Example: you audit 10,000 client domains once and then watch them weekly. The first run costs $20. If 150 domains change in a week, that week costs $0.30.
Change alerts
Turn on Only domains that changed, give the list a Monitor name, and schedule the Actor daily or weekly. Each run returns only the domains whose crawler rules, llms.txt or opt-out signals changed since the last run, with a changes list such as GPTBot allowed -> blocked or llms.txt added. Unchanged domains are free, so a quiet week costs only the run start.
Output
One row per domain. The Overview view shows the main columns; the full row has one entry per crawler.
{"domain": "news.example","status": "ok","verdict": "AI search only","aiTrainingAccess": "blocked","aiSearchAccess": "allowed","blockedBots": ["GPTBot", "ClaudeBot", "anthropic-ai", "Google-Extended", "Applebot-Extended", "CCBot", "Bytespider", "meta-externalagent", "cohere-ai", "Diffbot"],"partialBots": ["Googlebot", "Bingbot"],"robotsTxtStatus": "found","sitemaps": ["https://news.example/sitemap.xml"],"llmsTxtFound": true,"llmsTxtTitle": "Example Docs","llmsTxtLinks": 3,"llmsFullTxtFound": false,"tdmReservation": true,"tdmSource": "tdmrep.json","noaiDirective": true,"suggestedRobotsTxtLines": "","bots": [{ "bot": "GPTBot", "company": "OpenAI", "purpose": "training", "access": "blocked", "matchedGroup": "specific", "rule": "Disallow: /" },{ "bot": "OAI-SearchBot", "company": "OpenAI", "purpose": "search", "access": "allowed", "matchedGroup": "specific", "rule": "Allow: /" }],"changes": null,"checkedAt": "2026-09-27T10:00:00.000Z","charged": true}
robotsTxtStatus is found, not_found (no file, which means every crawler may read the site), forbidden (the site refused our request; verdict unknown, not charged), server_error (the standard treats this as a full block, and so does the audit) or unreachable (not charged). A STATS record in the key-value store counts domains audited, charged, unreachable and unchanged, and the verdicts.
Use it from AI agents (MCP)
For AI agents: pass a list of domains; get one JSON row per domain with the access of 23 AI and search crawlers (allowed, partial or blocked, with the deciding robots.txt rule), a verdict, llms.txt presence, and AI training opt-out signals (TDMRep, noai). Set onlyChanges for change alerts.
The Actor works through the Apify API and Apify's MCP server, so Claude, Cursor and other agents can call it as a tool with a list of domains and read the rows.
- Website Tech Stack Detector: the CMS, analytics, frameworks and hosting behind each site, for the same domain list.
Related Actors
- Sitemap URL Extractor API: Use it to list every page URL of a website from its XML sitemaps.
- Broken Link Checker: Use it to find broken internal and external links on a website.
- Web Page to Markdown for AI: Use it to turn web pages into clean Markdown for LLMs, RAG and AI agents.
- Domain Authority Checker: Use it to check domain authority for a bulk list of domains.
FAQ
Does it tell me whether a crawler really obeys robots.txt? No. It reports what the site asks each crawler to do. It does not send requests as any AI crawler.
Does it respect robots.txt itself? Yes. It reads robots.txt, and reads llms.txt, the TDMRep file and the home page only where the site's robots.txt allows the token DonMangu-AICrawlerAudit. A site that blocks it still gets its crawler table from robots.txt.
Is a TDM reservation the same as blocking crawlers? No. robots.txt controls access; a TDM reservation states that the site reserves its text and data mining rights. Sites that want to opt out of AI training often use both, which is why the audit returns both.
How current is the crawler list? botsListVersion in each row shows the date of the list. For a crawler that is not on it yet, add its token under Extra crawlers.
Can it check sites behind a login? No. It reads only public files and the public home page.
The crawler names are trademarks of their owners. This Actor is not affiliated with or endorsed by any of them.