# AI Crawler Access Audit: robots.txt, llms.txt, AI Opt-Out (`conserving_celerytop/ai-crawler-access-audit`) Actor

Check which AI crawlers each website allows or blocks: GPTBot, ClaudeBot, PerplexityBot, Google-Extended and 19 more, with the deciding robots.txt line. Plus llms.txt and AI training opt-out signals (TDMRep, noai). Bulk domains, change alerts, $2 per 1,000 domains. No login.

- **URL**: https://apify.com/conserving\_celerytop/ai-crawler-access-audit.md
- **Developed by:** [Don Mangu](https://apify.com/conserving_celerytop) (community)
- **Categories:** SEO tools, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.60 / 1,000 domain auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**AI Crawler Access Audit** checks which **AI crawlers** each website lets in: for GPTBot, ClaudeBot, PerplexityBot, Google-Extended and 19 more AI and search crawlers it reads the site's **robots.txt** and returns allowed, partial or blocked, with the line that decides it. It also checks **llms.txt** and the signals sites use to opt out of AI training. You pay **$2 per 1,000 domains**; unreachable sites, duplicates and unchanged domains in monitor mode are free.

### What does AI Crawler Access Audit do?

It checks, for each website you list, which AI crawlers the site lets in. For every crawler it reads the site's robots.txt the way the crawler should (the most specific group for that crawler, then the general group) and returns allowed, partial or blocked, with the exact line that decides it. It also looks for llms.txt and llms-full.txt, and for the signals sites use to opt out of AI training: a TDM reservation (the machine-readable opt-out of the EU text and data mining rules, from `/.well-known/tdmrep.json`, a header or a meta tag) and `noai` or `noimageai` in the robots meta tag or X-Robots-Tag header.

- **SEO and GEO agencies** audit client sites in bulk: is the site visible to ChatGPT search, Perplexity and Claude, or blocked by an old rule nobody remembers?
- **Publishers and legal teams** check that their AI training opt-out is in place on every domain they own, in both robots.txt and the TDM reservation.
- **Researchers and data teams** measure how many sites in a list block AI training, by crawler and by company.
- **AI agents** call it through the Apify API or Apify's MCP server.

**Try it now.** The form opens with three well-known sites. Click **Start**; it takes a few seconds and costs less than a cent.

### AI crawlers checked

| Company | AI training | AI search | Fetch for a user |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot, anthropic-ai | Claude-SearchBot | Claude-User |
| Perplexity | | PerplexityBot | Perplexity-User |
| Google | Google-Extended | | |
| Apple | Applebot-Extended | Applebot | |
| Meta | meta-externalagent | | meta-externalfetcher |
| Amazon | | Amazonbot | |
| DuckDuckGo | | DuckAssistBot | |
| Mistral | | | MistralAI-User |
| Others | CCBot (Common Crawl), Bytespider (ByteDance), cohere-ai, Diffbot | | |

Googlebot and Bingbot are included as a reference for classic search. Add any other user-agent token under **Extra crawlers**.

### How to audit robots.txt for AI crawlers, step by step

1. Paste your domains, website addresses or work emails into **Domains**, one per line (up to 5,000).
2. Optional: pick a policy under **Suggest robots.txt lines for this policy** to get the lines each site would need to add.
3. Click **Start**. A few hundred domains take a few minutes.
4. Open the **Overview** view for the verdict, AI training and AI search access and the blocked crawlers per domain, or export CSV, Excel or JSON.

#### Input

| Field | What it does |
|---|---|
| Domains | Domains, website addresses or work emails, up to 5,000 per run |
| Suggest robots.txt lines for this policy | None, search only, block all, or allow all |
| Check llms.txt and llms-full.txt | On by default |
| Check AI training opt-out signals | TDMRep file, header and meta tag, and noai directives. On by default |
| Only domains that changed, Monitor name | Monitor mode |
| Extra crawlers | More user-agent tokens to check |
| Sites at a time | 1 to 20, default 5 |

Example:

```json
{
  "domains": ["example.com", "https://www.example.org/blog", "press@example.net"],
  "targetPolicy": "search-only"
}
```

#### Verdicts

| Verdict | Meaning |
|---|---|
| open to AI | Every AI training and AI search crawler may read the site |
| AI search only | Training crawlers are blocked, search and user-request crawlers are allowed |
| AI training only | The reverse, which is rare and usually a mistake |
| closed to AI | Every AI crawler is blocked |
| mixed | Some crawlers of a kind are blocked and others are not |
| unknown | The site refused our request for robots.txt (HTTP 401, 403 or 429), so its rules could not be read. Not charged |

A crawler is **partial** when it may read the home page but some paths are kept from it, such as `/admin` or `/search`. For the verdict, partial counts as allowed, since most sites keep a few paths out of every crawler.

#### Suggested robots.txt lines

Pick a **target policy** (allow AI search and block AI training, block all AI crawlers, or allow all) and each row gets the robots.txt lines that site would need to add to reach it. Classic search crawlers are never touched.

### llms.txt and AI training opt-out checks

Besides robots.txt, each domain gets:

- **llms.txt** and **llms-full.txt**: whether the site has them, and the title and link count of llms.txt. A home page served at `/llms.txt` does not count.
- **TDM reservation**: the machine-readable opt-out of the EU text and data mining rules, read from `/.well-known/tdmrep.json`, the `tdm-reservation` header or the meta tag (`tdmReservation`, `tdmSource`, `tdmPolicy`).
- **noai and noimageai** in the robots meta tag or the `X-Robots-Tag` header (`noaiDirective`, `noimageaiDirective`).

Turn these checks off with **Check llms.txt and llms-full.txt** and **Check AI training opt-out signals** if you only need the robots.txt table.

### Input example

This is the input the form is filled in with when you open the Actor, as JSON. Paste it into the **JSON** tab of the input, or send it as the run input through the API. Fields you leave out keep their defaults.

```json
{
  "domains": [
    "nytimes.com",
    "stripe.com",
    "wikipedia.org"
  ]
}
```

### How much does it cost?

**$2 per 1,000 domains audited** ($0.002 per domain). Each domain gets at most 5 small requests (robots.txt, llms.txt, llms-full.txt, the TDMRep file and the home page). Apify adds its small standard fee per run start. Apify's free plan gives $5 of credit a month, which covers about 2,500 domains.

Example: you audit 10,000 client domains once and then watch them weekly. The first run costs $20. If 150 domains change in a week, that week costs $0.30.

### Change alerts

Turn on **Only domains that changed**, give the list a **Monitor name**, and schedule the Actor daily or weekly. Each run returns only the domains whose crawler rules, llms.txt or opt-out signals changed since the last run, with a `changes` list such as `GPTBot allowed -> blocked` or `llms.txt added`. Unchanged domains are free, so a quiet week costs only the run start.

### Output

One row per domain. The **Overview** view shows the main columns; the full row has one entry per crawler.

```json
{
  "domain": "news.example",
  "status": "ok",
  "verdict": "AI search only",
  "aiTrainingAccess": "blocked",
  "aiSearchAccess": "allowed",
  "blockedBots": ["GPTBot", "ClaudeBot", "anthropic-ai", "Google-Extended", "Applebot-Extended", "CCBot", "Bytespider", "meta-externalagent", "cohere-ai", "Diffbot"],
  "partialBots": ["Googlebot", "Bingbot"],
  "robotsTxtStatus": "found",
  "sitemaps": ["https://news.example/sitemap.xml"],
  "llmsTxtFound": true,
  "llmsTxtTitle": "Example Docs",
  "llmsTxtLinks": 3,
  "llmsFullTxtFound": false,
  "tdmReservation": true,
  "tdmSource": "tdmrep.json",
  "noaiDirective": true,
  "suggestedRobotsTxtLines": "",
  "bots": [
    { "bot": "GPTBot", "company": "OpenAI", "purpose": "training", "access": "blocked", "matchedGroup": "specific", "rule": "Disallow: /" },
    { "bot": "OAI-SearchBot", "company": "OpenAI", "purpose": "search", "access": "allowed", "matchedGroup": "specific", "rule": "Allow: /" }
  ],
  "changes": null,
  "checkedAt": "2026-09-27T10:00:00.000Z",
  "charged": true
}
```

`robotsTxtStatus` is `found`, `not_found` (no file, which means every crawler may read the site), `forbidden` (the site refused our request; verdict `unknown`, not charged), `server_error` (the standard treats this as a full block, and so does the audit) or `unreachable` (not charged). A STATS record in the key-value store counts domains audited, charged, unreachable and unchanged, and the verdicts.

### Use it from AI agents (MCP)

For AI agents: pass a list of domains; get one JSON row per domain with the access of 23 AI and search crawlers (allowed, partial or blocked, with the deciding robots.txt rule), a verdict, llms.txt presence, and AI training opt-out signals (TDMRep, noai). Set onlyChanges for change alerts.

The Actor works through the Apify API and Apify's MCP server, so Claude, Cursor and other agents can call it as a tool with a list of domains and read the rows.

- [Website Tech Stack Detector](https://apify.com/conserving_celerytop/website-tech-stack-detector): the CMS, analytics, frameworks and hosting behind each site, for the same domain list.

### Related Actors

- [Sitemap URL Extractor API](https://apify.com/conserving_celerytop/sitemap-url-extractor): Use it to list every page URL of a website from its XML sitemaps.
- [Broken Link Checker](https://apify.com/conserving_celerytop/broken-link-checker): Use it to find broken internal and external links on a website.
- [Web Page to Markdown for AI](https://apify.com/conserving_celerytop/web-page-to-markdown): Use it to turn web pages into clean Markdown for LLMs, RAG and AI agents.
- [Domain Authority Checker](https://apify.com/conserving_celerytop/domain-authority-checker): Use it to check domain authority for a bulk list of domains.

### FAQ

**Does it tell me whether a crawler really obeys robots.txt?** No. It reports what the site asks each crawler to do. It does not send requests as any AI crawler.

**Does it respect robots.txt itself?** Yes. It reads robots.txt, and reads llms.txt, the TDMRep file and the home page only where the site's robots.txt allows the token `DonMangu-AICrawlerAudit`. A site that blocks it still gets its crawler table from robots.txt.

**Is a TDM reservation the same as blocking crawlers?** No. robots.txt controls access; a TDM reservation states that the site reserves its text and data mining rights. Sites that want to opt out of AI training often use both, which is why the audit returns both.

**How current is the crawler list?** `botsListVersion` in each row shows the date of the list. For a crawler that is not on it yet, add its token under **Extra crawlers**.

**Can it check sites behind a login?** No. It reads only public files and the public home page.

*The crawler names are trademarks of their owners. This Actor is not affiliated with or endorsed by any of them.*

# Actor input Schema

## `domains` (type: `array`):

Domains or website addresses, one per line. Work emails such as info@acme.com also work. Up to 5,000 per run.

## `targetPolicy` (type: `string`):

Get the robots.txt lines each site would need to add to reach this policy. Classic search (Googlebot, Bingbot) always stays allowed.

## `checkLlmsTxt` (type: `boolean`):

Look for /llms.txt and /llms-full.txt and return the title and link count of llms.txt.

## `checkOptOuts` (type: `boolean`):

Read /.well-known/tdmrep.json (TDM reservation under the EU text and data mining rules), the tdm-reservation header and meta tag, and noai or noimageai in the robots meta tag and X-Robots-Tag header of the home page.

## `onlyChanges` (type: `boolean`):

Return only domains whose AI crawler rules, llms.txt or opt-out signals changed since the last run of this monitor. Unchanged domains are free. Schedule the Actor to get a change alert.

## `monitorName` (type: `string`):

Keeps separate memories for separate lists. Letters, digits and "-".

## `extraBots` (type: `array`):

More user-agent tokens to check, such as a new AI crawler. 23 AI and search crawlers are always checked.

## `maxConcurrency` (type: `integer`):

How many sites are checked in parallel (1 to 20). Each site gets at most 5 small requests.

## Actor input object example

```json
{
  "domains": [
    "cnn.com",
    "stripe.com",
    "wikipedia.org"
  ],
  "targetPolicy": "none",
  "checkLlmsTxt": true,
  "checkOptOuts": true,
  "onlyChanges": false,
  "monitorName": "default",
  "maxConcurrency": 5
}
```

# Actor output Schema

## `overview` (type: `string`):

Dataset items: verdict, AI training and AI search access, blocked bots, llms.txt and opt-out signals per domain.

## `stats` (type: `string`):

JSON record with domains audited and charged, unreachable and unchanged domains (free), verdict counts and errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "cnn.com",
        "stripe.com",
        "wikipedia.org"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("conserving_celerytop/ai-crawler-access-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "cnn.com",
        "stripe.com",
        "wikipedia.org",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("conserving_celerytop/ai-crawler-access-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "cnn.com",
    "stripe.com",
    "wikipedia.org"
  ]
}' |
apify call conserving_celerytop/ai-crawler-access-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,conserving_celerytop/ai-crawler-access-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/mx48HQfygmDyUBXaA/builds/otm0mZdKhef0aJBDc/openapi.json
