# AI Crawler Access Checker: robots.txt & llms.txt (`offerastudio/ai-crawler-access-audit`) Actor

Check which AI crawlers each domain allows in robots.txt: GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot and 20 more, each with the exact rule that matched. Adds a one-line AI visibility verdict (open to AI search, blocks training …), llms.txt and llms-full.txt checks and sitemap URLs.

- **URL**: https://apify.com/offerastudio/ai-crawler-access-audit.md
- **Developed by:** [Offera Studio](https://apify.com/offerastudio) (community)
- **Categories:** AI, SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 domain checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does AI Crawler Access Checker do?

**AI Crawler Access Checker** tells you, for any list of domains (up to 5,000 per run), **which AI crawlers the site lets in** and whether it publishes an **llms.txt** file. For every domain you get:

- 🤖 **26 AI crawlers checked against robots.txt**: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, CCBot, meta-externalagent, Amazonbot, DuckAssistBot, Mistral AI's crawlers, Bytespider, cohere-ai, Diffbot and more
- ✅ for each one: **allowed, partial or blocked**, the **group and rule that matched** (with its line number) and a **plain-English reason**
- 🧭 a one-line **AI visibility verdict**: *open to all AI crawlers*, *open to AI search, blocks AI training*, *blocks some AI search crawlers*, *blocks all AI crawlers* …
- 📄 **/llms.txt and /llms-full.txt**: found or not, H1 title, summary, link sections, size and format issues
- 🗺️ **sitemap URLs** from robots.txt (or `/sitemap.xml` when none are listed)
- 🔎 **Googlebot and Bingbot** access for context

Paste your domains, click **Start**, and export the results to CSV, Excel or JSON, or use them through the Apify API, Google Sheets, Make, Zapier or n8n.

### Who is this AI crawler checker for?

- **SEO and GEO (generative engine optimisation) teams**: check that ChatGPT search, Perplexity and Claude can actually read your pages, and that a CDN or plugin didn't block them.
- **Publishers and content owners**: verify that AI training crawlers are blocked while AI search crawlers stay allowed, across all your sites and subdomains.
- **Agencies**: audit clients' and prospects' AI visibility and llms.txt in bulk.
- **Researchers and journalists**: measure how many sites in a sector block GPTBot, ClaudeBot or CCBot, and track it over time with a schedule.
- **Developers of AI tools**: check before crawling whether a site allows your agent (add your own user agent token).

### Training, search and user-triggered crawlers

AI companies now run several crawlers with different jobs, and a site can allow one and block another:

| Purpose | What it means | Examples |
| --- | --- | --- |
| **AI training** | Collects content to train or improve AI models | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, meta-externalagent, Amazonbot, MistralAI-Training |
| **AI search** | Indexes pages so AI assistants can find and cite them in answers | OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot, meta-webindexer, Amzn-SearchBot, MistralAI-Index |
| **User-triggered** | Fetches a page when a user asks an assistant about it | ChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher, Amzn-User, MistralAI-User |
| **Other** | AI-related crawling of another kind | Diffbot, Google-CloudVertexBot |

Every token was checked against the vendor's own documentation on 30 September 2026, and each result links to it (`docs`). Bytespider, anthropic-ai and cohere-ai are not documented by their vendors today but are so common in robots.txt files that they are reported too (`documented: false`); they don't change the verdict. Some vendors say their user-triggered fetchers may not follow robots.txt; that is shown in `note`.

### How robots.txt is read

The Actor follows **RFC 9309**, the Robots Exclusion Protocol standard that Google, OpenAI, Anthropic and others follow:

- User-agent lines are matched case-insensitively; several groups for the same crawler are combined; a crawler without its own group follows `User-agent: *`; with no matching group, nothing is restricted.
- The **longest matching rule wins**, and **Allow wins a tie**. `*` matches anything and `$` anchors the end, so `Disallow: /*.pdf$` works as crawlers read it.
- A missing robots.txt (404) means no restrictions. Only the first 500 KiB are read, as the standard allows.

**What the three results mean:**

| Result | Meaning |
| --- | --- |
| **blocked** | The home page is disallowed and no Allow rule opens anything. |
| **partial** | Only some paths are open (for example `Disallow: /` with `Allow: /blog/`), or the site wrote rules for this crawler that close some paths. |
| **allowed** | Everything else. General `User-agent: *` rules that close a few paths for every crawler (like `/admin/`) don't count as a restriction on AI crawlers; they are listed in `disallowedPaths`. |

### How to check AI crawler access

1. Click **Try for free** and sign in to Apify.
2. Paste domains into **Domains**, one per line (up to 5,000 per run). `www.example.com` and `example.com` are checked separately because they can have different robots.txt files.
3. Optional: add your own crawler tokens under **Extra user agents to check**, or tick **Include the robots.txt text**.
4. Click **Start**. The **Overview** tab shows one row per domain; **AI crawlers** shows one row per crawler with the rule that matched; **llms.txt** and **Sitemaps** have their own tabs.

### Input example

```json
{
    "domains": ["example.com", "docs.example.com", "https://www.example.org/blog"],
    "checkLlmsTxt": true,
    "additionalUserAgents": ["YourBot"],
    "includeRobotsTxt": false
}
```

### Output example

One item per domain (shortened, made-up data):

```json
{
    "domain": "news.example",
    "robotsTxtUrl": "https://news.example/robots.txt",
    "robotsTxtStatus": "found",
    "aiVisibility": "blocks-training-only",
    "aiVisibilityLabel": "Open to AI search, blocks AI training",
    "aiVisibilitySummary": "AI training: 8 of 8 blocked (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and 4 more); AI search: 0 of 7 blocked; user-triggered fetchers: 0 of 6 blocked.",
    "trainingBlockedCount": 8,
    "searchBlockedCount": 0,
    "blockedAgents": ["GPTBot", "ClaudeBot", "Google-Extended", "Applebot-Extended", "CCBot", "meta-externalagent", "Amazonbot", "MistralAI-Training", "Bytespider", "anthropic-ai", "cohere-ai"],
    "partialAgents": [],
    "blocksAllCrawlers": false,
    "googlebotAccess": "allowed",
    "bingbotAccess": "allowed",
    "agents": [
        {
            "agent": "GPTBot",
            "vendor": "OpenAI",
            "purpose": "training",
            "documented": true,
            "access": "blocked",
            "group": "User-agent: GPTBot",
            "rule": "Disallow: /",
            "ruleLine": 16,
            "reason": "Blocked by \"Disallow: /\" (line 16) in the \"User-agent: GPTBot\" group.",
            "docs": "https://platform.openai.com/docs/bots"
        },
        {
            "agent": "OAI-SearchBot",
            "vendor": "OpenAI",
            "purpose": "search",
            "access": "allowed",
            "group": "User-agent: OAI-SearchBot",
            "rule": "Allow: /",
            "reason": "Allowed: the \"User-agent: OAI-SearchBot\" group closes no paths.",
            "note": "OpenAI says sites that block it are not shown in ChatGPT search answers."
        }
    ],
    "hasLlmsTxt": true,
    "hasLlmsFullTxt": false,
    "llmsTxt": {
        "url": "https://news.example/llms.txt",
        "found": true,
        "valid": true,
        "title": "News Example",
        "summary": "Independent local news since 1998.",
        "sectionsCount": 3,
        "linksCount": 24,
        "hasOptionalSection": true,
        "sizeBytes": 3120,
        "issues": []
    },
    "sitemaps": ["https://news.example/sitemap.xml", "https://news.example/sitemap-news.xml"],
    "sitemapSource": "robots.txt",
    "error": null
}
```

Domains that can't be checked get a row with an `error` code (`invalid-domain`, `dns-not-found`, `connection-failed`, `timeout`, `tls-error`, `robots-txt-forbidden`, `robots-txt-rate-limited`, `robots-txt-server-error`) and cost nothing.

### What is llms.txt?

[llms.txt](https://llmstxt.org) is a proposed standard: a Markdown file at `/llms.txt` that gives AI assistants a short, curated map of a site. It must start with an **H1 title**; it should have a **blockquote summary**, and **H2 sections with lists of links** (`- [name](url): notes`). An `## Optional` section holds links that can be skipped. `/llms-full.txt` is a common companion with the full content in one file. The Actor checks:

- presence (an HTML "not found" page served at the address doesn't count),
- the H1 title, summary, number of sections and links, and an Optional section,
- size, and plain-English `issues` such as a missing title or sections without links.

### How much does it cost?

This Actor uses **pay per event**:

| Event | Price |
| --- | --- |
| Domain checked | **$0.002** per domain |
| Domain that doesn't resolve, can't be reached or refuses the request | **free** |

- 1,000 domains cost **$2**; 5,000 domains cost **$10**.
- Checking llms.txt and llms-full.txt is included.
- Apify also charges a tiny standard start fee per run (about $0.000025 at the default 512 MB).
- Apify's free plan includes $5 of monthly usage, enough for about 2,500 domains a month.
- Set **Maximum cost per run** in the run options and the Actor stops when it is reached.

### Limitations

- **robots.txt is a request, not a lock.** It shows what a site asks crawlers to do. Some crawlers may ignore it, and some vendors say their user-triggered fetchers may not follow it (see `note`). Sites can also block AI crawlers at the firewall or CDN, which robots.txt can't show.
- **Google-Extended is not about Google Search.** It controls Gemini training and grounding. Google Search, including its AI features, uses Googlebot and Search's own controls such as `nosnippet`, which robots.txt tokens for AI crawlers don't change.
- **One hostname per row.** `example.com` and `www.example.com` can have different rules; enter the ones you care about.
- **Politeness:** the Actor makes at most four small requests per domain (robots.txt, llms.txt, llms-full.txt and, if robots.txt lists no sitemap, /sitemap.xml), half a second apart, and only fetches files that the site's robots.txt allows for crawlers. llms-full.txt is read up to 256 KB; its size comes from the server's Content-Length.
- A 401, 403 or 429 on robots.txt usually means the server blocks automated requests; such domains get a free error row instead of a guess.
- Not legal advice: whether AI companies may use content is a legal question that robots.txt alone doesn't answer.

### FAQ

#### Does blocking GPTBot keep my site out of ChatGPT?

Not entirely. GPTBot is OpenAI's training crawler. ChatGPT search uses **OAI-SearchBot**, and ChatGPT fetches pages that users ask about with **ChatGPT-User**. The verdict shows each group separately, so you can see, for example, "Open to AI search, blocks AI training".

#### Why is a crawler "partial" when I blocked it completely?

Check `reason` and `rule`: usually another rule opens some paths again, for example `Allow: /blog/` next to `Disallow: /`, or the crawler matches a group with only some paths disallowed. The line number points to the rule in your robots.txt.

#### Why does a domain return `robots-txt-forbidden`?

The server answered 401 or 403 to our request for robots.txt, which usually means a firewall blocks automated requests. Its rules can't be read, so the row is free.

#### Can I check my own crawler?

Yes. Add its robots.txt token under **Extra user agents to check**. It appears in `agents` with purpose `custom`.

#### Can I run it on a schedule?

Yes. Save a task with your domains and add a schedule in Apify Console, then compare runs to see when a site changes its rules.

### More tools from the same developer

All pay-per-result, no proxy or login needed, built and maintained by the same developer:

**Website audits**

- [Website Accessibility Checker: WCAG 2.2 & EAA](https://apify.com/offerastudio/website-accessibility-audit): accessibility issues with fixes, SEO basics and security headers.
- [Cookie & Tracker Audit: GDPR Consent Checker](https://apify.com/offerastudio/cookie-tracker-audit): cookies and tracking tags that load before consent.
- [Website Change Monitor: Diffs, Prices & Alerts](https://apify.com/offerastudio/website-change-monitor): get a row only when a page changes, with a clean diff.

**Company data and compliance**

- [Company Contact Finder: Emails, Phones & Socials](https://apify.com/offerastudio/company-contact-finder): contact details published on company websites.
- [UK New Companies Feed: Companies House Daily](https://apify.com/offerastudio/uk-new-companies-feed): newly incorporated UK companies with sector filters.
- [EU VAT Number Validator: Bulk VIES Checker](https://apify.com/offerastudio/eu-vat-number-validator): bulk VAT checks with name, address and consultation number.
- [LEI Corporate Tree: GLEIF Parents & Subsidiaries](https://apify.com/offerastudio/gleif-lei-corporate-tree): LEI lookup with parents, subsidiaries and a KYC summary.

**Market signals**

- [US WARN Layoff Notices: 12 States Daily Feed](https://apify.com/offerastudio/us-warn-layoff-notices): layoff and plant closure notices from official state sources.
- [US Product Recalls Monitor: FDA & CPSC Feed](https://apify.com/offerastudio/us-product-recalls-monitor): FDA and CPSC recalls in one feed, with severity.

### Feedback

A crawler missing from the list, or a result that looks wrong? Open an issue on the **Issues** tab with the domain. New AI crawlers are added when their vendors document them.

# Changelog

This Actor's version history is a separate document: https://apify.com/offerastudio/ai-crawler-access-audit/changelog.md

# Actor input Schema

## `domains` (type: `array`):

Websites to check, one per line, up to 5,000 per run. Bare domains (example.com), subdomains (docs.example.com) and full URLs all work; each hostname becomes one row. Note that example.com and www.example.com can have different robots.txt files.

## `checkLlmsTxt` (type: `boolean`):

Also look for /llms.txt and /llms-full.txt and check their basic format (H1 title, summary, link sections, size). Adds two small requests per domain. Same price.

## `additionalUserAgents` (type: `array`):

Optional robots.txt tokens to check besides the built-in list of AI crawlers, e.g. Googlebot-Image or YourOwnBot. Up to 20. They are reported per domain but don't change the AI visibility verdict.

## `includeRobotsTxt` (type: `boolean`):

Add the robots.txt content (first 20,000 characters) to each row, for your own review or archive.

## Actor input object example

```json
{
  "domains": [
    "apify.com",
    "crawlee.dev",
    "docs.apify.com"
  ],
  "checkLlmsTxt": true,
  "includeRobotsTxt": false
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `agents` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "apify.com",
        "crawlee.dev",
        "docs.apify.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("offerastudio/ai-crawler-access-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "apify.com",
        "crawlee.dev",
        "docs.apify.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("offerastudio/ai-crawler-access-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "apify.com",
    "crawlee.dev",
    "docs.apify.com"
  ]
}' |
apify call offerastudio/ai-crawler-access-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,offerastudio/ai-crawler-access-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/o8aeYnmb5ZkT9AqKa/builds/zh2kfJkvXNX9VwKdV/openapi.json
