# AI Crawler Access Checker (robots.txt & llms.txt) (`webintel/ai-crawler-access-checker`) Actor

See which of 33 AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot...) may crawl any website, per path, with an RFC 9309 robots.txt matcher. Also checks llms.txt, llms-full.txt, ai.txt, noai meta tags and TDMRep.

- **URL**: https://apify.com/webintel/ai-crawler-access-checker.md
- **Developed by:** [Deepak Ganesh](https://apify.com/webintel) (community)
- **Categories:** SEO tools, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 sites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## AI Crawler Access Checker (robots.txt & llms.txt)

![AI Crawler Access Checker (robots.txt & llms.txt)](https://api.apify.com/v2/key-value-stores/21SBwDwdNtlnapIdO/records/ai-crawler-access-checker.png?v=1403ae33)

Find out which AI crawlers can crawl any website. For **33 AI and search bots** (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, Google-Extended, PerplexityBot, CCBot, Bytespider, Applebot-Extended, Meta-ExternalAgent, Amazonbot and more) you get an **allowed / partial / blocked** verdict for each path you test, plus the exact group and robots.txt line behind it.

- ✅ **Correct robots.txt parsing (RFC 9309).** A purpose-built parser follows Google's documented behaviour: the most specific user-agent group wins, groups for the same bot are merged, the longest rule wins and Allow wins ties, `*` and `$` wildcards work, percent-encoding is normalized, and status codes are handled per spec (404 = all allowed, 5xx/429/timeout = disallow all).
- 🧭 **Tests any path, not just the homepage.** Add `/blog/` or `/products/` and see per-path verdicts. If an input URL has a path, that path is tested too.
- 🏷️ **Classifies each bot by purpose.** Every bot is tagged with its owner and purpose: *training*, *AI search*, *user-triggered* or *classic search engine*. The summary tells you at a glance whether a site blocks all AI training while staying visible in AI search.
- 📄 **Checks AI files and signals in one pass.** You also get `llms.txt` (title, summary, link and section counts, spec check), `llms-full.txt` (exact size via a Range request, without downloading it), `ai.txt`, `noai`/`noimageai` in `<meta name="robots">` and `X-Robots-Tag`, the **W3C TDMRep** opt-out (header, meta or `/.well-known/tdmrep.json`), Cloudflare **Content-Signal** lines, **RSL `License:`** URLs, sitemaps and crawl-delays.
- 🔀 **Follows redirects like a crawler.** `example.com` → `www.example.com` is followed, and the host that crawlers actually see is the one checked.
- ⚡ **Fast and cheap.** HTTP only, with no browser, and checks hundreds of sites per minute. Failed sites are free.

### What makes it different

Other AI-crawler checkers on the Store look at a fixed list of 6–30 bots and usually only test `/`. This Actor:

| | This Actor | Typical alternatives |
|---|---|---|
| Robots matcher | Full RFC 9309 implementation: longest match, Allow on ties, `*`/`$`, percent-encoding, group merging. Covered by Google's documented test cases | Often a plain prefix or "Disallow: /" check |
| Paths | Any number of paths per site, each with its own deciding rule | Homepage only |
| Custom bots | Add any user-agent token | Fixed list |
| 5xx / 429 / timeouts | Treated as *disallow all* (Google behaviour) and flagged | Often reported as "allowed" |
| `www`/canonical redirects | Checks the host crawlers actually reach | Often checks the bare domain |
| Extra AI opt-out signals | `noai`/`noimageai`, TDMRep, `ai.txt`, Content-Signal, RSL License | Rarely |
| `llms-full.txt` size | Exact size via HTTP Range, without downloading the file | Not checked or fully downloaded |

### Use cases

- **SEO / GEO agencies.** Audit clients to confirm they are visible to ChatGPT search, Perplexity and Claude, and not accidentally blocking OAI-SearchBot when they only wanted to block GPTBot.
- **Publishers and site owners.** Verify that AI training opt-outs (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot...) actually work on every section of the site.
- **AI companies and researchers.** Check consent signals (robots.txt, TDMRep, noai, Content-Signal, RSL) across many domains before collecting data.
- **Monitoring.** Schedule the Actor to catch robots.txt changes, such as a new block on your search bot or a 5xx robots.txt that silently stops all crawling.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `urls` | array of strings | (required) | Domains or URLs, e.g. `nytimes.com` or `https://example.com/blog/post`. Duplicates and URLs on the same host are merged into one result. |
| `testPaths` | array of strings | `["/"]` | Paths tested for every bot on every site. The input URL's own path is always added. Up to 25 paths per site. |
| `userAgents` | array of strings | `[]` | Extra robots.txt tokens to check, such as `MyCompanyBot`. A full user-agent string works too, because the product token is extracted. |
| `maxConcurrency` | integer | `10` | Websites checked in parallel. |
| `requestTimeoutSecs` | integer | `20` | Timeout for each request. Transient errors are retried. |

```json
{
  "urls": ["nytimes.com", "https://www.anthropic.com", "docs.apify.com/platform"],
  "testPaths": ["/", "/blog/"],
  "userAgents": ["MyCompanyBot"]
}
```

### Output

You get one dataset item per website. The dataset has these views: **Overview**, **Per-crawler verdicts** (one row per site × bot) and **llms.txt & AI signals**.

Example, trimmed from a real run (3 of the 33 bots shown):

```json
{
  "inputUrl": "https://www.nytimes.com/section/technology",
  "url": "https://www.nytimes.com/",
  "domain": "nytimes.com",
  "success": true,
  "robotsTxtUrl": "https://www.nytimes.com/robots.txt",
  "robotsTxtStatus": 200,
  "robotsTxtFound": true,
  "robotsTxtOutcome": "parsed",
  "robotsTxtSizeBytes": 8547,
  "testedPaths": ["/", "/section/technology"],
  "bots": [
    {
      "name": "GPTBot", "owner": "OpenAI", "purpose": "training", "userAgentToken": "GPTBot",
      "allowed": false, "access": "blocked", "matchedGroup": "GPTBot", "explicitlyListed": true,
      "matchedRule": "Disallow: / (line 233)",
      "paths": [
        { "path": "/", "allowed": false, "matchedRule": "Disallow: / (line 233)" },
        { "path": "/section/technology", "allowed": false, "matchedRule": "Disallow: / (line 233)" }
      ],
      "crawlDelay": null, "contentSignal": null, "note": null
    },
    {
      "name": "OAI-SearchBot", "owner": "OpenAI", "purpose": "search", "userAgentToken": "OAI-SearchBot",
      "allowed": false, "access": "blocked", "matchedGroup": "OAI-SearchBot", "explicitlyListed": true,
      "matchedRule": "Disallow: / (line 266)", "note": "ChatGPT search results"
    },
    {
      "name": "Googlebot", "owner": "Google", "purpose": "search-engine", "userAgentToken": "Googlebot",
      "allowed": true, "access": "allowed", "matchedGroup": "Googlebot", "explicitlyListed": true,
      "matchedRule": "No matching rule (allowed by default)", "note": "Google Search incl. AI Overviews"
    }
  ],
  "summary": {
    "allowedCount": 9,
    "partialCount": 0,
    "blockedCount": 24,
    "blocksAllAiTraining": false,
    "blocksAnyAiTraining": true,
    "blocksAiSearch": true,
    "blocksAiUserFetchers": true,
    "blockedBots": ["GPTBot", "OAI-SearchBot", "ChatGPT-User", "ClaudeBot", "Claude-SearchBot", "..."],
    "partiallyBlockedBots": [],
    "allowedBots": ["Googlebot", "Bingbot", "Applebot", "Amazonbot", "cohere-training-data-crawler", "MistralAI-User", "AI2Bot", "Webzio-Extended", "PanguBot"],
    "byPurpose": {
      "training": { "allowed": 5, "partial": 0, "blocked": 12 },
      "search": { "allowed": 0, "partial": 0, "blocked": 5 },
      "user-triggered": { "allowed": 1, "partial": 0, "blocked": 7 },
      "search-engine": { "allowed": 3, "partial": 0, "blocked": 0 }
    },
    "verdict": "Blocks 12 of 17 AI training bots; 5 of 5 AI search bots blocked; 7 of 8 user-triggered AI fetchers blocked."
  },
  "contentSignals": [],
  "llmsTxt": { "found": false, "url": "https://www.nytimes.com/llms.txt", "status": 404, "sizeBytes": null, "title": null, "linkCount": 0 },
  "llmsFullTxt": { "found": false, "url": "https://www.nytimes.com/llms-full.txt", "status": 404, "sizeBytes": null, "sizeIsLowerBound": false },
  "aiTxt": { "found": false, "url": "https://www.nytimes.com/ai.txt", "status": 404, "sizeBytes": null, "blocksAll": false },
  "metaRobotsAi": { "noai": false, "noimageai": false, "noindex": false, "nofollow": false, "directives": [] },
  "tdmRep": { "reserved": null, "source": null, "policyUrl": null, "wellKnownFound": false },
  "homepageStatus": 403,
  "sitemaps": ["https://www.nytimes.com/sitemaps/new/news.xml.gz", "https://www.nytimes.com/sitemaps/new/sitemap.xml.gz"],
  "rslLicenses": [],
  "crawlDelay": {},
  "issues": ["Homepage returned HTTP 403; meta robots tags may be missing from the response."],
  "checkedAt": "2026-10-07T16:30:26.278Z"
}
```

When a site has an `llms.txt`, the `llmsTxt` field looks like this (from docs.apify.com):
`{ "found": true, "sizeBytes": 96450, "title": "Apify Documentation", "linkCount": 460, "sectionCount": 17, "startsWithH1": true }`.

**Key fields**

- `access` can be `allowed` (all tested paths), `partial` (some paths) or `blocked` (no tested path).
- `matchedGroup` is the bot's own group, `*`, or `null` when no group applies.
- `robotsTxtOutcome` is `parsed`, `not-found-allow-all` (4xx) or `unreachable-disallow-all` (5xx, 429 or a network error).
- `issues` lists problems found: soft-404 HTML robots.txt or llms.txt, rules before any `User-agent`, unparseable lines, files over 500 KiB, redirects, `noai` set without matching robots.txt blocks, and more.
- Failed sites (invalid input, domains that don't resolve, sites that can't be reached) have `success: false` and an `error`.

### Pricing

This Actor uses pay-per-event pricing:

| Event | Price |
|---|---|
| Site checked (one website, all bots and paths) | **$0.002** ($2 per 1,000 sites) |
| Actor start | $0.00005 |

**Failed items are free.** You are not charged for invalid inputs, domains that don't resolve or sites that can't be reached. Testing extra paths or extra bots costs nothing more. Set a maximum cost per run and the Actor stops cleanly when it is reached.

### FAQ

**Why is a bot "blocked" when robots.txt never mentions it?**
The `*` group applies to it. Check `matchedGroup` (`*`) and `matchedRule`.

**robots.txt returns 503. Why is everything blocked?**
RFC 9309 and Google treat an unreachable robots.txt (5xx, 429 or a timeout) as "disallow all" until it recovers. The Actor flags this in `issues`, because it often means bot protection or a misconfiguration is silently stopping crawlers.

**robots.txt returns 403. Why is everything allowed?**
Per the spec, 4xx means "no robots.txt", so everything is allowed. A 401/403 is usually a firewall, so it is flagged in `issues`.

**Does Google-Extended block Google Search or AI Overviews?**
No. Google-Extended and Applebot-Extended are control tokens for AI training. They are not separate crawlers, and blocking them doesn't affect Googlebot or Applebot. That's why they are classified as *training*.

**Does the Actor obey robots.txt itself?**
It only fetches robots.txt, the homepage and the well-known AI files (`/llms.txt`, `/llms-full.txt`, `/ai.txt`, `/.well-known/tdmrep.json`), about 6 small requests per site. It never crawls content.

**Can I check my own crawler?**
Yes. Add its token to `userAgents`.

This Actor is not affiliated with, endorsed by or sponsored by OpenAI, Anthropic, Google, Microsoft, Perplexity, Apple, Meta, Amazon, ByteDance, Common Crawl, Cloudflare or any other company whose crawler is listed. Bot names are used only to identify them.

### Changelog

- **0.1** (2026-10): First release. Covers 33 crawlers, an RFC 9309 matcher with per-path tests, llms.txt / llms-full.txt / ai.txt checks, noai meta and X-Robots-Tag, TDMRep, Content-Signal and RSL License.

# Actor input Schema

## `urls` (type: `array`):

Websites to check, one per line. Plain domains ("example.com") or full URLs. If a URL has a path ("https://example.com/blog/post"), that path is tested too. URLs on the same host are merged into one result.

## `testPaths` (type: `array`):

Paths tested for every crawler on every website, e.g. "/", "/blog/", "/products/item?id=1". The path of each input URL is always tested as well.

## `userAgents` (type: `array`):

Optional extra robots.txt user-agent tokens to check besides the built-in list of 33 AI and search crawlers, e.g. "MyCompanyBot". Full user-agent strings work too; the product token is extracted.

## `maxConcurrency` (type: `integer`):

How many websites are checked in parallel.

## `requestTimeoutSecs` (type: `integer`):

Timeout for each HTTP request (robots.txt, homepage, llms.txt...). Transient errors are retried.

## Actor input object example

```json
{
  "urls": [
    "nytimes.com",
    "https://www.anthropic.com",
    "docs.apify.com"
  ],
  "testPaths": [
    "/"
  ],
  "userAgents": [],
  "maxConcurrency": 10,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

All checked websites (full JSON incl. bots\[], summary, llmsTxt, metaRobotsAi, sitemaps, issues).

## `overview` (type: `string`):

One row per website: robots.txt status, blocked/allowed counts, training/search verdicts and llms.txt.

## `bots` (type: `string`):

One row per website and crawler: owner, purpose, allowed/partial/blocked, matched group and rule.

## `aiFiles` (type: `string`):

llms.txt, llms-full.txt, ai.txt, noai meta tags, TDMRep and Content-Signal per website.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "nytimes.com",
        "https://www.anthropic.com",
        "docs.apify.com"
    ],
    "testPaths": [
        "/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("webintel/ai-crawler-access-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "nytimes.com",
        "https://www.anthropic.com",
        "docs.apify.com",
    ],
    "testPaths": ["/"],
}

# Run the Actor and wait for it to finish
run = client.actor("webintel/ai-crawler-access-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "nytimes.com",
    "https://www.anthropic.com",
    "docs.apify.com"
  ],
  "testPaths": [
    "/"
  ]
}' |
apify call webintel/ai-crawler-access-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,webintel/ai-crawler-access-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/FfGN4YDa4Z5cWS4UL/builds/mZUCJhnAuCSmNhE0Y/openapi.json
