# AI Crawler Checker - robots.txt Checker for GPTBot & AI Bots (`tidytools/ai-crawler-access-checker`) Actor

Which AI bots can read a site? Bulk-check up to 10,000 domains for 26 AI crawlers (GPTBot, ClaudeBot, PerplexityBot...) in robots.txt: AI search score, fix snippet, Content Signals. $2/1k sites.

- **URL**: https://apify.com/tidytools/ai-crawler-access-checker.md
- **Developed by:** [Yukai Lin](https://apify.com/tidytools) (community)
- **Categories:** SEO tools, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does AI Crawler Access Checker do?

It tells you **which AI crawlers are allowed to read a website** according to its `robots.txt`, scores it, gives you a **ready-to-paste robots.txt fix**, and can **track changes week to week**. It also reads **Cloudflare Content Signals** (`Content-Signal: search=yes, ai-train=no`), and checks for an **llms.txt** file and `noai` directives. Check one site or **up to 10,000 domains per run for $2 per 1,000 sites**.

Why it matters: if AI search crawlers such as OAI-SearchBot, Claude-SearchBot or PerplexityBot are blocked, the site is unlikely to be cited in ChatGPT, Claude or Perplexity answers. Blocking *training* crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) does not have that effect.

**Free daily benchmark:** see how 1,005 of the world's most visited websites treat these 29 crawlers, by category and for the top 100, in the [AI Crawler Index](https://tools.yukai.uk/ai-crawler-index). It is updated every day with the same checks as this Actor, and you can download the full table as CSV.

### What it checks

- 🤖 **29 AI crawlers** from OpenAI, Anthropic, Perplexity, Google, Apple, Meta, Amazon, DuckDuckGo, Mistral, Common Crawl, ByteDance and Cohere, each labelled as **training**, **search**, **user** (fetches made when a person asks an assistant) or **ads** (ad review, shown for information and not scored). **Crawler list last checked against each vendor's documentation: 2026-10-01** (see the list below)
- 🏷️ **Content Signals**: Cloudflare's `Content-Signal` line in robots.txt (`search`, `ai-input`, `ai-train` = yes/no) is returned as `contentSignals`, mentioned in the verdict, and tracked between runs
- 📜 **robots.txt rules** evaluated like search engines do: the most specific user-agent group, longest matching rule, `*` and `$` wildcards
- 🚦 **Per-bot status**: `allowed`, `partial` (home page allowed but a tested path is blocked, or the bot's own group has Disallow rules), `blocked`, or `unknown`
- 🔒 **Honest "unknown"**: if robots.txt answers 401, 403, 418, 429 or a bot challenge page, the result is `policy: "unknown"` with scores `null`, not "open", and **the site is not charged**
- 📊 **Scores and policy class**: `aiSearchScore` (AI search and assistant crawlers allowed, 0-100), `aiAccessScore` (all AI crawlers), `policy`: `open`, `search-only`, `blocks-search`, `restrictive`, `no-robots` or `unknown`
- 🛠️ **Fix snippets**: `recommendedRobotsSnippet` re-allows the blocked AI search and assistant crawlers (keeping your other rules); `suggestedPolicy` is a complete template that blocks training crawlers and allows AI search
- 🔁 **Change tracking**: give the run a **monitor name** and each site is compared with its previous result (`comparison.changedBots`, policy, llms.txt); optionally output only changed sites and get a Slack, Discord or JSON **webhook**
- 🛣️ **Any paths you choose** (e.g. `/blog/`, `/products/`), not only the home page
- 📄 **llms.txt and llms-full.txt**: present or not (HTML "not found" pages are not counted)
- 🚫 **noai / noimageai** in the home page's meta robots or `X-Robots-Tag` header
- 🗺️ **Sitemaps** listed in robots.txt, and a **plain-language verdict** for every site

### Which tool should I use?

| | **AI Crawler Access Checker** (this Actor) | **GEO Readiness Audit** (by TidyTools) |
|---|---|---|
| Question | "Do these 5,000 domains allow AI crawlers in robots.txt, and what changed since last week?" | "Why is my site not cited by AI answers, and what should I fix first?" |
| Depth | robots.txt policy only: custom paths, matched rules, fix snippet, change tracking | Live firewall test (catches CDN "block AI bots" rules and challenge pages), content without JavaScript, structured data, score, HTML report |
| Best for | Researchers, data teams, agencies watching many domains | Site owners and agencies auditing one site or a few competitors |
| Price | $0.002 per site | about $0.018 per site |

For a firewall diagnosis (robots.txt allows a bot but the CDN blocks it), use GEO Readiness Audit.

### How much does it cost?

| Event | Price |
|---|---|
| Checked website | **$2.00 / 1,000 websites** |
| Re-checked website, unchanged (monitoring) | **$0.50 / 1,000 websites** |

**No start fee.** Unreachable websites and sites whose robots.txt cannot be read are not charged. Scores, fix snippets and change tracking are included. With a **monitor name**, a site whose result did not change since the last run is charged at the re-check price ($0.0005); new and changed sites cost $0.002. Example: watching 1,000 domains weekly where 2% change costs about $0.53 per week. With *Only output changed sites* or *Only output sites with a problem*, skipped sites are still checked and charged. Higher Apify plans get volume discounts.

For comparison, other AI crawler checkers in the Apify Store charge $0.005 to $0.02 per site, some with a start fee, and some limit a run to 50 or 100 sites (checked September 2026).

### Control your cost

- **Charged**: each website whose robots.txt was read ($0.002), or $0.0005 for an unchanged site in a monitor.
- **Free**: invalid input lines, duplicate lines (merged into one site), unreachable websites, and robots.txt that could not be read (401, 403, 418, 429, bot challenge). Every row has `charged: true` or `false`, and the `error` text of free rows ends with "(not charged)".
- At the start, the log and status message show the plan: number of websites × price = the most the run can cost, compared with your **maximum charge per run** (set in the run options).
- When that maximum is reached, the run stops checking new sites. The status message says how many were not checked, and `SUMMARY.notProcessed` lists them (count and up to 100 inputs) so you can run them again. A monitor webhook sent from such a run carries `incomplete: true`.
- **If Apify restarts the run** (server migration or Resurrect), items already finished are skipped and not charged again (`SUMMARY.resumedSkipped`).
- **Time limit per website** (Advanced settings, default 120 seconds): a site whose robots.txt, llms.txt and home page take longer in total gets a row with `errorType: "timeout"` and is not charged.

### How to use it

1. Paste **websites**, one per line (`example.com` or full URLs). A line with several domains separated by commas or spaces is split. A line that is not a website gets its own error row (`errorType: "invalid_input"`, not charged) and the other sites are still checked. A URL with a path (e.g. `nytimes.com/section/world`) also tests that path.
2. Optional: add **paths to test**.
3. Optional, for monitoring: set a **monitor name**, schedule the Actor (e.g. weekly), and add a **webhook URL** (Slack and Discord URLs get a formatted message).
4. Click **Start** and export the results as JSON, CSV or Excel. The *Fix snippets* view lists the robots.txt lines to paste for each site.

#### Input example

```json
{
    "websites": ["nytimes.com", "python.org", "stripe.com, lowes.com"],
    "paths": ["/", "/search"],
    "monitorName": "clients-weekly",
    "outputOnlyChanges": false,
    "webhookUrl": "https://hooks.slack.com/services/..."
}
```

#### Output example (real result, September 2026, shortened)

```json
{
    "site": "https://nytimes.com",
    "verdict": "Some AI search/assistant crawlers are blocked (OAI-SearchBot, Claude-SearchBot, PerplexityBot, meta-webindexer, DuckAssistBot, ChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher): the site may not appear in those AI answers.",
    "policy": "blocks-search",
    "aiSearchScore": 47,
    "aiAccessScore": 35,
    "blockedBots": ["GPTBot", "OAI-SearchBot", "ChatGPT-User", "ClaudeBot", "..."],
    "partialBots": ["Google-GeminiNotebook", "Google-Agent", "Applebot", "Amazonbot", "..."],
    "contentSignals": null,
    "recommendedRobotsSnippet": "# Allow AI search and assistant crawlers (they decide whether the site can appear in AI answers)\n# Also remove \"Disallow: /\" from the existing group(s) for: OAI-SearchBot, ChatGPT-User, ...\nUser-agent: OAI-SearchBot\nUser-agent: ChatGPT-User\n...\nAllow: /\nDisallow: /ads/\n...",
    "suggestedPolicy": "# AI model training crawlers: blocked (does not affect AI search answers)\nUser-agent: GPTBot\n...\nDisallow: /\n\n# AI search and assistant crawlers: allowed\nUser-agent: OAI-SearchBot\n...",
    "robotsTxt": { "status": 200, "found": true, "state": "found", "reason": null, "sitemaps": ["https://www.nytimes.com/sitemaps/new/news.xml.gz", "..."] },
    "llmsTxt": { "found": false, "url": "https://nytimes.com/llms.txt", "status": 404 },
    "noAiDirective": false,
    "comparison": {
        "status": "changed",
        "changedBots": [{ "bot": "MistralAI-User", "before": "allowed", "after": "partial" }],
        "llmsTxtChanged": false,
        "policyChanged": null,
        "previousCheckedAt": "2026-09-29T02:18:46.163Z"
    },
    "bots": [
        { "bot": "GPTBot", "company": "OpenAI", "purpose": "training", "status": "blocked", "allowedAll": false, "matchedGroup": "GPTBot",
          "paths": [{ "path": "/", "allowed": false, "rule": "Disallow: /" }] }
    ],
    "charged": true
}
```

A site whose robots.txt refuses the request (lowes.com answered HTTP 403) is reported as unknown and not charged:

```json
{ "site": "https://lowes.com", "policy": "unknown", "aiSearchScore": null, "robotsTxt": { "status": 403, "state": "unavailable", "reason": "robots.txt answered HTTP 403" }, "charged": false }
```

The `SUMMARY` record in the key-value store has `status` (`SUCCESS`, `PARTIAL_RESULTS`, `FAILED`, `NO_RESULTS` or `LIMIT_REACHED`), counts (including `invalidInputs` and `duplicateInputs`), `changedSites`, the webhook result, `costPlan` and `notProcessed`. Unreachable sites get `success: false` and an `errorType`.

Every row carries `input`, your original line, so results map back to your spreadsheet. When several lines point to the same site (e.g. `example.com` and `EXAMPLE.com`), the site is checked and charged once and `inputs` lists all of them. `testedPaths` lists the paths checked for every bot. A line that is not a website:

```json
{ "input": "acme", "site": null, "success": false, "errorType": "invalid_input", "error": "\"acme\" is not a website: not a URL or domain (not charged)", "charged": false }
```

A line of words that are not domains (e.g. `Acme Inc`) stays one line and gives one error row; a line of domains separated by spaces (`a.com b.com`) is split.

A site that publishes Content Signals (real result, www.cloudflare.com, September 2026):

```json
{
    "site": "https://www.cloudflare.com",
    "verdict": "Open to all listed AI crawlers. Content-Signal: search allowed, AI answers (ai-input) allowed, AI training (ai-train) allowed.",
    "policy": "open",
    "contentSignals": { "search": "yes", "aiInput": "yes", "aiTrain": "yes", "userAgent": "*", "raw": "ai-train=yes, search=yes, ai-input=yes" },
    "comparison": { "status": "changed", "contentSignalChanged": { "before": null, "after": "ai-train=yes, search=yes, ai-input=yes" } }
}
```

`contentSignals` is `null` when robots.txt has no `Content-Signal` line. A missing signal means "no preference" (neither allowed nor refused). When different user-agent groups declare different signals, all of them are listed in `contentSignalGroups`.

#### Policy classes

| policy | Meaning |
|---|---|
| `open` | No listed AI crawler is blocked at the home page |
| `search-only` | Only training crawlers are blocked; AI search and assistants are allowed |
| `blocks-search` | Some AI search or assistant crawlers are blocked |
| `restrictive` | All AI search and assistant crawlers are blocked (or robots.txt returns a server error) |
| `no-robots` | No robots.txt: everything is allowed |
| `unknown` | robots.txt could not be read (401, 403, 418, 429 or a bot challenge) |

#### Crawlers checked (last checked 2026-10-01)

| Company | Training | AI search | User-triggered | Ads review |
|---|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User\* | OAI-AdsBot |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User | |
| Perplexity | | PerplexityBot | Perplexity-User\* | |
| Google | Google-Extended | Google-CloudVertexBot (Vertex AI Agents, crawls requested by the site owner) | Google-GeminiNotebook\*, Google-Agent\* | |
| Apple | Applebot-Extended | Applebot (follows Googlebot's rules when not named) | | |
| Meta | Meta-ExternalAgent | meta-webindexer | meta-externalfetcher\* | meta-externalads |
| Amazon | Amazonbot | Amzn-SearchBot | Amzn-User\* | |
| DuckDuckGo | | DuckAssistBot | | |
| Mistral | MistralAI-Training | MistralAI-Index | MistralAI-User | |
| Common Crawl | CCBot | | | |
| ByteDance, Cohere | Bytespider, cohere-ai (no vendor documentation; widely listed in robots.txt files) | | | |

Every token was checked against the company's own crawler documentation. \* The vendor says this user-triggered fetcher generally does not follow robots.txt; such bots carry `ignoresRobots: true` in `bots`, and robots.txt rules for them express your wish rather than a guarantee.

**Ads review** crawlers (`purpose: "ads"`) check ad landing pages or improve advertising products. They are listed so you can see your rules for them, but they count in no score and no policy class. **Amazonbot** is labelled training since 2026-10-01: Amazon's page says it "may be used to train Amazon AI models", while Amzn-SearchBot is the crawler for search experiences such as Alexa.

### Use with AI agents (MCP)

Connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/ai-crawler-access-checker) to Claude, Cursor or any MCP client, then ask e.g. "Which of these 50 domains block ChatGPT search in robots.txt, and what should they paste to fix it?"

```json
{ "websites": ["example.com", "example.org"] }
```

Failed items are not charged and carry an `errorType`.

### Use it from code and integrations

Run it from your own code with the Apify API. This call waits for the run and returns the results as JSON (replace `YOUR_TOKEN` with your Apify API token):

```bash
curl -X POST "https://api.apify.com/v2/acts/tidytools~ai-crawler-access-checker/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"websites":["https://www.nytimes.com","python.org","docs.anthropic.com"]}'
```

The synchronous endpoint waits up to 5 minutes. For bigger runs, start the run with `POST https://api.apify.com/v2/acts/tidytools~ai-crawler-access-checker/runs` and read the dataset when it finishes, or use the `apify-client` package for JavaScript or Python.

**Schedules and integrations:** run it daily or weekly with Apify Schedules, get a webhook when a run finishes, or send the results to Zapier, Make, n8n, Google Sheets, Slack and other apps with Apify integrations. Results can be exported as JSON, CSV, Excel or XML.

### FAQ

**Will blocking GPTBot remove my site from ChatGPT answers?** No. GPTBot collects training data. ChatGPT search uses OAI-SearchBot, and ChatGPT-User fetches pages when a person asks. You can block GPTBot and still allow the other two; `suggestedPolicy` does exactly that.

**What is Content-Signal?** An extension Cloudflare introduced for robots.txt: a line such as `Content-Signal: search=yes, ai-input=no, ai-train=no` inside a user-agent group states how content may be used after it is fetched: `search` (search index with links and snippets), `ai-input` (feeding AI answers, e.g. RAG or grounding) and `ai-train` (model training). It does not block crawling by itself; it states the site's preference. This Actor returns the signals for all crawlers (`*`) in `contentSignals`.

**My site has no llms.txt. Should it?** It is an emerging proposal, not yet used by the major AI search engines. If you want one, *llms.txt Generator* (by TidyTools) builds llms.txt and llms-full.txt from your sitemap.

**What does "partial" mean?** The home page is allowed, but one of your tested paths is blocked, or the bot has its own group with Disallow rules.

**I changed the paths to test (or the crawler list was updated). Will every site show as changed?** No. Each monitor snapshot stores the tested paths and the crawler list version. When the paths changed, only bots that became blocked or unblocked are reported; when the crawler list changed, new crawlers and the policy class are not compared that time. `comparison.note` says so, and a site with no other change is charged the re-check price.

**robots.txt allows AI search. Can the site still block AI crawlers?** Yes, a CDN or firewall rule can block them regardless of robots.txt. The verdict mentions this; GEO Readiness Audit (by TidyTools) runs a live firewall test.

### Use cases

- **Your own site**: confirm AI search crawlers can read it (and training crawlers are handled the way you want), and copy the fix snippet if not
- **SEO/GEO agencies**: audit client sites in bulk and get a weekly alert when a client's robots.txt starts blocking AI search
- **Research**: measure how many sites in a list block AI crawlers, and how that changes over time (the [AI Crawler Index](https://tools.yukai.uk/ai-crawler-index) does this daily for 1,005 popular sites)

### Tips

- **Advanced settings**: add a *Proxy* to read robots.txt and llms.txt through Apify Proxy (billed to your Apify account) for sites that refuse data-center requests. *Plain HTTP requests from* controls how the home page is read for the `noai` check.
- The webhook can be signed: set a *webhook signing secret* and verify the `X-Signature-256: sha256=<hex>` header (HMAC-SHA256 of the raw body).

### Limitations

- robots.txt is a request, not an enforcement: this Actor reports what the site asks crawlers to do.
- Firewalls or bot protection (e.g. blocking by IP or user agent at the CDN) are not detected here; use GEO Readiness Audit for that.
- The crawler list covers the major AI companies (checked 2026-10-01); others may exist.
- Adding crawlers to the list changes scores: a site that names only a few AI crawlers in robots.txt is judged on all of them (29 since 2026-10-01; ad-review crawlers are not scored).

### Support

Open an issue in the **Issues** tab with the website. Issues are checked regularly.

# Actor input Schema

## `websites` (type: `array`):

Domains or URLs to check, one per line (lines with several domains separated by commas or spaces are split), e.g. example.com or https://www.example.com. robots.txt is read from the site root; a path in a URL (e.g. nytimes.com/section/world) is tested too. Lines that are not websites get an error row and are not charged. Up to 10,000 sites per run.

## `paths` (type: `array`):

Optional paths to test for every bot, e.g. /blog/ or /products/. Default is the home page (/).

## `maxConcurrency` (type: `integer`):

How many websites are checked at the same time.

## `includeFixSnippets` (type: `boolean`):

Add recommendedRobotsSnippet (ready-to-paste lines that re-allow blocked AI search and assistant crawlers) and suggestedPolicy (a full template that blocks training crawlers and allows AI search) to each result.

## `blockedOnly` (type: `boolean`):

Output only sites where an AI search or assistant crawler is blocked or partly blocked, or robots.txt could not be read. All sites are still checked and charged (except unreadable robots.txt).

## `monitorName` (type: `string`):

Optional. Runs with the same name compare each site with its previous result and add a "comparison" field (changed bots, policy, llms.txt). Unchanged sites are charged the re-check price ($0.50 / 1,000 instead of $2 / 1,000). Changing the paths to test does not count as a site change. Schedule the Actor weekly to track robots.txt changes.

## `outputOnlyChanges` (type: `boolean`):

With a monitor name: skip sites whose AI crawler access did not change since the last run (new sites are still output). Skipped sites are still checked and charged the re-check price ($0.50 / 1,000).

## `webhookUrl` (type: `string`):

Optional. When at least one site changed, a summary is POSTed here. Slack and Discord webhook URLs get a formatted message; other URLs get JSON.

## `webhookSecret` (type: `string`):

Optional. Signs the webhook body with HMAC-SHA256 in the X-Signature-256 header (sha256=<hex>).

## `httpVia` (type: `string`):

Some sites block requests from data centers. Auto retries from a second network. Used for the home page (noai check); robots.txt and llms.txt are always read from Apify's network. If both are refused, Auto also tries our second server (Oracle, different IP) before a real browser.

## `siteTimeoutSecs` (type: `integer`):

A website whose robots.txt, llms.txt and home page take longer than this in total (all retries and fallbacks included) is reported with errorType "timeout" and not charged. 0 = no limit.

## `proxyConfiguration` (type: `object`):

Only used for requests sent from Apify's network (robots.txt, llms.txt and the home page in direct mode). Apify proxy usage is billed to your Apify account.

## Actor input object example

```json
{
  "websites": [
    "https://www.nytimes.com",
    "python.org",
    "docs.anthropic.com"
  ],
  "maxConcurrency": 10,
  "includeFixSnippets": true,
  "blockedOnly": false,
  "outputOnlyChanges": false,
  "httpVia": "auto",
  "siteTimeoutSecs": 120
}
```

# Actor output Schema

## `sites` (type: `string`):

No description

## `full` (type: `string`):

No description

## `fixes` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "https://www.nytimes.com",
        "python.org",
        "docs.anthropic.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tidytools/ai-crawler-access-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        "https://www.nytimes.com",
        "python.org",
        "docs.anthropic.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("tidytools/ai-crawler-access-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "https://www.nytimes.com",
    "python.org",
    "docs.anthropic.com"
  ]
}' |
apify call tidytools/ai-crawler-access-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tidytools/ai-crawler-access-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/fHmxS3tQMVb1EfQBu/builds/5ClgQJhNneyM82ezX/openapi.json
