# AI Readiness & SEO Audit: llms.txt, AI Crawlers, Schema (`locaihost/site-audit`) Actor

Check whether ChatGPT, Claude, Perplexity and Google AI can crawl and cite a website. Audits robots.txt AI-bot rules (GPTBot, ClaudeBot…), llms.txt, sitemap, schema.org, titles, meta, canonicals and alt text. Returns a 0–100 AI-readiness score and a prioritised fix list per site.

- **URL**: https://apify.com/locaihost/site-audit.md
- **Developed by:** [locaihost data](https://apify.com/locaihost) (community)
- **Categories:** SEO tools, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.50 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## AI Readiness & SEO Audit: llms.txt, AI Crawlers, Schema

Find out in one run whether **AI search engines and assistants can find, read and cite a website**, and fix the SEO basics at the same time.

ChatGPT search, Perplexity, Claude and Google's AI features now send real traffic. Many sites block them by accident in `robots.txt`, have no `llms.txt`, or lack the structured data AI answers rely on. This Actor checks all of that, page by page, and returns a **0–100 AI-readiness score**, page scores, and a plain-English fix list.

![Real site summaries from this Actor's output: AI-readiness score, llms.txt, robots.txt, AI crawler access and sitemap size for four public sites](https://locaihost.org/img/readme/site-audit.png)

### What it checks

**Per site**

- **AI crawler access** in robots.txt, bot by bot: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, Applebot-Extended, meta-externalagent, Amazonbot, CCBot, Bytespider. Each is reported as `allowed` or `blocked`.
- **llms.txt** present or not.
- **XML sitemap** found (via robots.txt or the default locations) and how many URLs it lists.
- **AI-readiness score** and **recommendations**, e.g. "Allow AI answer engines in robots.txt (PerplexityBot is blocked)".

**Per page** (crawled breadth-first from your start URL, same site only)

- Title and meta description (presence and length), `<h1>`, canonical, `lang`, Open Graph.
- **schema.org structured data** types (JSON-LD and microdata).
- **Indexability** (meta robots and X-Robots-Tag `noindex`), HTTP status, redirects, response time.
- Images without alt text, word count, internal and external link counts.
- A **page score** and a list of issues.

### Who is it for?

- **SEO and marketing agencies:** run it across every client site each month and send the fix list.
- **Site owners and content teams:** see whether your robots.txt is hiding you from AI search.
- **Developers:** add it to CI or a weekly schedule to catch accidental `noindex` or blocked bots after a deploy.

### How to use

1. Enter one or more sites (`example.com` or a start URL like `https://example.com/blog`).
2. Choose **Max pages per site** (10 is a good first audit).
3. Run it. The **Site summaries** view gives the scores and recommendations; **Page audits** gives the per-page detail.

#### Example output: site summary

```json
{
  "type": "site",
  "site": "https://example.com",
  "aiReadinessScore": 64,
  "averagePageScore": 78,
  "pagesAudited": 10,
  "aiCrawlers": { "GPTBot": "blocked", "OAI-SearchBot": "allowed", "ClaudeBot": "allowed", "PerplexityBot": "blocked", "Google-Extended": "allowed" },
  "aiCrawlersBlocked": ["GPTBot", "PerplexityBot"],
  "llmsTxt": false,
  "sitemap": "https://example.com/sitemap.xml",
  "sitemapUrlCount": 214,
  "recommendations": [
    "Add /llms.txt: a short Markdown guide to your key pages for AI assistants",
    "Allow AI answer engines in robots.txt (PerplexityBot are blocked) if you want to be cited in AI search",
    "Fix on 6 page(s): No schema.org structured data"
  ]
}
```

#### Example output: page audit

```json
{
  "type": "page",
  "url": "https://example.com/pricing",
  "status": 200,
  "score": 82,
  "issues": ["Meta description length 34 (aim for 50–170)", "No schema.org structured data"],
  "title": "Pricing — Example",
  "canonical": "https://example.com/pricing",
  "schemaTypes": [],
  "indexable": true,
  "imagesWithoutAlt": 0,
  "wordCount": 612,
  "responseTimeMs": 184
}
```

### How the AI-readiness score works

It is 60% the average page score, plus 10 points each for: an `llms.txt`, an XML sitemap, schema.org data on the homepage, and access for the AI **answer** engines (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot). Blocking **training** crawlers like GPTBot or CCBot is a legitimate choice and doesn't lower the score. It is reported so you can see exactly what you're blocking.

### Pricing

**Pay per page audited.** Site summaries are free. Robots.txt, llms.txt and sitemap checks are included. **Pages that fail are free:** a timeout, DNS error or HTTP 4xx/5xx still shows up as a row (`charged: false`) but costs nothing. A site that can't be reached at all gets a site row with `status: "unreachable"`, no score, and a `reason`.

| Apify plan | Price per 1,000 pages |
|---|---|
| Free and Starter | **$10.00** ($0.01 a page) |
| Scale | $7.50 |
| Business and above | $5.50 |

**Worked example.** An agency audits 20 client sites once a month at the default 10 pages each: 200 pages × $0.01 = **$2.00 a month** on Starter. A one-off 50-page deep audit of a single site costs $0.50. Apify's $0.00005 run-start fee comes on top.

**Free Apify plan:** up to **50 pages per run**. Any paid Apify plan removes the cap.

### Use with AI agents (MCP)

Use this Actor as a tool in Claude, Cursor, VS Code or any MCP client through the [Apify MCP server](https://mcp.apify.com). Agents find it by searching for "AI readiness" or "llms.txt".

1. Add the server. In Claude Code: `claude mcp add apify https://mcp.apify.com/ -t http`. In Cursor or Claude Desktop, add `https://mcp.apify.com` as a remote MCP server. Sign in with your Apify account when asked, or send your API token as `Authorization: Bearer <APIFY_TOKEN>`.
2. To expose only this Actor as a tool, use `https://mcp.apify.com/?tools=locaihost/site-audit`.
3. Ask in plain language, for example:

> Can ChatGPT search and Perplexity crawl our site acme.com? Audit 10 pages and give me the AI-readiness score and the top fixes.

The minimal input an agent should send:

```json
{
  "urls": ["acme.com"],
  "maxPagesPerSite": 10
}
```

The `site` row holds the score, the bot-by-bot `aiCrawlers` map and the `recommendations`, so an agent can answer from that single row.

### Integrations

Call it from any language with the Apify API. This request waits for the run and returns every row (site summaries and page audits) as JSON. Synchronous runs time out after 5 minutes, so use the async run endpoint for many sites:

```bash
curl -X POST "https://api.apify.com/v2/acts/locaihost~site-audit/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["example.com"], "maxPagesPerSite": 10}'
```

Python, with `pip install apify-client`:

```python
from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("locaihost/site-audit").call(run_input={
    "urls": ["example.com", "example.org"],
    "maxPagesPerSite": 10,
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    if row["type"] == "site":
        print(row["site"], row["aiReadinessScore"], row.get("aiCrawlersBlocked"), row["recommendations"][:3])
```

**Exports:** one dataset holds two kinds of rows. In the Console, open **Export** and pick the **Site summaries** or **Page audits** view, so your CSV or Excel file has only that row type's columns.

**Schedules:** save your site list as a task and add a weekly or monthly schedule in Apify Console → Schedules to catch an accidental `noindex` or a newly blocked bot after a deploy.

**Where the results go:** Apify's built-in integrations send each run's dataset to Slack, email, Google Sheets, Zapier, Make, n8n (the Apify node) or any webhook.

### FAQ

**How fresh are the results?**
Live. Every run fetches robots.txt, llms.txt, the sitemap and the pages at run time.

**Is it legal to audit a site with this?**
It reads only public pages, and it respects robots.txt for its own user agent (`locaihost-site-audit`). It's meant for sites you own or manage, and for your clients' sites.

**Does it collect personal data?**
No. It records technical facts about pages (titles, tags, status codes, counts), not people or form contents.

**Does blocking GPTBot lower my score?**
No. Blocking **training** crawlers such as GPTBot or CCBot is a legitimate choice and doesn't cost points. Blocking the **answer** engines (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot) does, because they decide whether AI search can cite you.

**Does it render JavaScript?**
Not in this version. Pages are analysed as the server sends them, which is also what most AI crawlers see.

**How many pages can it audit?**
Up to 500 per site (`maxPagesPerSite`), crawled breadth-first from the start URL. `maxResults` caps the whole run.

### Troubleshooting

| What you see | What it means and what to do |
|---|---|
| Site row with `status: "unreachable"` and a `reason` like `DNS lookup failed … (ENOTFOUND)` | The domain doesn't resolve or the server didn't answer. Check the spelling. No charge. |
| Site row with `status: "blocked-by-robots"` | The site's robots.txt disallows our crawler on the start URL. We respect that; there's no charge. |
| Site row with `status: "http-error"` | The start page and every page tried returned an HTTP error. No charge. |
| Page row with `charged: false` and an `error` | That page failed (timeout, 4xx/5xx). It's listed so you can fix it, and it's free. |
| `No valid public website in the input (… skipped — see the SUMMARY record)` | Every input was an IP address, `localhost`, a private network name or not http(s). Enter public domains like `example.com`. |
| `Done: 50 pages audited … (free-plan cap of 50 pages reached).` | The free plan's per-run cap. Lower `maxPagesPerSite` or use any paid Apify plan. |

### Good to know

- **Polite crawling:** the Actor reads robots.txt and **respects it for its own user agent** (`locaihost-site-audit`). It fetches one page at a time per site with a short pause, and only follows same-site links (no PDFs or images).
- **No JavaScript rendering in v1:** pages are analysed as served. That's what most AI crawlers see too.
- **Audit sites you own or manage.** It's built for site owners and their agencies.
- **Public websites only:** IP addresses, `localhost`, private/internal names (`.local`, `.internal`, …) and anything that resolves to a private network are skipped for free and listed in the run's `SUMMARY` record, as are non-http(s) inputs.

### Feedback

Want another check (Core Web Vitals, hreflang, broken-link checking)? Open an issue on the **Issues** tab.

# Actor input Schema

## `urls` (type: `array`):

Sites to audit: domains (example.com) or start URLs (https://example.com/blog). Each is crawled from that URL through same-site links.

## `maxPagesPerSite` (type: `integer`):

How many pages to audit per site (breadth-first from the start URL). You pay per page audited.

## `maxResults` (type: `integer`):

Overall page budget for the run. 0 = no limit.

## Actor input object example

```json
{
  "urls": [
    "apify.com",
    "example.com"
  ],
  "maxPagesPerSite": 10,
  "maxResults": 0
}
```

# Actor output Schema

## `sites` (type: `string`):

One record per site: AI-readiness score, AI crawler access, llms.txt, sitemap, top issues, recommendations.

## `pages` (type: `string`):

One record per audited page: score, issues, title, meta description, canonical, schema types.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "apify.com",
        "example.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("locaihost/site-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "apify.com",
        "example.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("locaihost/site-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "apify.com",
    "example.com"
  ]
}' |
apify call locaihost/site-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,locaihost/site-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/YgLt54U7WolUAbHPB/builds/cnpfBwqOFGsRXdoNm/openapi.json
