# SEO & GEO Audit (AI Search Readiness Checker) (`datamole/seo-geo-audit`) Actor

Audit any website for classic SEO and AI search readiness. Crawl its pages, score each one from 0 to 100 for SEO and GEO, and get prioritized issues with fixes, from titles, headings and broken links to structured data, llms.txt and which AI crawlers can read the site.

- **URL**: https://apify.com/datamole/seo-geo-audit.md
- **Developed by:** [DataMole](https://apify.com/datamole) (community)
- **Categories:** SEO tools, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $23.00 / 1,000 result-items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

![DataMole](https://raw.githubusercontent.com/YoiderD/datamole-assets/main/banner.png)

## 🔎 SEO & GEO Audit: AI Search Readiness Checker

> **Find out why a website is not ranking, and whether ChatGPT, Claude and Perplexity can read it.** Every page gets an SEO score and a GEO score from 0 to 100, plus a prioritized list of issues, each with a concrete fix.

> 🕒 **Last updated:** 2026-10-01 · **25+ page checks** · **9 site checks** · **13 AI crawlers tracked**

**SEO & GEO Audit** crawls a website and audits each page for classic SEO (titles, descriptions, headings, canonical tags, indexability, broken links, speed) and for **GEO, Generative Engine Optimization**: how ready the site is to be understood and cited by AI search tools. It checks structured data, `llms.txt`, whether content is readable without JavaScript, and which AI crawlers your `robots.txt` blocks. You get one row per page and one summary row per website, with site-wide problems such as duplicate titles, a missing sitemap or missing security headers.

| 🎯 Who uses it | 💡 What for |
|---|---|
| SEO agencies and freelancers, marketing teams, web developers, founders, content teams | Client audits and proposals, technical SEO checkups, AI search readiness reports, monitoring after a redesign or migration, prospecting (audit a lead's site before the sales call) |

### 📋 What it checks

**On every page (SEO)**

- 🏷️ **Title and meta description:** missing, too long, too short or duplicated across pages.
- 🧱 **Headings:** missing, empty or multiple H1, skipped heading levels.
- 🔗 **Canonical and indexability:** missing canonical, canonical pointing elsewhere, `noindex` in meta tags or the `X-Robots-Tag` header.
- 📱 **Mobile and language:** viewport tag and `lang` attribute.
- 🖼️ **Images:** images without alt text.
- 📝 **Content:** word count and thin content.
- ⚡ **Performance signals:** server response time and HTML size.
- 🔒 **Security:** HTTPS and mixed content.
- 💔 **Broken links:** links returning 404, 410, server errors or dead domains.
- 📣 **Social sharing:** Open Graph and Twitter card tags.

**On every page (GEO, AI search readiness)**

- 🧩 **Structured data:** JSON-LD and microdata types found, invalid JSON-LD.
- 🏢 **Organization schema** on the homepage, so AI assistants identify the brand.
- ❓ **FAQ content without FAQPage schema** (AI assistants often quote FAQs).
- 📰 **Articles without Article schema.**
- 🤖 **Content that requires JavaScript** (most AI crawlers don't run it).

**Once per website**

- 🗺️ `robots.txt`, XML sitemap (and how many URLs it lists), sitemap referenced in `robots.txt`.
- 🤖 `llms.txt` and `llms-full.txt`.
- 🚦 **AI crawler access:** GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, anthropic-ai, PerplexityBot, Google-Extended, Applebot-Extended, CCBot, Bytespider, Meta-ExternalAgent, Amazonbot.
- 🔐 HTTP to HTTPS redirect, HSTS and other security headers.
- 📊 Site SEO and GEO scores, and the most common issues across all pages.

### ⚙️ Input

| Field | What it does |
|---|---|
| `urls` | Domains or page URLs. Pages of the same website are grouped into one audit. |
| `maxPagesPerSite` | Pages to crawl per website (default 10). Use `1` to audit only the URL you entered. |
| `maxItems` | Hard cap of audited pages for the whole run. |
| `useSitemap` | Also discover pages from the XML sitemap (default on). |
| `checkBrokenLinks` | Test the links found on audited pages (default on). |
| `maxLinksToCheck` | Links to test per website (internal links first). |

**Example: quick audit of a client's homepage and 9 more pages**

```json
{
    "urls": ["https://www.pucp.edu.pe"],
    "maxPagesPerSite": 10
}
```

**Example: AI readiness check of many homepages (one page each, no link checks)**

```json
{
    "urls": ["apify.com", "example.com", "shopify.com"],
    "maxPagesPerSite": 1,
    "checkBrokenLinks": false
}
```

> ℹ️ **Free Apify plan:** runs return a preview of up to 10 results. Any paid Apify plan unlocks the full results.

### 📊 Output

Real page record from a cloud run (trimmed):

```json
{
    "type": "page",
    "url": "https://www.pucp.edu.pe/",
    "seoScore": 98,
    "geoScore": 90,
    "criticalIssues": 0,
    "warnings": 0,
    "notices": 2,
    "topIssue": "Meta description is 240 characters and may be truncated.",
    "httpStatus": 200,
    "responseTimeMs": 170,
    "title": "Pontificia Universidad Católica del Perú | PUCP",
    "h1": ["Pontificia Universidad Católica del Perú"],
    "indexable": true,
    "language": "es",
    "structuredDataTypes": ["CollectionPage", "BreadcrumbList", "WebSite"],
    "wordCount": 914,
    "internalLinks": 53,
    "externalLinks": 66,
    "brokenLinks": [],
    "issues": [
        {
            "code": "META_DESCRIPTION_TOO_LONG",
            "severity": "notice",
            "category": "seo",
            "message": "Meta description is 240 characters and may be truncated.",
            "fix": "Keep it under 160 characters."
        },
        {
            "code": "ORGANIZATION_SCHEMA_MISSING",
            "severity": "notice",
            "category": "geo",
            "message": "Homepage has no Organization or LocalBusiness structured data.",
            "fix": "Add Organization JSON-LD with name, logo, url and sameAs links so AI assistants identify the brand correctly."
        }
    ]
}
```

Real site summary record (trimmed):

```json
{
    "type": "site",
    "url": "https://example.com",
    "seoScore": 54,
    "geoScore": 40,
    "pagesAudited": 1,
    "robotsTxt": false,
    "sitemapUrlCount": 0,
    "llmsTxt": false,
    "aiCrawlers": { "GPTBot": "allowed", "ClaudeBot": "allowed", "PerplexityBot": "allowed" },
    "httpsRedirect": false,
    "mostCommonIssues": [
        { "code": "META_DESCRIPTION_MISSING", "pages": 1 },
        { "code": "NO_STRUCTURED_DATA", "pages": 1 }
    ],
    "issues": [
        { "code": "SITEMAP_MISSING", "severity": "warning", "category": "seo", "message": "No XML sitemap found in robots.txt or at /sitemap.xml.", "fix": "Publish an XML sitemap and reference it in robots.txt." },
        { "code": "LLMS_TXT_MISSING", "severity": "notice", "category": "geo", "message": "No /llms.txt file for AI assistants.", "fix": "Publish /llms.txt with a short summary of the site and links to key pages." }
    ]
}
```

Every issue has a `severity` (`critical`, `warning`, `notice`), a `category` (`seo` or `geo`), a plain-English `message` and a `fix`.

### ✨ Why this one

- ✅ **SEO and GEO in one audit.** Most audit tools stop at classic SEO. This one also tells you if AI search tools can read and cite the site.
- ✅ **Actionable.** Every issue comes with a severity and a concrete fix, ready to paste into a client report.
- ✅ **Site-wide view.** Duplicate titles and descriptions, sitemap coverage, AI crawler rules and security headers in one summary row.
- ✅ **Fast and light.** No browser needed: 11 pages across 3 websites took about 35 seconds in our tests.
- ✅ **Works on any website.** Domains or full URLs, one page or a full crawl.

### 🚀 How to use

1. Create a free Apify account.
2. Open the Actor and paste one or more websites.
3. Set how many pages per website to audit.
4. Click **Start**.
5. Sort the results by score or filter by `severity` to see what to fix first, or connect them to Google Sheets, Looker Studio, Make, Zapier or n8n.

### 💼 Use cases

- **SEO agencies:** run a first audit for every new lead and attach the top issues to the proposal.
- **In-house marketers:** schedule a weekly audit and catch new broken links, noindex mistakes or missing descriptions after each release.
- **Web developers:** check a site before and after a migration or redesign.
- **AI search optimization (GEO):** find out if ChatGPT, Claude, Perplexity and Gemini crawlers are blocked, and what structured data is missing.

### ❓ FAQ

**What is GEO?** Generative Engine Optimization: making a site easy for AI assistants and AI search engines to read, understand and cite. Structured data, server-rendered content, `llms.txt` and allowing AI crawlers all help.

**How is the score calculated?** Each page starts at 100 and loses points for each issue, weighted by severity. GEO issues have their own weights. The site score combines the average page score with the site-wide issues.

**Does it measure Core Web Vitals?** No. It reports server response time and HTML size. For lab Core Web Vitals use Google PageSpeed Insights.

**Is blocking AI crawlers always bad?** No. Some sites block them on purpose. The audit flags it so the decision is conscious.

**Does it render JavaScript?** No, it reads the HTML the server sends, which is also what most AI crawlers see. If a page needs JavaScript to show its content, the audit flags it.

**Will it audit pages behind a login?** No. Only public pages.

**How many pages can I audit?** Up to 1,000 per website per run. Large sitemaps are sampled.

### 🔌 Integrations

Works with the Apify API, schedules, webhooks, Google Sheets, Looker Studio, Make, Zapier, n8n, Slack and any tool that reads JSON.

### 🕳️ More from DataMole

| Actor | What it does |
|---|---|
| [Tech Stack Detector](https://apify.com/datamole/tech-stack-detector) | CMS, ecommerce platform, analytics, ad pixels and frameworks behind any website |
| [Domain Intelligence](https://apify.com/datamole/domain-intel-scraper) | WHOIS, DNS, email security grade, SSL expiry and hosting for any domain |

**Need a check we don't cover yet?** Open an issue on this Actor. We read everything.

### ⚖️ Disclaimer

This Actor reads publicly available pages, the same way a search engine crawler does. It is not affiliated with Google, OpenAI, Anthropic or Perplexity. Scores are guidance, not a guarantee of rankings.

# Actor input Schema

## `urls` (type: `array`):

Domains or page URLs to audit. Pages from the same website are grouped into one site audit.

## `maxPagesPerSite` (type: `integer`):

How many pages to crawl and audit per website. Use 1 to audit only the URL you entered.

## `maxItems` (type: `integer`):

Hard cap of audited pages for the whole run. Leave empty for no cap.

## `useSitemap` (type: `boolean`):

Add URLs from the XML sitemap to the crawl, in addition to links found on the pages.

## `checkBrokenLinks` (type: `boolean`):

Test the links found on audited pages and report the ones returning 404, 410, server errors or dead domains.

## `maxLinksToCheck` (type: `integer`):

Internal links are checked first, then external links.

## `proxyConfiguration` (type: `object`):

Not needed for most websites.

## Actor input object example

```json
{
  "urls": [
    "apify.com"
  ],
  "maxPagesPerSite": 10,
  "useSitemap": true,
  "checkBrokenLinks": true,
  "maxLinksToCheck": 100
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "apify.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("datamole/seo-geo-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["apify.com"] }

# Run the Actor and wait for it to finish
run = client.actor("datamole/seo-geo-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "apify.com"
  ]
}' |
apify call datamole/seo-geo-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datamole/seo-geo-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/rL9faYVZGtk9sMNtV/builds/JOsEU1t13kIMHiE8q/openapi.json
