# Website SEO Audit Crawler (`forevertools/website-seo-audit`) Actor

Crawl your website and audit every page for 25+ on-page and technical SEO issues: titles, meta descriptions, H1s, canonicals, noindex, broken links, duplicate content, Open Graph, JSON-LD, alt text, mixed content, speed.

- **URL**: https://apify.com/forevertools/website-seo-audit.md
- **Developed by:** [Forever Tools](https://apify.com/forevertools) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$5.00 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website SEO Audit Crawler

Crawl a whole website and get a page-by-page technical SEO audit plus a site-wide summary: **duplicate titles,
broken internal links, missing meta descriptions, noindex pages, canonical problems** and 20+ other checks.
Fast (plain HTTP, no browser), polite (respects robots.txt), and priced per page so small sites cost cents.

### What it checks (every page)

| Area | Checks |
|---|---|
| Indexing | HTTP status, redirects, `meta robots` / `X-Robots-Tag` noindex, canonical present / points elsewhere / invalid |
| Titles & meta | missing, too long (>60), too short (<20); meta description missing / >160 / <50; duplicates across the site |
| Content | H1 missing / multiple, word count & thin content (<200 words), `html lang`, viewport |
| Links | internal / external / nofollow counts, **broken internal links with the pages that link to them** |
| Media | images missing `alt` |
| Social & schema | Open Graph title/description/image, JSON-LD types (and invalid JSON-LD), hreflang alternates |
| Security & speed | mixed content on HTTPS, server response time, HTML size |

### Output

**Dataset** — one row per page:

```json
{
  "url": "https://example.com/pricing",
  "statusCode": 200,
  "title": "Pricing | Example",
  "titleLength": 17,
  "metaDescriptionLength": 0,
  "h1": ["Simple pricing"],
  "canonical": "https://example.com/pricing",
  "wordCount": 412,
  "imagesMissingAlt": 3,
  "jsonLdTypes": ["Organization"],
  "responseTimeMs": 184,
  "issues": ["title-too-short", "missing-meta-description", "images-missing-alt", "missing-open-graph"]
}
```

**Key-value store `SUMMARY`** — issue counts sorted by frequency, duplicate titles, duplicate meta descriptions,
broken internal links (with up to 20 referring pages each), average response time.

### Input

- **Start URLs** — your homepage is enough; links on the same host are followed.
- **Max pages** — hard cap; you are only charged for pages actually audited.
- **Max link depth**, **Include subdomains**, **Max concurrency** (lower it for small servers).

### Pricing

Pay per event: **$0.005 per audited page** ($5 per 1,000 pages) — no subscription.
A 200-page site costs about $1.

### Use it from AI agents / MCP

Clean, flat input and one row per page with an `issues` array make it easy for an LLM agent to call via the
Apify MCP server and summarise fixes.

### Notes

- Audit sites you own or are permitted to crawl. The crawler honours robots.txt.
- JavaScript-rendered content isn't executed (HTML as served to crawlers is audited, like most SEO bots see it).
- Built and maintained with AI assistance. Issues/feature requests: use the Issues tab — replies within a few days.

# Actor input Schema

## `startUrls` (type: `array`):

Homepage (or any page) of the site(s) to audit. Links on the same host are followed.

## `maxPages` (type: `integer`):

Stop after auditing this many pages. You are charged per audited page.

## `maxDepth` (type: `integer`):

How many clicks away from the start URL to follow.

## `includeSubdomains` (type: `boolean`):

Also crawl subdomains (e.g. blog.example.com when starting at example.com).

## `maxConcurrency` (type: `integer`):

Parallel requests. Lower it to be gentle on small servers.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://crawlee.dev"
    }
  ],
  "maxPages": 100,
  "maxDepth": 10,
  "includeSubdomains": false,
  "maxConcurrency": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://crawlee.dev"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("forevertools/website-seo-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://crawlee.dev" }] }

# Run the Actor and wait for it to finish
run = client.actor("forevertools/website-seo-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://crawlee.dev"
    }
  ]
}' |
apify call forevertools/website-seo-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,forevertools/website-seo-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cHxB3MGO3uj2xfvhl/builds/sFRIaHPg9XtuJf14P/openapi.json
