# SEO Audit Crawler: Site Score, Broken Links, Duplicate Titles (`swiftkit/seo-audit`) Actor

Crawl a website and audit every page for SEO: 0-100 score, titles, meta descriptions, H1s, canonicals, noindex, alt text, thin content, slow pages, broken links, duplicate titles. Free site summary with top issues. $2 per 1,000 pages.

- **URL**: https://apify.com/swiftkit/seo-audit.md
- **Developed by:** [SwiftKit](https://apify.com/swiftkit) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## SEO Audit Crawler: Site Score, Broken Links, Duplicate Titles

Give it a website and it crawls the pages and checks **every page for the SEO basics**, then gives
each page a **0–100 score** and the site a **free summary** of what to fix first.

**Checked on every page:**

- **Title:** missing, too long (over 60 characters), too short, duplicated on other pages
- **Meta description:** missing, too long or short, duplicated
- **Headings:** missing H1, more than one H1, H2 count
- **Indexing:** `noindex` in meta robots or the `X-Robots-Tag` header, canonical URL pointing elsewhere or missing
- **Broken links:** links to pages that return 4xx/5xx or don't load (internal by default, external optional)
- **Content:** word count and thin pages (under 200 words), images without alt text
- **Technical:** response time, HTML size, redirects, plain HTTP, mobile viewport tag, `lang` attribute
- **Extras:** structured data types (JSON-LD), Open Graph tags, hreflang count, internal and external link counts

**Site summary (free):** average score, issues ranked by how many pages they affect, every broken
link with the page it was found on, duplicate titles and descriptions grouped, non-indexable pages,
the slowest and lowest-scoring pages, and whether robots.txt and a sitemap exist.

No browser, no login, no API key: it reads pages the way a search engine bot does, gently
(a few requests at a time) and respecting robots.txt.

### Who it's for

- **SEO consultants and agencies:** a quick technical audit for a client or a prospect.
- **Site owners and marketers:** find broken links and missing titles before Google does.
- **Developers:** check a site after a migration or redesign; schedule it to catch regressions.

### Input

| Option | Default | What it does |
|---|---|---|
| Website | – | Start page(s); links on the same site are followed |
| Max pages | 100 | Stop after this many audited pages |
| Only / skip URLs containing | – | Limit the audit to e.g. `/blog/` |
| Check for broken links | on | Also check linked pages that weren't audited (up to 2,000) |
| Include links to other sites | off | Broken-link check for outgoing links too |
| Respect robots.txt | on | Skip pages the site asks bots not to visit |
| Pages at once | 4 | Keep it low to be gentle |

### Output

One row per page:

```json
{
  "url": "https://crawlee.dev/js/docs/guides",
  "status": "ok",
  "score": 86,
  "errors": 0,
  "warnings": 2,
  "title": "Guides | Crawlee for JavaScript · Build reliable crawlers. Fast.",
  "titleLength": 64,
  "indexable": true,
  "brokenLinks": [],
  "issues": [
    { "code": "title_too_long", "severity": "warning", "message": "Title is longer than 60 characters and may be cut off in search results.", "detail": 64 },
    { "code": "duplicate_title", "severity": "warning", "message": "Another page has the same title.", "detail": 10 },
    { "code": "duplicate_meta_description", "severity": "notice", "message": "Another page has the same meta description.", "detail": 15 }
  ]
}
```

Plus one `site_summary` row. Pages that return an error appear as `http_error` rows (not charged)
and in the broken-link list of the pages that link to them.

**Scoring:** each page starts at 100. Each error costs 15 points, each warning 6 and each notice 2.
It's a simple, transparent checklist score, not a guess at Google's ranking.

### Pricing

You pay **per audited page**. The site summary and failed pages are free. See the Pricing tab.

### Limits, honestly

- Pages are read as HTML without running JavaScript. Sites that build everything in the browser
  (some single-page apps) will show few links and thin content.
- It checks on-page and technical basics. It doesn't measure backlinks, rankings or keyword
  volumes. For speed and Core Web Vitals, use [Lighthouse Audit](https://apify.com/swiftkit/lighthouse-audit).
- Sites behind a bot check or a login can't be audited.

### More tools from SwiftKit

- [Lighthouse Audit](https://apify.com/swiftkit/lighthouse-audit): page speed, Core Web Vitals, accessibility and SEO scores in bulk
- [Sitemap Extractor](https://apify.com/swiftkit/sitemap-urls): every URL from a site's sitemaps, with broken-link check
- [Tech Stack Detector](https://apify.com/swiftkit/tech-stack): what a website is built with
- [Website to Markdown for AI](https://apify.com/swiftkit/web-to-markdown): clean page content for LLMs and RAG

### Questions?

Open an issue on the Issues tab.

# Actor input Schema

## `startUrls` (type: `array`):

Home page (or any start page) of the site to audit. Links on the same site are followed.

## `maxPages` (type: `integer`):

Stop after auditing this many pages.

## `includePatterns` (type: `array`):

Audit only URLs matching any of these texts or regular expressions, e.g. /blog/

## `excludePatterns` (type: `array`):

Skip URLs matching any of these, e.g. /tag/ or ?page=

## `checkLinks` (type: `boolean`):

Also check links to pages that weren't audited (up to 2,000 per run).

## `checkExternalLinks` (type: `boolean`):

Broken-link check for outgoing links too. Slower.

## `respectRobotsTxt` (type: `boolean`):

Skip pages the site asks bots not to visit.

## `maxConcurrency` (type: `integer`):

Parallel requests. Keep it low to be gentle with the site.

## Actor input object example

```json
{
  "startUrls": [
    "https://crawlee.dev"
  ],
  "maxPages": 100,
  "checkLinks": true,
  "checkExternalLinks": false,
  "respectRobotsTxt": true,
  "maxConcurrency": 4
}
```

# Actor output Schema

## `results` (type: `string`):

One row per audited page plus a site summary.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://crawlee.dev"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("swiftkit/seo-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": ["https://crawlee.dev"] }

# Run the Actor and wait for it to finish
run = client.actor("swiftkit/seo-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://crawlee.dev"
  ]
}' |
apify call swiftkit/seo-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,swiftkit/seo-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/iXgU8RIvtMOucacXa/builds/ljBW7uwMTBItLVGh9/openapi.json
