# Website SEO Audit & Broken Link Checker (`dima_kadirovich/website-seo-audit`) Actor

Crawl any website and audit every page: titles, meta descriptions, H1s, canonicals, noindex, alt text, mixed content, structured data, speed, and broken links. 0-100 score per page and a shareable HTML report.

- **URL**: https://apify.com/dima\_kadirovich/website-seo-audit.md
- **Developed by:** [Cronexa Data Tools](https://apify.com/dima_kadirovich) (community)
- **Categories:** SEO tools, MCP servers, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Website SEO Audit & Broken Link Checker do?

It **crawls a whole website and audits every page** for technical SEO problems, then gives you:

- 🎯 **A 0–100 SEO score for every page**, sorted worst first, so you know what to fix first
- 🔗 **Every broken link**, internal and external, with the page it was found on
- 📄 **A shareable HTML report** you can send to a client or your team
- 📊 **Structured data** (JSON, CSV or Excel) for every page: title, meta description, H1s, canonical, status code, redirect chain, response time, word count, images without alt text, Open Graph tags, JSON-LD types, and more

**$4 per 1,000 pages audited. Link checks are free**, and platform usage is included.

### What does it check?

| Check | Severity |
|---|---|
| Page returns 4xx/5xx error | 🔴 Error |
| Broken internal links | 🔴 Error |
| Missing `<title>` | 🔴 Error |
| Broken external links | 🟠 Warning |
| Missing meta description | 🟠 Warning |
| Missing H1 | 🟠 Warning |
| Page set to `noindex` | 🟠 Warning |
| Missing mobile viewport | 🟠 Warning |
| Images without alt text | 🟠 Warning |
| Mixed content (http resources on https pages) | 🟠 Warning |
| Slow server response (> 2 s) | 🟠 Warning |
| Title too long / too short | 🟠 / 🔵 |
| Meta description too long / too short | 🔵 Notice |
| Multiple H1s, missing canonical, canonical to another URL | 🔵 Notice |
| Missing `lang`, missing Open Graph, no structured data | 🔵 Notice |
| Thin content (< 200 words), large HTML, redirects | 🔵 Notice |
| **Duplicate titles and meta descriptions across pages** | Site summary |

**Score:** each page starts at 100 and loses 15 points per error, 5 per warning, and 1 per notice. Pages that return an error score 0.

### Who is it for?

- **SEO agencies and freelancers**: audit client sites and send them the HTML report
- **Website owners and marketers**: find broken links and quick wins before they cost you rankings
- **Developers**: run it on a **schedule** after every release to catch broken links and missing tags
- **Site migrations**: compare before and after (status codes, redirects, canonicals)

### How do I use it?

1. Enter your website's homepage in **Websites to audit**. You can add several sites.
2. Set **Max pages to audit** (default 100).
3. Click **Start**.
4. Open the **HTML report** from the Output tab, or download the **Pages** table as CSV/Excel.

The crawler finds pages by following links **and** reading your `sitemap.xml`. It respects `robots.txt` by default, and **automatically slows down** if a website says it's getting too many requests. Pages refused that way are skipped and **not charged**.

#### Example input

```json
{
  "startUrls": [{ "url": "https://www.example.com/" }],
  "maxPages": 500,
  "checkExternalLinks": true
}
```

### Output example

One row per page:

```json
{
  "finalUrl": "https://www.djangoproject.com/weblog/2026/sep/18/proposed-change-to-dsf-voting-membership/",
  "score": 93,
  "statusCode": 200,
  "title": "Proposed change to DSF voting membership | Weblog | Django",
  "titleLength": 58,
  "metaDescription": "Posted by Jeff Triplett and the DSF Board on Sept. 18, 2026",
  "h1": [
    "Proposed change to DSF voting membership"
  ],
  "canonical": null,
  "responseTimeMs": 29,
  "wordCount": 1282,
  "internalLinks": 280,
  "externalLinks": 29,
  "brokenLinks": [
    {
      "url": "https://forum.djangoproject.com/t/voting-members-distinction-proposal/45449",
      "status": 404,
      "internal": false
    }
  ],
  "issues": [
    {
      "severity": "notice",
      "code": "canonical_missing",
      "message": "No canonical link."
    },
    {
      "severity": "warning",
      "code": "broken_external_links",
      "message": "1 broken external link(s)."
    },
    {
      "severity": "notice",
      "code": "structured_data_missing",
      "message": "No JSON-LD structured data."
    }
  ]
}
```

The **Site summary** contains the average score, issue counts, duplicate titles and descriptions, and the report link.

### How much does it cost?

| What | Price |
|---|---|
| Page audited | **$0.004** ($4 per 1,000 pages) |
| Link checks, sitemap, robots.txt, report | **Free** |
| Platform usage | **Included** |

A 100-page website costs **$0.40**. Pages skipped because the site rate-limited the crawler are free. You can set a **maximum cost per run**, and the Actor stops exactly at your budget.

### Tips

- **Auditing your own site?** You can turn off *Respect robots.txt* to include pages hidden from search engines.
- **Large sites**: use *Exclude URLs matching* (for example `/tag/` or `\?page=`) to skip endless archive pages.
- **Subdomains**: turn on *Include subdomains* to audit `blog.example.com` together with `example.com`.
- **Monitoring**: schedule weekly runs and connect a Slack or email integration to get notified about new broken links.

### Limitations

- It reads the HTML the server sends. Content created only by JavaScript is not rendered, the same as most SEO crawlers' default mode.
- Core Web Vitals (LCP, CLS) need a real browser and are not measured. Server response time is measured.
- Some websites block all bots. Links to those sites are reported as unverified, not broken, to avoid false alarms.

### Questions or problems?

Open an issue in the **Issues** tab, and it will be answered quickly.

### Use it from AI assistants (Claude, ChatGPT, Cursor)

AI agents can run this Actor as a tool through the [Apify MCP server](https://mcp.apify.com). Add this to your MCP client (Claude Desktop, Claude Code, Cursor, VS Code…) and sign in with Apify in the browser when asked:

```json
{
  "mcpServers": {
    "website-seo-audit": { "url": "https://mcp.apify.com?tools=dima_kadirovich/website-seo-audit" }
  }
}
```

Then just ask, for example:

- *"Audit example.com (up to 50 pages) and list the 5 most important SEO problems to fix first."*
- *"Find all broken links on my website and tell me which pages they are on."*

The agent fills in the input, runs the Actor and reads the results. You pay the same per-result price.

### More tools from Dima Data Tools

- [Medium Articles Scraper & Monitor](https://apify.com/dima_kadirovich/medium-articles-scraper): Medium articles by tag, author, or publication as clean Markdown, with "only new" monitoring
- [Bulk Image Downloader](https://apify.com/dima_kadirovich/bulk-image-downloader): every image from any web page, as download links or ZIP, with duplicates and icons removed
- [Website SEO Audit & Broken Link Checker](https://apify.com/dima_kadirovich/website-seo-audit): crawl a site, score every page 0–100, find broken links, and get a shareable HTML report
- [Website to Markdown Crawler for AI](https://apify.com/dima_kadirovich/website-to-markdown): any website as clean main-content Markdown, with RAG chunks, llms.txt, and a cheap "only changed pages" refresh mode
- [Website Tech Stack & Domain Lookup](https://apify.com/dima_kadirovich/tech-stack-domain-lookup): technologies, email provider, SPF/DMARC, SaaS tools, SSL expiry and WHOIS for any list of domains

# Actor input Schema

## `startUrls` (type: `array`):

Homepage (or any page) of each website to audit. The Actor finds the other pages through links and the sitemap.

## `maxPages` (type: `integer`):

Upper limit of HTML pages audited (and charged) across all websites.

## `checkExternalLinks` (type: `boolean`):

Also check links that point to other websites. Internal links are always checked. Link checks are free.

## `useSitemap` (type: `boolean`):

Find pages from the site's sitemap (declared in robots.txt, or /sitemap.xml) in addition to following links.

## `respectRobotsTxt` (type: `boolean`):

Skip pages that robots.txt disallows. Turn off only when auditing your own site.

## `includeSubdomains` (type: `boolean`):

Also crawl subdomains such as blog.example.com when auditing example.com.

## `maxDepth` (type: `integer`):

How many clicks away from the start page to crawl.

## `excludeUrlPatterns` (type: `array`):

Regular expressions. URLs that match any of them are not crawled, for example `/tag/` or `\?page=`.

## `maxLinkChecks` (type: `integer`):

Upper limit of links checked for being broken (beyond the pages already crawled).

## `proxyConfiguration` (type: `object`):

Only needed if the website blocks the crawler.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://books.toscrape.com/"
    }
  ],
  "maxPages": 100,
  "checkExternalLinks": true,
  "useSitemap": true,
  "respectRobotsTxt": true,
  "includeSubdomains": false,
  "maxDepth": 20,
  "maxLinkChecks": 2000,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `report` (type: `string`):

Shareable report: average score, issues, pages worst-first, broken links, duplicate titles.

## `pages` (type: `string`):

One row per audited page with score, issues, SEO fields and broken links.

## `summary` (type: `string`):

Average score, issue counts, duplicate titles and descriptions, broken links total.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://books.toscrape.com/"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("dima_kadirovich/website-seo-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://books.toscrape.com/" }] }

# Run the Actor and wait for it to finish
run = client.actor("dima_kadirovich/website-seo-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://books.toscrape.com/"
    }
  ]
}' |
apify call dima_kadirovich/website-seo-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dima_kadirovich/website-seo-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/SEMt6Dn5r1SOK77KI/builds/ggaEeD7XllYxYfA5E/openapi.json
