# SEO Audit Crawler – Broken Links, Meta Tags, Redirects (`gazidev/seo-audit-crawler`) Actor

Technical SEO site audit and broken link checker. Finds broken internal and external links with source pages, redirect chains, title/meta issues, H1, canonical, hreflang, noindex, missing alt, duplicates, thin content and orphan sitemap URLs, plus a 0-100 site score.

- **URL**: https://apify.com/gazidev/seo-audit-crawler.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $8.00 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## SEO Audit Crawler – Broken Links, Meta Tags, Redirects

Give it your website. It crawls the site and runs a **full technical SEO audit** with a **broken link check** in one run. You get one row per page with every issue found, one row per external link checked, and a **site summary with a 0–100 SEO score** and the **top 10 fixes**.

- **Broken link checker:** finds broken internal and external links (404, 410, 5xx, DNS and timeout errors) and lists the **pages that link to them**, so you know exactly where to fix them.
- **Redirect checker:** records every hop of each redirect chain and flags chains longer than one hop, redirect loops and internal links that point to redirects.
- **On-page SEO:** title and meta description (missing, too long, too short, duplicate), H1 count, canonical tag (missing or pointing elsewhere), noindex (meta robots and `X-Robots-Tag`), hreflang, `<html lang>`, Open Graph, missing image alt text, thin content, mixed content on HTTPS, JSON-LD schema types, response time and page size.
- **Sitemap checks:** reads sitemap.xml from robots.txt (indexes and gzip included), finds **orphan pages** that are in the sitemap but not linked anywhere, sitemap URLs that redirect or fail, and noindex pages listed in the sitemap.
- **Issues-only mode** for scheduled monitoring: only pages with errors or warnings and only broken links are saved.
- **Cheap:** $8 per 1,000 audited pages. That is 5× cheaper than the most popular SEO audit Actor. Error pages (404, 5xx) cost nothing.
- **Polite:** HTTP only (no browser), robots.txt and Crawl-delay respected, 3 parallel requests per site by default, an honest User-Agent, and at most 1 request per second per external host.

### What can I use it for?

- **Website audit before a launch or migration:** catch broken links, redirect chains and missing tags before Google does.
- **Monthly SEO health check:** schedule the Actor with `outputMode: issuesOnly` and get only the new problems.
- **Agencies and freelancers:** produce a client site audit with a score, issue counts and the top fixes in minutes.
- **Content teams:** find duplicate titles and descriptions, thin pages and pages missing H1 or alt text.
- **Link rot monitoring:** find outgoing links to dead pages on blogs, docs and resource pages.
- **AI agents:** let an agent audit a site and propose fixes (see below).

### Input

| Field | Description |
|---|---|
| `startUrls` | One URL per site. The crawl stays on the same host (www and non-www count as the same site). |
| `maxPages` | Max HTML pages audited per site (default 50, max 10,000). Breadth-first, so the most important pages come first. |
| `checkExternalLinks` | Check outgoing links (HEAD, GET fallback, 8 s timeout). Default on. |
| `maxExternalLinks` | Max unique external URLs checked per site (default 500). The most-linked ones are checked first. |
| `useSitemap` | Use sitemap.xml to find more pages and run the sitemap checks. Default on. |
| `outputMode` | `allPages` (default) or `issuesOnly` |
| `respectRobots`, `maxConcurrency`, `requestTimeoutSecs`, `userAgent` | Advanced |

```json
{
  "startUrls": [{ "url": "https://www.example.com" }],
  "maxPages": 200,
  "checkExternalLinks": true,
  "maxExternalLinks": 500,
  "useSitemap": true,
  "outputMode": "allPages"
}
```

### Output

The Output tab has four tables: **Issues** (one row per issue: URL, severity, code, details), **Pages**, **External links** and the **Summary**. The summary is also saved in the key-value store as `SUMMARY`.

Sample summary (crawler-test.com, a site built with SEO problems on purpose, 40 pages):

| Site | Score | Grade | Pages | Errors | Warnings | Notices | Broken internal URLs | Broken external links |
|---|---|---|---|---|---|---|---|---|
| crawler-test.com | 71 | C | 40 | 8 | 60 | 132 | 5 | 2 |
| books.toscrape.com | 83 | B | 10 | 10 | 12 | 20 | 0 | 0 |

Page row (shortened):

```json
{
  "type": "page",
  "url": "https://crawler-test.com/links/broken_links_internal",
  "statusCode": 200,
  "redirectChain": [],
  "responseTimeMs": 163,
  "title": "Broken Links Internal",
  "titleLength": 21,
  "metaDescriptionLength": 40,
  "h1Count": 1,
  "wordCount": 30,
  "canonicalStatus": "missing",
  "indexable": true,
  "jsonLdTypes": [],
  "inSitemap": false,
  "linkedFrom": ["https://crawler-test.com/"],
  "issues": [
    {
      "severity": "error",
      "code": "BROKEN_INTERNAL",
      "message": "Page links to internal URLs that return 4xx/5xx or fail",
      "details": {
        "count": 5,
        "links": [{ "url": "https://crawler-test.com/links/not_found/foo1", "statusCode": 404 }]
      }
    },
    { "severity": "warn", "code": "THIN_CONTENT", "message": "Fewer than 200 words of visible text", "details": { "wordCount": 30 } }
  ],
  "errorCount": 1,
  "warningCount": 1
}
```

External link row:

```json
{
  "type": "externalLink",
  "url": "https://invalid.crawler-test.com/",
  "statusCode": null,
  "ok": false,
  "botBlocked": false,
  "error": "DNS resolution failed",
  "sourceCount": 1,
  "sourcePages": ["https://crawler-test.com/"]
}
```

The summary row has `score`, `grade`, `issueCounts` (pages per issue code), `topFixes` (code, severity, affected pages, example URLs), `statusCodes`, `brokenInternalExamples`, `sitemapUrlCount`, `crawlComplete` and `stoppedReason`. See `SAMPLE_OUTPUT.json` for full rows.

#### Issue codes

| Severity | Codes |
|---|---|
| error | `HTTP_4XX`, `HTTP_5XX`, `FETCH_FAILED`, `REDIRECT_LOOP`, `BROKEN_INTERNAL`, `BROKEN_EXTERNAL`, `TITLE_MISSING`, `MIXED_CONTENT` |
| warn | `META_DESC_MISSING`, `TITLE_TOO_LONG` (>60), `META_DESC_TOO_LONG` (>160), `H1_MISSING`, `MULTIPLE_H1`, `CANONICAL_TO_OTHER`, `NOINDEX`, `REDIRECT_CHAIN` (>1 hop), `THIN_CONTENT` (<200 words), `DUPLICATE_TITLE`, `DUPLICATE_META_DESC`, `IMG_MISSING_ALT`, `IN_SITEMAP_NON_200`, `NOINDEX_IN_SITEMAP`, `SLOW_RESPONSE` (>3 s) |
| info | `TITLE_TOO_SHORT`, `META_DESC_TOO_SHORT`, `CANONICAL_MISSING`, `INTERNAL_REDIRECT_LINK`, `IN_SITEMAP_NOT_LINKED`, `HREFLANG_NO_SELF`, `LANG_MISSING`, `OG_MISSING`, `LARGE_PAGE` |

**Score:** 100 minus a penalty for each issue code. Errors weigh 10, warnings 4 and notices 0.5, scaled by the share of pages affected. 90+ = A, 75+ = B, 60+ = C, 40+ = D.

### Pricing

Pay per event: you only pay for what was audited.

| Event | Price |
|---|---|
| Page audited (HTML page, status 200) | $0.008 ($8 per 1,000) |
| External link checked (unique URL) | $0.0003 ($0.30 per 1,000) |
| Site summary (one per site) | $0.01 |

Free: 404/5xx pages, redirects, non-HTML files and internal link status checks. Example: a 200-page site with 300 external links costs 200 × $0.008 + 300 × $0.0003 + $0.01 = **$1.70**.

| Actor | Price per 1,000 pages |
|---|---|
| smart-digital/complete-seo-audit-tool | $40 |
| automation-lab/website-lighthouse-seo-audit | $14 |
| logiover/website-seo-audit-crawler | $10 |
| **This Actor** | **$8**, plus broken external links with source pages, a summary score and an issues-only mode |

Set **Maximum cost per run** in the run options to cap your spend. The Actor stops cleanly when the limit is reached and always keeps enough budget to deliver the site summary.

### FAQ

**Does it render JavaScript?** No. It reads the HTML the server sends, which is what most technical SEO checks need, and is fast and cheap. Sites that build all content in the browser will show as thin content.

**Will it crawl other websites?** No. It only crawls the host you give it (www and non-www count as the same). External links get one HEAD request (or GET if HEAD is refused) to read their status.

**Why is an external link marked `botBlocked`?** Some sites (LinkedIn, G2, Cloudflare-protected pages) answer automated requests with 401, 403, 429 or 999. Those links usually work in a browser, so they are reported with `ok: true, botBlocked: true` and do not count as broken.

**Does it respect robots.txt?** Yes, by default. URLs disallowed for `SEOAuditCrawler` (or `*`) are skipped and Crawl-delay is honored. Turn `respectRobots` off only for your own site.

**Can I monitor my site every week?** Yes. Create a schedule with `outputMode: issuesOnly` and connect a Slack, email or webhook integration.

**What does "orphan page" mean here?** A URL that is in your sitemap but is not linked from any crawled page (`IN_SITEMAP_NOT_LINKED`). For a reliable result, set `maxPages` high enough to crawl the whole site (see `crawlComplete` in the summary).

### Use with AI agents / Apify MCP

The Actor works as a tool in the [Apify MCP server](https://mcp.apify.com) for Claude, ChatGPT, Cursor and other MCP clients. Ask something like *"Audit example.com (100 pages) and list the top 10 SEO fixes"*, and the agent calls the Actor and reads the `SUMMARY` record and the Issues view. Through the API:

```bash
curl -X POST "https://api.apify.com/v2/acts/gazidev~seo-audit-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"startUrls":[{"url":"https://www.example.com"}],"maxPages":100,"outputMode":"issuesOnly"}'
```

### Related Actors

- **Sitemap URL Extractor** (`gazidev/sitemap-url-extractor`): every URL from a site's sitemaps, with an optional status check.
- **Website to Markdown** (`gazidev/website-to-markdown`): turn pages into clean Markdown for LLMs and RAG.

# Actor input Schema

## `startUrls` (type: `array`):

One URL per site, e.g. `https://www.example.com`. The crawler stays on the same host (www and non-www count as the same site) and starts from this page. `https://` is added if missing.

## `maxPages` (type: `integer`):

Maximum number of HTML pages audited (and charged) per site. Pages are crawled breadth-first, so the most important (shallow) pages come first.

## `checkExternalLinks` (type: `boolean`):

Check every unique outgoing link (HEAD request, GET fallback, 8 s timeout, max 1 request per second per host) and report broken ones with the pages that link to them.

## `maxExternalLinks` (type: `integer`):

Upper limit of unique external URLs checked per site. The most-linked URLs are checked first. 0 = do not check external links.

## `useSitemap` (type: `boolean`):

Read sitemaps listed in robots.txt (or /sitemap.xml) to find more pages, flag sitemap URLs that redirect or error, and find pages that are in the sitemap but not linked from any crawled page (orphans).

## `outputMode` (type: `string`):

`allPages`: one row per page plus external link rows and the site summary. `issuesOnly`: only pages with errors or warnings, only broken external links, and the summary. Good for scheduled monitoring.

## `respectRobots` (type: `boolean`):

Skip URLs disallowed by robots.txt and honor Crawl-delay (up to 10 s). Turn off only for a site you own.

## `maxConcurrency` (type: `integer`):

How many pages of one site are fetched at the same time. Keep it low to stay polite to the server. Forced to 1 when robots.txt sets a Crawl-delay.

## `requestTimeoutSecs` (type: `integer`):

Timeout for one page request. Failed requests are retried once.

## `userAgent` (type: `string`):

Optional custom User-Agent header. Default: `Mozilla/5.0 (compatible; SEOAuditCrawler/0.1; +https://apify.com/gazidev/seo-audit-crawler)`. robots.txt rules are matched for `SEOAuditCrawler` (falling back to `*`).

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://books.toscrape.com"
    }
  ],
  "maxPages": 3,
  "checkExternalLinks": true,
  "maxExternalLinks": 20,
  "useSitemap": true,
  "outputMode": "allPages",
  "respectRobots": true,
  "maxConcurrency": 3,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `issues` (type: `string`):

No description

## `pages` (type: `string`):

No description

## `links` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://books.toscrape.com"
        }
    ],
    "maxPages": 3,
    "maxExternalLinks": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/seo-audit-crawler").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://books.toscrape.com" }],
    "maxPages": 3,
    "maxExternalLinks": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("gazidev/seo-audit-crawler").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://books.toscrape.com"
    }
  ],
  "maxPages": 3,
  "maxExternalLinks": 20
}' |
apify call gazidev/seo-audit-crawler --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/seo-audit-crawler"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ghsHRS6ZypxEkwzPf/builds/s37cQigPga36VIQFu/openapi.json
