# Website SEO Audit – Broken Links, Titles & Sitemap (`stevenkramp/website-seo-audit`) Actor

Technical SEO audit of whole websites: score 0–100 per page, issues with fix hints, broken links, redirect chains, titles, meta descriptions, headings, canonical, hreflang, JSON-LD, sitemap and robots.txt checks plus a site summary with the top 10 fixes. No browser.

- **URL**: https://apify.com/stevenkramp/website-seo-audit.md
- **Developed by:** [Steven Kramp](https://apify.com/stevenkramp) (community)
- **Categories:** SEO tools, Marketing, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$4.00 / 1,000 audited pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website SEO Audit – Broken Links, Titles, Sitemap & Score

Crawl any website and get a **technical SEO audit** in minutes: a **score from 0 to 100 for every page**, every issue with its **priority and a concrete fix hint**, broken internal and external links, redirect chains, duplicate titles, canonical and hreflang errors, structured data errors, sitemap and robots.txt checks – plus one **site summary with the top 10 fixes**.

**$4 per 1,000 audited pages.** You only pay for HTML pages that were actually audited. Error pages, redirects, pages blocked by bot protection or access rules and the site summary are free.

No browser, no setup: fast plain HTTP requests, polite by default (2 parallel requests, Crawl-delay honoured in full), and **robots.txt is always respected** – also for every redirect target.

### What you get for each page

| Field | Example |
|---|---|
| `url`, `finalUrl`, `statusCode`, `redirectChain` | https://example.com/old → 301 → https://example.com/new |
| `fetchStatus`, `error`, `redirectTarget` | ok · redirect · http-error · blocked · denied · error · robots-disallowed · robots-unreachable |
| `score` (0–100), `issueCount`, `issues[]` | 89 · `{code, priority, message, fixHint}` |
| `indexable`, `indexableReason` | false · noindex / canonical / http-404 |
| `title`, `titleLength`, `metaDescription`, `metaDescriptionLength` | "Pricing – Example" · 17 |
| `h1Count`, `h1`, `headingSkips[]` | 1 · `[{from: "h2", to: "h4"}]` |
| `canonical`, `canonicalSelf` | https://example.com/pricing · true |
| `hreflang[]` (with `returnLink` check), `lang` | `[{hreflang: "de", url: …, returnLink: false}]` |
| `schemaTypes[]`, `schemaErrors[]` | \["Organization", "WebSite"] · \["JSON-LD block 1: invalid JSON …"] |
| `ogTags`, `viewport` | og:title, og:description, og:image … |
| `imagesTotal`, `imagesWithoutAlt`, `wordCount` | 12 · 2 · 840 |
| `internalLinks`, `externalLinks`, `brokenLinks[]` | 48 · 6 · `[{url, statusCode, anchorText, internal}]` |
| `responseTimeMs`, `depth`, `foundVia`, `inSitemap` | 180 · 2 · link · true |

### The site summary (one item per website)

- **Average score**, pages audited, pages with issues, issues by priority and by code
- **Top 10 fixes** sorted by priority and number of affected pages, each with fix hint and example URLs
- **Duplicate titles and meta descriptions** (grouped)
- **Broken links** (internal and external) with the page they are on
- **Sitemap checks:** sitemap found (robots.txt, /sitemap.xml, index files, .gz), sitemap URLs that do not return 200, sitemap URLs blocked by robots.txt, sitemap pages that no crawled page links to
- **robots.txt state** and crawl statistics (success rate, blocked and denied pages, stop reason, links not checked)

The summary is the last item of the dataset and is also saved as `SUMMARY` in the key-value store.

### Checks

| Priority | Issue codes |
|---|---|
| High | `http-4xx`, `http-5xx`, `not-https`, `missing-title`, `broken-internal-links` |
| Medium | `missing-meta-description`, `missing-h1`, `duplicate-title`, `noindex-in-sitemap`, `redirect-chain`, `canonical-target-error`, `multiple-canonicals`, `hreflang-no-return-link`, `hreflang-invalid-code`, `hreflang-target-error`, `schema-invalid`, `broken-external-links`, `missing-viewport`, `slow-response` |
| Low | `title-too-long`, `title-too-short`, `meta-description-too-short`, `meta-description-too-long`, `duplicate-meta-description`, `multiple-h1`, `heading-skip`, `missing-canonical`, `missing-lang`, `missing-og-tags`, `images-missing-alt`, `low-word-count`, `noindex` |
| Site | `sitemap-missing`, `sitemap-urls-not-200`, `sitemap-urls-blocked-by-robots`, `sitemap-orphan-pages`, `robots-txt-missing`, `robots-txt-unreachable`, `robots-txt-no-sitemap` |

Score per page: 100 minus 20 per high, 8 per medium and 3 per low issue type. Search snippet checks (title and description length, missing description, canonical, word count) only apply to indexable pages, so intentional noindex pages are not punished twice.

### Use cases

- **SEO agencies and freelancers:** run a full audit for a new client in minutes and hand over the top 10 fixes.
- **Website launches and relaunches:** find broken links, redirect chains and missing titles before Google does.
- **Monitoring:** schedule a weekly run and watch the average score and broken links over time.
- **Content teams:** use *Only pages changed since* to check just the pages published or edited recently.
- **AI agents:** give an assistant a structured, prioritised to-do list for a website.

### How to use

1. Add one or more **start URLs** (usually the homepage). Each website is crawled on its own domain only.
2. Set **max pages per website** (default 200). Optional: include/exclude patterns, max click depth, extra sitemap URLs.
3. Run it, then export as JSON, CSV or Excel, or connect to Google Sheets, Make, Zapier or n8n.

```json
{
  "startUrls": [{"url": "https://www.example.com/"}],
  "maxPages": 200,
  "useSitemap": true,
  "checkExternalLinks": true,
  "excludePatterns": ["*/tag/*", "?replytocom="],
  "outputMode": "both"
}
```

The dataset has two ready-made views: **Pages** (score, status and key fields per page) and **Issues per page** (one row per issue with priority and fix hint).

### Pricing

Pay per event: **$0.004 per audited page** ($4 per 1,000 pages). Platform usage is included. Redirects, 4xx/5xx pages, blocked or denied pages, start URLs that could not be fetched and the site summary are not charged. If you set a maximum cost per run, the crawl stops before it.

### Use with AI agents (MCP)

This Actor works as a tool for AI assistants and agents – Claude, ChatGPT, Cursor, VS Code, n8n and other MCP clients – through Apify's hosted MCP server. Add this server URL to your client:

```
https://mcp.apify.com?tools=stevenkramp/website-seo-audit
```

Sign in with your Apify account when asked. Your agent can then call the Actor in plain language, for example: *"Audit https://www.example.com (up to 100 pages) and list the top 10 SEO fixes with the affected URLs."* – and gets clean, structured JSON back. Runs started by your agent are normal Actor runs on your Apify account at the same pay-per-event price.

### More from stevenkramp

Other Actors by the same developer – same quality standards, pay only for results:

**Search & trends**

- [Google Trends Scraper](https://apify.com/stevenkramp/google-trends-scraper) – interest over time, regions, rising queries
- [Keyword Trends Finder](https://apify.com/stevenkramp/keyword-trends-finder) – keyword ideas with trend direction
- [Google News Scraper](https://apify.com/stevenkramp/google-news-scraper) – news articles with real URLs
- [Google Images Scraper](https://apify.com/stevenkramp/google-images-scraper) – full-size image URLs
- [Google Shopping Scraper](https://apify.com/stevenkramp/google-shopping-scraper) – prices and merchants
- [Google Jobs Scraper](https://apify.com/stevenkramp/google-jobs-scraper) – job listings

**Apps**

- [Google Play Store Scraper](https://apify.com/stevenkramp/google-play-store-scraper) – Android app data and rankings
- [Google Play Reviews Scraper](https://apify.com/stevenkramp/google-play-reviews-scraper) – Play Store reviews and ratings
- [Apple App Store Scraper](https://apify.com/stevenkramp/apple-app-store-scraper) – iPhone, iPad and Mac app data
- [Shopify App Store Scraper](https://apify.com/stevenkramp/shopify-app-store-scraper) – Shopify apps and pricing plans

**Research & media**

- [arXiv Papers Scraper](https://apify.com/stevenkramp/arxiv-papers-scraper) – research papers and abstracts
- [Apple Podcasts Scraper](https://apify.com/stevenkramp/apple-podcasts-scraper) – podcasts with latest episodes

**Websites & places**

- [Germany Neighborhood Profile](https://apify.com/stevenkramp/germany-neighborhood-profile) – German neighborhood rents and vacancy

### FAQ

**Does it render JavaScript?** No. It reads the HTML the server sends, like most search engine first passes. Sites that build all content with JavaScript in the browser show little text and few links – the audit makes that visible.

**What about robots.txt?** It is always respected and cannot be switched off: disallowed URLs are never fetched – not even as the target of a redirect, because redirects are followed hop by hop and every target is checked first (they are counted in the summary). Crawl-delay is honoured in full; with a long delay the crawl stops in time before the run timeout instead of going faster. Outgoing links are only checked where the target site's robots.txt allows it. The crawler identifies itself as `StevenKrampBot` with a link to this page.

**A site blocks the crawler – what happens?** Pages behind a bot challenge (e.g. Cloudflare) are reported honestly as `fetchStatus: "blocked"` instead of empty data. Pages that refuse the crawler with 401, 403 or 429 (login, firewall, rate limit) are reported as `fetchStatus: "denied"` – not as errors of the website. Neither is charged, and neither counts as a success in `crawl.successRate`. After 30 blocked or denied requests in a row the crawl of that website stops. Every start URL gets an item, even when it cannot be fetched at all (robots.txt disallows it or is unreachable or behind a bot check, DNS error) – with `fetchStatus` and `error` saying why. The *Proxy* option can help for sites that block data center traffic.

**Which links count as broken?** 404, 410, other 4xx and 5xx responses, DNS, connection and SSL errors – also when the target site cannot even deliver its robots.txt because the domain or certificate is dead. 401/403/405/429 responses and bot challenges are not counted as broken, because the page usually works for people. Outgoing links that the target's robots.txt disallows, or whose robots.txt answers with a server error, are not fetched; they are counted in `crawl.externalLinksBlockedByRobots` and `crawl.externalLinksNotChecked`.

**Are all external links checked?** Each outgoing URL is checked once. Link checking gets a time budget per website (at least 2 minutes, or as long as the crawl itself took), so a page full of outgoing links cannot make a run explode; links left unchecked are counted in `crawl.externalLinksNotChecked`.

**How big can a crawl be?** Up to 10,000 pages per website. The audit keeps only the SEO data of each page, not its HTML, so memory stays small (tested on Apify: 2,000 pages peaked at 162 MB of the default 1 GB). The Actor also watches memory and the run timeout: it stops fetching in time (`crawl.stoppedBy`: `memory-limit` or `time-limit`) and still saves all results. For big sites raise the run timeout.

**Personal data?** The audit stores SEO data only: URLs, status codes, titles, descriptions, headings and counts – no page text. Contact links (mailto:, tel:) are ignored, and e-mail addresses (also written as name (at) domain) and phone numbers in international, German and North American formats are masked in titles, descriptions, headings and link texts.

**Something broken?** Open an issue in the Issues tab. We fix problems quickly.

# Actor input Schema

## `startUrls` (type: `array`):

Homepage (or any page) of each website to audit. Every website is crawled on its own domain only; several start URLs of the same domain are audited together.

## `maxPages` (type: `integer`):

Upper limit of internal URLs fetched per website (including redirects and error pages). You only pay for audited HTML pages.

## `maxDepth` (type: `integer`):

How many clicks away from the start URL links are followed (0 = start URLs and sitemap URLs only).

## `useSitemap` (type: `boolean`):

Find the sitemap via robots.txt, /sitemap.xml or /sitemap_index.xml (nested index files and .gz supported) and audit its URLs too. Also enables the sitemap checks in the site summary.

## `sitemapUrls` (type: `array`):

Optional sitemap URLs to read in addition to the ones found automatically.

## `lastmodAfter` (type: `string`):

Optional date (YYYY-MM-DD). Audits only the start URLs and the sitemap URLs with a lastmod on or after this date – handy for checking recently published or edited pages. Links are then not followed.

## `includePatterns` (type: `array`):

Optional. Glob patterns on the full URL, e.g. https://www.example.com/blog/\* or */docs/*. A pattern without \* matches URLs that contain it, e.g. /blog/. Start URLs are always audited.

## `excludePatterns` (type: `array`):

Optional. Same syntax as above, e.g. */tag/*, ?replytocom=, /wp-admin/.

## `checkExternalLinks` (type: `boolean`):

Check every outgoing link once (HEAD, GET if needed) and report broken ones. robots.txt of the target site is respected.

## `outputMode` (type: `string`):

pages = one item per URL, summary = one site-summary item per website, both = pages plus summary. The price per audited page is the same in every mode.

## `maxConcurrency` (type: `integer`):

Polite by default. Crawl-delay from robots.txt is always respected on top of this.

## `proxyConfiguration` (type: `object`):

Not needed for most websites. Turn on Apify Proxy if a site blocks data center traffic (some Cloudflare setups). Blocked pages are reported as blocked and not charged.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.stevenkramp.de/"
    }
  ],
  "maxPages": 200,
  "maxDepth": 10,
  "useSitemap": true,
  "checkExternalLinks": true,
  "outputMode": "both",
  "maxConcurrency": 2,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `pages` (type: `string`):

One item per crawled URL plus one site-summary item per website.

## `overview` (type: `string`):

Score, status and key SEO fields per page.

## `issues` (type: `string`):

One row per issue with priority and fix hint.

## `summary` (type: `string`):

Average score, top 10 fixes, duplicates, broken links, sitemap and robots.txt checks for each website.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.stevenkramp.de/"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("stevenkramp/website-seo-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://www.stevenkramp.de/" }] }

# Run the Actor and wait for it to finish
run = client.actor("stevenkramp/website-seo-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.stevenkramp.de/"
    }
  ]
}' |
apify call stevenkramp/website-seo-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,stevenkramp/website-seo-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KKYh9jeWKh8t3b4NJ/builds/OFcnH99muhQ7aCvoc/openapi.json
