# SEO Site Audit - Technical SEO & Broken Link Checker (`rod_analytics/seo-site-audit`) Actor

Technical website quality checker for site owners: crawls a site or checks a URL list for redirect chains, duplicate titles/descriptions, broken links, missing canonicals, hreflang errors, JSON-LD problems and more.

- **URL**: https://apify.com/rod\_analytics/seo-site-audit.md
- **Developed by:** [Rod Services](https://apify.com/rod_analytics) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## SEO Site Audit - Technical SEO & Broken Link Checker

Technical website quality checker for site owners, SEO consultants and dev teams. **SEO Site Audit** crawls a website (or checks a fixed list of URLs) and reports the technical signals that affect search visibility: redirect chains, duplicate titles and meta descriptions, broken internal links, missing canonicals, hreflang errors, invalid JSON-LD and images without alt text. Try it instantly with the prefilled example - it audits the [Apify Academy](https://docs.apify.com/academy) docs and returns one row per page with a clear list of issues.

Because it runs on the Apify platform, you get a scheduled job, an API/webhook trigger, and full run history for free - no server to maintain.

### What does SEO Site Audit do?

The Actor has two modes:

- **Crawl a site** - give it one start URL, and it follows same-host links plus the site's `sitemap.xml`, respecting `robots.txt`, up to a page limit you set.
- **Check a list of URLs** - give it a fixed list of URLs and it reports HTTP status, redirect chains and response time for each one, without crawling further. This is the cheap mode for monitoring known pages (e.g. after a migration).

For every audited page it records:

- HTTP status code and the full redirect chain (hop count, final URL, redirect-loop detection)
- Response time
- Title and meta description, their length, and duplicates across the site
- H1 count, canonical tag status, `robots` meta / `X-Robots-Tag` noindex signals
- Hreflang tags, including alternates that don't link back (missing return links)
- JSON-LD blocks that fail to parse
- Images missing an `alt` attribute
- Broken internal links found on the page
- Mixed content, Open Graph tags, word count and orphan-page detection (crawl mode)

Each row also carries a structured `issues` array (`severity`, `code`, `message`) so you can filter or alert on errors vs. warnings vs. notices. A `SUMMARY` record with aggregate issue counts is written to the key-value store at the end of every run.

#### Results are saved while the audit runs

Rows are saved to the dataset in **batches of 25 pages** while the crawl is still going, not only at the end. If a long run hits its timeout, is aborted or the platform migrates it, you keep every page audited up to that point, and an aborted run still writes a partial `SUMMARY` (`complete: false`, `stoppedBecause: "aborted"`). After a migration the run resumes and does not save or charge the same page twice. When your **Max cost per run** is reached, the audit stops and keeps exactly the pages that fit the budget (`stoppedBecause: "max-cost-reached"`).

#### Duplicates and orphan pages in incremental mode

Duplicate titles and meta descriptions need the whole site, but early rows are saved before later pages are seen. So:

- The **first** page with a given title or meta description is saved without a duplicate flag. Every **later** page with the same value gets a `duplicate-title` / `duplicate-meta-description` warning that names the first page.
- The complete groups, including the first page, are in `SUMMARY.duplicateTitles` and `SUMMARY.duplicateMetaDescriptions` (`value`, `count`, `urls`). `SUMMARY.issuesByCode` counts every page in a group.
- `inlinks` on a row is the number of internal links to that page found **so far** when the row was saved. Orphan status is only final at the end: rows of sitemap pages with no links yet have `isOrphan: null`, and the final list is `SUMMARY.orphanPageUrls`.

### Use cases

- **Pre-launch or post-migration QA** - catch broken links, redirect loops and missing canonicals before (or right after) a site relaunch.
- **Ongoing technical SEO monitoring** - schedule a crawl weekly and alert on new errors via Apify's webhook/integration options.
- **Duplicate content audits** - find pages sharing the same title or meta description, a common but easy-to-miss ranking issue.
- **International SEO checks** - verify hreflang alternates actually link back to each other.
- **Bulk redirect/status checks** - use URL-list mode to verify a list of URLs (e.g. from an old sitemap or a spreadsheet of legacy pages) still resolve correctly after a migration.

### How to use SEO Site Audit

1. Click **Try for free** (or **Start**) and open the **Input** tab.
2. Choose a mode: **Crawl a site** (set a **Start URL**) or **Check a list of URLs** (paste your URLs).
3. Optionally adjust **Max pages**, include/exclude URL patterns, or the proxy configuration.
4. Click **Start** and wait for the run to finish.
5. Open the **Output** tab to browse issues per page, or download the dataset as JSON/CSV/Excel.

### Input

Key input fields (see the **Input** tab for the full form and tooltips):

| Field | Description | Default |
| --- | --- | --- |
| `mode` | `crawl` or `urlList` | `crawl` |
| `startUrl` | Crawl mode: where the crawl begins | `https://docs.apify.com/academy` |
| `urlList` | URL-list mode: URLs to check | - |
| `maxPages` | Max pages (crawl) / URLs (list) to process | `20` |
| `respectRobotsTxt` | Skip pages disallowed by `robots.txt` | `true` |
| `includeGlobs` / `excludeGlobs` | Restrict crawl scope by glob pattern | `[]` |
| `maxConcurrency` | Parallel requests | `5` |
| `requestTimeoutSecs` | Per-request timeout | `30` |
| `userAgent` | User agent sent with requests | SEO Site Audit bot UA |
| `proxyConfiguration` | Apify datacenter proxy or your own proxy URLs (no residential) | disabled |

Example input for crawl mode:

```json
{
    "mode": "crawl",
    "startUrl": "https://docs.apify.com/academy",
    "maxPages": 20
}
```

### Output

One dataset item per audited page. Simplified example:

```json
{
    "url": "https://docs.apify.com/academy",
    "finalUrl": "https://docs.apify.com/academy",
    "mode": "crawl",
    "statusCode": 200,
    "finalStatusCode": 200,
    "redirectCount": 0,
    "redirectLoop": false,
    "responseTimeMs": 177,
    "title": "Apify Academy | Academy | Apify Documentation",
    "titleLength": 45,
    "metaDescription": "Learn everything about web scraping and automation...",
    "h1Count": 1,
    "canonicalUrl": "https://docs.apify.com/academy",
    "canonicalStatus": "self",
    "hreflangCount": 2,
    "jsonLdCount": 0,
    "imagesMissingAlt": 6,
    "brokenLinkCount": 0,
    "issueCount": 1,
    "errorCount": 0,
    "warningCount": 0,
    "issues": [
        { "severity": "notice", "code": "image-missing-alt", "message": "6 image(s) are missing an alt attribute" }
    ]
}
```

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. A `SUMMARY` record (issue counts by severity/code, duplicate title and meta description groups with their URLs, orphan page URLs, broken-link totals, average response time, whether the run completed) is written to the run's key-value store under the key `SUMMARY`.

#### Data table

| Field | Meaning |
| --- | --- |
| `statusCode` / `finalStatusCode` | Status of the first request / after following redirects |
| `redirectChain` | Every hop: URL, status, `Location`, timing |
| `canonicalStatus` | `self`, `other`, `missing`, `invalid`, or `broken` |
| `hreflang` | Each alternate with `returnLink: true/false/null` |
| `brokenLinks` | Internal links on this page that returned an error or restricted status |
| `inlinks` | Internal links pointing to this page found by the time the row was saved |
| `isOrphan` | `false` when the page has inlinks or was not found via the sitemap; `null` when undecided at save time (see `SUMMARY.orphanPageUrls`) |
| `issues` | `{ severity, code, message }[]` - the full list of findings for the page |

### Pricing

This Actor uses **Pay-Per-Event** pricing - you only pay for what gets checked, no compute-unit guessing:

| Event | When it's charged | Price |
| --- | --- | --- |
| `apify-actor-start` | Once per run | $0.001 |
| `page-audited` | Once per page in **crawl** mode (full SEO analysis) | $0.006 |
| `url-checked` | Once per URL in **urlList** mode (status/redirect check only) | $0.0005 |

Rough cost examples: auditing 1,000 pages in crawl mode costs about **$6.00**; checking 1,000 URLs in list mode costs about **$0.50**. Both include one flat $0.001 start fee per run. See [`PRICING.md`](./PRICING.md) in the source repository for the measured compute cost behind these prices and the resulting margin.

### FAQ

**Does this Actor modify my website?** No. It only makes read-only HTTP requests (`GET`/`HEAD`); it never submits forms or writes anything.

**Does it respect `robots.txt`?** Yes, by default, in crawl mode. Pages disallowed for the configured user agent are not fetched. You can disable this with `respectRobotsTxt: false` if you are auditing a staging site you own.

**Does it check external links?** It checks and reports **internal** broken links (same host). External link targets are counted but not fetched, to keep runs fast and cheap.

**Why is `urlList` mode so much cheaper?** It only checks status codes, redirects and response time - it doesn't download and parse full page content, so there's no title/meta/canonical/hreflang/JSON-LD analysis in that mode.

**What counts as "same site"?** `example.com` and `www.example.com` are treated as the same site so crawling isn't blocked by a `www` redirect.

**Which proxies can I use?** No proxy (the default), Apify datacenter proxy, or your own proxy URLs. Residential and SERP proxies are not supported. A run that asks for them stops at the start with a clear message and does no work.

**Something looks wrong or missing?** Please open an issue on the Actor's Issues tab with the input you used - happy to take a look. Custom variations of this audit (extra checks, different scoring, CMS-specific rules) are also available on request.

# Actor input Schema

## `mode` (type: `string`):

Crawl a site starting from one URL, or check a fixed list of URLs (status/redirects only, no page-level SEO analysis).

## `startUrl` (type: `string`):

Where the crawl begins. Only pages on the same host are followed. Required in "Crawl a site" mode.

## `urlList` (type: `array`):

URLs to check one by one. Required in "Check a list of URLs" mode. Each URL only gets a status/redirect check, not a full SEO audit.

## `maxPages` (type: `integer`):

Crawl mode: maximum number of pages to audit. URL list mode: maximum number of URLs to check (extra ones are ignored).

## `respectRobotsTxt` (type: `boolean`):

Crawl mode only. When enabled, pages disallowed by robots.txt for this Actor's user agent are not fetched.

## `includeGlobs` (type: `array`):

Crawl mode only. If set, only URLs matching at least one of these glob patterns are crawled, e.g. "https://example.com/blog/\*\*".

## `excludeGlobs` (type: `array`):

Crawl mode only. URLs matching any of these glob patterns are skipped, e.g. "\*\*/\*.pdf".

## `maxConcurrency` (type: `integer`):

Maximum number of pages/URLs processed at the same time.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for a response before treating a request as failed.

## `userAgent` (type: `string`):

User agent string sent with every request, and used to evaluate robots.txt rules.

## `proxyConfiguration` (type: `object`):

Optional. Use Apify datacenter proxy or your own proxy URLs if the target site blocks direct requests. Residential and SERP proxies are not supported.

## Actor input object example

```json
{
  "mode": "crawl",
  "startUrl": "https://docs.apify.com/academy",
  "urlList": [],
  "maxPages": 20,
  "respectRobotsTxt": true,
  "includeGlobs": [],
  "excludeGlobs": [],
  "maxConcurrency": 5,
  "requestTimeoutSecs": 30,
  "userAgent": "Mozilla/5.0 (compatible; SeoSiteAuditBot/1.0; +https://apify.com/apify/seo-site-audit)",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "crawl",
    "startUrl": "https://docs.apify.com/academy",
    "maxPages": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("rod_analytics/seo-site-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "crawl",
    "startUrl": "https://docs.apify.com/academy",
    "maxPages": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("rod_analytics/seo-site-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "crawl",
  "startUrl": "https://docs.apify.com/academy",
  "maxPages": 20
}' |
apify call rod_analytics/seo-site-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,rod_analytics/seo-site-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/iSKGEAPr7beDgfodf/builds/xbfhJ9tgRsUkieKig/openapi.json
