# SEO Audit & Broken Link Checker: Full Site Crawl (`digital_influx/seo-audit`) Actor

Crawl any website and get one row per page with its SEO issues: titles, meta descriptions, H1, noindex, canonical, hreflang, structured data, missing alt, broken links, redirects, duplicates and orphan pages. Monitor new and fixed issues. Respects robots.txt.

- **URL**: https://apify.com/digital\_influx/seo-audit.md
- **Developed by:** [Bruno Petrelli](https://apify.com/digital_influx) (community)
- **Categories:** SEO tools, Developer tools, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $20.00 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## SEO Audit & Broken Link Checker: Full Site Crawl

Give it a website and get back **one row per page with its SEO issues**, plus **one summary row per website**. The Actor follows the site's internal links and its sitemap, the way a search engine does, and checks:

- **Titles and meta descriptions:** missing, too long, too short, several `<title>` tags, and **duplicates across the site**.
- **Headings:** missing H1 or more than one H1.
- **Indexing:** `noindex` in meta robots or in the `X-Robots-Tag` header, canonical pointing to another page, hreflang, `lang` attribute.
- **Links:** **broken internal and external links** (404, 410, 5xx, dead domains) with the page they are on, redirects, and **orphan pages** that are in the sitemap but that no page links to.
- **Content:** word count (thin pages), **images without alt text**, JSON-LD structured data types and invalid JSON-LD blocks.
- **Technical:** status code, redirect chain, response time, page size, mixed content (http resources on https pages), mobile viewport.
- **Social previews:** Open Graph title, description and image, Twitter card.

**Respects robots.txt** on every host it reads, including the sites your links point to. No browser and no proxies: plain HTTP requests with an honest User-Agent (`seo-audit`).

### Use it for

- **SEO audits for clients:** a complete, sortable list of issues per page in minutes, exportable to CSV or Excel.
- **Broken link checks** before a launch or a migration, or every month on a schedule.
- **Site migrations:** find redirects, chains and pages that fell out of the sitemap.
- **Content inventories:** every page with its title, description, H1, word count and canonical.
- **AI agents:** plain JSON per page, callable through the Apify MCP server.

### Input

```json
{
  "websites": ["userpilot.com", "https://stripe.com/blog/"],
  "maxPagesPerSite": 200,
  "checkExternalLinks": true
}
```

### Output

Two kinds of rows. Use the **Pages** and **Site summaries** views of the dataset, or filter on `type`.

A page (a real result from a small business site on 2026-09-25; the address was replaced):

```json
{
  "type": "page",
  "url": "https://example.com/academy",
  "statusCode": 200,
  "depth": 1,
  "responseTimeMs": 2653,
  "title": "B2B Breakthrough Academy",
  "titleLength": 24,
  "metaDescriptionLength": 152,
  "h1": ["Become the strategic marketer your business needs.", ""],
  "h1Count": 2,
  "h2Count": 17,
  "canonical": "https://example.com/academy",
  "canonicalIsSelf": true,
  "noindex": false,
  "lang": "",
  "viewport": true,
  "wordCount": 2165,
  "imagesWithoutAlt": 0,
  "internalLinksCount": 8,
  "externalLinksCount": 13,
  "inlinks": 6,
  "inSitemap": true,
  "brokenLinks": [{ "url": "https://example.thrivecart.com/checkout/", "status": 404, "error": null }],
  "issues": ["h1-multiple", "lang-missing", "broken-links"]
}
```

A site summary (free) has `pagesCrawled`, `pagesWithIssues`, `issueCounts` (how many pages have each issue), `duplicateTitles`, `duplicateDescriptions`, `brokenLinks`, `blockedLinks`, `robotsTxt`, `sitemaps`, `sitemapUrlCount`, `blockedByRobots`, `notCrawled` (pages found beyond your page limit) and `avgResponseTimeMs`.

#### Issue codes

| Code | Meaning |
|---|---|
| `title-missing`, `title-too-long`, `title-too-short`, `title-multiple`, `title-duplicate` | Title absent, over 60 characters, under 10, more than one `<title>`, or shared with another page |
| `description-missing`, `description-too-long`, `description-too-short`, `description-duplicate` | Meta description absent, over 160 characters, under 50, or shared |
| `h1-missing`, `h1-multiple` | No H1, or more than one |
| `noindex` | The page asks search engines not to index it |
| `canonical-to-other-page` | The canonical URL is another page |
| `lang-missing`, `viewport-missing` | No `lang` on `<html>` (or empty), no mobile viewport |
| `images-missing-alt` | `<img>` without an `alt` attribute (`alt=""` counts as decorative, not an issue) |
| `thin-content` | Under 200 words |
| `structured-data-invalid-json` | A JSON-LD block that is not valid JSON |
| `mixed-content` | http:// images, scripts, iframes or stylesheets on an https page |
| `slow-response` | Over 3 seconds to download the page |
| `broken-links` | The page links to a URL that answers 404, 410, 5xx or does not resolve |
| `redirected` | The URL redirects (the target is audited as its own row) |
| `http-error`, `unreachable` | The page answers 4xx/5xx, or the server cannot be reached |
| `orphan-page` | Found only in the sitemap: no crawled page links to it |

The length limits follow what Google usually shows in results (about 60 characters of title and 155 to 160 of description). They are guidelines, not ranking rules.

### Monitoring: what changed since the last audit

Schedule the audit (weekly, monthly) with **Compare with the previous audit** on, and each run tells you what moved:

- every page row gets `changes`: `newIssues` (for example `["broken-links"]` after someone deleted a page it links to), `fixedIssues`, and `newPage: true` for pages the last audit did not have;
- the site row gets `changes` too: how many issues are new and how many were fixed, by issue code (`issuesAdded`, `issuesFixed`), the new pages and the pages that are gone (`missingPages`: removed, redirected elsewhere or no longer linked; `null` when the page limit cut the crawl, because pages left out this time are not gone), with the date of the audit it was compared with.

The first audit has nothing to compare with, so `changes` is `null`. The last audit of each site is kept in a key-value store named `seo-audit-state` in your own account; a *Memory name* keeps separate histories (for example one per client).

### Pricing

Pay per event: **one event per page audited**. The site summary rows are free. Set a maximum cost for the run and the Actor stops cleanly when it is reached.

### Good to know

- Only pages on the website's own host are crawled (`www.` and the bare domain count as the same site). Files such as PDFs and images are not audited as pages, but links to them are checked.
- Links that answer 401, 403, 429, LinkedIn's 999 or a non-standard 4xx code (such as the 419 Hacker News sends to scripts) are counted as `blockedLinks`, not as broken: the server refused the checker, which does not mean the page is gone. So are links that answer 405 (they exist but only take other methods, like an API endpoint).
- The Actor reads the HTML the server sends. Content that only appears after JavaScript runs is not seen (for example, images a page builder draws with JavaScript).
- Crawling is gentle by default: 3 requests at a time per website.

### Support

A false positive or a check you need? Open an issue on the Actor's Issues tab with the page URL. Issues get an answer within a few days.

# Actor input Schema

## `websites` (type: `array`):

One per line: a domain (example.com) or the URL where the crawl starts. Each website is audited separately, following its internal links and its sitemap.

## `maxPagesPerSite` (type: `integer`):

Pages audited per website (each one is charged). Pages found beyond this are counted in notCrawled.

## `useSitemap` (type: `boolean`):

Read the sitemaps listed in robots.txt (or /sitemap.xml) to find pages no link points to (orphan pages) and to mark which pages are in the sitemap.

## `checkExternalLinks` (type: `boolean`):

Also check links that point to other websites (a HEAD request each). Internal links are always checked.

## `maxExternalLinksPerSite` (type: `integer`):

Unique external links checked per website.

## `compareWithPreviousRun` (type: `boolean`):

For scheduled audits: every page row gets "changes" (issues that are new and issues that were fixed since the last audit of the site, and whether the page is new), and the site row sums it up, with the pages that disappeared. The last audit of each site is kept in a key-value store named "seo-audit-state" in your account.

## `stateKey` (type: `string`):

Give a name to keep a separate history, for example one per client, or a new name to start over. Letters, digits, dot, dash and underscore.

## `respectRobotsTxt` (type: `boolean`):

Recommended. Pages disallowed by robots.txt are not read (they are listed in the site summary), and external links on sites that disallow crawlers are not checked. Turn off only for your own website.

## `maxConcurrency` (type: `integer`):

Requests to one website at the same time (1 to 10). Keep it low to be gentle with small servers.

## Actor input object example

```json
{
  "websites": [
    "userpilot.com"
  ],
  "maxPagesPerSite": 30,
  "useSitemap": true,
  "checkExternalLinks": true,
  "maxExternalLinksPerSite": 200,
  "compareWithPreviousRun": false,
  "respectRobotsTxt": true,
  "maxConcurrency": 3
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "userpilot.com"
    ],
    "maxPagesPerSite": 30
};

// Run the Actor and wait for it to finish
const run = await client.actor("digital_influx/seo-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": ["userpilot.com"],
    "maxPagesPerSite": 30,
}

# Run the Actor and wait for it to finish
run = client.actor("digital_influx/seo-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "userpilot.com"
  ],
  "maxPagesPerSite": 30
}' |
apify call digital_influx/seo-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,digital_influx/seo-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/NiTzxzdkaNRzqvXEY/builds/iIgdPbjqpe8ZgK1As/openapi.json
