# Website SEO Audit & Broken Link Checker (`catalyst_prime/website-audit`) Actor

Crawls a list of start URLs and reports technical-SEO health per page: broken links, redirect chains, missing meta descriptions, alt-text coverage, structured data and mixed content — a broken link checker and SEO audit crawler in one.

- **URL**: https://apify.com/catalyst\_prime/website-audit.md
- **Developed by:** [Catalyst](https://apify.com/catalyst_prime) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website SEO Audit & Broken Link Checker

**Run it on Apify:
[apify.com/catalyst\_prime/website-audit](https://apify.com/catalyst_prime/website-audit)**: free,
no setup, runs in the browser.

Crawls a list of start URLs and reports the technical-SEO health of every page it finds: **broken
links**, **redirect chains**, **missing meta descriptions**, **alt-text coverage**, **structured
data** and **mixed content**. One row per page, so you can sort and filter the whole crawl in a
spreadsheet.

So it works as a broken link checker, a redirect chain checker and a meta description audit at
once, over a list of sites or a single site.

This repository is the Actor's source. The Actor itself runs on the
[Apify platform](https://apify.com/catalyst_prime/website-audit).

**Not a content extractor.** This does not convert pages to Markdown for an LLM/RAG pipeline.
That job is already well served on the Store. This tool answers a different question: *is this
page's plumbing broken?*

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `startUrls` | array of strings | none | Start URLs to crawl. Supply this or the alias below. Accepts a JSON array, a single string, or a comma/newline separated list. |
| `urls` | array | none | Alias of `startUrls`. |
| `maxPages` | integer | `20` | Total pages to crawl across every start URL combined. Clamped to 1-50. |
| `maxDepth` | integer | `2` | How many link-hops from each start URL to follow. `0` audits only the start URLs themselves. Clamped to 0-5. |
| `followExternalLinks` | boolean | `false` | When `true`, links to a different domain than the start URL are crawled too (still counted against `maxPages`/`maxDepth`). When `false`, off-site links are still checked and reported in `brokenLinks`, just not crawled as their own pages. |
| `timeoutSeconds` | integer | `15` | Per-request timeout, clamped to 3-60. |

### Output

One dataset row per page actually crawled:

```json
{
  "url": "https://example.com/",
  "ok": true,
  "httpStatus": 200,
  "depth": 0,
  "title": "Example Domain",
  "metaDescriptionLength": 0,
  "brokenLinks": ["https://example.com/old-page"],
  "redirectChain": [],
  "hasStructuredData": false,
  "imagesMissingAlt": 2,
  "mixedContentFound": false,
  "latencyMs": 143,
  "charged": false
}
```

- `url` is the page as requested, before any redirects; `httpStatus` and the parsed fields
  (`title`, `metaDescriptionLength`, etc.) reflect the **final** page after following them.
- `redirectChain` lists every URL hopped through after the first request, in order, ending with
  the final URL. Empty when there was no redirect.
- `brokenLinks` lists links found on the page (internal or external) that returned a network
  error or an HTTP status of 400+, checked up to 20 per page. A link already known from having
  been crawled or checked elsewhere in the same run is reused rather than re-requested.
- `hasStructuredData` is `true` if the page has a non-empty `application/ld+json` script or a
  microdata (`itemscope`) element.
- `imagesMissingAlt` counts `<img>` tags with no `alt` attribute or an empty one.
- `mixedContentFound` is `true` only for an HTTPS page that loads a resource (script, image,
  stylesheet, iframe, etc.) over plain `http://`. A protocol-relative URL (`//cdn...`) is not
  mixed content.
- `httpStatus` and the parsed fields are omitted from a row whose request never completed at
  all; `error` is set on that row instead. `charged` reports whether the row was billed (see
  Pricing): a row is charged whenever a response was actually received, even a 404 or 500,
  since the status itself is the data this Actor sells.

### How the crawl works

Breadth-first from each start URL. `maxPages` is a single global budget shared across every
start URL and every depth. Once it's spent, no further pages are fetched, wherever they were
discovered. `maxDepth` counts link-hops from whichever start URL began that branch, not from
the page that happens to link to it.

The crawler identifies itself as `CatalystWebsiteAuditBot` and reads `robots.txt` once per host
(cached for the rest of the run): a page disallowed there is skipped entirely: not fetched, not
reported, and its slot in `maxPages` is not spent. It never sends credentials, so it never
reaches anything behind a login; a page that requires one will show up as its own HTTP status
(401/403) rather than being silently skipped.

### Pricing

**Free.** This Actor has no price set, so a run costs you only your own Apify platform usage.

The code supports pay-per-event billing on a single `page-audit` event, priced per page
successfully fetched, if a price is ever set on the Apify Console. Nothing is hardcoded: it
reads its own current price from the platform at startup and runs unmetered when there isn't
one.

### Example

Crawling `https://example.com` with `maxDepth: 0` (audit just the one page, no following)
returns a single row confirming the page loads clean: `httpStatus: 200`, no broken links, no
mixed content, zero images missing alt text. Raise `maxDepth` and `maxPages` to audit a whole
section of a site in one run instead of one page at a time.

Four ready-to-run examples, each with real input and real output:

- [Audit a single page's technical SEO](https://apify.com/catalyst_prime/website-audit/examples/audit-a-single-page)
- [Find broken links across a site](https://apify.com/catalyst_prime/website-audit/examples/find-broken-links-across-a-site)
- [Check alt text coverage and structured data](https://apify.com/catalyst_prime/website-audit/examples/check-accessibility-and-structured-data)
- [Crawl a section of a site up to a page budget](https://apify.com/catalyst_prime/website-audit/examples/crawl-a-site-section)

### Other Actors from the same author

- [Email Verifier](https://apify.com/catalyst_prime/email-verifier): syntax, MX and
  disposable-domain checks, no target site to break.
  ([source](https://github.com/catalystprimeagent-bot/apify-email-verifier))
- [Tech Stack Lookup](https://apify.com/catalyst_prime/tech-stack-lookup): what a site runs on,
  plus TLS expiry, CDN and mail provider.
  ([source](https://github.com/catalystprimeagent-bot/apify-tech-stack-lookup))
- [Google Trends](https://apify.com/catalyst_prime/google-trends): interest over time and related
  queries for any term. ([source](https://github.com/catalystprimeagent-bot/apify-google-trends))

# Actor input Schema

## `startUrls` (type: `array`):

List of urls to process. Accepts a JSON array, a single string, or a comma/newline separated list.

## `urls` (type: `array`):

Alias of 'startUrls' for compatibility with tools that use this field name.

## `maxPages` (type: `integer`):

Total pages to crawl across every start URL combined. Clamped to 1-50.

## `maxDepth` (type: `integer`):

How many link-hops from each start URL to follow. 0 audits only the start URLs themselves. Clamped to 0-5.

## `followExternalLinks` (type: `boolean`):

When true, links to a different domain than the start URL are also crawled (still counted against maxPages/maxDepth). When false (default), off-site links are still checked for brokenLinks but not crawled as their own pages.

## `timeoutSeconds` (type: `integer`):

How long to wait for each page or link check before giving up. Clamped to 3-60.

## Actor input object example

```json
{
  "startUrls": [
    "example.com"
  ],
  "maxPages": 20,
  "maxDepth": 2,
  "followExternalLinks": false,
  "timeoutSeconds": 15
}
```

# Actor output Schema

## `results` (type: `string`):

The full set of results, one item per page crawled.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "example.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("catalyst_prime/website-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": ["example.com"] }

# Run the Actor and wait for it to finish
run = client.actor("catalyst_prime/website-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "example.com"
  ]
}' |
apify call catalyst_prime/website-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,catalyst_prime/website-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WVEdB5OryAODEU3Eh/builds/M4SJg3CgIWN07IzzS/openapi.json
