# Website SEO Audit - Redirects, Canonicals, Indexing (`maydit/website-seo-audit`) Actor

Audit supplied pages for redirects, canonical links, index directives and HTML metadata. Compare changes or verify expected migration targets. Static HTML, robots-aware, no API key.

- **URL**: https://apify.com/maydit/website-seo-audit.md
- **Developed by:** [Brandt May](https://apify.com/maydit) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.80 / 1,000 audited pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website SEO Audit

Audit supplied public pages for HTTP redirects, HTML canonical links, indexing directives and page metadata. Optionally check expected migration targets or compare selected fields with an earlier observation.

Each result shows the evidence and specific findings. This is an audit of downloaded HTML and HTTP responses; it does not claim to know whether a search engine has indexed a page or which canonical it selected.

### Quick start

Empty input audits the WordPress and Next.js home pages:

```json
{}
```

Supply your own public pages:

```json
{
  "urls": ["https://wordpress.org/", "https://nextjs.org/"],
  "comparePrevious": false,
  "maxRunSeconds": 240
}
```

No API key is required. The Actor fetches only the supplied pages and follows their HTTP redirects. It respects robots.txt at each destination, uses no login or proxy, and does not discover or crawl additional links.

### Inputs

| Field | Default | Meaning |
|---|---|---|
| `urls` | WordPress and Next.js home pages | Up to 100 public HTTP(S) URLs. Empty or omitted uses these samples. |
| `expectedFinalUrls` | none | Object mapping supplied source URLs to expected final migration targets. |
| `comparePrevious` | false | Opt in to storing and comparing observations. |
| `snapshotName` | `default` | Independent comparison namespace, up to 100 characters. |
| `maxRunSeconds` | 240 | 30–3,600 seconds, also bounded by the platform deadline. |

Public HTTP(S) URLs on ports 80 or 443 are supported. Credentials, local/private destinations and unsafe redirects are rejected. URL fragments are removed when normalizing inputs.

### Check a migration target

```json
{
  "urls": ["http://wordpress.org/"],
  "expectedFinalUrls": {
    "http://wordpress.org/": "https://wordpress.org/"
  }
}
```

Every mapping key must also appear in `urls`. The Actor compares the normalized observed final URL with the supplied expectation and checks the HTML canonical when present. This is an exact normalized URL comparison: a different path, query string, hostname or trailing slash can produce a mismatch. The target is not fetched separately unless it is reached through a redirect or also listed in `urls`.

A single redirect is recorded without a redirect-chain warning; multiple redirects generate a finding. Up to eight redirects are followed. Redirect loops or excessive redirects are source failures, reported in `SUMMARY`.

### Output

One dataset row represents one page with an actual HTTP response that could be audited. JSON preserves nested findings and comparisons; Apify can also export CSV and Excel.

| Fields | Meaning |
|---|---|
| `inputUrl`, `finalUrl`, `expectedFinalUrl` | Requested, observed and optionally expected URLs. |
| `httpStatus`, `contentType`, `redirectChain` | Final response status/type and every followed redirect. |
| `title`, `titleCount`, `metaDescription` | Metadata from the downloaded HTML. |
| `h1s` | H1 text values in the HTML. |
| `canonicalUrl`, `canonicalLinks` | Resolved single canonical when available and the raw canonical link values. |
| `robotsDirectives`, `xRobotsTag`, `hasNoindexDirective` | Observed HTML/HTTP directives, including raw crawler scope. |
| `hreflang` | Declared language links; targets are resolved but not fetched or validated. |
| `internalLinkCount` | Same-origin anchor occurrences, including duplicate links. |
| `imageCount`, `imagesMissingAlt` | Image elements and elements with no alt attribute; an empty alt attribute counts as present. |
| `wordCount` | Approximate whitespace-delimited body text count after selected noncontent elements are removed. |
| `issues` | Code, severity and explanation for each concrete finding. |
| `checkedAt`, `comparisonStatus`, `previousCheckedAt`, `comparison` | Observation time and optional historical comparison. |

Findings cover HTTP errors, multiple redirects, unexpected final targets, missing/duplicate titles or meta descriptions, missing H1, observed noindex/none directives, multiple or invalid canonical links, a canonical pointing elsewhere, and canonical migration mismatches.

A canonical pointing elsewhere or a noindex directive may be intentional. The Actor reports those as observations, not automatic evidence of a broken page. `hasNoindexDirective` can be true for a crawler-specific directive; inspect the raw directives to understand its scope. See Google's documentation on [robots directives](https://developers.google.com/search/docs/crawling-indexing/robots-meta-tag) and [canonical signals](https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls).

### Compare observations

```json
{
  "urls": ["https://wordpress.org/"],
  "comparePrevious": true,
  "snapshotName": "homepage-monitor"
}
```

The first successful observation has `comparisonStatus: "baseline_created"` with no prior comparison. Later results report `changedFields`, `newIssueCodes` and `resolvedIssueCodes`. Each page is still emitted and billed when unchanged.

The comparison covers HTTP status, final URL, title, meta description, H1 values, canonical URL, the observed noindex flag and issue codes. It is not a comparison of the entire document: changes to other fields, full page text, image URLs or raw directives without a changed flag are not tracked.

Snapshots live in the named key-value store `maydit-seo-audit-baselines`. The source URL, expected target and snapshot name identify a comparison scope. Changing an expected target starts a new scope. Avoid overlapping runs using the same namespace and scope.

Actual HTTP error pages are valid observations and can replace the previous baseline. Network failures, robots denial, oversized content and unsupported successful response types do not produce an observation and preserve previous history.

### Coverage, billing and failures

`SUMMARY` reports requested/emitted counts, failed URLs and reasons, robots-denied URLs, unprocessed inputs and deadline status. Usable page results are retained when another URL fails. If no page can be audited, the run fails with a diagnostic.

Launch price: **$3 per 1,000 audited pages** ($0.003 each) on Free/Bronze, plus an Actor Start event of $0.00005. Silver is 20% lower and Gold/Platinum/Diamond 40% lower. The live Pricing tab is authoritative.

**Actual HTTP error responses, such as 404 or 500, are billable audit results**, as are unchanged pages and first snapshots. Network or robots errors have no dataset row and no result charge. Each emitted page counts once, regardless of findings or redirect count. Consult the published pricing tab for platform resource charges.

### Limits

- JavaScript is not executed. Client-rendered metadata and content added after page load may be absent.
- This does not test real search-engine indexing, rankings, traffic, backlink quality, Core Web Vitals or mobile rendering.
- It checks supplied pages only. Internal links are counted; broken links elsewhere on the site are not checked.
- Only HTML and XHTML successful responses are audited. PDFs, JSON and other successful response types are reported as unsupported.
- Hreflang is extracted, not checked for reciprocity, language validity or reachable targets. Canonical destination pages are not fetched separately.
- HTML canonical links are inspected; HTTP `Link` header canonicals and XML sitemaps are not analyzed.
- Meta refresh and JavaScript redirects are not followed. Redirect tracking is HTTP only.
- Source content is bounded to 3 MiB per response. Authentication, anti-bot challenges and robots restrictions are not bypassed.
- Findings are descriptive checks, not a universal SEO score or ranking guarantee.

### Development

Run `npm test` for the fixture suite. Use isolated `CRAWLEE_STORAGE_DIR` directories for local execution. The three example inputs are drafts for verification. The comparison example explicitly opts into state storage; it does not create a schedule or send alerts.

# Actor input Schema

## `urls` (type: `array`):

Up to 100 public HTTP(S) page URLs. Omitted or empty uses https://wordpress.org/ and https://nextjs.org/. Only these pages and their redirects are fetched; links are counted but not crawled.

## `expectedFinalUrls` (type: `object`):

Optional object mapping each source URL to its expected final HTTP(S) URL, for example {"http://wordpress.org/":"https://wordpress.org/"}. Every key must also appear in urls. The target is compared with the observed final URL and declared canonical; it is not fetched separately.

## `comparePrevious` (type: `boolean`):

Opt in to named snapshot storage. First use creates an observation baseline; later runs report changed fields and new/resolved issue codes. Unchanged pages are still emitted and billed. Actual HTTP error responses can replace a baseline; network, robots and unsupported-content errors do not.

## `snapshotName` (type: `string`):

Optional namespace, up to 100 characters, for independent page monitors. Source URL and expected final URL also identify the scope. Avoid overlapping runs with the same namespace and scope.

## `maxRunSeconds` (type: `integer`):

Own wall-clock budget, shared by requests and bounded by the platform deadline with a safety margin. Work unfinished at the deadline is reported in SUMMARY.

## Actor input object example

```json
{
  "urls": [
    "https://wordpress.org/",
    "https://nextjs.org/"
  ],
  "comparePrevious": false,
  "snapshotName": "default",
  "maxRunSeconds": 240
}
```

# Actor output Schema

## `results` (type: `string`):

One audited page per row, including actual HTTP error responses, with redirects, raw indexing directives, HTML metadata, findings and optional changes.

## `summary` (type: `string`):

Requested and emitted counts, failures, skipped work and budget status. Use this to assess source coverage.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://wordpress.org/",
        "https://nextjs.org/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("maydit/website-seo-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://wordpress.org/",
        "https://nextjs.org/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("maydit/website-seo-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://wordpress.org/",
    "https://nextjs.org/"
  ]
}' |
apify call maydit/website-seo-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,maydit/website-seo-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wSBccF75cIEpqnD8u/builds/8hM0r5Z5uUecbp1YJ/openapi.json
