# Technical SEO Audit: Crawl, Issues & Scores per Page (`everyotherfriday/seo-audit`) Actor

Point it at a site and get an issue-coded audit for every page plus a site summary with duplicate titles, broken links and status counts. 18 issue checks with a 0-100 score, robots.txt respected, no browser needed. Built for agencies, in-house SEO and pre-launch checks.

- **URL**: https://apify.com/everyotherfriday/seo-audit.md
- **Developed by:** [Paul Vasquez](https://apify.com/everyotherfriday) (community)
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$10.00 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Technical SEO Audit

Audit public websites with a bounded breadth-first HTTP crawl. This Python 3.12 Apify Actor extracts technical SEO signals from server-delivered HTML, calculates a transparent issue score, and produces one dataset row per successfully fetched 2xx HTML response. It uses httpx and Beautiful Soup without a browser. JavaScript rendering, Core Web Vitals, keyword ranking, backlink discovery, and search-engine indexing verification are outside its scope.

### Quick start

Create a Python 3.12 environment, install dependencies, and run from this directory:

```powershell
python -m venv .venv
.venv/Scripts/python.exe -m pip install -r requirements.txt
apify validate-schema .actor/input_schema.json
.venv/Scripts/python.exe -m unittest discover -s tests -v
powershell -ExecutionPolicy Bypass -File validation/run_live.ps1
```

The script copies INPUT.json into fresh local Apify storage and starts `python -m src`. Logs and aggregate evidence go under validation/; dataset rows remain under ignored storage/. The supplied input crawls Python.org and Apify with 30 candidate URLs per site. Docker uses apify/actor-python:3.12 and the same entry point.

### Input and crawl scope

`startUrls` is a required nonempty array of absolute HTTP(S) URL strings. Seeds sharing a hostname share one crawl budget and summary. Fragments are removed; paths, queries, and trailing slashes remain significant. URLs containing credentials are rejected. `maxPagesPerSite` defaults to 100 and limits attempted candidate URLs, including robots exclusions, failures, and non-HTML responses. This bounds work even when no usable HTML exists.

`sameDomainOnly` defaults to true and restricts page requests and redirect targets to the seed hostname. `includeSubdomains`, default false, admits child hostnames, but not parent or sibling hosts. It does not infer registrable domains. Setting sameDomainOnly false admits external pages within the original seed budget. Internal/external classification still uses the seed hostname rules.

`maxDepth` defaults to 3; seeds have depth zero. `concurrency` defaults to 5, with a maximum of 20. Each depth completes before the next begins. Discovery retains at most twenty times the candidate budget. Sites run sequentially. `timeoutSecs` defaults to 30 per network operation. Redirects have a ten-hop limit. Optional `proxyConfiguration` passes through the SDK to httpx; direct access is the default.

`respectRobots` defaults to true. Cached per-origin robots.txt wildcard user-agent groups determine permission before page requests and redirect hops. Rules support wildcards, terminal dollar anchors, longest-match selection, and Allow precedence on ties. Missing robots files permit access; 401, 403, 429, server errors, and network failures conservatively deny access. Robots redirects are followed to retrieve policy. Sitemap directives populate hasSitemapRef; sitemaps are not traversed. Disabling enforcement still retrieves robots for metadata. Crawl-delay is not implemented.

### Results and interpretation

Rows include source/final URLs, redirect-hop objects, depth, status, milliseconds, decoded response size, content type, title/description lengths, canonical mismatch, robots directives, language, viewport, headings, word count, images, alt omissions, link counts, hreflang, nested JSON-LD types, Open Graph, Twitter card, sitemap reference, mixed content, issues, and score.

Empty alt attributes count as intentional decorative alternatives. Word counts exclude scripts, styles, templates, and noscript content. Mixed-content checks inspect HTML asset attributes, srcset, and inline CSS URLs, but not downloaded CSS or JavaScript. Canonicals use normalized absolute URLs. Indexability means successful HTML without noindex/none in robots meta or X-Robots-Tag; this heuristic does not guarantee search-engine indexing.

After crawling, exact nonempty titles and descriptions are compared within each site. Broken internal links require observed HTTP statuses of 400 or higher. Unvisited, robots-blocked, and network-failed targets are not declared broken. Redirect aliases can produce separate rows because they represent separate requested URLs. Link counts count occurrences; broken-link arrays deduplicate.

`checkExternalLinks`, default false, checks up to 200 distinct external URLs per site using HEAD. Redirects and robots rules apply. HEAD refusals have no GET fallback. Results populate summary externalLinkStatuses and brokenLinks; external failures do not lower page scores.

### Score and pricing

Start at 100, subtract each distinct issue's weight once, and clamp at zero:

| Weight | Issue codes |
|---|---|
| 20 | NOINDEX |
| 10 | MISSING\_TITLE, CANONICAL\_MISMATCH, BROKEN\_INTERNAL\_LINK, MIXED\_CONTENT |
| 5 | MISSING\_META\_DESCRIPTION, MISSING\_H1, IMAGES\_MISSING\_ALT, THIN\_CONTENT, MISSING\_VIEWPORT, DUPLICATE\_TITLE |
| 3 | TITLE\_TOO\_LONG, MULTIPLE\_H1, SLOW\_RESPONSE, REDIRECT\_CHAIN, MISSING\_LANG, DUPLICATE\_META\_DESCRIPTION |
| 2 | META\_DESCRIPTION\_TOO\_LONG |

Thresholds are title >60 characters, description >160, content <200 words, response >2000 milliseconds, and redirect chain >1 hop. Intentional noindex pages and short landing pages may legitimately trigger issues. Fetch timing includes permission checks and redirects.

`SUMMARY-<host>` contains pages crawled/skipped, average score, issue counts, broken links, duplicate-title URL groups, status counts, elapsed seconds, fetch failures, attempted URLs, external statuses, and written rows. Counts describe the audited sample, even if a hosted event limit stops dataset writes.

The sole PPE event is `page-audited` at **$0.01 per emitted audit row**. Skipped, failed, non-HTML responses and summaries are free. Configure this event in Console and disable synthetic charges before publication. Local runs do not bill. See VALIDATION.md for measured evidence and remaining deployment checks.

### Example output

One real saved dataset row, trimmed by omitting fields only. Source: `storage/live-20260926-053309/datasets/default/000000001.json`. This is historical validation evidence, not a live response.

```json
{
  "url": "https://www.python.org/",
  "finalUrl": "https://www.python.org/",
  "statusCode": 200,
  "depth": 0,
  "title": "Welcome to Python.org",
  "indexable": true,
  "h1Count": 5,
  "imagesMissingAlt": 0,
  "issues": [
    "DUPLICATE_META_DESCRIPTION",
    "MULTIPLE_H1",
    "SLOW_RESPONSE"
  ],
  "score": 91
}
```

This saved Python.org row illustrates the documented issue weights: three distinct three-point issues reduce 100 to 91. Other page measurements are omitted. Indexable is the actor heuristic, not confirmation that a search engine indexed this URL.

### Use cases

- An SEO agency can turn page-level issue rows into a client remediation queue, reviewing intentional noindex pages separately.
- A website migration manager can inspect redirects and canonical mismatches on a bounded sample after a deployment.
- An ecommerce content operations team can review missing titles, descriptions, and image alternatives across sampled product pages.
- A web development agency can compare repeated crawl exports for the same input scope when checking a server-rendered template change.

**Pricing example:** 1,000 successful `page-audited` events x $0.01 = **$10.00**, computed from `.actor/pay_per_event.json`. This is the declared event subtotal; it does not verify active hosted billing or include any separately applicable platform or proxy costs.

### Limitations

Results describe the fetched HTML sample within the candidate and depth limits. JavaScript-generated content and unvisited link targets can remain unseen. The score is an issue-weight calculation, not a ranking forecast. Compare runs with consistent settings and review flagged pages before deciding whether a template change is necessary.

# Actor input Schema

## `maxPagesPerSite` (type: `integer`):

maxPagesPerSite limit.

## `maxDepth` (type: `integer`):

maxDepth limit.

## `concurrency` (type: `integer`):

concurrency limit.

## `timeoutSecs` (type: `integer`):

timeoutSecs limit.

## `sameDomainOnly` (type: `boolean`):

sameDomainOnly option; see README for scope.

## `includeSubdomains` (type: `boolean`):

includeSubdomains option; see README for scope.

## `respectRobots` (type: `boolean`):

respectRobots option; see README for scope.

## `checkExternalLinks` (type: `boolean`):

checkExternalLinks option; see README for scope.

## `startUrls` (type: `array`):

Absolute HTTP(S) URLs, grouped by hostname.

## `proxyConfiguration` (type: `object`):

Optional proxy; direct HTTP by default.

## Actor input object example

```json
{
  "maxPagesPerSite": 100,
  "maxDepth": 3,
  "concurrency": 5,
  "timeoutSecs": 30,
  "sameDomainOnly": true,
  "includeSubdomains": false,
  "respectRobots": true,
  "checkExternalLinks": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `pages` (type: `string`):

pages

## `csv` (type: `string`):

csv

## `summaries` (type: `string`):

summaries

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("everyotherfriday/seo-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "proxyConfiguration": { "useApifyProxy": False } }

# Run the Actor and wait for it to finish
run = client.actor("everyotherfriday/seo-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call everyotherfriday/seo-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,everyotherfriday/seo-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/FX8bhVaEf6rzKCa0g/builds/bgKBca66OubgVJfzY/openapi.json
