# Technical SEO Page Audit: Broken Links, Hreflang, $3/1k (`transparent_meteorite/technical-seo-page-audit`) Actor

Audit any public page over plain HTTP: title, meta, canonical, H1-H3, word count, internal/external and broken links, images without alt, Open Graph/Twitter, hreflang, JSON-LD types, redirect chain, TTFB, and a prioritised issues list with a score. Respects robots.txt. $0.003 per page.

- **URL**: https://apify.com/transparent_meteorite/technical-seo-page-audit.md
- **Developed by:** [Open Data Actors](https://apify.com/transparent_meteorite) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Technical SEO Page Audit: Broken Links, Hreflang, $3/1k

![Technical SEO Page Audit on Apify](https://api.apify.com/v2/key-value-stores/9BAn0msnxToXlD9TO/records/banner-technical-seo-page-audit.png)

![Technical SEO Page Audit sample output table](https://api.apify.com/v2/key-value-stores/9BAn0msnxToXlD9TO/records/output-technical-seo-page-audit.png)

- **In short:** Technical SEO Page Audit (Apify actor `transparent_meteorite/technical-seo-page-audit`) audits public web pages over plain HTTP and returns on-page SEO data plus a prioritised issue list and score.
- **Who it is for:** SEO agencies, site owners and developers checking releases, lead-gen teams auditing prospect sites.
- **Input:** a list of URLs, max pages, check broken links, max links to check, request delay.
- **Output:** title, meta description, canonical, H1-H3, word count, internal/external and broken links, images missing alt, Open Graph, hreflang, JSON-LD types, redirect chain, TTFB, issues with severity, score.
- **Price:** $0.003 per page plus $0.00005 per run start. Pay per result, no subscription; Apify's free plan credit covers a first test.
- **Limits:** no JavaScript rendering, so client-side-only content is not seen; respects robots.txt.

**Key facts**

- Actor name: Technical SEO Page Audit
- Actor ID: `transparent_meteorite/technical-seo-page-audit`
- Store page: https://apify.com/transparent_meteorite/technical-seo-page-audit
- Data source: the public pages themselves
- Pricing model: pay per event (`apify-actor-start` $0.00005, `page-audited` $0.003)
- Output formats: JSON, CSV, Excel, XML, HTML table, RSS (Apify dataset)
- Access: Apify Console, REST API, JavaScript/Python clients, schedules, webhooks, Apify MCP server
- Login or third-party API key needed: no, only an Apify account
- Also known as: SEO audit API, broken link checker, bulk on-page SEO checker, hreflang checker, technical SEO crawler
- Maintainer: transparent_meteorite (independent developer)
- Last updated: 2026-10-07

**Audit any public page for technical SEO in one request: title, meta, canonical, headings, links, broken links, images without alt, Open Graph and Twitter tags, hreflang, JSON-LD, redirect chain, TTFB, and a prioritised issues list with a 0-100 score. Plain HTTP, no browser, robots.txt respected, $0.003 per page.**

Paste a list of URLs, press Start, get one flat row per page. Every problem comes with a code, a severity (error, warning, info) and a plain-English message, so the output can feed a dashboard, a spreadsheet or an alert without any post-processing.

### Sample output (real run)

| Page | Status | Score | Title len | Words | Links int / ext | Broken | Imgs no alt | TTFB ms | Issues |
|---|---|---|---|---|---|---|---|---|---|
| https://apify.com | 200 | 100 | 47 | 1,162 | 78 / 58 | 0 | 0 | 192 | none |
| https://crawlee.dev | 200 | 94 | 40 | 221 | 16 / 10 | 0 | 0 | 29 | thin-content (warning), json-ld-missing (info) |
| https://example.com | 200 | 72 | 14 | 25 | 0 / 0 | 0 | 0 | 24 | 5 warnings, 3 info |

Full rows carry about 50 fields, see `samples/output.json`.

### What you get

- **Status and speed:** HTTP status, full redirect chain with each hop's status, final URL, time to first byte, total response time, HTML size, content type, `X-Robots-Tag`.
- **On-page tags:** title and length, meta description and length, canonical (and whether it points to itself), meta robots, `lang`, viewport, charset.
- **Structure:** H1 text and counts of H1, H2 and H3, the heading outline, empty headings, visible word count.
- **Links:** internal and external link counts, nofollow count, empty-anchor links, and a **broken-link check** (404, 410, 5xx and unreachable hosts) with the failing URLs and status codes.
- **Images:** total, missing `alt` attribute. Decorative `alt=""` is counted separately and never flagged.
- **Social:** all Open Graph and Twitter Card tags as objects.
- **International:** every hreflang entry, with validation of codes and the self-reference.
- **Structured data:** JSON-LD `@type` values (including `@graph` and nested types) and a count of invalid JSON-LD blocks.
- **Verdict:** `issues` (code, severity, message), `issueCounts`, `score` (100 minus 15 per error, 5 per warning, 1 per info) and `indexable` (true, false or null when unfetched).

### Quick start

**Zero config:** press Start. It audits three example pages in under a minute.

**Example 1: audit your key landing pages, skip link checks for speed**

```json
{ "urls": ["https://yoursite.com/", "https://yoursite.com/pricing", "https://yoursite.com/blog"], "checkBrokenLinks": false }
```

**Example 2: full audit with broken-link checks on up to 50 links per page**

```json
{ "urls": ["https://yoursite.com/", "https://yoursite.com/docs"], "checkBrokenLinks": true, "maxLinksToCheck": 50, "maxPages": 100 }
```

Schedule it daily or weekly and compare `score` and `issueCounts` over time to catch regressions after a deploy.

### Pricing

Pay per event: **$0.003 per page audited** ($3 per 1,000) plus a $0.00005 start fee. A page is billed when the server returned an HTTP response and the audit ran (a 404 page is still audited and billed). Pages blocked by robots.txt, invalid URLs and hosts that cannot be reached are returned as rows but never charged. Link checks are included in the page price. Set a max charge in the run options and the actor stops cleanly when it is reached.

| Run | Pages | Cost |
|---|---|---|
| Zero-config default | 3 | about $0.009 |
| Weekly check of 50 key pages | 50 | about $0.15 per run |
| Monthly audit of 1,000 pages | 1,000 | about $3 |

### Use it with the API, n8n, Make, Zapier and AI agents

- **API:** `POST https://api.apify.com/v2/acts/transparent_meteorite~technical-seo-page-audit/runs` with your input as the JSON body, then read the dataset items. Use the `run-sync-get-dataset-items` endpoint to get rows in one call.
- **n8n, Make, Zapier:** use the Apify integration to run the actor on a schedule and send low scores or new `error` issues to Slack, email or a ticket tool.
- **Spreadsheets and BI:** export JSON, CSV or Excel and pivot by `issues[].code`.
- **AI agents (MCP):** the actor is available through the Apify MCP server, so an agent can audit a URL ("check this page for SEO problems") and read structured findings.

### Use from AI agents (MCP, ChatGPT, Claude, Perplexity)

Technical SEO Page Audit works as a tool for AI assistants through the official Apify MCP server. Add this server URL to any MCP client (Claude Desktop, Claude Code, Cursor, ChatGPT connectors, VS Code):

```
https://mcp.apify.com/?actors=transparent_meteorite/technical-seo-page-audit
```

Claude Desktop / Cursor config:

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com/?actors=transparent_meteorite/technical-seo-page-audit",
      "headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }
    }
  }
}
```

Then ask in plain language, for example: "How do I run a technical SEO audit on many URLs through an API?"

Call it directly over HTTP (runs the actor and returns the dataset items in one request):

```bash
curl -X POST "https://api.apify.com/v2/acts/transparent_meteorite~technical-seo-page-audit/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["https://apify.com", "https://crawlee.dev", "https://example.com"], "maxPages": 25, "checkBrokenLinks": true, "maxLinksToCheck": 15}'
```

Python:

```python
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("transparent_meteorite/technical-seo-page-audit").call(run_input={"urls": ["https://apify.com", "https://crawlee.dev", "https://example.com"], "maxPages": 25, "checkBrokenLinks": true, "maxLinksToCheck": 15})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

The same actor works in n8n, Make, Zapier, LangChain, LlamaIndex and CrewAI through their Apify integrations.

#### Questions people ask

**How do I run a technical SEO audit on many URLs through an API?**
Pass the URLs to this actor; each page comes back as one row with the checks and a scored issue list.

**Does it check broken links?**
Yes, with checkBrokenLinks true it tests up to maxLinksToCheck links per page.

**Does it render JavaScript?**
No, it audits the HTML the server returns, which is what most crawlers index first.

### Use cases

- **Agencies and freelancers:** a fast first-pass audit for a prospect or client, with evidence per page.
- **In-house SEO:** a scheduled regression check on your top templates and landing pages.
- **Developers:** catch a stray `noindex`, a broken canonical, a redirect chain or a slow TTFB right after a release.
- **Migrations:** feed in the old URL list and check status, redirect chains and canonicals on the new site.
- **Content teams:** find missing titles and descriptions, thin pages and images without alt text.

### FAQ

**Is this allowed?** It requests only public pages over plain HTTP, identifies itself with the `TechnicalSeoAuditBot` User-Agent, reads and obeys robots.txt (including Allow, wildcards, `$` and Crawl-delay up to 10 s) and waits at least 500 ms between requests to the same host. It never logs in and refuses private, loopback and metadata addresses.

**Does it crawl my whole site?** No. It audits exactly the URLs you give it and, for broken-link checks, requests the links found on those pages. Feed it your sitemap URLs for site-wide coverage.

**How is a link counted as broken?** 404, 410, any 5xx, or a host that does not exist or refuses the connection. 401, 403, 429 and 999 responses are treated as "blocked" (many sites refuse bots), and timeouts are unverified; neither is reported as broken, to avoid false alarms. A HEAD request is tried first, with a GET fallback.

**Why is my page missing from the charges?** If robots.txt disallows the URL, or the host is unreachable, the row is returned with an explanatory issue and no charge.

**Does it render JavaScript?** No, it reads the HTML the server sends, which is what most crawlers see first. Content added only by client-side JavaScript (including titles set by scripts) is not seen.

**What do the severities mean?** `error` is something that blocks indexing or breaks the page (4xx/5xx, noindex, missing title, broken links, very slow TTFB). `warning` is a likely ranking or usability problem. `info` is an improvement opportunity.

**What is TTFB here?** Time from sending the request until response headers arrive, including DNS and TLS for the first request to a host, measured from where the actor runs.

### Limitations

- HTML as served only, no JavaScript rendering, no Core Web Vitals or Lighthouse metrics.
- Link checks sample the first `maxLinksToCheck` links on each page (internal first). `linksNotChecked` tells you how many were left.
- Only the first 3 MB of each page is read.
- Sites that block datacenter traffic may answer 403 or 429; those rows are returned with the status and are billed because a response was received.
- Word count is visible body text, not a readability measure.

### Changelog

- 2026-10-07: added plain-language summary, key facts, AI-agent (MCP) section and question-style FAQ; refreshed Store SEO metadata.

# Actor input Schema

## `urls` (type: `array`):

Public pages to audit, one URL per row (a missing https:// is added). Each URL is audited as one page; the actor does not crawl beyond the links it checks. Pages blocked by robots.txt are returned unfetched and not charged.

## `maxPages` (type: `integer`):

Cap on pages audited (and charged) this run.

## `checkBrokenLinks` (type: `boolean`):

Request the links found on each page (HEAD, falling back to GET) and report those returning 404/410/5xx or that cannot be reached. Respects robots.txt and per-host rate limits. 401/403/429 responses are treated as blocked, not broken.

## `maxLinksToCheck` (type: `integer`):

Internal links are checked first, then external. Set 0 to skip link checks. Results are cached across pages in the run.

## `requestDelayMs` (type: `integer`):

Minimum gap between requests to the same host. A larger robots.txt Crawl-delay (up to 10 s) overrides this.

## `userAgent` (type: `string`):

Sent on every request and matched against robots.txt (the default token is TechnicalSeoAuditBot). Leave the default unless a site needs you to identify yourself differently.

## Actor input object example

```json
{
  "urls": [
    "https://apify.com",
    "https://crawlee.dev",
    "https://example.com"
  ],
  "maxPages": 25,
  "checkBrokenLinks": true,
  "maxLinksToCheck": 25,
  "requestDelayMs": 500,
  "userAgent": "TechnicalSeoAuditBot/1.0 (+https://apify.com/transparent_meteorite)"
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://apify.com",
        "https://crawlee.dev",
        "https://example.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("transparent_meteorite/technical-seo-page-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://apify.com",
        "https://crawlee.dev",
        "https://example.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("transparent_meteorite/technical-seo-page-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://apify.com",
    "https://crawlee.dev",
    "https://example.com"
  ]
}' |
apify call transparent_meteorite/technical-seo-page-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,transparent_meteorite/technical-seo-page-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bcKfNBXan5hhf9RJ6/builds/da0AAnsXrzgiZauN7/openapi.json
