# SEO Audit Crawler: Broken Links, llms.txt & AI Bot Check (`lindenwerk/seo-audit-crawler`) Actor

SEO audit crawler for any website: a full site audit of titles, meta descriptions, headings, canonical, noindex, hreflang, schema and broken links, plus duplicate titles, AI crawlers blocked by robots.txt and llms.txt. 0-100 SEO score and fix hints. No proxies, no login.

- **URL**: https://apify.com/lindenwerk/seo-audit-crawler.md
- **Developed by:** [Lindenwerk Data](https://apify.com/lindenwerk) (community)
- **Categories:** SEO tools, Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $8.00 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What is SEO Audit Crawler?

A bulk **website SEO audit** crawler and on-page SEO checker. Crawl **whole websites** and audit **every page** in one run: title, meta description, headings, canonical,
noindex / indexability, hreflang, Open Graph, structured data (JSON-LD and microdata), images without alt text,
mixed content, response time, and **broken internal and external links**. Each website also gets a **site summary**:
which **AI crawlers** (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot, Applebot-Extended,
Bytespider ...) robots.txt allows or blocks, **llms.txt**, XML sitemap, HTTP to HTTPS redirect, **duplicate titles
and descriptions**, average score and top issues.

Every finding comes with a severity (error / warning / notice) and a **plain-English fix**, so the output is ready
for a client report, a spreadsheet or an AI agent. No browser, no login, no proxies, no API key. **USD 8 per 1,000
pages.**

### Who it's for

- **SEO agencies and freelancers.** Audit 10 client sites overnight, export the issues view to Excel, and send each
  client a fix list. Schedule it monthly to catch new 404s and noindex accidents.
- **In-house marketing and web teams.** Check a site after a relaunch or CMS migration: redirect chains, lost
  canonicals, pages that turned `noindex`, broken internal links, missing meta descriptions.
- **Lead-gen and sales teams selling web services.** Score prospects' websites and open the conversation with
  concrete problems ("23 broken links, no meta descriptions on 40 pages").
- **Publishers and site owners deciding on AI crawlers.** See at a glance which AI training bots and AI search bots
  your robots.txt blocks, whether that is what you intended, and whether you publish an llms.txt.
- **AI agents and LLM workflows.** Flat, small JSON rows with issue ids and fix hints, callable over the Apify MCP
  server ("audit example.com and tell me the top 5 fixes").

### What it checks

**Per page (one row per URL, 40 rule types):**

| Area | Checks |
|---|---|
| Status & redirects | HTTP 4xx / 5xx, redirect chains (2+ hops), redirected internal URLs, HTTP instead of HTTPS, slow response (> 1.5 s), very large HTML |
| Title & description | Missing, too long / too short, multiple tags, duplicates across the site |
| Headings | Missing / empty / multiple H1, skipped heading levels |
| Indexability | `noindex` in meta robots or X-Robots-Tag, canonical missing / pointing elsewhere / cross-domain / multiple, robots.txt blocking Googlebot, not in XML sitemap |
| International | `lang` attribute, invalid hreflang codes, hreflang without self-reference |
| Social & schema | Open Graph tags, schema.org types from JSON-LD and microdata, invalid JSON-LD |
| Content & media | Thin content (< 200 words on indexable pages), images without alt, mixed content (http:// resources on https pages), viewport meta |
| Links | Internal / external / nofollow counts, **broken links** with URL, status and anchor text |

Each page gets a **score from 0 to 100** (100 minus 10 per error type, 4 per warning type, 1 per notice type) and a
sorted `issueIds` list for easy filtering.

**Per website (one summary row):**

- **robots.txt**: found, sitemaps declared, crawl-delay, whether it disallows us.
- **AI crawler matrix**: for 16 AI bots plus Googlebot and Bingbot, whether the homepage is allowed, whether the bot is
  named explicitly, and a one-line policy such as `blocks AI training, allows AI search` or `allows all AI crawlers
  (no AI-specific rules)`. Cloudflare-style `Content-Signal` lines (e.g. `search=yes, ai-train=no`) are reported
  too. We read the rules; we never pretend to be those bots.
- **llms.txt / llms-full.txt**: present, size, title, summary and link count. Reported as information only: Google
  says it doesn't use llms.txt for Search.
- **XML sitemap**: found (robots.txt or /sitemap.xml), sitemap indexes followed, URL count.
- **HTTPS**: does `http://` redirect to `https://`, and is it permanent?
- **Duplicates**: groups of pages sharing a title or meta description (indexable pages only).
- **Roll-up**: pages audited, status code counts, indexable pages, average score, issues by severity, top issues,
  broken link samples, what was skipped and why the crawl stopped.

### How to run an SEO site audit

1. Click **Start** with the prefilled example (10 pages of crawlee.dev) to see real output in about 15 seconds.
2. Paste your own websites into **Websites** and set **Max pages per site**.
3. Choose how pages are found: `crawl` follows internal links, `sitemap` reads the XML sitemap, `list` audits only
   the URLs you paste.
4. Look at the results in the **Output** tab (views such as "Overview", "Issues & fixes", "Broken links" and
   "Site summary & AI crawlers") or
   download them as JSON, CSV or Excel for a client report.
5. To re-audit regularly, save the input as a **task** and add a **schedule** (for example weekly).
6. Send results on with Apify **integrations** (Slack, Google Drive, Make, Zapier, n8n or a webhook), or call the Actor through the **Apify API** or from AI agents via MCP.

### Input

Only `startUrls` is required.

```json
{
  "startUrls": ["https://crawlee.dev"],
  "maxPagesPerSite": 10
}
```

**Agency batch: 5 client sites, 500 pages each, external links too:**

```json
{
  "startUrls": ["client-one.de", "client-two.com", "client-three.at", "client-four.ch", "client-five.fr"],
  "maxPagesPerSite": 500,
  "checkLinks": "all"
}
```

**Audit only the blog, from the sitemap:**

```json
{
  "startUrls": ["example.com"],
  "mode": "sitemap",
  "includeUrlPatterns": ["/blog/"],
  "maxPagesPerSite": 1000
}
```

**Exact URL list (e.g. your top landing pages), no site row:**

```json
{
  "startUrls": ["https://example.com/pricing", "https://example.com/features", "https://example.com/"],
  "mode": "list",
  "includeSiteSummary": false
}
```

| Field | Default | What it does |
|---|---|---|
| `startUrls` | none | Websites or URLs, one per line. Aliases: `urls`, `websites`, `domains`, `url`, so inputs from other SEO Actors work unchanged. |
| `mode` | `crawl` | `crawl` follows internal links breadth-first; `sitemap` audits the URLs in the XML sitemap; `list` audits exactly the URLs given |
| `maxPagesPerSite` | 100 | Audited pages per website (max 10,000) |
| `maxDepth` | 5 | Crawl mode: clicks from the start URL |
| `checkLinks` | `internal` | `internal`, `all` (adds external links) or `none` |
| `maxLinkChecksPerSite` | 300 | Cap on extra HEAD/GET link checks per site (link checks are free) |
| `includeSiteSummary` | true | Add the site row (AI crawlers, llms.txt, sitemap, HTTPS, duplicates, top issues) |
| `includeSubdomains` | false | Also crawl blog.example.com etc. |
| `includeUrlPatterns` / `excludeUrlPatterns` | none | Text or `*` wildcard filters, e.g. `/blog/` or `*/tag/*` |
| `skipUrlsWithQuery` | true | Ignore `?filter=` style URL variants while crawling |
| `maxConcurrencyPerSite` | 3 | Parallel requests per website (max 8; a robots.txt Crawl-delay forces 1) |
| `maxConcurrentSites` | 3 | Websites audited in parallel |
| `requestTimeoutSecs` | 20 | Per-request timeout |

### Output

Page rows (`rowType: "page"`) and one site row per website (`rowType: "site"`). Shortened sample from a real run:

```json
{
  "rowType": "page",
  "site": "crawlee.dev",
  "url": "https://crawlee.dev/blog",
  "httpStatus": 200,
  "title": "Crawlee Blog - learn how to build better scrapers | Crawlee for JavaScript · Build reliable crawlers. Fast.",
  "titleLength": 107,
  "metaDescriptionLength": 151,
  "h1Count": 0,
  "canonicalStatus": "self",
  "indexable": true,
  "structuredDataTypes": ["Blog"],
  "wordCount": 1492,
  "inSitemap": true,
  "brokenLinks": [],
  "issues": [
    {"id": "title-too-long", "severity": "warning", "message": "Title is 107 characters (over 60).", "fix": "Shorten the <title> to about 60 characters so it is not cut off in results."},
    {"id": "h1-missing", "severity": "warning", "message": "No <h1> heading.", "fix": "Add one visible <h1> that states the page topic."}
  ],
  "score": 92,
  "issueIds": ["h1-missing", "title-too-long"],
  "chargedEvent": "page-audited"
}
```

```json
{
  "rowType": "site",
  "site": "crawlee.dev",
  "title": "Site summary: 10 pages audited, average score 95.5, AI crawlers: allows all AI crawlers (no AI-specific rules)",
  "siteSummary": {
    "aiCrawlers": {"policy": "allows all AI crawlers (no AI-specific rules)", "trainingBlocked": [], "searchBlocked": [],
                   "googlebotAllowed": true, "contentSignals": []},
    "llms": {"llmsTxt": {"found": true, "bytes": 30542, "hasTitle": true, "linkCount": 361}},
    "sitemap": {"found": true, "urlCount": 4846},
    "averageScore": 95.5,
    "topIssues": [{"id": "title-too-long", "severity": "warning", "pages": 7}],
    "duplicateDescriptions": [{"count": 6, "urls": ["https://crawlee.dev/", "https://crawlee.dev/js"]}],
    "brokenLinks": {"total": 0, "linksChecked": 285},
    "crawlStoppedReason": "page limit reached"
  },
  "chargedEvent": "site-summary"
}
```

Dataset views: **Overview**, **On-page**, **Issues & fixes**, **Broken links** and **Site summary & AI crawlers**.
Export as CSV, Excel, JSON or XML, or read through the API. `RUN_SUMMARY` in the key-value store holds charged events,
free rows, HTTP requests, bytes downloaded and runtime.

`fetchStatus` values: `ok`, `http_error` (a 4xx/5xx page: a finding, charged), and free rows `not_html`, `timeout`,
`unreachable`, `refused` (non-public address) and `error` (e.g. redirect loop). URLs that robots.txt disallows for
us are not requested and produce no row; they are counted in the site summary.

### Pricing (pay per event)

| Event | Price | When |
|---|---|---|
| `page-audited` | **USD 0.008** (USD 8 per 1,000 pages) | A URL answered with an HTML page (including 404/500 pages, which are findings) and was audited |
| `site-summary` | **USD 0.01** | One summary row per website (optional, `includeSiteSummary`) |
| Free | 0 | Timeouts, unreachable hosts, non-HTML files, robots.txt-blocked URLs, all link checks |
| Actor start | USD 0.00005 | Apify platform minimum per run |

A 100-page site costs USD 0.81. A 1,000-page site costs USD 8.01. Set **Maximum cost per run** in Apify Console: the
Actor reserves budget before each page and stops cleanly at the limit, so it never charges more (parallel sites share
the limit).

### Use with AI agents (MCP)

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com/?tools=lindenwerk/seo-audit-crawler",
      "headers": {"Authorization": "Bearer <APIFY_TOKEN>"}
    }
  }
}
```

Example prompts:

- "Audit the first 50 pages of example.com and list the five fixes with the biggest impact."
- "Which AI crawlers does example.com block, and does it have an llms.txt?"
- "Find all broken internal links on example.de and the pages they are on."

For short agent answers, use `maxPagesPerSite: 20` and read the site row first. Over HTTP:
`POST https://api.apify.com/v2/acts/lindenwerk~seo-audit-crawler/run-sync-get-dataset-items?token=<APIFY_TOKEN>`
with the input JSON as the body.

### How it works and limits

- Plain HTTP requests (no headless browser), so it's fast and cheap: in our test, 300 pages on 3 sites took 37 s
  at 512 MB. Content injected only by client-side JavaScript isn't seen. Server-rendered and statically generated
  sites (WordPress, TYPO3, Shopify, Next.js/Docusaurus SSG ...) are audited fully.
- Link checks use HEAD, confirmed with GET when HEAD is refused or answers 4xx (many servers mishandle HEAD). Pages already crawled aren't requested twice. 401, 403 and 429
  answers are "restricted", not broken, because many sites refuse data-center requests.
- The AI crawler matrix evaluates robots.txt rules for each bot's token on the homepage. It doesn't test firewall or
  CDN bot blocking (e.g. Cloudflare's AI bot block), which can deny a bot even when robots.txt allows it.
- Some sites block all cloud-server traffic. Their pages come back as `http_error` 403 or as free timeouts. After 10
  timeouts or connection failures in a row the Actor stops crawling that site.

### Data, privacy and responsible use

- **Honours robots.txt** (RFC 9309) for the user agent token `LindenwerkSEOAudit`, including Crawl-delay (up to
  10 s). If robots.txt can't be fetched because of a server error, the site isn't crawled.
- **Polite by default**: 3 parallel requests per site, a clear user agent, no proxies, no login, no captcha solving,
  no bot-protection workarounds.
- **No personal data.** It reads technical page metadata only: tags, headings, link targets, word counts. It does not
  extract names, email addresses or phone numbers. `mailto:` and `tel:` links are ignored.
- Audit sites you own or are allowed to audit (clients, prospects' public pages at a polite rate). You are responsible
  for how you use the results.

### FAQ

**Is this a Screaming Frog alternative?** For the core HTTP audit, yes: crawl, status codes, titles, descriptions,
headings, canonicals, noindex, hreflang, structured data, broken links, duplicates. It runs in the cloud, on a
schedule, and through an API or AI agent. It doesn't render JavaScript or connect to Search Console.

**Why are some pages missing?** They are disallowed by robots.txt for our bot, outside the start URL's site, behind
a `?query` (see `skipUrlsWithQuery`), beyond `maxDepth`, or the page limit was reached. The site row tells you which.

**Does it change anything on my site?** No. It only sends GET and HEAD requests.

**Is llms.txt important for SEO?** No. We report it because many teams ask, but robots.txt is what controls AI
crawler access.

### Examples

- [Audit a website for SEO issues and broken links](https://apify.com/lindenwerk/seo-audit-crawler/examples/audit-website-seo-issues): a published example task with a ready-made input. Open it, adjust the input and run it in your own Apify account.

### Other Lindenwerk Data Actors

- [German & EU Tenders: Public Procurement Monitor](https://apify.com/lindenwerk/german-eu-tender-matcher): Ausschreibungen from oeffentlichevergabe.de and TED, deduplicated and scored to your CPV codes and keywords.
- [France Tenders Monitor: BOAMP, TED & Marchés Publics](https://apify.com/lindenwerk/france-tender-matcher): marchés publics from BOAMP and TED in one deduplicated, scored list.
- [UK Tenders Monitor: Government Contracts & Find a Tender Alerts](https://apify.com/lindenwerk/uk-tender-matcher): Find a Tender notices (above and below threshold) as daily tender alerts, scored to your profile.
- [SAM.gov Government Contract Opportunities Monitor](https://apify.com/lindenwerk/sam-gov-contract-matcher): US federal bids and RFPs from SAM.gov, scored to your NAICS and set-asides.
- [Website Technology & Tech Stack Detector: BuiltWith Alternative](https://apify.com/lindenwerk/tech-stack-detector): CMS, shop system, analytics, consent manager and email provider of any website, in bulk.
- [PDF to Markdown & RAG Chunks: Document Parser (DOCX/PPTX/XLSX)](https://apify.com/lindenwerk/pdf-to-markdown-rag): PDF, Word, PowerPoint and Excel to clean Markdown with inline tables and page-cited RAG chunks.

### Deutsch (Kurzfassung)

Dieser Actor crawlt ganze Websites und prüft jede Seite: Title, Meta-Description, Überschriften, Canonical, noindex,
hreflang, Open Graph, strukturierte Daten, Bilder ohne Alt-Text, Mixed Content und defekte interne und externe
Links. Jeder Fund kommt mit Schweregrad und konkretem Lösungshinweis. Pro Website gibt es eine Zusammenfassung:
welche KI-Crawler (GPTBot, ClaudeBot, PerplexityBot, Google-Extended ...) die robots.txt erlaubt oder sperrt,
llms.txt, XML-Sitemap, HTTP-zu-HTTPS-Weiterleitung, doppelte Titles und Descriptions und die häufigsten Probleme.
Ideal für SEO-Agenturen, Relaunch-Checks und Website-Leadlisten. Preis: USD 0,008 pro geprüfter Seite und USD 0,01
pro Website-Zusammenfassung; nicht erreichbare Seiten und Link-Checks sind kostenlos. Der Actor beachtet robots.txt,
nutzt keine Proxys und erhebt keine personenbezogenen Daten.

### Changelog

See the Changelog tab. Version 0.1 is the first public release.

*Made by Lindenwerk Data.*

# Changelog

This Actor's version history is a separate document: https://apify.com/lindenwerk/seo-audit-crawler/changelog.md

# Actor input Schema

## `startUrls` (type: `array`):

Websites to audit, one per line (example.com, www.example.com or https://example.com/blog/). In crawl and sitemap mode each site is audited once (www/non-www merged); in list mode every URL is audited exactly. Only public websites on ports 80/443; IP addresses and internal hostnames are refused.

## `mode` (type: `string`):

crawl = breadth-first through internal links (like Screaming Frog spider mode); sitemap = read robots.txt / sitemap.xml (incl. sitemap indexes) and audit those URLs; list = only the URLs you paste.

## `maxPagesPerSite` (type: `integer`):

Stop after this many audited pages per website (crawl and sitemap mode). The run also stops cleanly at your maximum charge.

## `maxDepth` (type: `integer`):

Crawl mode: how many clicks away from the start URL to follow (0 = start page only).

## `checkLinks` (type: `string`):

Check found links for 4xx/5xx and dead hosts with light HEAD requests (GET fallback). Pages already crawled are not requested twice. 401/403/429 answers are reported as "restricted", not broken.

## `maxLinkChecksPerSite` (type: `integer`):

Safety cap on extra link-check requests per website. Link checks are not charged.

## `includeSiteSummary` (type: `boolean`):

One extra row per website: robots.txt, AI crawler access matrix (GPTBot, ClaudeBot, PerplexityBot, Google-Extended ... allowed or blocked), llms.txt, XML sitemap, HTTP to HTTPS redirect, duplicate titles and descriptions, average score and top issues. Charged as one "site-summary" event.

## `includeSubdomains` (type: `boolean`):

Crawl mode: also follow links to subdomains (blog.example.com when you start at example.com).

## `includeUrlPatterns` (type: `array`):

Optional. Only audit URLs that contain one of these texts or match these \* wildcards (e.g. /blog/ or https://example.com/de/\*).

## `excludeUrlPatterns` (type: `array`):

Optional. Skip URLs that contain one of these texts or wildcards (e.g. /tag/, /cart, */page/*).

## `skipUrlsWithQuery` (type: `boolean`):

Crawl mode: ignore links with query strings (filters, sorting, session IDs) to avoid crawling endless variants.

## `maxConcurrencyPerSite` (type: `integer`):

Politeness: at most this many requests to one website at a time. A robots.txt Crawl-delay forces 1 (delay capped at 10 s).

## `maxConcurrentSites` (type: `integer`):

How many websites are audited at the same time.

## `requestTimeoutSecs` (type: `integer`):

Timeout per HTTP request. Pages that time out are reported as free rows.

## Actor input object example

```json
{
  "startUrls": [
    "https://crawlee.dev"
  ],
  "mode": "crawl",
  "maxPagesPerSite": 10,
  "maxDepth": 5,
  "checkLinks": "internal",
  "maxLinkChecksPerSite": 300,
  "includeSiteSummary": true,
  "includeSubdomains": false,
  "skipUrlsWithQuery": true,
  "maxConcurrencyPerSite": 3,
  "maxConcurrentSites": 3,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset with page rows and site summary rows

## `runSummary` (type: `string`):

Charged events, free rows, HTTP requests, bytes downloaded, runtime

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://crawlee.dev"
    ],
    "maxPagesPerSite": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("lindenwerk/seo-audit-crawler").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["https://crawlee.dev"],
    "maxPagesPerSite": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("lindenwerk/seo-audit-crawler").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://crawlee.dev"
  ],
  "maxPagesPerSite": 10
}' |
apify call lindenwerk/seo-audit-crawler --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,lindenwerk/seo-audit-crawler"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/D6vzVm6Kh0otiDWQ9/builds/zuudKgsRYJ6j88h8i/openapi.json
