# SEO Site Audit — Website Crawler & Lighthouse (`cheapapi/seo-site-audit`) Actor

SEO site audit crawler: status codes, titles, meta, H1, canonicals, indexability, broken links, duplicates, speed, onpage score & Lighthouse.

- **URL**: https://apify.com/cheapapi/seo-site-audit.md
- **Developed by:** [CheapAPI](https://apify.com/cheapapi) (community)
- **Categories:** SEO tools, Developer tools, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event + usage

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## SEO Site Audit — Website Crawler & Lighthouse

Crawl any website and get a full technical SEO audit for every page — status codes, titles, meta descriptions, H1s, canonicals, indexability, broken links, duplicate content, page speed, an onpage score and a plain-English list of issues — plus a site summary and optional Lighthouse scores.

**Who this is for:** SEO agencies auditing client sites, site owners and marketers checking their own website, and developers who need crawl data in a pipeline or CI check.

**Why this Actor**

- **From $3.40 per 1,000 audited pages** (Gold plan; $6.90 on Free), no monthly fee; Apify platform usage is billed separately — versus $30 (Gold) / $40 (Free) for the most-used SEO audit crawler in the Store and $5 (Gold) / $10 (Free) for a 28-field budget crawler. One per-page auditor is $0.40 cheaper on Gold ($3 vs $3.40) but finds no broken links, site-wide duplicates or Lighthouse scores (see the comparison below).
- **99 fields per page plus 60+ SEO checks** ([full field reference](#what-data-you-get)): onpage score 0–100, a priority (high / medium / low) and a fix hint for every issue, word count, readability, links, page size, timings, broken link targets, Core Web Vitals (optional) — and site-wide checks (sitemap, robots.txt, SSL, HTTPS/www redirects, 404 page, duplicates across the site).
- **Lighthouse on demand**: performance, accessibility, best practices and SEO scores with LCP, TBT, CLS and failed audits — from $0.007 per URL.
- **Crawl exactly what you need**: include-only and exclude URL patterns (`/blog/`, `*?sort=`, `*.pdf$`), crawl speed from gentle (5 s between requests) to fastest (0.1 s), depth and time limits.
- **Free change monitoring**: each page is marked `new`, `changed` or `unchanged` versus the previous run, with the changed fields — schedule it weekly and review only what moved. No browser, proxies or API keys to set up.

### Compared with alternatives

A typical run: **1,000 pages of one website**. Prices are the per-page charge × 1,000 on the Free and Gold plans (event fees only). For this Actor, Apify platform usage is billed separately by Apify (at the default 256 MB a small run typically uses about $0.001–$0.005), on top of the price shown.

| | 1,000 pages (Free / Gold) | Crawls the whole site | Fields per page (as listed) | Broken link targets | Site-wide duplicates | Lighthouse | JavaScript rendering | Change tracking |
|---|---|---|---|---|---|---|---|---|
| Most-used SEO audit crawler (~670 users) | $40 / $30 | ✓ | ~40 (as listed) | ✓ | titles, headings | lab speed metrics | ✗ | ✗ |
| Most-used SEO data extractor (~480 users) | $300 / $10 | ✗ (your URLs only) | 11 (as listed) | ✗ | ✗ | ✗ | ✓ | ✗ |
| Budget SEO audit crawler (~110 users) | $10 / $5 | ✓ | 28 (as listed) | ✗ | ✗ | ✗ | ✗ | ✗ |
| Low-cost per-page auditor (~190 users) | **$5 / $3** | optional | not stated (structured-data focus) | ✗ | ✗ | ✗ | ✗ | ✗ |
| **This Actor** | $6.90 / $3.40 **plus platform usage** | ✓ | **99** + 60+ checks | ✓ (URL, status, anchor) | titles, descriptions, content | ✓ ($0.007–$0.07 per URL) | ✓ (+$0.002 per page) | ✓ free |

The low-cost per-page auditor costs $1.90 (Free) / $0.40 (Gold) less per 1,000 pages. It checks structured data, social tags and headings of the URLs you give it; this Actor adds broken link targets, duplicate titles/descriptions/content across the site, indexability reasons, Lighthouse, JavaScript rendering, a site summary (sitemap, robots.txt, SSL, redirects) and free change tracking between runs. Our total includes the $0.002 website summary fee. Rival field counts are as listed on their Store pages, not independently verified.

**Cheaper still:** a very new crawler with about 4 users charges $0.50 per 1,000 pages (plus a small start fee and platform usage) for a basic audit — status, title/description, canonical, headings, word count, links, image alt gaps, indexability and structured data. If that is all you need, it costs less; it does not list broken link targets, site-wide duplicates, Lighthouse, JavaScript rendering or change tracking.

**Lighthouse only?** If you just need Lighthouse / PageSpeed scores for URLs you already know, a PageSpeed-API Actor (~90 users) is cheaper at about $0.002 per URL versus our $0.007 (Gold) – $0.07 (Free), and a free Lighthouse Actor (~200 users) charges only platform usage. Use this Actor when you also need the crawl, issues, broken links and duplicates.

Prices, user counts and feature lists from public Apify Store listings, checked September 2026.

#### Not included

- **Keyword rankings and search volumes** — this Actor audits pages, not Google positions. Use [Domain SEO Analyzer](https://apify.com/cheapapi/domain-seo-analyzer) for the keywords a domain ranks for, or [Keyword Research Tool](https://apify.com/cheapapi/keyword-research-tool) for volumes and ideas.
- **Backlinks** (who links to the site) — only links *on* the crawled pages are analysed. Use [Backlink Checker](https://apify.com/cheapapi/backlink-checker).
- **Real-user (field) Core Web Vitals** — Lighthouse and full browser rendering measure lab values from one test run, not visitor data.
- **Pages behind a login** — only publicly reachable pages are crawled.
- **Automatic fixes** — you get issues and data, not changes to your website.

### What data you get

One row per crawled page (HTML pages, broken URLs and redirects) in the dataset:

| Field | Type | Example |
|---|---|---|
| `id` | string | `3f9c1a0b7d2e4c61` (stable across runs: hash of website + URL) |
| `url` | string | `https://example-shop.com/` |
| `statusCode` | number | `200` |
| `pageType` | string | `html` / `broken` / `redirect` |
| `redirectUrl` | string | `https://example-shop.com/` (for redirects) |
| `onpageScore` | number (0–100) | `97.07` |
| `issuesCount`, `issues`, `issueCodes` | number, array | `4`, `["Images without alt text", …]`, `["noImageAlt", …]` |
| `issueDetails` | array | `[{"code": "noImageAlt", "priority": "medium", "fixHint": "Add descriptive alt text…"}]` — one entry per issue code, same order |
| `title`, `titleLength`, `titleMissing`, `titleTooShort`, `titleTooLong`, `duplicateTitle` | string, number, boolean | `"Example Shop – Running Shoes…"`, `55`, `false` |
| `metaDescription`, `metaDescriptionLength`, `metaDescriptionMissing`, `duplicateMetaDescription` | string, number, boolean | `"Shop running shoes…"`, `123` |
| `h1`, `h1Count`, `h1Missing`, `h1List`, `h2Count`, `h3Count` | string, number, boolean, array | `"Running shoes for every runner"`, `1` |
| `wordCount`, `textToHtmlRatio`, `readabilityScore` | number | `1165`, `0.047`, `31.26` |
| `canonicalUrl`, `canonicalIsSelf` | string, boolean | `https://example-shop.com/`, `true` |
| `indexable`, `nonIndexableReason` | boolean, string | `false`, `robots_txt` / `meta_tag` / `redirect` / `canonicalized_to_other_url` |
| `duplicateContent` | boolean | `false` |
| `hasBrokenLinks`, `brokenLinksCount`, `brokenLinks` | boolean, number, array | `true`, `2`, `[{"url": "…/missing", "statusCode": 404, "anchorText": "Old offer"}]` |
| `hasBrokenResources` | boolean | broken images, scripts or CSS |
| `clickDepth`, `internalLinksCount`, `externalLinksCount`, `inboundLinksCount` | number | `0`, `130`, `33`, `11` |
| `imagesCount`, `imagesWithoutAlt`, `scriptsCount`, `renderBlockingScripts` | number, boolean | `58`, `true`, `48`, `12` |
| `pageSizeBytes`, `totalTransferSizeBytes`, `domSize` | number | `148362`, `2184530`, `162918` |
| `loadTimeMs`, `timeToFirstByteMs`, `timeToInteractiveMs`, `largestContentfulPaintMs`, `cumulativeLayoutShift` | number (Core Web Vitals: null without Full browser rendering) | `486`, `112`, `512` |
| `isHttps`, `hasStructuredData`, `socialTags`, `misspelledWords`, `deprecatedTags` | mixed | `true`, `{"og:type": "website"}` |
| `lighthousePerformanceScore`, `lighthouseAccessibilityScore`, `lighthouseBestPracticesScore`, `lighthouseSeoScore`, `lighthouse` | number, object | `45`, `75`, `54`, `85` (when Lighthouse is on) |
| `checks` | object | all 60+ true/false checks |
| `changedSinceLastRun`, `changeType` | boolean, string | `true`, `new` / `changed` / `unchanged` / `baseline` (first run) |
| `changedFields`, `previousOnpageScore`, `previousStatusCode` | array, number, number | `["title", "onpageScore"]`, `92.5`, `200` |
| `website`, `fetchedAt`, `scrapedAt` | string | context |

<details>
<summary><b>Full field reference — all 99 fields of a page row</b> (generated from the dataset schema)</summary>

| # | Field | Type | Description |
|---|---|---|---|
| 1 | `id` \* | string | Stable row ID: first 16 hex characters of SHA-256 of website + page URL. Same page of the same website gets the same ID in every run. |
| 2 | `website` \* | string | The website exactly as you entered it in the input. |
| 3 | `url` \* | string | Crawled page URL. |
| 4 | `pageType` | string or null | "html" (normal page), "broken" (4xx/5xx or unreachable) or "redirect" (3xx). |
| 5 | `statusCode` | number or null | HTTP status code of the page. |
| 6 | `redirectUrl` | string or null | Redirect target (Location header) for redirect pages. |
| 7 | `onpageScore` | number or null | Onpage score 0–100 computed from the page checks (100 = no issues). |
| 8 | `issuesCount` \* | number | Number of issues found on the page. |
| 9 | `issues` \* | array | Issues in plain English, e.g. "Images without alt text". |
| 10 | `issueCodes` | array | Machine-readable issue codes matching "issues", e.g. "noImageAlt". |
| 11 | `issueDetails` | array | One entry per issue code, in the same order: {code, priority (high / medium / low), fixHint (one-sentence fix suggestion)}. |
| 12 | `title` | string or null | Content of the <title> tag. |
| 13 | `titleLength` | number or null | Title length in characters. |
| 14 | `titleMissing` | boolean or null | True if the page has no title (HTML pages only). |
| 15 | `titleTooShort` | boolean or null | True if the title is shorter than "Title too short below" (default 30). |
| 16 | `titleTooLong` | boolean or null | True if the title is longer than "Title too long above" (default 65). |
| 17 | `duplicateTitle` | boolean or null | True if another crawled page of the website has the same title. |
| 18 | `metaDescription` | string or null | Content of the meta description tag. |
| 19 | `metaDescriptionLength` | number or null | Meta description length in characters. |
| 20 | `metaDescriptionMissing` | boolean or null | True if the page has no meta description (HTML pages only). |
| 21 | `duplicateMetaDescription` | boolean or null | True if another crawled page has the same meta description. |
| 22 | `h1` | string or null | First H1 heading on the page. |
| 23 | `h1Count` | number or null | Number of H1 headings. |
| 24 | `h1Missing` | boolean or null | True if the page has no H1 (HTML pages only). |
| 25 | `h1List` | array or null | All H1 headings on the page. |
| 26 | `h2Count` | number or null | Number of H2 headings. |
| 27 | `h3Count` | number or null | Number of H3 headings. |
| 28 | `wordCount` | number or null | Words of visible text on the page. |
| 29 | `textToHtmlRatio` | number or null | Visible text size divided by HTML size (0–1). |
| 30 | `readabilityScore` | number or null | Flesch-Kincaid readability index of the page text. |
| 31 | `titleRelevance` | number or null | How well the title matches the page content (0–1). |
| 32 | `descriptionRelevance` | number or null | How well the meta description matches the page content (0–1). |
| 33 | `canonicalUrl` | string or null | URL in the canonical link tag. |
| 34 | `canonicalIsSelf` | boolean or null | True if the canonical URL points to the page itself. |
| 35 | `indexable` | boolean or null | True if search engines may index the page. |
| 36 | `nonIndexableReason` | string or null | Why the page is not indexable, e.g. robots\_txt, meta\_tag, redirect, canonicalized\_to\_other\_url, error\_status\_code. |
| 37 | `metaRobotsFollow` | boolean or null | False if the meta robots tag says nofollow. |
| 38 | `duplicateContent` | boolean or null | True if another crawled page has (nearly) the same content. |
| 39 | `hasBrokenLinks` | boolean or null | True if the page links to a broken URL. |
| 40 | `brokenLinksCount` | number | Number of broken links listed for the page. |
| 41 | `brokenLinks` | array or null | Broken link targets (URL, status code, anchor text), up to 100 per page. |
| 42 | `hasBrokenResources` | boolean or null | True if an image, script or stylesheet on the page is broken. |
| 43 | `clickDepth` | number or null | Clicks from the start page (0 = start page). |
| 44 | `internalLinksCount` | number or null | Links to pages of the same website. |
| 45 | `externalLinksCount` | number or null | Links to other websites. |
| 46 | `inboundLinksCount` | number or null | Links from other crawled pages to this page. |
| 47 | `imagesCount` | number or null | Number of images on the page. |
| 48 | `imagesSizeBytes` | number or null | Total size of the page images in bytes. |
| 49 | `imagesWithoutAlt` | boolean or null | True if at least one image has no alt text. |
| 50 | `scriptsCount` | number or null | Number of scripts on the page. |
| 51 | `stylesheetsCount` | number or null | Number of stylesheets on the page. |
| 52 | `renderBlockingScripts` | number or null | Number of render-blocking scripts. |
| 53 | `renderBlockingStylesheets` | number or null | Number of render-blocking stylesheets. |
| 54 | `pageSizeBytes` | number or null | HTML size in bytes (uncompressed). |
| 55 | `compressedSizeBytes` | number or null | HTML size in bytes as transferred (compressed). |
| 56 | `totalTransferSizeBytes` | number or null | Total transferred bytes including resources. |
| 57 | `domSize` | number or null | Size of the rendered DOM. |
| 58 | `loadTimeMs` | number or null | Total page load time in ms. |
| 59 | `timeToFirstByteMs` | number or null | Time to first byte (server waiting time) in ms. |
| 60 | `downloadTimeMs` | number or null | HTML download time in ms. |
| 61 | `connectionTimeMs` | number or null | Connection time in ms. |
| 62 | `timeToInteractiveMs` | number or null | Time to interactive in ms. |
| 63 | `domCompleteMs` | number or null | Time until the DOM is complete in ms. |
| 64 | `largestContentfulPaintMs` | number or null | Largest Contentful Paint in ms (with Full browser rendering only). |
| 65 | `firstInputDelayMs` | number or null | First Input Delay in ms (with Full browser rendering only). |
| 66 | `cumulativeLayoutShift` | number or null | Cumulative Layout Shift (with Full browser rendering only). |
| 67 | `isHttps` | boolean or null | True if the page is served over HTTPS. |
| 68 | `hasStructuredData` | boolean or null | True if the page has schema.org / microdata markup. |
| 69 | `metaKeywords` | string or null | Content of the meta keywords tag. |
| 70 | `charset` | any | Character set of the page (code or name). |
| 71 | `generator` | string or null | Content of the meta generator tag (CMS). |
| 72 | `faviconUrl` | string or null | Favicon URL. |
| 73 | `socialTags` | object or null | Open Graph and Twitter Card tags as key/value pairs. |
| 74 | `misspelledWords` | array or null | Misspelled words (with "Check spelling" on). |
| 75 | `deprecatedTags` | array or null | Deprecated HTML tags used on the page. |
| 76 | `duplicateMetaTags` | array or null | Meta tags that appear more than once. |
| 77 | `contentEncoding` | string or null | Content-Encoding header, e.g. gzip or br. |
| 78 | `mediaType` | string or null | Content-Type media type, e.g. text/html. |
| 79 | `server` | string or null | Server header. |
| 80 | `cacheable` | boolean or null | True if the page may be cached. |
| 81 | `cacheTtlSeconds` | number or null | Cache lifetime in seconds. |
| 82 | `urlLength` | number or null | URL length in characters. |
| 83 | `lastModified` | object or null | Last-modified dates from the HTTP header, sitemap and meta tag. |
| 84 | `customJsResult` | any | Result of your "Custom JavaScript" snippet. |
| 85 | `checks` | object or null | All 60+ page checks as true/false (with "Include all check results" on). |
| 86 | `lighthousePerformanceScore` | number or null | Lighthouse performance score 0–100 (Lighthouse pages only). |
| 87 | `lighthouseAccessibilityScore` | number or null | Lighthouse accessibility score 0–100. |
| 88 | `lighthouseBestPracticesScore` | number or null | Lighthouse best practices score 0–100. |
| 89 | `lighthouseSeoScore` | number or null | Lighthouse SEO score 0–100. |
| 90 | `lighthouse` | object or null | Lighthouse details: device, scores, FCP, LCP, TBT, CLS, Speed Index, TTI, server response time, failed audits, version. |
| 91 | `changedSinceLastRun` | boolean or null | True if the page is new or changed since the previous run (null on the first run). |
| 92 | `changeType` | string or null | new, changed, unchanged or baseline (first run). |
| 93 | `changedFields` | array | Which fields changed since the previous run, e.g. \["title", "onpageScore"]. |
| 94 | `previousOnpageScore` | number or null | Onpage score in the previous run. |
| 95 | `previousStatusCode` | number or null | Status code in the previous run. |
| 96 | `javascriptRendered` | boolean | True if the page was crawled with JavaScript rendering. |
| 97 | `browserRendered` | boolean | True if the page was crawled with full browser rendering. |
| 98 | `fetchedAt` | string or null | When the page was fetched (ISO 8601). |
| 99 | `scrapedAt` \* | string | When this run started (ISO 8601). |

\* always present and never null. Other fields are `null` (or an empty array) when the value does not apply to the page, e.g. Lighthouse fields on pages without a Lighthouse audit. `checks` is left out when **Include all check results** is off.

**Issue priorities and fix hints:** `issueDetails` gives every issue code a `priority` — `high` (hurts indexing, rankings or visitors directly, e.g. broken pages, `notIndexable`, duplicate content), `medium` (worth fixing soon, e.g. missing meta description, slow server) or `low` (minor, e.g. missing image title attributes) — and a one-sentence `fixHint`. They come from a fixed table per issue code, so they are free and consistent between runs; sort by priority to build a to-do list.

That is 99 fields per page, all declared in the dataset schema (expand the reference above). Plus a **site summary** per website in the key-value store record `SITE_SUMMARY`: overall onpage score, pages crawled, top issues with page counts, broken links/resources, duplicate titles/descriptions/content, non-indexable pages, redirect loops, CMS, server, SSL validity/expiry, sitemap/robots.txt/HTTPS/www-redirect/404 checks, and — with change monitoring — new/changed/unchanged page counts and the previous run's onpage score.

### How to use

1. Open the Actor and add one or more websites to **Websites** (e.g. `https://example.com`). For many sites, use **Bulk edit** to paste a list, or link a text file with one URL per line (up to 100 websites per run).
2. Set **Max pages per website** (prefilled with 20 for a quick first run; 100 if you leave it empty).
3. Optional: turn on **Run Lighthouse audit** or **JavaScript rendering** (for React/Vue/Angular sites).
4. Optional: under **Advanced: crawl scope**, limit the crawl with **Only URLs matching** / **Exclude URLs matching** and pick a **Crawl speed**.
5. Click **Start**. Pages appear in the **Output** tab; the site summary is in the key-value store (`SITE_SUMMARY`).
6. Export as JSON, CSV, Excel, XML or HTML, or read it through the API. Run it again later (or on a schedule) and use the **Changes since last run** view.

```json
{
    "startUrls": ["https://example.com", "https://shop.example.org/blog/"],
    "maxPagesPerSite": 500,
    "runLighthouse": true,
    "lighthousePagesPerSite": 3,
    "enableJavaScript": false,
    "includeUrlPatterns": ["/blog/", "/products/"],
    "excludeUrlPatterns": ["/cart/", "*?sort="],
    "crawlSpeed": "normal",
    "monitorChanges": true,
    "onlyPagesWithIssues": false
}
```

**API — curl**

```bash
curl -X POST "https://api.apify.com/v2/acts/cheapapi~seo-site-audit/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"startUrls": ["https://example.com"], "maxPagesPerSite": 100}'
```

**JavaScript (apify-client)**

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<YOUR_TOKEN>' });
const run = await client.actor('cheapapi/seo-site-audit').call({
    startUrls: ['https://example.com'],
    maxPagesPerSite: 100,
    runLighthouse: true,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
const siteSummary = await client.keyValueStore(run.defaultKeyValueStoreId).getRecord('SITE_SUMMARY');
console.log(items.length, siteSummary.value);
```

**Python (apify-client)**

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_TOKEN>")
run = client.actor("cheapapi/seo-site-audit").call(run_input={
    "startUrls": ["https://example.com"],
    "maxPagesPerSite": 100,
})
pages = client.dataset(run["defaultDatasetId"]).list_items().items
summary = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("SITE_SUMMARY")
print(len(pages), summary["value"])
```

### Use cases

- **Technical SEO audits** for clients or your own sites — export the issue list straight into a spreadsheet.
- **Pre- and post-migration checks**: find broken links, redirect chains, lost canonicals and non-indexable pages.
- **Content quality reviews**: thin content, missing or duplicate titles, descriptions and H1s, low readability.
- **Scheduled monitoring**: run weekly; `changeType` and `changedFields` show which pages are new or changed (title, meta description, H1, canonical, status, indexability, issues, score) since the last run.
- **Section audits**: audit only `/blog/` or `/products/` with include patterns, and skip faceted or cart URLs with exclude patterns.
- **Performance tracking** with Lighthouse and Core Web Vitals for your key landing pages.
- **Lead generation for agencies**: audit prospects' websites and send them the top issues.

### Advanced options

All options are optional. The first four fields (Websites, Max pages per website, Run Lighthouse audit, JavaScript rendering) are all most users need.

| Input | Name | Default | What it does |
|---|---|---|---|
| `includeUrlPatterns` | Only URLs matching | — | Crawl and deliver only URLs whose path matches one of these patterns (e.g. `/blog/`, `/products/*.html$`, or a full URL). `*` = anything, trailing `$` = end of URL. The start page is visited to discover links but only delivered if it matches. Max 50. |
| `excludeUrlPatterns` | Exclude URLs matching | — | Never crawl or deliver matching URLs (e.g. `/cart/`, `*?sort=`, `*.pdf$`). Exclusions always win. Max 50. |
| `crawlSpeed` | Crawl speed | normal | `gentle` = 5 s, `normal` = 2 s, `fast` = 0.5 s, `fastest` = 0.1 s between requests to the website. |
| `maxCrawlDepth` | Max crawl depth | — | How many clicks away from the start page to crawl (0 = start page only). Empty = no limit. |
| `crawlDelayMs` | Exact delay between requests (ms) | — | Exact pause in milliseconds (0–60,000). Overrides **Crawl speed**. |
| `maxCrawlMinutes` | Max crawl time (minutes) | 60 | If the crawl is not finished after this time, it is stopped and the pages crawled so far are delivered. |
| `priorityUrls` | Priority URLs | — | Up to 20 URLs that are crawled first (must belong to one of the websites). |
| `respectSitemap` | Follow sitemap order | off | Crawl pages in the order of the website's XML sitemap. |
| `customSitemapUrl` | Custom sitemap URL | — | Use this sitemap instead of the one listed in robots.txt (e.g. https://example.com/sitemap\_pages.xml). |
| `crawlSitemapOnly` | Crawl sitemap pages only | off | Only audit URLs listed in the sitemap; links on pages are not followed. |
| `includeSubdomains` | Include subdomains | off | Also crawl subdomains (blog.example.com, shop.example.com…). |
| `allowedSubdomains` | Only these subdomains | — | Crawl only these subdomains (e.g. blog.example.com). Leave empty for all. |
| `excludedSubdomains` | Excluded subdomains | — | Subdomains to skip. Needs "Include subdomains". |
| `respectRobotsTxt` | Respect robots.txt | on | Obey the website's robots.txt rules. Turn off to audit pages blocked for crawlers (only on sites you own or may audit). |
| `customRobotsTxt` | Custom robots.txt rules | — | Extra robots.txt rules for the crawl, e.g. "User-agent: \*\nDisallow: /cart/". |
| `robotsTxtMode` | Custom robots.txt mode | — | Merge your rules with the website's robots.txt or replace it. |
| `loadResources` | Load images, scripts & CSS | on | Load page resources to find broken images/scripts/CSS, measure image sizes and render-blocking resources. Included in the page price. |
| `enableBrowserRendering` | Full browser rendering (Core Web Vitals) | off | Render pages in a real browser to measure Core Web Vitals (LCP, CLS, FID) per page. Includes JavaScript rendering. Extra fee per page. |
| `enableXhr` | Allow XHR requests | off | Let pages load data with XHR/fetch during rendering. Needs JavaScript rendering. |
| `disableCookiePopup` | Hide cookie banners | off | Try to close cookie consent pop-ups during rendering. |
| `supportCookies` | Accept cookies | off | Keep cookies between requests (for sites that need a session cookie). |
| `customJavaScript` | Custom JavaScript | — | Snippet executed on every page (max 2,000 characters, 700 ms). Its result is returned in "customJsResult". Example: meta = {}; meta.url = document.URL; meta; |
| `userAgent` | Custom user agent | — | User-Agent header for the crawler and Lighthouse (max 254 characters). Empty = standard crawler user agent. |
| `acceptLanguage` | Accept-Language header | — | Language sent to the website, e.g. "en-US" or "de-DE,de;q=0.9". |
| `device` | Device | — | Screen preset used for rendering. Works with "Full browser rendering". |
| `screenWidth` | Screen width (px) | — | Custom screen width in pixels (240–9999) for Lighthouse, and for the crawl with "Full browser rendering". |
| `screenHeight` | Screen height (px) | — | Custom screen height in pixels (240–9999) for Lighthouse, and for the crawl with "Full browser rendering". |
| `screenScaleFactor` | Screen scale factor | — | Device pixel ratio, 0.5–3, for Lighthouse, and for the crawl with "Full browser rendering". |
| `useExtraCrawlerNetwork` | Use extra crawler network | off | Crawl through an additional pool of IP addresses — helps with sites that block or rate-limit crawlers. |
| `returnSlowPages` | Keep very slow pages | off | Deliver pages that take longer than 120 seconds to load (with the data collected so far) instead of marking them as failed. |
| `titleMinLength` | Title too short below (characters) | 30 | A page title shorter than this number of characters is reported as "Title too short". |
| `titleMaxLength` | Title too long above (characters) | 65 | A page title longer than this number of characters is reported as "Title too long", because search results usually cut it off. |
| `minPageSizeBytes` | Small page below (bytes) | 1024 | An HTML page smaller than this many bytes is reported as "Very small page", which often means an empty or error page. |
| `maxPageSizeBytes` | Large page above (bytes) | 1048576 | An HTML page larger than this many bytes is reported as "Large page size". |
| `minCharacterCount` | Thin content below (characters) | 1024 | A page with fewer visible text characters than this is reported as "Thin content". |
| `maxCharacterCount` | Very long content above (characters) | 256000 | A page with more visible text characters than this is reported as "Very long content". |
| `minTextToHtmlRatio` | Low text-to-HTML ratio below | 0.1 | A page whose visible text is a smaller share of its HTML than this ratio (0–1) is reported as "Low text-to-HTML ratio". |
| `maxTextToHtmlRatio` | High text-to-HTML ratio above | 0.9 | A page whose visible text is a larger share of its HTML than this ratio (0–1) is reported as "High text-to-HTML ratio". |
| `maxLoadTimeMs` | Slow page above (ms) | 3000 | A page that takes longer than this many milliseconds to load is reported as "Slow page load". |
| `maxWaitingTimeMs` | Slow server response above (ms) | 1500 | A page whose server takes longer than this many milliseconds to send the first byte (TTFB) is reported as "Slow server response". |
| `minReadabilityScore` | Low readability below | 15 | A page whose Flesch-Kincaid readability score is below this value is reported as "Low readability" (higher scores mean easier text). |
| `minTitleRelevance` | Title not relevant below | 0.3 | A title whose relevance to the page content (0–1) is below this value is reported as "Title not relevant to page content". |
| `minDescriptionRelevance` | Description not relevant below | 0.2 | A meta description whose relevance to the page content (0–1) is below this value is reported as "Meta description not relevant to page content". |
| `minKeywordsRelevance` | Meta keywords not relevant below | 0.6 | A meta keywords tag whose relevance to the page content (0–1) is below this value is reported as "Meta keywords not relevant to page content". |
| `disabledPageChecks` | Skip these page checks | — | Checks that are not run and do not affect the onpage score. |
| `disabledSiteChecks` | Skip these site-wide checks | — | Site-wide tests to skip. |
| `forceSiteWideChecks` | Run site-wide checks for 1-page audits | off | Also run site-wide checks (sitemap, robots.txt, redirects, 404 page) when only one page is audited. |
| `checkWwwRedirect` | Check www redirect | off | Test whether www and non-www versions redirect to one another. |
| `validateStructuredData` | Validate structured data | off | Check schema.org / microdata markup for errors. |
| `checkSpelling` | Check spelling | off | Find misspelled words on each page (see "misspelledWords"). |
| `spellingLanguage` | Spelling language | — | Language for the spelling check. Empty = detected automatically. |
| `spellingExceptions` | Spelling exceptions | — | Words that are never reported as misspelled (brand names…), max 1,000. |
| `includeBrokenLinkDetails` | List broken links per page | on | Add the list of broken link targets (URL, status code, anchor text) to each page. |
| `onlyPagesWithIssues` | Only pages with issues | off | Deliver (and charge) only pages that have at least one issue. |
| `includeAllChecks` | Include all check results | on | Add the full "checks" object (60+ true/false checks) to each page. |
| `lighthousePagesPerSite` | Lighthouse pages per website | 1 | How many pages per website get a Lighthouse audit (1 = the start page). |
| `lighthouseDevice` | Lighthouse device | mobile | Mobile (default, like Google's mobile-first index) or desktop. |
| `lighthouseCategories` | Lighthouse categories | — | Categories to test. Empty = all four. |
| `lighthouseAudits` | Only these Lighthouse audits | — | Run only specific audits, e.g. "metrics/largest-contentful-paint". Empty = all. |
| `lighthouseLanguage` | Lighthouse report language | en | Language code for audit titles, e.g. "en", "de", "es". Default "en". |
| `lighthouseVersion` | Lighthouse version | — | Specific Lighthouse version (e.g. "12.6.0"). Empty = latest. |
| `lighthouseThrottling` | Network throttling | — | Simulated network speed. Empty = Lighthouse default. |
| `lighthouseThrottlingMethod` | Throttling method | — | How throttling is applied. |
| `lighthouseCpuSlowdown` | CPU slowdown multiplier | — | 1–4, used with the DevTools throttling method. |
| `monitorChanges` | Track changes since the last run | on | Keep a compact fingerprint of every crawled page in a named key-value store and mark each page as new / changed / unchanged next time. Free. |
| `monitorStoreName` | Monitoring store name | seo-site-audit-monitor | Named key-value store for the fingerprints (one record per website). Use different names for separate histories. |

### Output example

Example output — one page row (shortened; values illustrative but consistent with each other and with the site summary below):

```json
{
    "website": "https://example-shop.com",
    "url": "https://example-shop.com/",
    "pageType": "html",
    "statusCode": 200,
    "redirectUrl": null,
    "onpageScore": 97.07,
    "issuesCount": 4,
    "issues": [
        "Links from HTTPS to HTTP pages",
        "Meta keywords not relevant to page content",
        "Low text-to-HTML ratio",
        "Images without alt text"
    ],
    "issueCodes": [
        "httpsToHttpLinks",
        "irrelevantMetaKeywords",
        "lowContentRate",
        "noImageAlt"
    ],
    "issueDetails": [
        { "code": "httpsToHttpLinks", "priority": "medium", "fixHint": "Change links on this HTTPS page that point to http:// URLs to their https:// versions." },
        { "code": "irrelevantMetaKeywords", "priority": "low", "fixHint": "Remove the meta keywords tag (search engines ignore it) or make it match the content." },
        { "code": "lowContentRate", "priority": "low", "fixHint": "Little visible text compared with HTML code. Add content or reduce inline code, scripts and markup." },
        { "code": "noImageAlt", "priority": "medium", "fixHint": "Add descriptive alt text to images that carry meaning; use alt=\"\" for purely decorative images." }
    ],
    "title": "Example Shop – Running Shoes, Trail Shoes & Sports Gear",
    "titleLength": 55,
    "titleMissing": false,
    "titleTooShort": false,
    "titleTooLong": false,
    "duplicateTitle": false,
    "metaDescription": "Shop running shoes, trail shoes and sports gear with free delivery and 30-day returns. Over 2,000 products from top brands.",
    "metaDescriptionLength": 123,
    "metaDescriptionMissing": false,
    "duplicateMetaDescription": false,
    "h1": "Running shoes for every runner",
    "h1Count": 1,
    "h1Missing": false,
    "h2Count": 3,
    "wordCount": 1165,
    "textToHtmlRatio": 0.047,
    "readabilityScore": 31.26,
    "canonicalUrl": "https://example-shop.com/",
    "canonicalIsSelf": true,
    "indexable": true,
    "nonIndexableReason": null,
    "duplicateContent": false,
    "hasBrokenLinks": false,
    "brokenLinksCount": 0,
    "brokenLinks": [],
    "hasBrokenResources": false,
    "clickDepth": 0,
    "internalLinksCount": 130,
    "imagesCount": 58,
    "imagesWithoutAlt": true,
    "renderBlockingScripts": 12,
    "pageSizeBytes": 148362,
    "totalTransferSizeBytes": 2184530,
    "domSize": 162918,
    "loadTimeMs": 486,
    "timeToFirstByteMs": 112,
    "timeToInteractiveMs": 512,
    "largestContentfulPaintMs": null,
    "cumulativeLayoutShift": null,
    "isHttps": true,
    "lighthousePerformanceScore": 45,
    "lighthouseAccessibilityScore": 75,
    "lighthouseBestPracticesScore": 54,
    "lighthouseSeoScore": 85,
    "lighthouse": {
        "device": "mobile",
        "performanceScore": 45,
        "accessibilityScore": 75,
        "bestPracticesScore": 54,
        "seoScore": 85,
        "firstContentfulPaintMs": 1266,
        "largestContentfulPaintMs": 2363,
        "totalBlockingTimeMs": 8342,
        "cumulativeLayoutShift": 0.002,
        "speedIndexMs": 5328,
        "timeToInteractiveMs": 13840,
        "serverResponseTimeMs": 95,
        "failedAudits": [
            {
                "id": "max-potential-fid",
                "title": "Max Potential First Input Delay",
                "score": 0,
                "value": "7,630 ms"
            }
        ],
        "lighthouseVersion": "13.4.0"
    },
    "changedSinceLastRun": true,
    "changeType": "changed",
    "changedFields": ["title", "onpageScore"],
    "previousOnpageScore": 95.2,
    "previousStatusCode": 200,
    "javascriptRendered": false,
    "browserRendered": false,
    "fetchedAt": "2026-09-27T01:52:24.000Z",
    "scrapedAt": "2026-09-28T16:04:14.261Z"
}
```

Example output — the site summary (`SITE_SUMMARY`, one entry per website) for the same 10-page crawl:

```json
{
    "website": "https://example-shop.com",
    "domain": "example-shop.com",
    "crawlFinished": true,
    "crawlStopReason": "page_limit_reached",
    "siteStatus": "no_errors",
    "pagesCrawled": 10,
    "pagesInQueue": 0,
    "pagesDelivered": 10,
    "onpageScore": 95.87,
    "averagePageScore": 95.86,
    "pagesWithIssues": 10,
    "brokenLinks": 0,
    "brokenResources": 0,
    "duplicateTitles": 0,
    "duplicateDescriptions": 0,
    "duplicateContent": 0,
    "nonIndexablePages": 1,
    "redirectLoops": 0,
    "internalLinks": 1248,
    "externalLinks": 188,
    "topIssues": [
        { "issue": "Images without title attribute", "pages": 10 },
        { "issue": "Low text-to-HTML ratio", "pages": 9 },
        { "issue": "Images without alt text", "pages": 7 },
        { "issue": "Links from HTTPS to HTTP pages", "pages": 3 },
        { "issue": "Not indexable", "pages": 1 }
    ],
    "cms": "WordPress 6.6",
    "server": "cloudflare",
    "ip": "203.0.113.10",
    "sslValid": true,
    "sslIssuer": "CN=WE1, O=Google Trust Services, C=US",
    "sslExpires": "2026-12-17T10:33:03.000Z",
    "notFoundStatusCode": 404,
    "wwwRedirectStatusCode": null,
    "crawlStartedAt": "2026-09-27T08:02:15.000Z",
    "crawlEndedAt": "2026-09-27T08:02:37.000Z",
    "lighthousePagesTested": 1,
    "scrapedAt": "2026-09-28T16:04:14.261Z",
    "changeMonitoring": {
        "previousRunAt": "2026-09-21T16:02:51.118Z",
        "previousOnpageScore": 95.4,
        "pagesNew": 1,
        "pagesChanged": 2,
        "pagesUnchanged": 7,
        "pagesNotCrawledAgain": 0
    }
}
```

The dataset has five table views: **Overview**, **Titles, meta & headings**, **Indexability & links**, **Page speed & Lighthouse** and **Changes since last run**. `RUN_SUMMARY` in the key-value store lists skipped websites and errors in plain words.

### Pricing

**Apify Free plan:** Apify does not pay developers for usage on its Free plan, so on the Free plan this Actor can be used for up to **$0.25 of results per calendar month** — enough to try it on a small input. When the allowance is used up, the run ends with a clear message (not an error). Any paid Apify plan removes the limit; prices are the same.

Pay per event — no monthly fee. Prices drop automatically on higher Apify plans (Platinum and Diamond get the Gold price):

| Event | Free | Bronze | Silver | Gold+ | When it is charged |
|---|---|---|---|---|---|
| Audited page | $0.0069 | $0.0055 | $0.0045 | **$0.0034** | Per page delivered to the dataset (loading of images/scripts/CSS included) |
| Website summary | $0.002 | $0.002 | $0.002 | $0.002 | Once per website that delivered at least one page |
| JavaScript rendering (add-on) | $0.002 | $0.002 | $0.002 | $0.002 | Per crawled page (delivered or not), only when JavaScript rendering is on |
| Full browser rendering (add-on) | $0.0065 | $0.0065 | $0.0065 | $0.0065 | Per crawled page (delivered or not), only when full browser rendering is on (JavaScript rendering is charged too) |
| Lighthouse audit | $0.07 | $0.021 | $0.014 | **$0.007** | Per URL with a successful Lighthouse result |
| Crawled page not delivered | $0.0009 | $0.0009 | $0.0009 | $0.0009 | Per page that was crawled but not delivered: left out by **Only pages with issues** or your URL patterns (including the start page visited for link discovery), or when the crawl results could not be read or the crawl start got no answer (the crawl may still have run; its page limit is charged). Never charged when every crawled page is delivered |
| Website crawled without delivered pages | $0.0003 | $0.0003 | $0.0003 | $0.0003 | Once per website whose crawl started but delivered no page (unreachable site, everything filtered out). Replaces the website summary fee |
| Lighthouse test without result | $0.0065 | $0.0065 | $0.0065 | $0.0065 | Per Lighthouse test that ran (or may have run, when no answer arrived) but returned no usable report, e.g. the page did not load |

Worked examples:

- **1,000 pages of one website** (Gold): 1,000 × $0.0034 + $0.002 = **$3.40**. On the Free plan: **$6.90**.
- **100 pages + Lighthouse on the 3 main pages** (Bronze): 100 × $0.0055 + $0.002 + 3 × $0.021 = **$0.615**.
- **200 pages of a React site with JavaScript rendering** (Silver): 200 × ($0.0045 + $0.002) + $0.002 = **$1.302**.
- **Only pages with issues, 500 crawled, 120 with issues** (Gold): 120 × $0.0034 + 380 × $0.0009 + $0.002 = **$0.752** (instead of $1.702 for all 500 pages).
- **A website that cannot be reached** (any plan): $0.0003 + 1 × $0.0009 = **$0.0012**.

Set **Maximum cost per run** in the run options: the Actor limits the number of crawled pages up front so the run never goes over your budget (it reserves the full price of every page a crawl may crawl). Change monitoring, crawls that were rejected (e.g. invalid input) and Lighthouse tests the service could not start are free. The small "not delivered" fees exist because every crawled page and every Lighthouse test is paid for when it runs, even if nothing is delivered; with default settings on a reachable website they are rarely charged.

Apify platform usage is billed separately by Apify (at the default 256 MB a small run typically uses about $0.001–$0.005); the prices and examples on this page are event fees only.

### Integrations

- **Schedules**: run the audit weekly or monthly from Apify Schedules — with change monitoring on, each run tells you what changed since the previous one.
- **Webhooks**: get notified when a run finishes and fetch the dataset automatically.
- **Make, Zapier, n8n**: use the Apify app/node to start audits and push issues to Slack, email or your CRM.
- **Google Sheets**: export the dataset with the Google Sheets integration, or download CSV/Excel.
- **API & MCP**: every run, dataset and the `SITE_SUMMARY` record are available over the Apify API, and AI agents can start audits through the Apify MCP server.

### FAQ

**Is there a limit on the Apify Free plan?** Yes: up to $0.25 of this Actor's results per calendar month, enough to try it. Apify pays developers nothing for Free-plan usage while our data costs are real, so this keeps the Actor sustainable. Runs that reach the allowance stop cleanly and keep everything collected so far; the allowance resets on the 1st of the month. Any paid Apify plan has no limit.

**Is it legal?** Crawling publicly available pages for SEO analysis is generally fine, but only audit websites you own or have permission to audit, respect the website's terms, and keep a gentle crawl speed for sites you don't control. Turning off **Respect robots.txt** should be limited to your own sites.

**How fresh is the data?** Every run crawls the website live, so results reflect the site at the time of the run (`fetchedAt` per page).

**Why did I get fewer pages than "Max pages per website"?** The website may have fewer reachable pages, robots.txt or your URL patterns may exclude parts of it, the crawl may have hit **Max crawl time** or **Max crawl depth**, or your **Maximum cost per run** capped it. Images, scripts and CSS files are checked but not returned as rows. See `RUN_SUMMARY` (including `pagesFilteredOutByUrlPatterns`) and `crawlStopReason` in `SITE_SUMMARY`.

**How do I audit only part of a website?** Put a section URL into **Websites** (e.g. `https://example.com/blog/`) and add `/blog/` to **Only URLs matching**. Use **Exclude URLs matching** for carts, filters, tag pages or PDFs.

**How does change monitoring work?** After each run, a compact fingerprint of every page (status, onpage score, title, meta description, H1, canonical, indexability, word count, issues) is saved in the named key-value store `seo-site-audit-monitor` in your account. The next run of the same website compares against it and fills `changedSinceLastRun`, `changeType` and `changedFields`. The first run is the baseline. It is free; switch it off with **Track changes since the last run**.

**How do I control the budget?** Set **Maximum cost per run** in the run options: the page limit is reduced up front so the run never goes over it. Rendering add-ons and Lighthouse are only charged when you switch them on.

**Why was I charged for "Crawled page not delivered" or "Website crawled without delivered pages"?** The crawler visited those pages, which costs the same whether or not you keep them. With **Only pages with issues** or URL patterns, pages that do not qualify cost $0.0009 each instead of the full page price. A website that could not be reached still costs $0.0003 + one checked page. If the crawl results could not be read, or the crawl start got no answer (the crawl may still have run), the pages up to the crawl's page limit are charged as checked pages; `RUN_SUMMARY` lists `pagesChecked`, `websitesChecked` and the error. With JavaScript or browser rendering on, the rendering add-on applies to checked pages too, because they were rendered.

**Why was I charged $0.0065 for a Lighthouse test without result?** The test ran (or may have run, when no answer arrived) but returned no usable report, for example because the page did not load in the test browser. Tests that could not be started are free. `RUN_SUMMARY.lighthouseChecked` counts them.

**Is Apify platform usage included?** No. Apify platform usage (compute) is billed separately by Apify, on top of the event fees. At the default 256 MB memory a small run typically uses about $0.001–$0.005; long crawls wait longer and use more. Each run's usage is shown in Apify Console.

**Can I schedule it?** Yes — create an Apify Schedule (e.g. every Monday) and a webhook or Make/Zapier scenario that sends you the rows where `changeType` is `new` or `changed`.

**Which export formats are available?** JSON, CSV, Excel, XML, HTML and RSS from the Output tab or the API, plus Google Sheets through the integration.

**How long does an audit take?** Roughly 1–3 minutes for 20 pages and 5–10 minutes for a few hundred pages at the normal crawl speed. Choose **Fast** only for sites that can handle it.

**Something doesn't work — where do I get help?** Open an issue on the Actor's **Issues** tab with the run link and your input (without private data) and we will look into it.

### Limitations

- Lighthouse runs only on crawled HTML pages with status 2xx (the start page first); URLs outside the crawl are not tested.
- The crawler identifies itself with its own user agent; sites with aggressive bot protection may block it (try **Use extra crawler network** or a **Custom user agent**).
- Core Web Vitals per page (`largestContentfulPaintMs`, `cumulativeLayoutShift`, `firstInputDelayMs`) are measured only with **Full browser rendering**; otherwise they are `null`. Lighthouse always measures them.
- Broken link details are listed for up to 5,000 broken links per website (100 per page).
- If the crawl reaches **Max crawl time**, it is stopped and the pages crawled so far are delivered (site-wide duplicate checks may then be incomplete).
- When several websites share a small budget, pages are allocated in input order, so the last websites may be skipped.
- URL patterns match the URL path and query from its start (robots.txt-style `*` and `$`), not full regular expressions. Pages behind excluded URLs can still be reached if other links point to them through included sections.
- Change monitoring compares with the last run of the same website in the same monitoring store; changing thresholds or skipped checks changes onpage scores and shows up as changes. `pagesNotCrawledAgain` can include pages that were simply outside this run's page limit. Up to 40,000 URLs per website are remembered.

### Privacy and personal data

The Actor audits the technical SEO of web pages; it does not look for, extract or build profiles of people. Page text is analysed for word counts and checks but not stored in the output (only titles, meta descriptions, headings and URLs). If a page you audit shows personal data (for example a name in a title), it can appear in those fields — you are responsible for processing it lawfully (e.g. under GDPR/CCPA). The change-monitoring store keeps only hashed fingerprints, scores and status codes per URL in your own Apify account; delete the store at any time to remove them.

# Actor input Schema

## `startUrls` (type: `array`):

Websites to audit — one per row. A domain (example.com) or a URL; a URL with a path (https://example.com/blog/) is used as the crawl starting point. To paste many sites at once use "Bulk edit", or link a text file with one URL per line (up to 100 websites per run).

## `maxPagesPerSite` (type: `integer`):

How many pages to crawl and audit per website. Delivered pages pay the page price; pages that are crawled but not delivered (filters, URL patterns) cost $0.0009 each.

## `runLighthouse` (type: `boolean`):

Also run a Google Lighthouse audit (performance, accessibility, best practices, SEO scores and Core Web Vitals) on the most important pages — the start page by default. Charged per tested page; a test that runs but returns no usable report (e.g. the page does not load) costs $0.0065 ("Lighthouse test without result").

## `enableJavaScript` (type: `boolean`):

Execute JavaScript like a browser before auditing. Turn on for single-page apps (React, Vue, Angular…). Extra fee per crawled page (also for crawled pages that are not delivered).

## `includeUrlPatterns` (type: `array`):

Crawl and deliver only URLs whose path matches one of these patterns, e.g. "/blog/", "/products/\*.html$" or a full URL like "https://example.com/docs/". The start page is still visited to discover links but is only delivered if it matches; crawled pages that do not match cost $0.0009 each ("Crawled page not delivered"). Max 50 patterns.

## `excludeUrlPatterns` (type: `array`):

Never crawl or deliver URLs matching these patterns, e.g. "/cart/", "/tag/", "*?sort=", "*.pdf$". Exclusions always win over "Only URLs matching". A matching page that is still reached (e.g. the start page) is not delivered and costs $0.0009 ("Crawled page not delivered"). Max 50 patterns.

## `crawlSpeed` (type: `string`):

How fast the crawler requests pages from the website. "Normal" pauses 2 seconds between requests; use "Gentle" for small or fragile servers and "Fast"/"Fastest" only for sites you own or that can handle the load. "Delay between requests" overrides this.

## `maxCrawlDepth` (type: `integer`):

How many clicks away from the start page to crawl (0 = start page only). Empty = no limit.

## `crawlDelayMs` (type: `integer`):

Exact pause between requests to the website, in milliseconds (0–60,000). Overrides "Crawl speed". Empty = use "Crawl speed".

## `maxCrawlMinutes` (type: `integer`):

If the crawl is not finished after this time, it is stopped and the pages crawled so far are delivered.

## `priorityUrls` (type: `array`):

Up to 20 URLs that are crawled first (must belong to one of the websites).

## `respectSitemap` (type: `boolean`):

Crawl pages in the order of the website's XML sitemap.

## `customSitemapUrl` (type: `string`):

Use this sitemap instead of the one listed in robots.txt (e.g. https://example.com/sitemap\_pages.xml).

## `crawlSitemapOnly` (type: `boolean`):

Only audit URLs listed in the sitemap; links on pages are not followed.

## `includeSubdomains` (type: `boolean`):

Also crawl subdomains (blog.example.com, shop.example.com…).

## `allowedSubdomains` (type: `array`):

Crawl only these subdomains (e.g. blog.example.com). Leave empty for all.

## `excludedSubdomains` (type: `array`):

Subdomains to skip. Needs "Include subdomains".

## `respectRobotsTxt` (type: `boolean`):

Obey the website's robots.txt rules. Turn off to audit pages blocked for crawlers (only on sites you own or may audit).

## `customRobotsTxt` (type: `string`):

Extra robots.txt rules for the crawl, e.g. "User-agent: \*\nDisallow: /cart/".

## `robotsTxtMode` (type: `string`):

Merge your rules with the website's robots.txt or replace it.

## `loadResources` (type: `boolean`):

Load page resources to find broken images/scripts/CSS, measure image sizes and render-blocking resources. Included in the page price.

## `enableBrowserRendering` (type: `boolean`):

Render pages in a real browser to measure Core Web Vitals (LCP, CLS, FID) per page. Includes JavaScript rendering. Extra fee per crawled page (also for crawled pages that are not delivered).

## `enableXhr` (type: `boolean`):

Let pages load data with XHR/fetch during rendering. Needs JavaScript rendering.

## `disableCookiePopup` (type: `boolean`):

Try to close cookie consent pop-ups during rendering.

## `supportCookies` (type: `boolean`):

Keep cookies between requests (for sites that need a session cookie).

## `customJavaScript` (type: `string`):

Snippet executed on every page (max 2,000 characters, 700 ms). Its result is returned in "customJsResult". Example: meta = {}; meta.url = document.URL; meta;

## `userAgent` (type: `string`):

User-Agent header for the crawler and Lighthouse (max 254 characters). Empty = standard crawler user agent.

## `acceptLanguage` (type: `string`):

Language sent to the website, e.g. "en-US" or "de-DE,de;q=0.9".

## `device` (type: `string`):

Screen preset used for rendering. Works with "Full browser rendering".

## `screenWidth` (type: `integer`):

Custom screen width in pixels (240–9999) for Lighthouse, and for the crawl with "Full browser rendering".

## `screenHeight` (type: `integer`):

Custom screen height in pixels (240–9999) for Lighthouse, and for the crawl with "Full browser rendering".

## `screenScaleFactor` (type: `number`):

Device pixel ratio, 0.5–3, for Lighthouse, and for the crawl with "Full browser rendering".

## `useExtraCrawlerNetwork` (type: `boolean`):

Crawl through an additional pool of IP addresses — helps with sites that block or rate-limit crawlers.

## `returnSlowPages` (type: `boolean`):

Deliver pages that take longer than 120 seconds to load (with the data collected so far) instead of marking them as failed.

## `titleMinLength` (type: `integer`):

A page title shorter than this number of characters is reported as "Title too short". The default of 30 characters suits most sites.

## `titleMaxLength` (type: `integer`):

A page title longer than this number of characters is reported as "Title too long", because search results usually cut it off. The default is 65 characters.

## `minPageSizeBytes` (type: `integer`):

An HTML page smaller than this many bytes is reported as "Very small page", which often means an empty or error page. The default is 1,024 bytes (1 KB).

## `maxPageSizeBytes` (type: `integer`):

An HTML page larger than this many bytes is reported as "Large page size". The default is 1,048,576 bytes (1 MB).

## `minCharacterCount` (type: `integer`):

A page with fewer visible text characters than this is reported as "Thin content". The default is 1,024 characters.

## `maxCharacterCount` (type: `integer`):

A page with more visible text characters than this is reported as "Very long content". The default is 256,000 characters.

## `minTextToHtmlRatio` (type: `number`):

A page whose visible text is a smaller share of its HTML than this ratio (0–1) is reported as "Low text-to-HTML ratio". The default is 0.1 (10%).

## `maxTextToHtmlRatio` (type: `number`):

A page whose visible text is a larger share of its HTML than this ratio (0–1) is reported as "High text-to-HTML ratio". The default is 0.9 (90%).

## `maxLoadTimeMs` (type: `integer`):

A page that takes longer than this many milliseconds to load is reported as "Slow page load". The default is 3,000 ms (3 seconds).

## `maxWaitingTimeMs` (type: `integer`):

A page whose server takes longer than this many milliseconds to send the first byte (TTFB) is reported as "Slow server response". The default is 1,500 ms.

## `minReadabilityScore` (type: `number`):

A page whose Flesch-Kincaid readability score is below this value is reported as "Low readability" (higher scores mean easier text). The default is 15.

## `minTitleRelevance` (type: `number`):

A title whose relevance to the page content (0–1) is below this value is reported as "Title not relevant to page content". The default is 0.3.

## `minDescriptionRelevance` (type: `number`):

A meta description whose relevance to the page content (0–1) is below this value is reported as "Meta description not relevant to page content". The default is 0.2.

## `minKeywordsRelevance` (type: `number`):

A meta keywords tag whose relevance to the page content (0–1) is below this value is reported as "Meta keywords not relevant to page content". The default is 0.6.

## `disabledPageChecks` (type: `array`):

Checks that are not run and do not affect the onpage score.

## `disabledSiteChecks` (type: `array`):

Site-wide tests to skip.

## `forceSiteWideChecks` (type: `boolean`):

Also run site-wide checks (sitemap, robots.txt, redirects, 404 page) when only one page is audited.

## `checkWwwRedirect` (type: `boolean`):

Test whether www and non-www versions redirect to one another.

## `validateStructuredData` (type: `boolean`):

Check schema.org / microdata markup for errors.

## `checkSpelling` (type: `boolean`):

Find misspelled words on each page (see "misspelledWords").

## `spellingLanguage` (type: `string`):

Language for the spelling check. Empty = detected automatically.

## `spellingExceptions` (type: `array`):

Words that are never reported as misspelled (brand names…), max 1,000.

## `includeBrokenLinkDetails` (type: `boolean`):

Add the list of broken link targets (URL, status code, anchor text) to each page.

## `onlyPagesWithIssues` (type: `boolean`):

Deliver only pages that have at least one issue. Crawled pages without issues are still crawled and cost $0.0009 each as "Crawled page not delivered" (instead of the full page price).

## `includeAllChecks` (type: `boolean`):

Add the full "checks" object (60+ true/false checks) to each page.

## `lighthousePagesPerSite` (type: `integer`):

How many pages per website get a Lighthouse audit (1 = the start page).

## `lighthouseDevice` (type: `string`):

Mobile (default, like Google's mobile-first index) or desktop.

## `lighthouseCategories` (type: `array`):

Categories to test. Empty = all four.

## `lighthouseAudits` (type: `array`):

Run only specific audits, e.g. "metrics/largest-contentful-paint". Empty = all.

## `lighthouseLanguage` (type: `string`):

Language code for audit titles, e.g. "en", "de", "es". Default "en".

## `lighthouseVersion` (type: `string`):

Specific Lighthouse version (e.g. "12.6.0"). Empty = latest.

## `lighthouseThrottling` (type: `string`):

Simulated network speed. Empty = Lighthouse default.

## `lighthouseThrottlingMethod` (type: `string`):

How throttling is applied.

## `lighthouseCpuSlowdown` (type: `number`):

1–4, used with the DevTools throttling method.

## `monitorChanges` (type: `boolean`):

Save a compact fingerprint of every crawled page (status, score, title, meta description, H1, canonical, indexability, word count, issues) in a named key-value store, and mark each page as new, changed or unchanged compared with the previous run ("changedSinceLastRun", "changeType", "changedFields"). The first run creates the baseline.

## `monitorStoreName` (type: `string`):

Named key-value store in your account that keeps the page fingerprints (lowercase letters, digits, hyphens). Use different names to keep separate histories, e.g. one per client.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://books.toscrape.com"
    }
  ],
  "maxPagesPerSite": 20,
  "runLighthouse": false,
  "enableJavaScript": false,
  "crawlSpeed": "normal",
  "maxCrawlMinutes": 60,
  "respectSitemap": false,
  "crawlSitemapOnly": false,
  "includeSubdomains": false,
  "respectRobotsTxt": true,
  "loadResources": true,
  "enableBrowserRendering": false,
  "enableXhr": false,
  "disableCookiePopup": false,
  "supportCookies": false,
  "useExtraCrawlerNetwork": false,
  "returnSlowPages": false,
  "titleMinLength": 30,
  "titleMaxLength": 65,
  "minPageSizeBytes": 1024,
  "maxPageSizeBytes": 1048576,
  "minCharacterCount": 1024,
  "maxCharacterCount": 256000,
  "minTextToHtmlRatio": 0.1,
  "maxTextToHtmlRatio": 0.9,
  "maxLoadTimeMs": 3000,
  "maxWaitingTimeMs": 1500,
  "minReadabilityScore": 15,
  "minTitleRelevance": 0.3,
  "minDescriptionRelevance": 0.2,
  "minKeywordsRelevance": 0.6,
  "forceSiteWideChecks": false,
  "checkWwwRedirect": false,
  "validateStructuredData": false,
  "checkSpelling": false,
  "includeBrokenLinkDetails": true,
  "onlyPagesWithIssues": false,
  "includeAllChecks": true,
  "lighthousePagesPerSite": 1,
  "lighthouseDevice": "mobile",
  "lighthouseLanguage": "en",
  "monitorChanges": true,
  "monitorStoreName": "seo-site-audit-monitor"
}
```

# Actor output Schema

## `pages` (type: `string`):

No description

## `changes` (type: `string`):

No description

## `siteSummary` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://books.toscrape.com"
        }
    ],
    "maxPagesPerSite": 20,
    "runLighthouse": false,
    "crawlSpeed": "normal"
};

// Run the Actor and wait for it to finish
const run = await client.actor("cheapapi/seo-site-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://books.toscrape.com" }],
    "maxPagesPerSite": 20,
    "runLighthouse": False,
    "crawlSpeed": "normal",
}

# Run the Actor and wait for it to finish
run = client.actor("cheapapi/seo-site-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://books.toscrape.com"
    }
  ],
  "maxPagesPerSite": 20,
  "runLighthouse": false,
  "crawlSpeed": "normal"
}' |
apify call cheapapi/seo-site-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cheapapi/seo-site-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/EWMxJLSLEEYDBaCXm/builds/i5zOk0WEKcZ5y8ddr/openapi.json
