# Technical SEO Audit Tool - Website SEO Checker & Site Audit (`tidytools/seo-audit-crawler`) Actor

SEO crawler for technical SEO audits, JS sites too: 0-100 SEO score and fix hints per page, site report (duplicate titles, broken pages, missing H1, meta, alt), optional Core Web Vitals. $5/1k pages.

- **URL**: https://apify.com/tidytools/seo-audit-crawler.md
- **Developed by:** [Yukai Lin](https://apify.com/tidytools) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does SEO Audit Tool do?

It crawls your website and checks **every page for on-page SEO problems**, including **sites built with JavaScript or protected against bots**: when a site blocks plain requests, the page is fetched from a second network or audited in a real browser automatically, and a page that loads almost empty until JavaScript runs (a client-rendered app shell) is rendered in a real browser too. Each page gets a **score from 0 to 100**, **category scores** and a list of issues with **a fix hint for every issue**. The whole site gets a **report**: score distribution, most common problems, crawl coverage (indexable / noindex / canonicalized / 4xx / 5xx), sitemap checks, canonical and redirect problems, duplicate titles, H1s and content, and broken internal pages.

**25+ checks per page**, including:

- 🏷️ **Title**: missing, too short, too long (search results cut titles at about 60 characters)
- 📝 **Meta description**: missing, too short, too long
- 🔠 **Headings**: missing, empty or multiple H1s, skipped heading levels (H2 → H4), full heading outline
- 🔗 **Canonical & indexing**: missing or conflicting canonical, `noindex` in meta robots or the `X-Robots-Tag` header, pages that robots.txt blocks for Googlebot
- 🖼️ **Images**: images without an alt attribute (`alt=""` on decorative images is correctly accepted); optional check for broken and oversized images
- 🌍 **hreflang**: invalid codes (`en-UK`, `jp`), missing `x-default`, missing self-reference, duplicate codes, broken or redirecting targets, missing return links
- 📱 **Mobile & language**: viewport tag, `lang` attribute
- 🌐 **Social & structured data**: Open Graph title/image, Twitter card, and **JSON-LD and Microdata validated** against Google's rich result rules (errors, eligible rich results; RDFa types listed)
- ⚙️ **Technical**: HTTP status, HTTPS, redirect chains, slow server response, very large HTML, thin content (word count)
- 💔 **Broken pages**: internal links that return 404/410/5xx, with the page that links to them

**Site-level checks** (in the report):

- 🗺️ **Sitemap**: no sitemap found, sitemap URLs that return errors, redirect or are `noindex`, crawled indexable pages missing from the sitemap
- ↪️ **Canonical tags** that point to an error page, a redirect or a `noindex` page
- 🔁 **Internal links that point to a redirect** instead of the final URL
- 👯 **Duplicates**: titles, meta descriptions, H1s and identical page text (duplicate content)
- 📊 **Crawl coverage**: indexable, noindex, canonicalized, redirected, 4xx, 5xx, blocked by robots.txt, deepest page

**Optional:**

- ⚡ **Core Web Vitals** with Google PageSpeed Insights, **no API key needed**: performance score, LCP, CLS, TBT, FCP and real-user INP for the pages you choose ($5 per 1,000 pages measured)
- 📈 **Monitoring**: compare with the previous run and see issues added or resolved per page and the score change

### Why use it?

- **Works on JavaScript sites and sites that block bots.** Many SEO crawlers analyze static HTML only; this one falls back to a real browser automatically when a site blocks plain requests or a page has almost no text until JavaScript runs.
- **A fix hint for every issue**, plus category scores (technical, indexing, meta, headings, content, links, images, social, schema, performance, mobile, international).
- **$5 per 1,000 pages**, no extra compute charges. See the price comparison below.
- **Pay only for audited pages**: broken pages are reported **for free**; blocked pages, pages skipped by robots.txt and downloads are not charged.
- **Built for recurring checks**: schedule it weekly with **Compare with the previous run** to get issues added and resolved per page, score changes and new broken pages, or feed the JSON into your dashboard.

### How much does it cost?

| Event | Price |
|---|---|
| Audited page | **$5.00 / 1,000 pages** |
| PageSpeed measurement (optional) | **$5.00 / 1,000 pages measured** |

**Example:** auditing a 1,000-page site costs 1,000 × $0.005 = **$5.00** on the Free plan; PageSpeed on 50 of those pages adds 50 × $0.005 = $0.25.

**No start fee.** Broken pages, failed pages and the site checks (sitemap, canonical targets, image checks, structured data, run-to-run comparison) are free. PageSpeed measurements are only charged when a score comes back. Your maximum charge limit is always respected.

#### Control your cost

- **Charged:** one `Audited page` event per HTML page audited (`charged: true` on the row); PageSpeed only when a score comes back.
- **Free** (`charged: false`, the `error` ends with "(not charged)"): invalid input lines (`errorType: "invalid_input"`), broken internal pages (a finding), blocked, failed and timed-out pages, downloads (PDF, images) and pages skipped by robots.txt.
- **Before it starts**, the run logs its plan: **Max pages** × price (plus PageSpeed pages if on) is the most it can cost, compared with your **maximum charge per run** (run options). Small sites have fewer pages and cost less. The plan is saved as `costPlan` in `SUMMARY`.
- **When the maximum is reached**, the run stops: the status message says so, `SUMMARY.status` is `LIMIT_REACHED` and `SUMMARY.notProcessed` has the number of pages found but not audited and the first 100 of them.
- **If Apify restarts the run** (server migration or Resurrect), pages audited before it are read again for the report but not charged again and not written to the dataset twice (`SUMMARY.resumed`).

#### Price comparison (checked September 2026)

| Actor | Price per 1,000 pages (free plan) | Crawls the site | Real-browser fallback |
|---|---|---|---|
| **SEO Audit Crawler (this Actor)** | **$5** | Yes | Yes |
| smart-digital/complete-seo-audit-tool | $40 | Yes | No (static HTML) |
| logiover/website-seo-audit-crawler | $10 | Yes | No (plain HTTP only) |
| autofacts/metadata-scraper | $5; the max charge per run must be set to at least $0.10 | Only with `followLinks` (off by default) | No (no JavaScript) |
| scrapesignal\_labs/website-seo-audit | **$0.50** + platform usage | Yes | No |

We are not the cheapest: scrapesignal\_labs costs less per page, but you also pay the Apify compute it uses, and it has no browser fallback. At $5 this Actor is the one with a real-browser fallback, a fix hint per issue and site-level checks (sitemap, canonical targets, duplicates, hreflang return links). For PageSpeed, lizaraco/site-audit charges $10 per 1,000 measurements; this Actor charges $5.

### How to use it

1. Enter your **website URL** in **URLs or domains**, e.g. `https://www.example.com/` or just `example.com` (one per line for several sites). A bad line never stops the run: it becomes a row with `errorType: "invalid_input"` (not charged).
2. Set **Max pages** (100 is a good start).
3. Optional: turn on **Also audit pages from the sitemap**, **Respect robots.txt**, **Check images**, **Measure Core Web Vitals** or **Compare with the previous run**, or limit the crawl with include/exclude URL patterns.
4. Click **Start**. Open **Site report** in the Output tab for the summary, or the **Pages**, **Issues and fixes** and **Category scores** tables for details.

#### Input example

```json
{
    "urls": ["https://www.python.org/"],
    "crawlScope": "domain",
    "maxPages": 100,
    "useSitemap": true,
    "excludeUrlPatterns": ["**/blog/tag/**"]
}
```

**Crawl scope**: *Whole website* includes subdomains (blog.example.com, shop.example.com); *Only this host* stays on the start URL's host (`www.` and the bare domain count as one); *Same folder* and *Only the start URLs* narrow it further.

Use `urls` (a plain list of strings) in API calls. The older `startUrls` field (`[{ "url": "https://…" }]`) still works and can hold uploaded or linked URL list files, but Apify rejects the whole request with HTTP 400 if one of its entries is a bare domain or a blank line.

**Several sites in one run (agencies).** Put one site per line in **URLs or domains**. **Max pages** is shared fairly: each site gets its share (Max pages ÷ number of sites) while the others are still being crawled, and a small site leaves its unused share to the others. Set **Max pages per site** for a fixed cap. Every row has `site` (host without `www.`) and `startUrl`, the sitemap of every site is read (up to 20 sites), `SUMMARY.sites` has one entry per site (`pagesAudited`, `averageScore`, `brokenPages`, `pagesFailed`, `sitemapFound`, `sitemapUrls`) and the REPORT has a **Per site** table. Real run (September 2026, `maxPages: 12`): python.org, books.toscrape.com and webscraper.io got 4 pages each.

```json
{
    "urls": ["https://www.python.org/", "books.toscrape.com", "https://webscraper.io/"],
    "maxPages": 12
}
```

#### Output: one item per page

Real output for https://www.python.org/ (shortened):

```json
{
    "url": "https://www.python.org/",
    "site": "python.org",
    "startUrl": "https://www.python.org/",
    "input": "https://www.python.org/",
    "success": true,
    "mode": "fast",
    "via": "direct",
    "score": 93,
    "issueCount": 2,
    "issueCodes": "h1-multiple, canonical-missing",
    "issues": [
        { "id": "h1-multiple", "severity": "warning", "message": "Page has 5 H1 headings.", "category": "headings",
          "fixHint": "Keep a single main H1 and turn the others into H2 headings.", "impact": "medium" },
        { "id": "canonical-missing", "severity": "notice", "message": "Canonical link is missing.", "category": "indexing",
          "fixHint": "Add <link rel=\"canonical\"> pointing to the preferred URL of this page (usually itself).", "impact": "low" }
    ],
    "categoryScores": { "technical": 100, "indexing": 94, "meta": 100, "headings": 85, "images": 100, "social": 100, "...": 100 },
    "httpStatus": 200,
    "title": "Welcome to Python.org",
    "titleLength": 21,
    "metaDescriptionLength": 52,
    "h1": ["Intuitive Interpretation", "Compound Data Types", "..."],
    "headings": [{ "level": 1, "text": "Intuitive Interpretation" }, "..."],
    "canonical": null,
    "xRobotsTag": null,
    "blockedByRobotsTxt": false,
    "viewport": true,
    "wordCount": 1182,
    "contentHash": "9f824c550ac6dbe88bd0688e",
    "bytes": 53229,
    "imagesMissingAlt": 0,
    "nofollowLinks": 0,
    "hreflangLinks": [],
    "twitterCard": null,
    "structuredDataTypes": ["WebSite"]
}
```

`via` tells you where the page was read from: `backend` (our servers), `direct` (Apify's network, used when a site blocks our servers), `vm` (our second server, Oracle, different IP) or `browser`. Failed pages carry an `errorType` (`blocked`, `not_found`, `timeout`, `network`…). `input` (on start pages) is your original line; `issueCodes` is the issue IDs as one text column for spreadsheets; `charged` says whether the row was charged.

#### Output: site report

The key-value store contains `REPORT` (Markdown, easy to read or share) and `SUMMARY` (JSON). From the same python.org run (20 pages):

```json
{
    "status": "SUCCESS",
    "pagesAudited": 20,
    "averageScore": 94.4,
    "scoreDistribution": { "80-100": 20, "60-79": 0, "40-59": 0, "0-39": 0 },
    "categoryAverages": { "technical": 100, "indexing": 94, "meta": 99.7, "headings": 89.5, "...": 100 },
    "site": {
        "coverage": { "crawled": 20, "indexable": 20, "noindex": 0, "canonicalized": 0, "redirected": 3, "clientErrors4xx": 0, "serverErrors5xx": 0, "blockedByRobotsTxt": 0, "maxDepth": 1 },
        "sitemap": { "found": false, "urls": 0 },
        "internalLinksToRedirects": [
            { "url": "https://www.python.org/psf/", "finalUrl": "https://www.python.org/psf-landing/", "linkedFrom": "https://www.python.org/" },
            { "url": "https://www.python.org/doc/av", "finalUrl": "https://www.python.org/doc/av/", "linkedFrom": "https://www.python.org/" }
        ],
        "canonicalProblems": [],
        "duplicateH1": [],
        "duplicateContent": [],
        "hreflangProblems": []
    }
}
```

`status` is `SUCCESS`, `PARTIAL_RESULTS` (some pages failed), `FAILED`, `NO_RESULTS` or `LIMIT_REACHED` (stopped at your spending limit). The summary also lists the most common issues (with fix hints), the lowest-scoring pages, duplicate titles and descriptions, and broken internal pages with the page that links to them. It also has `sites` (one entry per site), `invalidInputs`, `costPlan` and, when the run stopped at your limit, `notProcessed`.

#### Optional: PageSpeed Insights (Core Web Vitals)

Turn on **Measure Core Web Vitals (PageSpeed Insights)** to run Google's PageSpeed Insights on the start URL(s) and the next pages audited, up to **Max pages to measure** (default 5). **No setup is needed**: without your own key, our servers run Google PageSpeed Insights with our key, and when Google's quota is busy they measure with our own Lighthouse server instead. With your own key, the Actor calls Google directly.

**Where to measure from** (`pagespeedSource`): lab results depend on where Lighthouse runs, so you can pick the location. `auto` (default) uses Google, then our Lighthouse servers; `google` uses Google only; `lighthouse` uses our servers only (Taiwan first, then Phoenix, US); `lighthouse-tw` and `lighthouse-us` measure from that one location and are never swapped for another (if it is unavailable, the page is reported as not measured and is not charged). Choose Taiwan when your visitors are in Asia. Our servers give lab data only; real-user field data comes from Google. The price is the same for every option.

Each page gets a `pagespeed` object:

| Field | Meaning |
|---|---|
| `performanceScore` | Lighthouse performance score, 0-100 |
| `lcpMs`, `cls`, `tbtMs`, `fcpMs`, `speedIndexMs` | Lab measurements: Largest Contentful Paint, Cumulative Layout Shift, Total Blocking Time, First Contentful Paint, Speed Index |
| `inpMs` | Interaction to Next Paint from real-user data (lab tests cannot measure INP), when Google has enough traffic for the page or its site |
| `field` | Real-user (Chrome UX Report) 75th percentiles: `lcpMs`, `inpMs`, `cls`, `fcpMs`, `ttfbMs`, `category`, and `source` (`url` or `origin`) |
| `source` | Who measured: `google` (Google with our key), `vm-lighthouse` (our Lighthouse servers: lab data only, no `field`/`inpMs`) or `google (your key)` |
| `node`, `nodeLocation` | With `vm-lighthouse`: which of our servers measured, `tw` (Taiwan) or `us` (Phoenix, US) |
| `ttiMs`, `opportunities` | Time to interactive and the top savings suggestions (id, title, estimated ms/bytes saved); measured through our servers |
| `error`, `errorKind` | Why a page was not measured (`quota`, `key`, `page`, `api`, `skipped`); such pages are **not charged** |

Price: **$5 per 1,000 pages measured** (`pagespeed-audit` event), charged only when a score comes back.

Real result (September 2026, mobile, 3 of 5 audited python.org pages measured with `pagespeedMaxPages: 3`; charged 5 × page audit + 3 × PageSpeed):

```json
{
    "url": "https://www.python.org/",
    "score": 93,
    "pagespeed": {
        "strategy": "mobile",
        "performanceScore": 78,
        "lcpMs": 4802,
        "cls": 0.003,
        "tbtMs": 0,
        "fcpMs": 2718,
        "speedIndexMs": 2718,
        "inpMs": 88,
        "field": { "source": "url", "category": "FAST", "lcpMs": 1225, "inpMs": 88, "cls": 0, "fcpMs": 1158, "ttfbMs": 526 },
        "lighthouseVersion": "13.5.0",
        "measuredUrl": "https://www.python.org/"
    }
}
```

The lab LCP (4.8 s) is much slower than what real visitors see (`field.lcpMs` 1.2 s): Lighthouse simulates a slow phone on a throttled mobile network. Use the lab numbers to compare pages and track changes, and the `field` numbers for what your visitors actually experience. Lab scores also vary a few points between runs.

##### Optional: use your own Google API key (about 5 minutes)

You do not need a key. Add one only if you want the calls to count against your own Google quota (for example very large runs): put it in **Your own Google API key**. If our daily capacity is ever used up, the remaining pages are marked `errorKind: "quota"` and not charged, and adding your key lets you measure them.

1. Open [Google Cloud Console](https://console.cloud.google.com/) and create a project (no billing account needed).
2. Enable the [PageSpeed Insights API](https://console.cloud.google.com/apis/library/pagespeedonline.googleapis.com).
3. Go to **APIs & Services → Credentials → Create credentials → API key**.
4. Recommended: edit the key, keep **Application restrictions: None** (Apify servers have changing IPs) and set **API restrictions** to *PageSpeed Insights API* only.
5. Paste the key into the Actor input. It is stored encrypted.

The API is free; your daily quota is shown in Google Cloud Console under the PageSpeed Insights API's **Quotas** page.

After a quota or key error the Actor stops measuring for the rest of the run (those pages are not charged). `SUMMARY.pagespeed` has the average performance score and the pages sorted from slowest, and the REPORT gets a PageSpeed table.

#### Structured data (JSON-LD and Microdata)

Every page gets a `structuredData` object checked with the same rules as **Schema Markup Validator** (by TidyTools): `status`, `formats` (`json-ld`, `microdata`), `types`, eligible `richResults`, `errorCount`, `warningCount`, the first errors with code and path, and `rdfaTypes` (RDFa is detected, not validated). It does not change the page score. Real result from webscraper.io (September 2026), a page marked up with Microdata only:

```json
"structuredData": {
    "status": "errors",
    "formats": ["microdata"],
    "types": ["OfferCatalog", "VideoObject"],
    "richResults": [],
    "errorCount": 2,
    "warningCount": 2,
    "errors": [
        { "code": "REQUIRED_MISSING", "type": "VideoObject", "property": "thumbnailUrl", "path": "microdata[1]", "format": "microdata" },
        { "code": "REQUIRED_MISSING", "type": "VideoObject", "property": "uploadDate", "path": "microdata[1]", "format": "microdata" }
    ],
    "rdfaTypes": []
}
```

Turn off **Validate structured data** to skip it.

#### Monitoring: compare with the previous run

Turn on **Compare with the previous run** and schedule the Actor (for example weekly). Each page gets `change` (`new`, `changed`, `unchanged`), `previousScore`, `scoreDelta`, `issuesAdded` and `issuesResolved`. `SUMMARY.changes` and a **Since the last run** section in the REPORT show the average score change, issues added and resolved (by issue ID), the biggest score drops and gains, new broken pages, fixed broken pages and pages no longer found. Runs with the same **Monitor name** are compared (default: start URL + crawl scope). The first run is the baseline. If a run stops before the whole site is crawled (Max pages or spending limit), pages it did not reach are kept and not reported as removed.

Real second run on webscraper.io (September 2026; for the test, the previous snapshot of two pages was edited):

```json
[
    { "url": "https://webscraper.io/", "score": 96, "change": "changed", "previousScore": 90, "scoreDelta": 6,
      "issuesAdded": ["structured-data-missing"], "issuesResolved": ["title-missing"] },
    { "url": "https://webscraper.io/pricing", "score": 94, "change": "changed", "previousScore": 100, "scoreDelta": -6,
      "issuesAdded": ["canonical-other", "heading-skip", "structured-data-missing"], "issuesResolved": [] }
]
```

REPORT excerpt from the same run:

```
### Since the last run

- Average score: **97.3** (+2.3 from 95)
- Issues added: **4**, resolved: **1** on 6 page(s) seen in both runs
- New pages: **0**, pages no longer found: **0**, new broken pages: **0**, broken pages fixed: **1**
```

Pages are audited and charged the same way with or without monitoring.

### Use with AI agents (MCP)

Connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/seo-audit-crawler) to Claude, Cursor or any MCP client, then ask e.g. "Audit the first 50 pages of example.com and list the three fixes that would raise the score most."

```json
{ "urls": ["https://example.com/"], "maxPages": 50 }
```

Failed pages are not charged and carry an `errorType`, so the agent can tell a blocked site from a broken page.

### Use it from code and integrations

Run it from your own code with the Apify API. This call waits for the run and returns the results as JSON (replace `YOUR_TOKEN` with your Apify API token):

```bash
curl -X POST "https://api.apify.com/v2/acts/tidytools~seo-audit-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://www.python.org/"],"maxPages":20}'
```

The synchronous endpoint waits up to 5 minutes. For bigger runs, start the run with `POST https://api.apify.com/v2/acts/tidytools~seo-audit-crawler/runs` and read the dataset when it finishes, or use the `apify-client` package for JavaScript or Python.

**Schedules and integrations:** run it daily or weekly with Apify Schedules, get a webhook when a run finishes, or send the results to Zapier, Make, n8n, Google Sheets, Slack and other apps with Apify integrations. Results can be exported as JSON, CSV, Excel or XML.

### Scoring

Each page starts at 100. Critical issues (e.g. missing title, missing H1, `noindex`, HTTP errors, no HTTPS) cost 15 points, warnings (e.g. missing meta description, images without alt text, invalid hreflang) cost 5, and notices (e.g. missing canonical, Open Graph or structured data) cost 2. Each category score starts at 100 and loses three times those points for issues in that category.

Redirects that only add a language or country folder (`/pricing` → `/en-sg/pricing`) depend on where the request comes from, so they are counted separately and not reported as problems.

### Advanced settings

- **Plain HTTP requests from**: *Auto* (recommended) reads pages from our servers and retries from Apify's network when a site blocks data-center requests, then from our second server (Oracle, different IP), before falling back to a browser. You can force any one of them.
- **Time limit per page**: default 120 seconds for reading a page, all retries and the browser fallback included. A slower page gets `errorType: "timeout"` and is not charged. 0 = no limit.
- Image checks (*Check images*) count toward the same page time limit. When it runs out, the page is still audited, and `imageCheck` has `incomplete: true` and the number of images not checked.
- **Proxy**: optional Apify Proxy for requests sent from Apify's network (billed to your Apify account).

### Limitations

- On-page checks: no backlinks or keyword rankings. Core Web Vitals only with the optional PageSpeed Insights measurement (no key needed).
- JSON-LD and Microdata are validated; RDFa types are only listed.
- Only public `http`/`https` pages; no logins.
- Download files (PDF, ZIP, installers) are skipped.
- Sitemap URLs outside the crawl are checked as a sample (up to 100 per run).

### FAQ

**Does it measure Core Web Vitals?** Yes, optionally. Turn on PageSpeed Insights for the pages you choose to get the performance score, LCP, CLS, TBT, FCP and real-user INP. No API key is needed; it costs $5 per 1,000 pages measured.

**Can it audit JavaScript websites?** Yes. In *Auto* mode a page that comes back almost empty (under about 200 characters of visible text, typical for client-rendered React or Vue apps) is audited again in a real browser, and a site that blocks plain requests is fetched from a second network or a browser (`mode: "browser"` on the row). Text or tags that JavaScript adds to a page that already has content are only seen with **Page fetching (HTTP or browser)** set to *Browser*.

**Does it check backlinks or keyword rankings?** No. It is a technical and on-page SEO checker: titles, meta, headings, canonical, hreflang, structured data, broken pages, sitemap and duplicates.

**Can I track SEO issues over time?** Yes. With **Compare with the previous run** on, each run is compared with the previous one and shows the issues added or resolved per page and the score change.

### Support

Open an issue in the **Issues** tab with your start URL and input. Issues are checked regularly.

# Actor input Schema

## `urls` (type: `array`):

Main input (fill this, or `startUrls`). Start pages to audit, one per line: full URLs or bare domains like example.com. Several sites (e.g. one per client) share Max pages fairly and are reported site by site. Up to 100 start URLs. Invalid lines become a row with errorType "invalid\_input" (not charged) instead of failing the run.

## `startUrls` (type: `array`):

Older start URL field, kept for existing API calls. Use it to upload or link a .txt/.csv file of start URLs. Full URLs only here: bare domains such as example.com go in "URLs or domains" above. Both fields are combined (up to 100 start URLs).

## `crawlScope` (type: `string`):

Which discovered links to follow. "Whole website" includes subdomains such as blog. or shop.; "Only this host" stays on the start URL's host (www. and the bare domain count as one). Values: domain = Whole website (incl. subdomains, e.g. blog.example.com); host = Only this host (no other subdomains); path = Same folder as the start URL; page = Only the start URLs.

## `maxPages` (type: `integer`):

Stop after auditing this many pages in total (all sites together). You pay only for pages audited.

## `maxPagesPerSite` (type: `integer`):

With several sites: audit at most this many pages of each site. Leave empty to share Max pages evenly between the sites (a site with fewer pages leaves its share to the others).

## `maxDepth` (type: `integer`):

How many clicks away from the start URL to go.

## `mode` (type: `string`):

"Auto" audits the raw HTML (what search engines fetch first) and uses a real browser only for sites that block plain requests. "Browser" always audits the JavaScript-rendered page. Values: auto = Auto (recommended); fast = Raw HTML only; browser = Browser (rendered page).

## `includeUrlPatterns` (type: `array`):

Optional glob patterns; only matching URLs are audited, e.g. https://example.com/blog/\*\*

## `excludeUrlPatterns` (type: `array`):

Optional glob patterns to skip, e.g. **/tag/** or \**?page=*

## `useSitemap` (type: `boolean`):

Add the URLs listed in the site's XML sitemap to the crawl, so pages that are not linked from the navigation are audited too. The sitemap is always checked for the site report (missing sitemap, sitemap URLs that fail or redirect) even when this is off.

## `sitemapUrl` (type: `string`):

Leave empty to find the sitemap automatically (robots.txt, then /sitemap.xml). Set it if your sitemap lives elsewhere.

## `respectRobotsTxt` (type: `boolean`):

Skip pages that robots.txt disallows for crawlers (they are not charged). Every audited page also reports whether Googlebot is allowed to crawl it (blockedByRobotsTxt).

## `checkImages` (type: `boolean`):

Request every image (up to 30 per page, cached across pages) and report broken images and images larger than 200 KB. Makes the audit slower; no extra charge.

## `validateStructuredData` (type: `boolean`):

Check JSON-LD and Microdata with the Schema Markup Validator rules and add a structuredData object (errors, eligible rich results, RDFa types) to every page. Does not change the score. No extra charge.

## `maxConcurrency` (type: `integer`):

How many pages are audited at the same time. Lower it for small or slow sites.

## `pagespeed` (type: `boolean`):

Measure Core Web Vitals on the start URL(s) and the next pages audited, up to the limit below: performance score, LCP, CLS, TBT, FCP, Speed Index and time to interactive, plus real-user INP/LCP/CLS when Google has field data. Works without any setup: we run Google PageSpeed Insights for you (or our own Lighthouse server when Google's quota is busy; each page's result says which in `source`). Charged per page measured ($0.005); failed measurements are free.

## `pagespeedStrategy` (type: `string`):

Lighthouse device profile for the measurement.

## `pagespeedSource` (type: `string`):

Lab results depend on where and on what hardware Lighthouse runs, so measure from close to your users: choose Taiwan for sites whose audience is in Asia. A single location is never swapped for another: if it is unavailable, the page is reported as not measured (not charged). Each result says what was used in `source` and `node`. Real-user field data (CrUX) only comes from Google. Your own Google API key below is used only with Automatic or Google.

## `pagespeedMaxPages` (type: `integer`):

Cost cap: at most this many pages are measured (start URLs first). Each measurement takes 10-60 seconds.

## `pagespeedApiKey` (type: `string`):

Not needed: without a key, pages are measured through our servers. Add your free PageSpeed Insights API key from Google Cloud (APIs & Services > Credentials) only if you want the calls to use your own Google quota. Stored encrypted.

## `compareWithPrevious` (type: `boolean`):

Monitoring: remember each page's score and issues, and report issues added or resolved per page, the score change, new broken pages and the average score change in SUMMARY and REPORT. Use with a schedule.

## `monitorName` (type: `string`):

Runs with the same name are compared. Leave empty to use the start URL and crawl scope.

## `httpVia` (type: `string`):

Some sites block requests from data centers. Auto retries from a second network before falling back to a real browser. If both are refused, Auto also tries our second server (Oracle, different IP) before a real browser.

## `pageTimeoutSecs` (type: `integer`):

A page that takes longer to read (all retries and the browser fallback included) is reported with errorType "timeout" and not charged. 0 = no limit.

## `proxyConfiguration` (type: `object`):

Only used for requests sent from Apify's network. Apify proxy usage is billed to your Apify account.

## Actor input object example

```json
{
  "urls": [
    "https://www.python.org/"
  ],
  "crawlScope": "domain",
  "maxPages": 20,
  "maxDepth": 10,
  "mode": "auto",
  "useSitemap": false,
  "respectRobotsTxt": false,
  "checkImages": false,
  "validateStructuredData": true,
  "maxConcurrency": 10,
  "pagespeed": false,
  "pagespeedStrategy": "mobile",
  "pagespeedSource": "auto",
  "pagespeedMaxPages": 5,
  "compareWithPrevious": false,
  "httpVia": "auto",
  "pageTimeoutSecs": 120
}
```

# Actor output Schema

## `report` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `pages` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.python.org/"
    ],
    "maxPages": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("tidytools/seo-audit-crawler").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://www.python.org/"],
    "maxPages": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("tidytools/seo-audit-crawler").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.python.org/"
  ],
  "maxPages": 20
}' |
apify call tidytools/seo-audit-crawler --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tidytools/seo-audit-crawler"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/V1FNXB9GeyVft8LTH/builds/lapbqyBHawyzEoTko/openapi.json
