# SEO Site Audit - 0-100 Page Score, Broken Links & Fix Hints (`kadi_bence/seo-site-audit`) Actor

Crawl a website and score every page 0-100 for SEO with fix hints. Input: start URL(s); optional URL globs, link and image checks. Returns per page: seoScore, issues + fixes, status, redirects, title, meta, canonical, H1, broken links. Default: 20 pages. $1.50/1k pages.

- **URL**: https://apify.com/kadi_bence/seo-site-audit.md
- **Developed by:** [Bence Kadi](https://apify.com/kadi_bence) (community)
- **Categories:** SEO tools, Developer tools, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## SEO Site Audit: 0-100 page score, broken links, fix hints

Crawl your website and get a **technical SEO audit of every page**: broken links, redirects, titles, meta descriptions, canonical, hreflang, noindex and more. One row per page with a **0-100 `seoScore`**, an `issues` list and a plain-English **fix for each issue**, plus a site-wide **SUMMARY** with the average score and the worst pages. **$1.50 per 1,000 pages**.

### 🔎 What is SEO Site Audit Crawler?

It starts at your homepage, follows internal links and audits each HTML page. Then it checks every link target it found, so you see which links are broken and where. Each page gets issues with a code, a severity (`error`, `warning`, `notice`) and a message, e.g. `TITLE_TOO_LONG: Title is 74 characters (over 60)`.

- 💯 **SEO score and fixes per page:** a transparent 0-100 score (see [How the SEO score works](#-how-the-seo-score-works)) and one fix hint per issue, e.g. `Add a <title> of 30-60 characters describing the page and its main keyword.`
- 🔗 **Broken links, really checked:** internal targets (even beyond your page limit) and external links (HEAD, GET fallback), with status, anchor text and the page they are on.
- 🧭 **Redirects in full:** every hop (301/302/307/308), loops, temporary redirects, links to redirecting URLs.
- 🏷️ **On-page checks:** title, meta description, H1, canonical, meta robots / X-Robots-Tag, hreflang, Open Graph, structured data (JSON-LD, microdata, RDFa), word count, mixed content, viewport, lang, speed, size.
- 🤝 **Polite:** respects robots.txt (RFC 9309) and Crawl-delay, identifies itself as `SEOSiteAuditBot`.
- 💸 **Fair pricing:** only audited pages are charged; error pages, redirects and link checks are free.

**Use cases:** Analyze SEO & marketing (run website audits, get SEO & analytics data, extract website metadata & tech stacks) · Developer tools (monitor & alert on website changes).

#### Use cases

- **SEO agencies:** audit a client's site before a launch or redesign and hand over a spreadsheet.
- **Broken link checker:** 404s and dead external links, with the pages linking to them.
- **Redirect checker** after a migration.
- **Meta tag audit:** all titles, descriptions, H1s and canonicals in one table.
- **Monitoring:** weekly runs with alerts on new errors.
- **Developers and QA:** fail a release build if `errorCount` goes up.

#### Output fields

One row per crawled URL. Examples are from a real run on crawler-test.com (a site built for testing SEO crawlers).

| Field | Description | Example |
|---|---|---|
| `url` / `finalUrl` | Crawled URL / URL after redirects | `https://crawler-test.com/` |
| `status` | HTTP status (`null` if the request failed) | `200` |
| `redirectChain` / `redirectCount` | Redirect hops | `[]` / `0` |
| `depth` / `foundOn` | Clicks from start / page where first found | `1` / `https://crawler-test.com/` |
| `contentType` | Content type | `text/html` |
| `indexable` | No noindex, canonical to itself or missing | `true` |
| `title` / `titleLength` | Title and length | `Crawler Test Site` / `17` |
| `metaDescription` / `metaDescriptionLength` | Meta description and length | `Default description...` / `40` |
| `h1Count` / `h1` / `h2Count` | Headings | `1` / `["Crawler Test Site"]` / `0` |
| `canonical` | Canonical link | `https://crawler-test.com/mobile/separate_desktop` |
| `metaRobots` / `xRobotsTag` | Robots meta and header | `null` |
| `lang` / `hreflang` | `<html lang>` / hreflang alternates | `null` / `[]` |
| `openGraph` / `twitterCard` | Social tags | `{}` / `null` |
| `wordCount` | Visible words | `1576` |
| `internalLinks` / `externalLinks` / `nofollowLinks` | Link counts | `411` / `3` / `3` |
| `brokenLinks` / `brokenLinkCount` | Broken outgoing links (URL, status, error, anchor, internal) | see Output / `66` |
| `images` / `imagesMissingAlt` / `imagesMissingAltSamples` | Images and missing alt text | `0` / `0` / `[]` |
| `brokenImages` | Broken images (if checked) | `[]` |
| `structuredDataTypes` / `structuredDataFormats` | Schema types, formats (JSON-LD, microdata, RDFa) | `[]` |
| `responseTimeMs` / `sizeBytes` | Speed and HTML size | `328` / `42254` |
| `seoScore` | 0-100 page score from the issues (`0` = URL did not load, `null` = not audited) | `59` |
| `issues` | `{code, severity, message}` list | see Output |
| `fixes` | One plain-English fix hint per issue, same order as `issues` | see Output |
| `issueCount` / `errorCount` / `warningCount` | Issue totals | `10` / `1` / `3` |
| `error` | Why the URL was not audited | `null` |
| `crawledAt` | Crawl time (UTC) | `2026-10-05T03:24:23Z` |

### 🚀 How to use SEO Site Audit Crawler

1. Open the [Store page](https://apify.com/kadi_bence/seo-site-audit) and click **Try for free**.
2. Enter your homepage as the **Start URL** and set **Max pages** (default 20; raise it for a full site).
3. Click **Start**. The prefilled example (30 pages) takes about 1-2 minutes.
4. Download the rows as CSV, Excel or JSON and open the **SUMMARY** record.

#### Input guide

- **Start URLs (`startUrls`):** where the crawl starts; same-domain links are followed (`www.` counts as the same site), `https://` is added when missing. Good: `https://www.example.com/`. Bad: a deep blog post alone, because the crawler only finds pages linked from where it starts. Private and local addresses are refused.
- **Max pages (`maxPages`, default 20, max 10,000):** pages, redirects and error pages all count, but only audited pages are charged. Links to pages beyond the limit are still checked for 404s.
- **Max link depth (`maxDepth`, default 20):** clicks from the start URL; `0` = start URL(s) only.
- **Include subdomains (`includeSubdomains`):** also crawl `blog.example.com` when you start on `example.com`.
- **Check external links (`checkExternalLinks`, default on):** HEAD requests, max 1 per second and 25 links per external site. Turn off for a faster run. Internal links are always checked.
- **Check images (`checkImages`, default off):** checks up to 200 image URLs per page for 404s. Missing alt text is always reported.
- **URL patterns (`includeUrlGlobs` / `excludeUrlGlobs`):** globs matched against the full URL and the path, applied to discovered links (not start URLs). Good: `/blog/*`, `/tag/*`, `*?page=*`. Bad: `blog` (no wildcard).
- **Max link checks (`maxLinkChecks`, default 2000, max 50,000):** cap on link targets not crawled as pages. Free, but slow servers make them slow.
- **Parallel requests (`maxConcurrency`, default 5, max 20):** a robots.txt Crawl-delay always wins. Keep it low for small servers.
- **Request timeout (`requestTimeoutSecs`, default 20):** slower pages are reported as `Timed out`; 429 and 5xx answers get one retry.
- **Proxy (`proxyConfiguration`):** usually not needed. Datacenter Apify Proxy only; residential proxies make the run fail.

Other limits: HTML up to 5 MB, 10 redirects per chain; file links (PDF, images, ZIP...) are not crawled as pages.

Full example input:

```json
{
  "startUrls": [{ "url": "https://crawler-test.com/" }],
  "maxPages": 30,
  "maxDepth": 20,
  "includeSubdomains": false,
  "checkExternalLinks": true,
  "checkImages": false,
  "includeUrlGlobs": [],
  "excludeUrlGlobs": ["/tag/*", "*?page=*"],
  "maxLinkChecks": 2000,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 20,
  "proxyConfiguration": { "useApifyProxy": false }
}
```

### 📦 Output

Homepage row from a real run on crawler-test.com, 2026-10-05 (`brokenLinks`, `issues` and `fixes` shortened; 1 error, 3 warnings and 6 notices give 100 - 20 - 15 - 6 = 59):

```json
{
  "url": "https://crawler-test.com/",
  "finalUrl": "https://crawler-test.com/",
  "status": 200,
  "redirectChain": [],
  "redirectCount": 0,
  "depth": 0,
  "foundOn": null,
  "contentType": "text/html",
  "indexable": true,
  "title": "Crawler Test Site",
  "titleLength": 17,
  "metaDescription": "Default description XIbwNE7SSUJciq0/Jyty",
  "metaDescriptionLength": 40,
  "h1Count": 1,
  "h1": ["Crawler Test Site"],
  "h2Count": 0,
  "canonical": null,
  "metaRobots": null,
  "xRobotsTag": null,
  "lang": null,
  "hreflang": [],
  "openGraph": {},
  "twitterCard": null,
  "wordCount": 1576,
  "internalLinks": 411,
  "externalLinks": 3,
  "nofollowLinks": 3,
  "brokenLinks": [
    {"url": "http://www.søkbar.no/", "status": null, "error": "Domain not found (DNS lookup failed)",
     "anchor": "Foreign Character Domain", "internal": false},
    {"url": "https://crawler-test.com/redirects/redirect_to_404", "status": 404, "error": null,
     "anchor": "Redirect To 404 Http Status", "internal": true}
  ],
  "brokenLinkCount": 66,
  "images": 0,
  "imagesMissingAlt": 0,
  "imagesMissingAltSamples": [],
  "brokenImages": [],
  "structuredDataTypes": [],
  "structuredDataFormats": [],
  "responseTimeMs": 328,
  "sizeBytes": 42254,
  "seoScore": 59,
  "issues": [
    {"code": "BROKEN_INTERNAL_LINKS", "severity": "error", "message": "64 broken internal link(s)"},
    {"code": "MIXED_CONTENT", "severity": "warning", "message": "2 insecure http:// resource(s) on an https page"},
    {"code": "TITLE_TOO_SHORT", "severity": "notice", "message": "Title is only 17 characters (under 30)"},
    {"code": "LINKS_TO_REDIRECTS", "severity": "notice", "message": "10 internal link(s) point to redirecting URLs"}
  ],
  "fixes": [
    "Fix or remove the broken internal links listed in brokenLinks (or redirect their targets).",
    "Load all scripts, styles and images over https:// instead of http://.",
    "Expand the <title> to 30-60 characters with a descriptive keyword phrase.",
    "Update internal links to point directly at the final URL instead of a redirecting one."
  ],
  "issueCount": 10,
  "errorCount": 1,
  "warningCount": 3,
  "error": null,
  "crawledAt": "2026-10-05T03:24:23Z"
}
```

#### Dataset views

Table views in Console (API: `?view=<name>`):

- **Scores and issues per page** (`issues`): SEO score, status, indexable, issue counts, broken links, issue list, fix hints, error.
- **Titles and meta tags** (`meta`): title, meta description, H1, canonical, robots meta, word count, images without alt.
- **Links and redirects** (`links`): redirect chain, final URL, depth, link counts, broken links, response time.

#### Key-value store records

- **SUMMARY**: site-wide view. `averageSeoScore` (over all scored rows), `seoScoreBands` (`good` 90-100, `needsWork` 50-89, `poor` 0-49), `worstPages` (10 lowest scores with their top issues and fixes), issue counts by code and severity, status codes, top 50 broken link targets with how many pages link to them, duplicate titles, meta descriptions and content, slowest pages, robots.txt info (crawl-delay, sitemaps) and URLs blocked by robots.txt. In the run above: 30 pages, 27 indexable, 66 broken link targets, 7 missing meta descriptions.
- **RUN_SUMMARY**: run stats (pages crawled, audited and charged, requests, link checks, speed; the error if the input was invalid).

#### Export

Download as JSON, CSV, Excel, XML, RSS or HTML table, or via the API (`/v2/datasets/{datasetId}/items?format=csv&view=meta`). Duplicates are checked only between indexable pages (no noindex, canonical pointing to itself or missing). Links that answer 401, 403, 429 or 999 (common on social networks that refuse bots) are not counted as broken.

#### 💯 How the SEO score works

Every row gets `seoScore` = **100 minus a fixed deduction per issue**, by severity, floored at 0:

| Severity | Deduction | Examples |
|---|---|---|
| `error` | **-20** | `TITLE_MISSING`, `CANONICAL_BROKEN`, `BROKEN_INTERNAL_LINKS` |
| `warning` | **-5** | `META_DESCRIPTION_MISSING`, `H1_MISSING`, `TITLE_TOO_LONG`, `IMAGES_MISSING_ALT`, `NOINDEX`, `REDIRECT_CHAIN`, `DUPLICATE_TITLE` |
| `notice` | **-1** | `TITLE_TOO_SHORT`, `H1_MULTIPLE`, `CANONICAL_MISSING`, `OG_TAGS_MISSING`, `LANG_MISSING` |

- Each issue code counts once per page, however many links or images it covers (e.g. 64 broken internal links = one `-20`).
- A URL that did not load (`HTTP_4XX`, `HTTP_5XX`, `FETCH_FAILED`, `REDIRECT_LOOP`, `TOO_MANY_REDIRECTS`) scores **0**.
- Rows that were not audited (blocked by robots.txt, non-HTML files, redirects to a page that has its own row) get `null` and are left out of the averages.
- Rough reading: **90-100** good, **50-89** needs work, **0-49** poor. Recompute it yourself from `issues` if you want different weights.

`fixes[i]` is the fix for `issues[i]`, e.g. `META_DESCRIPTION_MISSING` → `Add a meta description of 70-160 characters summarising the page.`

#### Issue codes

- **Errors:** `HTTP_4XX`, `HTTP_5XX`, `FETCH_FAILED`, `REDIRECT_LOOP`, `TOO_MANY_REDIRECTS`, `TITLE_MISSING`, `CANONICAL_BROKEN`, `BROKEN_INTERNAL_LINKS`
- **Warnings:** `REDIRECT_CHAIN`, `TITLE_TOO_LONG`, `TITLE_MULTIPLE`, `META_DESCRIPTION_MISSING`, `META_DESCRIPTION_MULTIPLE`, `H1_MISSING`, `CANONICAL_MULTIPLE`, `NOINDEX`, `BROKEN_EXTERNAL_LINKS`, `BROKEN_IMAGES`, `IMAGES_MISSING_ALT`, `SLOW_RESPONSE` (> 2 s), `MIXED_CONTENT`, `HTTP_NOT_HTTPS`, `VIEWPORT_MISSING`, `HREFLANG_INVALID`, `HREFLANG_RELATIVE`, `STRUCTURED_DATA_INVALID`, `DUPLICATE_TITLE`, `DUPLICATE_META_DESCRIPTION`, `DUPLICATE_CONTENT`
- **Notices:** `TEMPORARY_REDIRECT`, `BLOCKED_BY_ROBOTS`, `NOT_HTML`, `TITLE_TOO_SHORT` (< 30), `META_DESCRIPTION_TOO_LONG` (> 160), `META_DESCRIPTION_TOO_SHORT` (< 70), `H1_MULTIPLE`, `CANONICAL_MISSING`, `CANONICAL_TO_OTHER_URL`, `NOFOLLOW_PAGE`, `LINKS_TO_REDIRECTS`, `LOW_WORD_COUNT` (< 200), `LARGE_PAGE` (> 1 MB), `LANG_MISSING`, `OG_TAGS_MISSING`, `HREFLANG_NO_SELF`, `META_REFRESH`, `URL_TOO_LONG` (> 115)

### 💵 Pricing: $1.50 per 1,000 pages

You pay **once per page that was audited** (event `page`: an HTML page that answered 2xx), plus Apify's $0.00005 per-run start fee. Platform compute is included. Error pages (404, 500), redirects, robots.txt-blocked URLs and all link checks are free.

| Run | Audited pages | Cost |
|---|---|---|
| Prefilled example | 30 | $0.045 |
| Small business site | 100 | $0.15 |
| Medium site | 1,000 | $1.50 |
| Large site (maximum per run) | 10,000 | $15.00 |

Set **Max charge per run** to cap spending: the crawl then stops at the number of pages your limit covers. Other SEO audit actors on the Store charge $2-40 per 1,000 pages.

### 🔌 Use it via API

**JavaScript** (`npm install apify-client`):

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });
const run = await client.actor('kadi_bence/seo-site-audit').call({
  startUrls: [{ url: 'https://www.example.com/' }],
  maxPages: 500,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
```

**Python** (`pip install apify-client`):

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("kadi_bence/seo-site-audit").call(run_input={
    "startUrls": [{"url": "https://www.example.com/"}],
    "maxPages": 500,
})
for page in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(page["url"], page["seoScore"], page["fixes"][:3])
summary = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("SUMMARY")
```

**cURL** (waits for the run and returns the rows):

```bash
curl -X POST "https://api.apify.com/v2/acts/kadi_bence~seo-site-audit/run-sync-get-dataset-items?token=YOUR_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"startUrls": [{"url": "https://www.example.com/"}], "maxPages": 100}'
```

**Apify CLI:**

```bash
apify call kadi_bence/seo-site-audit --input '{"startUrls": [{"url": "https://www.example.com/"}], "maxPages": 100}'
apify call kadi_bence/seo-site-audit -f input.json
```

**MCP for AI agents** (Claude, Cursor, VS Code and other MCP clients):

```json
{"mcpServers": {"apify": {"url": "https://mcp.apify.com/?actors=kadi_bence/seo-site-audit"}}}
```

### 🔁 Integrations & scheduling

- **Schedules:** run audits on a cron, e.g. `0 6 * * 1` = every Monday at 06:00.
- **Webhooks:** on "run succeeded", Apify calls your URL with the run ID.
- **Zapier, Make, n8n, Google Sheets, Slack, email:** pass results on with Apify integrations.

Recipes:

1. **Weekly SEO audit emailed to a client.** Save a task with the client's homepage and `maxPages: 1000`, schedule it every Monday, and let a Make or Zapier scenario email the **Scores and issues per page** view as CSV with the SUMMARY totals. About $1.50 a week.
2. **Migration check.** Run with `checkExternalLinks: false`, then filter the issues for `REDIRECT_CHAIN`, `REDIRECT_LOOP` and `LINKS_TO_REDIRECTS`.
3. **Broken link alert.** A webhook calls your script; if SUMMARY `brokenLinkTargets` is above zero, post `topBrokenLinks` to Slack.

### 🧰 Related tools by the same developer

Website due diligence:

- [Website Tech Stack Detector](https://apify.com/kadi_bence/tech-stack-detector): CMS, ecommerce, analytics and frameworks from 7,600+ fingerprints. $1.50 / 1,000 websites.
- [Website Screenshot & PDF API](https://apify.com/kadi_bence/website-screenshot): full-page and mobile PNG, JPEG, WebP and PDF captures. $2.50 / 1,000 screenshots.
- [Bulk WHOIS & RDAP Domain Lookup](https://apify.com/kadi_bence/domain-whois-lookup): registrar, age, expiry, DNS/MX/SPF/DMARC. $1.50 / 1,000 domains.

Content pipelines:

- [Document to Markdown](https://apify.com/kadi_bence/document-to-markdown): PDF, Word, PowerPoint, Excel and OCR to Markdown and RAG chunks. $2 / 1,000 documents.
- [Bulk Image Downloader](https://apify.com/kadi_bence/bulk-image-downloader): all images and files from pages into a ZIP. $0.80 / 1,000 files.

### ❓ FAQ

**How does it work?**
Plain HTTP, no browser: it crawls internal links breadth-first, then checks all link targets found, compares pages for duplicates and writes the SUMMARY.

**Which sites can I crawl?**
Your own sites, or sites you have permission to audit. It stays on the domain(s) of your start URLs; other sites only get single HEAD requests to check your outgoing links.

**Does it respect robots.txt?**
Always: RFC 9309 (group for `SEOSiteAuditBot`, else `*`; longest match wins; `*` and `$` wildcards) plus Crawl-delay. Disallowed URLs are never requested and are listed in the SUMMARY. If robots.txt answers 5xx or times out, the site is treated as "do not crawl". Pages with `nofollow` are audited, but their links are not followed. To audit blocked pages, add an `Allow` rule for `User-agent: SEOSiteAuditBot`.

**How fast is it?**
The crawl ran at about 700 pages per minute in our tests (5 parallel requests, server answering in about 0.4 s). Link checks come after; on a slow server 430 link targets took about 80 s in our 200-page test. Large sites need a longer run timeout; the Actor stops before it and still saves results.

**What are the limits?**
10,000 pages and 50,000 link checks per run, HTML up to 5 MB. No JavaScript rendering: single-page apps show what the server sends, as many search engines see it on the first pass. Sitemap-only and URL-list modes are on the roadmap.

**My site blocks the crawler. Do I need a proxy?**
Usually not. Allow the user agent `SEOSiteAuditBot` in your firewall or CDN, or try the datacenter Apify Proxy.

**Why is an external link OK here but broken in my browser?**
Some sites (LinkedIn, Instagram, some shops) answer 401/403/429/999 to every bot, so those are not reported. 404, 410, other 4xx, 5xx, dead domains, refused connections, timeouts and redirect loops are.

**Does it collect personal data? Is it legal?**
It reads SEO tags, not personal data. E-mails and phone numbers in titles, descriptions, headings or anchors become `[email removed]` / `[phone removed]`; `mailto:`/`tel:` links are ignored. Getting permission to crawl is your responsibility; this is not legal advice.

**What does a big job cost?**
10,000 pages cost $15. Use **Max charge per run** as a hard cap.

**What if it fails?**
Failed pages are free rows with `status`, an `error` text and an issue such as `FETCH_FAILED` or `HTTP_4XX`. Invalid input stops the run before any page is charged. Something wrong? Open an issue on the **Issues** tab with the URL. Fixes usually land within 24-48 hours.

**Can I use it from Python, integrations or an AI agent?**
Yes, see **Use it via API** and **Integrations** above. Version history is in the Changelog.

### 📝 Changelog

- **1.0 (2026-10-05):** First release. Full-site crawl with 47 issue codes, broken internal and external links (HEAD with GET fallback), redirect chains, duplicates, hreflang and structured data checks, robots.txt per RFC 9309 with Crawl-delay, SUMMARY record, pay only for audited pages.
- **2026-10-05:** Store page rewritten.
- **1.1 (2026-10-05):** 0-100 `seoScore` per page (transparent severity weights), a plain-English fix hint per issue (`fixes`), and `averageSeoScore`, `seoScoreBands` and `worstPages` in the SUMMARY.

### 🤝 Need a custom version?

I build custom scrapers, scheduled data feeds and integrations (CSV, Excel, Google Sheets, API, MCP for AI agents). Tell me the sites and fields you need:

- Email: bence.kadi@gmail.com
- Apify profile: https://apify.com/kadi_bence

### Related Actors

More low-cost Actors by the same developer, built on official APIs and public data:

- [PageSpeed Insights Bulk Checker](https://apify.com/kadi_bence/pagespeed-checker) — Core Web Vitals and Lighthouse scores in bulk
- [Website Tech Stack Detector](https://apify.com/kadi_bence/tech-stack-detector) — CMS, ecommerce, analytics and frameworks of any website
- [AI Visibility Tracker](https://apify.com/kadi_bence/ai-visibility-tracker) — how ChatGPT, Perplexity, Gemini and Claude mention your brand
- [Bulk WHOIS & RDAP Domain Lookup](https://apify.com/kadi_bence/domain-whois-lookup) — registrar, dates and DNS for many domains
- [Website Screenshot & PDF API](https://apify.com/kadi_bence/website-screenshot) — full-page screenshots and PDFs of any URL
- [Bulk Image Downloader](https://apify.com/kadi_bence/bulk-image-downloader) — download all images from URLs as a ZIP
- [Document to Markdown](https://apify.com/kadi_bence/document-to-markdown) — PDF, Word, PowerPoint and Excel to clean Markdown, with OCR

All my Actors: [apify.com/kadi_bence](https://apify.com/kadi_bence)

# Actor input Schema

## `startUrls` (type: `array`):

The page(s) where the crawl starts, usually your homepage, e.g. https://www.example.com/ . The crawler follows links on the same domain (www. counts as the same site). https:// is added when missing. Only audit sites you own or have permission to crawl.

## `includeSubdomains` (type: `boolean`):

Also crawl subdomains, e.g. blog.example.com when you start on example.com. www. is always treated as the same site.

## `checkExternalLinks` (type: `boolean`):

Check links to other websites for 404s and dead domains (light HEAD requests, max 1 per second and 25 links per external site, robots.txt respected). Turn off for a faster run. Internal links are always checked.

## `checkImages` (type: `boolean`):

Also check image URLs (up to 200 per page) for 404s. Missing alt texts are always reported, even when this is off.

## `includeUrlGlobs` (type: `array`):

Optional glob patterns matched against the full URL or the path, e.g. /blog/\* to audit only the blog. Leave empty to crawl the whole site. Start URLs are always crawled.

## `excludeUrlGlobs` (type: `array`):

Glob patterns of URLs not to crawl, e.g. /tag/\* or *?page=* to skip tag archives and pagination.

## `maxPages` (type: `integer`):

Stop after this many URLs, e.g. 500 (HTML pages, redirects and error pages all count; max 10,000). You pay only for pages that were audited successfully. Links to pages beyond the limit are still checked. Default 20.

## `maxDepth` (type: `integer`):

How many clicks away from the start URL the crawler goes, e.g. 3 for the main navigation only. 0 = only the start URL(s).

## `maxLinkChecks` (type: `integer`):

Upper limit of link targets checked that were not crawled as pages (internal links beyond Max pages, external links, images). Link checks are free, but slow servers make them take longer.

## `maxConcurrency` (type: `integer`):

How many pages of your site are fetched at the same time. A Crawl-delay in robots.txt always wins. Keep it low (e.g. 2-5) for small servers.

## `requestTimeoutSecs` (type: `integer`):

A page that does not answer in this time is reported as 'Timed out' (one retry on 429/5xx answers).

## `proxyConfiguration` (type: `object`):

Usually not needed for your own site. If your firewall blocks the crawler, allow the user agent SEOSiteAuditBot or try the datacenter Apify Proxy. Residential proxies are not supported.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://crawler-test.com/"
    }
  ],
  "includeSubdomains": false,
  "checkExternalLinks": true,
  "checkImages": false,
  "maxPages": 30,
  "maxDepth": 20,
  "maxLinkChecks": 200,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `issues` (type: `string`):

SEO score (0-100), status, issue counts, the issues and a fix hint for each.

## `meta` (type: `string`):

Title, meta description, H1, canonical and robots meta per page.

## `links` (type: `string`):

Redirect chains, link counts and broken outgoing links per page.

## `full` (type: `string`):

Every field incl. hreflang, Open Graph, structured data and images.

## `summary` (type: `string`):

Average SEO score, worst pages, issue counts, top broken links, duplicate titles and descriptions, slowest pages.

## `runSummary` (type: `string`):

Pages crawled and charged, requests, robots.txt info and speed.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://crawler-test.com/"
        }
    ],
    "maxPages": 30,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("kadi_bence/seo-site-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://crawler-test.com/" }],
    "maxPages": 30,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("kadi_bence/seo-site-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://crawler-test.com/"
    }
  ],
  "maxPages": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call kadi_bence/seo-site-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,kadi_bence/seo-site-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/go78524BUxyfJKC5I/builds/Ld3QilRTQAbWWxP1N/openapi.json
