# SEO Audit Tool - Site Crawler with 0-100 Score (`parseforge/seo-audit-tool`) Actor

Crawl a website and get a 0-100 SEO score for every page and the whole site: meta, headings, robots, sitemap, broken links, images, schema. Export to CSV, Excel, JSON or XML.

- **URL**: https://apify.com/parseforge/seo-audit-tool.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.40 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

![ParseForge Banner](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner-v4.webp)

## 🔎 SEO Website Audit Scraper

> 🚀 **Export a 0-100 SEO score for every page and for the whole site in minutes.** 30 checks in 7 groups, every issue with a fix hint, broken links verified by real requests, all in one flat table.

Point the Actor at a website and it crawls the site the way a search engine would: it reads robots.txt, follows internal links (and the XML sitemap when the links run out), and audits each page it fetches. You get one row per page with its score, grade, issues and the raw facts behind them, plus one summary row per site with the site score, the most common problems, sitemap and robots.txt findings, duplicate groups and every broken link.

Each row has up to 98 columns. In the reference runs (37 pages from 4 audits of 2 sites) 72 columns held data on at least one page row and 33 on the site summary rows; the others are the other row type's columns, or only appear when there is a problem to report (for example `brokenLinks` or `redirectChain`). The score method is fixed and documented below, so the same page always gets the same score and two audits can be compared over time.

| 🎯 Target Audience | 💡 Primary Use Cases |
|---|---|
| SEO freelancers, agencies, in-house marketers, web developers, site owners | Client audits, pre-launch checks, monitoring a score over time, finding broken links, spotting duplicate titles and thin pages |

### 📋 What the SEO Audit Tool does

> 💡 **Why it matters:** an SEO report is only useful if you can trust the number and act on it. Every score here is the sum of named checks with published weights, and every lost point comes with the issue, the URL and a fix hint.

- **Crawls politely.** It respects robots.txt by default, honours a Crawl-delay (up to 10 seconds), waits between requests and keeps to a page cap you set.
- **Grades each page.** Titles, meta descriptions, canonicals, headings, word count, language, HTTP status, noindex, HTTPS, redirects, mixed content, links, images, structured data, Open Graph, mobile viewport and speed hints.
- **Verifies links.** Links to pages it did not crawl are requested (HEAD first) and reported as broken, redirecting or unverified, attributed to the page they were found on.
- **Finds duplicates.** Pages sharing a title, a meta description, or identical and near-identical text are listed on both pages and grouped in the site summary.
- **Reads robots.txt and sitemaps.** Whether they exist, what they declare, how many URLs the sitemap holds and which crawled pages are missing from it.
- **Stays on public sites.** Requests to private, loopback and link-local addresses, and redirects into them, are refused before any connection is made.

### 🎬 Full Demo (🚧 Coming soon)

### 📊 Output

Two row types share one table, told apart by `type`: `page` and `site-summary`. Columns that do not apply to a row type hold `N/A`, `[]` or `{}`. Yes/No fields are the text `Yes` or `No`.

| Group | Fields |
|---|---|
| 🏷 Identity | `imageUrl` (og:image or first image), `type`, `url`, `finalUrl`, `siteUrl`, `host`, `crawlDepth`, `discoveredFrom` |
| 🎯 Score | `score`, `grade`, `siteScore`, `siteGrade`, `categoryScores` (metadata, content, indexability, links, images, markup, performance), `issueCount`, `criticalIssues`, `warningIssues`, `noticeIssues`, `issues` (id, severity, category, message, fixHint) |
| 📌 Metadata | `title`, `titleLength`, `metaDescription`, `metaDescriptionLength`, `canonicalUrl`, `canonicalStatus`, `language` |
| 🔎 Indexability | `statusCode`, `isIndexable`, `indexabilityReason`, `robotsMeta`, `xRobotsTag`, `isHttps`, `mixedContentCount`, `redirectChain`, `redirectCount` |
| 📝 Content | `h1Count`, `h1Text`, `headingCounts`, `headingOutline`, `headingLevelSkips`, `wordCount`, `textToHtmlRatio`, `likelyClientRendered` |
| 🧬 Duplicates | `duplicateTitleWith`, `duplicateDescriptionWith`, `duplicateContentWith` |
| 🔗 Links | `internalLinkCount`, `externalLinkCount`, `nofollowLinkCount`, `emptyAnchorCount`, `genericAnchorCount`, `linksChecked`, `brokenLinks`, `redirectingLinks`, `unverifiedLinkCount` |
| 🖼 Images | `imageCount`, `imagesMissingAlt`, `imagesMissingAltUrls`, `imagesMissingDimensions`, `lazyLoadedImages` |
| 🧩 Markup | `structuredDataFormats`, `structuredDataTypes`, `invalidJsonLdBlocks`, `openGraph`, `twitterCard`, `hreflangCount` |
| ⚡ Speed hints | `hasViewport`, `responseTimeMs`, `htmlBytes`, `isCompressed`, `cacheControl`, `renderBlockingScripts`, `stylesheetCount`, `scriptCount`, `domNodes`, `contentType` |
| 🌐 Site summary | `pageScoreAverage`, `siteChecksScore`, `siteChecks`, `pagesAudited`, `pagesOk`, `pagesWithErrors`, `pagesBlockedByRobots`, `scoreDistribution`, `categoryAverages`, `issueTotals`, `topIssues`, `robotsTxt`, `sitemap`, `duplicateTitleGroups`, `duplicateDescriptionGroups`, `duplicateContentGroups`, `brokenLinkCount`, `redirectingLinkCount`, `crawlStopReason`, `notes`, `crawlSeconds` |
| 🕒 Run | `scrapedAt`, `error` |

Three real rows from a cloud run: the audit of a one-page site, its site summary, and a page from a documentation site.

```json
[
  {
    "imageUrl": "N/A",
    "type": "page",
    "url": "https://example.com/",
    "finalUrl": "https://example.com/",
    "siteUrl": "https://example.com/",
    "host": "example.com",
    "statusCode": 200,
    "score": 60,
    "grade": "D",
    "siteScore": 61,
    "siteGrade": "D",
    "categoryScores": {"metadata":36,"content":31,"indexability":100,"links":50,"images":100,"markup":0,"performance":100},
    "issueCount": 10,
    "criticalIssues": 1,
    "warningIssues": 4,
    "noticeIssues": 5,
    "issues": [{"id":"H1_MISSING","severity":"critical","category":"content","message":"The page has no H1 heading.","fixHint":"Add one H1 that states what the page is about."},{"id":"TITLE_TOO_SHORT","severity":"warning","category":"metadata","message":"Title is 14 characters (best practice: 30-60).","fixHint":"Expand the title with the page topic and a qualifier."},{"id":"DESCRIPTION_MISSING","severity":"warning","category":"metadata","message":"The page has no meta description.","fixHint":"Add a meta description of 70-160 characters that summarises the page and invites the click."},{"id":"THIN_CONTENT","severity":"warning","category":"content","message":"The page has 26 words of text (300+ is a healthy floor for a content page).","fixHint":"Add useful, original text, or noindex pages that are not meant to rank."},{"id":"FEW_INTERNAL_LINKS","severity":"warning","category":"links","message":"The page links to only 0 internal page(s).","fixHint":"Link to related pages so visitors and crawlers can move on."},{"id":"CANONICAL_MISSING","severity":"notice","category":"metadata","message":"No canonical URL is declared.","fixHint":"Add <link rel=\"canonical\"> pointing at the preferred version of this page."},{"id":"WEAK_ANCHOR_TEXT","severity":"notice","category":"links","message":"1 link(s) have empty or generic anchor text (\"click here\", \"read more\").","fixHint":"Use anchor text that describes the destination."},{"id":"STRUCTURED_DATA_MISSING","severity":"notice","category":"markup","message":"No structured data (JSON-LD, microdata or RDFa) was found.","fixHint":"Add schema.org JSON-LD that fits the page (Article, Product, Organization, FAQ...)."},{"id":"OPEN_GRAPH_INCOMPLETE","severity":"notice","category":"markup","message":"Open Graph is missing og:title, og:description, og:image.","fixHint":"Add og:title, og:description and og:image so shared links look right."},{"id":"TWITTER_CARD_MISSING","severity":"notice","category":"markup","message":"No twitter:card tag.","fixHint":"Add <meta name=\"twitter:card\" content=\"summary_large_image\">."}],
    "title": "Example Domain",
    "titleLength": 14,
    "metaDescription": "N/A",
    "metaDescriptionLength": 0,
    "canonicalUrl": "N/A",
    "canonicalStatus": "missing",
    "robotsMeta": "N/A",
    "xRobotsTag": "N/A",
    "isIndexable": "Yes",
    "indexabilityReason": "Indexable",
    "language": "en",
    "h1Count": 0,
    "h1Text": "N/A",
    "headingCounts": {"h1":0,"h2":0,"h3":0,"h4":0,"h5":0,"h6":0},
    "headingOutline": [],
    "headingLevelSkips": 0,
    "wordCount": 26,
    "textToHtmlRatio": 23.3,
    "duplicateTitleWith": [],
    "duplicateDescriptionWith": [],
    "duplicateContentWith": [],
    "internalLinkCount": 0,
    "externalLinkCount": 1,
    "nofollowLinkCount": 0,
    "emptyAnchorCount": 0,
    "genericAnchorCount": 1,
    "linksChecked": 1,
    "brokenLinks": [],
    "redirectingLinks": [],
    "unverifiedLinkCount": 0,
    "imageCount": 0,
    "imagesMissingAlt": 0,
    "imagesMissingAltUrls": [],
    "imagesMissingDimensions": 0,
    "lazyLoadedImages": 0,
    "structuredDataFormats": [],
    "structuredDataTypes": [],
    "invalidJsonLdBlocks": 0,
    "openGraph": {"title":"N/A","description":"N/A","image":"N/A","type":"N/A","url":"N/A"},
    "twitterCard": "N/A",
    "hasViewport": "Yes",
    "hreflangCount": 0,
    "isHttps": "Yes",
    "mixedContentCount": 0,
    "redirectChain": [],
    "redirectCount": 0,
    "contentType": "text/html; charset=utf-8",
    "responseTimeMs": 42,
    "htmlBytes": 713,
    "isCompressed": "Yes",
    "cacheControl": "N/A",
    "renderBlockingScripts": 0,
    "stylesheetCount": 0,
    "scriptCount": 1,
    "domNodes": 11,
    "likelyClientRendered": "No",
    "crawlDepth": 0,
    "discoveredFrom": "N/A",
    "pageScoreAverage": "N/A",
    "siteChecksScore": "N/A",
    "siteChecks": [],
    "pagesAudited": "N/A",
    "pagesOk": "N/A",
    "pagesWithErrors": "N/A",
    "pagesBlockedByRobots": "N/A",
    "scoreDistribution": {},
    "categoryAverages": {},
    "issueTotals": {},
    "topIssues": [],
    "robotsTxt": {},
    "sitemap": {},
    "duplicateTitleGroups": [],
    "duplicateDescriptionGroups": [],
    "duplicateContentGroups": [],
    "brokenLinkCount": "N/A",
    "redirectingLinkCount": "N/A",
    "crawlStopReason": "N/A",
    "notes": [],
    "crawlSeconds": "N/A",
    "scrapedAt": "2026-09-29T04:02:03.560Z",
    "error": null
  },
  {
    "imageUrl": "N/A",
    "type": "site-summary",
    "url": "https://example.com/",
    "finalUrl": "https://example.com/",
    "siteUrl": "https://example.com/",
    "host": "example.com",
    "statusCode": "N/A",
    "score": "N/A",
    "grade": "N/A",
    "siteScore": 61,
    "siteGrade": "D",
    "categoryScores": {},
    "issueCount": "N/A",
    "criticalIssues": "N/A",
    "warningIssues": "N/A",
    "noticeIssues": "N/A",
    "issues": [],
    "title": "N/A",
    "titleLength": "N/A",
    "metaDescription": "N/A",
    "metaDescriptionLength": "N/A",
    "canonicalUrl": "N/A",
    "canonicalStatus": "N/A",
    "robotsMeta": "N/A",
    "xRobotsTag": "N/A",
    "isIndexable": "N/A",
    "indexabilityReason": "N/A",
    "language": "N/A",
    "h1Count": "N/A",
    "h1Text": "N/A",
    "headingCounts": {},
    "headingOutline": [],
    "headingLevelSkips": "N/A",
    "wordCount": "N/A",
    "textToHtmlRatio": "N/A",
    "duplicateTitleWith": [],
    "duplicateDescriptionWith": [],
    "duplicateContentWith": [],
    "internalLinkCount": "N/A",
    "externalLinkCount": "N/A",
    "nofollowLinkCount": "N/A",
    "emptyAnchorCount": "N/A",
    "genericAnchorCount": "N/A",
    "linksChecked": 1,
    "brokenLinks": [],
    "redirectingLinks": [],
    "unverifiedLinkCount": 0,
    "imageCount": "N/A",
    "imagesMissingAlt": "N/A",
    "imagesMissingAltUrls": [],
    "imagesMissingDimensions": "N/A",
    "lazyLoadedImages": "N/A",
    "structuredDataFormats": [],
    "structuredDataTypes": [],
    "invalidJsonLdBlocks": "N/A",
    "openGraph": {},
    "twitterCard": "N/A",
    "hasViewport": "N/A",
    "hreflangCount": "N/A",
    "isHttps": "N/A",
    "mixedContentCount": "N/A",
    "redirectChain": [],
    "redirectCount": "N/A",
    "contentType": "N/A",
    "responseTimeMs": "N/A",
    "htmlBytes": "N/A",
    "isCompressed": "N/A",
    "cacheControl": "N/A",
    "renderBlockingScripts": "N/A",
    "stylesheetCount": "N/A",
    "scriptCount": "N/A",
    "domNodes": "N/A",
    "likelyClientRendered": "N/A",
    "crawlDepth": "N/A",
    "discoveredFrom": "N/A",
    "pageScoreAverage": 60,
    "siteChecksScore": 65,
    "siteChecks": [{"id":"ROBOTS_TXT","earned":0,"max":10,"note":"no robots.txt (HTTP 404 or unreadable)"},{"id":"ROBOTS_ALLOWS","earned":10,"max":10,"note":"robots.txt does not block the whole site"},{"id":"SITEMAP","earned":0,"max":20,"note":"no XML sitemap found"},{"id":"SITEMAP_IN_ROBOTS","earned":0,"max":5,"note":"no Sitemap: line in robots.txt"},{"id":"HTTPS_SHARE","earned":10,"max":10,"note":"100% of audited pages are served over HTTPS"},{"id":"BROKEN_LINK_SHARE","earned":15,"max":15,"note":"0 of 1 checked link(s) are broken"},{"id":"DUP_TITLES","earned":10,"max":10,"note":"0 page(s) share a title with another page"},{"id":"DUP_DESCRIPTIONS","earned":5,"max":5,"note":"0 page(s) share a meta description with another page"},{"id":"DUP_CONTENT","earned":10,"max":10,"note":"0 page(s) have duplicate or near-duplicate text"},{"id":"PAGES_OK","earned":5,"max":5,"note":"1 of 1 audited pages answered with HTTP 2xx"}],
    "pagesAudited": 1,
    "pagesOk": 1,
    "pagesWithErrors": 0,
    "pagesBlockedByRobots": 0,
    "scoreDistribution": {"90-100":0,"80-89":0,"70-79":0,"60-69":1,"0-59":0},
    "categoryAverages": {"metadata":36,"content":31,"indexability":100,"links":50,"images":100,"markup":0,"performance":100},
    "issueTotals": {"critical":1,"warning":4,"notice":5},
    "topIssues": [{"id":"H1_MISSING","severity":"critical","category":"content","fixHint":"Add one H1 that states what the page is about.","pagesAffected":1,"examples":["https://example.com/"],"exampleMessage":"The page has no H1 heading."},{"id":"TITLE_TOO_SHORT","severity":"warning","category":"metadata","fixHint":"Expand the title with the page topic and a qualifier.","pagesAffected":1,"examples":["https://example.com/"],"exampleMessage":"Title is 14 characters (best practice: 30-60)."},{"id":"DESCRIPTION_MISSING","severity":"warning","category":"metadata","fixHint":"Add a meta description of 70-160 characters that summarises the page and invites the click.","pagesAffected":1,"examples":["https://example.com/"],"exampleMessage":"The page has no meta description."},{"id":"THIN_CONTENT","severity":"warning","category":"content","fixHint":"Add useful, original text, or noindex pages that are not meant to rank.","pagesAffected":1,"examples":["https://example.com/"],"exampleMessage":"The page has 26 words of text (300+ is a healthy floor for a content page)."},{"id":"FEW_INTERNAL_LINKS","severity":"warning","category":"links","fixHint":"Link to related pages so visitors and crawlers can move on.","pagesAffected":1,"examples":["https://example.com/"],"exampleMessage":"The page links to only 0 internal page(s)."},{"id":"CANONICAL_MISSING","severity":"notice","category":"metadata","fixHint":"Add <link rel=\"canonical\"> pointing at the preferred version of this page.","pagesAffected":1,"examples":["https://example.com/"],"exampleMessage":"No canonical URL is declared."},{"id":"WEAK_ANCHOR_TEXT","severity":"notice","category":"links","fixHint":"Use anchor text that describes the destination.","pagesAffected":1,"examples":["https://example.com/"],"exampleMessage":"1 link(s) have empty or generic anchor text (\"click here\", \"read more\")."},{"id":"STRUCTURED_DATA_MISSING","severity":"notice","category":"markup","fixHint":"Add schema.org JSON-LD that fits the page (Article, Product, Organization, FAQ...).","pagesAffected":1,"examples":["https://example.com/"],"exampleMessage":"No structured data (JSON-LD, microdata or RDFa) was found."},{"id":"OPEN_GRAPH_INCOMPLETE","severity":"notice","category":"markup","fixHint":"Add og:title, og:description and og:image so shared links look right.","pagesAffected":1,"examples":["https://example.com/"],"exampleMessage":"Open Graph is missing og:title, og:description, og:image."},{"id":"TWITTER_CARD_MISSING","severity":"notice","category":"markup","fixHint":"Add <meta name=\"twitter:card\" content=\"summary_large_image\">.","pagesAffected":1,"examples":["https://example.com/"],"exampleMessage":"No twitter:card tag."}],
    "robotsTxt": {"found":false,"url":"https://example.com/robots.txt","blocksAll":false,"crawlDelaySeconds":0,"sitemaps":[]},
    "sitemap": {"found":false,"sources":[],"urlCount":0,"crawledPagesMissingFromSitemap":0},
    "duplicateTitleGroups": [],
    "duplicateDescriptionGroups": [],
    "duplicateContentGroups": [],
    "brokenLinkCount": 0,
    "redirectingLinkCount": 0,
    "crawlStopReason": "complete",
    "notes": [],
    "crawlSeconds": 0.9,
    "scrapedAt": "2026-09-29T04:02:03.560Z",
    "error": null
  },
  {
    "imageUrl": "https://crawlee.dev/img/crawlee-js-og.png",
    "type": "page",
    "url": "https://crawlee.dev/js/docs/guides",
    "finalUrl": "https://crawlee.dev/js/docs/guides",
    "siteUrl": "https://crawlee.dev/",
    "host": "crawlee.dev",
    "statusCode": 200,
    "score": 91,
    "grade": "A",
    "siteScore": 89,
    "siteGrade": "B",
    "categoryScores": {"metadata":93,"content":84,"indexability":100,"links":100,"images":60,"markup":90,"performance":100},
    "issueCount": 5,
    "criticalIssues": 0,
    "warningIssues": 1,
    "noticeIssues": 4,
    "issues": [{"id":"TITLE_TOO_LONG","severity":"warning","category":"metadata","message":"Title is 64 characters (best practice: 30-60).","fixHint":"Shorten the title so search results do not truncate it."},{"id":"LOW_WORD_COUNT","severity":"notice","category":"content","message":"The page has 234 words of text (300+ is a healthy floor for a content page).","fixHint":"Add useful, original text, or noindex pages that are not meant to rank."},{"id":"IMG_DIMENSIONS_MISSING","severity":"notice","category":"images","message":"12 image(s) have no width and height attributes, which can shift the layout while loading.","fixHint":"Set width and height on every <img>."},{"id":"IMG_NOT_LAZY","severity":"notice","category":"images","message":"9 below-the-fold image(s) are not lazy-loaded.","fixHint":"Add loading=\"lazy\" to images that are not visible on first paint."},{"id":"OPEN_GRAPH_INCOMPLETE","severity":"notice","category":"markup","message":"Open Graph is missing og:description.","fixHint":"Add og:title, og:description and og:image so shared links look right."}],
    "title": "Guides | Crawlee for JavaScript · Build reliable crawlers. Fast.",
    "titleLength": 64,
    "metaDescription": "Crawlee helps you build and maintain your crawlers. It's open source, but built by developers who scrape millions of pages every day for a living.",
    "metaDescriptionLength": 146,
    "canonicalUrl": "https://crawlee.dev/js/docs/guides",
    "canonicalStatus": "self",
    "robotsMeta": "N/A",
    "xRobotsTag": "N/A",
    "isIndexable": "Yes",
    "indexabilityReason": "Indexable",
    "language": "en",
    "h1Count": 1,
    "h1Text": "Guides",
    "headingCounts": {"h1":1,"h2":19,"h3":0,"h4":0,"h5":0,"h6":0},
    "headingOutline": ["h1: Guides","h2: 📄️Request Storage","h2: 📄️Result Storage","h2: 📄️HTTP clients","h2: 📄️Configuration","h2: 📄️CheerioCrawler","h2: 📄️JavaScript rendering","h2: 📄️Proxy Management","h2: 📄️Session Management","h2: 📄️Scaling our crawlers","h2: 📄️Avoid getting blocked","h2: 📄️JSDOMCrawler","h2: 📄️Impit HTTP Client","h2: 📄️Got Scraping","h2: 📄️TypeScript Projects","h2: 📄️Running in Docker","h2: 📄️StagehandCrawler","h2: 📄️Running in web server","h2: 📄️Parallel Scraping","h2: 📄️Using a custom HTTP client (Experimental)"],
    "headingLevelSkips": 0,
    "wordCount": 234,
    "textToHtmlRatio": 5.9,
    "duplicateTitleWith": [],
    "duplicateDescriptionWith": ["https://crawlee.dev/","https://crawlee.dev/js"],
    "duplicateContentWith": [],
    "internalLinkCount": 50,
    "externalLinkCount": 9,
    "nofollowLinkCount": 0,
    "emptyAnchorCount": 0,
    "genericAnchorCount": 0,
    "linksChecked": 54,
    "brokenLinks": [],
    "redirectingLinks": [],
    "unverifiedLinkCount": 1,
    "imageCount": 12,
    "imagesMissingAlt": 0,
    "imagesMissingAltUrls": [],
    "imagesMissingDimensions": 12,
    "lazyLoadedImages": 0,
    "structuredDataFormats": ["json-ld"],
    "structuredDataTypes": ["BreadcrumbList"],
    "invalidJsonLdBlocks": 0,
    "openGraph": {"title":"Guides | Crawlee for JavaScript · Build reliable crawlers. Fast.","description":"N/A","image":"https://crawlee.dev/img/crawlee-js-og.png","type":"N/A","url":"https://crawlee.dev/js/docs/guides"},
    "twitterCard": "summary_large_image",
    "hasViewport": "Yes",
    "hreflangCount": 2,
    "isHttps": "Yes",
    "mixedContentCount": 0,
    "redirectChain": [],
    "redirectCount": 0,
    "contentType": "text/html; charset=utf-8",
    "responseTimeMs": 86,
    "htmlBytes": 41677,
    "isCompressed": "Yes",
    "cacheControl": "max-age=600",
    "renderBlockingScripts": 1,
    "stylesheetCount": 1,
    "scriptCount": 3,
    "domNodes": 468,
    "likelyClientRendered": "No",
    "crawlDepth": 1,
    "discoveredFrom": "https://crawlee.dev/",
    "pageScoreAverage": "N/A",
    "siteChecksScore": "N/A",
    "siteChecks": [],
    "pagesAudited": "N/A",
    "pagesOk": "N/A",
    "pagesWithErrors": "N/A",
    "pagesBlockedByRobots": "N/A",
    "scoreDistribution": {},
    "categoryAverages": {},
    "issueTotals": {},
    "topIssues": [],
    "robotsTxt": {},
    "sitemap": {},
    "duplicateTitleGroups": [],
    "duplicateDescriptionGroups": [],
    "duplicateContentGroups": [],
    "brokenLinkCount": "N/A",
    "redirectingLinkCount": "N/A",
    "crawlStopReason": "N/A",
    "notes": [],
    "crawlSeconds": "N/A",
    "scrapedAt": "2026-09-29T04:02:35.695Z",
    "error": null
  }
]
```

Fill rates from the reference runs (37 page rows from 4 audits of 2 sites) so you know what to expect: `title`, `statusCode`, `score`, `wordCount`, `canonicalStatus`, `responseTimeMs` and the link and image counts were filled on every page. `metaDescription`, `canonicalUrl`, `twitterCard`, `cacheControl` and `imageUrl` were filled on 95%. `structuredDataTypes` on 76%. `duplicateTitleWith` and `duplicateContentWith` on 14%. These depend heavily on the site. Columns that only appear when there is something to report (`brokenLinks`, `imagesMissingAltUrls`, `redirectChain`, `robotsMeta`, `xRobotsTag`) stay empty on clean pages.

### 🧮 How the score works

**Page score (0-100).** 30 checks in 7 groups; the weights add up to exactly 100. Each check earns its full weight, a fraction of it, or nothing, and the page score is the sum, rounded. The group scores in `categoryScores` are points earned divided by points possible, as a percentage.

| Group (points) | Checks and weights |
|---|---|
| Metadata (22) | title present 6, title 30-60 characters 3 (half credit for 15-70), description present 5, description 70-160 characters 3 (half for 50-200), canonical declared 3, canonical not pointing to another URL 2 |
| Content (16) | exactly one H1 6 (half for several), no skipped heading levels 3, at least 300 words 5 (half for 150+, one fifth for 50+), `lang` attribute 2 |
| Indexability (22) | HTTP 2xx 8, no noindex in meta robots or X-Robots-Tag 6, HTTPS 4, at most one redirect 2 (half for two), no mixed content 2 |
| Links (10) | broken links 5 (2 points off per broken link found on the page), at least 3 internal links 3, descriptive anchor text 2 |
| Images (10) | alt attribute on every image 6 (proportional), width and height set 2 (proportional), lazy loading below the first 3 images 2 |
| Markup (10) | structured data present 4, JSON-LD valid 2, Open Graph title, description and image 3, twitter:card 1 |
| Performance hints (10) | mobile viewport 3, server answer under 800 ms 2 (half under 1.8 s), compressed HTML 2, at most 2 render-blocking scripts 2 (half up to 5), HTML under 300 KB 1 |

A page that answers with an HTTP error is still audited and scores 0 with one critical issue. A page the site refuses to serve (HTTP 401, 403, 429 or 503) is not scored and not counted: it is reported as an error row.

**Site score (0-100).** 85% of the average page score plus 15% of the site checks, rounded. The site checks are worth 100 points: robots.txt exists 10, robots.txt does not block the whole site 10, XML sitemap found 20, sitemap declared in robots.txt 5, share of audited pages on HTTPS 10, broken links 15 (reaches zero when one in five checked links is broken), no duplicate titles 10, no duplicate descriptions 5, no duplicate or near-duplicate text 10, share of audited pages answering 2xx 5. `siteChecks` lists the points earned for each one with a note.

**Grades.** A is 90 and up, B 80, C 70, D 60, F below.

Scores describe what a crawler can read in the HTML. They are a consistent yardstick, not Google's own ranking factors.

### ✨ Why choose this Actor

- **One flat table.** Pages and site summaries in one dataset, ready for a spreadsheet, a dashboard or an agent. Every issue has an id, a severity, a message and a fix hint.
- **A score you can explain.** The weights are published above and in the summary rows, so a client can see why a page got 74.
- **Safe to run on any input.** Only public sites are fetched. Private, loopback and link-local addresses, and redirects into them, are refused. The crawler identifies itself and respects robots.txt.
- **Honest failures.** A dead domain, a site that refuses the audit, or a start URL that is not a web page fails the run with a plain message. Errors are reported in uncharged rows, never as fake results.
- **Cheap to run.** Plain HTTP requests, no browser. In a reference run, 30 pages of a documentation site plus a one-page site, with link checking, took 59 seconds at 1 GB.

### 📈 How it compares to alternatives

| | This Actor | Browser-based audit Actors |
|---|---|---|
| How pages are read | Plain HTTP, the HTML the server sends | Headless browser, JavaScript executed |
| Core Web Vitals (LCP, CLS, INP) | Not measured; server answer time, compression, render-blocking scripts and page weight as hints | Measured |
| Client-rendered apps | Flagged with `likelyClientRendered`; scored on the HTML only | Rendered content scored |
| robots.txt | Respected by default, switchable | Varies |
| Duplicate title, description and text detection | Yes, exact and near-duplicate text | Varies |
| Link verification | Yes, with a per-site cap | Varies |
| Score method | Published weights | Often not published |

**Ceilings, stated plainly.** Public pages only, no login. HTML only: pages that need JavaScript to show their content are scored on what the server sends. Response time is measured from the Apify data center, not from your visitors. Each page is read up to 2 MB. Link checking is capped per site (150 links by default and never more than 20 per audited page) and the summary says when the cap cut it short. A site can hold at most 5000 audited pages per run. Duplicate detection compares only the pages that were audited.

### 🚀 How to use

1. Create a free Apify account with $5 credit: https://console.apify.com/sign-up?fpr=vmoqkp
2. Open the Actor, paste the website addresses (a bare domain like `example.com` works) and set **Max pages per site** and **Max Items**.
3. Run it and download the dataset as CSV, Excel, JSON or XML. Filter `type` to `site-summary` for the overview or `page` for the per-page audit.

A minimal input:

```json
{ "startUrls": [{ "url": "https://example.com" }], "maxPagesPerSite": 25, "respectRobotsTxt": true }
```

### 💼 Business use cases

#### 🧑‍💼 Agencies and freelancers

Run the audit before a pitch or at the start of a retainer, sort pages by `score`, and hand the client the `topIssues` list with fix hints. Re-run monthly and compare `siteScore`.

#### 🚀 Pre-launch and migrations

Crawl staging or the new domain, check that every page has a title, one H1, a canonical and no noindex, and that `brokenLinks` and `redirectChain` are empty before the switch.

#### 🛠 Developers and CI

Schedule a run after each deploy and fail the job when `criticalIssues` on any page is above zero or `siteScore` drops. The output is plain JSON.

#### 📈 In-house marketing

Find thin pages, duplicate titles and descriptions, images without alt text and pages missing from the sitemap, then work the list from the most common issue down.

### 🔌 Automating the SEO Audit Tool

Schedule the Actor in Apify and connect the dataset to Make, Zapier, Slack, Airbyte, GitHub or Google Drive. A common setup runs weekly, posts the `siteScore` and the top three issues to Slack, and saves the full dataset to Google Drive.

### 🌟 Beyond business use cases

- **Research:** compare the technical health of many sites in one niche.
- **Personal:** keep a blog or portfolio tidy without a paid suite.
- **Non-profit:** audit a volunteer-run site and fix the biggest problems first.
- **Experimentation:** measure how a change to templates moves the average page score.

### 🤖 Ask an AI assistant about this scraper

Copy a few rows into an AI assistant and ask: "Which five issues cost this site the most points, and which page should I fix first?" Because every issue carries an id, severity and fix hint, the answer stays specific.

### ❓ Frequently Asked Questions

#### ❓ Does it respect robots.txt?

Yes, by default. URLs the file disallows for all crawlers are skipped and counted in `pagesBlockedByRobots`, and a Crawl-delay is honoured up to 10 seconds. Turn **Respect robots.txt** off only for sites you own.

#### ❓ Can I audit any website?

Audit your own sites and sites you have permission to audit. The Actor fetches only public pages, identifies itself as ParseForgeSEOAudit, and refuses private, loopback and link-local addresses.

#### ❓ Does it run JavaScript?

No. It reads the HTML the server returns, like a basic crawler. A page that builds its content in the browser is flagged with `likelyClientRendered` and scored on the HTML it sends.

#### ❓ Does it measure Core Web Vitals?

No, that needs a real browser. The speed group scores hints that come from the HTTP response: time to first byte, compression, render-blocking scripts and page weight.

#### ❓ How are broken links found?

Pages the crawl visited already have a status. Other links are requested with HEAD (then GET if the server dislikes HEAD). 404, 410, 400, most 5xx errors and unresolvable domains count as broken. 401, 403, 429, timeouts and blocked requests count as unverified (`unverifiedLinkCount`), not broken.

#### ❓ Why did a site fail with HTTP 403?

Some sites block unknown crawlers or data-center traffic. The run reports the refusal instead of scoring an error page. Switch on the Apify proxy, set a browser-like user agent, or allow ParseForgeSEOAudit in your firewall.

#### ❓ How many pages can it audit?

Up to 5000 per site and 1,000,000 in total per run. The free plan is limited to 10 pages per run. Set **Max pages per site** and **Max Items** to control the size.

#### ❓ Why are there rows with an error message?

A start URL that does not resolve, is not a web page, is disallowed by robots.txt or points at a private address gets a row with `error` set and no score. These rows are not results.

#### ❓ Can I skip parts of a site?

Yes. **Skip URLs matching** takes text such as `/tag/` or a regular expression wrapped in slashes. **Only crawl URLs matching** restricts the crawl to a section. Tracking parameters like `utm_source` are ignored, and query-string traps are capped.

#### ❓ Are the scores comparable between runs?

Yes, the method and weights are fixed. Scores change only when the site changes, the crawl covers different pages (for example a different page cap), or a server answers slowly on one run.

#### ❓ Does it collect personal data?

No. It reads page structure and public metadata: titles, headings, links, image attributes and headers. It does not extract emails, phone numbers or people.

### 🔌 Integrate with any app

Use the Apify API, webhooks or integrations to trigger a run from your deploy pipeline, send the summary row to a chat channel, or load the dataset into a data warehouse.

### ➡️ Next step

After an audit, hand the URLs to the next tool. Each of these takes URLs as input:

- [Website Content Crawler](https://apify.com/parseforge/website-content-crawler): paste `finalUrl` values into its **Website URLs to crawl** to get the text of the pages you just scored.
- [Contact Info Scraper](https://apify.com/parseforge/contact-info-scraper): paste `siteUrl` values into its **Website URLs to scan** to collect the contact details each site publishes.
- [Screenshot URL](https://apify.com/parseforge/screenshot-url): paste the `finalUrl` of low-scoring pages into its **URLs to capture** for a visual record, including a mobile variant.
- [Broken Link Checker](https://apify.com/parseforge/broken-link-checker): paste the `url` values from `brokenLinks` into its **URLs to check** to confirm fixes after a deploy.

### 🔗 Recommended Actors

- [Website Content Crawler](https://apify.com/parseforge/website-content-crawler)
- [Broken Link Checker](https://apify.com/parseforge/broken-link-checker)
- [Google Search Scraper](https://apify.com/parseforge/google-search-scraper)
- [Screenshot URL](https://apify.com/parseforge/screenshot-url)

> 💡 **Pro Tip:** browse the complete [ParseForge collection](https://apify.com/parseforge).

**🆘 Need Help?** [Open our contact form](https://tally.so/r/BzdKgA)

> **⚠️ Disclaimer:** independent tool, not affiliated with any website it audits or with any search engine; it reads only publicly available pages and respects robots.txt by default. Audit sites you own or have permission to audit.

# Actor input Schema

## `startUrls` (type: `array`):

The sites to audit. Each one gets its own crawl, one row per page and one site summary with a 0-100 score. Bare domains like example.com work too.

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000. Counts audited pages across all sites; the site summary rows are free.

## `maxPagesPerSite` (type: `integer`):

Most pages audited for each website. The crawl starts at the URL you gave and follows internal links breadth first.

## `crawlLinks` (type: `boolean`):

On: crawl the site from each start URL. Off: audit only the URLs you listed.

## `maxDepth` (type: `integer`):

How many clicks away from the start URL the crawl may go. 0 audits only the start URL.

## `useSitemap` (type: `boolean`):

When the links run out before the page limit, continue with the URLs listed in the site's XML sitemap.

## `includeSubdomains` (type: `boolean`):

Also audit subdomains of the start URL (for example blog.example.com when you start at example.com).

## `includeUrlPatterns` (type: `array`):

Optional. Text the URL must contain, or a regular expression wrapped in slashes such as //blog//. The start URL is always audited.

## `excludeUrlPatterns` (type: `array`):

Text the URL contains (for example /tag/), or a regular expression wrapped in slashes such as /?page=\d+/. Useful for archives and filters.

## `respectRobotsTxt` (type: `boolean`):

Skip URLs the site's robots.txt disallows and honour its Crawl-delay (up to 10 seconds). Turn off only for sites you own.

## `requestDelayMs` (type: `integer`):

Minimum pause between two requests to the same site.

## `maxConcurrency` (type: `integer`):

How many pages are fetched at the same time. Lower it for small or slow servers.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for a page before giving up.

## `userAgent` (type: `string`):

Leave empty to identify as ParseForgeSEOAudit. Set your own string if your firewall only lets a specific agent through.

## `checkLinks` (type: `boolean`):

Request the links found on the pages (HEAD first) and report the broken ones per page. Pages the crawl already visited are not requested twice.

## `checkExternalLinks` (type: `boolean`):

Also check outgoing links to other websites. Off: only internal links.

## `maxLinkChecks` (type: `integer`):

Cap on distinct links requested per site, internal ones first. Pages that were crawled are always counted for free.

## `proxyConfiguration` (type: `object`):

Apify proxy settings.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://crawlee.dev"
    }
  ],
  "maxItems": 10,
  "maxPagesPerSite": 25,
  "crawlLinks": true,
  "maxDepth": 6,
  "useSitemap": true,
  "includeSubdomains": false,
  "respectRobotsTxt": true,
  "requestDelayMs": 300,
  "maxConcurrency": 3,
  "requestTimeoutSecs": 20,
  "checkLinks": true,
  "checkExternalLinks": true,
  "maxLinkChecks": 150,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

Scores, status and issue counts per page

## `fullData` (type: `string`):

Complete dataset with all 98 fields

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://crawlee.dev"
        }
    ],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/seo-audit-tool").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://crawlee.dev" }],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/seo-audit-tool").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://crawlee.dev"
    }
  ],
  "maxItems": 10
}' |
apify call parseforge/seo-audit-tool --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/seo-audit-tool"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CAnGq98hlVs4ENcZA/builds/QJntUgZPWlNZZcRgQ/openapi.json
