Technical SEO Audit Tool - Website SEO Checker & Site Audit avatar

Technical SEO Audit Tool - Website SEO Checker & Site Audit

Pricing

Pay per event

Go to Apify Store
Technical SEO Audit Tool - Website SEO Checker & Site Audit

Technical SEO Audit Tool - Website SEO Checker & Site Audit

SEO crawler for technical SEO audits, JS sites too: 0-100 SEO score and fix hints per page, site report (duplicate titles, broken pages, missing H1, meta, alt), optional Core Web Vitals. $5/1k pages.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Yukai Lin

Yukai Lin

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

10 minutes ago

Last modified

Share

What does SEO Audit Tool do?

It crawls your website and checks every page for on-page SEO problems, including sites built with JavaScript or protected against bots: when a site blocks plain requests, the page is fetched from a second network or audited in a real browser automatically, and a page that loads almost empty until JavaScript runs (a client-rendered app shell) is rendered in a real browser too. Each page gets a score from 0 to 100, category scores and a list of issues with a fix hint for every issue. The whole site gets a report: score distribution, most common problems, crawl coverage (indexable / noindex / canonicalized / 4xx / 5xx), sitemap checks, canonical and redirect problems, duplicate titles, H1s and content, and broken internal pages.

25+ checks per page, including:

  • 🏷️ Title: missing, too short, too long (search results cut titles at about 60 characters)
  • πŸ“ Meta description: missing, too short, too long
  • πŸ”  Headings: missing, empty or multiple H1s, skipped heading levels (H2 β†’ H4), full heading outline
  • πŸ”— Canonical & indexing: missing or conflicting canonical, noindex in meta robots or the X-Robots-Tag header, pages that robots.txt blocks for Googlebot
  • πŸ–ΌοΈ Images: images without an alt attribute (alt="" on decorative images is correctly accepted); optional check for broken and oversized images
  • 🌍 hreflang: invalid codes (en-UK, jp), missing x-default, missing self-reference, duplicate codes, broken or redirecting targets, missing return links
  • πŸ“± Mobile & language: viewport tag, lang attribute
  • 🌐 Social & structured data: Open Graph title/image, Twitter card, and JSON-LD and Microdata validated against Google's rich result rules (errors, eligible rich results; RDFa types listed)
  • βš™οΈ Technical: HTTP status, HTTPS, redirect chains, slow server response, very large HTML, thin content (word count)
  • πŸ’” Broken pages: internal links that return 404/410/5xx, with the page that links to them

Site-level checks (in the report):

  • πŸ—ΊοΈ Sitemap: no sitemap found, sitemap URLs that return errors, redirect or are noindex, crawled indexable pages missing from the sitemap
  • β†ͺ️ Canonical tags that point to an error page, a redirect or a noindex page
  • πŸ” Internal links that point to a redirect instead of the final URL
  • πŸ‘― Duplicates: titles, meta descriptions, H1s and identical page text (duplicate content)
  • πŸ“Š Crawl coverage: indexable, noindex, canonicalized, redirected, 4xx, 5xx, blocked by robots.txt, deepest page

Optional:

  • ⚑ Core Web Vitals with Google PageSpeed Insights, no API key needed: performance score, LCP, CLS, TBT, FCP and real-user INP for the pages you choose ($5 per 1,000 pages measured)
  • πŸ“ˆ Monitoring: compare with the previous run and see issues added or resolved per page and the score change

Why use it?

  • Works on JavaScript sites and sites that block bots. Many SEO crawlers analyze static HTML only; this one falls back to a real browser automatically when a site blocks plain requests or a page has almost no text until JavaScript runs.
  • A fix hint for every issue, plus category scores (technical, indexing, meta, headings, content, links, images, social, schema, performance, mobile, international).
  • $5 per 1,000 pages, no extra compute charges. See the price comparison below.
  • Pay only for audited pages: broken pages are reported for free; blocked pages, pages skipped by robots.txt and downloads are not charged.
  • Built for recurring checks: schedule it weekly with Compare with the previous run to get issues added and resolved per page, score changes and new broken pages, or feed the JSON into your dashboard.

How much does it cost?

EventPrice
Audited page$5.00 / 1,000 pages
PageSpeed measurement (optional)$5.00 / 1,000 pages measured

Example: auditing a 1,000-page site costs 1,000 Γ— $0.005 = $5.00 on the Free plan; PageSpeed on 50 of those pages adds 50 Γ— $0.005 = $0.25.

No start fee. Broken pages, failed pages and the site checks (sitemap, canonical targets, image checks, structured data, run-to-run comparison) are free. PageSpeed measurements are only charged when a score comes back. Your maximum charge limit is always respected.

Control your cost

  • Charged: one Audited page event per HTML page audited (charged: true on the row); PageSpeed only when a score comes back.
  • Free (charged: false, the error ends with "(not charged)"): invalid input lines (errorType: "invalid_input"), broken internal pages (a finding), blocked, failed and timed-out pages, downloads (PDF, images) and pages skipped by robots.txt.
  • Before it starts, the run logs its plan: Max pages Γ— price (plus PageSpeed pages if on) is the most it can cost, compared with your maximum charge per run (run options). Small sites have fewer pages and cost less. The plan is saved as costPlan in SUMMARY.
  • When the maximum is reached, the run stops: the status message says so, SUMMARY.status is LIMIT_REACHED and SUMMARY.notProcessed has the number of pages found but not audited and the first 100 of them.

Price comparison (checked September 2026)

ActorPrice per 1,000 pages (free plan)Crawls the siteReal-browser fallback
SEO Audit Crawler (this Actor)$5YesYes
smart-digital/complete-seo-audit-tool$40YesNo (static HTML)
logiover/website-seo-audit-crawler$10YesNo (plain HTTP only)
autofacts/metadata-scraper$5; the max charge per run must be set to at least $0.10Only with followLinks (off by default)No (no JavaScript)
scrapesignal_labs/website-seo-audit$0.50 + platform usageYesNo

We are not the cheapest: scrapesignal_labs costs less per page, but you also pay the Apify compute it uses, and it has no browser fallback. At $5 this Actor is the one with a real-browser fallback, a fix hint per issue and site-level checks (sitemap, canonical targets, duplicates, hreflang return links). For PageSpeed, lizaraco/site-audit charges $10 per 1,000 measurements; this Actor charges $5.

How to use it

  1. Enter your website URL in URLs or domains, e.g. https://www.example.com/ or just example.com (one per line for several sites). A bad line never stops the run: it becomes a row with errorType: "invalid_input" (not charged).
  2. Set Max pages (100 is a good start).
  3. Optional: turn on Also audit pages from the sitemap, Respect robots.txt, Check images, Measure Core Web Vitals or Compare with the previous run, or limit the crawl with include/exclude URL patterns.
  4. Click Start. Open Site report in the Output tab for the summary, or the Pages, Issues and fixes and Category scores tables for details.

Input example

{
"urls": ["https://www.python.org/"],
"crawlScope": "domain",
"maxPages": 100,
"useSitemap": true,
"excludeUrlPatterns": ["**/blog/tag/**"]
}

Crawl scope: Whole website includes subdomains (blog.example.com, shop.example.com); Only this host stays on the start URL's host (www. and the bare domain count as one); Same folder and Only the start URLs narrow it further.

Use urls (a plain list of strings) in API calls. The older startUrls field ([{ "url": "https://…" }]) still works and can hold uploaded or linked URL list files, but Apify rejects the whole request with HTTP 400 if one of its entries is a bare domain or a blank line.

Several sites in one run (agencies). Put one site per line in URLs or domains. Max pages is shared fairly: each site gets its share (Max pages Γ· number of sites) while the others are still being crawled, and a small site leaves its unused share to the others. Set Max pages per site for a fixed cap. Every row has site (host without www.) and startUrl, the sitemap of every site is read (up to 20 sites), SUMMARY.sites has one entry per site (pagesAudited, averageScore, brokenPages, pagesFailed, sitemapFound, sitemapUrls) and the REPORT has a Per site table. Real run (September 2026, maxPages: 12): python.org, books.toscrape.com and webscraper.io got 4 pages each.

{
"urls": ["https://www.python.org/", "books.toscrape.com", "https://webscraper.io/"],
"maxPages": 12
}

Output: one item per page

Real output for https://www.python.org/ (shortened):

{
"url": "https://www.python.org/",
"site": "python.org",
"startUrl": "https://www.python.org/",
"input": "https://www.python.org/",
"success": true,
"mode": "fast",
"via": "direct",
"score": 93,
"issueCount": 2,
"issueCodes": "h1-multiple, canonical-missing",
"issues": [
{ "id": "h1-multiple", "severity": "warning", "message": "Page has 5 H1 headings.", "category": "headings",
"fixHint": "Keep a single main H1 and turn the others into H2 headings.", "impact": "medium" },
{ "id": "canonical-missing", "severity": "notice", "message": "Canonical link is missing.", "category": "indexing",
"fixHint": "Add <link rel=\"canonical\"> pointing to the preferred URL of this page (usually itself).", "impact": "low" }
],
"categoryScores": { "technical": 100, "indexing": 94, "meta": 100, "headings": 85, "images": 100, "social": 100, "...": 100 },
"httpStatus": 200,
"title": "Welcome to Python.org",
"titleLength": 21,
"metaDescriptionLength": 52,
"h1": ["Intuitive Interpretation", "Compound Data Types", "..."],
"headings": [{ "level": 1, "text": "Intuitive Interpretation" }, "..."],
"canonical": null,
"xRobotsTag": null,
"blockedByRobotsTxt": false,
"viewport": true,
"wordCount": 1182,
"contentHash": "9f824c550ac6dbe88bd0688e",
"bytes": 53229,
"imagesMissingAlt": 0,
"nofollowLinks": 0,
"hreflangLinks": [],
"twitterCard": null,
"structuredDataTypes": ["WebSite"]
}

via tells you where the page was read from: backend (our servers), direct (Apify's network, used when a site blocks our servers), vm (our second server, Oracle, different IP) or browser. Failed pages carry an errorType (blocked, not_found, timeout, network…). input (on start pages) is your original line; issueCodes is the issue IDs as one text column for spreadsheets; charged says whether the row was charged.

Output: site report

The key-value store contains REPORT (Markdown, easy to read or share) and SUMMARY (JSON). From the same python.org run (20 pages):

{
"status": "SUCCESS",
"pagesAudited": 20,
"averageScore": 94.4,
"scoreDistribution": { "80-100": 20, "60-79": 0, "40-59": 0, "0-39": 0 },
"categoryAverages": { "technical": 100, "indexing": 94, "meta": 99.7, "headings": 89.5, "...": 100 },
"site": {
"coverage": { "crawled": 20, "indexable": 20, "noindex": 0, "canonicalized": 0, "redirected": 3, "clientErrors4xx": 0, "serverErrors5xx": 0, "blockedByRobotsTxt": 0, "maxDepth": 1 },
"sitemap": { "found": false, "urls": 0 },
"internalLinksToRedirects": [
{ "url": "https://www.python.org/psf/", "finalUrl": "https://www.python.org/psf-landing/", "linkedFrom": "https://www.python.org/" },
{ "url": "https://www.python.org/doc/av", "finalUrl": "https://www.python.org/doc/av/", "linkedFrom": "https://www.python.org/" }
],
"canonicalProblems": [],
"duplicateH1": [],
"duplicateContent": [],
"hreflangProblems": []
}
}

status is SUCCESS, PARTIAL_RESULTS (some pages failed), FAILED, NO_RESULTS or LIMIT_REACHED (stopped at your spending limit). The summary also lists the most common issues (with fix hints), the lowest-scoring pages, duplicate titles and descriptions, and broken internal pages with the page that links to them. It also has sites (one entry per site), invalidInputs, costPlan and, when the run stopped at your limit, notProcessed.

Optional: PageSpeed Insights (Core Web Vitals)

Turn on Measure Core Web Vitals (PageSpeed Insights) to run Google's PageSpeed Insights on the start URL(s) and the next pages audited, up to Max pages to measure (default 5). No setup is needed: without your own key, our servers run Google PageSpeed Insights with our key, and when Google's quota is busy they measure with our own Lighthouse server instead. With your own key, the Actor calls Google directly.

Where to measure from (pagespeedSource): lab results depend on where Lighthouse runs, so you can pick the location. auto (default) uses Google, then our Lighthouse servers; google uses Google only; lighthouse uses our servers only (Taiwan first, then Phoenix, US); lighthouse-tw and lighthouse-us measure from that one location and are never swapped for another (if it is unavailable, the page is reported as not measured and is not charged). Choose Taiwan when your visitors are in Asia. Our servers give lab data only; real-user field data comes from Google. The price is the same for every option.

Each page gets a pagespeed object:

FieldMeaning
performanceScoreLighthouse performance score, 0-100
lcpMs, cls, tbtMs, fcpMs, speedIndexMsLab measurements: Largest Contentful Paint, Cumulative Layout Shift, Total Blocking Time, First Contentful Paint, Speed Index
inpMsInteraction to Next Paint from real-user data (lab tests cannot measure INP), when Google has enough traffic for the page or its site
fieldReal-user (Chrome UX Report) 75th percentiles: lcpMs, inpMs, cls, fcpMs, ttfbMs, category, and source (url or origin)
sourceWho measured: google (Google with our key), vm-lighthouse (our Lighthouse servers: lab data only, no field/inpMs) or google (your key)
node, nodeLocationWith vm-lighthouse: which of our servers measured, tw (Taiwan) or us (Phoenix, US)
ttiMs, opportunitiesTime to interactive and the top savings suggestions (id, title, estimated ms/bytes saved); measured through our servers
error, errorKindWhy a page was not measured (quota, key, page, api, skipped); such pages are not charged

Price: $5 per 1,000 pages measured (pagespeed-audit event), charged only when a score comes back.

Real result (September 2026, mobile, 3 of 5 audited python.org pages measured with pagespeedMaxPages: 3; charged 5 Γ— page audit + 3 Γ— PageSpeed):

{
"url": "https://www.python.org/",
"score": 93,
"pagespeed": {
"strategy": "mobile",
"performanceScore": 78,
"lcpMs": 4802,
"cls": 0.003,
"tbtMs": 0,
"fcpMs": 2718,
"speedIndexMs": 2718,
"inpMs": 88,
"field": { "source": "url", "category": "FAST", "lcpMs": 1225, "inpMs": 88, "cls": 0, "fcpMs": 1158, "ttfbMs": 526 },
"lighthouseVersion": "13.5.0",
"measuredUrl": "https://www.python.org/"
}
}

The lab LCP (4.8 s) is much slower than what real visitors see (field.lcpMs 1.2 s): Lighthouse simulates a slow phone on a throttled mobile network. Use the lab numbers to compare pages and track changes, and the field numbers for what your visitors actually experience. Lab scores also vary a few points between runs.

Optional: use your own Google API key (about 5 minutes)

You do not need a key. Add one only if you want the calls to count against your own Google quota (for example very large runs): put it in Your own Google API key. If our daily capacity is ever used up, the remaining pages are marked errorKind: "quota" and not charged, and adding your key lets you measure them.

  1. Open Google Cloud Console and create a project (no billing account needed).
  2. Enable the PageSpeed Insights API.
  3. Go to APIs & Services β†’ Credentials β†’ Create credentials β†’ API key.
  4. Recommended: edit the key, keep Application restrictions: None (Apify servers have changing IPs) and set API restrictions to PageSpeed Insights API only.
  5. Paste the key into the Actor input. It is stored encrypted.

The API is free; your daily quota is shown in Google Cloud Console under the PageSpeed Insights API's Quotas page.

After a quota or key error the Actor stops measuring for the rest of the run (those pages are not charged). SUMMARY.pagespeed has the average performance score and the pages sorted from slowest, and the REPORT gets a PageSpeed table.

Structured data (JSON-LD and Microdata)

Every page gets a structuredData object checked with the same rules as Schema Markup Validator (by TidyTools): status, formats (json-ld, microdata), types, eligible richResults, errorCount, warningCount, the first errors with code and path, and rdfaTypes (RDFa is detected, not validated). It does not change the page score. Real result from webscraper.io (September 2026), a page marked up with Microdata only:

"structuredData": {
"status": "errors",
"formats": ["microdata"],
"types": ["OfferCatalog", "VideoObject"],
"richResults": [],
"errorCount": 2,
"warningCount": 2,
"errors": [
{ "code": "REQUIRED_MISSING", "type": "VideoObject", "property": "thumbnailUrl", "path": "microdata[1]", "format": "microdata" },
{ "code": "REQUIRED_MISSING", "type": "VideoObject", "property": "uploadDate", "path": "microdata[1]", "format": "microdata" }
],
"rdfaTypes": []
}

Turn off Validate structured data to skip it.

Monitoring: compare with the previous run

Turn on Compare with the previous run and schedule the Actor (for example weekly). Each page gets change (new, changed, unchanged), previousScore, scoreDelta, issuesAdded and issuesResolved. SUMMARY.changes and a Since the last run section in the REPORT show the average score change, issues added and resolved (by issue ID), the biggest score drops and gains, new broken pages, fixed broken pages and pages no longer found. Runs with the same Monitor name are compared (default: start URL + crawl scope). The first run is the baseline. If a run stops before the whole site is crawled (Max pages or spending limit), pages it did not reach are kept and not reported as removed.

Real second run on webscraper.io (September 2026; for the test, the previous snapshot of two pages was edited):

[
{ "url": "https://webscraper.io/", "score": 96, "change": "changed", "previousScore": 90, "scoreDelta": 6,
"issuesAdded": ["structured-data-missing"], "issuesResolved": ["title-missing"] },
{ "url": "https://webscraper.io/pricing", "score": 94, "change": "changed", "previousScore": 100, "scoreDelta": -6,
"issuesAdded": ["canonical-other", "heading-skip", "structured-data-missing"], "issuesResolved": [] }
]

REPORT excerpt from the same run:

## Since the last run
- Average score: **97.3** (+2.3 from 95)
- Issues added: **4**, resolved: **1** on 6 page(s) seen in both runs
- New pages: **0**, pages no longer found: **0**, new broken pages: **0**, broken pages fixed: **1**

Pages are audited and charged the same way with or without monitoring.

Use with AI agents (MCP)

Connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/seo-audit-crawler) to Claude, Cursor or any MCP client, then ask e.g. "Audit the first 50 pages of example.com and list the three fixes that would raise the score most."

{ "urls": ["https://example.com/"], "maxPages": 50 }

Failed pages are not charged and carry an errorType, so the agent can tell a blocked site from a broken page.

Use it from code and integrations

Run it from your own code with the Apify API. This call waits for the run and returns the results as JSON (replace YOUR_TOKEN with your Apify API token):

curl -X POST "https://api.apify.com/v2/acts/tidytools~seo-audit-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls":["https://www.python.org/"],"maxPages":20}'

The synchronous endpoint waits up to 5 minutes. For bigger runs, start the run with POST https://api.apify.com/v2/acts/tidytools~seo-audit-crawler/runs and read the dataset when it finishes, or use the apify-client package for JavaScript or Python.

Schedules and integrations: run it daily or weekly with Apify Schedules, get a webhook when a run finishes, or send the results to Zapier, Make, n8n, Google Sheets, Slack and other apps with Apify integrations. Results can be exported as JSON, CSV, Excel or XML.

Scoring

Each page starts at 100. Critical issues (e.g. missing title, missing H1, noindex, HTTP errors, no HTTPS) cost 15 points, warnings (e.g. missing meta description, images without alt text, invalid hreflang) cost 5, and notices (e.g. missing canonical, Open Graph or structured data) cost 2. Each category score starts at 100 and loses three times those points for issues in that category.

Redirects that only add a language or country folder (/pricing β†’ /en-sg/pricing) depend on where the request comes from, so they are counted separately and not reported as problems.

Advanced settings

  • Plain HTTP requests from: Auto (recommended) reads pages from our servers and retries from Apify's network when a site blocks data-center requests, then from our second server (Oracle, different IP), before falling back to a browser. You can force any one of them.
  • Time limit per page: default 120 seconds for reading a page, all retries and the browser fallback included. A slower page gets errorType: "timeout" and is not charged. 0 = no limit.
  • Proxy: optional Apify Proxy for requests sent from Apify's network (billed to your Apify account).

Limitations

  • On-page checks: no backlinks or keyword rankings. Core Web Vitals only with the optional PageSpeed Insights measurement (no key needed).
  • JSON-LD and Microdata are validated; RDFa types are only listed.
  • Only public http/https pages; no logins.
  • Download files (PDF, ZIP, installers) are skipped.
  • Sitemap URLs outside the crawl are checked as a sample (up to 100 per run).

FAQ

Does it measure Core Web Vitals? Yes, optionally. Turn on PageSpeed Insights for the pages you choose to get the performance score, LCP, CLS, TBT, FCP and real-user INP. No API key is needed; it costs $5 per 1,000 pages measured.

Can it audit JavaScript websites? Yes. In Auto mode a page that comes back almost empty (under about 200 characters of visible text, typical for client-rendered React or Vue apps) is audited again in a real browser, and a site that blocks plain requests is fetched from a second network or a browser (mode: "browser" on the row). Text or tags that JavaScript adds to a page that already has content are only seen with Page fetching (HTTP or browser) set to Browser.

Does it check backlinks or keyword rankings? No. It is a technical and on-page SEO checker: titles, meta, headings, canonical, hreflang, structured data, broken pages, sitemap and duplicates.

Can I track SEO issues over time? Yes. With Compare with the previous run on, each run is compared with the previous one and shows the issues added or resolved per page and the score change.

Support

Open an issue in the Issues tab with your start URL and input. Issues are checked regularly.