Website SEO Audit – Broken Links, Titles & Sitemap avatar

Website SEO Audit – Broken Links, Titles & Sitemap

Pricing

$4.00 / 1,000 audited pages

Go to Apify Store
Website SEO Audit – Broken Links, Titles & Sitemap

Website SEO Audit – Broken Links, Titles & Sitemap

Technical SEO audit of whole websites: score 0–100 per page, issues with fix hints, broken links, redirect chains, titles, meta descriptions, headings, canonical, hreflang, JSON-LD, sitemap and robots.txt checks plus a site summary with the top 10 fixes. No browser.

Pricing

$4.00 / 1,000 audited pages

Rating

0.0

(0)

Developer

Steven Kramp

Steven Kramp

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Website SEO Audit – Broken Links, Titles, Sitemap & Score

Crawl any website and get a technical SEO audit in minutes: a score from 0 to 100 for every page, every issue with its priority and a concrete fix hint, broken internal and external links, redirect chains, duplicate titles, canonical and hreflang errors, structured data errors, sitemap and robots.txt checks – plus one site summary with the top 10 fixes.

$4 per 1,000 audited pages. You only pay for HTML pages that were actually audited. Error pages, redirects, pages blocked by bot protection or access rules and the site summary are free.

No browser, no setup: fast plain HTTP requests, polite by default (2 parallel requests, Crawl-delay honoured in full), and robots.txt is always respected – also for every redirect target.

What you get for each page

FieldExample
url, finalUrl, statusCode, redirectChainhttps://example.com/old → 301 → https://example.com/new
fetchStatus, error, redirectTargetok · redirect · http-error · blocked · denied · error · robots-disallowed · robots-unreachable
score (0–100), issueCount, issues[]89 · {code, priority, message, fixHint}
indexable, indexableReasonfalse · noindex / canonical / http-404
title, titleLength, metaDescription, metaDescriptionLength"Pricing – Example" · 17
h1Count, h1, headingSkips[]1 · [{from: "h2", to: "h4"}]
canonical, canonicalSelfhttps://example.com/pricing · true
hreflang[] (with returnLink check), lang[{hreflang: "de", url: …, returnLink: false}]
schemaTypes[], schemaErrors[]["Organization", "WebSite"] · ["JSON-LD block 1: invalid JSON …"]
ogTags, viewportog:title, og:description, og:image …
imagesTotal, imagesWithoutAlt, wordCount12 · 2 · 840
internalLinks, externalLinks, brokenLinks[]48 · 6 · [{url, statusCode, anchorText, internal}]
responseTimeMs, depth, foundVia, inSitemap180 · 2 · link · true

The site summary (one item per website)

  • Average score, pages audited, pages with issues, issues by priority and by code
  • Top 10 fixes sorted by priority and number of affected pages, each with fix hint and example URLs
  • Duplicate titles and meta descriptions (grouped)
  • Broken links (internal and external) with the page they are on
  • Sitemap checks: sitemap found (robots.txt, /sitemap.xml, index files, .gz), sitemap URLs that do not return 200, sitemap URLs blocked by robots.txt, sitemap pages that no crawled page links to
  • robots.txt state and crawl statistics (success rate, blocked and denied pages, stop reason, links not checked)

The summary is the last item of the dataset and is also saved as SUMMARY in the key-value store.

Checks

PriorityIssue codes
Highhttp-4xx, http-5xx, not-https, missing-title, broken-internal-links
Mediummissing-meta-description, missing-h1, duplicate-title, noindex-in-sitemap, redirect-chain, canonical-target-error, multiple-canonicals, hreflang-no-return-link, hreflang-invalid-code, hreflang-target-error, schema-invalid, broken-external-links, missing-viewport, slow-response
Lowtitle-too-long, title-too-short, meta-description-too-short, meta-description-too-long, duplicate-meta-description, multiple-h1, heading-skip, missing-canonical, missing-lang, missing-og-tags, images-missing-alt, low-word-count, noindex
Sitesitemap-missing, sitemap-urls-not-200, sitemap-urls-blocked-by-robots, sitemap-orphan-pages, robots-txt-missing, robots-txt-unreachable, robots-txt-no-sitemap

Score per page: 100 minus 20 per high, 8 per medium and 3 per low issue type. Search snippet checks (title and description length, missing description, canonical, word count) only apply to indexable pages, so intentional noindex pages are not punished twice.

Use cases

  • SEO agencies and freelancers: run a full audit for a new client in minutes and hand over the top 10 fixes.
  • Website launches and relaunches: find broken links, redirect chains and missing titles before Google does.
  • Monitoring: schedule a weekly run and watch the average score and broken links over time.
  • Content teams: use Only pages changed since to check just the pages published or edited recently.
  • AI agents: give an assistant a structured, prioritised to-do list for a website.

How to use

  1. Add one or more start URLs (usually the homepage). Each website is crawled on its own domain only.
  2. Set max pages per website (default 200). Optional: include/exclude patterns, max click depth, extra sitemap URLs.
  3. Run it, then export as JSON, CSV or Excel, or connect to Google Sheets, Make, Zapier or n8n.
{
"startUrls": [{"url": "https://www.example.com/"}],
"maxPages": 200,
"useSitemap": true,
"checkExternalLinks": true,
"excludePatterns": ["*/tag/*", "?replytocom="],
"outputMode": "both"
}

The dataset has two ready-made views: Pages (score, status and key fields per page) and Issues per page (one row per issue with priority and fix hint).

Pricing

Pay per event: $0.004 per audited page ($4 per 1,000 pages). Platform usage is included. Redirects, 4xx/5xx pages, blocked or denied pages, start URLs that could not be fetched and the site summary are not charged. If you set a maximum cost per run, the crawl stops before it.

Use with AI agents (MCP)

This Actor works as a tool for AI assistants and agents – Claude, ChatGPT, Cursor, VS Code, n8n and other MCP clients – through Apify's hosted MCP server. Add this server URL to your client:

https://mcp.apify.com?tools=stevenkramp/website-seo-audit

Sign in with your Apify account when asked. Your agent can then call the Actor in plain language, for example: "Audit https://www.example.com (up to 100 pages) and list the top 10 SEO fixes with the affected URLs." – and gets clean, structured JSON back. Runs started by your agent are normal Actor runs on your Apify account at the same pay-per-event price.

More from stevenkramp

Other Actors by the same developer – same quality standards, pay only for results:

Search & trends

Apps

Research & media

Websites & places

FAQ

Does it render JavaScript? No. It reads the HTML the server sends, like most search engine first passes. Sites that build all content with JavaScript in the browser show little text and few links – the audit makes that visible.

What about robots.txt? It is always respected and cannot be switched off: disallowed URLs are never fetched – not even as the target of a redirect, because redirects are followed hop by hop and every target is checked first (they are counted in the summary). Crawl-delay is honoured in full; with a long delay the crawl stops in time before the run timeout instead of going faster. Outgoing links are only checked where the target site's robots.txt allows it. The crawler identifies itself as StevenKrampBot with a link to this page.

A site blocks the crawler – what happens? Pages behind a bot challenge (e.g. Cloudflare) are reported honestly as fetchStatus: "blocked" instead of empty data. Pages that refuse the crawler with 401, 403 or 429 (login, firewall, rate limit) are reported as fetchStatus: "denied" – not as errors of the website. Neither is charged, and neither counts as a success in crawl.successRate. After 30 blocked or denied requests in a row the crawl of that website stops. Every start URL gets an item, even when it cannot be fetched at all (robots.txt disallows it or is unreachable or behind a bot check, DNS error) – with fetchStatus and error saying why. The Proxy option can help for sites that block data center traffic.

Which links count as broken? 404, 410, other 4xx and 5xx responses, DNS, connection and SSL errors – also when the target site cannot even deliver its robots.txt because the domain or certificate is dead. 401/403/405/429 responses and bot challenges are not counted as broken, because the page usually works for people. Outgoing links that the target's robots.txt disallows, or whose robots.txt answers with a server error, are not fetched; they are counted in crawl.externalLinksBlockedByRobots and crawl.externalLinksNotChecked.

Are all external links checked? Each outgoing URL is checked once. Link checking gets a time budget per website (at least 2 minutes, or as long as the crawl itself took), so a page full of outgoing links cannot make a run explode; links left unchecked are counted in crawl.externalLinksNotChecked.

How big can a crawl be? Up to 10,000 pages per website. The audit keeps only the SEO data of each page, not its HTML, so memory stays small (tested on Apify: 2,000 pages peaked at 162 MB of the default 1 GB). The Actor also watches memory and the run timeout: it stops fetching in time (crawl.stoppedBy: memory-limit or time-limit) and still saves all results. For big sites raise the run timeout.

Personal data? The audit stores SEO data only: URLs, status codes, titles, descriptions, headings and counts – no page text. Contact links (mailto:, tel:) are ignored, and e-mail addresses (also written as name (at) domain) and phone numbers in international, German and North American formats are masked in titles, descriptions, headings and link texts.

Something broken? Open an issue in the Issues tab. We fix problems quickly.