SEO Audit & Broken Link Checker: Full Site Crawl avatar

SEO Audit & Broken Link Checker: Full Site Crawl

Pricing

from $20.00 / 1,000 page auditeds

Go to Apify Store
SEO Audit & Broken Link Checker: Full Site Crawl

SEO Audit & Broken Link Checker: Full Site Crawl

Crawl any website and get one row per page with its SEO issues: titles, meta descriptions, H1, noindex, canonical, hreflang, structured data, missing alt, broken links, redirects, duplicates and orphan pages. Monitor new and fixed issues. Respects robots.txt.

Pricing

from $20.00 / 1,000 page auditeds

Rating

0.0

(0)

Developer

Bruno Petrelli

Bruno Petrelli

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

12 hours ago

Last modified

Share

Give it a website and get back one row per page with its SEO issues, plus one summary row per website. The Actor follows the site's internal links and its sitemap, the way a search engine does, and checks:

  • Titles and meta descriptions: missing, too long, too short, several <title> tags, and duplicates across the site.
  • Headings: missing H1 or more than one H1.
  • Indexing: noindex in meta robots or in the X-Robots-Tag header, canonical pointing to another page, hreflang, lang attribute.
  • Links: broken internal and external links (404, 410, 5xx, dead domains) with the page they are on, redirects, and orphan pages that are in the sitemap but that no page links to.
  • Content: word count (thin pages), images without alt text, JSON-LD structured data types and invalid JSON-LD blocks.
  • Technical: status code, redirect chain, response time, page size, mixed content (http resources on https pages), mobile viewport.
  • Social previews: Open Graph title, description and image, Twitter card.

Respects robots.txt on every host it reads, including the sites your links point to. No browser and no proxies: plain HTTP requests with an honest User-Agent (seo-audit).

Use it for

  • SEO audits for clients: a complete, sortable list of issues per page in minutes, exportable to CSV or Excel.
  • Broken link checks before a launch or a migration, or every month on a schedule.
  • Site migrations: find redirects, chains and pages that fell out of the sitemap.
  • Content inventories: every page with its title, description, H1, word count and canonical.
  • AI agents: plain JSON per page, callable through the Apify MCP server.

Input

{
"websites": ["userpilot.com", "https://stripe.com/blog/"],
"maxPagesPerSite": 200,
"checkExternalLinks": true
}

Output

Two kinds of rows. Use the Pages and Site summaries views of the dataset, or filter on type.

A page (a real result from a small business site on 2026-09-25; the address was replaced):

{
"type": "page",
"url": "https://example.com/academy",
"statusCode": 200,
"depth": 1,
"responseTimeMs": 2653,
"title": "B2B Breakthrough Academy",
"titleLength": 24,
"metaDescriptionLength": 152,
"h1": ["Become the strategic marketer your business needs.", ""],
"h1Count": 2,
"h2Count": 17,
"canonical": "https://example.com/academy",
"canonicalIsSelf": true,
"noindex": false,
"lang": "",
"viewport": true,
"wordCount": 2165,
"imagesWithoutAlt": 0,
"internalLinksCount": 8,
"externalLinksCount": 13,
"inlinks": 6,
"inSitemap": true,
"brokenLinks": [{ "url": "https://example.thrivecart.com/checkout/", "status": 404, "error": null }],
"issues": ["h1-multiple", "lang-missing", "broken-links"]
}

A site summary (free) has pagesCrawled, pagesWithIssues, issueCounts (how many pages have each issue), duplicateTitles, duplicateDescriptions, brokenLinks, blockedLinks, robotsTxt, sitemaps, sitemapUrlCount, blockedByRobots, notCrawled (pages found beyond your page limit) and avgResponseTimeMs.

Issue codes

CodeMeaning
title-missing, title-too-long, title-too-short, title-multiple, title-duplicateTitle absent, over 60 characters, under 10, more than one <title>, or shared with another page
description-missing, description-too-long, description-too-short, description-duplicateMeta description absent, over 160 characters, under 50, or shared
h1-missing, h1-multipleNo H1, or more than one
noindexThe page asks search engines not to index it
canonical-to-other-pageThe canonical URL is another page
lang-missing, viewport-missingNo lang on <html> (or empty), no mobile viewport
images-missing-alt<img> without an alt attribute (alt="" counts as decorative, not an issue)
thin-contentUnder 200 words
structured-data-invalid-jsonA JSON-LD block that is not valid JSON
mixed-contenthttp:// images, scripts, iframes or stylesheets on an https page
slow-responseOver 3 seconds to download the page
broken-linksThe page links to a URL that answers 404, 410, 5xx or does not resolve
redirectedThe URL redirects (the target is audited as its own row)
http-error, unreachableThe page answers 4xx/5xx, or the server cannot be reached
orphan-pageFound only in the sitemap: no crawled page links to it

The length limits follow what Google usually shows in results (about 60 characters of title and 155 to 160 of description). They are guidelines, not ranking rules.

Monitoring: what changed since the last audit

Schedule the audit (weekly, monthly) with Compare with the previous audit on, and each run tells you what moved:

  • every page row gets changes: newIssues (for example ["broken-links"] after someone deleted a page it links to), fixedIssues, and newPage: true for pages the last audit did not have;
  • the site row gets changes too: how many issues are new and how many were fixed, by issue code (issuesAdded, issuesFixed), the new pages and the pages that are gone (missingPages: removed, redirected elsewhere or no longer linked; null when the page limit cut the crawl, because pages left out this time are not gone), with the date of the audit it was compared with.

The first audit has nothing to compare with, so changes is null. The last audit of each site is kept in a key-value store named seo-audit-state in your own account; a Memory name keeps separate histories (for example one per client).

Pricing

Pay per event: one event per page audited. The site summary rows are free. Set a maximum cost for the run and the Actor stops cleanly when it is reached.

Good to know

  • Only pages on the website's own host are crawled (www. and the bare domain count as the same site). Files such as PDFs and images are not audited as pages, but links to them are checked.
  • Links that answer 401, 403, 429, LinkedIn's 999 or a non-standard 4xx code (such as the 419 Hacker News sends to scripts) are counted as blockedLinks, not as broken: the server refused the checker, which does not mean the page is gone. So are links that answer 405 (they exist but only take other methods, like an API endpoint).
  • The Actor reads the HTML the server sends. Content that only appears after JavaScript runs is not seen (for example, images a page builder draws with JavaScript).
  • Crawling is gentle by default: 3 requests at a time per website.

Support

A false positive or a check you need? Open an issue on the Actor's Issues tab with the page URL. Issues get an answer within a few days.