Website SEO Audit & Broken Link Checker avatar

Website SEO Audit & Broken Link Checker

Pricing

Pay per usage

Go to Apify Store
Website SEO Audit & Broken Link Checker

Website SEO Audit & Broken Link Checker

Crawls a list of start URLs and reports technical-SEO health per page: broken links, redirect chains, missing meta descriptions, alt-text coverage, structured data and mixed content — a broken link checker and SEO audit crawler in one.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Catalyst

Catalyst

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

an hour ago

Last modified

Share

Run it on Apify: apify.com/catalyst_prime/website-audit: free, no setup, runs in the browser.

Crawls a list of start URLs and reports the technical-SEO health of every page it finds: broken links, redirect chains, missing meta descriptions, alt-text coverage, structured data and mixed content. One row per page, so you can sort and filter the whole crawl in a spreadsheet.

So it works as a broken link checker, a redirect chain checker and a meta description audit at once, over a list of sites or a single site.

This repository is the Actor's source. The Actor itself runs on the Apify platform.

Not a content extractor. This does not convert pages to Markdown for an LLM/RAG pipeline. That job is already well served on the Store. This tool answers a different question: is this page's plumbing broken?

Input

FieldTypeDefaultDescription
startUrlsarray of stringsnoneStart URLs to crawl. Supply this or the alias below. Accepts a JSON array, a single string, or a comma/newline separated list.
urlsarraynoneAlias of startUrls.
maxPagesinteger20Total pages to crawl across every start URL combined. Clamped to 1-50.
maxDepthinteger2How many link-hops from each start URL to follow. 0 audits only the start URLs themselves. Clamped to 0-5.
followExternalLinksbooleanfalseWhen true, links to a different domain than the start URL are crawled too (still counted against maxPages/maxDepth). When false, off-site links are still checked and reported in brokenLinks, just not crawled as their own pages.
timeoutSecondsinteger15Per-request timeout, clamped to 3-60.

Output

One dataset row per page actually crawled:

{
"url": "https://example.com/",
"ok": true,
"httpStatus": 200,
"depth": 0,
"title": "Example Domain",
"metaDescriptionLength": 0,
"brokenLinks": ["https://example.com/old-page"],
"redirectChain": [],
"hasStructuredData": false,
"imagesMissingAlt": 2,
"mixedContentFound": false,
"latencyMs": 143,
"charged": false
}
  • url is the page as requested, before any redirects; httpStatus and the parsed fields (title, metaDescriptionLength, etc.) reflect the final page after following them.
  • redirectChain lists every URL hopped through after the first request, in order, ending with the final URL. Empty when there was no redirect.
  • brokenLinks lists links found on the page (internal or external) that returned a network error or an HTTP status of 400+, checked up to 20 per page. A link already known from having been crawled or checked elsewhere in the same run is reused rather than re-requested.
  • hasStructuredData is true if the page has a non-empty application/ld+json script or a microdata (itemscope) element.
  • imagesMissingAlt counts <img> tags with no alt attribute or an empty one.
  • mixedContentFound is true only for an HTTPS page that loads a resource (script, image, stylesheet, iframe, etc.) over plain http://. A protocol-relative URL (//cdn...) is not mixed content.
  • httpStatus and the parsed fields are omitted from a row whose request never completed at all; error is set on that row instead. charged reports whether the row was billed (see Pricing): a row is charged whenever a response was actually received, even a 404 or 500, since the status itself is the data this Actor sells.

How the crawl works

Breadth-first from each start URL. maxPages is a single global budget shared across every start URL and every depth. Once it's spent, no further pages are fetched, wherever they were discovered. maxDepth counts link-hops from whichever start URL began that branch, not from the page that happens to link to it.

The crawler identifies itself as CatalystWebsiteAuditBot and reads robots.txt once per host (cached for the rest of the run): a page disallowed there is skipped entirely: not fetched, not reported, and its slot in maxPages is not spent. It never sends credentials, so it never reaches anything behind a login; a page that requires one will show up as its own HTTP status (401/403) rather than being silently skipped.

Pricing

Free. This Actor has no price set, so a run costs you only your own Apify platform usage.

The code supports pay-per-event billing on a single page-audit event, priced per page successfully fetched, if a price is ever set on the Apify Console. Nothing is hardcoded: it reads its own current price from the platform at startup and runs unmetered when there isn't one.

Example

Crawling https://example.com with maxDepth: 0 (audit just the one page, no following) returns a single row confirming the page loads clean: httpStatus: 200, no broken links, no mixed content, zero images missing alt text. Raise maxDepth and maxPages to audit a whole section of a site in one run instead of one page at a time.

Four ready-to-run examples, each with real input and real output:

Other Actors from the same author