SEO Audit & Broken Link Checker: Full Site Crawl
Pricing
from $20.00 / 1,000 page auditeds
SEO Audit & Broken Link Checker: Full Site Crawl
Crawl any website and get one row per page with its SEO issues: titles, meta descriptions, H1, noindex, canonical, hreflang, structured data, missing alt, broken links, redirects, duplicates and orphan pages. Monitor new and fixed issues. Respects robots.txt.
Pricing
from $20.00 / 1,000 page auditeds
Rating
0.0
(0)
Developer
Bruno Petrelli
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
12 hours ago
Last modified
Categories
Share
Give it a website and get back one row per page with its SEO issues, plus one summary row per website. The Actor follows the site's internal links and its sitemap, the way a search engine does, and checks:
- Titles and meta descriptions: missing, too long, too short, several
<title>tags, and duplicates across the site. - Headings: missing H1 or more than one H1.
- Indexing:
noindexin meta robots or in theX-Robots-Tagheader, canonical pointing to another page, hreflang,langattribute. - Links: broken internal and external links (404, 410, 5xx, dead domains) with the page they are on, redirects, and orphan pages that are in the sitemap but that no page links to.
- Content: word count (thin pages), images without alt text, JSON-LD structured data types and invalid JSON-LD blocks.
- Technical: status code, redirect chain, response time, page size, mixed content (http resources on https pages), mobile viewport.
- Social previews: Open Graph title, description and image, Twitter card.
Respects robots.txt on every host it reads, including the sites your links point to. No browser and no proxies: plain HTTP requests with an honest User-Agent (seo-audit).
Use it for
- SEO audits for clients: a complete, sortable list of issues per page in minutes, exportable to CSV or Excel.
- Broken link checks before a launch or a migration, or every month on a schedule.
- Site migrations: find redirects, chains and pages that fell out of the sitemap.
- Content inventories: every page with its title, description, H1, word count and canonical.
- AI agents: plain JSON per page, callable through the Apify MCP server.
Input
{"websites": ["userpilot.com", "https://stripe.com/blog/"],"maxPagesPerSite": 200,"checkExternalLinks": true}
Output
Two kinds of rows. Use the Pages and Site summaries views of the dataset, or filter on type.
A page (a real result from a small business site on 2026-09-25; the address was replaced):
{"type": "page","url": "https://example.com/academy","statusCode": 200,"depth": 1,"responseTimeMs": 2653,"title": "B2B Breakthrough Academy","titleLength": 24,"metaDescriptionLength": 152,"h1": ["Become the strategic marketer your business needs.", ""],"h1Count": 2,"h2Count": 17,"canonical": "https://example.com/academy","canonicalIsSelf": true,"noindex": false,"lang": "","viewport": true,"wordCount": 2165,"imagesWithoutAlt": 0,"internalLinksCount": 8,"externalLinksCount": 13,"inlinks": 6,"inSitemap": true,"brokenLinks": [{ "url": "https://example.thrivecart.com/checkout/", "status": 404, "error": null }],"issues": ["h1-multiple", "lang-missing", "broken-links"]}
A site summary (free) has pagesCrawled, pagesWithIssues, issueCounts (how many pages have each issue), duplicateTitles, duplicateDescriptions, brokenLinks, blockedLinks, robotsTxt, sitemaps, sitemapUrlCount, blockedByRobots, notCrawled (pages found beyond your page limit) and avgResponseTimeMs.
Issue codes
| Code | Meaning |
|---|---|
title-missing, title-too-long, title-too-short, title-multiple, title-duplicate | Title absent, over 60 characters, under 10, more than one <title>, or shared with another page |
description-missing, description-too-long, description-too-short, description-duplicate | Meta description absent, over 160 characters, under 50, or shared |
h1-missing, h1-multiple | No H1, or more than one |
noindex | The page asks search engines not to index it |
canonical-to-other-page | The canonical URL is another page |
lang-missing, viewport-missing | No lang on <html> (or empty), no mobile viewport |
images-missing-alt | <img> without an alt attribute (alt="" counts as decorative, not an issue) |
thin-content | Under 200 words |
structured-data-invalid-json | A JSON-LD block that is not valid JSON |
mixed-content | http:// images, scripts, iframes or stylesheets on an https page |
slow-response | Over 3 seconds to download the page |
broken-links | The page links to a URL that answers 404, 410, 5xx or does not resolve |
redirected | The URL redirects (the target is audited as its own row) |
http-error, unreachable | The page answers 4xx/5xx, or the server cannot be reached |
orphan-page | Found only in the sitemap: no crawled page links to it |
The length limits follow what Google usually shows in results (about 60 characters of title and 155 to 160 of description). They are guidelines, not ranking rules.
Monitoring: what changed since the last audit
Schedule the audit (weekly, monthly) with Compare with the previous audit on, and each run tells you what moved:
- every page row gets
changes:newIssues(for example["broken-links"]after someone deleted a page it links to),fixedIssues, andnewPage: truefor pages the last audit did not have; - the site row gets
changestoo: how many issues are new and how many were fixed, by issue code (issuesAdded,issuesFixed), the new pages and the pages that are gone (missingPages: removed, redirected elsewhere or no longer linked;nullwhen the page limit cut the crawl, because pages left out this time are not gone), with the date of the audit it was compared with.
The first audit has nothing to compare with, so changes is null. The last audit of each site is kept in a key-value store named seo-audit-state in your own account; a Memory name keeps separate histories (for example one per client).
Pricing
Pay per event: one event per page audited. The site summary rows are free. Set a maximum cost for the run and the Actor stops cleanly when it is reached.
Good to know
- Only pages on the website's own host are crawled (
www.and the bare domain count as the same site). Files such as PDFs and images are not audited as pages, but links to them are checked. - Links that answer 401, 403, 429, LinkedIn's 999 or a non-standard 4xx code (such as the 419 Hacker News sends to scripts) are counted as
blockedLinks, not as broken: the server refused the checker, which does not mean the page is gone. So are links that answer 405 (they exist but only take other methods, like an API endpoint). - The Actor reads the HTML the server sends. Content that only appears after JavaScript runs is not seen (for example, images a page builder draws with JavaScript).
- Crawling is gentle by default: 3 requests at a time per website.
Support
A false positive or a check you need? Open an issue on the Actor's Issues tab with the page URL. Issues get an answer within a few days.