SEO Site Audit - Technical SEO & Broken Link Checker avatar

SEO Site Audit - Technical SEO & Broken Link Checker

Pricing

Pay per event

Go to Apify Store
SEO Site Audit - Technical SEO & Broken Link Checker

SEO Site Audit - Technical SEO & Broken Link Checker

Technical website quality checker for site owners: crawls a site or checks a URL list for redirect chains, duplicate titles/descriptions, broken links, missing canonicals, hreflang errors, JSON-LD problems and more.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Rod Services

Rod Services

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Technical website quality checker for site owners, SEO consultants and dev teams. SEO Site Audit crawls a website (or checks a fixed list of URLs) and reports the technical signals that affect search visibility: redirect chains, duplicate titles and meta descriptions, broken internal links, missing canonicals, hreflang errors, invalid JSON-LD and images without alt text. Try it instantly with the prefilled example - it audits the Apify Academy docs and returns one row per page with a clear list of issues.

Because it runs on the Apify platform, you get a scheduled job, an API/webhook trigger, and full run history for free - no server to maintain.

What does SEO Site Audit do?

The Actor has two modes:

  • Crawl a site - give it one start URL, and it follows same-host links plus the site's sitemap.xml, respecting robots.txt, up to a page limit you set.
  • Check a list of URLs - give it a fixed list of URLs and it reports HTTP status, redirect chains and response time for each one, without crawling further. This is the cheap mode for monitoring known pages (e.g. after a migration).

For every audited page it records:

  • HTTP status code and the full redirect chain (hop count, final URL, redirect-loop detection)
  • Response time
  • Title and meta description, their length, and duplicates across the site
  • H1 count, canonical tag status, robots meta / X-Robots-Tag noindex signals
  • Hreflang tags, including alternates that don't link back (missing return links)
  • JSON-LD blocks that fail to parse
  • Images missing an alt attribute
  • Broken internal links found on the page
  • Mixed content, Open Graph tags, word count and orphan-page detection (crawl mode)

Each row also carries a structured issues array (severity, code, message) so you can filter or alert on errors vs. warnings vs. notices. A SUMMARY record with aggregate issue counts is written to the key-value store at the end of every run.

Results are saved while the audit runs

Rows are saved to the dataset in batches of 25 pages while the crawl is still going, not only at the end. If a long run hits its timeout, is aborted or the platform migrates it, you keep every page audited up to that point, and an aborted run still writes a partial SUMMARY (complete: false, stoppedBecause: "aborted"). After a migration the run resumes and does not save or charge the same page twice. When your Max cost per run is reached, the audit stops and keeps exactly the pages that fit the budget (stoppedBecause: "max-cost-reached").

Duplicates and orphan pages in incremental mode

Duplicate titles and meta descriptions need the whole site, but early rows are saved before later pages are seen. So:

  • The first page with a given title or meta description is saved without a duplicate flag. Every later page with the same value gets a duplicate-title / duplicate-meta-description warning that names the first page.
  • The complete groups, including the first page, are in SUMMARY.duplicateTitles and SUMMARY.duplicateMetaDescriptions (value, count, urls). SUMMARY.issuesByCode counts every page in a group.
  • inlinks on a row is the number of internal links to that page found so far when the row was saved. Orphan status is only final at the end: rows of sitemap pages with no links yet have isOrphan: null, and the final list is SUMMARY.orphanPageUrls.

Use cases

  • Pre-launch or post-migration QA - catch broken links, redirect loops and missing canonicals before (or right after) a site relaunch.
  • Ongoing technical SEO monitoring - schedule a crawl weekly and alert on new errors via Apify's webhook/integration options.
  • Duplicate content audits - find pages sharing the same title or meta description, a common but easy-to-miss ranking issue.
  • International SEO checks - verify hreflang alternates actually link back to each other.
  • Bulk redirect/status checks - use URL-list mode to verify a list of URLs (e.g. from an old sitemap or a spreadsheet of legacy pages) still resolve correctly after a migration.

How to use SEO Site Audit

  1. Click Try for free (or Start) and open the Input tab.
  2. Choose a mode: Crawl a site (set a Start URL) or Check a list of URLs (paste your URLs).
  3. Optionally adjust Max pages, include/exclude URL patterns, or the proxy configuration.
  4. Click Start and wait for the run to finish.
  5. Open the Output tab to browse issues per page, or download the dataset as JSON/CSV/Excel.

Input

Key input fields (see the Input tab for the full form and tooltips):

FieldDescriptionDefault
modecrawl or urlListcrawl
startUrlCrawl mode: where the crawl beginshttps://docs.apify.com/academy
urlListURL-list mode: URLs to check-
maxPagesMax pages (crawl) / URLs (list) to process20
respectRobotsTxtSkip pages disallowed by robots.txttrue
includeGlobs / excludeGlobsRestrict crawl scope by glob pattern[]
maxConcurrencyParallel requests5
requestTimeoutSecsPer-request timeout30
userAgentUser agent sent with requestsSEO Site Audit bot UA
proxyConfigurationApify datacenter proxy or your own proxy URLs (no residential)disabled

Example input for crawl mode:

{
"mode": "crawl",
"startUrl": "https://docs.apify.com/academy",
"maxPages": 20
}

Output

One dataset item per audited page. Simplified example:

{
"url": "https://docs.apify.com/academy",
"finalUrl": "https://docs.apify.com/academy",
"mode": "crawl",
"statusCode": 200,
"finalStatusCode": 200,
"redirectCount": 0,
"redirectLoop": false,
"responseTimeMs": 177,
"title": "Apify Academy | Academy | Apify Documentation",
"titleLength": 45,
"metaDescription": "Learn everything about web scraping and automation...",
"h1Count": 1,
"canonicalUrl": "https://docs.apify.com/academy",
"canonicalStatus": "self",
"hreflangCount": 2,
"jsonLdCount": 0,
"imagesMissingAlt": 6,
"brokenLinkCount": 0,
"issueCount": 1,
"errorCount": 0,
"warningCount": 0,
"issues": [
{ "severity": "notice", "code": "image-missing-alt", "message": "6 image(s) are missing an alt attribute" }
]
}

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. A SUMMARY record (issue counts by severity/code, duplicate title and meta description groups with their URLs, orphan page URLs, broken-link totals, average response time, whether the run completed) is written to the run's key-value store under the key SUMMARY.

Data table

FieldMeaning
statusCode / finalStatusCodeStatus of the first request / after following redirects
redirectChainEvery hop: URL, status, Location, timing
canonicalStatusself, other, missing, invalid, or broken
hreflangEach alternate with returnLink: true/false/null
brokenLinksInternal links on this page that returned an error or restricted status
inlinksInternal links pointing to this page found by the time the row was saved
isOrphanfalse when the page has inlinks or was not found via the sitemap; null when undecided at save time (see SUMMARY.orphanPageUrls)
issues{ severity, code, message }[] - the full list of findings for the page

Pricing

This Actor uses Pay-Per-Event pricing - you only pay for what gets checked, no compute-unit guessing:

EventWhen it's chargedPrice
apify-actor-startOnce per run$0.001
page-auditedOnce per page in crawl mode (full SEO analysis)$0.006
url-checkedOnce per URL in urlList mode (status/redirect check only)$0.0005

Rough cost examples: auditing 1,000 pages in crawl mode costs about $6.00; checking 1,000 URLs in list mode costs about $0.50. Both include one flat $0.001 start fee per run. See ./PRICING.md in the source repository for the measured compute cost behind these prices and the resulting margin.

FAQ

Does this Actor modify my website? No. It only makes read-only HTTP requests (GET/HEAD); it never submits forms or writes anything.

Does it respect robots.txt? Yes, by default, in crawl mode. Pages disallowed for the configured user agent are not fetched. You can disable this with respectRobotsTxt: false if you are auditing a staging site you own.

Does it check external links? It checks and reports internal broken links (same host). External link targets are counted but not fetched, to keep runs fast and cheap.

Why is urlList mode so much cheaper? It only checks status codes, redirects and response time - it doesn't download and parse full page content, so there's no title/meta/canonical/hreflang/JSON-LD analysis in that mode.

What counts as "same site"? example.com and www.example.com are treated as the same site so crawling isn't blocked by a www redirect.

Which proxies can I use? No proxy (the default), Apify datacenter proxy, or your own proxy URLs. Residential and SERP proxies are not supported. A run that asks for them stops at the start with a clear message and does no work.

Something looks wrong or missing? Please open an issue on the Actor's Issues tab with the input you used - happy to take a look. Custom variations of this audit (extra checks, different scoring, CMS-specific rules) are also available on request.