Technical SEO Audit: Crawl, Issues & Scores per Page avatar

Technical SEO Audit: Crawl, Issues & Scores per Page

Under maintenance

Pricing

$10.00 / 1,000 page auditeds

Go to Apify Store
Technical SEO Audit: Crawl, Issues & Scores per Page

Technical SEO Audit: Crawl, Issues & Scores per Page

Under maintenance

Point it at a site and get an issue-coded audit for every page plus a site summary with duplicate titles, broken links and status counts. 18 issue checks with a 0-100 score, robots.txt respected, no browser needed. Built for agencies, in-house SEO and pre-launch checks.

Pricing

$10.00 / 1,000 page auditeds

Rating

0.0

(0)

Developer

Paul Vasquez

Paul Vasquez

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

a day ago

Last modified

Categories

Share

Technical SEO Audit

Audit public websites with a bounded breadth-first HTTP crawl. This Python 3.12 Apify Actor extracts technical SEO signals from server-delivered HTML, calculates a transparent issue score, and produces one dataset row per successfully fetched 2xx HTML response. It uses httpx and Beautiful Soup without a browser. JavaScript rendering, Core Web Vitals, keyword ranking, backlink discovery, and search-engine indexing verification are outside its scope.

Quick start

Create a Python 3.12 environment, install dependencies, and run from this directory:

python -m venv .venv
.venv/Scripts/python.exe -m pip install -r requirements.txt
apify validate-schema .actor/input_schema.json
.venv/Scripts/python.exe -m unittest discover -s tests -v
powershell -ExecutionPolicy Bypass -File validation/run_live.ps1

The script copies INPUT.json into fresh local Apify storage and starts python -m src. Logs and aggregate evidence go under validation/; dataset rows remain under ignored storage/. The supplied input crawls Python.org and Apify with 30 candidate URLs per site. Docker uses apify/actor-python:3.12 and the same entry point.

Input and crawl scope

startUrls is a required nonempty array of absolute HTTP(S) URL strings. Seeds sharing a hostname share one crawl budget and summary. Fragments are removed; paths, queries, and trailing slashes remain significant. URLs containing credentials are rejected. maxPagesPerSite defaults to 100 and limits attempted candidate URLs, including robots exclusions, failures, and non-HTML responses. This bounds work even when no usable HTML exists.

sameDomainOnly defaults to true and restricts page requests and redirect targets to the seed hostname. includeSubdomains, default false, admits child hostnames, but not parent or sibling hosts. It does not infer registrable domains. Setting sameDomainOnly false admits external pages within the original seed budget. Internal/external classification still uses the seed hostname rules.

maxDepth defaults to 3; seeds have depth zero. concurrency defaults to 5, with a maximum of 20. Each depth completes before the next begins. Discovery retains at most twenty times the candidate budget. Sites run sequentially. timeoutSecs defaults to 30 per network operation. Redirects have a ten-hop limit. Optional proxyConfiguration passes through the SDK to httpx; direct access is the default.

respectRobots defaults to true. Cached per-origin robots.txt wildcard user-agent groups determine permission before page requests and redirect hops. Rules support wildcards, terminal dollar anchors, longest-match selection, and Allow precedence on ties. Missing robots files permit access; 401, 403, 429, server errors, and network failures conservatively deny access. Robots redirects are followed to retrieve policy. Sitemap directives populate hasSitemapRef; sitemaps are not traversed. Disabling enforcement still retrieves robots for metadata. Crawl-delay is not implemented.

Results and interpretation

Rows include source/final URLs, redirect-hop objects, depth, status, milliseconds, decoded response size, content type, title/description lengths, canonical mismatch, robots directives, language, viewport, headings, word count, images, alt omissions, link counts, hreflang, nested JSON-LD types, Open Graph, Twitter card, sitemap reference, mixed content, issues, and score.

Empty alt attributes count as intentional decorative alternatives. Word counts exclude scripts, styles, templates, and noscript content. Mixed-content checks inspect HTML asset attributes, srcset, and inline CSS URLs, but not downloaded CSS or JavaScript. Canonicals use normalized absolute URLs. Indexability means successful HTML without noindex/none in robots meta or X-Robots-Tag; this heuristic does not guarantee search-engine indexing.

After crawling, exact nonempty titles and descriptions are compared within each site. Broken internal links require observed HTTP statuses of 400 or higher. Unvisited, robots-blocked, and network-failed targets are not declared broken. Redirect aliases can produce separate rows because they represent separate requested URLs. Link counts count occurrences; broken-link arrays deduplicate.

checkExternalLinks, default false, checks up to 200 distinct external URLs per site using HEAD. Redirects and robots rules apply. HEAD refusals have no GET fallback. Results populate summary externalLinkStatuses and brokenLinks; external failures do not lower page scores.

Score and pricing

Start at 100, subtract each distinct issue's weight once, and clamp at zero:

WeightIssue codes
20NOINDEX
10MISSING_TITLE, CANONICAL_MISMATCH, BROKEN_INTERNAL_LINK, MIXED_CONTENT
5MISSING_META_DESCRIPTION, MISSING_H1, IMAGES_MISSING_ALT, THIN_CONTENT, MISSING_VIEWPORT, DUPLICATE_TITLE
3TITLE_TOO_LONG, MULTIPLE_H1, SLOW_RESPONSE, REDIRECT_CHAIN, MISSING_LANG, DUPLICATE_META_DESCRIPTION
2META_DESCRIPTION_TOO_LONG

Thresholds are title >60 characters, description >160, content <200 words, response >2000 milliseconds, and redirect chain >1 hop. Intentional noindex pages and short landing pages may legitimately trigger issues. Fetch timing includes permission checks and redirects.

SUMMARY-<host> contains pages crawled/skipped, average score, issue counts, broken links, duplicate-title URL groups, status counts, elapsed seconds, fetch failures, attempted URLs, external statuses, and written rows. Counts describe the audited sample, even if a hosted event limit stops dataset writes.

The sole PPE event is page-audited at $0.01 per emitted audit row. Skipped, failed, non-HTML responses and summaries are free. Configure this event in Console and disable synthetic charges before publication. Local runs do not bill. See VALIDATION.md for measured evidence and remaining deployment checks.

Example output

One real saved dataset row, trimmed by omitting fields only. Source: storage/live-20260926-053309/datasets/default/000000001.json. This is historical validation evidence, not a live response.

{
"url": "https://www.python.org/",
"finalUrl": "https://www.python.org/",
"statusCode": 200,
"depth": 0,
"title": "Welcome to Python.org",
"indexable": true,
"h1Count": 5,
"imagesMissingAlt": 0,
"issues": [
"DUPLICATE_META_DESCRIPTION",
"MULTIPLE_H1",
"SLOW_RESPONSE"
],
"score": 91
}

This saved Python.org row illustrates the documented issue weights: three distinct three-point issues reduce 100 to 91. Other page measurements are omitted. Indexable is the actor heuristic, not confirmation that a search engine indexed this URL.

Use cases

  • An SEO agency can turn page-level issue rows into a client remediation queue, reviewing intentional noindex pages separately.
  • A website migration manager can inspect redirects and canonical mismatches on a bounded sample after a deployment.
  • An ecommerce content operations team can review missing titles, descriptions, and image alternatives across sampled product pages.
  • A web development agency can compare repeated crawl exports for the same input scope when checking a server-rendered template change.

Pricing example: 1,000 successful page-audited events x $0.01 = $10.00, computed from .actor/pay_per_event.json. This is the declared event subtotal; it does not verify active hosted billing or include any separately applicable platform or proxy costs.

Limitations

Results describe the fetched HTML sample within the candidate and depth limits. JavaScript-generated content and unvisited link targets can remain unseen. The score is an issue-weight calculation, not a ranking forecast. Compare runs with consistent settings and review flagged pages before deciding whether a template change is necessary.