Technical SEO Audit: Crawl, Issues & Scores per Page
Under maintenancePricing
$10.00 / 1,000 page auditeds
Technical SEO Audit: Crawl, Issues & Scores per Page
Under maintenancePoint it at a site and get an issue-coded audit for every page plus a site summary with duplicate titles, broken links and status counts. 18 issue checks with a 0-100 score, robots.txt respected, no browser needed. Built for agencies, in-house SEO and pre-launch checks.
Pricing
$10.00 / 1,000 page auditeds
Rating
0.0
(0)
Developer
Paul Vasquez
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
a day ago
Last modified
Categories
Share
Technical SEO Audit
Audit public websites with a bounded breadth-first HTTP crawl. This Python 3.12 Apify Actor extracts technical SEO signals from server-delivered HTML, calculates a transparent issue score, and produces one dataset row per successfully fetched 2xx HTML response. It uses httpx and Beautiful Soup without a browser. JavaScript rendering, Core Web Vitals, keyword ranking, backlink discovery, and search-engine indexing verification are outside its scope.
Quick start
Create a Python 3.12 environment, install dependencies, and run from this directory:
python -m venv .venv.venv/Scripts/python.exe -m pip install -r requirements.txtapify validate-schema .actor/input_schema.json.venv/Scripts/python.exe -m unittest discover -s tests -vpowershell -ExecutionPolicy Bypass -File validation/run_live.ps1
The script copies INPUT.json into fresh local Apify storage and starts python -m src. Logs and aggregate evidence go under validation/; dataset rows remain under ignored storage/. The supplied input crawls Python.org and Apify with 30 candidate URLs per site. Docker uses apify/actor-python:3.12 and the same entry point.
Input and crawl scope
startUrls is a required nonempty array of absolute HTTP(S) URL strings. Seeds sharing a hostname share one crawl budget and summary. Fragments are removed; paths, queries, and trailing slashes remain significant. URLs containing credentials are rejected. maxPagesPerSite defaults to 100 and limits attempted candidate URLs, including robots exclusions, failures, and non-HTML responses. This bounds work even when no usable HTML exists.
sameDomainOnly defaults to true and restricts page requests and redirect targets to the seed hostname. includeSubdomains, default false, admits child hostnames, but not parent or sibling hosts. It does not infer registrable domains. Setting sameDomainOnly false admits external pages within the original seed budget. Internal/external classification still uses the seed hostname rules.
maxDepth defaults to 3; seeds have depth zero. concurrency defaults to 5, with a maximum of 20. Each depth completes before the next begins. Discovery retains at most twenty times the candidate budget. Sites run sequentially. timeoutSecs defaults to 30 per network operation. Redirects have a ten-hop limit. Optional proxyConfiguration passes through the SDK to httpx; direct access is the default.
respectRobots defaults to true. Cached per-origin robots.txt wildcard user-agent groups determine permission before page requests and redirect hops. Rules support wildcards, terminal dollar anchors, longest-match selection, and Allow precedence on ties. Missing robots files permit access; 401, 403, 429, server errors, and network failures conservatively deny access. Robots redirects are followed to retrieve policy. Sitemap directives populate hasSitemapRef; sitemaps are not traversed. Disabling enforcement still retrieves robots for metadata. Crawl-delay is not implemented.
Results and interpretation
Rows include source/final URLs, redirect-hop objects, depth, status, milliseconds, decoded response size, content type, title/description lengths, canonical mismatch, robots directives, language, viewport, headings, word count, images, alt omissions, link counts, hreflang, nested JSON-LD types, Open Graph, Twitter card, sitemap reference, mixed content, issues, and score.
Empty alt attributes count as intentional decorative alternatives. Word counts exclude scripts, styles, templates, and noscript content. Mixed-content checks inspect HTML asset attributes, srcset, and inline CSS URLs, but not downloaded CSS or JavaScript. Canonicals use normalized absolute URLs. Indexability means successful HTML without noindex/none in robots meta or X-Robots-Tag; this heuristic does not guarantee search-engine indexing.
After crawling, exact nonempty titles and descriptions are compared within each site. Broken internal links require observed HTTP statuses of 400 or higher. Unvisited, robots-blocked, and network-failed targets are not declared broken. Redirect aliases can produce separate rows because they represent separate requested URLs. Link counts count occurrences; broken-link arrays deduplicate.
checkExternalLinks, default false, checks up to 200 distinct external URLs per site using HEAD. Redirects and robots rules apply. HEAD refusals have no GET fallback. Results populate summary externalLinkStatuses and brokenLinks; external failures do not lower page scores.
Score and pricing
Start at 100, subtract each distinct issue's weight once, and clamp at zero:
| Weight | Issue codes |
|---|---|
| 20 | NOINDEX |
| 10 | MISSING_TITLE, CANONICAL_MISMATCH, BROKEN_INTERNAL_LINK, MIXED_CONTENT |
| 5 | MISSING_META_DESCRIPTION, MISSING_H1, IMAGES_MISSING_ALT, THIN_CONTENT, MISSING_VIEWPORT, DUPLICATE_TITLE |
| 3 | TITLE_TOO_LONG, MULTIPLE_H1, SLOW_RESPONSE, REDIRECT_CHAIN, MISSING_LANG, DUPLICATE_META_DESCRIPTION |
| 2 | META_DESCRIPTION_TOO_LONG |
Thresholds are title >60 characters, description >160, content <200 words, response >2000 milliseconds, and redirect chain >1 hop. Intentional noindex pages and short landing pages may legitimately trigger issues. Fetch timing includes permission checks and redirects.
SUMMARY-<host> contains pages crawled/skipped, average score, issue counts, broken links, duplicate-title URL groups, status counts, elapsed seconds, fetch failures, attempted URLs, external statuses, and written rows. Counts describe the audited sample, even if a hosted event limit stops dataset writes.
The sole PPE event is page-audited at $0.01 per emitted audit row. Skipped, failed, non-HTML responses and summaries are free. Configure this event in Console and disable synthetic charges before publication. Local runs do not bill. See VALIDATION.md for measured evidence and remaining deployment checks.
Example output
One real saved dataset row, trimmed by omitting fields only. Source: storage/live-20260926-053309/datasets/default/000000001.json. This is historical validation evidence, not a live response.
{"url": "https://www.python.org/","finalUrl": "https://www.python.org/","statusCode": 200,"depth": 0,"title": "Welcome to Python.org","indexable": true,"h1Count": 5,"imagesMissingAlt": 0,"issues": ["DUPLICATE_META_DESCRIPTION","MULTIPLE_H1","SLOW_RESPONSE"],"score": 91}
This saved Python.org row illustrates the documented issue weights: three distinct three-point issues reduce 100 to 91. Other page measurements are omitted. Indexable is the actor heuristic, not confirmation that a search engine indexed this URL.
Use cases
- An SEO agency can turn page-level issue rows into a client remediation queue, reviewing intentional noindex pages separately.
- A website migration manager can inspect redirects and canonical mismatches on a bounded sample after a deployment.
- An ecommerce content operations team can review missing titles, descriptions, and image alternatives across sampled product pages.
- A web development agency can compare repeated crawl exports for the same input scope when checking a server-rendered template change.
Pricing example: 1,000 successful page-audited events x $0.01 = $10.00, computed from .actor/pay_per_event.json. This is the declared event subtotal; it does not verify active hosted billing or include any separately applicable platform or proxy costs.
Limitations
Results describe the fetched HTML sample within the candidate and depth limits. JavaScript-generated content and unvisited link targets can remain unseen. The score is an issue-weight calculation, not a ranking forecast. Compare runs with consistent settings and review flagged pages before deciding whether a template change is necessary.