Website SEO Audit - Redirects, Canonicals, Indexing
Pricing
from $1.80 / 1,000 audited pages
Website SEO Audit - Redirects, Canonicals, Indexing
Audit supplied pages for redirects, canonical links, index directives and HTML metadata. Compare changes or verify expected migration targets. Static HTML, robots-aware, no API key.
Pricing
from $1.80 / 1,000 audited pages
Rating
0.0
(0)
Developer
Brandt May
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Website SEO Audit
Audit supplied public pages for HTTP redirects, HTML canonical links, indexing directives and page metadata. Optionally check expected migration targets or compare selected fields with an earlier observation.
Each result shows the evidence and specific findings. This is an audit of downloaded HTML and HTTP responses; it does not claim to know whether a search engine has indexed a page or which canonical it selected.
Quick start
Empty input audits the WordPress and Next.js home pages:
{}
Supply your own public pages:
{"urls": ["https://wordpress.org/", "https://nextjs.org/"],"comparePrevious": false,"maxRunSeconds": 240}
No API key is required. The Actor fetches only the supplied pages and follows their HTTP redirects. It respects robots.txt at each destination, uses no login or proxy, and does not discover or crawl additional links.
Inputs
| Field | Default | Meaning |
|---|---|---|
urls | WordPress and Next.js home pages | Up to 100 public HTTP(S) URLs. Empty or omitted uses these samples. |
expectedFinalUrls | none | Object mapping supplied source URLs to expected final migration targets. |
comparePrevious | false | Opt in to storing and comparing observations. |
snapshotName | default | Independent comparison namespace, up to 100 characters. |
maxRunSeconds | 240 | 30–3,600 seconds, also bounded by the platform deadline. |
Public HTTP(S) URLs on ports 80 or 443 are supported. Credentials, local/private destinations and unsafe redirects are rejected. URL fragments are removed when normalizing inputs.
Check a migration target
{"urls": ["http://wordpress.org/"],"expectedFinalUrls": {"http://wordpress.org/": "https://wordpress.org/"}}
Every mapping key must also appear in urls. The Actor compares the normalized observed final URL with the supplied expectation and checks the HTML canonical when present. This is an exact normalized URL comparison: a different path, query string, hostname or trailing slash can produce a mismatch. The target is not fetched separately unless it is reached through a redirect or also listed in urls.
A single redirect is recorded without a redirect-chain warning; multiple redirects generate a finding. Up to eight redirects are followed. Redirect loops or excessive redirects are source failures, reported in SUMMARY.
Output
One dataset row represents one page with an actual HTTP response that could be audited. JSON preserves nested findings and comparisons; Apify can also export CSV and Excel.
| Fields | Meaning |
|---|---|
inputUrl, finalUrl, expectedFinalUrl | Requested, observed and optionally expected URLs. |
httpStatus, contentType, redirectChain | Final response status/type and every followed redirect. |
title, titleCount, metaDescription | Metadata from the downloaded HTML. |
h1s | H1 text values in the HTML. |
canonicalUrl, canonicalLinks | Resolved single canonical when available and the raw canonical link values. |
robotsDirectives, xRobotsTag, hasNoindexDirective | Observed HTML/HTTP directives, including raw crawler scope. |
hreflang | Declared language links; targets are resolved but not fetched or validated. |
internalLinkCount | Same-origin anchor occurrences, including duplicate links. |
imageCount, imagesMissingAlt | Image elements and elements with no alt attribute; an empty alt attribute counts as present. |
wordCount | Approximate whitespace-delimited body text count after selected noncontent elements are removed. |
issues | Code, severity and explanation for each concrete finding. |
checkedAt, comparisonStatus, previousCheckedAt, comparison | Observation time and optional historical comparison. |
Findings cover HTTP errors, multiple redirects, unexpected final targets, missing/duplicate titles or meta descriptions, missing H1, observed noindex/none directives, multiple or invalid canonical links, a canonical pointing elsewhere, and canonical migration mismatches.
A canonical pointing elsewhere or a noindex directive may be intentional. The Actor reports those as observations, not automatic evidence of a broken page. hasNoindexDirective can be true for a crawler-specific directive; inspect the raw directives to understand its scope. See Google's documentation on robots directives and canonical signals.
Compare observations
{"urls": ["https://wordpress.org/"],"comparePrevious": true,"snapshotName": "homepage-monitor"}
The first successful observation has comparisonStatus: "baseline_created" with no prior comparison. Later results report changedFields, newIssueCodes and resolvedIssueCodes. Each page is still emitted and billed when unchanged.
The comparison covers HTTP status, final URL, title, meta description, H1 values, canonical URL, the observed noindex flag and issue codes. It is not a comparison of the entire document: changes to other fields, full page text, image URLs or raw directives without a changed flag are not tracked.
Snapshots live in the named key-value store maydit-seo-audit-baselines. The source URL, expected target and snapshot name identify a comparison scope. Changing an expected target starts a new scope. Avoid overlapping runs using the same namespace and scope.
Actual HTTP error pages are valid observations and can replace the previous baseline. Network failures, robots denial, oversized content and unsupported successful response types do not produce an observation and preserve previous history.
Coverage, billing and failures
SUMMARY reports requested/emitted counts, failed URLs and reasons, robots-denied URLs, unprocessed inputs and deadline status. Usable page results are retained when another URL fails. If no page can be audited, the run fails with a diagnostic.
Launch price: $3 per 1,000 audited pages ($0.003 each) on Free/Bronze, plus an Actor Start event of $0.00005. Silver is 20% lower and Gold/Platinum/Diamond 40% lower. The live Pricing tab is authoritative.
Actual HTTP error responses, such as 404 or 500, are billable audit results, as are unchanged pages and first snapshots. Network or robots errors have no dataset row and no result charge. Each emitted page counts once, regardless of findings or redirect count. Consult the published pricing tab for platform resource charges.
Limits
- JavaScript is not executed. Client-rendered metadata and content added after page load may be absent.
- This does not test real search-engine indexing, rankings, traffic, backlink quality, Core Web Vitals or mobile rendering.
- It checks supplied pages only. Internal links are counted; broken links elsewhere on the site are not checked.
- Only HTML and XHTML successful responses are audited. PDFs, JSON and other successful response types are reported as unsupported.
- Hreflang is extracted, not checked for reciprocity, language validity or reachable targets. Canonical destination pages are not fetched separately.
- HTML canonical links are inspected; HTTP
Linkheader canonicals and XML sitemaps are not analyzed. - Meta refresh and JavaScript redirects are not followed. Redirect tracking is HTTP only.
- Source content is bounded to 3 MiB per response. Authentication, anti-bot challenges and robots restrictions are not bypassed.
- Findings are descriptive checks, not a universal SEO score or ranking guarantee.
Development
Run npm test for the fixture suite. Use isolated CRAWLEE_STORAGE_DIR directories for local execution. The three example inputs are drafts for verification. The comparison example explicitly opts into state storage; it does not create a schedule or send alerts.