Website SEO Audit - Redirects, Canonicals, Indexing avatar

Website SEO Audit - Redirects, Canonicals, Indexing

Pricing

from $1.80 / 1,000 audited pages

Go to Apify Store
Website SEO Audit - Redirects, Canonicals, Indexing

Website SEO Audit - Redirects, Canonicals, Indexing

Audit supplied pages for redirects, canonical links, index directives and HTML metadata. Compare changes or verify expected migration targets. Static HTML, robots-aware, no API key.

Pricing

from $1.80 / 1,000 audited pages

Rating

0.0

(0)

Developer

Brandt May

Brandt May

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Share

Website SEO Audit

Audit supplied public pages for HTTP redirects, HTML canonical links, indexing directives and page metadata. Optionally check expected migration targets or compare selected fields with an earlier observation.

Each result shows the evidence and specific findings. This is an audit of downloaded HTML and HTTP responses; it does not claim to know whether a search engine has indexed a page or which canonical it selected.

Quick start

Empty input audits the WordPress and Next.js home pages:

{}

Supply your own public pages:

{
"urls": ["https://wordpress.org/", "https://nextjs.org/"],
"comparePrevious": false,
"maxRunSeconds": 240
}

No API key is required. The Actor fetches only the supplied pages and follows their HTTP redirects. It respects robots.txt at each destination, uses no login or proxy, and does not discover or crawl additional links.

Inputs

FieldDefaultMeaning
urlsWordPress and Next.js home pagesUp to 100 public HTTP(S) URLs. Empty or omitted uses these samples.
expectedFinalUrlsnoneObject mapping supplied source URLs to expected final migration targets.
comparePreviousfalseOpt in to storing and comparing observations.
snapshotNamedefaultIndependent comparison namespace, up to 100 characters.
maxRunSeconds24030–3,600 seconds, also bounded by the platform deadline.

Public HTTP(S) URLs on ports 80 or 443 are supported. Credentials, local/private destinations and unsafe redirects are rejected. URL fragments are removed when normalizing inputs.

Check a migration target

{
"urls": ["http://wordpress.org/"],
"expectedFinalUrls": {
"http://wordpress.org/": "https://wordpress.org/"
}
}

Every mapping key must also appear in urls. The Actor compares the normalized observed final URL with the supplied expectation and checks the HTML canonical when present. This is an exact normalized URL comparison: a different path, query string, hostname or trailing slash can produce a mismatch. The target is not fetched separately unless it is reached through a redirect or also listed in urls.

A single redirect is recorded without a redirect-chain warning; multiple redirects generate a finding. Up to eight redirects are followed. Redirect loops or excessive redirects are source failures, reported in SUMMARY.

Output

One dataset row represents one page with an actual HTTP response that could be audited. JSON preserves nested findings and comparisons; Apify can also export CSV and Excel.

FieldsMeaning
inputUrl, finalUrl, expectedFinalUrlRequested, observed and optionally expected URLs.
httpStatus, contentType, redirectChainFinal response status/type and every followed redirect.
title, titleCount, metaDescriptionMetadata from the downloaded HTML.
h1sH1 text values in the HTML.
canonicalUrl, canonicalLinksResolved single canonical when available and the raw canonical link values.
robotsDirectives, xRobotsTag, hasNoindexDirectiveObserved HTML/HTTP directives, including raw crawler scope.
hreflangDeclared language links; targets are resolved but not fetched or validated.
internalLinkCountSame-origin anchor occurrences, including duplicate links.
imageCount, imagesMissingAltImage elements and elements with no alt attribute; an empty alt attribute counts as present.
wordCountApproximate whitespace-delimited body text count after selected noncontent elements are removed.
issuesCode, severity and explanation for each concrete finding.
checkedAt, comparisonStatus, previousCheckedAt, comparisonObservation time and optional historical comparison.

Findings cover HTTP errors, multiple redirects, unexpected final targets, missing/duplicate titles or meta descriptions, missing H1, observed noindex/none directives, multiple or invalid canonical links, a canonical pointing elsewhere, and canonical migration mismatches.

A canonical pointing elsewhere or a noindex directive may be intentional. The Actor reports those as observations, not automatic evidence of a broken page. hasNoindexDirective can be true for a crawler-specific directive; inspect the raw directives to understand its scope. See Google's documentation on robots directives and canonical signals.

Compare observations

{
"urls": ["https://wordpress.org/"],
"comparePrevious": true,
"snapshotName": "homepage-monitor"
}

The first successful observation has comparisonStatus: "baseline_created" with no prior comparison. Later results report changedFields, newIssueCodes and resolvedIssueCodes. Each page is still emitted and billed when unchanged.

The comparison covers HTTP status, final URL, title, meta description, H1 values, canonical URL, the observed noindex flag and issue codes. It is not a comparison of the entire document: changes to other fields, full page text, image URLs or raw directives without a changed flag are not tracked.

Snapshots live in the named key-value store maydit-seo-audit-baselines. The source URL, expected target and snapshot name identify a comparison scope. Changing an expected target starts a new scope. Avoid overlapping runs using the same namespace and scope.

Actual HTTP error pages are valid observations and can replace the previous baseline. Network failures, robots denial, oversized content and unsupported successful response types do not produce an observation and preserve previous history.

Coverage, billing and failures

SUMMARY reports requested/emitted counts, failed URLs and reasons, robots-denied URLs, unprocessed inputs and deadline status. Usable page results are retained when another URL fails. If no page can be audited, the run fails with a diagnostic.

Launch price: $3 per 1,000 audited pages ($0.003 each) on Free/Bronze, plus an Actor Start event of $0.00005. Silver is 20% lower and Gold/Platinum/Diamond 40% lower. The live Pricing tab is authoritative.

Actual HTTP error responses, such as 404 or 500, are billable audit results, as are unchanged pages and first snapshots. Network or robots errors have no dataset row and no result charge. Each emitted page counts once, regardless of findings or redirect count. Consult the published pricing tab for platform resource charges.

Limits

  • JavaScript is not executed. Client-rendered metadata and content added after page load may be absent.
  • This does not test real search-engine indexing, rankings, traffic, backlink quality, Core Web Vitals or mobile rendering.
  • It checks supplied pages only. Internal links are counted; broken links elsewhere on the site are not checked.
  • Only HTML and XHTML successful responses are audited. PDFs, JSON and other successful response types are reported as unsupported.
  • Hreflang is extracted, not checked for reciprocity, language validity or reachable targets. Canonical destination pages are not fetched separately.
  • HTML canonical links are inspected; HTTP Link header canonicals and XML sitemaps are not analyzed.
  • Meta refresh and JavaScript redirects are not followed. Redirect tracking is HTTP only.
  • Source content is bounded to 3 MiB per response. Authentication, anti-bot challenges and robots restrictions are not bypassed.
  • Findings are descriptive checks, not a universal SEO score or ranking guarantee.

Development

Run npm test for the fixture suite. Use isolated CRAWLEE_STORAGE_DIR directories for local execution. The three example inputs are drafts for verification. The comparison example explicitly opts into state storage; it does not create a schedule or send alerts.