Hreflang Cluster Regression Audit avatar

Hreflang Cluster Regression Audit

Pricing

$0.25 / completed report

Go to Apify Store
Hreflang Cluster Regression Audit

Hreflang Cluster Regression Audit

Audit supplied hreflang graphs for missing return links, self references, conflicting targets and regression changes. Export evidence, clusters, coverage, CSV and HTML.

Pricing

$0.25 / completed report

Rating

0.0

(0)

Developer

Gilad Ronen

Gilad Ronen

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What does Hreflang Cluster Regression Audit do?

Audit a supplied hreflang graph and compare release snapshots with traceable evidence. Submit page observations from your own crawler or export, then download the missing-link findings, cluster edge matrix, coverage, and baseline changes. This Actor runs offline: it does not crawl websites, translate content, detect page language, or predict indexing or rankings.

Use it after a locale rollout, multilingual template change, or crawl refresh. Findings retain exact source and target URLs, original annotation language, customer source labels, and JSON pointers back to the supplied observation. The rules follow Google's localized-page guidance; the report is a reproducible audit of your data, not a Google certification.

How to use the audit

  1. Export one observation per page URL. Keep IDs as strings, including leading zeros.
  2. Paste the observations into Pages on the Input tab. Each page must include an alternates array. For partial exports, explicitly set alternatesComplete: false.
  3. Optionally specify the languages required on every supplied page in policy.expectedLanguages. Set policy.requireXDefault: true only when your own policy requires a fallback annotation.
  4. Run the Actor and download the JSON, issue CSV, cluster edge CSV, or standalone HTML report.
  5. For a later release, pass the previous unchanged full JSON object as previousReport along with the new pages. API access, saved tasks, scheduling, and integrations can automate this handoff.

Input and coverage

The Input tab accepts JSON only. Convert CSV crawl exports to the documented page structure before submission. Up to 1,000 pages, 10,000 alternate edges, and 4 MB of total JSON, including a previous report, are supported. Long URLs or dense findings may hit the output cap first; split such graphs into smaller complete groups. Unknown fields, duplicate page URLs, duplicate supplied IDs, numeric IDs, status strings, and unsupported policy values are rejected rather than silently coerced.

{
"pages": [
{
"id": "0001",
"url": "https://example.com/en",
"status": 200,
"canonical": "https://example.com/en",
"alternatesComplete": true,
"alternates": [
{"language": "en-GB", "url": "https://example.com/en", "source": "html:head/link[1]"},
{"language": "de", "url": "https://example.de/de", "source": "html:head/link[2]"}
]
}
]
}

url must be an absolute HTTP(S) URL. status, when supplied, is an observed integer from 100 through 599. canonical is optional; this version compares absolute HTTP(S) canonical values and marks unsupported values for review. indexability can be indexable, noindex, or unknown; this is your observation, never an inferred crawl result. alternatesComplete defaults to true, meaning the provided list contains every annotation in that observation. Missing status, canonical, indexability, or target pages remain unknown.

Alternate targets retain their original values. Relative and protocol-relative alternate URLs are findings; any displayed resolution is explanatory and is excluded from reciprocity checks. Comparison preserves URL path case, trailing slashes, query order, encoded values, and fragments. Standard URL parsing normalizes scheme/host case, default ports, and dot segments. No request is made to any URL.

Checks and interpretation

The audit checks self references, conflicting destinations for the same language, repeated annotations, known missing returns, supplied non-200 pages and targets, supplied noindex, and canonical alignment warnings. Non-self canonical values require review: a different canonical is not automatically an invalid SEO configuration. Content equivalence and canonical language are not assessed.

Language values are compared without case sensitivity. The supported form is an assigned ISO 639-1 language, optional registered ISO 15924 script, and optional assigned ISO 3166-1 alpha-2 region, plus x-default. Examples include en-GB, zh-Hans, and zh-Hans-US; en-gb is equivalent to en-GB. Numeric regions such as es-419, three-letter languages, extensions, and reserved region values such as en-UK are unsupported. be and uk are valid language codes for Belarusian and Ukrainian. Cross-domain alternates are allowed.

Cluster membership means weak graph connectivity, including referenced but unsupplied absolute targets. It does not prove that pages are equivalent translations. The audit checks self and return links; it does not require identical all-to-all language sets across a connected cluster. Reciprocal subsets can have no findings. Supply policy.expectedLanguages if every page must contain a specific complete language set. Missing x-default produces no issue unless explicitly required by that policy or requireXDefault.

A missing return is proven only when the target page and its complete alternate list are supplied. Partial lists and absent targets produce unknown findings. “No findings in supplied graph” describes the tested annotations; it does not certify unobserved metadata or live pages.

Output and regression comparison

One dataset item contains the complete report. Download the dataset as JSON; use the dedicated flat CSV and standalone HTML files for human review.

FieldContents
summaryPage, edge, cluster, severity, coverage, and baseline counts
rowsStable issue ID, severity, reason, exact URLs, language, and source references
pageResultsPage ID, observed fields, alternate-list coverage, and assessment
edgesExact supplied annotations, normalized language, target coverage, return-link outcome
clustersWeakly connected URLs, languages, supplied and unsupplied counts
changesNew, unchanged, resolved, or unobserved baseline findings

OUTPUT stores the complete JSON; issues.csv is the flat issue queue; clusters.csv is the flat annotated edge matrix. Isolated pages and cluster totals are retained in JSON and HTML. report.html is self-contained, contains no scripts, and is downloaded as an attachment. CSV cells are protected against spreadsheet formula execution. IDs and source evidence are stable for repeat findings, although cluster IDs change when member URLs change.

Baseline resolutions require relevant pages and fields to be covered in both reports. Removed pages, incomplete lists, omitted statuses, unsupported canonical observations, and changed requirements are conservatively marked unobserved when they prevent comparison. A missing finding is not sufficient proof of a fix. Keep the baseline JSON unchanged; integrity and version validation prevent accidental reuse of edited or incompatible results. The input size limit may require retaining smaller cluster-specific baselines.

Price and recovery

The intended Store price is $0.25 per completed report using the report-completed event, including platform usage. There is no startup or per-issue fee. Invalid input, insufficient budget, and reports rejected by the export limits are not charged a report event. The runtime verifies the actual positive configured report price and rejects any other positive-priced event.

The dataset is the primary deliverable. It is persisted before the report charge. Convenience exports are written afterward, so an interrupted run can have a complete charged dataset with missing files. Resurrecting that same run validates and regenerates its exports while preserving one dataset item and one charge; an interrupted charge can be completed once. Starting a new run is a new billable report. If charged data is missing or the saved input/result was changed, recovery stops for inspection. Local SDK simulation does not prove cloud billing or cloud UI behavior.

Questions and support

Provide page observations you are authorized to process; no credentials, network access, proxies, or external model service are required. This Actor does not modify your site. Share a small synthetic reproduction through the Issues tab for unexpected results, and use the API tab for programmatic integration. There are no translation, indexing, accuracy, or ranking guarantees. Language/script/region snapshots are versioned with the engine and may require a future update when standards change.