# Hreflang Cluster Regression Audit (`h_murdock/hreflang-cluster-audit`) Actor

Audit supplied hreflang graphs for missing return links, self references, conflicting targets and regression changes. Export evidence, clusters, coverage, CSV and HTML.

- **URL**: https://apify.com/h\_murdock/hreflang-cluster-audit.md
- **Developed by:** [Gilad Ronen](https://apify.com/h_murdock) (community)
- **Categories:** Marketing, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.25 / completed report

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Hreflang Cluster Regression Audit do?

**Audit a supplied hreflang graph and compare release snapshots with traceable evidence.** Submit page observations from your own crawler or export, then download the missing-link findings, cluster edge matrix, coverage, and baseline changes. This Actor runs offline: it does not crawl websites, translate content, detect page language, or predict indexing or rankings.

Use it after a locale rollout, multilingual template change, or crawl refresh. Findings retain exact source and target URLs, original annotation language, customer source labels, and JSON pointers back to the supplied observation. The rules follow [Google's localized-page guidance](https://developers.google.com/search/docs/specialty/international/localized-versions); the report is a reproducible audit of your data, not a Google certification.

### How to use the audit

1. Export one observation per page URL. Keep IDs as strings, including leading zeros.
2. Paste the observations into **Pages** on the Input tab. Each page must include an `alternates` array. For partial exports, explicitly set `alternatesComplete: false`.
3. Optionally specify the languages required on every supplied page in `policy.expectedLanguages`. Set `policy.requireXDefault: true` only when your own policy requires a fallback annotation.
4. Run the Actor and download the JSON, issue CSV, cluster edge CSV, or standalone HTML report.
5. For a later release, pass the previous unchanged full JSON object as `previousReport` along with the new pages. API access, saved tasks, scheduling, and integrations can automate this handoff.

### Input and coverage

The Input tab accepts JSON only. Convert CSV crawl exports to the documented page structure before submission. Up to **1,000 pages, 10,000 alternate edges, and 4 MB of total JSON**, including a previous report, are supported. Long URLs or dense findings may hit the output cap first; split such graphs into smaller complete groups. Unknown fields, duplicate page URLs, duplicate supplied IDs, numeric IDs, status strings, and unsupported policy values are rejected rather than silently coerced.

```json
{
  "pages": [
    {
      "id": "0001",
      "url": "https://example.com/en",
      "status": 200,
      "canonical": "https://example.com/en",
      "alternatesComplete": true,
      "alternates": [
        {"language": "en-GB", "url": "https://example.com/en", "source": "html:head/link[1]"},
        {"language": "de", "url": "https://example.de/de", "source": "html:head/link[2]"}
      ]
    }
  ]
}
```

`url` must be an absolute HTTP(S) URL. `status`, when supplied, is an observed integer from 100 through 599. `canonical` is optional; this version compares absolute HTTP(S) canonical values and marks unsupported values for review. `indexability` can be `indexable`, `noindex`, or `unknown`; this is your observation, never an inferred crawl result. `alternatesComplete` defaults to `true`, meaning the provided list contains every annotation in that observation. Missing status, canonical, indexability, or target pages remain unknown.

Alternate targets retain their original values. Relative and protocol-relative alternate URLs are findings; any displayed resolution is explanatory and is excluded from reciprocity checks. Comparison preserves URL path case, trailing slashes, query order, encoded values, and fragments. Standard URL parsing normalizes scheme/host case, default ports, and dot segments. No request is made to any URL.

### Checks and interpretation

The audit checks self references, conflicting destinations for the same language, repeated annotations, known missing returns, supplied non-200 pages and targets, supplied noindex, and canonical alignment warnings. Non-self canonical values require review: a different canonical is not automatically an invalid SEO configuration. Content equivalence and canonical language are not assessed.

Language values are compared without case sensitivity. The supported form is an assigned ISO 639-1 language, optional registered ISO 15924 script, and optional assigned ISO 3166-1 alpha-2 region, plus `x-default`. Examples include `en-GB`, `zh-Hans`, and `zh-Hans-US`; `en-gb` is equivalent to `en-GB`. Numeric regions such as `es-419`, three-letter languages, extensions, and reserved region values such as `en-UK` are unsupported. `be` and `uk` are valid language codes for Belarusian and Ukrainian. Cross-domain alternates are allowed.

**Cluster membership means weak graph connectivity**, including referenced but unsupplied absolute targets. It does not prove that pages are equivalent translations. The audit checks self and return links; it does not require identical all-to-all language sets across a connected cluster. Reciprocal subsets can have no findings. Supply `policy.expectedLanguages` if every page must contain a specific complete language set. Missing `x-default` produces no issue unless explicitly required by that policy or `requireXDefault`.

A missing return is proven only when the target page and its complete alternate list are supplied. Partial lists and absent targets produce unknown findings. “No findings in supplied graph” describes the tested annotations; it does not certify unobserved metadata or live pages.

### Output and regression comparison

One dataset item contains the complete report. Download the dataset as JSON; use the dedicated flat CSV and standalone HTML files for human review.

| Field | Contents |
| --- | --- |
| `summary` | Page, edge, cluster, severity, coverage, and baseline counts |
| `rows` | Stable issue ID, severity, reason, exact URLs, language, and source references |
| `pageResults` | Page ID, observed fields, alternate-list coverage, and assessment |
| `edges` | Exact supplied annotations, normalized language, target coverage, return-link outcome |
| `clusters` | Weakly connected URLs, languages, supplied and unsupplied counts |
| `changes` | New, unchanged, resolved, or unobserved baseline findings |

`OUTPUT` stores the complete JSON; `issues.csv` is the flat issue queue; `clusters.csv` is the flat annotated edge matrix. Isolated pages and cluster totals are retained in JSON and HTML. `report.html` is self-contained, contains no scripts, and is downloaded as an attachment. CSV cells are protected against spreadsheet formula execution. IDs and source evidence are stable for repeat findings, although cluster IDs change when member URLs change.

Baseline resolutions require relevant pages and fields to be covered in both reports. Removed pages, incomplete lists, omitted statuses, unsupported canonical observations, and changed requirements are conservatively marked **unobserved** when they prevent comparison. A missing finding is not sufficient proof of a fix. Keep the baseline JSON unchanged; integrity and version validation prevent accidental reuse of edited or incompatible results. The input size limit may require retaining smaller cluster-specific baselines.

### Price and recovery

The intended Store price is **$0.25 per completed report** using the `report-completed` event, including platform usage. There is no startup or per-issue fee. Invalid input, insufficient budget, and reports rejected by the export limits are not charged a report event. The runtime verifies the actual positive configured report price and rejects any other positive-priced event.

The dataset is the primary deliverable. It is persisted before the report charge. Convenience exports are written afterward, so an interrupted run can have a complete charged dataset with missing files. Resurrecting that same run validates and regenerates its exports while preserving one dataset item and one charge; an interrupted charge can be completed once. Starting a new run is a new billable report. If charged data is missing or the saved input/result was changed, recovery stops for inspection. Local SDK simulation does not prove cloud billing or cloud UI behavior.

### Questions and support

Provide page observations you are authorized to process; no credentials, network access, proxies, or external model service are required. This Actor does not modify your site. Share a small synthetic reproduction through the Issues tab for unexpected results, and use the API tab for programmatic integration. There are no translation, indexing, accuracy, or ranking guarantees. Language/script/region snapshots are versioned with the engine and may require a future update when standards change.

# Actor input Schema

## `pages` (type: `array`):

Required array of {id?,url,status?,canonical?,indexability?,alternatesComplete?,alternates:\[{language,url,source?}]}. URL must be absolute HTTP(S). Alternates are required; set alternatesComplete=false for partial exports. Status is an integer 100–599; missing status stays unknown. IDs must be strings. See README.

## `policy` (type: `object`):

requireXDefault defaults false. expectedLanguages is a list applied to every supplied page, e.g. \["en-GB","de","zh-Hans"]. Missing required languages are proved only on complete alternate lists.

## `previousReport` (type: `object`):

Optional unchanged OUTPUT object from this engine version. Do not paste the dataset array. Resolutions require comparable covered fields; total input including baseline must fit 4 MB.

## Actor input object example

```json
{
  "pages": [
    {
      "id": "001",
      "url": "https://example.com/en",
      "status": 200,
      "canonical": "https://example.com/en",
      "alternates": [
        {
          "language": "en-GB",
          "url": "https://example.com/en",
          "source": "html:head/link[1]"
        },
        {
          "language": "de",
          "url": "https://example.de/de",
          "source": "html:head/link[2]"
        },
        {
          "language": "zh-Hans",
          "url": "https://example.com/zh",
          "source": "html:head/link[3]"
        }
      ]
    },
    {
      "id": "002",
      "url": "https://example.de/de",
      "status": 200,
      "canonical": "https://example.de/de",
      "alternates": [
        {
          "language": "EN-GB",
          "url": "https://example.com/en",
          "source": "html:head/link[1]"
        },
        {
          "language": "DE",
          "url": "https://example.de/de",
          "source": "html:head/link[2]"
        },
        {
          "language": "ZH-HANS",
          "url": "https://example.com/zh",
          "source": "html:head/link[3]"
        }
      ]
    },
    {
      "id": "003",
      "url": "https://example.com/zh",
      "status": 200,
      "canonical": "https://example.com/zh",
      "alternates": [
        {
          "language": "en-GB",
          "url": "https://example.com/en",
          "source": "html:head/link[1]"
        },
        {
          "language": "de",
          "url": "https://example.de/de",
          "source": "html:head/link[2]"
        },
        {
          "language": "zh-Hans",
          "url": "https://example.com/zh",
          "source": "html:head/link[3]"
        }
      ]
    },
    {
      "id": "004",
      "url": "https://shop.example.com/en/item",
      "status": 200,
      "alternates": [
        {
          "language": "en",
          "url": "https://shop.example.com/en/item"
        },
        {
          "language": "fr",
          "url": "https://shop.example.fr/fr/item"
        }
      ]
    },
    {
      "id": "005",
      "url": "https://shop.example.fr/fr/item",
      "status": 200,
      "alternates": [
        {
          "language": "fr",
          "url": "https://shop.example.fr/fr/item"
        }
      ]
    },
    {
      "id": "006",
      "url": "https://example.org/catalog",
      "alternates": [
        {
          "language": "en",
          "url": "https://example.org/catalog"
        },
        {
          "language": "de",
          "url": "https://example.org/de/catalog"
        },
        {
          "language": "zh-Hans",
          "url": "/zh/catalog"
        }
      ]
    }
  ],
  "policy": {
    "requireXDefault": false,
    "expectedLanguages": []
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

One dataset item containing the full report and rows.

## `json` (type: `string`):

Use this unchanged JSON as previousReport.

## `issues` (type: `string`):

No description

## `clusters` (type: `string`):

No description

## `html` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pages": [
        {
            "id": "001",
            "url": "https://example.com/en",
            "status": 200,
            "canonical": "https://example.com/en",
            "alternates": [
                {
                    "language": "en-GB",
                    "url": "https://example.com/en",
                    "source": "html:head/link[1]"
                },
                {
                    "language": "de",
                    "url": "https://example.de/de",
                    "source": "html:head/link[2]"
                },
                {
                    "language": "zh-Hans",
                    "url": "https://example.com/zh",
                    "source": "html:head/link[3]"
                }
            ]
        },
        {
            "id": "002",
            "url": "https://example.de/de",
            "status": 200,
            "canonical": "https://example.de/de",
            "alternates": [
                {
                    "language": "EN-GB",
                    "url": "https://example.com/en",
                    "source": "html:head/link[1]"
                },
                {
                    "language": "DE",
                    "url": "https://example.de/de",
                    "source": "html:head/link[2]"
                },
                {
                    "language": "ZH-HANS",
                    "url": "https://example.com/zh",
                    "source": "html:head/link[3]"
                }
            ]
        },
        {
            "id": "003",
            "url": "https://example.com/zh",
            "status": 200,
            "canonical": "https://example.com/zh",
            "alternates": [
                {
                    "language": "en-GB",
                    "url": "https://example.com/en",
                    "source": "html:head/link[1]"
                },
                {
                    "language": "de",
                    "url": "https://example.de/de",
                    "source": "html:head/link[2]"
                },
                {
                    "language": "zh-Hans",
                    "url": "https://example.com/zh",
                    "source": "html:head/link[3]"
                }
            ]
        },
        {
            "id": "004",
            "url": "https://shop.example.com/en/item",
            "status": 200,
            "alternates": [
                {
                    "language": "en",
                    "url": "https://shop.example.com/en/item"
                },
                {
                    "language": "fr",
                    "url": "https://shop.example.fr/fr/item"
                }
            ]
        },
        {
            "id": "005",
            "url": "https://shop.example.fr/fr/item",
            "status": 200,
            "alternates": [
                {
                    "language": "fr",
                    "url": "https://shop.example.fr/fr/item"
                }
            ]
        },
        {
            "id": "006",
            "url": "https://example.org/catalog",
            "alternates": [
                {
                    "language": "en",
                    "url": "https://example.org/catalog"
                },
                {
                    "language": "de",
                    "url": "https://example.org/de/catalog"
                },
                {
                    "language": "zh-Hans",
                    "url": "/zh/catalog"
                }
            ]
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("h_murdock/hreflang-cluster-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "pages": [
        {
            "id": "001",
            "url": "https://example.com/en",
            "status": 200,
            "canonical": "https://example.com/en",
            "alternates": [
                {
                    "language": "en-GB",
                    "url": "https://example.com/en",
                    "source": "html:head/link[1]",
                },
                {
                    "language": "de",
                    "url": "https://example.de/de",
                    "source": "html:head/link[2]",
                },
                {
                    "language": "zh-Hans",
                    "url": "https://example.com/zh",
                    "source": "html:head/link[3]",
                },
            ],
        },
        {
            "id": "002",
            "url": "https://example.de/de",
            "status": 200,
            "canonical": "https://example.de/de",
            "alternates": [
                {
                    "language": "EN-GB",
                    "url": "https://example.com/en",
                    "source": "html:head/link[1]",
                },
                {
                    "language": "DE",
                    "url": "https://example.de/de",
                    "source": "html:head/link[2]",
                },
                {
                    "language": "ZH-HANS",
                    "url": "https://example.com/zh",
                    "source": "html:head/link[3]",
                },
            ],
        },
        {
            "id": "003",
            "url": "https://example.com/zh",
            "status": 200,
            "canonical": "https://example.com/zh",
            "alternates": [
                {
                    "language": "en-GB",
                    "url": "https://example.com/en",
                    "source": "html:head/link[1]",
                },
                {
                    "language": "de",
                    "url": "https://example.de/de",
                    "source": "html:head/link[2]",
                },
                {
                    "language": "zh-Hans",
                    "url": "https://example.com/zh",
                    "source": "html:head/link[3]",
                },
            ],
        },
        {
            "id": "004",
            "url": "https://shop.example.com/en/item",
            "status": 200,
            "alternates": [
                {
                    "language": "en",
                    "url": "https://shop.example.com/en/item",
                },
                {
                    "language": "fr",
                    "url": "https://shop.example.fr/fr/item",
                },
            ],
        },
        {
            "id": "005",
            "url": "https://shop.example.fr/fr/item",
            "status": 200,
            "alternates": [{
                    "language": "fr",
                    "url": "https://shop.example.fr/fr/item",
                }],
        },
        {
            "id": "006",
            "url": "https://example.org/catalog",
            "alternates": [
                {
                    "language": "en",
                    "url": "https://example.org/catalog",
                },
                {
                    "language": "de",
                    "url": "https://example.org/de/catalog",
                },
                {
                    "language": "zh-Hans",
                    "url": "/zh/catalog",
                },
            ],
        },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("h_murdock/hreflang-cluster-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pages": [
    {
      "id": "001",
      "url": "https://example.com/en",
      "status": 200,
      "canonical": "https://example.com/en",
      "alternates": [
        {
          "language": "en-GB",
          "url": "https://example.com/en",
          "source": "html:head/link[1]"
        },
        {
          "language": "de",
          "url": "https://example.de/de",
          "source": "html:head/link[2]"
        },
        {
          "language": "zh-Hans",
          "url": "https://example.com/zh",
          "source": "html:head/link[3]"
        }
      ]
    },
    {
      "id": "002",
      "url": "https://example.de/de",
      "status": 200,
      "canonical": "https://example.de/de",
      "alternates": [
        {
          "language": "EN-GB",
          "url": "https://example.com/en",
          "source": "html:head/link[1]"
        },
        {
          "language": "DE",
          "url": "https://example.de/de",
          "source": "html:head/link[2]"
        },
        {
          "language": "ZH-HANS",
          "url": "https://example.com/zh",
          "source": "html:head/link[3]"
        }
      ]
    },
    {
      "id": "003",
      "url": "https://example.com/zh",
      "status": 200,
      "canonical": "https://example.com/zh",
      "alternates": [
        {
          "language": "en-GB",
          "url": "https://example.com/en",
          "source": "html:head/link[1]"
        },
        {
          "language": "de",
          "url": "https://example.de/de",
          "source": "html:head/link[2]"
        },
        {
          "language": "zh-Hans",
          "url": "https://example.com/zh",
          "source": "html:head/link[3]"
        }
      ]
    },
    {
      "id": "004",
      "url": "https://shop.example.com/en/item",
      "status": 200,
      "alternates": [
        {
          "language": "en",
          "url": "https://shop.example.com/en/item"
        },
        {
          "language": "fr",
          "url": "https://shop.example.fr/fr/item"
        }
      ]
    },
    {
      "id": "005",
      "url": "https://shop.example.fr/fr/item",
      "status": 200,
      "alternates": [
        {
          "language": "fr",
          "url": "https://shop.example.fr/fr/item"
        }
      ]
    },
    {
      "id": "006",
      "url": "https://example.org/catalog",
      "alternates": [
        {
          "language": "en",
          "url": "https://example.org/catalog"
        },
        {
          "language": "de",
          "url": "https://example.org/de/catalog"
        },
        {
          "language": "zh-Hans",
          "url": "/zh/catalog"
        }
      ]
    }
  ]
}' |
apify call h_murdock/hreflang-cluster-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,h_murdock/hreflang-cluster-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/jxjyARLv7ofL6JTgg/builds/XdvonqnVa5Kbtw7aR/openapi.json
