# Website Migration QA — Redirects, Canonicals & Content (`elfajad/website-migration-qa`) Actor

Audit up to 100 old/new URL pairs: redirect targets and traces, missing destinations, HTML canonicals, indexing directives and content differences. Prioritized checklist, printable HTML and JSON. Free fictional demo. No AI key.

- **URL**: https://apify.com/elfajad/website-migration-qa.md
- **Developed by:** [El Fajad](https://apify.com/elfajad) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $15.00 / migration qa report

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Migration QA

Turn an **old → new URL map** into a prioritized repair and review checklist. Check up to **100 URL pairs** for wrong redirect targets, missing destinations, redirect chains and loops, title/heading differences, HTML canonical conflicts and indexing directives. Get **printable HTML, JSON, redirect traces and reusable page snapshots**.

Built for agencies and site teams checking a launch. This Actor checks the URLs you supply; it does not discover or certify an entire website. No AI API key, external paid scraper, browser automation or VPS is required.

### Try the fictional demo

```json
{"mode":"demo"}
```

Demo uses fixed synthetic responses for four example mappings: a correct redirect, a wrong destination, missing pages and changed static text. It **does not fetch real websites** or trigger the `migration-report` event. The startup event in Pricing still applies. Custom URLs are ignored in demo mode.

Open **Migration brief → Open printable report**. Evidence sections open by default. Print the page to save a PDF, or use the JSON for a team checklist.

### Audit your URL map

```json
{
  "mode": "audit",
  "phase": "postlaunch",
  "reportLabel": "Client website launch QA",
  "urlPairs": [
    {
      "oldUrl": "https://old.example.com/about",
      "newUrl": "https://new.example.com/about",
      "label": "About page"
    },
    {
      "oldUrl": "https://old.example.com/services",
      "newUrl": "https://new.example.com/services",
      "label": "Services page",
      "baseline": {
        "title": "Earlier services title",
        "h1": "Earlier services heading",
        "noindex": false,
        "observedAt": "2026-09-28T12:00:00Z"
      }
    }
  ]
}
```

The domains above are placeholders. Replace them with the public pages you intend to check.

1. Export your old/new page mapping from the migration plan. Paste **1–100 pairs** into **Old → new URL mapping**.
2. Choose **Audit my URL map**. Choose **After launch** when old URLs should redirect, or **Before launch** to compare pages without a missing-redirect finding.
3. Add pre-migration baseline fields where available. If omitted, the Actor may compare two distinct live pages, explicitly labeled as a current comparison.
4. Set **maximum charge to at least $15.01**, then run. Account spending limits must allow that budget. The saved small default is intended for the demo.
5. Review the checklist, pair traces and unknowns. Download output before your account's retention period expires.

In postlaunch mode, an old URL identical to its new URL does not require a redirect. Each old URL may have only one destination. Exact duplicate pairs are merged; conflicting destinations or baselines are rejected.

### What is checked

| Check | Evidence and interpretation |
|---|---|
| Old URL mapping | Its terminal redirect URL compared with the mapped new URL's terminal URL |
| Redirect behavior | Server response codes, Location targets, multiple hops, loops and missing Location |
| Missing destinations | Terminal HTTP 404/410; intentionally retired pages still need owner review |
| Server/access failures | Server errors are review items; 403/429 and network failures remain unknown |
| Destination HTML | Title, description, H1, HTML canonical targets, robots/googlebot metadata and X-Robots-Tag |
| Canonical differences | Conflicting targets or targets different from the fetched final page; intentional canonicalization may be valid |
| Baseline fields | Only supplied fields are compared; absent baseline fields remain unknown |
| Static text changes | Bounded word-retention and text-length heuristic, for sufficiently large usable text |

Canonical checks cover **HTML link declarations**, not HTTP Link headers or a search engine's selected canonical. A noindex directive may be intentional. Changed titles/headings are review items, not automatic regressions.

**An HTTP 200 alone does not prove the mapping is correct.** The old URL must reach the intended destination. Conversely, robots restrictions, TLS/DNS errors, bot challenges, unsupported formats and thin JavaScript shells are not called missing pages.

### Before/after content needs a baseline

After launch, the old URL often redirects to the same page as the new URL. The Actor **does not infer removed content by comparing that page with itself**. Supply an earlier snapshot when you need historical field comparisons.

Optional `baseline` fields for each pair:

| Field | Limit |
|---|---|
| `title`, `description`, `h1` | Strings, up to 1,000 characters each; an explicitly empty string is a known empty value |
| `text` | Plain page text, up to 30,000 characters |
| `noindex` | Explicit boolean |
| `canonicalUrl` | Public HTTP(S) URL or null |
| `observedAt` | Non-future ISO timestamp with timezone, e.g. `2026-09-28T12:00:00Z` |

Baseline contents, provenance and dates are supplied by you and are not independently verified. Ordinary live old/new comparisons are labeled **live-old-page**, not historical evidence.

Text comparison uses distinct Unicode words of at least two characters. A reduction review appears only when fewer than 50% of comparison words remain **and** destination text length is under 60% of comparison text length. This is not semantic analysis or proof of accidental deletion. Navigation, headers, footers and scripts are removed; main/article text is preferred. JavaScript content is not rendered.

**Reusable snapshots:** download `NEW-SNAPSHOTS`. Each usable destination includes a `baseline` object you can place directly into a later pair. Match the correct page yourself. If extracted text exceeds 30,000 characters, the snapshot labels `textOmitted: true` and omits text rather than presenting a truncated full-content baseline.

For a pre-migration snapshot, use prelaunch mode with each current URL supplied as both `oldUrl` and `newUrl`. This is a normal billable audit, including when no findings are observed.

### Scope and resource limits

- 1–100 URL pairs, maximum 8 MiB input.
- Only public HTTP(S) pages on standard ports. Credential/signed links, private/reserved DNS/IP targets and account/action URLs are rejected.
- HTTP GET only; no forms, purchases, login, JavaScript, discovered-link crawl or browser sessions.
- Each connection pins the validated DNS answer; redirects are validated again.
- `robots.txt` is respected for the Actor's named user agent. An unavailable/ambiguous robots policy stops that origin's checks.
- Up to five followed redirects per URL, 600 total HTTP requests, approximately ten minutes of collection. Unchecked pairs remain explicit.
- Sequential requests, at least 400 ms between requests to the same origin; robots crawl delays are respected within the bounded budget.
- Requests have an 8-second DNS/response time budget; responses are limited to 2 MiB and UTF-8/ASCII static HTML. Unsupported encodings/compressed responses remain unknown.
- No ranking, traffic, backlinks, sitemap, Search Console, actual search-indexing or complete-migration certification.

Reports may be **partial** when one side is available but another is blocked, or the collection limit is reached. The report lists attempted and unknown pairs. This does not imply every supplied URL was successfully checked.

### Pricing

Launch price: **$15 per delivered migration report**, covering up to 100 supplied URL pairs. The **Pricing tab is authoritative**.

- `migration-report`: once per usable non-demo report, including partial reports and reports with no findings.
- Fixed fictional demo: no report event.
- Wholly unknown/unusable collection: no report event; failure details are saved, and the run fails clearly.
- Invalid input or insufficient report budget: no report event.
- Startup event: **$0.00005** at supported memory sizes, including demos and failures.
- No separate dataset-row charge or customer platform-usage pass-through.

At least one pair with usable static HTML, an observed missing terminal, a confirmed redirect loop or a missing redirect Location is needed for a paid report. A successful fetch of the old page alone can support a partial billable report even if the new page is unavailable; review the coverage before using it.

The Actor checks the report-event budget before live requests, saves HTML before billing its report row, and skips already-delivered results if the same run is resurrected. A cap is not a minimum charge. Free-plan or remaining-credit restrictions may prevent a $15 real report; the fictional demo remains available with a small budget.

### Outputs

The default dataset contains one report: label, phase, `isDemo`, timestamps, observation status, counts, grouped `findings`, detailed `pairs`, `coverage`, and `reportUrl`.

Default key-value-store records:

- `migration-<analysisId>.html`: printable report.
- `PAIR-RESULTS`: full pair observations and unknowns, also saved when all pairs are unusable.
- `NEW-SNAPSHOTS`: reusable destination baseline objects.
- `RUN-SUMMARY`: delivery, usable/unknown counts and demo/budget state.

### API and MCP

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({token: process.env.APIFY_TOKEN});
const run = await client.actor('elfajad/website-migration-qa').call({
  mode: 'audit',
  phase: 'postlaunch',
  urlPairs: [{oldUrl: 'https://old.example.com/about', newUrl: 'https://new.example.com/about'}]
}, {memory: 256, timeout: 900, maxTotalChargeUsd: 15.01});
const {items} = await client.dataset(run.defaultDatasetId).listItems();
```

Replace placeholder URLs with your actual mapping. The published Actor is also accessible through Apify MCP. No third-party API key is required.

### Development and support

Node.js 22+, [Apify SDK](https://github.com/apify/apify-sdk-js) (Apache-2.0), [Cheerio](https://github.com/cheeriojs/cheerio) (MIT), [ipaddr.js](https://github.com/whitequark/ipaddr.js) (MIT) and [robots-parser](https://github.com/samclarke/robots-parser) (MIT). Deterministic checks, no proprietary model.

```sh
npm ci
npm test
npm run sample
```

Use the Issues tab for problems, with the run ID and non-sensitive URLs/field names. Do not publish credentials, private baseline text or payout information. For migration planning context, see [Google's site-move guidance](https://developers.google.com/search/docs/crawling-indexing/site-move-with-url-changes) and [Screaming Frog's redirect-audit workflow](https://www.screamingfrog.co.uk/seo-spider/tutorials/audit-redirects/).

# Actor input Schema

## `mode` (type: `string`):

Demo ignores custom URLs and uses only fixed fictional fixtures. Audit performs public HTTP checks and costs $15 per usable report.

## `urlPairs` (type: `array`):

Audit mode requires 1–100 objects: {"oldUrl":"https://old.example.com/page","newUrl":"https://new.example.com/page"}. Optional label and baseline fields: title, description, h1, text (max 30,000 characters), noindex (boolean), canonicalUrl, observedAt (ISO timestamp). One old URL cannot map to different destinations.

## `phase` (type: `string`):

Postlaunch expects changed old URLs to redirect. Prelaunch skips the missing-redirect rule. Other HTTP and destination HTML checks still run.

## `reportLabel` (type: `string`):

Heading for your printable client/team report. Up to 120 characters.

## Actor input object example

```json
{
  "mode": "demo",
  "phase": "postlaunch"
}
```

# Actor output Schema

## `report` (type: `string`):

One report with a prioritized checklist, pair evidence and printable HTML link.

## `pairs` (type: `string`):

Detailed URL-pair observations, including failures when no paid report is delivered.

## `snapshots` (type: `string`):

Each usable destination has a baseline object for use in a later audit. Oversized text is omitted and labeled.

## `summary` (type: `string`):

Delivered, usable, partial and unknown counts; demo and budget status.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "demo"
};

// Run the Actor and wait for it to finish
const run = await client.actor("elfajad/website-migration-qa").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "mode": "demo" }

# Run the Actor and wait for it to finish
run = client.actor("elfajad/website-migration-qa").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "demo"
}' |
apify call elfajad/website-migration-qa --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,elfajad/website-migration-qa"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/b5rZ4nQcRu0CueYgy/builds/6sUrMYMFAfJNJx836/openapi.json
