# URL Status Checker: Broken Links & Redirects (`axiorasolutions/url-status-checker`) Actor

Bulk-check any list of URLs. Returns HTTP status, the full redirect chain with each hop, final URL, cross-domain redirect detection, response time, content type, page title and meta robots. Built for broken-link reports, migration QA and link inventory checks.

- **URL**: https://apify.com/axiorasolutions/url-status-checker.md
- **Developed by:** [Axiora Solutions](https://apify.com/axiorasolutions) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.83 / 1,000 url checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## URL Status Checker — bulk broken link checker and redirect audit

**URL status checker** and broken link checker for bulk link audits: paste thousands of URLs and get each one's HTTP status, redirect chain, response time and page facts in a single dataset. No API key and no configuration — run it on Apify and export to JSON, CSV, Excel or Google Sheets. Built for broken-link reports, migration QA and link-inventory checks.

### What you get

- `httpStatus` and `statusCategory` — the raw status code plus `2xx-success`, `3xx-redirect`, `4xx-client-error` or `5xx-server-error` for one-click pivots.
- `redirectChain` — every hop with its status and destination, plus `isCrossDomainRedirect` when the trail changes domain.
- `isBroken` — `true` for 404 and 410, so a single filter gives you the dead-link report.
- `responseMs` — per-URL timing, including redirect hops, so slow endpoints surface alongside broken ones.
- `title`, `h1`, `canonicalUrl` and `metaRobots` — page facts parsed from the same response, with `canonicalIsSelf` and `isNoindex` derived from them.
- `errorCode` on unreachable hosts — DNS failures and timeouts become data rows instead of crashing the run.

### Quick start

1. Open the Actor on Apify and paste your URLs into **URLs to check** — bare domains are upgraded to `https://` automatically.
2. Keep the defaults: `GET`, follow redirects, collect page facts. Nothing else is required.
3. Click **Start**. The run stops cleanly at **Max URLs for the whole run** and reports filtered URLs as not billed.
4. Read the **URL checks**, **Broken links** and **Redirects** dataset tabs, or export or schedule the run.

Minimal input — the only required field is `urls`:

```json
{
  "urls": [
    "https://apify.com",
    "https://apify.com/this-page-does-not-exist",
    "http://github.com"
  ]
}
```

### Example output

One dataset row per URL. Here is a cross-domain redirect — the migration case:

```json
{
  "ok": true,
  "errorCode": null,
  "requestedUrl": "https://oldbrand.com/pricing",
  "method": "GET",
  "httpStatus": 200,
  "statusCategory": "2xx-success",
  "isOk": true,
  "isRedirect": false,
  "isError": false,
  "isBroken": false,
  "finalUrl": "https://newbrand.com/pricing",
  "redirectCount": 2,
  "redirectChain": [
    { "status": 301, "to": "https://www.oldbrand.com/pricing" },
    { "status": 301, "to": "https://newbrand.com/pricing" }
  ],
  "isCrossDomainRedirect": true,
  "finalDomain": "newbrand.com",
  "responseMs": 612,
  "contentType": "text/html",
  "contentBytes": 14234,
  "title": "Pricing",
  "h1": "Simple pricing",
  "metaRobots": "index, follow",
  "canonicalUrl": "https://newbrand.com/pricing",
  "canonicalIsSelf": true,
  "isNoindex": false,
  "scrapedAt": "2026-10-02T12:00:00.000Z"
}
```

### What this URL checker returns

**Check thousands of URLs in one run.** Every URL gets its HTTP status, the **full redirect chain with each hop**, the final URL, cross-domain redirect detection, response time, content type and — for HTML pages — title, h1, canonical and meta robots.

- 🚦 **Every status, categorised** — `statusCategory` groups rows into `2xx-success`, `3xx-redirect`, `4xx-client-error` and `5xx-server-error`, so you can pivot a 20,000-row dataset in one click.
- 🔗 **The whole redirect chain, not just the endpoint** — `redirectChain` lists each hop with its status code and destination. More than two hops is a real finding: it wastes crawl budget and slows real users.
- 🌐 **Cross-domain redirect detection** — `isCrossDomainRedirect` and `finalDomain` flag the signature of a migration, a domain handover or an acquired brand being folded in. This is how you find the 301s you forgot about.
- 💀 **`isBroken` for the list that matters** — true for 404 and 410 specifically, the two statuses that mean content is genuinely gone. Filter on it and you have a dead-link report.
- ⏱️ **Response times** — `responseMs` for every URL, including the redirect hops, so slow endpoints surface alongside broken ones.
- 📄 **Page facts without extra requests** — title, first h1, canonical and meta robots are parsed from the same response. `canonicalIsSelf` and `isNoindex` tell you when a reachable page is nonetheless excluded from search.
- 🧱 **Handles tens of thousands of URLs** — bounded concurrency scaled to your allocated memory, per-host pacing so you do not hammer one server, and a graceful stop at your run-cost ceiling.
- 🛟 **Invalid URLs are data, not crashes** — a malformed entry or an unreachable host produces an error row with a stable code, and every other URL still gets checked.

Running on Apify adds scheduling, webhooks, monitoring, API and SDK access, and one-click export to JSON, CSV, Excel, Google Sheets and 20+ integrations.

### How to use it

1. Paste your URLs into **URLs to check**. Bare domains are upgraded to `https://` automatically.
2. Leave **Follow redirects** on to get the full chain, or turn it off to see the raw first response.
3. Leave **Collect page title and meta robots** on unless you only want statuses — it adds no requests.
4. To get a clean broken-link list, put `404` and `410` in **Report only these status codes**.
5. Click **Start**, then use the **URL checks**, **Broken links** and **Redirects** dataset tabs.

#### How do I check every link on my site?

Run the **Sitemap & Indexability Audit** Actor first and export the `internalLinks` column, or just the `url` column for every page. Paste that list here. One Actor maps your URLs, the other verifies them.

### Full input example

Every option, with the defaults shown:

```json
{
  "urls": [
    "https://example.com",
    "example.com/old-page",
    "http://example.com/redirect-me",
    "https://example.com/removed-product"
  ],
  "method": "GET",
  "followRedirects": true,
  "fetchPageFacts": true,
  "filterByStatus": [],
  "maxUrlsTotal": 10000,
  "requestTimeoutSecs": 20
}
```

### What an error row looks like

An unreachable host still produces a row — the rest of the run continues:

```json
{
  "ok": false,
  "errorCode": "NETWORK_ERROR",
  "requestedUrl": "https://this-host-does-not-resolve.example",
  "error": {
    "code": "NETWORK_ERROR",
    "message": "DNS lookup failed for this-host-does-not-resolve.example: ENOTFOUND",
    "httpStatus": null,
    "hint": "The host could not be reached. Check the hostname and DNS."
  }
}
```

### Use cases

- **Broken-link reports** — filter to `isBroken` and hand the list to whoever owns the content.
- **Site migration QA** — check every legacy URL and confirm each one lands on the right new page with no cross-domain surprise.
- **Redirect-chain cleanup** — find chains longer than two hops and collapse them to a single 301.
- **Link inventory and hygiene** — verify outbound links, affiliate parameters and partner URLs still resolve.
- **Uptime spot checks** — schedule the Actor over your critical URLs and alert on any non-2xx.
- **Noindex audits** — `isNoindex` finds pages that respond 200 but are silently excluded from search.
- **Vendor and citation checking** — validate a list of source URLs before publishing.

### How much does it cost

Pricing is **pay per event** with one event:

| Event | What triggers it | Billed |
|---|---|---|
| URL checked | One URL requested and a result row written | per URL |
| Actor start | Once per run, platform fee | per run |

URLs removed by **Report only these status codes** or **Report only URLs matching** are **not** billed, because they are never written. A URL that cannot be resolved at all (DNS failure, timeout, private address) is reported but is not billed as a check.

10,000 URLs is 10,000 check events — that is the whole calculation. Compute, bandwidth and storage are included; there is no separate platform-usage charge on top.

Set **Max cost per run** in the run options for a hard ceiling. Higher Apify plans get progressively lower per-URL pricing through Apify Store tier discounts.

**Zero-cost verification:** leave the prefilled rows in place and run once. You will see a live 200, a live 404 and a live redirect from your own first run, for a fraction of a cent.

### Frequently asked questions

#### GET or HEAD?

**GET** is the default and the honest choice: it is what a browser and a crawler actually do. **HEAD** is roughly twice as fast and downloads nothing, but a minority of servers return 405 or 403 to HEAD while answering GET perfectly well, so you may see false failures. Use HEAD only for very large sweeps where you do not need page facts.

#### Why is `contentBytes` not the real page weight?

Because the response is capped at 512 KB. That keeps a 20,000-URL run affordable and stops one enormous page from consuming your transfer budget. Treat the number as a lower bound and a rough signal, not a measurement.

#### How do I get only the broken links?

Put `404` and `410` in **Report only these status codes**. Those are the two statuses that mean content is gone. If you also want to catch soft failures, add `500` and `503` for server-side problems. Everything else is filtered out and not billed.

#### What counts as a cross-domain redirect?

The chain ending on a different registrable domain from the one you requested. `www.example.com` to `example.com` is same-domain and correctly reported as such; `oldbrand.com` to `newbrand.com` is cross-domain. Subdomain changes within one registrable domain do not trigger it.

#### How many URLs can one run handle?

Tens of thousands, bounded by **Max URLs for the whole run** and your run-cost ceiling. Concurrency scales with the memory you allocate, and requests to a single host are paced so you do not trip rate limits. For very large lists, split by domain across a few runs.

#### Should I use a proxy?

Usually not. A status check should report what an ordinary visitor receives. Enable datacenter proxy rotation only when one host rate-limits a large sweep, and remember that a proxy can change what the server returns — which is the opposite of what you want from an availability check.

#### Does it run JavaScript?

No. It reports the server's response. Because page facts are read from the raw HTML, a client-rendered page may show no title here even though a browser displays one. For indexability of rendered content you need a browser-based crawler; that is a different trade-off and a different Actor.

#### Something looks wrong — how do I report it?

Open the **Issues** tab on this Actor page with the URL and the status you expected. A wrong status is treated as a bug.

### Related Actors by Axiora Solutions

| Actor | Use it for |
|---|---|
| **Sitemap & Indexability Audit** | Produce the URL list from sitemaps, with full on-page SEO findings |
| **Domain Contact Enricher** | Contact and technology enrichment for the domains you are checking |
| **Shopify Product & Variant Scraper** | Track competitor catalogues and pricing |

***

Runnable examples and how-to guides for these Actors: [github.com/batow133/axiora-apify-actors](https://github.com/batow133/axiora-apify-actors)

# Actor input Schema

## `urls` (type: `array`):

One URL per entry. Bare domains are upgraded to https automatically. Paste up to tens of thousands of URLs; anything invalid is reported as an error row rather than failing the run.

## `method` (type: `string`):

GET downloads the body and is the only way to get page facts. HEAD is roughly twice as fast and much lighter, but some servers reject it with a 405 even when GET works.

## `followRedirects` (type: `boolean`):

Follow the full redirect chain and report the final URL. When off, a 301 or 302 is reported as the final answer with no chain walked.

## `fetchPageFacts` (type: `boolean`):

Record the title, first h1, canonical and meta robots for responses that are HTML. Adds no requests, only a little bandwidth. Turn off for the fastest possible status-only sweep.

## `filterByStatus` (type: `array`):

Use this to get a broken-link list and nothing else. 404 alone gives you dead pages; 404 and 410 together catches removed content. Leave empty to report every URL.

## `filterByUrlPattern` (type: `string`):

Keep only URLs matching this regular expression, for example /blog/ to check just one section.

## `maxUrlsTotal` (type: `integer`):

Hard ceiling on billed rows. The run stops cleanly when it is reached.

## `requestTimeoutSecs` (type: `integer`):

Give up on a single URL after this many seconds. Slow URLs are reported with a TIMEOUT error row, not a failed run.

## `proxyConfiguration` (type: `object`):

Optional. A status check should report what the server actually returns to ordinary visitors, so running without a proxy gives the truest answer. Enable datacenter proxy rotation only if a host rate-limits a very large sweep.

## Actor input object example

```json
{
  "urls": [
    "https://example.com",
    "example.com/old-page",
    "http://example.com/redirect-me"
  ],
  "method": "GET",
  "followRedirects": true,
  "fetchPageFacts": true,
  "filterByStatus": [
    "404",
    "410",
    "500"
  ],
  "filterByUrlPattern": "/blog/",
  "maxUrlsTotal": 10000,
  "requestTimeoutSecs": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `checks` (type: `string`):

One row per checked URL with status, redirects, timing and page facts.

## `runSummary` (type: `string`):

Status distribution across 2xx, 3xx, 4xx, 5xx and network errors, plus filter effects, billing and network totals.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://apify.com",
        "https://apify.com/this-page-does-not-exist",
        "http://github.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("axiorasolutions/url-status-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://apify.com",
        "https://apify.com/this-page-does-not-exist",
        "http://github.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("axiorasolutions/url-status-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://apify.com",
    "https://apify.com/this-page-does-not-exist",
    "http://github.com"
  ]
}' |
apify call axiorasolutions/url-status-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,axiorasolutions/url-status-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/T3CUrihJSKC9HYlXU/builds/rLiwhx8dkVlim3J5d/openapi.json
