# Bulk URL Status & Redirect Checker (`kernfetch/url-status-redirect-checker`) Actor

Check HTTP status codes and full redirect chains for thousands of URLs. Get the final URL, hops, loops, 302s and noindex headers, and verify expected redirects for site migrations. Works with Sitemap URL Extractor output. Respects robots.txt. $1 per 1,000 URLs.

- **URL**: https://apify.com/kernfetch/url-status-redirect-checker.md
- **Developed by:** [kernfetch](https://apify.com/kernfetch) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk URL Status & Redirect Checker – status codes and redirect chains in bulk

Check **thousands of URLs** in one run and get, for each one, the **HTTP status code**, the **full redirect chain** (every hop with its status), the **final URL** and a ready-made list of **issues**: redirect chains, loops, temporary (302/307) redirects, HTTPS→HTTP downgrades, cross-host redirects, `X-Robots-Tag: noindex`, 4xx and 5xx errors.

Built for **SEO audits, site migrations, broken link checks, sitemap QA and monitoring**.

### Why this Actor

- ⚡ **Fast and cheap** – headers only (HEAD, with automatic GET fallback): page bodies are never downloaded.
- 🔁 **Full redirect chains** – every hop with its status code, loop detection, configurable max hops, final URL and final status.
- 🚚 **Migration QA** – paste `old URL -> expected new URL` pairs and get `matchesExpected` for each one, plus `allRedirectsPermanent` (only 301/308 in the chain).
- 🧩 **Works with Sitemap URL Extractor** – pass the dataset ID of a [Sitemap URL Extractor](https://apify.com/kernfetch/sitemap-url-extractor) run and check every URL of a site.
- 🩺 **Issues ready to filter** – `redirect_chain`, `temporary_redirect`, `https_downgrade`, `cross_host_redirect`, `redirect_loop`, `noindex_header`, `final_4xx`, `final_5xx`, `expected_mismatch`.
- 📋 **Run report** – a `SUMMARY` record with counts by outcome, status code and issue, plus skipped URLs, blocked hosts and unreachable hosts.
- 💸 **Fair billing** – URLs that are not requested (robots.txt, host that refused access) are **not charged**.
- ✅ **Reliable and gentle** – robots.txt respected on every hop, low load per site, protections never bypassed.

### How to use

1. Paste your URLs in **URLs to check** (one per line; *Bulk edit* accepts thousands). URLs without `http(s)://` get `https://`.
2. Optionally add **Expected redirects** (`https://old.com/page -> https://new.com/page`) or a **Dataset ID** from another run.
3. Click **Start** and download the results as JSON, CSV, Excel or via API. Use the **Overview** and **Redirect chains** views.

#### Input example

```json
{
  "urls": ["https://apify.com", "http://example.com", "example.org/old-page"],
  "redirectMap": ["https://old.example.com/a -> https://www.example.com/a"],
  "method": "HEAD_THEN_GET",
  "maxRedirects": 10,
  "maxUrls": 1000
}
```

#### Output example

```json
{
  "url": "http://example.com/old",
  "outcome": "redirected_ok",
  "statusCode": 301,
  "finalStatusCode": 200,
  "finalUrl": "https://www.example.com/new",
  "redirectCount": 2,
  "redirectChain": [
    { "url": "http://example.com/old", "status": 301, "location": "https://example.com/old" },
    { "url": "https://example.com/old", "status": 302, "location": "https://www.example.com/new" },
    { "url": "https://www.example.com/new", "status": 200 }
  ],
  "allRedirectsPermanent": false,
  "issues": ["redirect_chain", "temporary_redirect"],
  "expectedUrl": "https://www.example.com/new",
  "matchesExpected": true,
  "contentType": "text/html; charset=UTF-8",
  "contentLength": 1256,
  "lastModified": null,
  "xRobotsTag": null,
  "errorType": null,
  "method": "HEAD",
  "responseTimeMs": 412,
  "checkedAt": "2026-09-29T10:00:00+00:00"
}
```

| Field | Description |
|---|---|
| `url` | URL you submitted (normalized) |
| `outcome` | `ok`, `redirected_ok`, `client_error`, `server_error`, `unreachable`, `redirect_loop`, `too_many_redirects`, `redirect_without_location`, `invalid_redirect`, `blocked`, `robots_disallowed` |
| `statusCode` / `finalStatusCode` | Status of the first response / of the last response in the chain |
| `finalUrl` | Where the URL finally lands |
| `redirectCount` / `redirectChain` | Number of redirects / every hop with status and `location` |
| `allRedirectsPermanent` | `true` if every redirect is 301 or 308 |
| `issues` | List of detected problems (see above) |
| `expectedUrl` / `matchesExpected` | Only with Expected redirects: `true` if the URL lands on the expected URL with a 2xx status |
| `contentType`, `contentLength`, `lastModified`, `xRobotsTag` | Headers of the final response |
| `errorType` | For `unreachable`: `dns`, `connect`, `ssl`, `timeout`, `protocol` |
| `method` | `HEAD` or `GET` (fallback) |
| `responseTimeMs` | Network time only (queueing and politeness delays excluded) |
| `checkedAt` | Check timestamp (UTC) |

### Use with Sitemap URL Extractor

1. Run [Sitemap URL Extractor](https://apify.com/kernfetch/sitemap-url-extractor) on a domain.
2. Copy the run's **Dataset ID** (Storage tab).
3. Paste it in **Dataset ID** here and start: every sitemap URL is checked. Sitemaps should list only URLs that return 200 without redirects: anything else is a finding.

### Use with AI agents

Give your agent a list of URLs and get back **structured, stable fields** it can reason on: status, final URL, redirect hops and issues. Useful to validate links before citing them, clean URL lists before scraping, or verify that pages still exist. Works with the Apify API, Apify MCP server, Make, Zapier, n8n and LangChain.

### Pricing

Pay only for results: **$1.00 per 1,000 checked URLs**. No subscription. A run of 10,000 URLs costs about $10. URLs skipped because of robots.txt or a host that refused access are **not charged**. Set **Maximum cost per run** in Run options to cap your spend: the Actor never checks more URLs than your limit allows.

### FAQ

**Why is a URL "blocked" when it works in my browser?**
The site answered 403, 429 or an anti-bot challenge to automated requests. The Actor respects that and never tries to bypass protections: the remaining URLs of that host are skipped, listed in the `SUMMARY` and not charged.

**Why is `finalStatusCode` empty?**
The URL was unreachable (see `errorType`), or a redirect pointed to a URL disallowed by robots.txt, which is not requested.

**Does it download the pages?**
No. It only reads response headers, so it is fast, cheap and gentle on websites, and no page content is collected.

**How fast is it?**
Lists spread over many hosts run in parallel. On a single host the Actor waits at least 500 ms between requests (about 2 URLs per second), so 5,000 URLs on one site take about 40 minutes.

**HEAD or GET?**
The default *HEAD then GET* is the lightest option and retries with GET when a server answers HEAD with an error. Choose *GET only* for servers that answer HEAD incorrectly.

**Can I run it on a schedule?**
Yes. Create a Task with your URLs and a daily schedule, then filter rows where `issues` is not empty or connect an alert via Integrations (Slack, email, webhook).

**Is it legal?**
The Actor only requests public URLs you provide, reads HTTP headers, follows robots.txt and stops when a site refuses access. Site owners can block it with `User-agent: kernfetch` in robots.txt. You are responsible for the URLs you submit.

### Related Actors by kernfetch

- [Sitemap URL Extractor](https://apify.com/kernfetch/sitemap-url-extractor) – every URL of a website from its sitemaps, with lastmod dates.
- [RSS Feed Finder & Reader](https://apify.com/kernfetch/rss-feed-finder) – discover the RSS, Atom and JSON feeds of any site and get the latest articles.

### Support

Found a URL that doesn't behave as expected? Open an issue on the **Issues** tab with the URL: fixes are usually shipped within days.

# Actor input Schema

## `urls` (type: `array`):

One URL per line (use Bulk edit to paste thousands). URLs without a scheme get https://. Duplicates are removed.

## `datasetId` (type: `string`):

Check the URLs stored in an Apify dataset, e.g. the output of Sitemap URL Extractor (kernfetch/sitemap-url-extractor). SUMMARY records are ignored.

## `datasetUrlField` (type: `string`):

Name of the field that contains the URL in the dataset items.

## `redirectMap` (type: `array`):

One pair per line: old URL and expected final URL, separated by space, tab, comma, semicolon or '->'. Each old URL is checked and the result says whether it lands on the expected URL with a 2xx status.

## `method` (type: `string`):

HEAD then GET: lightest for the target site, falls back to GET when HEAD returns an error status. GET only: most accurate for servers that answer HEAD badly. The response body is never downloaded in either mode.

## `maxRedirects` (type: `integer`):

Chains longer than this are reported as too\_many\_redirects.

## `maxUrls` (type: `integer`):

Safety cap on the number of URLs checked (and charged) in one run.

## `maxConcurrency` (type: `integer`):

Maximum number of requests running at the same time across all hosts. Each host is still limited by the per-host settings below.

## `perHostConcurrency` (type: `integer`):

Kept low on purpose to be polite to each site.

## `perHostDelayMs` (type: `integer`):

Minimum pause between two requests to the same host. Default 500 ms is about 2 URLs per second per host.

## `requestTimeoutSecs` (type: `integer`):

Maximum time to wait for a response. URLs that time out are reported as unreachable (errorType timeout).

## Actor input object example

```json
{
  "urls": [
    "https://apify.com",
    "http://example.com",
    "https://httpbin.org/redirect/3"
  ],
  "datasetUrlField": "url",
  "method": "HEAD_THEN_GET",
  "maxRedirects": 10,
  "maxUrls": 1000,
  "maxConcurrency": 10,
  "perHostConcurrency": 2,
  "perHostDelayMs": 500,
  "requestTimeoutSecs": 15
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `skipped` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://apify.com",
        "http://example.com",
        "https://httpbin.org/redirect/3"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("kernfetch/url-status-redirect-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://apify.com",
        "http://example.com",
        "https://httpbin.org/redirect/3",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("kernfetch/url-status-redirect-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://apify.com",
    "http://example.com",
    "https://httpbin.org/redirect/3"
  ]
}' |
apify call kernfetch/url-status-redirect-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,kernfetch/url-status-redirect-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gRzzXxUBzhk3OOSJe/builds/mVLHRYgbpqIwAz3OX/openapi.json
