# Bulk URL Status Checker (`palgenius/bulk-url-status-checker`) Actor

Check thousands of URLs for broken links: status code, full redirect chain, response time, content type and size. Chains straight off the Sitemap URL Extractor.

- **URL**: https://apify.com/palgenius/bulk-url-status-checker.md
- **Developed by:** [Ali Alsaidi](https://apify.com/palgenius) (community)
- **Categories:** Developer tools, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.50 / 1,000 url checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk URL Status Checker

Give it a list of URLs and get back what each one actually does — the status code, every hop
of the redirect chain, how long it took, the content type and the size.

### What does Bulk URL Status Checker do?

- **Finds broken links at scale.** Built to run over tens of thousands of URLs in one go
  without hammering anyone.
- **Reports the whole redirect chain**, hop by hop, with a loop guard. Most checkers give you
  a final status code; the chain is what tells you whether a migration kept its redirects,
  where a loop is, and which links cost users two extra round trips.
- **Reads URLs straight from another run.** Point it at a dataset ID — for example the output
  of the Sitemap URL Extractor — instead of pasting a list.
- **Falls back from HEAD to GET automatically** when a server refuses HEAD with 405, 501 or
  403, and tells you which method answered.
- **Polite by default.** Optional minimum gap per host, backoff on 429 and 5xx that honours
  `Retry-After`, and `robots.txt` respected.

### What data can you get?

| Field | What it means |
|---|---|
| `url` | The URL you asked for |
| `finalUrl` | Where it ended up after redirects |
| `status` | HTTP status of the final response |
| `ok` | `true` for a 2xx final status |
| `redirects` | Number of hops |
| `redirectChain` | Every hop: `url`, `status`, `location` |
| `responseTimeMs` | Whole check, including redirects and retries |
| `contentType` | `Content-Type` of the final response |
| `sizeBytes` | Size from `Content-Length`, or measured if you ask |
| `server` | `Server` header, when the site sends one |
| `method` | `HEAD` or `GET` — which one actually answered |
| `attempts` | How many tries it took |
| `error` | Error type, or `null` |

### How to use it

1. Create a free Apify account.
2. Open the Actor and click **Try for free**.
3. Paste your URLs into **URLs**, or leave that empty and put a **Dataset ID** from a previous
   run instead.
4. Set **Maximum URLs** to cap the run, then click **Start**.
5. Open the **Dataset** tab and download as JSON, CSV, Excel or XML, or pull it through the
   Apify API.

### Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `urls` | array | — | URLs to check. Leave empty if you use `datasetId` |
| `datasetId` | string | — | Read URLs from another run's dataset instead |
| `datasetField` | string | `url` | Which field of that dataset holds the URL |
| `maxUrls` | integer | 10000 | Stop after this many URLs. A hard ceiling on run cost |
| `followRedirects` | boolean | true | Follow each redirect and report the chain |
| `measureSize` | boolean | false | Download the body to measure size when the server sends no `Content-Length` |
| `concurrency` | integer | 10 | How many URLs to check at once |
| `minHostIntervalSecs` | integer | 0 | Smallest gap between two requests to the same host |
| `requestTimeoutSecs` | integer | 20 | Per-request timeout |
| `maxRetries` | integer | 2 | Retries with backoff on timeouts and 429/5xx |
| `maxRedirects` | integer | 10 | Hops before calling it a redirect loop |
| `respectRobotsTxt` | boolean | true | Skip URLs that `robots.txt` disallows |

```json
{
  "urls": ["https://apify.com", "https://apify.com/store"],
  "maxUrls": 1000,
  "followRedirects": true
}
```

To chain it onto a sitemap run instead:

```json
{ "urls": [], "datasetId": "PASTE_DATASET_ID", "datasetField": "url" }
```

### Output

One item per URL. This is a real item from a run, a URL that redirected twice:

```json
{
  "url": "https://httpbin.org/redirect/2",
  "finalUrl": "https://httpbin.org/get",
  "status": 200,
  "ok": true,
  "redirects": 2,
  "redirectChain": [
    { "url": "https://httpbin.org/redirect/2", "status": 302, "location": "/relative-redirect/1" },
    { "url": "https://httpbin.org/relative-redirect/1", "status": 302, "location": "/get" },
    { "url": "https://httpbin.org/get", "status": 200, "location": null }
  ],
  "responseTimeMs": 3754,
  "contentType": "application/json",
  "sizeBytes": 383,
  "server": "gunicorn/19.9.0",
  "method": "HEAD",
  "attempts": 1,
  "error": null
}
```

Every run also writes a `RUN_SUMMARY` record to the key-value store — `urlsChecked`, a
`byStatus` breakdown (`2xx`, `3xx`, `4xx`, `5xx`), `averageResponseMs`, `requests`, `retries`
and `blockedByRobotsTxt` — so you can see the shape of a big run without reading the dataset:

```json
{ "urlsChecked": 4, "byStatus": { "2xx": 3, "4xx": 1 },
  "averageResponseMs": 1708, "requests": 10, "retries": 0, "blockedByRobotsTxt": 0 }
```

### How much does it cost?

**$0.50 per 1,000 URLs checked** ($0.0005 each). **10,000 URLs = $5.** Apify platform usage —
compute, storage, transfer — is included in that price, not billed on top.

A URL skipped for `robots.txt` is **not charged**, and `maxUrls` is a hard ceiling on what a
run can cost.

### Integrations

Connect it to webhooks, Make, Zapier, Google Sheets, Slack, Airbyte or GitHub from the
**Integrations** tab — for example, run it weekly and get a Slack message when the 4xx count
goes up.

Call it from your own code with the Apify API:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("palgenius/bulk-url-status-checker").call(
    run_input={"urls": ["https://apify.com", "https://apify.com/store"], "maxUrls": 1000}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    if not item["ok"]:
        print(item["url"], item["status"], item["error"])
```

AI agents can call it through the **Apify MCP server** at `https://mcp.apify.com`, e.g.
`https://mcp.apify.com?tools=palgenius/bulk-url-status-checker`.

### Error items and limits

A URL that could not be checked still comes back as an item, with `ok: false` and the reason
in `error` — never a silent gap:

```json
{ "url": "https://example.com/private", "finalUrl": "https://example.com/private",
  "status": null, "ok": false, "redirects": 0, "redirectChain": [],
  "responseTimeMs": null, "error": "blocked by robots.txt" }
```

Typical `error` values: `blocked by robots.txt` (skipped, and not charged), a timeout, a DNS
failure, or a redirect loop past `maxRedirects`.

Honest limits:

- **A status code is not a promise.** Some sites return 200 with an error page, and some
  return 403 to any non-browser client, including this one. `server` and `contentType` help
  you spot both.
- **HEAD can lie.** A few servers answer HEAD differently from GET. That is why a refusal
  triggers a GET, and why `method` is reported.
- **Timing is indicative.** `responseTimeMs` covers the whole check from one cloud region.
  It is not a substitute for real user monitoring.
- **Size comes from the server** unless you set `measureSize`, and servers are not always
  truthful about `Content-Length`.

### FAQ

**Is this legal?** Yes. It requests public URLs the same way any browser or link checker
does, obeys `robots.txt` by default, never logs in, and collects no personal data.

**Can I check a whole website?** Run the Sitemap URL Extractor first, then paste its dataset
ID here.

**Will it overload a small server?** Set `minHostIntervalSecs` to put a gap between requests
to the same host, and lower `concurrency`.

**Why is `sizeBytes` null?** The server sent no `Content-Length`. Turn on `measureSize` to
download the body and measure it.

**Does it find redirect loops?** Yes — it stops at `maxRedirects` and reports the chain it
followed.

**Do I pay for URLs that are skipped?** No. Only URLs actually checked are charged.

### You might also like

- [Sitemap URL Extractor](https://apify.com/palgenius/sitemap-url-extractor) — get every page
  URL a site publishes, then feed its dataset ID straight into this Actor.
- [Tech Stack Detector](https://apify.com/palgenius/tech-stack-detector) — find the platform,
  CMS and frameworks behind any domain, with evidence.

### Feedback

Hit a site it reports wrongly, or want a field added? Open an issue on the Actor's **Issues**
tab and include the URL. Fixes ship within days.

# Actor input Schema

## `urls` (type: `array`):

URLs to check. You can paste them here, or leave this empty and use a dataset ID below.

## `datasetId` (type: `string`):

Read URLs from another run's dataset, for example the output of the Sitemap URL Extractor. Leave empty if you pasted URLs above.

## `datasetField` (type: `string`):

Which field of that dataset holds the URL.

## `maxUrls` (type: `integer`):

Stop after this many URLs. Each URL checked is one charged result; URLs skipped for robots.txt are not charged.

## `followRedirects` (type: `boolean`):

Follow each redirect and report the whole chain. Turn off to see only the first response.

## `measureSize` (type: `boolean`):

Download the body to measure size when the server sends no Content-Length. Slower and uses more traffic.

## `concurrency` (type: `integer`):

How many URLs to check at once across all hosts.

## `minHostIntervalSecs` (type: `integer`):

Politeness: the smallest gap between two requests to the same host.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for a response before giving up and retrying.

## `maxRetries` (type: `integer`):

Retries with backoff on timeouts and 429/5xx responses, honouring Retry-After.

## `maxRedirects` (type: `integer`):

How many hops to follow before calling it a redirect loop.

## `respectRobotsTxt` (type: `boolean`):

Skip URLs that robots.txt disallows. Those are reported but never charged.

## Actor input object example

```json
{
  "urls": [
    "https://apify.com",
    "https://apify.com/store"
  ],
  "datasetField": "url",
  "maxUrls": 10,
  "followRedirects": true,
  "measureSize": false,
  "concurrency": 10,
  "minHostIntervalSecs": 0,
  "requestTimeoutSecs": 20,
  "maxRetries": 2,
  "maxRedirects": 10,
  "respectRobotsTxt": true
}
```

# Actor output Schema

## `urlStatus` (type: `string`):

One item per URL, with the final status and URL, every redirect hop, response time, content type, size, server and any error.

## `runSummary` (type: `string`):

Counts for this run: URLs checked, a tally by status code, the average response time, requests, retries and robots.txt blocks.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://apify.com",
        "https://apify.com/store"
    ],
    "maxUrls": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("palgenius/bulk-url-status-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://apify.com",
        "https://apify.com/store",
    ],
    "maxUrls": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("palgenius/bulk-url-status-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://apify.com",
    "https://apify.com/store"
  ],
  "maxUrls": 10
}' |
apify call palgenius/bulk-url-status-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,palgenius/bulk-url-status-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/hX9BFg2yz4OXBUmzx/builds/dnSj4icczLmjyzajm/openapi.json
