# Bulk URL Status Checker — 404s & Redirects (`ingenious_quip_bxq/url-status-checker`) Actor

Bulk-check HTTP status with HEAD→GET fallback, final URL, timing, and error class — chain from Sitemap Actor output. Error class dns/timeout/ssl/http; accept URL list, sitemap datasetId, or DOC\_TO\_MARKDOWN\_INPUT.

- **URL**: https://apify.com/ingenious\_quip\_bxq/url-status-checker.md
- **Developed by:** [新世紀書僮](https://apify.com/ingenious_quip_bxq) (community)
- **Categories:** SEO tools, Developer tools, Open source
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.30 / 1,000 url checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk URL Status Checker — Broken Links & Redirects

**Check a list of URLs for HTTP status, redirects, content-type and failure class — then get a SUMMARY of what is broken.**
Paste URLs, point at another Actor's dataset, or load a `DOC_TO_MARKDOWN_INPUT`-style key-value record (for example from the Sitemap URL Extractor). Each URL is probed with HEAD, falling back to GET when the server blocks HEAD. Default memory: 256 MB. No browser, no AI keys.

### What you get

- 🔗 **Bulk status check** — status code, final URL after redirects, content-type, content-length, timing
- 🧭 **HEAD → GET fallback** — many CDNs reject HEAD; the Actor retries with a streamed GET (body discarded)
- 🏷️ **Error class** — `dns`, `timeout`, `ssl`, `http`, `network`, or `none` when the request succeeded
- 📊 **SUMMARY report** — counts by status class (`2xx` / `3xx` / `4xx` / `5xx` / `error`) and by error class
- 🧺 **Output filter** — optionally save only broken / 3xx / 4xx / 5xx / network errors to the dataset (SUMMARY still covers every check)
- 🔌 **Chain from Sitemap** — `urls` accepts `{"url":…}` objects; or pass a Sitemap run's `datasetId` / `DOC_TO_MARKDOWN_INPUT` KV record
- 🐢 **Polite** — concurrency + per-host caps, optional delay, retries with backoff for 429 / 5xx
- 💾 **Light** — HTTP only, 256 MB default

### Measured results

Local + private cloud benches (2026-09-30 Asia/Taipei). PPE locked from the 1000-URL settled cost — see `docs/PRICING.md`.

| Test | Result |
|---|---|
| Local smoke: example.com + httpbingo 200/404/redirect + invalid DNS host | **5/5 checked in 0.3 s**, peak **75 MB**; classes `2xx×3`, `4xx×1`, `error×1` (DNS) |
| Local filter=`broken` (example.com + 404 + DNS-fail) | checked 3, **saved 2** (404 + DNS); charged 3 |
| Cloud smoke `4cmiGOIr4daLNLJFa` (5 URLs, 256 MB, build 0.1.1) | **SUCCEEDED**; wall 4.5 s; peak **83 MB**; settled **$0.000245** |
| Cloud bench `cKB1gqsrunccLI7Kx` (100 public URLs, build **0.1.2**) | **SUCCEEDED**; wall 8.9 s; peak **86 MB**; settled **$0.000798** (~$0.000008 / URL) |
| Cloud bench `9knsf0XkoShrTGXnM` (1000 public URLs, build **0.1.2**) | **SUCCEEDED**; wall 83 s; peak **94 MB**; settled **$0.006448** (~$0.00000645 / URL); ok 938 / issues 62 |

### Use cases

- **Broken-link audit** after a sitemap crawl
- **URL hygiene before RAG / crawling** — drop 4xx/5xx and dead hosts
- **Redirect inventory** — see where URLs finally land
- **Batch health check** for a known URL list

### How to use

1. Add URLs in **URLs to check**, and/or a **Source dataset ID** / **key-value store** from another run.
2. Optional: set **Max URLs**, **Dataset filter**, concurrency and timeout.
3. Click **Start**. Results appear in the **Dataset**; `SUMMARY` and `OUTPUT` are in the **Key-value store**.

#### Input example

```json
{
  "urls": [
    { "url": "https://example.com/" },
    { "url": "https://httpbin.org/status/404" }
  ],
  "maxUrls": 100,
  "httpMethod": "head_then_get",
  "outputFilter": "all"
}
```

Chain from a Sitemap Actor dataset (same `url` field):

```json
{
  "datasetId": "<sitemap-run-default-dataset-id>",
  "maxUrls": 500,
  "outputFilter": "broken"
}
```

Or from `DOC_TO_MARKDOWN_INPUT`:

```json
{
  "keyValueStoreId": "<sitemap-run-default-kv-id>",
  "keyValueRecordKey": "DOC_TO_MARKDOWN_INPUT",
  "outputFilter": "all"
}
```

#### Output example (one dataset item)

```json
{
  "url": "https://httpbin.org/status/404",
  "finalUrl": "https://httpbin.org/status/404",
  "httpStatus": 404,
  "statusClass": "4xx",
  "ok": false,
  "contentType": "text/html; charset=utf-8",
  "contentLength": null,
  "durationMs": 180,
  "methodUsed": "HEAD",
  "redirectCount": 0,
  "errorClass": "http",
  "error": "HTTP 404",
  "attempts": 1,
  "host": "httpbin.org"
}
```

#### Key-value store records

| Key | Content |
|---|---|
| `SUMMARY` | Counts: `totalChecked`, `saved`, `ok`, `notOk`, `byStatusClass`, `byErrorClass`, filter, duration, peak memory |
| `OUTPUT` | Same summary (kept for consistency with sibling Actors) |

### Pricing

Pay per event:

| Event | Price |
|---|---|
| URL checked (primary) | **$0.0003** per URL (= $0.30 per 1,000) |
| Actor start (Apify synthetic) | $0.00005 per GB (platform default) |

**Worked example:** 1,000 URLs ≈ **$0.30** + one start event. Platform cost measured on a 1,000-URL public bench was about **$0.00645** (own run; not what you pay as revenue share).

Every URL that is actually checked is charged once — including 3xx/4xx/5xx and network errors, and including URLs filtered out of the dataset (`outputFilter`). Failed inputs that never become a checked row are not charged. Details: `docs/PRICING.md`.

### Chaining: Sitemap → URL status

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")

## 1) discover URLs
run = client.actor("ingenious_quip_bxq/sitemap-url-discovery").call(run_input={
    "startUrls": [{"url": "https://www.example.com"}],
    "maxUrls": 100,
})

## 2) check their HTTP status (broken only in the dataset)
run2 = client.actor("ingenious_quip_bxq/url-status-checker").call(run_input={
    "datasetId": run["defaultDatasetId"],
    "outputFilter": "broken",
    "maxUrls": 100,
})
for item in client.dataset(run2["defaultDatasetId"]).iterate_items():
    print(item["url"], item["httpStatus"], item["errorClass"])
```

### Known limits

- Does **not** render JavaScript or detect soft-404s (short error pages that return HTTP 200).
- Does **not** crawl HTML for in-page links; it only checks the URLs you give it (or load from another Actor).
- Some hosts rate-limit or block data-center IPs — use the proxy input if needed.
- `contentLength` is whatever the server sends in `Content-Length` (often missing on chunked responses).
- Soft-404 detection and HTML in-page link crawling are out of scope (see above).

### FAQ

**Why HEAD then GET?** Many servers (and WAFs) answer HEAD with 403/405. Falling back to GET reads headers only and discards the body.

**Are redirects “broken”?** No. With the default filter they are saved as `statusClass: "3xx"` (or as the final 2xx if redirects are followed). Use filter `redirects_and_errors` or `3xx` if you only want those rows.

**DNS / timeout / SSL?** Recorded with `statusClass: "error"` and `errorClass` set; the run continues.

### License & source code

This Actor is open source under the **GNU Affero General Public License v3.0 (AGPL-3.0)** — see `LICENSE`. The full source code is public: https://github.com/xbox002000/url-status-checker

See `CHANGELOG.md` for version history.

# Changelog

This Actor's version history is a separate document: https://apify.com/ingenious\_quip\_bxq/url-status-checker/changelog.md

# Actor input Schema

## `urls` (type: `array`):

One or more URLs. Accepts `{"url": "..."}` objects (same shape as Sitemap Actor dataset rows and DOC\_TO\_MARKDOWN\_INPUT). You can also paste plain URL strings or use a remote text list (`requestsFromUrl`).

## `datasetId` (type: `string`):

Optional. Apify dataset ID from another Actor run (e.g. Sitemap URL Extractor). Reads the `url` field from each item. Combined with `urls` if both are set.

## `keyValueStoreId` (type: `string`):

Optional. Apify key-value store ID. Used with Record key below to load a JSON list such as DOC\_TO\_MARKDOWN\_INPUT (`{"urls":[{"url":...}]}`) or a plain URL array.

## `keyValueRecordKey` (type: `string`):

Key inside the source key-value store (e.g. `DOC_TO_MARKDOWN_INPUT`). Ignored unless a store ID is set.

## `maxUrls` (type: `integer`):

Stop after checking this many unique URLs (0 = no limit). Also caps cost.

## `outputFilter` (type: `string`):

Which checked URLs to save in the dataset. SUMMARY always includes counts for every URL checked. `all` = every URL; `broken` = not OK (4xx/5xx/network error); `redirects_and_errors` = 3xx + broken; `3xx` / `4xx` / `5xx` = that status class only; `errors_only` = network/DNS/timeout/SSL (no HTTP status).

## `httpMethod` (type: `string`):

`head_then_get` (default): try HEAD, fall back to a streamed GET when the server returns 403/405/501 or an empty response. `head`: HEAD only. `get`: GET only (headers; body is not stored).

## `followRedirects` (type: `boolean`):

Follow HTTP redirects and record the final URL. When off, the first response status is kept (often 3xx).

## `maxConcurrency` (type: `integer`):

Parallel requests overall.

## `maxConcurrencyPerHost` (type: `integer`):

Parallel requests to any single host (kept low to be polite).

## `minDelayMsPerHost` (type: `integer`):

Minimum milliseconds between starting requests to the same host (simple rate limit). 0 = no extra delay.

## `requestTimeoutSecs` (type: `integer`):

Timeout for each HTTP request (connect + read).

## `maxRetries` (type: `integer`):

Retries for timeouts, network blips, HTTP 429 and selected 5xx (exponential backoff, honours Retry-After up to 30 s). DNS and certificate failures are not retried.

## `userAgent` (type: `string`):

Optional custom User-Agent header.

## `proxyConfiguration` (type: `object`):

Optional. Use a proxy if a site blocks data-center IPs. Not needed for most public sites.

## Actor input object example

```json
{
  "urls": [
    {
      "url": "https://example.com/"
    },
    {
      "url": "https://httpbin.org/status/404"
    }
  ],
  "keyValueRecordKey": "DOC_TO_MARKDOWN_INPUT",
  "maxUrls": 100,
  "outputFilter": "all",
  "httpMethod": "head_then_get",
  "followRedirects": true,
  "maxConcurrency": 8,
  "maxConcurrencyPerHost": 2,
  "minDelayMsPerHost": 0,
  "requestTimeoutSecs": 20,
  "maxRetries": 2,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Default dataset items, one per checked URL (after optional filter): url, finalUrl, httpStatus, statusClass, ok, contentType, contentLength, durationMs, methodUsed, redirectCount, errorClass, error, attempts, host. Views: overview, broken.

## `summary` (type: `string`):

Key-value store record SUMMARY (JSON): totalChecked, saved, counts by statusClass and errorClass, filter applied, durationSecs, peakMemoryMb.

## `output` (type: `string`):

Key-value store record OUTPUT (JSON): same run summary as SUMMARY plus charged event counts and input source notes. Kept for consistency with sibling Actors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        {
            "url": "https://example.com/"
        },
        {
            "url": "https://httpbin.org/status/404"
        }
    ],
    "maxUrls": 100,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("ingenious_quip_bxq/url-status-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        { "url": "https://example.com/" },
        { "url": "https://httpbin.org/status/404" },
    ],
    "maxUrls": 100,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("ingenious_quip_bxq/url-status-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    {
      "url": "https://example.com/"
    },
    {
      "url": "https://httpbin.org/status/404"
    }
  ],
  "maxUrls": 100,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call ingenious_quip_bxq/url-status-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ingenious_quip_bxq/url-status-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yV8iCN2SrK4NyPb1r/builds/s25mcITewDCmxQqu3/openapi.json
