# Bulk URL Checker – HTTP Status & Broken Links (`datascoutapi/bulk-url-status-checker`) Actor

Check 10000's of URLs in bulk for HTTP status codes, broken links, 404 errors, redirects, redirect chains, final URLs, and response times.

- **URL**: https://apify.com/datascoutapi/bulk-url-status-checker.md
- **Developed by:** [halam](https://apify.com/datascoutapi) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk URL Status Checker & Redirect Tracer — HTTP Status Codes, Broken Links & Redirect Chains at Scale

Bulk-check thousands of URLs for HTTP status codes, full redirect chains, broken links, response times, and content type — all in one run. No browser, no login, no API key required. The go-to tool for SEO audits, site migrations, link monitoring, and QA pipelines.

Powered by [RedirectChecker.com](https://redirectchecker.com) — built for speed, accuracy, and scale.

***

### 🏆 Why choose this Bulk URL Status Checker?

Check thousands of URLs per run · full hop-by-hop redirect chains · per-hop latency · 14 output fields per URL · 29 user agent options including Googlebot, GPTBot, and SEMrush · export to JSON / CSV / Excel.

The most data-rich bulk HTTP status checker on Apify — built for SEO professionals, developers, and QA teams who need more than just a status code.

***

### ✨ Key features

- 🔗 **Full redirect chain tracing** — records every hop with URL, status code, latency, and complete response headers. Not just the final destination — every intermediate step.
- 🚦 **Complete HTTP status detection** — captures 200 OK, 301/302 redirects, 403 Forbidden, 404 Not Found, 410 Gone, 429 Too Many Requests, 500 Server Error, and network errors (DNS failures, timeouts).
- 🔴 **Broken link detection** — `is_broken` flag instantly identifies 4xx, 5xx, and network failures. Filter your entire dataset in one click.
- ⏱️ **Per-hop latency measurement** — see exactly which redirect hop is slow, not just the total round-trip time. Millisecond precision.
- 🤖 **29 user agent options** — check how Chrome, Googlebot, GPTBot, Bingbot, AhrefsBot, SEMrushBot, and 23 more browsers and crawlers see your URLs.
- 📍 **Check location reporting** — shows which server location (e.g. Washington DC, Frankfurt, Tokyo) performed the check.
- 📋 **Content type & length** — know what the final response returns, not just whether it responded.
- ⚡ **Concurrent batch processing** — 3 batches of 100 URLs run in parallel. Fast even for large lists.
- 🔄 **Automatic retries** — transient failures are retried automatically. Failed URLs always get a result, never silently dropped.
- 📤 **Export-ready** — download every result as JSON, CSV, Excel, or JSONL from the Apify Dataset.

***

### 🚀 Quick start (3 steps)

1. **Add your URLs** — paste them into the *URLs to check* field, one per line
2. **Choose your user agent** — leave as `chrome` for standard checks, or pick `googlebot` to see what Google sees
3. **Click Start** — results fill the dataset within seconds, ready to export as JSON / CSV / Excel

No API key. No login. No code to write.

***

### 📥 Input

| Field | Type | Default | Description |
|---|---|---|---|
| `urls` | `string[]` | — | **Required.** List of URLs to check. No hard limit. |
| `userAgent` | `string` | `chrome` | Browser or bot user agent — see full list below |
| `timeout` | `integer` | `10000` | Per-URL timeout in ms (1,000–60,000) |
| `includeHeaders` | `boolean` | `true` | Include full response headers for every hop |

#### Example input — broken link audit

```json
{
  "urls": [
    "https://example.com",
    "https://example.com/old-page",
    "https://example.com/contact"
  ],
  "userAgent": "chrome",
  "timeout": 10000,
  "includeHeaders": true
}
```

#### Example input — Googlebot redirect check

```json
{
  "urls": [
    "https://oldsite.com/legacy-path",
    "https://oldsite.com/page-1",
    "https://oldsite.com/page-2"
  ],
  "userAgent": "googlebot",
  "timeout": 10000,
  "includeHeaders": true
}
```

#### Supported user agents

**Desktop browsers**

| Key | Browser |
|---|---|
| `chrome` | Chrome 124 — Windows (default) |
| `chrome_mac` | Chrome 124 — macOS |
| `firefox` | Firefox 125 — Windows |
| `safari` | Safari 17.4 — macOS |
| `edge` | Edge 124 — Windows |
| `opera` | Opera 110 — Windows |
| `brave` | Brave (Chrome engine) |

**Mobile**

| Key | Device |
|---|---|
| `mobile` | iPhone Safari — iOS 17 |
| `android` | Pixel 8 Chrome — Android |
| `samsung` | Samsung Browser — Android |

**Search engine bots**

| Key | Bot |
|---|---|
| `googlebot` | Googlebot Desktop |
| `googlebot_mobile` | Googlebot Mobile |
| `bingbot` | Bingbot |
| `yandexbot` | YandexBot |
| `duckduckbot` | DuckDuckBot |
| `baidubot` | Baiduspider |
| `applebot` | Applebot |

**Social media bots**

| Key | Platform |
|---|---|
| `twitterbot` | Twitter / X |
| `facebookbot` | Facebook |
| `linkedinbot` | LinkedIn |
| `slackbot` | Slack |
| `whatsapp` | WhatsApp |
| `telegram` | Telegram |

**AI bots**

| Key | Bot |
|---|---|
| `gptbot` | OpenAI GPTBot |
| `claudebot` | Anthropic ClaudeBot |
| `perplexitybot` | Perplexity AI |

**SEO tools**

| Key | Tool |
|---|---|
| `ahrefsbot` | Ahrefs |
| `semrushbot` | SEMrush |
| `screaming` | Screaming Frog |

***

### 📤 Output

One row per URL pushed to the Apify Dataset — 14 fields per result.

| Field | Type | Description |
|---|---|---|
| `input_url` | string | Original URL submitted |
| `final_url` | string | Final URL after all redirects |
| `final_status` | integer | HTTP status code at the final destination |
| `status_message` | string | Status text — e.g. `OK`, `Not Found`, `Too Many Requests` |
| `is_broken` | boolean | `true` for 4xx, 5xx, and network errors |
| `redirect_count` | integer | Number of 3xx redirects followed |
| `has_redirect` | boolean | `true` if at least one redirect occurred |
| `total_time_ms` | integer | Total round-trip time in milliseconds |
| `content_type` | string | Content-Type of the final response |
| `content_length` | integer | Content-Length of the final response in bytes |
| `check_location` | object | Server location that performed the check — `{ iata, city }` |
| `chain` | array | Every hop: `url`, `status_code`, `status_text`, `latency_ms`, `headers` |
| `checked_at` | string | ISO 8601 timestamp |
| `error` | string | Error message if request failed, otherwise `null` |

#### Example — successful redirect (GitHub HTTP → HTTPS)

```json
{
  "input_url": "http://github.com",
  "final_url": "https://github.com/",
  "final_status": 200,
  "status_message": "OK",
  "is_broken": false,
  "redirect_count": 1,
  "has_redirect": true,
  "total_time_ms": 31,
  "content_type": "text/html",
  "content_length": null,
  "check_location": { "iata": "IAD", "city": "Washington DC" },
  "chain": [
    { "hop": 1, "url": "http://github.com/", "status_code": 301, "status_text": "Moved Permanently", "latency_ms": 21, "headers": { "location": "https://github.com/" } },
    { "hop": 2, "url": "https://github.com/", "status_code": 200, "status_text": "OK", "latency_ms": 10, "headers": {} }
  ],
  "checked_at": "2026-09-25T12:26:15.931Z",
  "error": null
}
```

#### Example — broken link (404)

```json
{
  "input_url": "https://example.com/old-page",
  "final_url": "https://example.com/old-page",
  "final_status": 404,
  "status_message": "Not Found",
  "is_broken": true,
  "redirect_count": 0,
  "has_redirect": false,
  "total_time_ms": 210,
  "content_type": "text/html",
  "content_length": 1256,
  "check_location": { "iata": "LHR", "city": "London" },
  "chain": [],
  "checked_at": "2026-09-25T12:26:16.000Z",
  "error": null
}
```

#### Example — network error

```json
{
  "input_url": "https://this-domain-does-not-exist.io/page",
  "final_url": null,
  "final_status": 0,
  "status_message": "Network Error",
  "is_broken": true,
  "redirect_count": 0,
  "has_redirect": false,
  "total_time_ms": null,
  "content_type": null,
  "content_length": null,
  "check_location": null,
  "chain": [],
  "checked_at": "2026-09-25T12:26:17.000Z",
  "error": "getaddrinfo ENOTFOUND this-domain-does-not-exist.io"
}
```

#### HTTP status code reference

| Code | Message | `is_broken` | Meaning |
|---|---|---|---|
| 200 | OK | ❌ | Page loaded successfully |
| 301 | Moved Permanently | ❌ | Permanent redirect (followed) |
| 302 | Found | ❌ | Temporary redirect (followed) |
| 308 | Permanent Redirect | ❌ | Permanent redirect (followed) |
| 401 | Unauthorized | ✅ | Authentication required |
| 403 | Forbidden | ✅ | Access denied |
| 404 | Not Found | ✅ | Page does not exist |
| 410 | Gone | ✅ | Page permanently removed |
| 429 | Too Many Requests | ✅ | Rate limited by target server |
| 500 | Internal Server Error | ✅ | Server-side error |
| 503 | Service Unavailable | ✅ | Server temporarily down |
| 0 | Network Error | ✅ | DNS failure, timeout, connection refused |

***

### 💡 Use cases

**SEO broken-link audits**
Export your full site URL list, run the actor, filter `is_broken = true`. Prioritize fixing 404s and 5xx errors by page importance to protect crawl budget and search rankings.

**Redirect chain optimization**
Use `chain[]` to spot multi-hop redirect chains (`redirect_count > 1`) and collapse them into single direct redirects — saving latency and preserving link equity.

**Site migration validation**
Confirm every old URL returns 301/308 and that `final_url` points to the correct new destination. Catch missing redirects before they cost you rankings.

**Googlebot vs user comparison**
Run the same URL list twice — once with `userAgent: googlebot` and once with `userAgent: chrome` — to detect cloaking, soft 404s, or bot-specific behavior.

**AI bot access auditing**
Use `gptbot`, `claudebot`, or `perplexitybot` to verify your robots.txt and server rules are correctly blocking or allowing AI crawlers.

**Affiliate & backlink link health checks**
Verify that inbound affiliate links and earned backlinks still resolve to 200 OK. A 404 on a linked page means lost link equity and lost commissions.

**API & webhook URL validation**
Run all endpoint URLs through the actor before deploying an integration. Confirm they return the expected 2xx rather than a surprise 4xx or 5xx.

**Uptime & QA monitoring**
Schedule recurring runs via the Apify Scheduler and fire webhooks on completion to alert on newly broken links. Drop it into CI/CD as a link-health gate.

**Content migration QA**
After a CMS migration, batch-check every URL in the old sitemap to confirm redirects are in place and no pages returned 404.

***

### 👥 Who uses this actor

- **SEO teams & auditors** — identify 404s, redirect chains, and crawl budget leaks across large sites
- **Web developers & QA engineers** — validate links after migrations, CMS changes, or deployments
- **Growth & affiliate marketers** — monitor that earned backlinks and affiliate URLs resolve correctly
- **Site-migration engineers** — verify redirect maps end-to-end before and after cutover
- **Content managers** — keep internal and external link health clean across articles
- **Security researchers** — inspect redirect chains and response headers for misconfigurations

***

### ⏰ Scheduling & integration

- **Schedule** recurring runs (daily, weekly, or any cron interval) from the Apify Console
- **API-first** — start runs and pull results programmatically via the Apify API; filter broken URLs with `GET /v2/datasets/{id}/items?filter=is_broken%3Dtrue`
- **Webhooks** — trigger Slack, email, or your own endpoint on run completion
- **Pipeline it** — chain with a Sitemap Crawler actor for a fully automated whole-site health check

***

### ❓ Frequently asked questions

**Do I need an API key or login?**
No. Just paste your URLs and click Start.

**How many URLs can I check per run?**
No hard limit. The actor batches URLs internally and processes them concurrently. Check an entire sitemap in a single run.

**Does it follow redirects?**
Yes — up to 15 hops per URL. Every intermediate URL, status code, and latency is recorded in `chain[]`.

**How do I find all broken links?**
Filter the dataset by `is_broken = true`. This covers 4xx, 5xx, and network errors in one step.

**How do I check HTTP (non-HTTPS) URLs?**
Both HTTP and HTTPS are fully supported.

**Can I see how Googlebot crawls my URLs?**
Yes — set `userAgent: googlebot` or `googlebot_mobile`. You can also check with `bingbot`, `ahrefsbot`, `semrushbot`, `gptbot`, and 24 more user agents.

**What does `check_location` mean?**
It shows which server location performed the check — for example `{ "iata": "IAD", "city": "Washington DC" }`. Useful for detecting geo-based redirects or location-specific responses.

**How do I export to CSV?**
Click **Export** on the dataset page and choose CSV. The `chain[]` array is serialized as JSON within the CSV cell; all scalar fields are separate columns.

**Can I use it for scheduled link monitoring?**
Yes. Use the Apify Scheduler to trigger daily or weekly runs and connect webhooks or the API to alert when new broken links appear.

**Is it legal?**
The actor sends standard HTTP requests to URLs you supply — the same as a browser or any link checker. You are responsible for checking only URLs you are authorized to access and for complying with applicable laws and the target sites' terms of service.

***

### ⚖️ Legal

This actor sends standard HTTP GET requests to URLs you provide and reports the returned status metadata. No authentication bypass, no content harvesting, no scraping. You are responsible for checking only URLs you are authorized to access and for complying with target websites' terms of service, robots directives, and all applicable laws and regulations in your jurisdiction.

# Actor input Schema

## `urls` (type: `array`):

List of URLs to check. Paste one URL per line. Max 500 per batch.

## `userAgent` (type: `string`):

Browser, bot, or crawler user agent to use for all requests. Useful for checking how Google, Bing, or social media bots see your URLs vs real users.

## `timeout` (type: `integer`):

Maximum time in milliseconds to wait for each URL. Applies per URL, not per batch.

## `includeHeaders` (type: `boolean`):

When enabled, full response headers for each hop in the redirect chain are included in the output. Useful for inspecting Location, Cache-Control, Content-Type, and security headers.

## Actor input object example

```json
{
  "urls": [
    "https://example.com",
    "http://github.com"
  ],
  "userAgent": "chrome",
  "timeout": 10000,
  "includeHeaders": true
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset containing one row per checked URL with status code, redirect chain, is\_broken flag, response time, content type, and Cloudflare check location.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com",
        "http://github.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("datascoutapi/bulk-url-status-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://example.com",
        "http://github.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("datascoutapi/bulk-url-status-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com",
    "http://github.com"
  ]
}' |
apify call datascoutapi/bulk-url-status-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datascoutapi/bulk-url-status-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/p0CpcBz6syXpf08hG/builds/12t5E3Ozeaj4tiijJ/openapi.json
