# Bulk Broken Link Checker — 404 & Dead Link Audit (`weio/broken-link-checker-bulk`) Actor

Crawl a list of websites and find broken links and images: 404s, 5xx errors, DNS failures and timeouts, with the page each sits on, its anchor text and internal/external. HEAD failures re-checked by GET; bot walls (401/403/429) listed separately. Dead link checker for SEO audits and site migrations.

- **URL**: https://apify.com/weio/broken-link-checker-bulk.md
- **Developed by:** [Weio, Inc.](https://apify.com/weio) (community)
- **Categories:** SEO tools, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 site checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Broken Link Checker (bulk, crawls each site)

Audit one website or a list of websites for dead links before a migration, during an SEO audit, or as part of an agency's recurring client checks. For each supplied site, this Actor crawls same-host pages (up to 200; default 50) and tests every link and image it finds, internal and external.

### Best fit

- **Site migrations:** find links that need redirect or replacement work before and after a move.
- **SEO and content audits:** identify 404s, server errors, timeouts, and broken redirects with the page and anchor text that reference each link.
- **Agencies:** run the same repeatable check across a client list and hand the result dataset to the next reporting or repair step.

### Quick start

Enter the websites to audit and choose a crawl limit that fits the size of the site. The Actor follows public, static HTML links only; it does not render JavaScript or submit forms. A link assembled by JavaScript in a visitor's browser therefore will not be discovered by this crawl.

Every failed HEAD request is confirmed with GET before a link is called broken.

### Results

The dataset has one row per supplied website:

- `pagesCrawled`, `linksChecked`, and `brokenCount`
- `brokenLinks`: URL, HTTP `status` (404, 410, 500...), or network `error` (DNS failure, timeout, TLS); the source page (`foundOn`), anchor text, and internal/external flag
- `blockedLinks`: links that answered 401, 403, 406, or 429, and links the crawler was not allowed to check (the target's robots.txt, a private network address, or a host asking us to slow down), with the reason in `error`

Those 401/403/406/429 responses are listed separately, not counted as broken. They usually mean authentication, rate limiting, or a bot wall prevented verification; they do not prove that the destination is gone.

#### Real result example

This row is from an existing successful run on a small company website, shown without its address (no new run was made for this example):

| website | pagesCrawled | linksChecked | brokenCount | blockedCount |
| --- | ---: | ---: | ---: | ---: |
| (address left out) | 1 | 12 | 0 | 0 |

### Pricing

One `site-checked` event is charged per website that can be crawled. Current examples: **$0.01 per site** on Free and Starter, **$0.0075** on Scale, and **$0.005** on Business and Enterprise. A site that is unreachable, blocks the crawler at the start, or answers with a bot check, a queue page or an empty or non-HTML start page returns a free `error` row; it is not charged.

### API, schedules, and integrations

Use the [Apify API](https://docs.apify.com/api/v2) to start runs and consume the resulting dataset from your own code. For recurring audits, create an Actor task and use [Apify schedules](https://docs.apify.com/actors/running/schedules). Apify's [integrations documentation](https://docs.apify.com/integrations) covers workflow tools, webhooks, data destinations, and AI clients.

### FAQ

**Why are some links missing?** The Actor reads static HTML and does not run browser JavaScript, so it cannot see links that a page creates only after JavaScript executes.

**Why is a 401, 403, 406, or 429 not called broken?** Those statuses say access was refused or limited, not that the target no longer exists. They remain in `blockedLinks` so you can review them separately without inflating the broken-link count. Links whose robots.txt does not allow our crawler, or that point to a private network address, are listed there too: they were not checked, so they are not called broken.

**Are private pages or forms checked?** No. The Actor checks public pages only and never submits forms.

# Actor input Schema

## `websites` (type: `array`):

Domains or start URLs, one per line (max 200 per run).

## `maxPagesPerSite` (type: `integer`):

Same-host HTML pages followed from the start URL.

## `checkExternal` (type: `boolean`):

Also test links pointing to other domains (one request each).

## `checkImages` (type: `boolean`):

Also test <img> sources.

## Actor input object example

```json
{
  "websites": [
    "anchorbrewing.com"
  ],
  "maxPagesPerSite": 50,
  "checkExternal": true,
  "checkImages": true
}
```

# Actor output Schema

## `siteRows` (type: `string`):

The default dataset with one row per checked website.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "anchorbrewing.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("weio/broken-link-checker-bulk").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": ["anchorbrewing.com"] }

# Run the Actor and wait for it to finish
run = client.actor("weio/broken-link-checker-bulk").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "anchorbrewing.com"
  ]
}' |
apify call weio/broken-link-checker-bulk --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,weio/broken-link-checker-bulk"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/dFy5qWbZvT4frdpXO/builds/1I0ucFDyuxkmZtsQX/openapi.json
