# Broken Link Checker: Find Dead Links on Any Site (`digital_influx/broken-links`) Actor

Crawl a website and get every broken link (404, 410, 5xx, dead domains), internal and external, with the exact pages it appears on. Schedule it to see only newly broken and fixed links. Respects robots.txt. Pay per page crawled.

- **URL**: https://apify.com/digital\_influx/broken-links.md
- **Developed by:** [Bruno Petrelli](https://apify.com/digital_influx) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 page crawleds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Broken Link Checker: Find Dead Links on Any Site

Give it a website and get **every broken link, internal and external, with the exact pages it appears on**, the most used first, so you fix what hurts most. The Actor crawls the site through its links and its sitemap, then checks every link it found:

- **Broken:** 404, 410, 5xx, domains that no longer exist, refused connections, bad certificates, timeouts.
- **Blocked:** the server refused the checker (401, 403, 429, LinkedIn's 999, or a non-standard 4xx code such as the 419 Hacker News sends to scripts), or the link only takes other methods (405, like an API endpoint). Listed apart, because the page may well exist.
- **Working:** optional, for a full link inventory.

**Respects robots.txt** on every host it reads, including the sites your links point to. No browser and no proxies; the User-Agent says `broken-links`.

### Use it for

- **Before a launch or after a migration:** catch every 404 your own pages link to.
- **Monthly site maintenance:** schedule it and get only the new problems in a CSV.
- **SEO and UX:** broken links waste crawl budget and lose visitors. Fix the ones on the most pages first.
- **Agencies:** check all your clients' sites in one run, one summary row per site.

### Input

```json
{
  "websites": ["example.com"],
  "maxPagesPerSite": 500,
  "checkExternalLinks": true
}
```

### Output

One row per problem link (a real result on 2026-09-25; the addresses were replaced):

```json
{
  "type": "link",
  "site": "example.com",
  "url": "https://example.thrivecart.com/checkout/",
  "state": "broken",
  "status": 404,
  "error": null,
  "internal": false,
  "foundOnCount": 1,
  "foundOn": ["https://example.com/academy"]
}
```

- `state`: `broken`, `blocked` or `ok`. `error` explains network failures (for example `ENOTFOUND` for a dead domain).
- `foundOn` lists up to 25 pages that contain the link; `foundOnCount` is the total.
- One free summary row per website: `pagesCrawled`, `linksChecked`, `brokenLinks`, `blockedLinks`, `pagesWithBrokenLinks`, `notCrawled` (pages found beyond your limit), `blockedByRobots`.

That same run crawled 16 pages in 21 seconds and checked 39 unique links.

### Monitoring: only what changed

Schedule the check (daily, weekly) with **Compare with the previous check** on. Each broken link then says `newlyBroken: true` when it was fine (or not linked) at the last check, and `false` when it was already broken, so a report can show only what broke since. The site row adds `changes`: how many links are newly broken, the links that were fixed (still linked and working now) and the broken links that no page links to any more, with the date of the check it was compared with. When the page limit cut the crawl (`notCrawled` above 0), that last list is `null`: a link found last time on a page not read now is unknown, not gone. The first check has nothing to compare with (`changes: null`).

### Pricing

Pay per event: **one event per page crawled**, whatever the number of links on it. Link rows and summary rows are free. Set a maximum cost for the run and the Actor stops cleanly when it is reached.

Need the full SEO picture per page (titles, descriptions, headings, noindex, alt text, duplicates)? Use **SEO Audit & Broken Link Checker** by the same author, which includes this link check.

### Good to know

- Only pages on the website's own host are crawled (`www.` and the bare domain count as the same site); links to anywhere are checked.
- Links are checked with a HEAD request; servers that refuse HEAD get one GET.
- Links inside content that only appears after JavaScript runs are not seen.
- Crawling is gentle by default: 3 requests at a time per website.

### Support

A link reported wrongly? Open an issue on the Actor's Issues tab with the link and the page. Issues get an answer within a few days.

# Actor input Schema

## `websites` (type: `array`):

One per line: a domain (example.com) or the URL where the crawl starts. Each website is crawled separately, following its internal links and its sitemap.

## `maxPagesPerSite` (type: `integer`):

Pages crawled per website (each one is charged). The links on those pages are all checked.

## `checkExternalLinks` (type: `boolean`):

Also check links to other websites. Internal links are always checked.

## `maxExternalLinksPerSite` (type: `integer`):

Unique external links checked per website.

## `includeBlockedLinks` (type: `boolean`):

Links that answered 401, 403, 429 or 999 (LinkedIn). The server refused the checker; the page may well exist. Listed with state "blocked".

## `includeOkLinks` (type: `boolean`):

Output every checked link, not only the problems (a full link inventory). Rows are free either way.

## `useSitemap` (type: `boolean`):

Also crawl pages listed in the sitemap that no link points to.

## `respectRobotsTxt` (type: `boolean`):

Recommended. Pages disallowed by robots.txt are not read, and external links on sites that disallow crawlers are not checked. Turn off only for your own website.

## `maxConcurrency` (type: `integer`):

Requests to one website at the same time (1 to 10).

## `compareWithPreviousRun` (type: `boolean`):

For scheduled checks: every broken link gets "newlyBroken" (true if it was not broken at the last check of the site), and the site row lists the links that were fixed and the broken links no page links to any more. The broken links of the last check are kept in a key-value store named "broken-links-state" in your account.

## `stateKey` (type: `string`):

Give a name to keep a separate history, for example one per client, or a new name to start over.

## Actor input object example

```json
{
  "websites": [
    "37signals.com"
  ],
  "maxPagesPerSite": 30,
  "checkExternalLinks": true,
  "maxExternalLinksPerSite": 500,
  "includeBlockedLinks": true,
  "includeOkLinks": false,
  "useSitemap": true,
  "respectRobotsTxt": true,
  "maxConcurrency": 3,
  "compareWithPreviousRun": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "37signals.com"
    ],
    "maxPagesPerSite": 30
};

// Run the Actor and wait for it to finish
const run = await client.actor("digital_influx/broken-links").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": ["37signals.com"],
    "maxPagesPerSite": 30,
}

# Run the Actor and wait for it to finish
run = client.actor("digital_influx/broken-links").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "37signals.com"
  ],
  "maxPagesPerSite": 30
}' |
apify call digital_influx/broken-links --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,digital_influx/broken-links"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/iILLVg0jr0Dz8LnX0/builds/VWkdGhvap7yBSySuf/openapi.json
