# Broken Link Checker: Find Dead Links on Any Website (`jtpalms/broken-link-checker`) Actor

Crawl a website and find every broken link, image, script and stylesheet: 404s, 5xx, DNS and TLS errors, redirect loops and redirected links, each with the page it is on and its link text. Checks external links too. USD 1 per 1,000 pages plus USD 0.20 per 1,000 links.

- **URL**: https://apify.com/jtpalms/broken-link-checker.md
- **Developed by:** [JT Palms](https://apify.com/jtpalms) (community)
- **Categories:** SEO tools, Developer tools, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 page crawleds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Broken Link Checker: find dead links on any website

Give it your home page and it crawls the site, collects every link, image, script and stylesheet on every page, and checks each unique target once. You get one row per problem: the page it is on, the link, its text, the tag, the status code and what went wrong. Broken links (404, 410, 5xx, DNS, TLS, timeouts, redirect loops, malformed URLs), links blocked by bot protection (401, 403, 429, 999) and redirected links (with where they end up) are all reported.

**USD 1 per 1,000 pages crawled plus USD 0.20 per 1,000 links checked.** A 500-page site with 1,500 other links (images, scripts, external links) costs about USD 0.80.

### What people use it for

- **Website QA before and after launches.** Catch the 404s, missing images and broken scripts a redesign or CMS migration left behind.
- **SEO audits.** Broken internal links waste crawl budget and leak link equity; internal links that redirect add a hop for every crawler. Fix them from the list, page by page.
- **Agencies.** Run it on every client site on a schedule and send the results to Google Sheets or Slack.
- **Content maintenance.** Find outbound links in old blog posts and docs that now point to dead or moved pages.

### Sample output

Real rows from a 30-page crawl of scrapethissite.com, trimmed:

```json
{
  "sourcePage": "https://www.scrapethissite.com/pages/simple/",
  "linkUrl": "https://lipis.github.io/flag-icon-css/css/flag-icon.css",
  "linkText": "stylesheet",
  "element": "link",
  "issue": "BROKEN",
  "status": 404,
  "finalStatus": 404,
  "redirectTo": null,
  "redirectCount": 0,
  "isExternal": true,
  "errorType": null,
  "error": null
}
```

```json
{
  "sourcePage": "https://quotes.toscrape.com/",
  "linkUrl": "https://quotes.toscrape.com/author/Albert-Einstein",
  "linkText": "(about)",
  "element": "a",
  "issue": "REDIRECT",
  "status": 308,
  "finalStatus": 200,
  "redirectTo": "http://quotes.toscrape.com/author/Albert-Einstein/",
  "redirectCount": 1,
  "isExternal": false
}
```

That run crawled 30 pages, found 52 unique links, checked them in about 2 seconds and reported a missing stylesheet on 4 pages, two HTTP 400 links and two redirected external links. The second row, from quotes.toscrape.com, shows an internal link that redirects from https to http.

| Field | What it is |
|---|---|
| `sourcePage` | The page the link is on. A broken link on 40 pages gives 40 rows, so you know every place to fix. |
| `linkUrl`, `linkText`, `element` | The link target, its visible text (or image alt, `rel` for `<link>`), and the tag: `a`, `img`, `link` or `script`. |
| `issue` | `BROKEN`, `BLOCKED` (401/403/429/999, often bot protection: check it in a browser) or `REDIRECT`. |
| `status`, `finalStatus` | First response status and the status after following redirects. |
| `redirectTo`, `redirectCount` | Where a redirected link ends up and how many hops it took. |
| `isExternal` | `true` for links to other websites. |
| `errorType`, `error` | For links with no usable response: `DNS`, `TLS`, `TIMEOUT`, `CONNECTION_REFUSED`, `REDIRECT_LOOP`, `TOO_MANY_REDIRECTS`, `INVALID_URL` and so on. |

The run's key-value store gets a `SUMMARY` record: pages crawled, links checked, unique links found, counts of broken, blocked and redirected links, links skipped by robots.txt or your patterns, and the top 100 problem links with how many pages each is on.

### How to use it

1. Put your home page (or any start page) in **Website to check**.
2. Set **Max pages to crawl** (500 by default) and, if you like, **Max link depth**.
3. Keep **Check external links** on to catch dead outbound links, or turn it off to check only your own site.
4. Add **Skip URLs containing** patterns for endless or private sections, such as `/logout`, `/cart` or `?sort=`.
5. Click **Start**, then sort the table by `issue` or export CSV, JSON or Excel.

### How it crawls

- It starts at your URL and follows `<a href>` links (plus canonical, hreflang and next/prev `<link>` tags) on the same site, breadth first. `www` and non-`www` count as the same site; subdomains too if you turn that on.
- It collects every `a[href]`, `img[src]`, `link[href]` and `script[src]` on each crawled page, resolves them against `<base href>`, and ignores `mailto:`, `tel:`, `javascript:` and `#anchors`.
- Each unique target is checked once, however many pages link to it. Internal pages are checked by crawling them; everything else gets a HEAD request, with GET as a fallback for servers that refuse HEAD.
- It reads the site's robots.txt and does not request pages it disallows (you can turn this off for your own site).
- External sites get at most 2 requests at a time each, so nobody else's server is hammered.

### Pricing

| What | Price |
|---|---|
| Page crawled (an HTML page of the site downloaded and scanned for links) | USD 0.001 (USD 1 per 1,000) |
| Link checked (each unique external link, image, script, stylesheet, or internal URL that is not crawled) | USD 0.0002 (USD 0.20 per 1,000) |
| Result rows | Free |
| Start URL that cannot be reached | Free |

Links found on crawled pages are all checked, including internal pages beyond **Max pages**, so a site with many product images costs more in link checks than in pages. Set a maximum cost per run in the run options; the actor stops cleanly when it is reached and keeps every result found so far.

### Limits

- It reads the HTML the server sends and does not run JavaScript, so links added only by JavaScript (some single-page apps) are not found.
- Up to 5 MB of HTML per page. Links inside iframes, CSS files and `srcset` are not checked.
- Sites behind bot protection may answer with 403 or 429. Those rows are marked `BLOCKED`, not `BROKEN`, so you can check them by hand; a residential proxy often helps.
- Crawl-delay in robots.txt is not applied; lower **Parallel requests** for small or slow sites.

### Related actors

- [Website Screenshot API: Full Page, Mobile, PDF](https://apify.com/JTPalms/website-screenshot): Screenshot any list of URLs as full-page or viewport PNG, JPEG, WebP or PDF, on desktop, laptop, tablet or mobile.

### FAQ

**Why is a link marked BLOCKED?** The server refused the automated check with 401, 403, 429 or 999 (LinkedIn). The link may work fine in a browser. Open it to confirm.

**Why are redirects listed?** A link that redirects still works, but updating it to the final URL saves a hop for users and search engines. Turn off **Report redirected links** if you only want broken ones.

**Can it check a site that needs a login?** No. It sees what a logged-out visitor sees.

**Is it legal?** It requests public pages of the website you give it and the public URLs they link to, the same way a browser would, and it respects robots.txt by default. It does not log in and does not save page text; email and phone links are skipped, and the results are only link addresses and status codes.

**Something missing or wrong?** Open an issue with the site.

# Actor input Schema

## `startUrls` (type: `array`):

The page to start crawling from, usually the home page (https://example.com). The crawl stays on this site. Add more lines to check several sites in one run; Max pages is shared between them.

## `maxPages` (type: `integer`):

Stop crawling after this many pages of the site. Links found on those pages are still all checked, including internal pages beyond this limit (with a quick HEAD request, not a crawl).

## `maxDepth` (type: `integer`):

How many clicks away from the start page to crawl. 0 checks only the links on the start page. 1 also crawls the pages it links to, and so on.

## `checkExternalLinks` (type: `boolean`):

Also check links to other websites (each one once, at most 2 requests at a time per external site). Turn off to check only your own pages, images, scripts and stylesheets.

## `reportRedirects` (type: `boolean`):

Also list links that work but go through a redirect (for example http:// to https:// or a moved page), with where they end up. Updating them saves a hop and helps SEO.

## `excludeUrlPatterns` (type: `array`):

Links containing any of these texts are neither crawled nor checked, for example /logout, /cart or ?sort=. Wrap in slashes for a regular expression, for example //tag/.+/.

## `includeSubdomains` (type: `boolean`):

Crawl subdomains such as blog.example.com and shop.example.com as part of example.com. www and non-www always count as the same site.

## `respectRobotsTxt` (type: `boolean`):

Do not request pages of the site that its robots.txt disallows; they are counted in the run summary. Turn off only for a site you own or are allowed to audit.

## `maxConcurrency` (type: `integer`):

How many requests to make at the same time to the site being crawled. External sites always get at most 2 at a time. Lower it for small or slow sites.

## `timeoutSecs` (type: `integer`):

How long to wait for one request before reporting the link as timed out.

## `proxyConfiguration` (type: `object`):

Optional. Only needed if the site blocks requests from data centers.

## Actor input object example

```json
{
  "startUrls": [
    "https://www.scrapethissite.com/"
  ],
  "maxPages": 500,
  "maxDepth": 20,
  "checkExternalLinks": true,
  "reportRedirects": true,
  "includeSubdomains": false,
  "respectRobotsTxt": true,
  "maxConcurrency": 10,
  "timeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

All output rows in the default dataset (JSON, CSV, Excel via the format parameter).

## `summary` (type: `string`):

Counts and per-input status for this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://www.scrapethissite.com/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("jtpalms/broken-link-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": ["https://www.scrapethissite.com/"] }

# Run the Actor and wait for it to finish
run = client.actor("jtpalms/broken-link-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://www.scrapethissite.com/"
  ]
}' |
apify call jtpalms/broken-link-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jtpalms/broken-link-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/qZoHeOh61QXFrIu1Y/builds/0xSyE5HAQWobYbCXA/openapi.json
