# Broken Link Checker: 404s, No Login (`conserving_celerytop/broken-link-checker`) Actor

Find broken links on your website with a bulk broken link checker. $1.00 per 1,000 pages checked, links free. Crawls your pages and checks every internal and external link, image, script and stylesheet once. Returns status code, redirect chain, source pages and anchor text. MCP-ready, no login.

- **URL**: https://apify.com/conserving\_celerytop/broken-link-checker.md
- **Developed by:** [Don Mangu](https://apify.com/conserving_celerytop) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 page checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Broken Link Checker

Broken Link Checker finds broken links on your website. It crawls your pages over plain HTTP, checks every internal and external link once, and also checks images, scripts, stylesheets and icons. For each problem link you get the HTTP status code, the redirect chain, the pages that contain it and the anchor text, so you know exactly what to fix and where.

Use it for a one-off website audit, before and after a migration or redesign, or on a schedule so a dead link never sits on your site for weeks.

### What does Broken Link Checker find?

| Category | What it means |
| --- | --- |
| `broken` | 404, 410, 500 and other 4xx or 5xx answers, redirect loops, malformed link addresses and domains that no longer exist |
| `unreachable` | The server timed out or refused the connection three times |
| `restricted` | 401, 403, 429 or 999: the server refused an automated check. The link may still work in a browser, so check it by hand |
| `redirected` | The link works only through one or more redirects. Point it to the final address to save a hop |
| `not_checked` | Not requested: the target's robots.txt disallows it, or the target site does not allow automated checks |
| `ok` | Answered 2xx directly (listed only when you turn on "List working links too") |

### How to check a website for broken links

1. Enter your website in **Start URLs**, as a domain (`example.com`) or a full URL.
2. Set **Maximum pages per website**. Start with 10 to see the output, then raise it.
3. Choose whether to check external links and images, scripts and stylesheets. Both are on by default.
4. Click **Start**. A 100-page website usually finishes in 5 to 15 minutes, because the Actor paces its requests so your server is never under load.
5. Open the **Problem links** view and sort by `foundOnCount` to fix the links that appear on the most pages first. Export to CSV, Excel or JSON, or read the dataset through the API.

To keep a site clean, save the input as a task and add a weekly schedule in Apify.

### How much does it cost to check links?

Broken Link Checker uses pay-per-event pricing. You pay **$0.001 per page crawled** on the Free plan ($1.00 per 1,000 pages; $0.0009 on Starter, $0.0008 on Scale, $0.0007 on Business), and every link on that page is checked at no extra cost. Link rows and the site summary are free.

Worked example: a website with 100 pages and about 600 unique links costs 100 x $0.001 = **$0.10** on the Free plan, plus the small Apify start fee. A tool that charges $0.001 per link would charge about $0.60 for the same site.

You can set a spending limit on the run. The Actor reads it before crawling and never crawls more pages than you pay for.

### Input example

```json
{
    "startUrls": ["https://www.example.com/"],
    "maxPagesPerSite": 200,
    "checkExternalLinks": true,
    "checkAssets": true,
    "reportRedirects": true
}
```

### Output example

One row per problem link:

```json
{
    "rowType": "link",
    "url": "https://www.example.com/old-pricing",
    "category": "broken",
    "status": 404,
    "finalUrl": "https://www.example.com/old-pricing",
    "redirectChain": [],
    "isInternal": true,
    "linkType": "link",
    "error": null,
    "foundOnCount": 12,
    "foundOn": [{ "pageUrl": "https://www.example.com/", "anchorText": "See pricing" }]
}
```

You also get one row per crawled page (`rowType: "page"`) with its link counts and up to 50 problem links, and one summary row per website (`rowType: "site"`) with totals and the reason the check stopped, if it stopped early.

### How it works

- It starts at your URL and follows links on the same host (with or without `www.`), up to your page limit.
- Links found on those pages are each checked once per run, even if they appear on hundreds of pages. Pages beyond the limit are still checked as links.
- External links get a HEAD request first and a GET only when HEAD is refused, so the check is light on other websites.
- It obeys robots.txt on every host, including the Crawl-delay of your own site (up to 10 seconds), and waits at least 0.25 s between requests to your site and 1 s between requests to any external host.
- It identifies itself with an honest user agent that includes a contact address.

### Related Actors

- [Sitemap URL Extractor API](https://apify.com/conserving_celerytop/sitemap-url-extractor): Use it to list every page URL of a website from its XML sitemaps.
- [AI Crawler Access Audit](https://apify.com/conserving_celerytop/ai-crawler-access-audit): Use it to check which AI crawlers a website allows or blocks in robots.txt.
- [Domain Authority Checker](https://apify.com/conserving_celerytop/domain-authority-checker): Use it to check domain authority for a bulk list of domains.
- [Website Screenshot API](https://apify.com/conserving_celerytop/website-screenshot-api): Use it to capture full-page screenshots or PDFs of a list of URLs.

### FAQ

**Is it legal to check links on a website?**
Checking the links on your own website, or on a public website you have permission to audit, is a normal part of site maintenance. The Actor reads only public pages, needs no login, collects no personal data, and respects robots.txt. It does not crawl sites whose terms forbid automated access, such as LinkedIn, Facebook, Instagram, X, Reddit, TikTok, Amazon and Google. Links pointing to those sites are listed as `not_checked` without a request.

**Why is a link "restricted" when it works in my browser?**
Some servers refuse automated requests with 403 or 999. The Actor does not try to get around that. Open those links by hand once.

**Does it render JavaScript?**
No. It reads the HTML your server sends. Links that exist only after JavaScript runs are not found. Most websites put their navigation and content links in the HTML.

**Can it check a page list instead of crawling?**
Not in this version. Set Maximum pages per website to 1 to check only the links on the start page.

**The run stopped early. Why?**
Read `stoppedReason` in the site row: the page limit, your spending limit, or the website refusing requests (three block signals in a row). The Actor stops instead of pushing harder.

Found a problem or need a feature? Open an issue on the Issues tab and we will reply within a day.

# Actor input Schema

## `startUrls` (type: `array`):

Enter the websites to check, one per line, as a domain (example.com) or a URL. Each website is crawled on its own host from this address. Up to 10 websites per run.

## `maxPagesPerSite` (type: `integer`):

Crawl at most this many pages per website. Every link on those pages is checked. Each crawled page is charged once.

## `checkExternalLinks` (type: `boolean`):

Check links that point to other websites. Each external link is requested once, with a pause between requests to the same host.

## `checkAssets` (type: `boolean`):

Check image, script, stylesheet, icon and media files as well as page links.

## `reportRedirects` (type: `boolean`):

List links that work only through a redirect, so you can point them to the final address.

## `reportOkLinks` (type: `boolean`):

Save a row for every checked link, including the ones that work.

## `excludeUrlPatterns` (type: `array`):

Do not crawl pages whose URL matches one of these patterns. Use \* as a wildcard (/blog/\*) or a regular expression in slashes. Links to them are still checked.

## `maxConcurrency` (type: `integer`):

Requests in flight at once. Each host still gets at most one request every 0.25 s (1 s for external hosts), and the website's robots.txt Crawl-delay is honoured up to 10 s.

## Actor input object example

```json
{
  "startUrls": [
    "https://www.gov.uk/"
  ],
  "maxPagesPerSite": 3,
  "checkExternalLinks": true,
  "checkAssets": true,
  "reportRedirects": true,
  "reportOkLinks": false,
  "excludeUrlPatterns": [],
  "maxConcurrency": 4
}
```

# Actor output Schema

## `links` (type: `string`):

Each broken, redirected, restricted or unreachable link with status code and the pages it is on.

## `pages` (type: `string`):

One row per crawled page with its link counts and problem links.

## `site` (type: `string`):

Totals per website: pages, links, broken, redirected and why the check stopped.

## `all` (type: `string`):

Every row with every field.

## `stats` (type: `string`):

Requests, retries, charges and time per website.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://www.gov.uk/"
    ],
    "maxPagesPerSite": 3
};

// Run the Actor and wait for it to finish
const run = await client.actor("conserving_celerytop/broken-link-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["https://www.gov.uk/"],
    "maxPagesPerSite": 3,
}

# Run the Actor and wait for it to finish
run = client.actor("conserving_celerytop/broken-link-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://www.gov.uk/"
  ],
  "maxPagesPerSite": 3
}' |
apify call conserving_celerytop/broken-link-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,conserving_celerytop/broken-link-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/xj21Sl2HwsJnuCgP1/builds/KIMV3BXeWi53o2xHg/openapi.json
