# Broken Link Checker: Every 404 Billed Once (`pradio/broken-link-checker`) Actor

Find broken links and 404s on a page, a list, a whole site or its sitemap, billed per link the site answered for. An error on HEAD is re-checked with GET, and a site that refuses the checker is never called broken. One row per link with its status, verdict and page.

- **URL**: https://apify.com/pradio/broken-link-checker.md
- **Developed by:** [Pradio Actors](https://apify.com/pradio) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.43 / 1,000 link checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Broken Link Checker: Every 404 Billed Once

### What does Broken Link Checker do?

Broken Link Checker finds the broken links on a page, a site or its sitemap and returns one row per link it checks. Each row carries the link's final HTTP status, an `ok`, `broken` or `error` verdict, the anchor text when it had any, the page it sat on and every redirect it took. Every link the site answered for is billed once at $0.0008. Paste URLs or bare domains, pick a mode and press Start; what comes back is your fix list.

On 40 sites it had never seen, run 2026-10-01, every one of the 40 produced a row, 80% of the sites answered and 30 produced checked-link rows. A link it could not judge is a free `error` row with the reason on it.

### Who uses Broken Link Checker

| Buyer | What they run it for |
|---|---|
| Website owners | Finding dead links on their own pages before visitors or crawlers do. |
| SEO practitioners | A technical SEO audit in one run: dead links leak ranking, and each row names the page and the anchor to fix. |
| Teams after a migration or redesign | Checking that old URLs still redirect somewhere instead of dying. |
| Content maintainers | Auditing outbound links that rot quietly over time. |

### Features

- **Every 404 re-checked with GET**. You never see `broken` on a HEAD answer alone: an error code that is not a refusal is asked once more with GET, the request a browser makes. The GET decides. Some servers answer HEAD with an error and serve the page to GET; those links come back `ok`.
- **A 5xx is asked once more, after a pause**. A server error is often a passing failure, not a dead link. Before a link is called `broken` on one, it is re-checked once after a short wait. A healthy re-check comes back `ok` and `status_text` names the transient first answer. Only a 5xx that stands on re-check, or one whose re-check never answers, lands as `broken`.
- **A page, a list, a whole site or its sitemap**. `mode` picks what is read: only the pages you list, a crawl of each site from the page you give, or every page the site's sitemap lists.
- **Crawls politely, in every mode**. Every mode, Pages included, reads the site's robots.txt first and never reads a page it disallows. The file is read under the checker's own name and under the browser identity its requests carry; a Disallow under either stops the read. Its crawl delay is waited between every request to that site, link checks included. No run reads more than 2,000 pages in any mode, and a site that refuses the checker (403, 429 or 999) is read no further.
- **A refusal is not a dead link**. When a site refuses the request (401, 403, 429 or 999) it has not said the link is dead. You get a free `error` row with the diagnosis `blocked_by_site`, never a billed `broken`.
- **A script gate is named for what it is**. A site that answers the read with a script gate, an Incapsula or Cloudflare challenge stub instead of the real page, is reported on a free `blocked` row. `reason` names the gate. It is never mistaken for a page that carries no links, and the read is never retried past the gate.
- **A login wall is a working link**. A link that redirects to a sign-in page, like a social profile that sends logged-out visitors to its login screen, resolves: what is behind it needs an account. It is reported `ok` with the diagnosis `login_required`; the sign-in page itself is never requested.
- **The whole redirect chain**. Up to 10 redirects are followed per link, with no cookies sent. `redirect_chain` lists each hop with its status, and `final_url` and `redirect_count` are read, not assumed. A chain that never reaches a page is a free `error` row with the diagnosis `redirect_loop`.
- **Three honest verdicts and a diagnosis**. `broken` when the site says the link is dead (404, 410, a 4xx that is not a refusal, a 5xx that stood on re-check). `ok` on a healthy answer, `error` when it could not be judged. The reason sits in `error_message`, empty on every link the site answered for, and `diagnosis` says what happened in one word, from `not_found` and `server_error` to `timeout`, `dns_error` and `slow`.
- **Internal or external**. `is_external` marks each link that points off the site it was found on, so your own dead pages and dead outbound links sort apart. `checkExternal` off keeps the check to your own site.
- **Links, and assets when you ask**. `checkAssets` adds the images, scripts and stylesheets each page loads, each a checked row of its own.
- **Checked once across pages**. A footer link repeated across pages is checked once; `all_sources` lists every place it appeared.
- **Every entry answered, even a clean site**. `no_broken_links` means its links were already checked under an earlier entry or all came back healthy, `no_links_found` means the page carried nothing to check, and `not_reached` means the run's budget ended first. A clean site never reads as silence.
- **Unusable entries answered, not dropped**. A `queries` entry that is not a fetchable URL gets its own free `bad_url` row with the reason.
- **Ordinary browser headers on every request**. Every request, the first one included, goes out with the same ordinary browser headers, a Chrome user-agent and an accept-language. The identity is uniform: never switched, never rotated, and never swapped in after a refusal.
- **Plain HTTP, no browser, no login, no cookies**. Up to 8 links are checked at once, each on a different site, and several sites run in parallel. A 401 is one page asking for a login, so the rest of that site is still read; a 403, 429 or 999 closes the site for the run.
- **Tune the check to the server**. `requestTimeoutMs` waits longer for a slow server or less for a fast sweep. `maxConcurrency` checks fewer sites at once. `maxRedirects` stops following redirects early, or at 0 reports each redirect itself.
- **Caps you control**. `maxResultsPerQuery` limits the links checked per page; `maxItems` limits the whole run, 100 by default.

### What you can count on

- You pay only for rows the run judged; a row it could not judge is pushed as an uncharged `ITEM_STATUS` row with the reason on it.
- Every row is charged only after it is written to your dataset; a row you cannot see is never billed.
- A run that finds nothing returns one `NO_RESULTS` row that says so, never an empty dataset, and it is not charged.
- A spending limit you set stops the run cleanly: one `STOPPED_EARLY` row reports how many rows were returned and how many were not.
- You always know a short run from a broken one: every run writes a `RUN_SUMMARY` entry with `rowsFetched`, `rowsPushed`, `rowsCharged` and `duplicatesDropped`.
- If the read itself fails, the run fails with the error in the log; it never returns rows full of nulls and calls it success.
- No value is invented: a field the page does not show comes back empty, and the field table says which fields can be.

### Why this one

- The most-used alternative on this platform bills $0.001 for every link it checks (its pricing read on 2026-09-17). This Actor bills $0.0008 for the same unit, one event per link the site answered for, and a link it cannot judge is free.
- Run on the same crawler-test.com page on 2026-09-16, the most-used alternative's crawl kept going past the page it was given and returned 19 rows. This Actor, in its Pages mode, checked the page it was given and returned 11. It reads what you name, not what it can find.
- A link is never called broken on a HEAD answer alone: an unhealthy answer is asked once more with GET. A server error is re-checked once after a pause, and a refusal is never billed as broken.
- Misses are free and explained: a link that never answered, one that looped or one the site refused is an uncharged `error` row with the reason on it.
- The fill rate is measured, not promised: 30 of 40 sites it had never seen produced checked-link rows. Every one of the 40 produced a row that says what happened (2026-10-01).
- Every row names the page the link was found on, so a dead link is a fix, not a riddle.

### What data does Broken Link Checker return?

One real row from a run over the default input's three example pages:

```json
{
  "url": "https://crawler-test.com/links/not_found/foo1",
  "final_url": "https://crawler-test.com/links/not_found/foo1",
  "status": 404,
  "status_text": "Not Found",
  "classification": "broken",
  "is_broken": true,
  "source_domain": "crawler-test.com",
  "source_url": "https://crawler-test.com/links/broken_links_internal",
  "anchor_text": "Broken Internal Link 1",
  "element": "a",
  "all_sources": [
    {
      "source_url": "https://crawler-test.com/links/broken_links_internal",
      "anchor_text": "Broken Internal Link 1",
      "element": "a"
    }
  ],
  "method": "GET",
  "redirect_count": 0,
  "redirect_chain": [],
  "content_type": "text/html",
  "error_message": null,
  "duration_ms": 345,
  "checked_at": "2026-09-29T13:49:50.580Z",
  "diagnosis": "not_found",
  "is_external": false,
  "is_redirect": false,
  "is_slow": false,
  "row_type": "ROW"
}
```

Every field below sits on each row about a link or an entry, so the dead links and the healthy ones read the same. The run's own bookkeeping, the `RUN_SUMMARY` entry and the `STOPPED_EARLY` or `NO_RESULTS` row, carries the counters; a stopped or empty run's row also carries `reason`, saying why. When a check cannot produce a value the cell is left empty; nothing is guessed.

| Field | What it is |
|---|---|
| `url` | The link as it was found, made absolute. This is the row's key: one row per unique link URL. |
| `final_url` | Where the link ended after its redirects; the same as `url` when it never redirected, null when it never answered. On a link that redirects to a sign-in page, the sign-in page's address, which is not requested. |
| `status` | The HTTP code the link's final address answered with. On a link that redirects to a sign-in page, the redirect's own code. On a free status row this cell carries the miss word instead; `error`, `bad_url` and the rest are listed below. |
| `status_text` | The server's status line, like `Not Found`; empty when the server sent none. On an `error` row where the site answered (a refusal, a redirect loop) it keeps the code, like `HTTP 403 Forbidden`; when nothing answered, the failure reason. When a first answer was a server error and a re-check settled the link, it names the transient error too. |
| `classification` | The verdict on the link: `ok` on a healthy answer (a link that redirects to a sign-in page included), `broken` when the site says the link is dead, `error` when it could not be checked. |
| `is_broken` | true when the site says the link is dead (a 4xx other than a refusal, or a 5xx answered again on re-check after a pause, confirmed with GET), false on a healthy answer, null when the link could not be checked or the site refused the checker. |
| `diagnosis` | What happened, in one word: `ok`, `redirected`, `slow`, `redirect_unfollowed`, `login_required` (the link redirects to a sign-in page: it works, and what is behind it needs an account), `not_found`, `gone`, `client_error`, `server_error`, `blocked_by_site`, `redirect_loop`, `timeout`, `dns_error`, `tls_error` or `connection_error`. Null on an entry that was never checked. |
| `source_url` | The page this link was first found on; for a sitemap-listed page that failed, the sitemap that listed it. |
| `is_external` | true when the link points to another site than the page it was found on (a leading `www.` ignored), false when it stays on that site; null on an entry that was never checked. |
| `is_redirect` | true when the link redirected before its final answer, or answered with a redirect it was not asked to follow; false when it answered directly; null when it never answered. |
| `is_slow` | true when the check took longer than `slowThresholdMs`, redirects included; false when it did not; null when the link never answered. |
| `anchor_text` | The clickable text of the link on that page; an image's alt text when assets are checked. Empty when the link carried no text. |
| `element` | The HTML element the link came from: `a` for a link, `img`, `script` or `link` for an asset, `sitemap` for a page a sitemap listed. |
| `all_sources` | Every place the link appeared: one entry per anchor that carries it, so the same page linking it again lists it again. Empty on a row for an entry that could not be read. |
| `method` | `HEAD` or `GET`: which request gave the verdict. `GET` means HEAD was not healthy or failed and the link was asked again. |
| `redirect_count` | How many redirects were followed before the final answer; zero means the link answered directly. |
| `redirect_chain` | Every redirect in order: the address that answered, its status code and where it pointed. Empty when the link answered directly. |
| `content_type` | The media type the final address answered with, like `text/html` or `image/png`; null when the server sent none or never answered. |
| `error_message` | Why the link could not be judged, on an `error` row: a timeout, a refused connection, a name that does not resolve, a redirect loop, or a site that refused the checker. Null on every link the site answered for. |
| `duration_ms` | How long the check took in milliseconds, redirects included. |
| `checked_at` | The ISO timestamp of the check. |
| `source_domain` | The host of the start page the link was found on (one of the pages you named); the link's own host is in `url`. |
| `row_type` | `ROW` on every link the site answered for, `ITEM_STATUS` on an `error` row or an entry that could not be read; `NO_RESULTS` and `STOPPED_EARLY` are the run's own messages. |

The Overview tab shows the per-link fields; the All fields tab adds `redirect_chain`, `content_type`, `error_message`, `diagnosis`, `is_external`, `is_redirect`, `is_slow`, `duration_ms`, `checked_at` and `source_domain`. `error_message` is filled only on `error` rows, so on a sweep where every link answered it is empty on every row.

A free row is a miss, not a checked link; its `status` carries one of these words in place of an HTTP code:

| `status` value | On which rows | What it means |
|---|---|---|
| An integer HTTP code, like `200` or `404` | `ROW` | The status the link's final address answered with. |
| `error` | `ITEM_STATUS` | The link, or a start page you named, could not be judged: it never answered (a timeout, a refused connection, a name that does not resolve), it redirected in a loop, or the site refused the checker. `classification` is `error`, `diagnosis` says which, the reason sits in `error_message`, and the row is free. |
| `bad_url` | `ITEM_STATUS` | The `queries` entry was not a usable URL. The entry is named in `url`, the reason sits in `reason`, and the row is pushed free: a malformed entry is answered, never dropped. |
| `not_found` | `ITEM_STATUS` | Sitemap mode found no sitemap listing a page for this site. The entry is in `url`, what was tried is in `reason`, and the row is free. A different scope from the `diagnosis` `not_found` on a link row, which is the link's own 404. |
| `blocked` | `ITEM_STATUS` | The page was not read, in any mode, and `reason` says why: the site's robots.txt does not allow it, the site refused the checker earlier in the same run, it is a sign-in page, its site is on this Actor's exclusion list, the run had already read 2,000 pages, or the site answered the read with a script gate, a challenge stub such as an Incapsula or Cloudflare interstitial, instead of the page. A gate is named for what it is and never retried past. Free. |
| `no_broken_links` | `ITEM_STATUS` | The site answered and there was nothing new to report: every link it carried was already checked earlier in the run, or every checked link came back healthy. The entry is named in `url`, the detail sits in `reason`, and the row is free. |
| `no_links_found` | `ITEM_STATUS` | The page answered but carried nothing this run could check: not a readable HTML page, no links a plain read can see, or links all outside the scope you asked for. A page that answered with a recognised script gate is `blocked`, not this. The entry is named in `url`, the row is free. |
| `not_reached` | `ITEM_STATUS` | The run's row cap or page budget ended before this entry could be read: nothing was fetched for it and nothing was billed. |

You are billed for `ROW` rows, one per link the site answered for. `ITEM_STATUS` rows are pushed so you see them but never billed: a link that could not be judged, or an entry that could not be read. `NO_RESULTS` means no page produced a checkable link. `STOPPED_EARLY` means your spending limit ended the run. Status rows also carry the run's bookkeeping:

| Field | What it is |
|---|---|
| `reason` | On a `bad_url`, `not_found`, `blocked`, `no_broken_links`, `no_links_found` or `not_reached` row, why that entry was not read or what its read found. On `NO_RESULTS` or `STOPPED_EARLY`, why the run returned no link rows or stopped early. An `error` row carries its reason in `error_message` instead. |
| `rowsFetched` | How many link rows the check produced, counted before duplicates were dropped and any cap applied; on status rows only. |
| `rowsReturned` | On a status row: the link rows in the dataset. On `STOPPED_EARLY`, the rows returned before the spending limit stopped the run; 0 on an empty-result row. |
| `rowsRemaining` | On a status row: the rows not returned when the spending limit stopped the run (`STOPPED_EARLY`), and 0 on an empty-result row. |

### HTTP status code cheat sheet

What the `status` on a checked link means, and what to do about it. Redirects (`301`, `302`, `303`, `307`, `308`) are followed. `status` is the code the final address answered with. Each redirect's own code sits in `redirect_chain`:

| Status | Verdict | What it means | What to do |
|---|---|---|---|
| `200` | `ok` | The page answered normally. | Nothing. |
| `204`, `206` | `ok` | Answered with no content, or part of a file. | Usually nothing. |
| `301`, `308` in `redirect_chain` | set by the final status | Moved permanently. `final_url` is where it ended. | Update the link to `final_url` so visitors skip the hop. |
| `302`, `303`, `307` in `redirect_chain` | set by the final status | Moved temporarily. | Fine for a login or a tracking link; check a long chain. |
| A `3xx` in `status` | `ok` | Not followed to its end: the redirect named no new address, or the checker does not follow that code, such as `300` or `304`. `diagnosis` is `redirect_unfollowed`. | Open `final_url` by hand. |
| A `3xx` in `status`, diagnosis `login_required` | `ok` | The link redirects to a sign-in page, shown in `final_url`. The link works; what is behind it needs an account. Not broken. The sign-in page is not requested. | Nothing, unless the page should be public. |
| `error`, diagnosis `redirect_loop` | `error` | More than 10 redirects without reaching a page. `status_text` keeps the last code, `redirect_chain` shows the hops. The row is free. | Open the link in a browser; a loop there too needs fixing. |
| `400` | `broken` | The server rejected the request as malformed. | Check the URL for typos or bad characters. |
| `401`, `403`, `999` | `error` | The site refused the checker or wants a login. That does not say the link is dead, so the row is free, with `diagnosis` `blocked_by_site` and the code in `status_text`. | Open it in a browser; fine if the page is private. |
| `404` | `broken` | Not found. | Fix the link, or redirect the old address. |
| `410` | `broken` | Gone on purpose. A site can also answer 410 to a checker while still serving the page to a browser. | Remove the link; open it in a browser first if it matters. |
| `429` | `error` | Too many requests. The checker does not ask again, and asks that site nothing more in the run, so the row is free, with `diagnosis` `blocked_by_site`. | Re-run later, or with `maxPages` lower. |
| `500`, `502`, `503`, `504` and other 5xx | `broken` | The server failed or was down when checked, and it is not believed on one answer: the link is asked once more after a pause. A healthy re-check comes back `ok` with the first error named in `status_text`; a refusal on the re-check is a free `error` row. Only a 5xx that stands, or a re-check that never answers, lands `broken`. | Re-run later; a link that stays 5xx is broken. |
| `error` (no HTTP code) | `error` | No answer: a timeout, a refused connection, a name that does not resolve. The row is free. | Read `error_message`; re-run to rule out a passing outage. |

### How much does it cost?

Every checked-link row costs $0.0008, the `link-checked` event, charged only after the row is written. Apify also bills its own `apify-actor-start` event once per run, $0.00005 at this Actor's memory size. Status rows, miss rows (`error`, `bad_url`, `not_found`, `blocked`, `no_broken_links`, `no_links_found`, `not_reached`) and dropped duplicate links are free. You pay for every link the site answered for, `ok` and `broken` alike. A link is billed once however its check ran: HEAD alone, HEAD then GET, or a server error asked once more after a pause.

| Checked-link rows | Link charges | Start event | Total |
|---|---|---|---|
| 100 | $0.08 | $0.00005 | about $0.08 |
| 1,000 | $0.80 | $0.00005 | about $0.80 |
| 10,000 | $8.00 | $0.00005 | about $8.00 |

To budget a sweep of your own sites, the measured run gives an estimate. On 40 sites it had never seen, 80% answered and 30 produced link rows, about 157 checked-link rows each (measured 2026-10-01). At that rate, and with `maxItems` raised past the default 100:

- 100 sites: about 80 answer and about 75 produce link rows, about 11,745 checked-link rows, about $9.40 plus the start event.
- 1,000 sites: about 800 answer and about 750 produce link rows, about 117,450 checked-link rows, about $93.96 plus the start event.
- 10,000 sites: about 8,000 answer and about 7,500 produce link rows, about 1,174,500 checked-link rows, about $939.60 plus the start event.

Your sites will differ, and `error` rows are free, so a real sweep can come in under those totals.

On a paid Apify plan the per-row price steps down with your tier: $0.00067 on Bronze, $0.00054 on Silver and $0.00043 on Gold and above. That is $0.67, $0.54 and $0.43 per 1,000 rows.

Duplicates are billed once, not once per sighting. If 50 link sightings across your pages deduplicate to 47 unique URLs, you pay for 47 checks; the repeats join `all_sources` free.

### How do I use Broken Link Checker?

1. Open the Actor and press **Try for free**.
2. Paste your pages or sites into **Pages or sites to check**, one URL or bare domain per line. The input carries three example pages from a public test site; replace them with yours.
3. Pick a **Mode**: Pages for just those pages, Crawl for the whole site from each page, Sitemap for every page the sitemap lists.
4. Set **Max pages** and **Maximum items** if you need to; every link the site answers for comes back as a billed row carrying its verdict.
5. Press **Start**. Rows land in the dataset as links are checked, and the run ends when every page is read.

Example input:

```json
{
  "queries": ["https://crawler-test.com/links/broken_links_internal", "https://crawler-test.com/links/broken_links_external", "https://example.com/"],
  "maxItems": 100
}
```

#### Worked examples

Use it to get every link on a few pages you just published, healthy ones included.

```json
{
  "queries": ["https://crawler-test.com/links/broken_links_internal", "https://crawler-test.com/links/broken_links_external"],
  "mode": "pages"
}
```

Use it to sweep a whole site and get every checked link back with its verdict.

```json
{
  "queries": ["crawler-test.com"],
  "mode": "crawl",
  "maxPages": 200,
  "maxItems": 1000
}
```

Use it to check every page in a sitemap, the site's own links only, as a technical SEO audit.

```json
{
  "queries": ["https://crawler-test.com/test_sitemap.xml"],
  "mode": "sitemap",
  "maxPages": 500,
  "maxItems": 5000,
  "checkExternal": false
}
```

Use it to find missing images, scripts and stylesheets on a landing page.

```json
{
  "queries": ["https://crawler-test.com/"],
  "checkAssets": true
}
```

Or start a run over the API:

```
curl -X POST "https://api.apify.com/v2/acts/Pradio~broken-link-checker/runs?token=YOUR_APIFY_TOKEN" -H "Content-Type: application/json" -d "{\"queries\":[\"https://example.com/\"]}"
```

### Input

| Input | Default | What it does |
|---|---|---|
| `queries` | three example pages | The pages or sites to check, one URL or bare domain per line, at most 2,000. An entry that is not a fetchable URL gets its own free `bad_url` row with the reason, so a typo is answered, never dropped. |
| `mode` | `pages` | `pages` reads only the pages you list. `crawl` starts at each and follows links on the same site. `sitemap` reads the pages the site's sitemap lists. |
| `maxPages` | 50 | In `crawl` and `sitemap` modes, the most pages one run reads for links, across all entries. `pages` mode reads the pages you list, never more than 2,000 in one run. |
| `checkExternal` | true | Check links to other sites too. Off keeps the check to the site you named. |
| `checkAssets` | false | Also check the images, scripts and stylesheets each page loads. |
| `slowThresholdMs` | 5000 | A healthy link whose check takes longer than this, in milliseconds, gets the diagnosis `slow` and `is_slow` true. It changes no verdict and no charge. |
| `requestTimeoutMs` | 15000 | How long one link check waits for an answer, in milliseconds, from 1,000 to 60,000. A link that does not answer in time is a free `error` row with the diagnosis `timeout`. Pages read for links always wait at least 15 seconds. |
| `maxRedirects` | 10 | How many redirects a link check follows, from 0 to 10. At 0 each redirect is reported itself with the diagnosis `redirect_unfollowed`; a lower cap reports where it stopped. Only a chain past 10 is a `redirect_loop`. |
| `maxConcurrency` | 8 | How many links are checked at the same time, from 1 to 8, each on a different site: one site never gets more than one request at a time. Lower it to check fewer sites at once. |
| `maxResultsPerQuery` | none | The most new links checked from one page read. Unset means every link found is checked. |
| `maxItems` | 100 | The most link rows one run returns in total. Rows past the cap are dropped, and the run summary shows how many links were checked. |

#### queries

Each entry is a full URL like `https://example.com/blog`, or a bare domain like `example.com`, which is fetched over HTTPS. In `pages` mode the page is fetched once, its anchors are collected and each unique link is checked. An entry nothing can be fetched for, such as a typo or an unrecognisable line, gets its own uncharged row. That row carries `status` `bad_url`, the entry in `url` and the reason in `reason`.

#### mode

- **`pages`** (the default) reads exactly the pages you list and nothing else, at most 2,000 in one run. It reads each site's robots.txt first: a page robots.txt disallows, or one answered by a script gate, is not read and comes back as a free `blocked` row.
- **`crawl`** reads each entry, then the pages on the same site it links to, then theirs, breadth first, until `maxPages` pages are read or `maxItems` rows are found. It never leaves the site, reads the site's robots.txt first, skips any page robots.txt disallows and waits the crawl delay it asks for. Links to disallowed pages are still checked with one request each; their own links are not read.
- **`sitemap`** reads the site's sitemap: the entry itself when it is a `.xml` sitemap URL, else the sitemaps robots.txt names, else `/sitemap.xml`. A sitemap index is followed into its child sitemaps on the same site. Up to `maxPages` listed pages are read for links, and a listed page that fails is itself a row, with the sitemap as its `source_url`. A site with no readable sitemap gets one free `not_found` row.

Raise `maxItems` for a crawl: at the default of 100 rows a crawl stops at the first hundred links.

```json
{
  "queries": ["https://example.com/"],
  "mode": "crawl",
  "maxPages": 50,
  "maxItems": 2000
}
```

### Output

Rows land in the default dataset as links are checked. The default input's three example pages return 12 link rows. A page whose checked links all come back healthy also gets a free `no_broken_links` row saying so. Four extras tell you how a run went:

- `ITEM_STATUS`: a link or start page that never answered, looped or was refused. It carries `status` `error`, the kind in `diagnosis` and the reason in `error_message`. The same kind marks a `queries` entry that could not be read or had nothing to report. It is named in `url` with `status` `bad_url`, `not_found`, `blocked`, `no_broken_links`, `no_links_found` or `not_reached` and the reason in `reason`. It is pushed so you see it, and never billed.
- `NO_RESULTS`: no start page produced a checkable link. One uncharged row carries `reason` and `rowsFetched`, so an empty answer is an answer, not silence.
- `STOPPED_EARLY`: your spending limit ended the run. One uncharged row carries `rowsReturned` and `rowsRemaining`; raise the limit and re-run for the rest.
- `RUN_SUMMARY`: an entry in the run's key-value store carrying the counts: links fetched, rows pushed, rows charged, duplicates dropped, stopped early or not.

A start page that cannot be reached is itself a free `error` row with the reason in `error_message`. A run over pages that are all down still tells you so, page by page, and costs only the start event.

### What can you do with the data?

**Sweep your own site before a launch**. Run the checker over the pages, filter `classification` to `broken`, and hand the list to whoever fixes it. Each row already carries the dead URL, the page it sits on and the anchor text to search for.

**Audit a migration**. Old URLs that still redirect show every hop in `redirect_chain` and where they landed in `final_url`. The ones that answer 404 instead are the ones to remap.

**Watch outbound links on a schedule**. Reference, partner and affiliate links rot quietly. Put the run on an Apify schedule, export the broken rows to a sheet, and the dataset becomes a monthly fix list.

**Qualify a site you are evaluating**. A page full of dead links says something about how it is maintained. One run counts them without a manual click-through.

### Use Broken Link Checker with AI agents

Paste this line to give an agent this Actor through Apify's MCP server:

```
claude mcp add --transport http apify "https://mcp.apify.com?tools=Pradio/broken-link-checker"
```

### Personal data

- A row carries only the declared link-health fields: the checked URL, its HTTP status and verdict, the anchor text and the page it appeared on. No field names or identifies a person.
- A row keeps facts and the link's own anchor text, cut to at most 200 characters, never the body or a substantial part of a page.
- `mailto:` and `tel:` links are never requested, and the email addresses and phone numbers in them are never recorded. `javascript:` links are never requested either.
- Every row links back to the page it was found on, so what was collected is easy to check.
- A page that refuses the read is reported as a free `error` row with the diagnosis `blocked_by_site`. A link whose HEAD request is refused (401, 403, 429 or 999) is not asked again: the refusal is its answer. A link that redirects to a sign-in page is reported as working; the sign-in page is never requested. An error code that is not a refusal is asked once with an ordinary GET. A server error is asked once more after a pause, at most one re-check per link. Every request, the first one included, carries the same ordinary browser headers, a Chrome user-agent and an accept-language. It is never switched or rotated, and never after a refusal. A page that answers the read with a script gate instead of the page is reported `blocked`, free. The verdict is a report, never a trigger to re-request with different headers, a proxy or a solver. The read is never retried past the gate. Once a site refuses the checker (403, 429 or 999), nothing more is asked of that site in the run. One request at a time goes to any one site. Nothing retries past a block or works around a refusal, and there is no proxy.
- Personal data is not what a row is about. It can still appear incidentally inside anchor text or a URL path, such as a name in a link label or a profile slug in a path. Run it against sites you operate or are permitted to check.
- You choose the pages and sites, so you are the controller for the list you supply. In `pages` mode the Actor reads only what you list. In `crawl` and `sitemap` modes it reads further pages of the same site, never another site. In every mode it never reads a page the site's robots.txt disallows. To exclude a section, disallow it in robots.txt.
- A site owner can ask for their site to be left out entirely. Open an issue on this Actor's Issues tab naming the domain, and it goes on the Actor's exclusion list. Every later run reads that list before any request and sends the site nothing. A page on it comes back as a free `blocked` row, a link to it as a free `error` row.

### Release notes

- 2026-10-01: every request now carries ordinary browser headers, a Chrome user-agent and an accept-language, from the first request on; the identity is never switched after a refusal. A link whose check ends in a server error is asked once more after a pause before it is called `broken`; a healthy re-check comes back `ok`. The transient first answer stays named in `status_text`. And a site that answers a read with a script gate is reported as a free `blocked` row. It names the gate (an Incapsula or Cloudflare challenge stub instead of the page) and is never called a page with no links.
- 2026-10-01: every entry in the list now lands a row of its own. An entry whose links were already checked under another entry, or that carried nothing to check, gets a free row saying so; before, a quiet entry could come back with no row when entries ran at the same time.
- 2026-09-29: billing moves to per link checked. The `link-checked` event bills every link the site answered for, `ok` and `broken` alike, and every checked link lands a row carrying its verdict. The `onlyBroken` input is gone: billed rows and dataset rows are the same set. Entries on different sites now run in parallel; each site still gets one request at a time at its own pace.
- 2026-09-29: an entry that answered with nothing to report now says so on a free row. `no_broken_links` means every checked link was healthy, `no_links_found` means the page carried nothing to check, `not_reached` means the run's row cap or page budget ended first.
- 2026-09-28: new `requestTimeoutMs`, `maxRedirects` and `maxConcurrency` inputs to tune the check to the server, and new `is_redirect` and `is_slow` columns. `onlyBroken` is added, on by default, so a run can return and bill only the broken links (superseded by the 2026-09-29 billing change, which removed it). Anchor text is kept to a short label of at most 200 characters. One request at a time goes to any one site. A 429 is no longer asked again after a wait, and a site that refuses the checker is asked nothing more in the run. Pages mode now reads robots.txt too, and reads at most 2,000 pages in one run. A robots.txt crawl delay now paces every request to its site, link checks included. No cookies are sent: a redirect that loops without one is a free `redirect_loop` row. A page that answers with a sign-in page is not read, and a link that redirects to one is reported as working, with the new diagnosis `login_required`. Site owners can ask to be excluded. Existing columns and inputs keep their names.
- 2026-09-27: a link is never called broken on a HEAD answer alone. An error code on HEAD that is not a refusal, or a failure a GET could answer, is asked again with GET, and the GET decides. A DNS or TLS failure is not re-asked: no method reaches it. A refusal on HEAD is not asked again. A site that refuses the checker (401, 403, 429, 999, a sign-in redirect) is now a free `error` row, and so is a redirect loop. Commented-out links in a page are no longer read. New `diagnosis` and `is_external` columns and a `slowThresholdMs` input. Existing columns and inputs keep their names.
- `0.1.22` (2026-09-25): an `error` row, a link or start page that never answered, is now free; `status` carries `error` on it and `row_type` is `ITEM_STATUS`.
- `0.1.20` (2026-09-24): `mode` adds `crawl` and `sitemap` beside `pages`; new `checkExternal`, `checkAssets` and `maxPages` inputs; new `redirect_chain`, `content_type` and `error_message` columns. Existing inputs work as before.

### Limits

- It checks links and reports them; it does not fix links: repair stays with you.
- A crawl stays on the start page's site and reads at most `maxPages` pages; a site larger than that is read in part. Pages robots.txt disallows are not read, so links that only appear on them are not found.
- Anchor tags are checked by default; images, scripts and stylesheets only with `checkAssets`. `javascript:`, `mailto:`, `tel:`, `data:`, `sms:` and `ftp:` links are skipped: they are not requested, and no email address or phone number from them reaches a row.
- Links a page builds with JavaScript after load are not seen: pages are fetched over plain HTTP, with no browser. A page that answers with a recognised script gate, an Incapsula or Cloudflare challenge stub, is not mistaken for an empty one. The entry comes back a free `blocked` row that names the gate; the read is never retried past it. Any other page whose links exist only after scripts run reads as `no_links_found`.
- Requests are logged-out and public, sent with no cookies and through no proxy. Every request, the first one included, carries the same ordinary browser headers, a Chrome user-agent and an accept-language. That identity is uniform: it is never switched, rotated or swapped in after a refusal. A site can still answer automated traffic with a dead status it would not serve a signed-in browser. That row lands `broken` on the site's own answer, kept in `status_text`, and a browser could disagree with it. A redirect that only reaches its page when a cookie is sent back reads as a free `redirect_loop` row. There is no retry past a block: a site that refuses is a free `error` row with the diagnosis `blocked_by_site`, not a workaround. A 429 is the answer, never asked again, and a site that refused is asked nothing more in that run.
- A start page that cannot be read is its own row, and it never disappears from the results. It is an `error` row with the reason in `error_message` when nothing answered or the site refused. It is a `broken` row with the status code when it answered a 4xx that is not a refusal or a 5xx; the pause re-check is for link checks, not page reads.
- A certificate the checker cannot verify is an `error` row with the diagnosis `tls_error`, even where a browser completes the chain on its own.
- A link that never answers within the request timeout, 15 seconds by default, is an `error` row, suspected but not proven broken.
- The verdict is the HTTP answer. A page that answers `200` with a "not found" message of its own (a soft 404) reads `ok`, whatever its link text says.
- A redirect chain is followed for at most 10 redirects; a longer chain is a free `error` row with the diagnosis `redirect_loop`. `maxRedirects` can lower the cap, never raise it.
- Sparse fields, measured on the 4,698 rows from the sites it had never seen: `anchor_text` was filled on 87.5% (an anchor with no text leaves it empty). `content_type` was filled on 99.8% (a server that sends no media type), `status_text` on 99.5%. `redirect_chain` is filled only when a redirect happened, which was 7.9% of rows.
- Every link the site answered for is a billed row, `ok` or `broken`. An `error` row is free, so a run of nothing but unanswered or refused links costs only the start event.

### Troubleshooting

**The run returned fewer rows than `maxItems`.**
`maxItems` is a ceiling, not a target. Fewer rows means the pages ran out of links first; the run log shows how many links were checked.

**The dataset has one row saying `NO_RESULTS`.**
No start page produced a checkable link: a clean result. That is the answer to this run, not a bug, and the row is free.

**A row says `no_broken_links` or `no_links_found` in `status`.**
The site answered. `no_broken_links` means its links were already checked under an earlier entry or all came back healthy. `no_links_found` means the page carried nothing this run could check. A site that answered with a recognised script gate comes back `blocked` instead, not `no_links_found`. All are complete answers, and all such rows are free.

**An entry comes back `not_reached` or `blocked`.**
`not_reached` means the run's row cap or page budget ended before that entry was read: nothing was fetched for it and nothing billed. `blocked` means the page was not read, and `reason` says why: robots.txt disallows it, the site refused the checker earlier in the run, it is a sign-in page, or its site is on the exclusion list. A site that answered with a script gate instead of the page is also `blocked`. Both rows are free.

**A row says `bad_url` in `status`.**
One `queries` entry was not a URL the Actor could fetch, a typo or an unrecognisable line. The row names the entry in `url`, gives the reason in `reason`, and is free. The other entries ran normally.

**A row says `error` with `is_broken` empty.**
The link could not be judged, and `diagnosis` says why: `timeout` (no answer within the request timeout), `dns_error`, `tls_error`, `connection_error`, `redirect_loop` or `blocked_by_site`. It is not proven dead, and the row is free; re-run later to retry it.

**An `ok` row's `status_text` names a server error, like HTTP 503.**
The first answer was a transient server error and the re-check after a pause answered healthy, so the link is fine. The row keeps the first answer for you to see; only a 5xx that stands on re-check is called `broken`.

**A link I can open in my browser shows `blocked_by_site`.**
The site refused this checker's plain request while it serves your browser. The row is free and is not counted as broken. The checker does not work around a refusal.

**A link to a social profile shows `ok` with the diagnosis `login_required`.**
The link redirected to a sign-in page, as Instagram and others do for logged-out visitors. The link works and the page behind it needs an account, so it is not broken. The sign-in page is not requested.

**Every row is an `error` row for a start page I listed.**
None of the pages you listed answered at the network level. Each row names its page in `url` and the reason in `error_message`: a timeout, a refused connection, a name that does not resolve. Check the URLs you pasted.

### FAQ

**Can I use integrations with Broken Link Checker?**
Yes. The dataset plugs into Apify integrations like Zapier, Make and webhooks, and Apify schedules can re-run the sweep weekly or monthly without you touching it.

**Can I use Broken Link Checker with the Apify API?**
Yes. Start runs and read the dataset back over the API; the curl example above is the whole call.

**Can I use Broken Link Checker through an MCP server?**
Yes. The line under "Use Broken Link Checker with AI agents" adds it to an agent through Apify's MCP server.

**Will it call a link broken that still works?**
Not on a first answer alone. A link whose HEAD request comes back with an error code is asked once more with GET, the request a browser makes, and the GET decides. A server error is asked once more after a pause before it counts: a healthy re-check comes back `ok`. A refused request (401, 403, 429 or 999) is a free `error` row with the diagnosis `blocked_by_site`, never a billed `broken`. So is a link that never answered. A site that answers dead statuses to automated traffic can still produce a `broken` row a browser would disagree with; the row keeps the site's own answer in `status_text`. The one case it cannot see is a page that answers `200` with its own "not found" message; that reads `ok`.

**Can I use it for scheduled link monitoring?**
Yes. Save your input as a task and put it on an Apify schedule, daily, weekly or monthly. Each run is a fresh sweep, and the dataset holds every link the sites answered for at that moment, each billed as a checked link.

**How do I export results to CSV or JSON?**
Open the run's dataset and export it as CSV, JSON, Excel, XML or HTML, or read it over the API. The Overview tab is the view most exports want; All fields adds the timing and redirect detail.

**How many URLs can I check?**
As many pages as you paste into **Pages or sites to check**, up to 2,000 entries. `maxItems` caps the link rows one run returns, 100 by default, so raise it for a big sweep. Crawl and Sitemap modes read at most `maxPages` pages per run. No mode reads more than 2,000 pages in one run, and every mode stops reading a site that refuses a request.

**Do I need an API key or login?**
No API key and no login on the sites you check: it reads public pages with plain requests. You need only your Apify account to run it. Pages behind a login are not read; a link that sends the checker to a sign-in page is reported as working, with the diagnosis `login_required`.

**Is it legal to check links this way?**
The Actor is for checking links. It makes plain HTTP requests that carry ordinary browser headers, the same on every request from the first; it bypasses no login and keeps only link facts. It honours robots.txt and its crawl delay in every mode, sends one request at a time to any one site and stops on a refusal. It does not claim a site's terms permit the checks; that permission stays with you. You choose the pages, so point it at sites you are responsible for or allowed to test. A site owner can ask to be excluded on the Issues tab, and an excluded site is sent no request at all. You are responsible for having the right to check the pages and sites you give it, and for how you use the rows. It makes no promise to get past blocks, logins or CAPTCHAs: a page that refuses it is reported, never worked around. A page that answers with a script gate instead of the page is reported `blocked`, never retried past.

### Feedback

Found a problem or a missing field? Open an issue on the Issues tab; it is answered within two days. If this Actor saved you time, a review on the Store helps other buyers find it.

### Not affiliated

Broken Link Checker is an independent tool, not affiliated with or endorsed by any site you point it at. The example pages in the default input are public test pages on crawler-test.com and example.com.

# Actor input Schema

## `queries` (type: `array`):

The pages or sites to check: a full URL or a bare domain per line, at most 2,000. In Pages mode each is read once and every link on it is checked. In Crawl mode each is a start page. In Sitemap mode each names a site (or a sitemap .xml file). Every mode reads the site's robots.txt first and never reads a page it disallows.

## `mode` (type: `string`):

Which pages a run reads. Pages: only the pages you list. Crawl: each listed page and the pages on the same site it links to, breadth first, up to Max pages. Sitemap: the pages the site's sitemap lists, up to Max pages. Every mode honours the site's robots.txt and its crawl delay.

## `maxPages` (type: `integer`):

In Crawl and Sitemap modes, the most pages one run reads for links, across every entry (at most 2,000). Pages mode reads the pages you list, never more than 2,000 in one run. A site that refuses any request (403, 429 or 999) is read no further in the run.

## `checkExternal` (type: `boolean`):

Check links that point to other sites as well as links on the same site. Turn off to check only links on the site you named.

## `checkAssets` (type: `boolean`):

Also check the images, scripts and stylesheets each page loads, not only its links. Each one found is a checked row.

## `slowThresholdMs` (type: `integer`):

A healthy link whose check takes longer than this many milliseconds, redirects included, gets the diagnosis slow. It changes no verdict and no charge.

## `requestTimeoutMs` (type: `integer`):

How long one link check waits for an answer, in milliseconds, before the link comes back as a free error row with the diagnosis timeout. Raise it for slow servers, lower it for a fast sweep. Pages read for links always wait at least 15 seconds.

## `maxRedirects` (type: `integer`):

How many redirects a link check follows, at most 10. Set 0 to report each redirect itself without following it (the diagnosis redirect\_unfollowed). A link that still redirects after a lower cap is reported the same way; only a chain past 10 is a redirect loop.

## `maxConcurrency` (type: `integer`):

How many links are checked at the same time, at most 8, each on a different site: one site never gets more than one request at a time. Lower it to check fewer sites at once.

## `maxResultsPerQuery` (type: `integer`):

The most new links checked from one page read; leave unset to check every link the page carries.

## `maxItems` (type: `integer`):

The most link rows one run returns in total; rows past the cap are dropped and the run summary shows how many links were checked.

## Actor input object example

```json
{
  "queries": [
    "https://crawler-test.com/links/broken_links_internal",
    "https://crawler-test.com/links/broken_links_external",
    "https://example.com/"
  ],
  "mode": "pages",
  "maxPages": 50,
  "checkExternal": true,
  "checkAssets": false,
  "slowThresholdMs": 5000,
  "requestTimeoutMs": 15000,
  "maxRedirects": 10,
  "maxConcurrency": 8,
  "maxItems": 100
}
```

# Actor output Schema

## `rows` (type: `string`):

The checked-link rows for this run, one row per unique link found on the pages given.

## `summary` (type: `string`):

Counts for this run: links fetched, rows pushed, rows charged, duplicates dropped, stopped early.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("pradio/broken-link-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("pradio/broken-link-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call pradio/broken-link-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pradio/broken-link-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pDKqvReedkPnaPEhA/builds/tDAxDu1AZksZbbRxq/openapi.json
