Broken Link Checker — Find 404s & Dead Links avatar

Broken Link Checker — Find 404s & Dead Links

Pricing

from $1.00 / 1,000 page scanneds

Go to Apify Store
Broken Link Checker — Find 404s & Dead Links

Broken Link Checker — Find 404s & Dead Links

Crawl a website and check every link it points to — pages, images, scripts, stylesheets, internal and external. Get status code, redirect chain and the page each link was found on.

Pricing

from $1.00 / 1,000 page scanneds

Rating

0.0

(0)

Developer

KeyMan98

KeyMan98

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Share

Broken Link Checker — Find 404s & Dead Links on Any Website

Crawl a website and check every link it points to — pages, images, scripts, and stylesheets, internal or external — and get a clean report of what's broken: 404s, server errors, timeouts, DNS failures. Plain HTTP requests only, no browser.

What it does

Starting from the URL(s) you give it, the Actor crawls the site (following only links on the same host, up to a page limit you set) and collects every link it finds along the way: page links (<a>), images, scripts, and stylesheets. Each unique link is then checked once — is it reachable, does it return an error, where does it redirect to — and written to the dataset.

Input

  • Start URLs — pages to start crawling from.
  • Max pages to scan — how many same-site HTML pages to crawl (default 50). This is also what you're charged for.
  • Check external links — when on (default), links pointing to other websites are checked too, not just the site being crawled.
  • Only output broken links — when on (default), the dataset only contains broken links. Turn off to get a row for every link checked, working or not.
  • Max concurrency — how many pages/links are checked in parallel (default 10).

Output example (JSON)

{
"url": "https://example.com/old-page.html",
"sourcePage": "https://example.com/",
"sourcePagesCount": 1,
"linkText": "Read more",
"linkType": "a",
"isExternal": false,
"statusCode": 404,
"finalUrl": "https://example.com/old-page.html",
"redirectChain": [],
"isBroken": true,
"note": null,
"error": null,
"checkedAt": "2026-09-24T12:00:00+00:00"
}
  • url — the link that was checked.
  • sourcePage — first crawled page where this link was found. null for a start URL.
  • sourcePagesCount — number of distinct crawled pages this link was found on.
  • linkText — visible text of the link (up to 200 characters). Empty for images/scripts/stylesheets.
  • linkType — a (page link), img, script, or css (stylesheet).
  • isExternal — true if the link points to a different host than the site being crawled.
  • statusCode — HTTP status received, or null if the link could not be reached at all.
  • finalUrl — URL after following redirects, or null if unreachable.
  • redirectChain — URLs visited before reaching finalUrl, in order.
  • isBroken — true for an error status (4xx/5xx) or a link that could not be reached at all. A link still returning 429 (rate limited) after retries is not counted as broken — see note.
  • note — set to "rate limited (429)" when the link is still answering 429 after retries; that is a server asking to slow down, not a dead link, so isBroken stays false. null otherwise.
  • error — network error (DNS, timeout, SSL, refused URL), or null if the link responded.
  • checkedAt — UTC timestamp of the check.

Pricing

Pay only for pages actually crawled — checking the links found on them costs nothing extra. Pricing model: pay-per-event.

EventWhen it's chargedPrice
page-scanneda same-site HTML page was successfully crawled and scanned for links0.001 USD

Example: a 50-page site costs 0.05 USD to crawl, however many links (broken or not, internal or external) it turns up.

Not charged: a page that could not be reached, a non-HTML page, or any individual link check.

Limitations

  • Same-host crawl only. example.com and www.example.com count as the same host; a link to any other host is only checked, never crawled further.
  • JavaScript-rendered links are not seen. No browser is used; links only added to the page by client-side JavaScript won't be found.
  • robots.txt is honored for crawling (which pages get followed to find more links), not for checking a link that was already found. If robots.txt itself cannot be reached, crawling is allowed.
  • Login-gated sites are not supported. No authentication of any kind.
  • Up to 5,000 unique links are tracked per run, regardless of maxPages — a handful of pages can reference far more assets than that.
  • Internal/private/loopback addresses and non-http(s) URLs (e.g. file://) are refused before any request is made.

FAQ

Does this use a headless browser?

No — plain HTTP requests only, which keeps runs fast and cheap. Links that only appear after JavaScript runs are not found.

Yes, by default — every link found, including ones to other websites, is checked (unless "Check external links" is off). Only links on the site being crawled are followed to find more pages.

What if a page requires a login?

Not supported. The Actor makes plain, unauthenticated requests.

Does it respect robots.txt?

Yes, for deciding which pages to crawl. It does not skip checking a link just because robots.txt disallows it — the link is still reported.

Can I check a fixed list of URLs instead of crawling?

Not today. This Actor only crawls starting from the "Start URLs" you give it, following same-host links to discover pages; there's no input for a fixed URL list or a sitemap.

Can I use this through the Apify API or an MCP server?

Yes. Like any Apify Actor, you can run it and read results through the standard Apify API, or through the Apify MCP server if you use Claude, Cursor, or another MCP-enabled client.

Export

Results can be downloaded from the Apify dataset as JSON, CSV, or Excel, or accessed via the Apify API.