Broken Link Checker — Find 404s & Dead Links
Pricing
from $1.00 / 1,000 page scanneds
Broken Link Checker — Find 404s & Dead Links
Crawl a website and check every link it points to — pages, images, scripts, stylesheets, internal and external. Get status code, redirect chain and the page each link was found on.
Pricing
from $1.00 / 1,000 page scanneds
Rating
0.0
(0)
Developer
KeyMan98
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
Broken Link Checker — Find 404s & Dead Links on Any Website
Crawl a website and check every link it points to — pages, images, scripts, and stylesheets, internal or external — and get a clean report of what's broken: 404s, server errors, timeouts, DNS failures. Plain HTTP requests only, no browser.
What it does
Starting from the URL(s) you give it, the Actor crawls the site (following only links on the same host, up to a page limit you set) and collects every link it finds along the way: page links (<a>), images, scripts, and stylesheets. Each unique link is then checked once — is it reachable, does it return an error, where does it redirect to — and written to the dataset.
Input
- Start URLs — pages to start crawling from.
- Max pages to scan — how many same-site HTML pages to crawl (default 50). This is also what you're charged for.
- Check external links — when on (default), links pointing to other websites are checked too, not just the site being crawled.
- Only output broken links — when on (default), the dataset only contains broken links. Turn off to get a row for every link checked, working or not.
- Max concurrency — how many pages/links are checked in parallel (default 10).
Output example (JSON)
{"url": "https://example.com/old-page.html","sourcePage": "https://example.com/","sourcePagesCount": 1,"linkText": "Read more","linkType": "a","isExternal": false,"statusCode": 404,"finalUrl": "https://example.com/old-page.html","redirectChain": [],"isBroken": true,"note": null,"error": null,"checkedAt": "2026-09-24T12:00:00+00:00"}
url— the link that was checked.sourcePage— first crawled page where this link was found.nullfor a start URL.sourcePagesCount— number of distinct crawled pages this link was found on.linkText— visible text of the link (up to 200 characters). Empty for images/scripts/stylesheets.linkType—a(page link),img,script, orcss(stylesheet).isExternal—trueif the link points to a different host than the site being crawled.statusCode— HTTP status received, ornullif the link could not be reached at all.finalUrl— URL after following redirects, ornullif unreachable.redirectChain— URLs visited before reachingfinalUrl, in order.isBroken—truefor an error status (4xx/5xx) or a link that could not be reached at all. A link still returning429(rate limited) after retries is not counted as broken — seenote.note— set to"rate limited (429)"when the link is still answering429after retries; that is a server asking to slow down, not a dead link, soisBrokenstaysfalse.nullotherwise.error— network error (DNS, timeout, SSL, refused URL), ornullif the link responded.checkedAt— UTC timestamp of the check.
Pricing
Pay only for pages actually crawled — checking the links found on them costs nothing extra. Pricing model: pay-per-event.
| Event | When it's charged | Price |
|---|---|---|
page-scanned | a same-site HTML page was successfully crawled and scanned for links | 0.001 USD |
Example: a 50-page site costs 0.05 USD to crawl, however many links (broken or not, internal or external) it turns up.
Not charged: a page that could not be reached, a non-HTML page, or any individual link check.
Limitations
- Same-host crawl only.
example.comandwww.example.comcount as the same host; a link to any other host is only checked, never crawled further. - JavaScript-rendered links are not seen. No browser is used; links only added to the page by client-side JavaScript won't be found.
robots.txtis honored for crawling (which pages get followed to find more links), not for checking a link that was already found. Ifrobots.txtitself cannot be reached, crawling is allowed.- Login-gated sites are not supported. No authentication of any kind.
- Up to 5,000 unique links are tracked per run, regardless of
maxPages— a handful of pages can reference far more assets than that. - Internal/private/loopback addresses and non-http(s) URLs (e.g.
file://) are refused before any request is made.
FAQ
Does this use a headless browser?
No — plain HTTP requests only, which keeps runs fast and cheap. Links that only appear after JavaScript runs are not found.
Will it check links outside the site I'm crawling?
Yes, by default — every link found, including ones to other websites, is checked (unless "Check external links" is off). Only links on the site being crawled are followed to find more pages.
What if a page requires a login?
Not supported. The Actor makes plain, unauthenticated requests.
Does it respect robots.txt?
Yes, for deciding which pages to crawl. It does not skip checking a link just because robots.txt disallows it — the link is still reported.
Can I check a fixed list of URLs instead of crawling?
Not today. This Actor only crawls starting from the "Start URLs" you give it, following same-host links to discover pages; there's no input for a fixed URL list or a sitemap.
Can I use this through the Apify API or an MCP server?
Yes. Like any Apify Actor, you can run it and read results through the standard Apify API, or through the Apify MCP server if you use Claude, Cursor, or another MCP-enabled client.
Export
Results can be downloaded from the Apify dataset as JSON, CSV, or Excel, or accessed via the Apify API.