Broken Link Checker: Every 404 Billed Once
Pricing
from $0.43 / 1,000 link checkeds
Broken Link Checker: Every 404 Billed Once
Find broken links and 404s on a page, a list, a whole site or its sitemap, billed per link the site answered for. An error on HEAD is re-checked with GET, and a site that refuses the checker is never called broken. One row per link with its status, verdict and page.
Pricing
from $0.43 / 1,000 link checkeds
Rating
0.0
(0)
Developer
Pradio Actors
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
16 minutes ago
Last modified
Categories
Share
What does Broken Link Checker do?
Broken Link Checker finds the broken links on a page, a site or its sitemap and returns one row per link it checks. Each row carries the link's final HTTP status, an ok, broken or error verdict, the anchor text when it had any, the page it sat on and every redirect it took. Every link the site answered for is billed once at $0.0008. Paste URLs or bare domains, pick a mode and press Start; what comes back is your fix list.
On 40 sites it had never seen, run 2026-10-01, every one of the 40 produced a row, 80% of the sites answered and 30 produced checked-link rows. A link it could not judge is a free error row with the reason on it.
Who uses Broken Link Checker
| Buyer | What they run it for |
|---|---|
| Website owners | Finding dead links on their own pages before visitors or crawlers do. |
| SEO practitioners | A technical SEO audit in one run: dead links leak ranking, and each row names the page and the anchor to fix. |
| Teams after a migration or redesign | Checking that old URLs still redirect somewhere instead of dying. |
| Content maintainers | Auditing outbound links that rot quietly over time. |
Features
- Every 404 re-checked with GET. You never see
brokenon a HEAD answer alone: an error code that is not a refusal is asked once more with GET, the request a browser makes. The GET decides. Some servers answer HEAD with an error and serve the page to GET; those links come backok. - A 5xx is asked once more, after a pause. A server error is often a passing failure, not a dead link. Before a link is called
brokenon one, it is re-checked once after a short wait. A healthy re-check comes backokandstatus_textnames the transient first answer. Only a 5xx that stands on re-check, or one whose re-check never answers, lands asbroken. - A page, a list, a whole site or its sitemap.
modepicks what is read: only the pages you list, a crawl of each site from the page you give, or every page the site's sitemap lists. - Crawls politely, in every mode. Every mode, Pages included, reads the site's robots.txt first and never reads a page it disallows. The file is read under the checker's own name and under the browser identity its requests carry; a Disallow under either stops the read. Its crawl delay is waited between every request to that site, link checks included. No run reads more than 2,000 pages in any mode, and a site that refuses the checker (403, 429 or 999) is read no further.
- A refusal is not a dead link. When a site refuses the request (401, 403, 429 or 999) it has not said the link is dead. You get a free
errorrow with the diagnosisblocked_by_site, never a billedbroken. - A script gate is named for what it is. A site that answers the read with a script gate, an Incapsula or Cloudflare challenge stub instead of the real page, is reported on a free
blockedrow.reasonnames the gate. It is never mistaken for a page that carries no links, and the read is never retried past the gate. - A login wall is a working link. A link that redirects to a sign-in page, like a social profile that sends logged-out visitors to its login screen, resolves: what is behind it needs an account. It is reported
okwith the diagnosislogin_required; the sign-in page itself is never requested. - The whole redirect chain. Up to 10 redirects are followed per link, with no cookies sent.
redirect_chainlists each hop with its status, andfinal_urlandredirect_countare read, not assumed. A chain that never reaches a page is a freeerrorrow with the diagnosisredirect_loop. - Three honest verdicts and a diagnosis.
brokenwhen the site says the link is dead (404, 410, a 4xx that is not a refusal, a 5xx that stood on re-check).okon a healthy answer,errorwhen it could not be judged. The reason sits inerror_message, empty on every link the site answered for, anddiagnosissays what happened in one word, fromnot_foundandserver_errortotimeout,dns_errorandslow. - Internal or external.
is_externalmarks each link that points off the site it was found on, so your own dead pages and dead outbound links sort apart.checkExternaloff keeps the check to your own site. - Links, and assets when you ask.
checkAssetsadds the images, scripts and stylesheets each page loads, each a checked row of its own. - Checked once across pages. A footer link repeated across pages is checked once;
all_sourceslists every place it appeared. - Every entry answered, even a clean site.
no_broken_linksmeans its links were already checked under an earlier entry or all came back healthy,no_links_foundmeans the page carried nothing to check, andnot_reachedmeans the run's budget ended first. A clean site never reads as silence. - Unusable entries answered, not dropped. A
queriesentry that is not a fetchable URL gets its own freebad_urlrow with the reason. - Ordinary browser headers on every request. Every request, the first one included, goes out with the same ordinary browser headers, a Chrome user-agent and an accept-language. The identity is uniform: never switched, never rotated, and never swapped in after a refusal.
- Plain HTTP, no browser, no login, no cookies. Up to 8 links are checked at once, each on a different site, and several sites run in parallel. A 401 is one page asking for a login, so the rest of that site is still read; a 403, 429 or 999 closes the site for the run.
- Tune the check to the server.
requestTimeoutMswaits longer for a slow server or less for a fast sweep.maxConcurrencychecks fewer sites at once.maxRedirectsstops following redirects early, or at 0 reports each redirect itself. - Caps you control.
maxResultsPerQuerylimits the links checked per page;maxItemslimits the whole run, 100 by default.
What you can count on
- You pay only for rows the run judged; a row it could not judge is pushed as an uncharged
ITEM_STATUSrow with the reason on it. - Every row is charged only after it is written to your dataset; a row you cannot see is never billed.
- A run that finds nothing returns one
NO_RESULTSrow that says so, never an empty dataset, and it is not charged. - A spending limit you set stops the run cleanly: one
STOPPED_EARLYrow reports how many rows were returned and how many were not. - You always know a short run from a broken one: every run writes a
RUN_SUMMARYentry withrowsFetched,rowsPushed,rowsChargedandduplicatesDropped. - If the read itself fails, the run fails with the error in the log; it never returns rows full of nulls and calls it success.
- No value is invented: a field the page does not show comes back empty, and the field table says which fields can be.
Why this one
- The most-used alternative on this platform bills $0.001 for every link it checks (its pricing read on 2026-09-17). This Actor bills $0.0008 for the same unit, one event per link the site answered for, and a link it cannot judge is free.
- Run on the same crawler-test.com page on 2026-09-16, the most-used alternative's crawl kept going past the page it was given and returned 19 rows. This Actor, in its Pages mode, checked the page it was given and returned 11. It reads what you name, not what it can find.
- A link is never called broken on a HEAD answer alone: an unhealthy answer is asked once more with GET. A server error is re-checked once after a pause, and a refusal is never billed as broken.
- Misses are free and explained: a link that never answered, one that looped or one the site refused is an uncharged
errorrow with the reason on it. - The fill rate is measured, not promised: 30 of 40 sites it had never seen produced checked-link rows. Every one of the 40 produced a row that says what happened (2026-10-01).
- Every row names the page the link was found on, so a dead link is a fix, not a riddle.
What data does Broken Link Checker return?
One real row from a run over the default input's three example pages:
{"url": "https://crawler-test.com/links/not_found/foo1","final_url": "https://crawler-test.com/links/not_found/foo1","status": 404,"status_text": "Not Found","classification": "broken","is_broken": true,"source_domain": "crawler-test.com","source_url": "https://crawler-test.com/links/broken_links_internal","anchor_text": "Broken Internal Link 1","element": "a","all_sources": [{"source_url": "https://crawler-test.com/links/broken_links_internal","anchor_text": "Broken Internal Link 1","element": "a"}],"method": "GET","redirect_count": 0,"redirect_chain": [],"content_type": "text/html","error_message": null,"duration_ms": 345,"checked_at": "2026-09-29T13:49:50.580Z","diagnosis": "not_found","is_external": false,"is_redirect": false,"is_slow": false,"row_type": "ROW"}
Every field below sits on each row about a link or an entry, so the dead links and the healthy ones read the same. The run's own bookkeeping, the RUN_SUMMARY entry and the STOPPED_EARLY or NO_RESULTS row, carries the counters; a stopped or empty run's row also carries reason, saying why. When a check cannot produce a value the cell is left empty; nothing is guessed.
| Field | What it is |
|---|---|
url | The link as it was found, made absolute. This is the row's key: one row per unique link URL. |
final_url | Where the link ended after its redirects; the same as url when it never redirected, null when it never answered. On a link that redirects to a sign-in page, the sign-in page's address, which is not requested. |
status | The HTTP code the link's final address answered with. On a link that redirects to a sign-in page, the redirect's own code. On a free status row this cell carries the miss word instead; error, bad_url and the rest are listed below. |
status_text | The server's status line, like Not Found; empty when the server sent none. On an error row where the site answered (a refusal, a redirect loop) it keeps the code, like HTTP 403 Forbidden; when nothing answered, the failure reason. When a first answer was a server error and a re-check settled the link, it names the transient error too. |
classification | The verdict on the link: ok on a healthy answer (a link that redirects to a sign-in page included), broken when the site says the link is dead, error when it could not be checked. |
is_broken | true when the site says the link is dead (a 4xx other than a refusal, or a 5xx answered again on re-check after a pause, confirmed with GET), false on a healthy answer, null when the link could not be checked or the site refused the checker. |
diagnosis | What happened, in one word: ok, redirected, slow, redirect_unfollowed, login_required (the link redirects to a sign-in page: it works, and what is behind it needs an account), not_found, gone, client_error, server_error, blocked_by_site, redirect_loop, timeout, dns_error, tls_error or connection_error. Null on an entry that was never checked. |
source_url | The page this link was first found on; for a sitemap-listed page that failed, the sitemap that listed it. |
is_external | true when the link points to another site than the page it was found on (a leading www. ignored), false when it stays on that site; null on an entry that was never checked. |
is_redirect | true when the link redirected before its final answer, or answered with a redirect it was not asked to follow; false when it answered directly; null when it never answered. |
is_slow | true when the check took longer than slowThresholdMs, redirects included; false when it did not; null when the link never answered. |
anchor_text | The clickable text of the link on that page; an image's alt text when assets are checked. Empty when the link carried no text. |
element | The HTML element the link came from: a for a link, img, script or link for an asset, sitemap for a page a sitemap listed. |
all_sources | Every place the link appeared: one entry per anchor that carries it, so the same page linking it again lists it again. Empty on a row for an entry that could not be read. |
method | HEAD or GET: which request gave the verdict. GET means HEAD was not healthy or failed and the link was asked again. |
redirect_count | How many redirects were followed before the final answer; zero means the link answered directly. |
redirect_chain | Every redirect in order: the address that answered, its status code and where it pointed. Empty when the link answered directly. |
content_type | The media type the final address answered with, like text/html or image/png; null when the server sent none or never answered. |
error_message | Why the link could not be judged, on an error row: a timeout, a refused connection, a name that does not resolve, a redirect loop, or a site that refused the checker. Null on every link the site answered for. |
duration_ms | How long the check took in milliseconds, redirects included. |
checked_at | The ISO timestamp of the check. |
source_domain | The host of the start page the link was found on (one of the pages you named); the link's own host is in url. |
row_type | ROW on every link the site answered for, ITEM_STATUS on an error row or an entry that could not be read; NO_RESULTS and STOPPED_EARLY are the run's own messages. |
The Overview tab shows the per-link fields; the All fields tab adds redirect_chain, content_type, error_message, diagnosis, is_external, is_redirect, is_slow, duration_ms, checked_at and source_domain. error_message is filled only on error rows, so on a sweep where every link answered it is empty on every row.
A free row is a miss, not a checked link; its status carries one of these words in place of an HTTP code:
status value | On which rows | What it means |
|---|---|---|
An integer HTTP code, like 200 or 404 | ROW | The status the link's final address answered with. |
error | ITEM_STATUS | The link, or a start page you named, could not be judged: it never answered (a timeout, a refused connection, a name that does not resolve), it redirected in a loop, or the site refused the checker. classification is error, diagnosis says which, the reason sits in error_message, and the row is free. |
bad_url | ITEM_STATUS | The queries entry was not a usable URL. The entry is named in url, the reason sits in reason, and the row is pushed free: a malformed entry is answered, never dropped. |
not_found | ITEM_STATUS | Sitemap mode found no sitemap listing a page for this site. The entry is in url, what was tried is in reason, and the row is free. A different scope from the diagnosis not_found on a link row, which is the link's own 404. |
blocked | ITEM_STATUS | The page was not read, in any mode, and reason says why: the site's robots.txt does not allow it, the site refused the checker earlier in the same run, it is a sign-in page, its site is on this Actor's exclusion list, the run had already read 2,000 pages, or the site answered the read with a script gate, a challenge stub such as an Incapsula or Cloudflare interstitial, instead of the page. A gate is named for what it is and never retried past. Free. |
no_broken_links | ITEM_STATUS | The site answered and there was nothing new to report: every link it carried was already checked earlier in the run, or every checked link came back healthy. The entry is named in url, the detail sits in reason, and the row is free. |
no_links_found | ITEM_STATUS | The page answered but carried nothing this run could check: not a readable HTML page, no links a plain read can see, or links all outside the scope you asked for. A page that answered with a recognised script gate is blocked, not this. The entry is named in url, the row is free. |
not_reached | ITEM_STATUS | The run's row cap or page budget ended before this entry could be read: nothing was fetched for it and nothing was billed. |
You are billed for ROW rows, one per link the site answered for. ITEM_STATUS rows are pushed so you see them but never billed: a link that could not be judged, or an entry that could not be read. NO_RESULTS means no page produced a checkable link. STOPPED_EARLY means your spending limit ended the run. Status rows also carry the run's bookkeeping:
| Field | What it is |
|---|---|
reason | On a bad_url, not_found, blocked, no_broken_links, no_links_found or not_reached row, why that entry was not read or what its read found. On NO_RESULTS or STOPPED_EARLY, why the run returned no link rows or stopped early. An error row carries its reason in error_message instead. |
rowsFetched | How many link rows the check produced, counted before duplicates were dropped and any cap applied; on status rows only. |
rowsReturned | On a status row: the link rows in the dataset. On STOPPED_EARLY, the rows returned before the spending limit stopped the run; 0 on an empty-result row. |
rowsRemaining | On a status row: the rows not returned when the spending limit stopped the run (STOPPED_EARLY), and 0 on an empty-result row. |
HTTP status code cheat sheet
What the status on a checked link means, and what to do about it. Redirects (301, 302, 303, 307, 308) are followed. status is the code the final address answered with. Each redirect's own code sits in redirect_chain:
| Status | Verdict | What it means | What to do |
|---|---|---|---|
200 | ok | The page answered normally. | Nothing. |
204, 206 | ok | Answered with no content, or part of a file. | Usually nothing. |
301, 308 in redirect_chain | set by the final status | Moved permanently. final_url is where it ended. | Update the link to final_url so visitors skip the hop. |
302, 303, 307 in redirect_chain | set by the final status | Moved temporarily. | Fine for a login or a tracking link; check a long chain. |
A 3xx in status | ok | Not followed to its end: the redirect named no new address, or the checker does not follow that code, such as 300 or 304. diagnosis is redirect_unfollowed. | Open final_url by hand. |
A 3xx in status, diagnosis login_required | ok | The link redirects to a sign-in page, shown in final_url. The link works; what is behind it needs an account. Not broken. The sign-in page is not requested. | Nothing, unless the page should be public. |
error, diagnosis redirect_loop | error | More than 10 redirects without reaching a page. status_text keeps the last code, redirect_chain shows the hops. The row is free. | Open the link in a browser; a loop there too needs fixing. |
400 | broken | The server rejected the request as malformed. | Check the URL for typos or bad characters. |
401, 403, 999 | error | The site refused the checker or wants a login. That does not say the link is dead, so the row is free, with diagnosis blocked_by_site and the code in status_text. | Open it in a browser; fine if the page is private. |
404 | broken | Not found. | Fix the link, or redirect the old address. |
410 | broken | Gone on purpose. A site can also answer 410 to a checker while still serving the page to a browser. | Remove the link; open it in a browser first if it matters. |
429 | error | Too many requests. The checker does not ask again, and asks that site nothing more in the run, so the row is free, with diagnosis blocked_by_site. | Re-run later, or with maxPages lower. |
500, 502, 503, 504 and other 5xx | broken | The server failed or was down when checked, and it is not believed on one answer: the link is asked once more after a pause. A healthy re-check comes back ok with the first error named in status_text; a refusal on the re-check is a free error row. Only a 5xx that stands, or a re-check that never answers, lands broken. | Re-run later; a link that stays 5xx is broken. |
error (no HTTP code) | error | No answer: a timeout, a refused connection, a name that does not resolve. The row is free. | Read error_message; re-run to rule out a passing outage. |
How much does it cost?
Every checked-link row costs $0.0008, the link-checked event, charged only after the row is written. Apify also bills its own apify-actor-start event once per run, $0.00005 at this Actor's memory size. Status rows, miss rows (error, bad_url, not_found, blocked, no_broken_links, no_links_found, not_reached) and dropped duplicate links are free. You pay for every link the site answered for, ok and broken alike. A link is billed once however its check ran: HEAD alone, HEAD then GET, or a server error asked once more after a pause.
| Checked-link rows | Link charges | Start event | Total |
|---|---|---|---|
| 100 | $0.08 | $0.00005 | about $0.08 |
| 1,000 | $0.80 | $0.00005 | about $0.80 |
| 10,000 | $8.00 | $0.00005 | about $8.00 |
To budget a sweep of your own sites, the measured run gives an estimate. On 40 sites it had never seen, 80% answered and 30 produced link rows, about 157 checked-link rows each (measured 2026-10-01). At that rate, and with maxItems raised past the default 100:
- 100 sites: about 80 answer and about 75 produce link rows, about 11,745 checked-link rows, about $9.40 plus the start event.
- 1,000 sites: about 800 answer and about 750 produce link rows, about 117,450 checked-link rows, about $93.96 plus the start event.
- 10,000 sites: about 8,000 answer and about 7,500 produce link rows, about 1,174,500 checked-link rows, about $939.60 plus the start event.
Your sites will differ, and error rows are free, so a real sweep can come in under those totals.
On a paid Apify plan the per-row price steps down with your tier: $0.00067 on Bronze, $0.00054 on Silver and $0.00043 on Gold and above. That is $0.67, $0.54 and $0.43 per 1,000 rows.
Duplicates are billed once, not once per sighting. If 50 link sightings across your pages deduplicate to 47 unique URLs, you pay for 47 checks; the repeats join all_sources free.
How do I use Broken Link Checker?
- Open the Actor and press Try for free.
- Paste your pages or sites into Pages or sites to check, one URL or bare domain per line. The input carries three example pages from a public test site; replace them with yours.
- Pick a Mode: Pages for just those pages, Crawl for the whole site from each page, Sitemap for every page the sitemap lists.
- Set Max pages and Maximum items if you need to; every link the site answers for comes back as a billed row carrying its verdict.
- Press Start. Rows land in the dataset as links are checked, and the run ends when every page is read.
Example input:
{"queries": ["https://crawler-test.com/links/broken_links_internal", "https://crawler-test.com/links/broken_links_external", "https://example.com/"],"maxItems": 100}
Worked examples
Use it to get every link on a few pages you just published, healthy ones included.
{"queries": ["https://crawler-test.com/links/broken_links_internal", "https://crawler-test.com/links/broken_links_external"],"mode": "pages"}
Use it to sweep a whole site and get every checked link back with its verdict.
{"queries": ["crawler-test.com"],"mode": "crawl","maxPages": 200,"maxItems": 1000}
Use it to check every page in a sitemap, the site's own links only, as a technical SEO audit.
{"queries": ["https://crawler-test.com/test_sitemap.xml"],"mode": "sitemap","maxPages": 500,"maxItems": 5000,"checkExternal": false}
Use it to find missing images, scripts and stylesheets on a landing page.
{"queries": ["https://crawler-test.com/"],"checkAssets": true}
Or start a run over the API:
curl -X POST "https://api.apify.com/v2/acts/Pradio~broken-link-checker/runs?token=YOUR_APIFY_TOKEN" -H "Content-Type: application/json" -d "{\"queries\":[\"https://example.com/\"]}"
Input
| Input | Default | What it does |
|---|---|---|
queries | three example pages | The pages or sites to check, one URL or bare domain per line, at most 2,000. An entry that is not a fetchable URL gets its own free bad_url row with the reason, so a typo is answered, never dropped. |
mode | pages | pages reads only the pages you list. crawl starts at each and follows links on the same site. sitemap reads the pages the site's sitemap lists. |
maxPages | 50 | In crawl and sitemap modes, the most pages one run reads for links, across all entries. pages mode reads the pages you list, never more than 2,000 in one run. |
checkExternal | true | Check links to other sites too. Off keeps the check to the site you named. |
checkAssets | false | Also check the images, scripts and stylesheets each page loads. |
slowThresholdMs | 5000 | A healthy link whose check takes longer than this, in milliseconds, gets the diagnosis slow and is_slow true. It changes no verdict and no charge. |
requestTimeoutMs | 15000 | How long one link check waits for an answer, in milliseconds, from 1,000 to 60,000. A link that does not answer in time is a free error row with the diagnosis timeout. Pages read for links always wait at least 15 seconds. |
maxRedirects | 10 | How many redirects a link check follows, from 0 to 10. At 0 each redirect is reported itself with the diagnosis redirect_unfollowed; a lower cap reports where it stopped. Only a chain past 10 is a redirect_loop. |
maxConcurrency | 8 | How many links are checked at the same time, from 1 to 8, each on a different site: one site never gets more than one request at a time. Lower it to check fewer sites at once. |
maxResultsPerQuery | none | The most new links checked from one page read. Unset means every link found is checked. |
maxItems | 100 | The most link rows one run returns in total. Rows past the cap are dropped, and the run summary shows how many links were checked. |
queries
Each entry is a full URL like https://example.com/blog, or a bare domain like example.com, which is fetched over HTTPS. In pages mode the page is fetched once, its anchors are collected and each unique link is checked. An entry nothing can be fetched for, such as a typo or an unrecognisable line, gets its own uncharged row. That row carries status bad_url, the entry in url and the reason in reason.
mode
pages(the default) reads exactly the pages you list and nothing else, at most 2,000 in one run. It reads each site's robots.txt first: a page robots.txt disallows, or one answered by a script gate, is not read and comes back as a freeblockedrow.crawlreads each entry, then the pages on the same site it links to, then theirs, breadth first, untilmaxPagespages are read ormaxItemsrows are found. It never leaves the site, reads the site's robots.txt first, skips any page robots.txt disallows and waits the crawl delay it asks for. Links to disallowed pages are still checked with one request each; their own links are not read.sitemapreads the site's sitemap: the entry itself when it is a.xmlsitemap URL, else the sitemaps robots.txt names, else/sitemap.xml. A sitemap index is followed into its child sitemaps on the same site. Up tomaxPageslisted pages are read for links, and a listed page that fails is itself a row, with the sitemap as itssource_url. A site with no readable sitemap gets one freenot_foundrow.
Raise maxItems for a crawl: at the default of 100 rows a crawl stops at the first hundred links.
{"queries": ["https://example.com/"],"mode": "crawl","maxPages": 50,"maxItems": 2000}
Output
Rows land in the default dataset as links are checked. The default input's three example pages return 12 link rows. A page whose checked links all come back healthy also gets a free no_broken_links row saying so. Four extras tell you how a run went:
ITEM_STATUS: a link or start page that never answered, looped or was refused. It carriesstatuserror, the kind indiagnosisand the reason inerror_message. The same kind marks aqueriesentry that could not be read or had nothing to report. It is named inurlwithstatusbad_url,not_found,blocked,no_broken_links,no_links_foundornot_reachedand the reason inreason. It is pushed so you see it, and never billed.NO_RESULTS: no start page produced a checkable link. One uncharged row carriesreasonandrowsFetched, so an empty answer is an answer, not silence.STOPPED_EARLY: your spending limit ended the run. One uncharged row carriesrowsReturnedandrowsRemaining; raise the limit and re-run for the rest.RUN_SUMMARY: an entry in the run's key-value store carrying the counts: links fetched, rows pushed, rows charged, duplicates dropped, stopped early or not.
A start page that cannot be reached is itself a free error row with the reason in error_message. A run over pages that are all down still tells you so, page by page, and costs only the start event.
What can you do with the data?
Sweep your own site before a launch. Run the checker over the pages, filter classification to broken, and hand the list to whoever fixes it. Each row already carries the dead URL, the page it sits on and the anchor text to search for.
Audit a migration. Old URLs that still redirect show every hop in redirect_chain and where they landed in final_url. The ones that answer 404 instead are the ones to remap.
Watch outbound links on a schedule. Reference, partner and affiliate links rot quietly. Put the run on an Apify schedule, export the broken rows to a sheet, and the dataset becomes a monthly fix list.
Qualify a site you are evaluating. A page full of dead links says something about how it is maintained. One run counts them without a manual click-through.
Use Broken Link Checker with AI agents
Paste this line to give an agent this Actor through Apify's MCP server:
claude mcp add --transport http apify "https://mcp.apify.com?tools=Pradio/broken-link-checker"
Personal data
- A row carries only the declared link-health fields: the checked URL, its HTTP status and verdict, the anchor text and the page it appeared on. No field names or identifies a person.
- A row keeps facts and the link's own anchor text, cut to at most 200 characters, never the body or a substantial part of a page.
mailto:andtel:links are never requested, and the email addresses and phone numbers in them are never recorded.javascript:links are never requested either.- Every row links back to the page it was found on, so what was collected is easy to check.
- A page that refuses the read is reported as a free
errorrow with the diagnosisblocked_by_site. A link whose HEAD request is refused (401, 403, 429 or 999) is not asked again: the refusal is its answer. A link that redirects to a sign-in page is reported as working; the sign-in page is never requested. An error code that is not a refusal is asked once with an ordinary GET. A server error is asked once more after a pause, at most one re-check per link. Every request, the first one included, carries the same ordinary browser headers, a Chrome user-agent and an accept-language. It is never switched or rotated, and never after a refusal. A page that answers the read with a script gate instead of the page is reportedblocked, free. The verdict is a report, never a trigger to re-request with different headers, a proxy or a solver. The read is never retried past the gate. Once a site refuses the checker (403, 429 or 999), nothing more is asked of that site in the run. One request at a time goes to any one site. Nothing retries past a block or works around a refusal, and there is no proxy. - Personal data is not what a row is about. It can still appear incidentally inside anchor text or a URL path, such as a name in a link label or a profile slug in a path. Run it against sites you operate or are permitted to check.
- You choose the pages and sites, so you are the controller for the list you supply. In
pagesmode the Actor reads only what you list. Incrawlandsitemapmodes it reads further pages of the same site, never another site. In every mode it never reads a page the site's robots.txt disallows. To exclude a section, disallow it in robots.txt. - A site owner can ask for their site to be left out entirely. Open an issue on this Actor's Issues tab naming the domain, and it goes on the Actor's exclusion list. Every later run reads that list before any request and sends the site nothing. A page on it comes back as a free
blockedrow, a link to it as a freeerrorrow.
Release notes
- 2026-10-01: every request now carries ordinary browser headers, a Chrome user-agent and an accept-language, from the first request on; the identity is never switched after a refusal. A link whose check ends in a server error is asked once more after a pause before it is called
broken; a healthy re-check comes backok. The transient first answer stays named instatus_text. And a site that answers a read with a script gate is reported as a freeblockedrow. It names the gate (an Incapsula or Cloudflare challenge stub instead of the page) and is never called a page with no links. - 2026-10-01: every entry in the list now lands a row of its own. An entry whose links were already checked under another entry, or that carried nothing to check, gets a free row saying so; before, a quiet entry could come back with no row when entries ran at the same time.
- 2026-09-29: billing moves to per link checked. The
link-checkedevent bills every link the site answered for,okandbrokenalike, and every checked link lands a row carrying its verdict. TheonlyBrokeninput is gone: billed rows and dataset rows are the same set. Entries on different sites now run in parallel; each site still gets one request at a time at its own pace. - 2026-09-29: an entry that answered with nothing to report now says so on a free row.
no_broken_linksmeans every checked link was healthy,no_links_foundmeans the page carried nothing to check,not_reachedmeans the run's row cap or page budget ended first. - 2026-09-28: new
requestTimeoutMs,maxRedirectsandmaxConcurrencyinputs to tune the check to the server, and newis_redirectandis_slowcolumns.onlyBrokenis added, on by default, so a run can return and bill only the broken links (superseded by the 2026-09-29 billing change, which removed it). Anchor text is kept to a short label of at most 200 characters. One request at a time goes to any one site. A 429 is no longer asked again after a wait, and a site that refuses the checker is asked nothing more in the run. Pages mode now reads robots.txt too, and reads at most 2,000 pages in one run. A robots.txt crawl delay now paces every request to its site, link checks included. No cookies are sent: a redirect that loops without one is a freeredirect_looprow. A page that answers with a sign-in page is not read, and a link that redirects to one is reported as working, with the new diagnosislogin_required. Site owners can ask to be excluded. Existing columns and inputs keep their names. - 2026-09-27: a link is never called broken on a HEAD answer alone. An error code on HEAD that is not a refusal, or a failure a GET could answer, is asked again with GET, and the GET decides. A DNS or TLS failure is not re-asked: no method reaches it. A refusal on HEAD is not asked again. A site that refuses the checker (401, 403, 429, 999, a sign-in redirect) is now a free
errorrow, and so is a redirect loop. Commented-out links in a page are no longer read. Newdiagnosisandis_externalcolumns and aslowThresholdMsinput. Existing columns and inputs keep their names. 0.1.22(2026-09-25): anerrorrow, a link or start page that never answered, is now free;statuscarrieserroron it androw_typeisITEM_STATUS.0.1.20(2026-09-24):modeaddscrawlandsitemapbesidepages; newcheckExternal,checkAssetsandmaxPagesinputs; newredirect_chain,content_typeanderror_messagecolumns. Existing inputs work as before.
Limits
- It checks links and reports them; it does not fix links: repair stays with you.
- A crawl stays on the start page's site and reads at most
maxPagespages; a site larger than that is read in part. Pages robots.txt disallows are not read, so links that only appear on them are not found. - Anchor tags are checked by default; images, scripts and stylesheets only with
checkAssets.javascript:,mailto:,tel:,data:,sms:andftp:links are skipped: they are not requested, and no email address or phone number from them reaches a row. - Links a page builds with JavaScript after load are not seen: pages are fetched over plain HTTP, with no browser. A page that answers with a recognised script gate, an Incapsula or Cloudflare challenge stub, is not mistaken for an empty one. The entry comes back a free
blockedrow that names the gate; the read is never retried past it. Any other page whose links exist only after scripts run reads asno_links_found. - Requests are logged-out and public, sent with no cookies and through no proxy. Every request, the first one included, carries the same ordinary browser headers, a Chrome user-agent and an accept-language. That identity is uniform: it is never switched, rotated or swapped in after a refusal. A site can still answer automated traffic with a dead status it would not serve a signed-in browser. That row lands
brokenon the site's own answer, kept instatus_text, and a browser could disagree with it. A redirect that only reaches its page when a cookie is sent back reads as a freeredirect_looprow. There is no retry past a block: a site that refuses is a freeerrorrow with the diagnosisblocked_by_site, not a workaround. A 429 is the answer, never asked again, and a site that refused is asked nothing more in that run. - A start page that cannot be read is its own row, and it never disappears from the results. It is an
errorrow with the reason inerror_messagewhen nothing answered or the site refused. It is abrokenrow with the status code when it answered a 4xx that is not a refusal or a 5xx; the pause re-check is for link checks, not page reads. - A certificate the checker cannot verify is an
errorrow with the diagnosistls_error, even where a browser completes the chain on its own. - A link that never answers within the request timeout, 15 seconds by default, is an
errorrow, suspected but not proven broken. - The verdict is the HTTP answer. A page that answers
200with a "not found" message of its own (a soft 404) readsok, whatever its link text says. - A redirect chain is followed for at most 10 redirects; a longer chain is a free
errorrow with the diagnosisredirect_loop.maxRedirectscan lower the cap, never raise it. - Sparse fields, measured on the 4,698 rows from the sites it had never seen:
anchor_textwas filled on 87.5% (an anchor with no text leaves it empty).content_typewas filled on 99.8% (a server that sends no media type),status_texton 99.5%.redirect_chainis filled only when a redirect happened, which was 7.9% of rows. - Every link the site answered for is a billed row,
okorbroken. Anerrorrow is free, so a run of nothing but unanswered or refused links costs only the start event.
Troubleshooting
The run returned fewer rows than maxItems.
maxItems is a ceiling, not a target. Fewer rows means the pages ran out of links first; the run log shows how many links were checked.
The dataset has one row saying NO_RESULTS.
No start page produced a checkable link: a clean result. That is the answer to this run, not a bug, and the row is free.
A row says no_broken_links or no_links_found in status.
The site answered. no_broken_links means its links were already checked under an earlier entry or all came back healthy. no_links_found means the page carried nothing this run could check. A site that answered with a recognised script gate comes back blocked instead, not no_links_found. All are complete answers, and all such rows are free.
An entry comes back not_reached or blocked.
not_reached means the run's row cap or page budget ended before that entry was read: nothing was fetched for it and nothing billed. blocked means the page was not read, and reason says why: robots.txt disallows it, the site refused the checker earlier in the run, it is a sign-in page, or its site is on the exclusion list. A site that answered with a script gate instead of the page is also blocked. Both rows are free.
A row says bad_url in status.
One queries entry was not a URL the Actor could fetch, a typo or an unrecognisable line. The row names the entry in url, gives the reason in reason, and is free. The other entries ran normally.
A row says error with is_broken empty.
The link could not be judged, and diagnosis says why: timeout (no answer within the request timeout), dns_error, tls_error, connection_error, redirect_loop or blocked_by_site. It is not proven dead, and the row is free; re-run later to retry it.
An ok row's status_text names a server error, like HTTP 503.
The first answer was a transient server error and the re-check after a pause answered healthy, so the link is fine. The row keeps the first answer for you to see; only a 5xx that stands on re-check is called broken.
A link I can open in my browser shows blocked_by_site.
The site refused this checker's plain request while it serves your browser. The row is free and is not counted as broken. The checker does not work around a refusal.
A link to a social profile shows ok with the diagnosis login_required.
The link redirected to a sign-in page, as Instagram and others do for logged-out visitors. The link works and the page behind it needs an account, so it is not broken. The sign-in page is not requested.
Every row is an error row for a start page I listed.
None of the pages you listed answered at the network level. Each row names its page in url and the reason in error_message: a timeout, a refused connection, a name that does not resolve. Check the URLs you pasted.
FAQ
Can I use integrations with Broken Link Checker? Yes. The dataset plugs into Apify integrations like Zapier, Make and webhooks, and Apify schedules can re-run the sweep weekly or monthly without you touching it.
Can I use Broken Link Checker with the Apify API? Yes. Start runs and read the dataset back over the API; the curl example above is the whole call.
Can I use Broken Link Checker through an MCP server? Yes. The line under "Use Broken Link Checker with AI agents" adds it to an agent through Apify's MCP server.
Will it call a link broken that still works?
Not on a first answer alone. A link whose HEAD request comes back with an error code is asked once more with GET, the request a browser makes, and the GET decides. A server error is asked once more after a pause before it counts: a healthy re-check comes back ok. A refused request (401, 403, 429 or 999) is a free error row with the diagnosis blocked_by_site, never a billed broken. So is a link that never answered. A site that answers dead statuses to automated traffic can still produce a broken row a browser would disagree with; the row keeps the site's own answer in status_text. The one case it cannot see is a page that answers 200 with its own "not found" message; that reads ok.
Can I use it for scheduled link monitoring? Yes. Save your input as a task and put it on an Apify schedule, daily, weekly or monthly. Each run is a fresh sweep, and the dataset holds every link the sites answered for at that moment, each billed as a checked link.
How do I export results to CSV or JSON? Open the run's dataset and export it as CSV, JSON, Excel, XML or HTML, or read it over the API. The Overview tab is the view most exports want; All fields adds the timing and redirect detail.
How many URLs can I check?
As many pages as you paste into Pages or sites to check, up to 2,000 entries. maxItems caps the link rows one run returns, 100 by default, so raise it for a big sweep. Crawl and Sitemap modes read at most maxPages pages per run. No mode reads more than 2,000 pages in one run, and every mode stops reading a site that refuses a request.
Do I need an API key or login?
No API key and no login on the sites you check: it reads public pages with plain requests. You need only your Apify account to run it. Pages behind a login are not read; a link that sends the checker to a sign-in page is reported as working, with the diagnosis login_required.
Is it legal to check links this way?
The Actor is for checking links. It makes plain HTTP requests that carry ordinary browser headers, the same on every request from the first; it bypasses no login and keeps only link facts. It honours robots.txt and its crawl delay in every mode, sends one request at a time to any one site and stops on a refusal. It does not claim a site's terms permit the checks; that permission stays with you. You choose the pages, so point it at sites you are responsible for or allowed to test. A site owner can ask to be excluded on the Issues tab, and an excluded site is sent no request at all. You are responsible for having the right to check the pages and sites you give it, and for how you use the rows. It makes no promise to get past blocks, logins or CAPTCHAs: a page that refuses it is reported, never worked around. A page that answers with a script gate instead of the page is reported blocked, never retried past.
Feedback
Found a problem or a missing field? Open an issue on the Issues tab; it is answered within two days. If this Actor saved you time, a review on the Store helps other buyers find it.
Not affiliated
Broken Link Checker is an independent tool, not affiliated with or endorsed by any site you point it at. The example pages in the default input are public test pages on crawler-test.com and example.com.