Broken Link Checker: Find Dead Links on Any Website avatar

Broken Link Checker: Find Dead Links on Any Website

Pricing

from $1.00 / 1,000 page crawleds

Go to Apify Store
Broken Link Checker: Find Dead Links on Any Website

Broken Link Checker: Find Dead Links on Any Website

Crawl a website and find every broken link, image, script and stylesheet: 404s, 5xx, DNS and TLS errors, redirect loops and redirected links, each with the page it is on and its link text. Checks external links too. USD 1 per 1,000 pages plus USD 0.20 per 1,000 links.

Pricing

from $1.00 / 1,000 page crawleds

Rating

0.0

(0)

Developer

JT Palms

JT Palms

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

29 minutes ago

Last modified

Share

Give it your home page and it crawls the site, collects every link, image, script and stylesheet on every page, and checks each unique target once. You get one row per problem: the page it is on, the link, its text, the tag, the status code and what went wrong. Broken links (404, 410, 5xx, DNS, TLS, timeouts, redirect loops, malformed URLs), links blocked by bot protection (401, 403, 429, 999) and redirected links (with where they end up) are all reported.

USD 1 per 1,000 pages crawled plus USD 0.20 per 1,000 links checked. A 500-page site with 1,500 other links (images, scripts, external links) costs about USD 0.80.

What people use it for

  • Website QA before and after launches. Catch the 404s, missing images and broken scripts a redesign or CMS migration left behind.
  • SEO audits. Broken internal links waste crawl budget and leak link equity; internal links that redirect add a hop for every crawler. Fix them from the list, page by page.
  • Agencies. Run it on every client site on a schedule and send the results to Google Sheets or Slack.
  • Content maintenance. Find outbound links in old blog posts and docs that now point to dead or moved pages.

Sample output

Real rows from a 30-page crawl of scrapethissite.com, trimmed:

{
"sourcePage": "https://www.scrapethissite.com/pages/simple/",
"linkUrl": "https://lipis.github.io/flag-icon-css/css/flag-icon.css",
"linkText": "stylesheet",
"element": "link",
"issue": "BROKEN",
"status": 404,
"finalStatus": 404,
"redirectTo": null,
"redirectCount": 0,
"isExternal": true,
"errorType": null,
"error": null
}
{
"sourcePage": "https://quotes.toscrape.com/",
"linkUrl": "https://quotes.toscrape.com/author/Albert-Einstein",
"linkText": "(about)",
"element": "a",
"issue": "REDIRECT",
"status": 308,
"finalStatus": 200,
"redirectTo": "http://quotes.toscrape.com/author/Albert-Einstein/",
"redirectCount": 1,
"isExternal": false
}

That run crawled 30 pages, found 52 unique links, checked them in about 2 seconds and reported a missing stylesheet on 4 pages, two HTTP 400 links and two redirected external links. The second row, from quotes.toscrape.com, shows an internal link that redirects from https to http.

FieldWhat it is
sourcePageThe page the link is on. A broken link on 40 pages gives 40 rows, so you know every place to fix.
linkUrl, linkText, elementThe link target, its visible text (or image alt, rel for <link>), and the tag: a, img, link or script.
issueBROKEN, BLOCKED (401/403/429/999, often bot protection: check it in a browser) or REDIRECT.
status, finalStatusFirst response status and the status after following redirects.
redirectTo, redirectCountWhere a redirected link ends up and how many hops it took.
isExternaltrue for links to other websites.
errorType, errorFor links with no usable response: DNS, TLS, TIMEOUT, CONNECTION_REFUSED, REDIRECT_LOOP, TOO_MANY_REDIRECTS, INVALID_URL and so on.

The run's key-value store gets a SUMMARY record: pages crawled, links checked, unique links found, counts of broken, blocked and redirected links, links skipped by robots.txt or your patterns, and the top 100 problem links with how many pages each is on.

How to use it

  1. Put your home page (or any start page) in Website to check.
  2. Set Max pages to crawl (500 by default) and, if you like, Max link depth.
  3. Keep Check external links on to catch dead outbound links, or turn it off to check only your own site.
  4. Add Skip URLs containing patterns for endless or private sections, such as /logout, /cart or ?sort=.
  5. Click Start, then sort the table by issue or export CSV, JSON or Excel.

How it crawls

  • It starts at your URL and follows <a href> links (plus canonical, hreflang and next/prev <link> tags) on the same site, breadth first. www and non-www count as the same site; subdomains too if you turn that on.
  • It collects every a[href], img[src], link[href] and script[src] on each crawled page, resolves them against <base href>, and ignores mailto:, tel:, javascript: and #anchors.
  • Each unique target is checked once, however many pages link to it. Internal pages are checked by crawling them; everything else gets a HEAD request, with GET as a fallback for servers that refuse HEAD.
  • It reads the site's robots.txt and does not request pages it disallows (you can turn this off for your own site).
  • External sites get at most 2 requests at a time each, so nobody else's server is hammered.

Pricing

WhatPrice
Page crawled (an HTML page of the site downloaded and scanned for links)USD 0.001 (USD 1 per 1,000)
Link checked (each unique external link, image, script, stylesheet, or internal URL that is not crawled)USD 0.0002 (USD 0.20 per 1,000)
Result rowsFree
Start URL that cannot be reachedFree

Links found on crawled pages are all checked, including internal pages beyond Max pages, so a site with many product images costs more in link checks than in pages. Set a maximum cost per run in the run options; the actor stops cleanly when it is reached and keeps every result found so far.

Limits

  • It reads the HTML the server sends and does not run JavaScript, so links added only by JavaScript (some single-page apps) are not found.
  • Up to 5 MB of HTML per page. Links inside iframes, CSS files and srcset are not checked.
  • Sites behind bot protection may answer with 403 or 429. Those rows are marked BLOCKED, not BROKEN, so you can check them by hand; a residential proxy often helps.
  • Crawl-delay in robots.txt is not applied; lower Parallel requests for small or slow sites.

FAQ

Why is a link marked BLOCKED? The server refused the automated check with 401, 403, 429 or 999 (LinkedIn). The link may work fine in a browser. Open it to confirm.

Why are redirects listed? A link that redirects still works, but updating it to the final URL saves a hop for users and search engines. Turn off Report redirected links if you only want broken ones.

Can it check a site that needs a login? No. It sees what a logged-out visitor sees.

Is it legal? It requests public pages of the website you give it and the public URLs they link to, the same way a browser would, and it respects robots.txt by default. It does not log in and does not save page text; email and phone links are skipped, and the results are only link addresses and status codes.

Something missing or wrong? Open an issue with the site.