Broken Link Checker: Find Dead Links on Any Site avatar

Broken Link Checker: Find Dead Links on Any Site

Pricing

from $5.00 / 1,000 page crawleds

Go to Apify Store
Broken Link Checker: Find Dead Links on Any Site

Broken Link Checker: Find Dead Links on Any Site

Crawl a website and get every broken link (404, 410, 5xx, dead domains), internal and external, with the exact pages it appears on. Schedule it to see only newly broken and fixed links. Respects robots.txt. Pay per page crawled.

Pricing

from $5.00 / 1,000 page crawleds

Rating

0.0

(0)

Developer

Bruno Petrelli

Bruno Petrelli

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

12 hours ago

Last modified

Share

Give it a website and get every broken link, internal and external, with the exact pages it appears on, the most used first, so you fix what hurts most. The Actor crawls the site through its links and its sitemap, then checks every link it found:

  • Broken: 404, 410, 5xx, domains that no longer exist, refused connections, bad certificates, timeouts.
  • Blocked: the server refused the checker (401, 403, 429, LinkedIn's 999, or a non-standard 4xx code such as the 419 Hacker News sends to scripts), or the link only takes other methods (405, like an API endpoint). Listed apart, because the page may well exist.
  • Working: optional, for a full link inventory.

Respects robots.txt on every host it reads, including the sites your links point to. No browser and no proxies; the User-Agent says broken-links.

Use it for

  • Before a launch or after a migration: catch every 404 your own pages link to.
  • Monthly site maintenance: schedule it and get only the new problems in a CSV.
  • SEO and UX: broken links waste crawl budget and lose visitors. Fix the ones on the most pages first.
  • Agencies: check all your clients' sites in one run, one summary row per site.

Input

{
"websites": ["example.com"],
"maxPagesPerSite": 500,
"checkExternalLinks": true
}

Output

One row per problem link (a real result on 2026-09-25; the addresses were replaced):

{
"type": "link",
"site": "example.com",
"url": "https://example.thrivecart.com/checkout/",
"state": "broken",
"status": 404,
"error": null,
"internal": false,
"foundOnCount": 1,
"foundOn": ["https://example.com/academy"]
}
  • state: broken, blocked or ok. error explains network failures (for example ENOTFOUND for a dead domain).
  • foundOn lists up to 25 pages that contain the link; foundOnCount is the total.
  • One free summary row per website: pagesCrawled, linksChecked, brokenLinks, blockedLinks, pagesWithBrokenLinks, notCrawled (pages found beyond your limit), blockedByRobots.

That same run crawled 16 pages in 21 seconds and checked 39 unique links.

Monitoring: only what changed

Schedule the check (daily, weekly) with Compare with the previous check on. Each broken link then says newlyBroken: true when it was fine (or not linked) at the last check, and false when it was already broken, so a report can show only what broke since. The site row adds changes: how many links are newly broken, the links that were fixed (still linked and working now) and the broken links that no page links to any more, with the date of the check it was compared with. When the page limit cut the crawl (notCrawled above 0), that last list is null: a link found last time on a page not read now is unknown, not gone. The first check has nothing to compare with (changes: null).

Pricing

Pay per event: one event per page crawled, whatever the number of links on it. Link rows and summary rows are free. Set a maximum cost for the run and the Actor stops cleanly when it is reached.

Need the full SEO picture per page (titles, descriptions, headings, noindex, alt text, duplicates)? Use SEO Audit & Broken Link Checker by the same author, which includes this link check.

Good to know

  • Only pages on the website's own host are crawled (www. and the bare domain count as the same site); links to anywhere are checked.
  • Links are checked with a HEAD request; servers that refuse HEAD get one GET.
  • Links inside content that only appears after JavaScript runs are not seen.
  • Crawling is gentle by default: 3 requests at a time per website.

Support

A link reported wrongly? Open an issue on the Actor's Issues tab with the link and the page. Issues get an answer within a few days.