Broken Link Checker — Find 404s & Redirects Site-Wide avatar

Broken Link Checker — Find 404s & Redirects Site-Wide

Pricing

Pay per event

Go to Apify Store
Broken Link Checker — Find 404s & Redirects Site-Wide

Broken Link Checker — Find 404s & Redirects Site-Wide

Crawl a website and check every link: broken (404, 5xx, DNS, timeout), redirected links and their chains, with the pages and anchor text where each link appears. Internal and external links, optional images/scripts/CSS. Pay per link checked.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Mohamed T.

Mohamed T.

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Crawl a website and get every broken link, every redirect, and the exact pages (with anchor text) where each one appears. Pay only per link checked, from $0.0005.

Who it's for

  • SEO agencies and consultants running technical audits or post-migration checks.
  • Site owners and content teams who want a weekly broken-link report (schedule it on Apify).
  • Developers adding a link check to CI or to an n8n / Make / Zapier workflow through the Apify API.

What it does

  1. Starts from the pages you give (usually the homepage) and crawls the same site (www. and bare domain count as one site), up to Max pages to crawl per site.
  2. Extracts every <a href> link (optionally <img>, <script> and stylesheets too).
  3. Checks each unique link once, internal and external, following redirects by hand to record the full chain.
  4. Outputs one row per link with its verdict and up to 5 pages that link to it, plus the anchor text and nofollow flag.

Verdicts (the status field)

statusmeaning
ok2xx/3xx final response, no redirect
redirectedworks, but through one or more redirects (update the link to the final URL)
broken404/410/5xx, DNS failure, timeout, TLS error, redirect loop
blocked401/403/429/999: the site refuses crawlers or requires a login. It usually works in a browser, so it is not counted as broken (no false alarms from LinkedIn, npm, Stack Overflow…)
skippedrobots.txt disallows fetching it. Reported but not billed

Input example

{
"startUrls": ["https://books.toscrape.com"],
"maxPagesPerSite": 10,
"checkExternalLinks": true,
"checkResources": false,
"onlyProblems": false
}

Output example

{
"url": "https://example.com/old-pricing",
"status": "broken",
"broken": true,
"linkType": "link",
"internal": true,
"statusCode": 404,
"finalUrl": "https://example.com/old-pricing",
"finalStatusCode": 404,
"redirectCount": 0,
"redirectChain": [],
"responseTimeMs": 87,
"error": "HTTP 404",
"foundOnCount": 3,
"foundOn": [
{"page": "https://example.com/", "anchorText": "Pricing", "nofollow": false},
{"page": "https://example.com/blog/launch", "anchorText": "see our plans", "nofollow": false}
],
"checkedAt": "2026-09-27T12:00:00+00:00"
}

The Links table view shows the verdict, status code, final URL and source pages. Export to CSV, Excel or JSON, or fetch it through the API.

Pricing

Pay per event: link-checked = $0.0005 per unique link checked (crawled pages count as links; each link is billed once per run, however many pages contain it). URLs skipped because of robots.txt are free. Examples: a 10-page crawl with about 200 unique links costs about $0.10. 1,000 links cost $0.50. Your run's maximum cost setting is respected: the crawl stops at what it allows.

Tips

  • Want a short report? Turn on Only output broken or redirected links. Every checked link is still billed; healthy ones are just hidden.
  • Big site? Raise Max pages to crawl per site and Max links to check. For your own site you can lower the per-host delay.
  • Schedule weekly runs and send the dataset to Slack or email with an Apify integration.

Known limits

  • HTTP only, no browser: links that exist only after JavaScript runs (single-page apps) are not seen.
  • blocked links are not verified; check them manually if they matter.
  • Some servers return 200 for missing pages ("soft 404"); these can't be detected from the status code.
  • robots.txt is respected by default. Turn it off only for sites you own or are authorized to audit.

FAQ

Does it check links to other websites? Yes, with one request each (no crawling of other sites). Turn off Check external links to skip them.

Why is a link marked blocked and not broken? Many large sites answer 403 or 999 to any bot. Flagging those as broken would bury your real 404s in false alarms.

Does it collect personal data? No. It records URLs, status codes and anchor text only. mailto: and tel: links are ignored.

Is it polite to the sites it crawls? Yes. It sends an identifying user agent, waits between requests to the same host, retries with exponential backoff and respects robots.txt.

Related Actors: Bulk URL Status & Indexability Checker (check a list of URLs you already have) · Sitemap Extractor & Monitor.