Website Broken Link Checker: 404s, Redirects, SEO Checks
Pricing
Pay per event
Website Broken Link Checker: 404s, Redirects, SEO Checks
Returns every link on the websites or sitemaps you list with HTTP status code, redirect chain, final URL, broken flag and anchor text, plus title, meta description, H1 count and canonical per page. Optional image check. Respects robots.txt, 2 requests per second per domain. onlyNew for monitoring.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Viktor Wiberg
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 hours ago
Last modified
Categories
Share
A website broken link checker: it crawls the sites you give it, checks every link on every crawled page and reports broken links, dead links to other websites and redirect chains. Each row is one link on one page: the HTTP status code, the full redirect chain, the final URL, whether the link is broken, the anchor text and whether it points inside the site or to another website. Each row also carries four SEO basics of the page the link sits on (title, meta description, number of H1 headings, canonical URL) and a short list of page issues such as a missing meta description or two H1s.
Use it before and after a site migration, as a weekly check of a client's site, to find 404s that waste crawl budget, or as a daily monitor that only reports new broken links. Give it a sitemap.xml instead of a start page and it checks exactly the pages you publish; turn on checkImages and broken images are reported too.
How it works
- Starts at each URL in
startUrlsand follows<a href>links breadth first, up tomaxPagesHTML pages in total. - With
sitemapUrlsit does not crawl: the pages listed in the sitemaps (and in any start URLs) are checked as a fixed list, up tomaxPages. See "Sitemap mode". - With
checkImagesevery<img src>on a checked page is checked as well and returned as a row withlinkTypeimage. - Only the websites you list are crawled. Links to other websites are checked with a single request (HEAD, and GET when HEAD fails) and never followed further. With
sameDomainOnly(default) only the start host is crawled;www.example.comandexample.comcount as the same site. - robots.txt is read for every host before the first request to it and followed for every URL, including external links (RFC 9309: a missing robots.txt allows everything, a robots.txt that answers with a server error blocks the host, and a host that cannot be reached at all is reported as a broken link). A
Crawl-delayis honoured up to 10 seconds. - At most 2 requests per second per domain by default (
requestsPerSecond, 0.2 to 5). Different domains are checked in parallel, so external links do not slow down the crawl of your own site. - Redirects are followed by hand, up to 10 hops, so the whole chain is recorded. Redirect loops and too many hops are reported as broken.
- Each URL is requested once per run, however many pages link to it.
- The User-Agent is
Mozilla/5.0 (compatible; NightwaveLinkChecker/1.0; +https://apify.com/nightwave-owner/website-link-checker). Site owners can address it in robots.txt asNightwaveLinkChecker.
Example from a real run
These runs were made on the Apify platform on 3 October 2026 against crawler-test.com, a public website built for testing crawlers, with pages that are broken, redirected or slow on purpose. Nothing in the rows is edited.
Input (run CbcOyZGC7m6gILp7F, 2 pages, 12 rows, 9 seconds):
{"startUrls": [{ "url": "https://crawler-test.com/links/broken_links_internal" },{ "url": "https://crawler-test.com/links/broken_links_external" }],"maxPages": 2}
One of the five broken internal links it found:
{"pageUrl": "https://crawler-test.com/links/broken_links_internal","linkUrl": "https://crawler-test.com/links/not_found/foo1","anchorText": "Broken Internal Link 1","type": "internal","rel": null,"statusCode": 404,"linkStatus": "broken","isBroken": true,"redirectChain": [],"finalUrl": "https://crawler-test.com/links/not_found/foo1","error": null,"title": "Broken Links Internal","metaDescription": "Default description m+vb5k7BpjDA9lNDOYXv","h1Count": 1,"canonical": null,"pageIssues": ["missing-canonical"],"startUrl": "https://crawler-test.com/links/broken_links_internal","checkedAt": "2026-10-03T13:31:04.974Z"}
A larger run (oUHPPtvHAlom0gXYQ, input {"startUrls": [{"url": "https://crawler-test.com/"}], "maxPages": 15}) crawled 15 pages and checked 426 links in 4 minutes: 325 ok, 14 redirect, 63 broken, 3 restricted, 2 timeout and 19 blocked-by-robots. This is one of its redirect rows, a link that answers after two 301 hops:
{"pageUrl": "https://crawler-test.com/","linkUrl": "https://crawler-test.com/redirects/redirect_2","anchorText": "Redirect Double 301","type": "internal","rel": null,"statusCode": 200,"linkStatus": "redirect","isBroken": false,"redirectChain": [{"url": "https://crawler-test.com/redirects/redirect_2","statusCode": 301},{"url": "https://crawler-test.com/redirects/redirect_1","statusCode": 301}],"finalUrl": "https://crawler-test.com/redirects/redirect_target","error": null,"title": "Crawler Test Site","metaDescription": "Default description XIbwNE7SSUJciq0/Jyty","h1Count": 1,"canonical": null,"pageIssues": ["missing-canonical"],"startUrl": "https://crawler-test.com/","checkedAt": "2026-10-03T13:15:22.342Z"}
Sitemap mode
Set sitemapUrls to one or more sitemap URLs, for example ["https://www.example.com/sitemap.xml"]. The actor then reads the sitemaps and checks the pages they list, in sitemap order, up to maxPages. It does not follow links to find more pages, so a weekly scheduled run covers the same pages every time: the ones you have told search engines about.
- Sitemap index files are followed into the sitemaps they list, depth first. Gzip sitemaps (
sitemap.xml.gz) are unpacked. - Reading stops as soon as
maxPagespage URLs are found, so a 10-page run on a site with a large sitemap index reads only the first sitemap file. At most 50 sitemap files are read per run. - A bare domain such as
example.commeanshttps://example.com/sitemap.xml. - URLs in a sitemap that point to another website are skipped, as the sitemap protocol only allows URLs on the sitemap's own site.
- Every
<a href>link on each listed page is checked as in a normal crawl. Links to pages that are not in the sitemap are checked but not crawled. - A sitemap should list working, final URLs. A listed URL that is broken, blocked by robots.txt, not HTML or redirected gets a row of its own with the sitemap file as
pageUrl, so you can see which sitemap entries to fix. startUrlsandsitemapUrlscan be combined. The start URLs are then checked as single pages first, followed by the sitemap pages, all withinmaxPages.- Every row has
foundOnSitemap:truewhen the page the link sits on is listed in one of the sitemaps.
Real run on 7 October 2026 (cod1yk9opSv8Wfjwc, 10 pages, 12 rows, 12 seconds) against the test sitemap of crawler-test.com:
{"sitemapUrls": ["https://crawler-test.com/test_sitemap.xml"],"maxPages": 10,"checkImages": true}
Three of the ten sitemap entries redirect. This is one of them, reported with the sitemap as pageUrl and the full chain of three hops:
{"pageUrl": "https://crawler-test.com/test_sitemap.xml","linkUrl": "https://crawler-test.com/redirects/redirect_3_302","linkType": "link","anchorText": null,"type": "internal","rel": null,"statusCode": 200,"linkStatus": "redirect","isBroken": false,"redirectChain": [{"url": "https://crawler-test.com/redirects/redirect_3_302","statusCode": 302},{"url": "https://crawler-test.com/redirects/redirect_2","statusCode": 301},{"url": "https://crawler-test.com/redirects/redirect_1","statusCode": 301}],"finalUrl": "https://crawler-test.com/redirects/redirect_target","error": null,"title": null,"metaDescription": null,"h1Count": null,"canonical": null,"pageIssues": [],"startUrl": "https://crawler-test.com/test_sitemap.xml","foundOnSitemap": true,"checkedAt": "2026-10-07T05:16:16.966Z"}
The same run found a broken link on a page that the test site lists only in its sitemap:
{"pageUrl": "https://crawler-test.com/sitemap/only_in_sitemap/1","linkUrl": "https://crawler-test.com/sitemap/only_in_sitemap_link","linkType": "link","anchorText": "Link in Sitemap only page","type": "internal","statusCode": 404,"linkStatus": "broken","isBroken": true,"foundOnSitemap": true}
(Fields shortened; the other fields are as in the rows above.)
Image check
With checkImages: true every <img src> on a checked page is checked with one HEAD request, and GET when HEAD fails, at the same requestsPerSecond as links. Each image URL appears once per page and is requested once per run. Inline data: images are skipped. Images on other websites, such as a CDN, are checked only when checkExternal is on. Image rows have linkType image and the image's alt text as anchorText (null when the alt text is empty). A missing image gives isBroken: true exactly like a broken link, so brokenOnly and onlyNew work for images too.
Real image row from run K6VavQ8hMq9HcqdNF (input {"startUrls": [{"url": "https://crawler-test.com/links/image_links"}], "maxPages": 1, "checkImages": true}, 3 rows, 4 seconds):
{"pageUrl": "https://crawler-test.com/links/image_links","linkUrl": "https://crawler-test.com/image_link.png","linkType": "image","anchorText": "Image alt tag that is not empty","type": "internal","rel": null,"statusCode": 200,"linkStatus": "ok","isBroken": false,"redirectChain": [],"finalUrl": "https://crawler-test.com/image_link.png","error": null,"title": "Image Links","metaDescription": "Default description IarCjBS9KSdEtfxXQqFU","h1Count": 1,"canonical": null,"pageIssues": ["missing-canonical"],"startUrl": "https://crawler-test.com/links/image_links","foundOnSitemap": false,"checkedAt": "2026-10-07T05:14:34.423Z"}
How long a sitemap run takes depends on how many links the pages have per domain. Run rNcq35zsLYODwktfy with {"sitemapUrls": ["https://butik.nightwave.se/sitemap.xml"], "maxPages": 10, "checkImages": true} checked our own shop's first 10 sitemap pages in 107 seconds and returned 330 rows: one of the pages links to 141 notices on ted.europa.eu, and at 2 requests per second for that one domain those alone take about 70 seconds. The same input with "checkExternal": false (run jCnSDBhBbqZ3Dd7LB) took 38 seconds and returned 183 rows.
Input
| Field | Type | Default | Description |
|---|---|---|---|
startUrls | array | https://butik.nightwave.se/ when sitemapUrls is empty too | Start pages of the sites to crawl. Strings or { "url": ... } objects. A bare domain gets https://. With sitemapUrls they are checked as single pages. |
sitemapUrls | array | none | Sitemap URLs, for example ["https://www.example.com/sitemap.xml"]. When set, the listed pages are checked instead of a crawl. Sitemap indexes and gzip sitemaps are read. |
maxPages | integer | 10 | Maximum HTML pages to check across all start URLs and sitemaps. 1 to 5 000. The low default keeps a first run short; raise it to check a whole site. |
sameDomainOnly | boolean | true | Crawl only the start host. Set to false to include subdomains such as shop.example.com. Other websites are never crawled. |
checkExternal | boolean | true | Also check links that point to other websites. |
brokenOnly | boolean | false | Return only rows where isBroken is true. |
checkImages | boolean | false | Also check every <img src> and return it as a row with linkType image. |
requestsPerSecond | number | 2 | Highest request rate per domain, 0.2 to 5. |
onlyNew | boolean | false | Return only broken links that earlier runs with the same input did not report. See "Monitoring and scheduling". |
A run with empty input crawls https://butik.nightwave.se/ (our own shop) so you can see the output format in a few seconds.
Output
| Field | Description |
|---|---|
pageUrl | The crawled page the link was found on (after redirects). For a sitemap entry that is broken or redirects, the sitemap file |
linkUrl | The link or image, resolved to an absolute URL without #fragment. null on the single row of a page without links |
linkType | link for <a href> links (and sitemap entries), image for <img src> with checkImages |
anchorText | Link text, or the aria-label or image alt text when the link has no text. For image rows, the image's alt text |
type | internal (same site) or external |
rel | The link's rel attribute, for example nofollow |
statusCode | Final HTTP status code after redirects. null when there was no answer |
linkStatus | ok, redirect (works after one or more redirects), broken, restricted (401, 403, 429 or 999: the server refuses automated checks, the page may work in a browser), timeout (no answer in 15 seconds) or blocked-by-robots (not checked because robots.txt disallows it) |
isBroken | true for 4xx and 5xx answers (except the restricted codes), DNS errors, refused connections, TLS errors, redirect loops and more than 10 redirects |
redirectChain | Every redirect hop as { url, statusCode }. Empty when the link answered directly |
finalUrl | Where the link ends up after redirects |
error | Short code when there was no HTTP answer: dns-not-found, connection-refused, connection-reset, tls-error, timeout, redirect-loop, too-many-redirects, robots |
title | <title> of the page |
metaDescription | Meta description of the page |
h1Count | Number of <h1> elements on the page. null on sitemap entry rows |
canonical | Canonical URL of the page (<link rel="canonical">), absolute |
pageIssues | SEO findings for the page: missing-title, title-too-long (over 60 characters), missing-meta-description, meta-description-too-long (over 160), missing-h1, multiple-h1, missing-canonical, noindex |
startUrl | The start URL whose crawl found the page, or the sitemap file that lists it |
foundOnSitemap | true when the page is listed in one of the sitemapUrls. Always false in a crawl without sitemaps |
checkedAt | When the page's links were checked (ISO 8601, UTC) |
Monitoring and scheduling
Set onlyNew to true to use the actor as a broken link monitor. The actor then remembers which broken links it has reported for the same input, in a named key-value store in your Apify account (nightwave-state-website-link-checker, one record per input). A finding is the combination of page, link and status code, so a link that changes from 404 to 500 is reported again. The key is the same whether the page was found by a crawl or listed in a sitemap, and image findings get a key of their own. Each run returns only broken links that earlier runs did not report. A page with at least one new broken link is charged as a normal page (event page), every other crawled page only the monitoring fee (event page-monitored). The first run returns every broken link the crawl finds. A run without news finishes successfully with 0 rows.
onlyNew is not part of the remembered input. Changing any other field, maxPages included, starts a fresh state. To start over with the same input, delete the record in the key-value store.
A second run straight after the first, with onlyNew and the same input, returns 0 rows and charges no page events. The pages it crawls to look for new broken links are still charged the monitoring fee (event page-monitored, see Pricing).
Example: every morning at 06:00 Swedish time, check up to 200 pages of your site and get the new broken links. In Apify Console, open Schedules, create a schedule with the cron expression 0 6 * * * and add this actor with the input below.
{"startUrls": [{ "url": "https://www.example.com/" }],"maxPages": 200,"onlyNew": true}
For a weekly check of every page in your sitemap, including images, use the cron expression 0 6 * * 1 and this input:
{"sitemapUrls": ["https://www.example.com/sitemap.xml"],"maxPages": 500,"checkImages": true,"onlyNew": true}
The same schedule through the Apify API:
curl -X POST "https://api.apify.com/v2/schedules?token=<YOUR_TOKEN>" \-H "Content-Type: application/json" \-d '{"name": "daily-link-check", "cronExpression": "0 6 * * *", "timezone": "Europe/Stockholm", "isEnabled": true, "isExclusive": true,"actions": [{"type": "RUN_ACTOR", "actorId": "nightwave-owner~website-link-checker","runInput": {"contentType": "application/json; charset=utf-8", "body": "<the input above as a JSON string>"}}]}'
Connect a webhook or an integration (Slack, e-mail, Google Sheets) to the actor in Apify Console if you want the new broken links sent somewhere when the run finishes.
Limitations
- Only check sites you own or have permission to check. The actor is built for your own websites and your clients' websites. It follows robots.txt and keeps the request rate low, but you are responsible for having the right to crawl the sites you enter.
- Links rendered by JavaScript are not seen. The actor reads the HTML the server sends, like a search engine's first pass. Single page apps that build their menus in the browser show fewer links.
- Only
<a href>links, and<img src>withcheckImages, are checked.srcsetvariants, CSS background images, scripts, stylesheets and links in PDFs are not. - Plain text sitemaps (
sitemap.txt) and sitemaps listed only in robots.txt are not read. Give the XML sitemap URL insitemapUrls. - Restricted is not broken. LinkedIn (999), many social networks (403, 429) and login pages (401) refuse automated requests. They are reported as
restrictedwithisBroken: false, so your broken link list stays free of false alarms. - Pages larger than 5 MB are checked but not parsed for links.
- Each request has a 15 second timeout. Slow links are reported as
timeout, not as broken. - The crawl stops at
maxPagescrawled HTML pages. Links found on the last pages are still checked, but their targets are not crawled.
Use cases
- Site migration: run before and after, compare the redirect chains and find old URLs that now give 404.
- SEO audits: broken internal links, long redirect chains, pages without meta description or with several H1s, in one table you can filter and export to CSV or Excel.
- Agency maintenance: one scheduled run per client site that reports only new broken links.
- Content teams: find external links in old articles that now point to dead pages.
- Sitemap hygiene: find sitemap entries that redirect or give 404, and broken links on pages that only the sitemap points to.
- Image checks: find images that were deleted from the media library but are still used on a page.
- AI agents (Apify MCP): "check example.com for broken links" works with only
startUrls, andbrokenOnlykeeps the answer short.
FAQ
How long does a run take? About as many seconds as there are unique links on your site divided by 2 (the default rate), since every unique URL is requested once. A 50-page site with 300 unique internal links takes 2-3 minutes. External links are checked in parallel with your own site.
What does a run cost?
0.02 USD per crawled page (event page), so the default 10-page run costs 0.20 USD and a 50-page check costs 1 USD. With onlyNew, pages without a new broken link cost 0.005 USD (event page-monitored). Links and images are not charged separately, however many a page has.
Will it put load on my server?
No more than 2 requests per second per domain by default. Lower requestsPerSecond for small servers, or set a Crawl-delay in robots.txt.
Can I use the results commercially? Yes. The output is information about your own website (status codes and page metadata), produced by your own run. The actor does not copy page content beyond titles, descriptions and anchor texts.
Data source and license
The data is what the websites you enter answer, read live during the run. There is no third-party database behind it, so no data license applies: the output describes your own websites and is yours to use. The rules the actor follows:
- robots.txt according to RFC 9309, the Robots Exclusion Protocol, read on 3 October 2026. Section 2.3.1.3: "If a server status code indicates that the robots.txt file is unavailable to the crawler, then the crawler MAY access any resources on the server." Section 2.3.1.4: "If the robots.txt file is unreachable due to server or network errors, this means the robots.txt file is undefined and the crawler MUST assume complete disallow."
- Apify's General Terms and Conditions, read on 3 October 2026: "You are solely responsible for the legality, accuracy, quality, appropriateness, and use of all Customer Data." The websites you point the actor at, and what you do with the results, are your responsibility.
This actor is not affiliated with any of the websites it checks.
Pricing
Pay per page: 0.02 USD per crawled HTML page with all its links checked (event page), which is 20 USD per 1 000 pages. Links are included, whether a page has 5 or 500. With onlyNew, a crawled page without a new broken link costs 0.005 USD (event page-monitored), so a quiet daily check of a 100-page site costs 0.50 USD. Apify bills platform usage on top as usual; the 15-page test run above used 0.0065 USD. maxPages caps how many pages a run crawls, so you always know the highest possible cost.
Contact
Built and maintained by Nightwave AB. Questions, bugs and feature requests: kontakt@nightwave.se
På svenska
Actorn går igenom de webbplatser du anger och kontrollerar varje länk på varje sida: statuskod, hela kedjan av omdirigeringar, slutadress, om länken är trasig, länktext och om den är intern eller extern. Varje rad har också sidans titel, metabeskrivning, antal H1 och kanonisk adress, plus en lista över SEO-brister på sidan.
- Bara sajterna du anger genomsöks. Länkar till andra sajter kontrolleras med ett enda anrop och följs aldrig vidare.
- Med
sitemapUrls(till exempel["https://www.example.se/sitemap.xml"]) kontrolleras sidorna i sitemapen, upp tillmaxPages, i stället för att sajten genomsöks. Sitemap-index och gzip-sitemaps läses. En adress i sitemapen som är trasig eller omdirigeras får en egen rad med sitemapfilen sompageUrl, och varje rad harfoundOnSitemap. - Med
checkImages: truekontrolleras också varje<img src>och rapporteras som rader medlinkTypeimage. - robots.txt följs för varje adress, också externa länkar (RFC 9309), och Crawl-delay respekteras upp till 10 sekunder.
- Högst 2 anrop per sekund per domän som standard (0,2-5 med
requestsPerSecond). - 401, 403, 429 och 999 räknas som
restricted, inte trasiga, eftersom servern bara vägrar automatiska anrop. - Med
onlyNew: truereturneras bara trasiga länkar som tidigare körningar med samma input inte har rapporterat, och bara sidor med nya fel debiteras (se "Monitoring and scheduling"). - Kontrollera bara sajter som du äger eller har tillstånd att kontrollera.
- Pris: 0,02 USD per genomsökt sida (20 USD per 1 000), alla länkar och bilder på sidan ingår. Med
onlyNewkostar sidor utan nya fel 0,005 USD. - Kontakt: kontakt@nightwave.se