# Website Broken Link Checker: 404s, Redirects, SEO Checks (`nightwave-owner/website-link-checker`) Actor

Returns every link on the websites or sitemaps you list with HTTP status code, redirect chain, final URL, broken flag and anchor text, plus title, meta description, H1 count and canonical per page. Optional image check. Respects robots.txt, 2 requests per second per domain. onlyNew for monitoring.

- **URL**: https://apify.com/nightwave-owner/website-link-checker.md
- **Developed by:** [Viktor Wiberg](https://apify.com/nightwave-owner) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Broken Link Checker: 404s, Redirects, SEO Checks

A website broken link checker: it crawls the sites you give it, checks every link on every crawled page and reports broken links, dead links to other websites and redirect chains. Each row is one link on one page: the HTTP status code, the full redirect chain, the final URL, whether the link is broken, the anchor text and whether it points inside the site or to another website. Each row also carries four SEO basics of the page the link sits on (title, meta description, number of H1 headings, canonical URL) and a short list of page issues such as a missing meta description or two H1s.

Use it before and after a site migration, as a weekly check of a client's site, to find 404s that waste crawl budget, or as a daily monitor that only reports new broken links. Give it a sitemap.xml instead of a start page and it checks exactly the pages you publish; turn on `checkImages` and broken images are reported too.

### How it works

- Starts at each URL in `startUrls` and follows `<a href>` links breadth first, up to `maxPages` HTML pages in total.
- With `sitemapUrls` it does not crawl: the pages listed in the sitemaps (and in any start URLs) are checked as a fixed list, up to `maxPages`. See "Sitemap mode".
- With `checkImages` every `<img src>` on a checked page is checked as well and returned as a row with `linkType` `image`.
- Only the websites you list are crawled. Links to other websites are checked with a single request (HEAD, and GET when HEAD fails) and never followed further. With `sameDomainOnly` (default) only the start host is crawled; `www.example.com` and `example.com` count as the same site.
- robots.txt is read for every host before the first request to it and followed for every URL, including external links (RFC 9309: a missing robots.txt allows everything, a robots.txt that answers with a server error blocks the host, and a host that cannot be reached at all is reported as a broken link). A `Crawl-delay` is honoured up to 10 seconds.
- At most 2 requests per second per domain by default (`requestsPerSecond`, 0.2 to 5). Different domains are checked in parallel, so external links do not slow down the crawl of your own site.
- Redirects are followed by hand, up to 10 hops, so the whole chain is recorded. Redirect loops and too many hops are reported as broken.
- Each URL is requested once per run, however many pages link to it.
- The User-Agent is `Mozilla/5.0 (compatible; NightwaveLinkChecker/1.0; +https://apify.com/nightwave-owner/website-link-checker)`. Site owners can address it in robots.txt as `NightwaveLinkChecker`.

### Example from a real run

These runs were made on the Apify platform on 3 October 2026 against [crawler-test.com](https://crawler-test.com/), a public website built for testing crawlers, with pages that are broken, redirected or slow on purpose. Nothing in the rows is edited.

Input (run `CbcOyZGC7m6gILp7F`, 2 pages, 12 rows, 9 seconds):

```json
{
  "startUrls": [
    { "url": "https://crawler-test.com/links/broken_links_internal" },
    { "url": "https://crawler-test.com/links/broken_links_external" }
  ],
  "maxPages": 2
}
```

One of the five broken internal links it found:

```json
{
  "pageUrl": "https://crawler-test.com/links/broken_links_internal",
  "linkUrl": "https://crawler-test.com/links/not_found/foo1",
  "anchorText": "Broken Internal Link 1",
  "type": "internal",
  "rel": null,
  "statusCode": 404,
  "linkStatus": "broken",
  "isBroken": true,
  "redirectChain": [],
  "finalUrl": "https://crawler-test.com/links/not_found/foo1",
  "error": null,
  "title": "Broken Links Internal",
  "metaDescription": "Default description m+vb5k7BpjDA9lNDOYXv",
  "h1Count": 1,
  "canonical": null,
  "pageIssues": [
    "missing-canonical"
  ],
  "startUrl": "https://crawler-test.com/links/broken_links_internal",
  "checkedAt": "2026-10-03T13:31:04.974Z"
}
```

A larger run (`oUHPPtvHAlom0gXYQ`, input `{"startUrls": [{"url": "https://crawler-test.com/"}], "maxPages": 15}`) crawled 15 pages and checked 426 links in 4 minutes: 325 `ok`, 14 `redirect`, 63 `broken`, 3 `restricted`, 2 `timeout` and 19 `blocked-by-robots`. This is one of its redirect rows, a link that answers after two 301 hops:

```json
{
  "pageUrl": "https://crawler-test.com/",
  "linkUrl": "https://crawler-test.com/redirects/redirect_2",
  "anchorText": "Redirect Double 301",
  "type": "internal",
  "rel": null,
  "statusCode": 200,
  "linkStatus": "redirect",
  "isBroken": false,
  "redirectChain": [
    {
      "url": "https://crawler-test.com/redirects/redirect_2",
      "statusCode": 301
    },
    {
      "url": "https://crawler-test.com/redirects/redirect_1",
      "statusCode": 301
    }
  ],
  "finalUrl": "https://crawler-test.com/redirects/redirect_target",
  "error": null,
  "title": "Crawler Test Site",
  "metaDescription": "Default description XIbwNE7SSUJciq0/Jyty",
  "h1Count": 1,
  "canonical": null,
  "pageIssues": [
    "missing-canonical"
  ],
  "startUrl": "https://crawler-test.com/",
  "checkedAt": "2026-10-03T13:15:22.342Z"
}
```

### Sitemap mode

Set `sitemapUrls` to one or more sitemap URLs, for example `["https://www.example.com/sitemap.xml"]`. The actor then reads the sitemaps and checks the pages they list, in sitemap order, up to `maxPages`. It does not follow links to find more pages, so a weekly scheduled run covers the same pages every time: the ones you have told search engines about.

- Sitemap index files are followed into the sitemaps they list, depth first. Gzip sitemaps (`sitemap.xml.gz`) are unpacked.
- Reading stops as soon as `maxPages` page URLs are found, so a 10-page run on a site with a large sitemap index reads only the first sitemap file. At most 50 sitemap files are read per run.
- A bare domain such as `example.com` means `https://example.com/sitemap.xml`.
- URLs in a sitemap that point to another website are skipped, as the sitemap protocol only allows URLs on the sitemap's own site.
- Every `<a href>` link on each listed page is checked as in a normal crawl. Links to pages that are not in the sitemap are checked but not crawled.
- A sitemap should list working, final URLs. A listed URL that is broken, blocked by robots.txt, not HTML or redirected gets a row of its own with the sitemap file as `pageUrl`, so you can see which sitemap entries to fix.
- `startUrls` and `sitemapUrls` can be combined. The start URLs are then checked as single pages first, followed by the sitemap pages, all within `maxPages`.
- Every row has `foundOnSitemap`: `true` when the page the link sits on is listed in one of the sitemaps.

Real run on 7 October 2026 (`cod1yk9opSv8Wfjwc`, 10 pages, 12 rows, 12 seconds) against the test sitemap of crawler-test.com:

```json
{
  "sitemapUrls": ["https://crawler-test.com/test_sitemap.xml"],
  "maxPages": 10,
  "checkImages": true
}
```

Three of the ten sitemap entries redirect. This is one of them, reported with the sitemap as `pageUrl` and the full chain of three hops:

```json
{
  "pageUrl": "https://crawler-test.com/test_sitemap.xml",
  "linkUrl": "https://crawler-test.com/redirects/redirect_3_302",
  "linkType": "link",
  "anchorText": null,
  "type": "internal",
  "rel": null,
  "statusCode": 200,
  "linkStatus": "redirect",
  "isBroken": false,
  "redirectChain": [
    {
      "url": "https://crawler-test.com/redirects/redirect_3_302",
      "statusCode": 302
    },
    {
      "url": "https://crawler-test.com/redirects/redirect_2",
      "statusCode": 301
    },
    {
      "url": "https://crawler-test.com/redirects/redirect_1",
      "statusCode": 301
    }
  ],
  "finalUrl": "https://crawler-test.com/redirects/redirect_target",
  "error": null,
  "title": null,
  "metaDescription": null,
  "h1Count": null,
  "canonical": null,
  "pageIssues": [],
  "startUrl": "https://crawler-test.com/test_sitemap.xml",
  "foundOnSitemap": true,
  "checkedAt": "2026-10-07T05:16:16.966Z"
}
```

The same run found a broken link on a page that the test site lists only in its sitemap:

```json
{
  "pageUrl": "https://crawler-test.com/sitemap/only_in_sitemap/1",
  "linkUrl": "https://crawler-test.com/sitemap/only_in_sitemap_link",
  "linkType": "link",
  "anchorText": "Link in Sitemap only page",
  "type": "internal",
  "statusCode": 404,
  "linkStatus": "broken",
  "isBroken": true,
  "foundOnSitemap": true
}
```

(Fields shortened; the other fields are as in the rows above.)

### Image check

With `checkImages: true` every `<img src>` on a checked page is checked with one HEAD request, and GET when HEAD fails, at the same `requestsPerSecond` as links. Each image URL appears once per page and is requested once per run. Inline `data:` images are skipped. Images on other websites, such as a CDN, are checked only when `checkExternal` is on. Image rows have `linkType` `image` and the image's `alt` text as `anchorText` (`null` when the alt text is empty). A missing image gives `isBroken: true` exactly like a broken link, so `brokenOnly` and `onlyNew` work for images too.

Real image row from run `K6VavQ8hMq9HcqdNF` (input `{"startUrls": [{"url": "https://crawler-test.com/links/image_links"}], "maxPages": 1, "checkImages": true}`, 3 rows, 4 seconds):

```json
{
  "pageUrl": "https://crawler-test.com/links/image_links",
  "linkUrl": "https://crawler-test.com/image_link.png",
  "linkType": "image",
  "anchorText": "Image alt tag that is not empty",
  "type": "internal",
  "rel": null,
  "statusCode": 200,
  "linkStatus": "ok",
  "isBroken": false,
  "redirectChain": [],
  "finalUrl": "https://crawler-test.com/image_link.png",
  "error": null,
  "title": "Image Links",
  "metaDescription": "Default description IarCjBS9KSdEtfxXQqFU",
  "h1Count": 1,
  "canonical": null,
  "pageIssues": [
    "missing-canonical"
  ],
  "startUrl": "https://crawler-test.com/links/image_links",
  "foundOnSitemap": false,
  "checkedAt": "2026-10-07T05:14:34.423Z"
}
```

How long a sitemap run takes depends on how many links the pages have per domain. Run `rNcq35zsLYODwktfy` with `{"sitemapUrls": ["https://butik.nightwave.se/sitemap.xml"], "maxPages": 10, "checkImages": true}` checked our own shop's first 10 sitemap pages in 107 seconds and returned 330 rows: one of the pages links to 141 notices on ted.europa.eu, and at 2 requests per second for that one domain those alone take about 70 seconds. The same input with `"checkExternal": false` (run `jCnSDBhBbqZ3Dd7LB`) took 38 seconds and returned 183 rows.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `startUrls` | array | `https://butik.nightwave.se/` when `sitemapUrls` is empty too | Start pages of the sites to crawl. Strings or `{ "url": ... }` objects. A bare domain gets `https://`. With `sitemapUrls` they are checked as single pages. |
| `sitemapUrls` | array | none | Sitemap URLs, for example `["https://www.example.com/sitemap.xml"]`. When set, the listed pages are checked instead of a crawl. Sitemap indexes and gzip sitemaps are read. |
| `maxPages` | integer | `10` | Maximum HTML pages to check across all start URLs and sitemaps. 1 to 5 000. The low default keeps a first run short; raise it to check a whole site. |
| `sameDomainOnly` | boolean | `true` | Crawl only the start host. Set to `false` to include subdomains such as `shop.example.com`. Other websites are never crawled. |
| `checkExternal` | boolean | `true` | Also check links that point to other websites. |
| `brokenOnly` | boolean | `false` | Return only rows where `isBroken` is `true`. |
| `checkImages` | boolean | `false` | Also check every `<img src>` and return it as a row with `linkType` `image`. |
| `requestsPerSecond` | number | `2` | Highest request rate per domain, 0.2 to 5. |
| `onlyNew` | boolean | `false` | Return only broken links that earlier runs with the same input did not report. See "Monitoring and scheduling". |

A run with empty input crawls https://butik.nightwave.se/ (our own shop) so you can see the output format in a few seconds.

### Output

| Field | Description |
|---|---|
| `pageUrl` | The crawled page the link was found on (after redirects). For a sitemap entry that is broken or redirects, the sitemap file |
| `linkUrl` | The link or image, resolved to an absolute URL without `#fragment`. `null` on the single row of a page without links |
| `linkType` | `link` for `<a href>` links (and sitemap entries), `image` for `<img src>` with `checkImages` |
| `anchorText` | Link text, or the `aria-label` or image `alt` text when the link has no text. For image rows, the image's `alt` text |
| `type` | `internal` (same site) or `external` |
| `rel` | The link's `rel` attribute, for example `nofollow` |
| `statusCode` | Final HTTP status code after redirects. `null` when there was no answer |
| `linkStatus` | `ok`, `redirect` (works after one or more redirects), `broken`, `restricted` (401, 403, 429 or 999: the server refuses automated checks, the page may work in a browser), `timeout` (no answer in 15 seconds) or `blocked-by-robots` (not checked because robots.txt disallows it) |
| `isBroken` | `true` for 4xx and 5xx answers (except the restricted codes), DNS errors, refused connections, TLS errors, redirect loops and more than 10 redirects |
| `redirectChain` | Every redirect hop as `{ url, statusCode }`. Empty when the link answered directly |
| `finalUrl` | Where the link ends up after redirects |
| `error` | Short code when there was no HTTP answer: `dns-not-found`, `connection-refused`, `connection-reset`, `tls-error`, `timeout`, `redirect-loop`, `too-many-redirects`, `robots` |
| `title` | `<title>` of the page |
| `metaDescription` | Meta description of the page |
| `h1Count` | Number of `<h1>` elements on the page. `null` on sitemap entry rows |
| `canonical` | Canonical URL of the page (`<link rel="canonical">`), absolute |
| `pageIssues` | SEO findings for the page: `missing-title`, `title-too-long` (over 60 characters), `missing-meta-description`, `meta-description-too-long` (over 160), `missing-h1`, `multiple-h1`, `missing-canonical`, `noindex` |
| `startUrl` | The start URL whose crawl found the page, or the sitemap file that lists it |
| `foundOnSitemap` | `true` when the page is listed in one of the `sitemapUrls`. Always `false` in a crawl without sitemaps |
| `checkedAt` | When the page's links were checked (ISO 8601, UTC) |

### Monitoring and scheduling

Set `onlyNew` to `true` to use the actor as a broken link monitor. The actor then remembers which broken links it has reported for the same input, in a named key-value store in your Apify account (`nightwave-state-website-link-checker`, one record per input). A finding is the combination of page, link and status code, so a link that changes from 404 to 500 is reported again. The key is the same whether the page was found by a crawl or listed in a sitemap, and image findings get a key of their own. Each run returns only broken links that earlier runs did not report. A page with at least one new broken link is charged as a normal page (event `page`), every other crawled page only the monitoring fee (event `page-monitored`). The first run returns every broken link the crawl finds. A run without news finishes successfully with 0 rows.

`onlyNew` is not part of the remembered input. Changing any other field, `maxPages` included, starts a fresh state. To start over with the same input, delete the record in the key-value store.

A second run straight after the first, with `onlyNew` and the same input, returns 0 rows and charges no `page` events. The pages it crawls to look for new broken links are still charged the monitoring fee (event `page-monitored`, see Pricing).

Example: every morning at 06:00 Swedish time, check up to 200 pages of your site and get the new broken links. In Apify Console, open **Schedules**, create a schedule with the cron expression `0 6 * * *` and add this actor with the input below.

```json
{
  "startUrls": [{ "url": "https://www.example.com/" }],
  "maxPages": 200,
  "onlyNew": true
}
```

For a weekly check of every page in your sitemap, including images, use the cron expression `0 6 * * 1` and this input:

```json
{
  "sitemapUrls": ["https://www.example.com/sitemap.xml"],
  "maxPages": 500,
  "checkImages": true,
  "onlyNew": true
}
```

The same schedule through the Apify API:

```sh
curl -X POST "https://api.apify.com/v2/schedules?token=<YOUR_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"name": "daily-link-check", "cronExpression": "0 6 * * *", "timezone": "Europe/Stockholm", "isEnabled": true, "isExclusive": true,
       "actions": [{"type": "RUN_ACTOR", "actorId": "nightwave-owner~website-link-checker",
                    "runInput": {"contentType": "application/json; charset=utf-8", "body": "<the input above as a JSON string>"}}]}'
```

Connect a webhook or an integration (Slack, e-mail, Google Sheets) to the actor in Apify Console if you want the new broken links sent somewhere when the run finishes.

### Limitations

- **Only check sites you own or have permission to check.** The actor is built for your own websites and your clients' websites. It follows robots.txt and keeps the request rate low, but you are responsible for having the right to crawl the sites you enter.
- **Links rendered by JavaScript are not seen.** The actor reads the HTML the server sends, like a search engine's first pass. Single page apps that build their menus in the browser show fewer links.
- **Only `<a href>` links, and `<img src>` with `checkImages`, are checked.** `srcset` variants, CSS background images, scripts, stylesheets and links in PDFs are not.
- **Plain text sitemaps** (`sitemap.txt`) and sitemaps listed only in robots.txt are not read. Give the XML sitemap URL in `sitemapUrls`.
- **Restricted is not broken.** LinkedIn (999), many social networks (403, 429) and login pages (401) refuse automated requests. They are reported as `restricted` with `isBroken: false`, so your broken link list stays free of false alarms.
- **Pages larger than 5 MB** are checked but not parsed for links.
- Each request has a 15 second timeout. Slow links are reported as `timeout`, not as broken.
- The crawl stops at `maxPages` crawled HTML pages. Links found on the last pages are still checked, but their targets are not crawled.

### Use cases

- Site migration: run before and after, compare the redirect chains and find old URLs that now give 404.
- SEO audits: broken internal links, long redirect chains, pages without meta description or with several H1s, in one table you can filter and export to CSV or Excel.
- Agency maintenance: one scheduled run per client site that reports only new broken links.
- Content teams: find external links in old articles that now point to dead pages.
- Sitemap hygiene: find sitemap entries that redirect or give 404, and broken links on pages that only the sitemap points to.
- Image checks: find images that were deleted from the media library but are still used on a page.
- AI agents (Apify MCP): "check example.com for broken links" works with only `startUrls`, and `brokenOnly` keeps the answer short.

### FAQ

**How long does a run take?**
About as many seconds as there are unique links on your site divided by 2 (the default rate), since every unique URL is requested once. A 50-page site with 300 unique internal links takes 2-3 minutes. External links are checked in parallel with your own site.

**What does a run cost?**
0.02 USD per crawled page (event `page`), so the default 10-page run costs 0.20 USD and a 50-page check costs 1 USD. With `onlyNew`, pages without a new broken link cost 0.005 USD (event `page-monitored`). Links and images are not charged separately, however many a page has.

**Will it put load on my server?**
No more than 2 requests per second per domain by default. Lower `requestsPerSecond` for small servers, or set a `Crawl-delay` in robots.txt.

**Can I use the results commercially?**
Yes. The output is information about your own website (status codes and page metadata), produced by your own run. The actor does not copy page content beyond titles, descriptions and anchor texts.

### Data source and license

The data is what the websites you enter answer, read live during the run. There is no third-party database behind it, so no data license applies: the output describes your own websites and is yours to use. The rules the actor follows:

- robots.txt according to [RFC 9309, the Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309), read on 3 October 2026. Section 2.3.1.3: "If a server status code indicates that the robots.txt file is unavailable to the crawler, then the crawler MAY access any resources on the server." Section 2.3.1.4: "If the robots.txt file is unreachable due to server or network errors, this means the robots.txt file is undefined and the crawler MUST assume complete disallow."
- Apify's [General Terms and Conditions](https://docs.apify.com/legal/general-terms-and-conditions), read on 3 October 2026: "You are solely responsible for the legality, accuracy, quality, appropriateness, and use of all Customer Data." The websites you point the actor at, and what you do with the results, are your responsibility.

This actor is not affiliated with any of the websites it checks.

### Pricing

Pay per page: 0.02 USD per crawled HTML page with all its links checked (event `page`), which is 20 USD per 1 000 pages. Links are included, whether a page has 5 or 500. With `onlyNew`, a crawled page without a new broken link costs 0.005 USD (event `page-monitored`), so a quiet daily check of a 100-page site costs 0.50 USD. Apify bills platform usage on top as usual; the 15-page test run above used 0.0065 USD. `maxPages` caps how many pages a run crawls, so you always know the highest possible cost.

### Contact

Built and maintained by Nightwave AB. Questions, bugs and feature requests: kontakt@nightwave.se

### På svenska

Actorn går igenom de webbplatser du anger och kontrollerar varje länk på varje sida: statuskod, hela kedjan av omdirigeringar, slutadress, om länken är trasig, länktext och om den är intern eller extern. Varje rad har också sidans titel, metabeskrivning, antal H1 och kanonisk adress, plus en lista över SEO-brister på sidan.

- Bara sajterna du anger genomsöks. Länkar till andra sajter kontrolleras med ett enda anrop och följs aldrig vidare.
- Med `sitemapUrls` (till exempel `["https://www.example.se/sitemap.xml"]`) kontrolleras sidorna i sitemapen, upp till `maxPages`, i stället för att sajten genomsöks. Sitemap-index och gzip-sitemaps läses. En adress i sitemapen som är trasig eller omdirigeras får en egen rad med sitemapfilen som `pageUrl`, och varje rad har `foundOnSitemap`.
- Med `checkImages: true` kontrolleras också varje `<img src>` och rapporteras som rader med `linkType` `image`.
- robots.txt följs för varje adress, också externa länkar (RFC 9309), och Crawl-delay respekteras upp till 10 sekunder.
- Högst 2 anrop per sekund per domän som standard (0,2-5 med `requestsPerSecond`).
- 401, 403, 429 och 999 räknas som `restricted`, inte trasiga, eftersom servern bara vägrar automatiska anrop.
- Med `onlyNew: true` returneras bara trasiga länkar som tidigare körningar med samma input inte har rapporterat, och bara sidor med nya fel debiteras (se "Monitoring and scheduling").
- Kontrollera bara sajter som du äger eller har tillstånd att kontrollera.
- Pris: 0,02 USD per genomsökt sida (20 USD per 1 000), alla länkar och bilder på sidan ingår. Med `onlyNew` kostar sidor utan nya fel 0,005 USD.
- Kontakt: kontakt@nightwave.se

# Actor input Schema

## `startUrls` (type: `array`):

Start pages of the websites to crawl, for example \[{"url": "https://www.example.com/"}]. Only these sites are crawled; links to other sites are checked once but never crawled. A bare domain such as example.com gets https://. With sitemapUrls, the start URLs are checked as single pages and not crawled. When both fields are empty the actor checks https://butik.nightwave.se/ so an empty run shows how the output looks.

## `sitemapUrls` (type: `array`):

Sitemap URLs, for example \["https://butik.nightwave.se/sitemap.xml"]. When set, the pages listed in the sitemaps are checked (up to maxPages) instead of crawling from startUrls, so a scheduled check covers exactly the pages you publish. Sitemap index files and gzip sitemaps (.xml.gz) are read too. A bare domain such as example.com means https://example.com/sitemap.xml. Sitemap entries that are broken or redirect are reported as rows with the sitemap as pageUrl. Leave empty to crawl from startUrls.

## `maxPages` (type: `integer`):

Maximum number of HTML pages to check across all start URLs and sitemaps, for example 10. Every link on a checked page is checked, and each checked page is one billable event. Defaults to 10 so a first run finishes in well under a minute; raise it for a full site check.

## `sameDomainOnly` (type: `boolean`):

When true, only pages on the start URL's host are crawled (www.example.com and example.com count as the same). When false, subdomains such as shop.example.com are crawled too. Other websites are never crawled either way. Defaults to true.

## `checkExternal` (type: `boolean`):

Also check the status of links that point to other websites, for example true. Each external URL gets one HEAD request (GET if HEAD fails) and is not crawled. Defaults to true.

## `brokenOnly` (type: `boolean`):

Return only rows where isBroken is true (broken links, and broken images with checkImages), for example true. Pages are still checked and charged. Defaults to false, which returns every link with its status.

## `requestsPerSecond` (type: `number`):

Highest request rate against any one domain, for example 2. Keep it low on small servers. A Crawl-delay in robots.txt slows it further. 0.2 to 5, defaults to 2.

## `onlyNew` (type: `boolean`):

For scheduled runs. When true, only broken links that an earlier run with the same input did not report are returned. Pages with a new broken link are charged as usual, other crawled pages only a small monitoring fee. The first run returns every broken link. Defaults to false.

## `checkImages` (type: `boolean`):

Also check every <img src> on the checked pages, for example true. Each image URL gets one HEAD request (GET if HEAD fails) at the same requestsPerSecond, and is returned as a row with linkType "image" and the alt text as anchorText. Images on other websites are checked only with checkExternal. Defaults to false.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.example.com/"
    }
  ],
  "sitemapUrls": [
    "https://butik.nightwave.se/sitemap.xml"
  ],
  "maxPages": 10,
  "sameDomainOnly": true,
  "checkExternal": true,
  "brokenOnly": true,
  "requestsPerSecond": 2,
  "onlyNew": true,
  "checkImages": true
}
```

# Actor output Schema

## `results` (type: `string`):

One row per link per crawled page, as JSON. Open in Apify Console or download via the dataset API.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://butik.nightwave.se/"
        }
    ],
    "maxPages": 10,
    "sameDomainOnly": true,
    "checkExternal": true,
    "brokenOnly": false,
    "requestsPerSecond": 2,
    "onlyNew": false,
    "checkImages": false
};

// Run the Actor and wait for it to finish
const run = await client.actor("nightwave-owner/website-link-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://butik.nightwave.se/" }],
    "maxPages": 10,
    "sameDomainOnly": True,
    "checkExternal": True,
    "brokenOnly": False,
    "requestsPerSecond": 2,
    "onlyNew": False,
    "checkImages": False,
}

# Run the Actor and wait for it to finish
run = client.actor("nightwave-owner/website-link-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://butik.nightwave.se/"
    }
  ],
  "maxPages": 10,
  "sameDomainOnly": true,
  "checkExternal": true,
  "brokenOnly": false,
  "requestsPerSecond": 2,
  "onlyNew": false,
  "checkImages": false
}' |
apify call nightwave-owner/website-link-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nightwave-owner/website-link-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/8Q8WRdLhhzwtkvFUg/builds/Ad4zw80q4NJvFLPZN/openapi.json
