# Broken Link Checker — SEO Crawl & Monitor (`zenomastro/broken-link-checker-pro`) Actor

Broken link checker and technical SEO crawler for dead links, 4xx/5xx errors, redirects, assets, canonicals and hreflang. Crawl sites or sitemaps, check URLs concurrently, flag slow responses, rank issues, and persist snapshots to surface new, regressed, resolved and changed links.

- **URL**: https://apify.com/zenomastro/broken-link-checker-pro.md
- **Developed by:** [Rosario Vitale](https://apify.com/zenomastro) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 link checks

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Broken Link Checker API — SEO Crawl, Assets & Change Monitor

### Why use this Actor?

Broken link checker and technical SEO crawler for dead links, 4xx/5xx errors, redirects, assets, canonicals and hreflang. Crawl sites or sitemaps, check URLs concurrently, flag slow responses, rank issues, and persist snapshots to surface new, regressed, resolved and changed links.

### Features

- **Start URLs** — Public website pages to crawl for links. Optional when checkUrls or sitemapUrls are supplied.
- **Maximum pages** — Maximum number of HTML pages to crawl across the run.
- **Maximum unique links** — Maximum number of unique HTTP/HTTPS links to check.
- **Crawl same domain only** — When enabled, discovered pages are crawled only on the starting hostname; external links are still checked.
- **Include healthy 2xx links** — Include healthy 2xx link rows in the dataset; otherwise the output focuses on redirects and problems.
- **Request timeout seconds** — Maximum seconds allowed for each network request.
- **User agent** — HTTP User-Agent header sent to target websites.
- **URLs to check directly** — Optional URL list to validate without crawling source pages.
- **Sitemap URLs to import** — Optional XML sitemap URLs whose <loc> targets will be checked.
- **Maximum redirect hops** — Follow and report complete redirect chains up to this limit.
- **Link check concurrency** — Parallel HTTP link checks after discovery.
- **Check page assets** — Also validate images, JavaScript, stylesheets, media, source and iframe URLs discovered on crawled pages.

### Use cases

- Technical seo audits.
- Site migrations.
- Content qa.
- Recurring 404 and redirect monitoring.

### Example input

```json
{
  "startUrls": [
    "https://example.com"
  ],
  "maxPages": 200,
  "maxLinks": 5000,
  "sameDomainOnly": true,
  "includeOk": false,
  "requestTimeoutSecs": 20
}
```

### Pricing & cost control

Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.

### FAQ

**What is this Actor for?**\
It is designed for technical SEO audits, site migrations, content QA.

**Can I run it on a schedule?**\
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

**How do I control cost and run size?**\
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

### Search keywords

broken link checker, broken link checker free, broken link checker extension, broken link checker wordpress, broken link checker tool, broken link checker chrome extension, broken link checker online, broken link checker aioseo, broken link checker ahrefs, broken link checker plugin, 404 checker, 404 checker bulk, 404 checker tool, 404 checker online

Crawl one or more public websites and find broken links, redirects, HTTP errors, and unreachable URLs. The Actor is designed for technical SEO checks, website migrations, QA, content audits, monitoring pipelines, and agency reporting.

### What it checks

The Actor discovers HTTP/HTTPS links from HTML pages, deduplicates them, checks each URL with a lightweight HEAD request and automatically falls back to GET when a server rejects HEAD. Results include source page, target URL, HTTP status, state, redirect target when exposed, whether the link is internal, error text, and timestamps.

### Reliability and cost controls

Crawling is bounded by `maxPages` and `maxLinks`. Requests have explicit timeouts, non-HTTP links are ignored, fragments are removed before deduplication, and duplicate targets are checked only once per run. Multi-page crawls keep going when individual pages or targets fail.

By default only redirects and problematic links are emitted. Enable `includeOk` when you want a full link inventory.

### Input example

```json
{"startUrls":["https://example.com"],"maxPages":50,"maxLinks":1000,"sameDomainOnly":true,"includeOk":false,"requestTimeoutSecs":20}
```

### Output

`link_check` rows contain the individual checks. `page_error` rows describe a source page that could not be crawled. A final `summary` row reports pages scanned and unique links checked.

### Pricing

Target launch price is **$0.00065 per checked emitted link** plus the small Actor start event. Healthy links suppressed by the default `includeOk=false` setting are checked but are not emitted/billed as result rows, making the default mode efficient for finding issues.

### Responsible use

Use this Actor only on public websites and respect applicable terms, robots policies, rate limits, and legal requirements. Keep crawl limits reasonable for the target site.

### Support

For reproducible problems provide the public start URL, relevant input settings and Apify run ID. Never include secrets or private credentials.

### Extended capabilities

- Crawl HTML links, direct URL lists, and XML sitemap targets with redirect-chain reporting.
- Optionally validate assets, canonical URLs, and hreflang targets.
- Flag slow responses, multi-hop redirects, broken assets, and sitemap-only internal orphan candidates.

# Actor input Schema

## `startUrls` (type: `array`):

Public website pages to crawl for links. Optional when checkUrls or sitemapUrls are supplied.

## `maxPages` (type: `integer`):

Maximum number of HTML pages to crawl across the run.

## `maxLinks` (type: `integer`):

Maximum number of unique HTTP/HTTPS links to check.

## `sameDomainOnly` (type: `boolean`):

When enabled, discovered pages are crawled only on the starting hostname; external links are still checked.

## `includeOk` (type: `boolean`):

Include healthy 2xx link rows in the dataset; otherwise the output focuses on redirects and problems.

## `requestTimeoutSecs` (type: `integer`):

Maximum seconds allowed for each network request.

## `userAgent` (type: `string`):

HTTP User-Agent header sent to target websites.

## `checkUrls` (type: `array`):

Optional URL list to validate without crawling source pages.

## `sitemapUrls` (type: `array`):

Optional XML sitemap URLs whose <loc> targets will be checked.

## `maxRedirects` (type: `integer`):

Follow and report complete redirect chains up to this limit.

## `checkConcurrency` (type: `integer`):

Parallel HTTP link checks after discovery.

## `checkAssets` (type: `boolean`):

Also validate images, JavaScript, stylesheets, media, source and iframe URLs discovered on crawled pages.

## `checkSeoLinks` (type: `boolean`):

Validate canonical and hreflang link targets as part of the technical SEO audit.

## `excludePatterns` (type: `array`):

Case-insensitive URL substrings to skip, useful for logout, tracking or intentionally unreachable endpoints.

## `slowResponseMs` (type: `integer`):

Links at or above this end-to-end response time are flagged and summarized as slow.

## `monitorKey` (type: `string`):

Reuse the same key on scheduled runs to compare each checked URL with its previous state.

## `onlyChanges` (type: `boolean`):

With monitorKey, emit only new, regressed, resolved or otherwise changed links. Healthy links that recovered are emitted even when healthy 2xx output is disabled.

## Actor input object example

```json
{
  "startUrls": [
    "https://example.com"
  ],
  "maxPages": 200,
  "maxLinks": 5000,
  "sameDomainOnly": true,
  "includeOk": false,
  "requestTimeoutSecs": 20,
  "userAgent": "Mozilla/5.0 (compatible; ApifyBrokenLinkChecker/1.0)",
  "checkUrls": [],
  "sitemapUrls": [],
  "maxRedirects": 8,
  "checkConcurrency": 12,
  "checkAssets": false,
  "checkSeoLinks": true,
  "excludePatterns": [],
  "slowResponseMs": 3000,
  "monitorKey": "",
  "onlyChanges": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("zenomastro/broken-link-checker-pro").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("zenomastro/broken-link-checker-pro").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call zenomastro/broken-link-checker-pro --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zenomastro/broken-link-checker-pro"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/PA9BeGTAM2WlLjAfb/builds/ZiFzNCjdEHTbz8FYn/openapi.json
