# Sitemap and Broken Link Checker with On-Page SEO Audit (`madrasco/sitemap-health-check`) Actor

Check your sitemap URLs, any list of URLs, or every link on your pages: status, redirect chain, noindex, canonical, robots.txt, and broken links with the page and anchor text they sit on. On-page SEO rules (title, meta description, H1, canonical, alt text, schema, hreflang) with fix hints.

- **URL**: https://apify.com/madrasco/sitemap-health-check.md
- **Developed by:** [Jack Valmadre](https://apify.com/madrasco) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap and Broken Link Checker with On-Page SEO Audit

Find the URLs and links on your site that search engines and visitors trip over, as data you can sort and filter:

- **Broken links on your pages**, each with **the page it sits on and its anchor text**, so you know exactly where to fix it.
- **Sitemap problems**: listed URLs that are broken, redirected, `noindex`, canonicalised to another URL, or blocked by your own robots.txt.
- **Status and redirect chains for any list of URLs**, for example before and after a site migration.
- **On-page SEO issues** on every HTML page it reads: missing or duplicate titles, meta descriptions, H1s and canonicals, missing `lang`, viewport, image `alt` text and structured data, invalid JSON-LD and hreflang values. Each issue comes with a severity and a one-line fix hint.

No login to your site, no browser extension. Give it a sitemap, a list of URLs, or the pages whose links you want checked.

### Pick a mode

| You want to… | Set `mode` to | Give it | You get |
|---|---|---|---|
| Check the URLs in your sitemap | `sitemap` (default) | a sitemap URL or just your site's address; it finds the sitemaps itself (robots.txt `Sitemap:` lines, `/sitemap.xml`, sitemap index files, `.gz` sitemaps) | one row per URL listed in the sitemaps |
| Check a list of URLs (redirect maps, migrations, old links) | `urlList` | the URLs in `urls` (or one per line in `url`) | one row per URL |
| Find broken links on pages | `pageLinks` | page URLs, or a sitemap URL to check the links on every page it lists | one row per unique link or image found, and one row per page with its broken and redirected links |

`pageLinks` does not follow links to discover more pages of your site. To check the links on your whole site, give it your sitemap URL; the pages it lists are read and every unique link on them is checked once.

### What it flags

**URLs and pages** (`problems` on `sitemap` and `urlList` rows, and on page rows in `pageLinks`):

| Code | Meaning |
|---|---|
| `broken` | the URL ends in a 4xx or 5xx status |
| `redirected` | the URL redirects; every hop is listed in `redirectChain` with its status |
| `redirect-loop`, `too-many-redirects` | the redirects go round in a circle, or run past 10 hops |
| `noindex` | `noindex` or `none` in `<meta name="robots">` (or `googlebot` / `bingbot`), or in an `X-Robots-Tag` header |
| `canonical-elsewhere` | the page's `<link rel="canonical">`, or a `Link: rel="canonical"` header (used on PDFs and feeds), points to a different URL |
| `blocked-by-robots`, `redirect-target-blocked-by-robots` | the site's own robots.txt disallows the URL, or a redirect leads into a disallowed path; the URL is reported but not requested |
| `fetch-error` | no usable answer: the name doesn't resolve, the connection fails or times out, the site's security certificate fails (expired, self-signed, incomplete chain), or the address has an invalid port (a URL that cannot be read at all is `invalid-url`) |
| `has-broken-links` | `pageLinks` page rows: at least one link on the page is broken (see `brokenLinks`) |

Also possible: `robots-unreachable` (the site's robots.txt answered with a server error, so its URLs were skipped), `other-host-not-checked`, `invalid-url`, `unexpected-status`.

A sitemap should list only final, indexable, canonical URLs, so every flagged sitemap row is a URL you probably want to fix or take out of the sitemap.

**Links** (`problems` on `pageLinks` link rows):

| Code | Meaning |
|---|---|
| `broken` | 4xx (404, 410 and so on) |
| `server-error` | 5xx |
| `access-denied` | 401, 403 or 429: the linked site refused our checker. Many large sites do this to all automated clients, so check these by hand before removing them |
| `fetch-error` | no usable answer, as above (a site that no longer exists usually shows up here) |
| `redirected`, `redirect-loop`, `too-many-redirects` | as above; `finalUrl` is where the link ends up |
| `blocked-by-robots` | the linked site's robots.txt disallows our checker, so the link was not requested |

`mailto:`, `tel:`, `javascript:` and same-page `#` links are skipped, and so is a link whose address cannot be read as a URL at all. On a page row, `links.broken` counts links that are `broken`, `server-error`, `fetch-error`, `redirect-loop` or `too-many-redirects`; `access-denied` links are counted separately in `links.accessDenied`.

### On-page rules

On every HTML page it reads (in all modes), `seoIssues` lists the rules the page fails, each with a severity, a detail (such as the length found) and a fix hint, and `seo` holds what was found: title, meta description, H1s, `lang`, JSON-LD / microdata types, image and missing-`alt` counts, hreflang codes.

| Rule | Severity | Fails when |
|---|---|---|
| `title-missing` | error | no `<title>`, or it is empty (a `<title>` inside an inline SVG doesn't count) |
| `title-multiple` | warning | more than one `<title>` |
| `title-too-long` / `title-too-short` | notice | over 60 / under 30 characters |
| `meta-description-missing` | warning | no `<meta name="description">`, or it is empty |
| `meta-description-multiple` | warning | more than one |
| `meta-description-too-long` / `meta-description-too-short` | notice | over 155 / under 70 characters |
| `h1-missing` / `h1-multiple` | warning / notice | no `<h1>` / more than one |
| `canonical-missing` / `canonical-multiple` | notice / warning | no `<link rel="canonical">` / several different ones |
| `html-lang-missing` | notice | no `lang` on `<html>` |
| `viewport-missing` | notice | no `<meta name="viewport">` |
| `img-alt-missing` | notice | an `<img>` has no `alt` attribute (`alt=""` is fine) |
| `structured-data-missing` | notice | no JSON-LD, microdata or RDFa |
| `json-ld-invalid` | error | a JSON-LD block is not valid JSON |
| `hreflang-invalid` | warning | an hreflang value is not a language code or `x-default` |

The length limits are the ones common audit tools use; search engines publish no fixed limits. Notices are worth a look, not necessarily a fix: several `<h1>` elements, for example, are valid HTML. Turn the rules off with `checkOnPage: false`.

### Input

| Field | What it does | Default |
|---|---|---|
| `mode` | `sitemap`, `urlList` or `pageLinks` (see above) | `sitemap` |
| `url` | A sitemap, site or page URL; several allowed, separated by spaces or new lines. Also accepted: `urls`, `startUrls`, `sitemapUrls`, `sitemapUrl` | required |
| `maxUrls` | Stop after this many unique URLs (in `pageLinks`: pages) | 5,000 |
| `maxLinks` | `pageLinks`: stop after this many unique links | 2,000 |
| `checkExternalLinks` | `pageLinks`: also check links to other sites (their robots.txt is obeyed) | true |
| `checkImages` | `pageLinks`: also check `<img src>` | true |
| `onlyProblems` | Leave healthy rows out of the dataset | false |
| `checkPageTags` | Read each HTML page (first 1 MB) for meta robots and canonical | true |
| `checkOnPage` | Run the on-page rules on every HTML page read | true |
| `maxSitemaps` | Limit on sitemap files read (indexes included) | 100 |
| `respectRobotsTxt` | Don't request URLs robots.txt disallows (they are still reported) | true |
| `sameHostOnly` | `sitemap` mode: don't request URLs on other hosts than the site | true |
| `delaySecs` | Spacing between requests to one host; a larger robots.txt `Crawl-delay`, up to 10 s, wins | 1 |
| `timeoutSecs` | Give up on a request that takes longer than this | 20 |
| `maxConcurrencyPerHost` | 1 or 2 requests at once per host | 2 |

The prefilled example checks a small demo sitemap we host in Apify storage (a sitemap index, a plain and a gzip sitemap, 6 listed URLs): one page marked noindex, one whose canonical points elsewhere, one deleted page (404), one http:// URL that redirects, and one duplicate. The storage host sends `X-Robots-Tag: none` on every file, so every demo page is also flagged `noindex`. Replace it with your own sitemap or site.

Examples:

```json
{ "url": "https://example.com/sitemap_index.xml", "onlyProblems": true }
```

```json
{ "mode": "urlList", "urls": ["http://example.com/old-page", "https://example.com/pricing"] }
```

```json
{ "mode": "pageLinks", "url": "https://example.com/sitemap.xml", "maxLinks": 5000, "onlyProblems": true }
```

### Output

A broken link found in `pageLinks` mode (a real row from one of our test runs on a public blog post):

```json
{
  "type": "link",
  "url": "http://stackoverflow.com/jobs",
  "ok": false,
  "problems": ["redirected", "broken"],
  "finalStatus": 404,
  "finalUrl": "https://stackoverflow.com/jobs",
  "redirectCount": 1,
  "redirectChain": [{"url": "http://stackoverflow.com/jobs", "status": 301, "location": "https://stackoverflow.com/jobs"}],
  "kind": "link",
  "internal": false,
  "foundOn": [{"page": "https://www.joelonsoftware.com/2000/08/09/the-joel-test-12-steps-to-better-code/", "anchorText": "Stack Overflow Jobs", "nofollow": false}],
  "foundOnCount": 1,
  "error": null
}
```

The page row for that post lists `links` (`found`, `checked`, `broken`, `redirected`, `accessDenied`, `notChecked`), `brokenLinks` and `redirectedLinks` (each with URL, anchor text and final status), plus its own status, canonical, noindex and `seoIssues`.

A `sitemap` or `urlList` row:

```json
{
  "type": "url",
  "url": "http://jekyllrb.com/docs",
  "ok": false,
  "problems": ["redirected", "canonical-elsewhere"],
  "finalStatus": 200,
  "finalUrl": "http://jekyllrb.com/docs/",
  "redirectCount": 1,
  "redirectChain": [{"url": "http://jekyllrb.com/docs", "status": 301, "location": "http://jekyllrb.com/docs/"}],
  "noindex": false,
  "canonicalUrl": "https://jekyllrb.com/docs/",
  "canonicalElsewhere": true,
  "blockedByRobots": false,
  "seoIssues": [
    {"rule": "meta-description-too-long", "severity": "notice", "detail": "222 characters", "fix": "Descriptions over 155 characters are usually cut off."},
    {"rule": "h1-multiple", "severity": "notice", "detail": "2 <h1> elements", "fix": "Several <h1> elements: fine in HTML5, but one clear main heading is the usual advice."}
  ]
}
```

Rows also carry `sitemap` and `lastmod` (sitemap mode), `contentType`, `method`, `error` and `checkedAt`. The last row (`"type": "summary"`, also saved as the `OUTPUT` record) gives the totals: sitemaps read (status, URL count, gzip, parse errors), URLs and links checked, duplicates, caps reached, and counts per problem, per on-page rule and per final status.

### How it checks, and how we test it

- Each URL or link gets a `HEAD` request, or a `GET` if the server refuses `HEAD`. Redirects are followed one hop at a time (up to 10) so every hop is recorded. HTML pages that answer 200 are read (first 1 MB) for meta robots, canonical, the on-page rules and, in `pageLinks`, their links.
- Each unique link is checked once, however many pages it appears on; `foundOn` lists up to 20 of those pages with the anchor text used on each.
- Broken sitemap XML is reported, and the URLs read before the error are still checked.
- Polite by design: robots.txt is read once per site and obeyed, also for links to other sites. At most 2 requests at a time per host, spaced by `delaySecs` or the site's `Crawl-delay`. User agent `MadrascoSitemapHealth`, so site owners can allow or block it by name.
- We test it on public sites we didn't build and compare each result with a second tool: `curl` for statuses and redirects, and a separate implementation of each on-page rule on a different HTML parser; every difference is checked by hand. Latest check (27 Sep 2026, 7 sites on WordPress, Shopify, Ghost, Jekyll, Hugo and Drupal): the on-page rules agreed on 85 of 86 pages (on the other, the site sent a different page to each tool); link statuses and redirect chains agreed with curl on 91 of 97 links (the other 6: two where curl skipped a redirect hop that we recorded, three on a site that rate-limits automated clients and answered each tool differently, one we did not request because robots.txt disallowed it).

### Limits

- It checks what the server returns to an automated HTTP client. It does not run JavaScript, so titles, canonicals, `noindex` or links added by scripts are not seen.
- Some sites answer automated clients differently from browsers: they refuse them (`access-denied`) or send them somewhere else. Check `access-denied` links by hand before removing them.
- A site whose security certificate is incomplete may still open in a browser but is reported as `fetch-error`; the `error` field says which certificate check failed.
- It reports the signals a search engine reads; it can't tell you how a search engine will treat a page, and it gives no ranking advice or SEO score.
- `pageLinks` checks the pages you give it (or the pages your sitemap lists); it does not crawl your site to find pages.
- Pages behind logins or bot walls come back as errors (such as 403); it does not try to get past them.

### Use with AI agents

Call it with `{"url": "<sitemap or site>"}`, `{"mode": "urlList", "urls": [...]}` or `{"mode": "pageLinks", "url": "<page or sitemap>"}`. Read `problems` on each row; rows with `"ok": true` need nothing. The `type: "summary"` row gives the totals.

### Pricing

No charge from us for now: you pay only Apify's platform usage of your run, which is small at the default 256 MB memory. We may add a per-URL price in a later release; any price is shown on this page and by Apify before you start a run.

### Support

Use the Issues tab on this actor's page. Built and maintained by Madrasco, with AI assistance.

# Actor input Schema

## `mode` (type: `string`):

sitemap: find the sitemaps of the site(s) in 'url' and check every URL listed. urlList: check exactly the URLs given (any site, one per line). pageLinks: read the pages given (a sitemap URL is expanded into its pages) and check every link and image on them once: broken, redirected, refused, unreachable, with the page and anchor text each sits on. On-page rules run on every HTML page read in all modes.

## `url` (type: `string`):

sitemap mode: a sitemap (https://example.com/sitemap.xml, .xml.gz, index files welcome) or a website (then sitemaps come from robots.txt Sitemap lines, else /sitemap.xml or /sitemap\_index.xml). urlList mode: the URLs to check. pageLinks mode: the pages whose links to check, or a sitemap of them. Separate several with spaces or new lines.

## `sitemapUrls` (type: `array`):

Optional: more URLs as a list (up to 20 sitemaps/sites, 100 pages in pageLinks mode, 50,000 URLs in urlList mode). Also accepted as 'urls' or 'startUrls'.

## `maxUrls` (type: `integer`):

Stop after this many unique URLs (sitemap and urlList modes) or pages read (pageLinks mode). The summary says when the cap was reached.

## `maxSitemaps` (type: `integer`):

Limit on sitemap files, counting index files and their children.

## `onlyProblems` (type: `boolean`):

Leave healthy URLs out of the dataset (they are still counted in the summary).

## `checkPageTags` (type: `boolean`):

Download the page HTML (first 1 MB) to read <meta name=robots> and <link rel=canonical>. Off: status, redirects and headers only (HEAD requests, faster).

## `respectRobotsTxt` (type: `boolean`):

Don't request URLs the site's robots.txt disallows for our crawler; they are reported as blocked (a real sitemap problem). Turn off only for sites you own.

## `sameHostOnly` (type: `boolean`):

URLs on other hosts are reported but not requested (the www and non-www variants count as the same site).

## `delaySecs` (type: `integer`):

Minimum spacing between request starts to one host; a larger robots.txt Crawl-delay (up to 10 s) wins.

## `maxConcurrencyPerHost` (type: `integer`):

1 or 2. Never more than 2, to stay gentle on the site.

## `timeoutSecs` (type: `integer`):

Per request.

## `maxLinks` (type: `integer`):

pageLinks mode: stop after this many unique links. Each link is checked once however many pages it is on.

## `checkExternalLinks` (type: `boolean`):

Off: only links to the page's own site are checked. Each other site's robots.txt is read and obeyed too.

## `checkImages` (type: `boolean`):

Also check <img src> addresses.

## `checkOnPage` (type: `boolean`):

Title, meta description, H1, canonical, html lang, viewport, image alt, structured data (JSON-LD, microdata, RDFa) and hreflang checks on every HTML page read. Needs 'Check meta robots and canonical tags' on (the page must be downloaded).

## Actor input object example

```json
{
  "mode": "sitemap",
  "url": "https://api.apify.com/v2/key-value-stores/PBsuTupDt43RyGCyc/records/demo-sitemap-index.xml",
  "maxUrls": 5000,
  "maxSitemaps": 100,
  "onlyProblems": false,
  "checkPageTags": true,
  "respectRobotsTxt": true,
  "sameHostOnly": true,
  "delaySecs": 1,
  "maxConcurrencyPerHost": 2,
  "timeoutSecs": 20,
  "maxLinks": 2000,
  "checkExternalLinks": true,
  "checkImages": true,
  "checkOnPage": true
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `links` (type: `string`):

No description

## `pages` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "url": "https://api.apify.com/v2/key-value-stores/PBsuTupDt43RyGCyc/records/demo-sitemap-index.xml"
};

// Run the Actor and wait for it to finish
const run = await client.actor("madrasco/sitemap-health-check").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "url": "https://api.apify.com/v2/key-value-stores/PBsuTupDt43RyGCyc/records/demo-sitemap-index.xml" }

# Run the Actor and wait for it to finish
run = client.actor("madrasco/sitemap-health-check").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "url": "https://api.apify.com/v2/key-value-stores/PBsuTupDt43RyGCyc/records/demo-sitemap-index.xml"
}' |
apify call madrasco/sitemap-health-check --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,madrasco/sitemap-health-check"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/261VfAowYzyhNuxvY/builds/mYmUP2mS2bTbCrIm1/openapi.json
