# Sitemap SEO Audit & Monitor - Broken Links & Hreflang (`prevailing_glow/sitemap-seo-audit-monitor`) Actor

Audit every page in your sitemap: broken links, redirects, noindex, canonicals, titles, meta descriptions, hreflang, AI-crawler robots.txt. Alerts on changes.

- **URL**: https://apify.com/prevailing\_glow/sitemap-seo-audit-monitor.md
- **Developed by:** [Dave West](https://apify.com/prevailing_glow) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 page audits

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap SEO Audit & Monitor: broken links, redirects, meta tags, hreflang and AI crawlers

**Run a technical SEO audit on every page in your sitemap and get alerts when something breaks.** Give it a website, a sitemap URL or a list of pages. It finds the sitemap, checks each page for the issues that hurt indexing and rankings, and scores the site from 0 to 100. It also checks which AI crawlers robots.txt allows and whether `/llms.txt` exists. Schedule it to see on every run which issues are **new**, which you **fixed**, and which pages **started returning 404** or **became noindex**.

- One dataset row per page, listing its issues with a severity and a plain-English message.
- A site summary (JSON) and a self-contained **HTML report**.
- A change log between runs, so you can use it as an **SEO monitoring** tool.
- Pay only for pages audited: **$0.003 per page** plus **$0.02 per run** for the site report.

### Use cases

- **Technical SEO audit** of a client site in minutes, with no desktop crawler to install.
- **Sitemap hygiene**: find sitemap URLs that return 404, redirect, are noindex or canonicalised elsewhere. These waste crawl budget and confuse search engines.
- **Website monitoring after releases and migrations**: schedule a daily or weekly run and catch new 404s, accidental `noindex` tags, broken canonicals and redirect chains before Google does.
- **Broken link checker**: deduplicated broken internal links (and external links if you turn that on), with the pages that link to them.
- **International SEO**: hreflang code validation, `x-default`, return links (reciprocity) and hreflang targets that don't return 200.
- **AI search readiness**: see whether GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Google-Extended, CCBot and Applebot-Extended may crawl the site, and whether `/llms.txt` exists.
- **Agencies and freelancers**: one run per client site, with an HTML report you can share.

### What is checked

| Area | Checks (issue codes) |
|---|---|
| Status & speed | 4xx / 5xx (`HTTP_4XX`, `HTTP_5XX`), network errors (`FETCH_FAILED`), `RATE_LIMITED`, slow responses over 3 s (`SLOW_RESPONSE`), HTML over 1 MB (`PAGE_TOO_LARGE`), non-HTML URLs (`NOT_HTML`) |
| Redirects | `REDIRECTED`, chains of 2+ hops (`REDIRECT_CHAIN`), `REDIRECT_LOOP`, `TOO_MANY_REDIRECTS`, HTTPS to HTTP (`HTTPS_TO_HTTP_REDIRECT`) |
| Indexability | noindex from meta robots or `X-Robots-Tag` (`NOINDEX`), blocked for Googlebot by robots.txt (`BLOCKED_BY_ROBOTS_TXT`) |
| Canonical | missing, multiple, pointing elsewhere, canonical target not 200 (`CANONICAL_*`) |
| Title & meta description | missing, too short (<30 / <70), too long (>60 / >160), duplicated across the site (`TITLE_*`, `META_DESCRIPTION_*`) |
| Headings & language | `H1_MISSING`, `H1_MULTIPLE`, `HTML_LANG_MISSING`, `HTML_LANG_INVALID` |
| Hreflang | invalid codes (e.g. `en-UK`, `jp`), missing `x-default`, missing self-reference, conflicting codes, missing return links between audited pages, targets that don't return 200 (`HREFLANG_*`) |
| Links | broken internal links, broken external links (optional), internal links to redirecting URLs |
| Content | images without `alt`, mixed content (http:// resources on https pages) |
| Sitemap hygiene | sitemap URLs that aren't 200, redirect, are noindex, are canonicalised elsewhere or are blocked by robots.txt (`SITEMAP_URL_*`) |
| Site level | robots.txt (found, unreachable, Crawl-delay), AI crawler access table, `/llms.txt`, sitemap discovery (robots.txt `Sitemap:` lines, common locations, sitemap indexes, gzip) |

Every issue has a fixed **severity**:

- **error**: breaks indexing or the user experience. Fix these first.
- **warning**: likely hurts rankings, snippets or crawl efficiency.
- **notice**: worth a look, but often intentional.

### Health score (0-100)

```
pageScore   = 100 - 15 x errors - 5 x warnings - 1 x notices   (distinct issue codes on the page, floored at 0)
healthScore = round( average pageScore of all fetched pages - site penalties ), clamped to 0..100
site penalties: -5 robots.txt unreachable, -5 no sitemap found, -5 homepage blocked for Googlebot
```

Pages that weren't fetched because robots.txt blocks this crawler are left out of the average. The formula is also stored in every `SUMMARY` (`healthScoreFormula`).

### Quick start

1. Enter your website (e.g. `https://www.example.com`) in **Website or sitemap URL**.
2. Set **Max pages**. Large sitemaps are sampled evenly across their sitemap files.
3. Click **Start**. A 30-page audit usually finishes in 10-40 seconds.
4. Open the **HTML report** in the Output tab, or the dataset for per-page details.
5. To monitor a site, **schedule** the Actor (daily or weekly). From the second run on, every row has a `change` flag and the summary lists new and fixed issues.

Leave the input empty for a **free preview** on a small example site. Preview runs are never charged.

Want just the list of URLs in a sitemap, or alerts when pages are added or removed? Use [Sitemap Extractor & Change Monitor](https://apify.com/prevailing_glow/sitemap-extractor-monitor).

### Input example

```json
{
    "startUrl": "https://crawlee.dev",
    "maxPages": 30,
    "checkExternalLinks": false,
    "maxLinkChecks": 1000,
    "compareWithPreviousRun": true
}
```

| Field | Default | Notes |
|---|---|---|
| `startUrl` | - | A homepage or a sitemap / sitemap index URL. One site per run. |
| `urlList` | - | Audit exactly these pages instead of the sitemap. |
| `maxPages` | 500 | 1-5,000. Each audited page is one billable event. |
| `checkExternalLinks` | false | Also check links to other sites. |
| `maxLinkChecks` | 1000 | Distinct link targets fetched for the broken-link check. Pages that were already audited are free. |
| `compareWithPreviousRun` | true | Store state per site in a named key-value store and report changes. |
| `stateStoreName` | `sitemap-seo-audit-state` | Use different names to keep separate histories. |
| `followLinksIfNoSitemap` | true | Crawl internal links from the homepage when there is no sitemap. |
| `respectRobotsTxt` | true | Skip disallowed URLs (reported for free) and honour Crawl-delay up to 10 s. |
| `maxConcurrency` / `minDelayBetweenRequestsMs` | 4 / 150 | Polite defaults. |

### Output

**Dataset**: one row per page. Views: Overview, Indexability, Content, Links & performance, Changes. Below is a real row from crawlee.dev, shortened:

```json
{
    "url": "https://crawlee.dev/blog/crawlee-for-python-v1",
    "statusCode": 200,
    "responseTimeMs": 569,
    "pageSizeBytes": 156820,
    "indexable": true,
    "canonicalStatus": "self",
    "title": "Crawlee for Python v1 | Crawlee for JavaScript · Build reliable crawlers. Fast.",
    "titleLength": 79,
    "metaDescriptionLength": 47,
    "h1Count": 1,
    "htmlLang": "en",
    "hreflang": [{ "lang": "en", "href": "https://crawlee.dev/blog/crawlee-for-python-v1", "validCode": true, "targetStatus": 200, "reciprocal": true }],
    "brokenInternalLinkCount": 1,
    "brokenInternalLinks": [{ "url": "https://www.crawlee.dev/python/docs/examples/playwright-crawler-with-fingeprint-generator", "status": 404 }],
    "imagesMissingAlt": 0,
    "issues": [
        { "code": "BROKEN_INTERNAL_LINKS", "severity": "error", "message": "1 internal link(s) are broken, e.g. https://www.crawlee.dev/python/docs/examples/playwright-crawler-with-fingeprint-generator (404)." },
        { "code": "TITLE_TOO_LONG", "severity": "warning", "message": "Title is 79 characters and may be cut off in results (aim for at most 60)." },
        { "code": "META_DESCRIPTION_TOO_SHORT", "severity": "notice", "message": "Meta description is 47 characters (aim for 70-160)." }
    ],
    "pageScore": 78,
    "change": { "status": "existing", "changed": false, "newIssues": [], "fixedIssues": [], "new404": false, "newlyNoindexed": false, "previousStatusCode": 200 }
}
```

**Key-value store**:

- `SUMMARY`: health score and formula, page counts and status codes, issue counts by code and severity, top issues, duplicate titles and descriptions, the AI crawler table, llms.txt, robots.txt, sitemap files, link-check stats, and `changes` (new and fixed issues, new 404s, newly noindexed pages, new and removed pages, change in score).
- `OUTPUT.html`: a self-contained report you can open in a browser or send to a client.

Excerpt from a real `SUMMARY` (crawlee.dev, 30 pages, second run):

```json
{
    "healthScore": 92,
    "pages": { "audited": 30, "indexable": 30, "statusCodes": { "200": 30 }, "avgResponseTimeMs": 549 },
    "issuesBySeverity": { "error": 1, "warning": 39, "notice": 17 },
    "topIssues": [{ "code": "BROKEN_INTERNAL_LINKS", "severity": "error", "pages": 1 }, { "code": "TITLE_TOO_LONG", "severity": "warning", "pages": 30 }],
    "aiCrawlers": [{ "bot": "GPTBot", "owner": "OpenAI", "access": "allowed", "hasSpecificRules": false }],
    "llmsTxt": { "exists": true, "httpStatus": 200 },
    "changes": { "comparedWithPreviousRun": true, "newIssues": [], "fixedIssues": [], "new404s": [], "newlyNoindexed": [] }
}
```

### Pricing

Pay per event, with no subscription:

| Event | Price | When |
|---|---|---|
| `page-audited` | **$0.003** ($3 per 1,000 pages) | each audited page saved to the dataset |
| `site-report` | **$0.02** | once per run, after the summary and report are built. Only charged if at least one page was billed |

- **Examples**: 30 pages = $0.11. 500 pages = $1.52. 5,000 pages = $15.02.
- **Free rows**: pages that robots.txt blocks (not fetched) are listed but not charged. Preview runs (empty input) cost nothing.
- **Max cost per run**: the Actor respects the limit you set. It audits only as many pages as fit, always keeps room for the site report, and stops gracefully.

### FAQ

**How is the sitemap found?** Through the `Sitemap:` lines in robots.txt, then common locations (`/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml`, and more). Sitemap indexes, gzip sitemaps, text sitemaps and RSS/Atom feeds are supported. If no sitemap exists, the Actor follows internal links from the homepage.

**What happens with very large sitemaps?** Up to `maxPages` URLs are sampled evenly across all sitemap files, so every section of the site is represented. The summary shows how many URLs the sitemaps contain.

**Does it render JavaScript?** No. It audits the HTML your server sends, which is what search engines index first. Pages that build titles or links only in the browser may show missing titles or few links.

**Will it overload my server?** No. Defaults are 4 parallel requests and at least 150 ms between requests. It honours robots.txt (including Crawl-delay) and backs off on 429 and 503 responses.

**How does change monitoring work?** After each run, a small state record per site is stored in a named key-value store. The next run compares each page's status code, noindex and issue codes with it. Use a different `stateStoreName` for separate histories.

**Why is a page "blocked by robots.txt" but not charged?** When robots.txt disallows this crawler, the page isn't fetched. It is still listed so you can see it, but you don't pay for it.

**What if I abort a run?** Pages audited before the abort are still saved and appear in the summary and report. Link checks may be incomplete. The site report isn't charged and the monitoring state isn't updated.

**Can I audit several sites?** Yes, with one run per site. Create a task per site and schedule them.

### Support

Found a false positive or missing a check? Open an issue on the Actor's Issues tab and include the URL and the issue code.

# Actor input Schema

## `startUrl` (type: `string`):

A homepage (e.g. <code>https://example.com</code>; the sitemap is found via robots.txt and common locations) or a sitemap / sitemap index URL. One site per run.

## `urlList` (type: `array`):

Audit exactly these pages instead of the sitemap. All URLs must be on one site; others are ignored.

## `maxPages` (type: `integer`):

Maximum number of pages to audit. Large sitemaps are sampled evenly across their sitemap files. Each audited page is one billable event.

## `checkExternalLinks` (type: `boolean`):

Also check links to other websites (slower). Internal links are always checked.

## `maxLinkChecks` (type: `integer`):

Maximum number of distinct link targets fetched for the broken-link check (targets that were audited as pages are free). 0 disables extra link requests.

## `compareWithPreviousRun` (type: `boolean`):

Store this run's results per site and report new issues, fixed issues, new 404s and newly noindexed pages compared with the previous run. Schedule the Actor to monitor a site.

## `stateStoreName` (type: `string`):

Named key-value store that keeps the per-site state between runs. Use different names to keep separate histories.

## `followLinksIfNoSitemap` (type: `boolean`):

When the site has no sitemap, discover pages by following internal links from the homepage.

## `respectRobotsTxt` (type: `boolean`):

Do not fetch pages that robots.txt disallows for this crawler, and honour Crawl-delay (up to 10 s). Blocked pages are reported for free.

## `maxConcurrency` (type: `integer`):

Parallel requests to the site. Keep it low to be polite.

## `minDelayBetweenRequestsMs` (type: `integer`):

Minimum time between two requests to the same host.

## `requestTimeoutSecs` (type: `integer`):

Timeout for a single request.

## `maxRetries` (type: `integer`):

Retries for network errors, 429 and 503 responses.

## `userAgent` (type: `string`):

User-Agent header. robots.txt rules are matched against its product token. Leave empty for the default.

## `proxyConfiguration` (type: `object`):

Optional proxy. Usually not needed for your own site.

## Actor input object example

```json
{
  "startUrl": "https://crawlee.dev",
  "maxPages": 10,
  "checkExternalLinks": false,
  "maxLinkChecks": 100,
  "compareWithPreviousRun": true,
  "stateStoreName": "sitemap-seo-audit-state",
  "followLinksIfNoSitemap": true,
  "respectRobotsTxt": true,
  "maxConcurrency": 4,
  "minDelayBetweenRequestsMs": 150,
  "requestTimeoutSecs": 20,
  "maxRetries": 2,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `report` (type: `string`):

Health score, top issues, AI crawlers, duplicates and changes in one page.

## `overview` (type: `string`):

One row per audited page: status, indexability, score and issue codes.

## `pages` (type: `string`):

Full per-page audit rows.

## `summary` (type: `string`):

Health score, issue counts, AI-crawler table, duplicates and changes since the previous run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrl": "https://crawlee.dev",
    "maxPages": 10,
    "maxLinkChecks": 100
};

// Run the Actor and wait for it to finish
const run = await client.actor("prevailing_glow/sitemap-seo-audit-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrl": "https://crawlee.dev",
    "maxPages": 10,
    "maxLinkChecks": 100,
}

# Run the Actor and wait for it to finish
run = client.actor("prevailing_glow/sitemap-seo-audit-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrl": "https://crawlee.dev",
  "maxPages": 10,
  "maxLinkChecks": 100
}' |
apify call prevailing_glow/sitemap-seo-audit-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,prevailing_glow/sitemap-seo-audit-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/DfmCkOemAmNGF3LuN/builds/91jjkPz10q7QYuIja/openapi.json
