# SEO Audit Tool: Sitemap & Indexability (`axiorasolutions/sitemap-seo-audit`) Actor

Extract every URL from a site's sitemaps and audit indexability. Reads robots.txt, follows sitemap index files, then checks HTTP status, redirects, canonical tags, noindex, hreflang, titles, meta descriptions, headings, image alt text and internal links. Every finding carries severity and evidence.

- **URL**: https://apify.com/axiorasolutions/sitemap-seo-audit.md
- **Developed by:** [Axiora Solutions](https://apify.com/axiorasolutions) (community)
- **Categories:** SEO tools, Marketing, Automation
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.80 / 1,000 page audits

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap SEO Audit — extract sitemaps & check indexability

Sitemap SEO audit in one run: this Actor reads a site's `robots.txt`, discovers and recursively follows its sitemaps, then audits every URL for indexability — with no Search Console access, signup or API key required. You get one dataset row per URL with `isIndexable`, HTTP status and redirect chains, canonical checks, `noindex` directives, title and meta lengths, hreflang sets, image alt coverage and internal links. It works on competitor sites and client sites you have not been granted access to.

### What you get

- `isIndexable` — true only when the page returns 200, carries no `noindex`, and its canonical resolves to itself.
- `findings[]` — every issue with `id`, `severity` and an `evidence` string, sorted most severe first.
- `httpStatus`, `redirectChain`, `responseMs` and `contentBytes` — the full HTTP layer for each URL.
- `canonicalUrl` / `canonicalIsSelf`, `title`, `metaDescription`, `h1Count` and `hreflang` — content signals.
- `imageCount` / `imagesMissingAltCount`, `internalLinkCount` / `externalLinkCount`, `structuredDataTypes`.
- `recordHash` plus `highSeverityCount` / `mediumSeverityCount` / `lowSeverityCount` for diffing and triage.

### Quick start

1. Open the Actor and leave the prefilled `apify.com` in **Sites**, or enter your own domain, full URL or explicit sitemap URL.
2. Choose **Mode**: `extract-urls` to size a site cheaply first, `audit` to fetch and audit every URL.
3. Optionally set **Max URLs per site** and **Max URLs for the whole run** to control your bill, then click **Start**.
4. Open the **Audit results**, **Findings** and **Indexability** dataset tabs.

Minimal input:

```json
{
  "sites": ["apify.com"],
  "mode": "audit"
}
```

### Example output

One representative dataset row for one audited page:

```json
{
  "recordType": "page",
  "ok": true,
  "errorCode": null,
  "site": "example.com",
  "url": "https://example.com/products/widget",
  "sitemapLastmod": "2026-09-30",
  "httpStatus": 200,
  "finalUrl": "https://example.com/products/widget",
  "redirectCount": 1,
  "redirectChain": [{ "status": 301, "to": "https://example.com/products/widget" }],
  "contentType": "text/html",
  "responseMs": 412,
  "contentBytes": 142331,
  "isIndexable": false,
  "findingCount": 3,
  "highSeverityCount": 1,
  "mediumSeverityCount": 1,
  "lowSeverityCount": 1,
  "findings": [
    { "id": "canonical-differs", "severity": "high", "message": "Canonical points to a different domain (www.example.com), so this URL will not be indexed.", "evidence": "https://www.example.com/products/widget" },
    { "id": "h1-multiple", "severity": "medium", "message": "3 h1 elements found.", "evidence": "Widget | Buy now | Related" },
    { "id": "images-missing-alt", "severity": "low", "message": "2 of 9 images have no alt text (22%).", "evidence": null }
  ],
  "metaRobots": "index, follow",
  "xRobotsTag": null,
  "canonicalUrl": "https://www.example.com/products/widget",
  "canonicalIsSelf": false,
  "title": "Widget - Example",
  "titleLength": 16,
  "metaDescription": "Buy the Widget, shipped worldwide.",
  "metaDescriptionLength": 34,
  "h1Count": 3,
  "h1": "Widget",
  "hreflang": [],
  "imageCount": 9,
  "imagesMissingAltCount": 2,
  "internalLinkCount": 118,
  "externalLinkCount": 12,
  "structuredDataTypes": ["Product", "BreadcrumbList"],
  "mainEntity": "Product",
  "recordHash": "91ab7c3e0d5f2846"
}
```

### Why this is different from a page-speed or crawl tool

Most SEO tools either need your own verified property or give you a score you cannot audit. This one gives you **findings with evidence strings** — every issue cites the actual markup or header that caused it:

```json
{
  "id": "canonical-differs",
  "severity": "high",
  "message": "Canonical points to a different domain (example.com), so this URL will not be indexed.",
  "evidence": "https://example.com/store"
}
```

### What this SEO audit tool checks

- 🔎 **Sitemap discovery** — `robots.txt` `Sitemap:` declarations first, then eight conventional paths (`/sitemap.xml`, `/wp-sitemap.xml`, `/sitemap_index.xml`, …). Sitemap **index** files are followed recursively, bounded by your file limit. Plain-text and HTML sitemaps are handled too.
- 🚦 **HTTP layer** — status codes, redirect chains with each hop, cross-domain redirects, response time, response size and truncation.
- 🛑 **Indexability, the headline metric** — `isIndexable` is true **only** when the page returns 200, carries no `noindex`, and its canonical resolves to itself. A `noindex` URL that is *also listed in the sitemap* raises `noindex-in-sitemap`, because those two signals contradict each other and Google resolves the conflict against you.
- 🔗 **Canonical analysis** — missing canonicals, cross-domain canonicals, and self-referencing checks that ignore `utm_` parameters and trailing slashes.
- 🏷️ **Meta and content** — title presence and length, meta description presence and length, h1 count, multiple or missing h1, JSON-LD `@type` inventory.
- 🌍 **Hreflang** — declared alternates with resolved URLs, duplicate codes, missing `x-default`.
- 🖼️ **Images** — total count, images missing alt text, with the offending URLs.
- 🔗 **Links** — deduplicated internal and external links per page, ready to feed into a broken-link report.
- 🔒 **Hygiene** — missing HSTS and Content-Security-Policy headers.

Every finding is sorted most severe first, and every row carries `highSeverityCount` / `mediumSeverityCount` / `lowSeverityCount` so you can triage a 50,000-URL audit with one sort.

### Two modes

| Mode | What it does | Cost |
|---|---|---|
| **Extract URLs only** | Collects the sitemap URL list without fetching pages | Very cheap, per URL |
| **Audit every URL** | Fetches and audits each URL | Per page |

Start with **Extract URLs only** to size a site, then run **Audit every URL** against a filtered subset. The `urlPattern` and `excludeUrlPattern` fields let you audit just `/blog/`, or skip everything with query parameters.

### How to use it

1. Add sites to **Sites**. A domain, a full URL, or an explicit sitemap URL all work. Passing a sitemap URL skips discovery.
2. Choose **Mode**.
3. Set **Max URLs per site** and **Max URLs for the whole run** — these control your bill.
4. Optionally narrow with **Include URL pattern** (`/blog/`) and **Exclude URL pattern** (`\?`).
5. Click **Start**, then use the **Audit results**, **Findings** and **Indexability** dataset tabs.
6. Schedule it weekly and diff `recordHash` to catch regressions.

### How much does it cost to audit a sitemap?

Two events, and you only ever pay one of them per URL:

| Event | What triggers it | Billed |
|---|---|---|
| Page audit | One URL fetched and audited in **Audit every URL** mode | per page |
| URL discovered | One URL extracted from a sitemap without being fetched | per URL |
| Actor start | Once per run, platform fee | per run |

A URL is charged **either** as a discovery **or** as an audit, never both. URLs skipped because of `robots.txt`, a failed fetch, or your include/exclude patterns are **not** billed as page audits. Sites where no sitemap was found are **not** billed.

1,000 URLs audited is 1,000 page-audit events — that is the whole calculation. Compute, bandwidth and storage are included; there is no separate platform-usage charge on top.

Set **Max cost per run** in the run options for a hard ceiling. Higher Apify plans get progressively lower per-page pricing through Apify Store tier discounts.

Evaluating? Run **Extract URLs only** first — it costs a fraction of a full audit and tells you exactly how many pages you would be buying.

### Example input

```json
{
  "sites": ["example.com", "https://www.example.com/sitemap_index.xml"],
  "mode": "audit",
  "maxUrlsPerSite": 1000,
  "maxUrlsTotal": 3000,
  "maxSitemapsPerSite": 30,
  "excludeUrlPattern": "\\?",
  "respectRobotsTxt": true,
  "checkInternalLinks": true,
  "includeImageAlt": true
}
```

### Use cases

- **Agency technical audits** — a client-ready **Findings** view with severity and evidence, generated on demand.
- **Migration QA** — run before and after a replatform and diff `recordHash` per URL to catch canonical and noindex mistakes.
- **Competitor teardown** — indexability and content-gap analysis on sites you cannot access in Search Console.
- **Programmatic SEO monitoring** — schedule weekly, alert when `isIndexable` flips to false on any templated page.
- **Broken-link hunting** — take `internalLinks` from this dataset and feed them into the **URL Status Checker**.
- **Content audits at scale** — title and meta description length distributions across an entire site.

### Related Actors by Axiora Solutions

| Actor | Use it for |
|---|---|
| **URL Status Checker** | Bulk-check the `internalLinks` from this dataset for 404s and redirect chains |
| **Domain Contact Enricher** | Turn a prospect list into contactable records before pitching the audit |
| **Shopify Product & Variant Scraper** | Catalogue and pricing data for the storefronts you audit |

### Frequently asked questions

#### Does it need Google Search Console?

No, and that is the point. Everything comes from public signals: `robots.txt`, sitemaps and the pages themselves. It works on any site on the internet, including competitors and prospects.

#### How does it decide `isIndexable`?

Three conditions, all required: the final HTTP status is 200, no `noindex` appears in `meta robots`, `meta googlebot` or the `X-Robots-Tag` header, and the canonical — if one is declared — resolves to the same URL after normalising `www`, tracking parameters, a trailing slash and `/index.html`. All three components are on the row as separate fields, so you can override the verdict if your definition differs.

#### What if the site has no sitemap?

You get an `ok: false` row with `EMPTY_RESULT` explaining that `robots.txt` and eight conventional paths were checked, and nothing is billed. Go back a step with **Extract URLs only** off and confirm by opening `/sitemap.xml`. If the site genuinely has none, this Actor is not the right tool — a link-following crawler is.

#### Does it run JavaScript?

No. It reads server-rendered HTML, which is what search engines see first and what makes a 10,000-URL audit affordable. Sites that render content entirely client-side will show short titles and empty meta descriptions; that is itself a finding worth knowing.

#### What does `respectRobotsTxt` cost me?

Coverage. URLs that `robots.txt` disallows are counted in `robotsSkipped` and are not audited or billed as page audits. Leave it on unless you own the site or have written permission. It is the safer default and the honest one.

#### Why is `sitemapLastmod` not validated?

Because a large share of sites publish incorrect `lastmod` values, including dates in the future. The Actor reports what the sitemap claims and separately computes `responseMs` and `recordHash` from live observation. Trust the hash, not the claim.

#### Can it check page speed?

It records `responseMs` (server response time) and `contentBytes` as observed. It does not run Lighthouse: real Core Web Vitals need a headless browser, which would multiply the cost of an audit by orders of magnitude. If you need a performance audit as well, run a dedicated tool alongside this one.

#### Something looks wrong — how do I report it?

Open the **Issues** tab on this Actor page with the site, the URL and the finding ID you disagreed with. A false-positive finding is treated as a bug.

***

Runnable examples and how-to guides for these Actors: [github.com/batow133/axiora-apify-actors](https://github.com/batow133/axiora-apify-actors)

# Actor input Schema

## `sites` (type: `array`):

One site per entry. A bare domain, a full URL or an explicit sitemap URL all work: example.com, https://example.com, https://example.com/sitemap\_index.xml. A sitemap URL is used directly and skips discovery.

## `mode` (type: `string`):

extract-urls only collects the URL list and is very cheap. audit fetches every URL and reports indexability, canonicals, meta and links. Use extract-urls first to size a site before paying for a full audit.

## `maxUrlsPerSite` (type: `integer`):

Cap the number of URLs taken from a site's sitemaps. This is the number that ultimately controls your bill in audit mode.

## `maxUrlsTotal` (type: `integer`):

Hard ceiling across all sites. The run stops cleanly when it is reached.

## `maxSitemapsPerSite` (type: `integer`):

Limit how many sitemap files are followed, including nested sitemap index files. Large publishers can have thousands.

## `urlPattern` (type: `string`):

Keep only URLs matching this regular expression, for example /blog/ to audit just the blog. Leave empty to audit everything.

## `excludeUrlPattern` (type: `string`):

Drop URLs matching this regular expression, for example ? to skip anything with query parameters.

## `respectRobotsTxt` (type: `boolean`):

Skip URLs the site disallows for automated clients. Leave this on unless you own the site or have permission. URLs skipped this way are reported, not audited, and are not billed as page audits.

## `checkInternalLinks` (type: `boolean`):

Record internal and external links found on each audited page, so you can build a broken-link report from the same dataset. Adds no requests.

## `includeImageAlt` (type: `boolean`):

Count images and list those missing alt text, up to your chosen limit.

## `maxImagesReported` (type: `integer`):

Cap the list of image URLs on each page so the dataset stays manageable.

## `requestTimeoutSecs` (type: `integer`):

Give up on a single page after this many seconds. Slow pages are reported as TIMEOUT findings rather than failing the run.

## `proxyConfiguration` (type: `object`):

Optional. Auditing reports the status the server actually returns, so running without a proxy gives the truest result. Enable datacenter proxy rotation only if a site starts rate-limiting a large audit.

## Actor input object example

```json
{
  "sites": [
    "example.com",
    "https://example.com",
    "https://www.example.com/sitemap_index.xml"
  ],
  "mode": "audit",
  "maxUrlsPerSite": 500,
  "maxUrlsTotal": 2000,
  "maxSitemapsPerSite": 30,
  "urlPattern": "/blog/",
  "excludeUrlPattern": "\\?",
  "respectRobotsTxt": true,
  "checkInternalLinks": true,
  "includeImageAlt": true,
  "maxImagesReported": 50,
  "requestTimeoutSecs": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `pages` (type: `string`):

One row per audited page, or per sitemap URL in extraction mode. Filter on recordType.

## `runSummary` (type: `string`):

Per-site sitemaps parsed, URLs found, URLs audited, robots.txt skips, fetch failures, billing and network totals.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sites": [
        "apify.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("axiorasolutions/sitemap-seo-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "sites": ["apify.com"] }

# Run the Actor and wait for it to finish
run = client.actor("axiorasolutions/sitemap-seo-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sites": [
    "apify.com"
  ]
}' |
apify call axiorasolutions/sitemap-seo-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,axiorasolutions/sitemap-seo-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/UG6vZTi6C3mNSstER/builds/l7bbuy8rWw3FQ0jY3/openapi.json
