# Sitemap Analyzer - URL Structure, Freshness & SEO Issues (`apisight/sitemap-structure-analyzer`) Actor

Analyse any site's XML sitemaps: find every sitemap, count URLs, map site structure and depth, measure lastmod freshness, and get a ranked list of sitemap problems.

- **URL**: https://apify.com/apisight/sitemap-structure-analyzer.md
- **Developed by:** [Apisight](https://apify.com/apisight) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $50.00 / 1,000 analysed sites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap Analyzer — URL Structure, Freshness & SEO Issues

Point it at a domain. It finds every XML sitemap the site publishes, reads them all
(including gzipped ones and nested sitemap indexes), and returns a structural analysis of
the site plus a ranked list of problems worth fixing.

No sitemap URL needed — discovery happens automatically via `robots.txt` and well-known
paths. You can also pass a sitemap URL directly if you already know it.

### What you get per site

**Inventory** — every sitemap file found, its type (index or urlset), URL count, byte size,
whether it was gzipped, and how it was discovered.

**Scale and structure**

- Total and unique URL counts, and how many URLs are duplicated
- URL count per top-level section — the actual shape of the site
- Depth distribution, max depth and average depth

**Content freshness**

- `lastmod` coverage as a percentage
- URLs modified in the last 7 / 30 / 90 / 365 days, older than a year, or dated in the future
- Oldest and newest `lastmod` dates

**Media and internationalisation** — image, video, Google News and `hreflang` alternate counts.

**Ranked issues** — errors first, then warnings, then informational notes. Each has a stable
`code` so you can filter or track them over time:

| Code | Severity | Meaning |
|---|---|---|
| `sitemap-unreadable` | error | A sitemap returned an error or malformed XML |
| `too-many-urls` | error | Over the sitemaps.org limit of 50,000 URLs in one file |
| `file-too-large` | error | Over the 50 MB uncompressed limit |
| `no-robots-txt` | warning | `robots.txt` unreachable, so crawlers can't discover the sitemap there |
| `sitemap-not-in-robots` | warning | `robots.txt` exists but has no `Sitemap:` line |
| `cross-host-urls` | warning | Sitemap lists URLs on other hosts |
| `mixed-schemes` | warning | Mixes `http://` and `https://` URLs |
| `no-lastmod` | warning | No `<lastmod>` anywhere |
| `duplicate-urls` | warning | The same URL appears more than once |
| `future-lastmod` | warning | `<lastmod>` dated in the future |
| `uniform-priority` | info | Every URL has identical `<priority>`, making it meaningless |
| `inconsistent-trailing-slash` | info | Mixed trailing-slash conventions risk duplicate content |
| `truncated` | info | The `maxUrls` limit stopped the analysis early |

### Input

```json
{
  "startUrls": ["ghost.org", "wordpress.org"],
  "maxUrls": 50000,
  "maxSitemaps": 50,
  "includeUrlList": false,
  "proxyConfiguration": { "useApifyProxy": true }
}
```

| Field | Type | Default | Notes |
|---|---|---|---|
| `startUrls` | array | — | Domains, or a direct sitemap URL |
| `maxUrls` | integer | `50000` | Per site. Results flag whether analysis was truncated |
| `maxSitemaps` | integer | `50` | Caps files fetched when a sitemap index is used |
| `includeUrlList` | boolean | `false` | Adds every URL with its `lastmod`, `changefreq`, `priority` |
| `maxConcurrency` | integer | `5` | Sites analysed in parallel |

### Output (abridged)

```json
{
  "domain": "ghost.org",
  "status": "ok",
  "discoveryMethod": "robots.txt",
  "sitemapCount": 8,
  "sitemapIndexCount": 1,
  "totalUrls": 1003,
  "uniqueUrls": 1003,
  "duplicateUrlCount": 0,
  "lastmodCoveragePct": 100.0,
  "freshness": {
    "last7Days": 41, "last30Days": 88, "last90Days": 122,
    "last365Days": 402, "olderThan1Year": 350, "inTheFuture": 0
  },
  "newestLastmod": "2026-09-30",
  "oldestLastmod": "2019-02-22",
  "maxDepth": 3,
  "avgDepth": 1.98,
  "topSections": { "resources": 428, "themes": 260, "help": 140 },
  "mediaCounts": { "images": 684, "videos": 0, "newsItems": 0, "hreflangAlternates": 0 },
  "errorCount": 0,
  "warningCount": 0,
  "issues": []
}
```

Three output views are provided: **Overview** (one row per site), **Issues** (one row per
problem), and **Sitemap files** (one row per sitemap document).

### Pricing and what you are charged for

| Event | Price | When |
|---|---|---|
| `site-analysed` | $0.05 | Once per site where at least one sitemap was read and analysed |
| `apify-actor-start` | $0.00005 | Once per run |

**You are only charged when a sitemap was actually analysed.** Three outcomes cost you
nothing:

- `status: "no-sitemap"` — the site genuinely publishes no sitemap. That is a real finding,
  and you still get it, but we don't think you should pay for an empty result when you're
  screening a list of domains.
- `status: "blocked"` — bot protection served a challenge instead of the sitemap.
- `status: "error"` — the host never responded.

### Honest limitations

- **Bot protection can block sitemap access.** Some sites put `robots.txt` and sitemaps
  behind a challenge. These are reported as `blocked`, with the reason, and never billed.
- **A site with no discoverable sitemap may still have one** at an unconventional path that
  is neither in `robots.txt` nor one of the well-known locations. Pass the URL directly if
  you know it.
- **`lastmod` is self-reported.** Some CMSs stamp every URL with the build date, which makes
  freshness look better than it is. Treat 100% coverage with identical dates sceptically —
  the freshness buckets make that pattern easy to spot.
- **Nested indexes are followed three levels deep**, which covers essentially every real
  site while bounding runtime.
- **No personal data.** This Actor reads `robots.txt` and XML sitemaps only — public
  technical files. It collects no personal data and is not a contact-scraping tool.

### Common uses

- **SEO audits** — inventory a site's indexable URLs and find sitemap problems before they
  cost you crawl budget.
- **Migration checks** — compare structure and URL counts before and after a replatform.
- **Competitive research** — see how large a competitor's site is and which sections dominate.
- **Content freshness monitoring** — track how much of a site is genuinely being updated.
- **Crawl seeding** — enable `includeUrlList` to get a clean, deduplicated URL list to feed
  another tool.

# Actor input Schema

## `startUrls` (type: `array`):

Domains to analyse. Sitemaps are discovered automatically from robots.txt and well-known paths. You can also pass a sitemap URL directly (e.g. https://example.com/sitemap\_index.xml).

## `maxUrls` (type: `integer`):

Caps how many sitemap URLs are read per site. Raise it for very large sites; results say whether the analysis was truncated.

## `maxSitemaps` (type: `integer`):

Caps how many sitemap documents are fetched when a site uses a sitemap index.

## `includeUrlList` (type: `boolean`):

Adds every extracted URL (with lastmod, changefreq and priority) to the result. Off by default because it makes records very large.

## `proxyConfiguration` (type: `object`):

Routes requests through Apify Proxy. Recommended: some sites rate-limit shared cloud IPs, and blocked sites are never charged.

## `maxConcurrency` (type: `integer`):

How many sites to analyse at once.

## `requestTimeoutSecs` (type: `integer`):

Seconds to wait for each request.

## Actor input object example

```json
{
  "startUrls": [
    "ghost.org",
    "wordpress.org"
  ],
  "maxUrls": 50000,
  "maxSitemaps": 50,
  "includeUrlList": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "maxConcurrency": 5,
  "requestTimeoutSecs": 25
}
```

# Actor output Schema

## `results` (type: `string`):

Per-site sitemap analysis: every sitemap found, total and unique URL counts, lastmod freshness buckets, URL depth and section breakdown, media counts, and a severity-ranked issue list. Sites with no sitemap, blocked sites and unreachable sites are included for transparency and are not charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "ghost.org",
        "wordpress.org"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("apisight/sitemap-structure-analyzer").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        "ghost.org",
        "wordpress.org",
    ],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("apisight/sitemap-structure-analyzer").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "ghost.org",
    "wordpress.org"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call apisight/sitemap-structure-analyzer --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,apisight/sitemap-structure-analyzer"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/qr7h5M53gdiP20eFh/builds/2EGD13xrnXPdBQ80N/openapi.json
