# Audit Website URL Structure and Sitemap

**Use case:** 

Compare crawled WordPress URLs with XML sitemap evidence, HTTP status, canonicals, indexability, and orphan candidates.

## Input

```json
{
  "startUrls": [
    {
      "url": "https://wordpress.org/news/"
    }
  ],
  "crawlScope": "same-domain",
  "includeUrlPatterns": [],
  "excludeUrlPatterns": [],
  "queryParameterPolicy": "drop-tracking",
  "assetPolicy": "pages-and-documents",
  "discoverSitemaps": true,
  "maxSitemapsPerStartUrl": 8,
  "maxSitemapUrlsPerStartUrl": 25,
  "renderMode": "http",
  "browserFallbackMaxPages": 10,
  "followNofollow": true,
  "maxResults": 25,
  "maxPagesPerStartUrl": 10,
  "maxDepth": 1,
  "maxLinksPerPage": 500,
  "maxEdges": 100000,
  "maxParentUrlsPerRecord": 50,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 30,
  "maxRedirects": 10,
  "maxResponseBytes": 5000000,
  "requestDelayMillis": 1000,
  "maxRunSeconds": 150,
  "retryCount": 2,
  "saveUrlInventoryCsv": true,
  "saveLinkGraphCsv": true,
  "saveSitemapDiffCsv": true,
  "generateXmlSitemap": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

## Output

```json
{
  "start_url": {
    "label": "Start URL",
    "format": "string"
  },
  "url": {
    "label": "URL",
    "format": "string"
  },
  "source_classification": {
    "label": "Source Classification",
    "format": "string"
  },
  "crawl_status": {
    "label": "Crawl Status",
    "format": "string"
  },
  "depth": {
    "label": "Depth",
    "format": "integer"
  },
  "status_code": {
    "label": "HTTP Status",
    "format": "integer"
  },
  "indexability": {
    "label": "Indexability",
    "format": "string"
  },
  "inbound_link_count": {
    "label": "Inbound Link Count",
    "format": "integer"
  },
  "parent_url": {
    "label": "Primary Parent URL",
    "format": "string"
  },
  "canonical_url": {
    "label": "Canonical URL",
    "format": "string"
  },
  "sitemap_lastmod": {
    "label": "Sitemap Last Modified",
    "format": "string"
  },
  "issue_codes": {
    "label": "Issue Codes",
    "format": "array"
  },
  "scraped_at": {
    "label": "Scraped At",
    "format": "string"
  }
}
```

## About this Actor

This example demonstrates how to use [Website Crawl Map & Sitemap Diff](https://apify.com/harvestlab/website-crawl-map.md) with a specific input configuration. Visit the [Actor detail page](https://apify.com/harvestlab/website-crawl-map.md) to learn more, explore other use cases, and run it yourself.


## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
This Task's input is already configured above — use it as-is rather than inventing a new one.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For full API examples (JavaScript, Python, CLI, MCP, OpenAPI), see this Task's Actor page: https://apify.com/harvestlab/website-crawl-map.md

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).
