# Every page of a SaaS company site

**Use case:** 

One row per URL: the address, its host and path, how many levels deep it is, the file extension, and the last-modified date, change frequency and priority the sitemap states. Measured on 2026-09-07: Stripe declares its sitemap in robots.txt at a path that is not /sitemap.xml, and it fans out into nine separate files. This Actor follows that automatically.

## Input

```json
{
  "sites": [
    "stripe.com",
    "vercel.com",
    "openai.com",
    "www.cloudflare.com"
  ],
  "sitesText": "",
  "sitemapUrls": [],
  "useCommonPaths": true,
  "maxUrls": 2000,
  "maxUrlsPerSite": 500,
  "maxSitemaps": 50,
  "maxSitemapDepth": 3,
  "timeoutSecs": 20,
  "pathIncludes": [],
  "pathExcludes": [],
  "extensions": [],
  "lastmodAfter": "",
  "lastmodBefore": "",
  "minPriority": 0,
  "changefreqIn": [],
  "keywords": [],
  "keywordMatch": "any",
  "excludeKeywords": [],
  "monitoringMode": false,
  "resetMonitoringState": false
}
```

## Output

```json
{
  "site": {
    "label": "site",
    "format": "string"
  },
  "url": {
    "label": "url",
    "format": "string"
  },
  "host": {
    "label": "host",
    "format": "string"
  },
  "path": {
    "label": "path",
    "format": "string"
  },
  "depth": {
    "label": "depth",
    "format": "string"
  },
  "extension": {
    "label": "extension",
    "format": "string"
  },
  "lastmod": {
    "label": "lastmod",
    "format": "string"
  },
  "lastmodAt": {
    "label": "lastmodAt",
    "format": "string"
  },
  "changefreq": {
    "label": "changefreq",
    "format": "string"
  },
  "priority": {
    "label": "priority",
    "format": "string"
  },
  "foundVia": {
    "label": "foundVia",
    "format": "string"
  },
  "sitemapUrl": {
    "label": "sitemapUrl",
    "format": "string"
  },
  "status": {
    "label": "status",
    "format": "string"
  },
  "source": {
    "label": "source",
    "format": "string"
  },
  "sitemapIsGzipped": {
    "label": "sitemapIsGzipped",
    "format": "string"
  },
  "sitemapDepth": {
    "label": "sitemapDepth",
    "format": "string"
  },
  "scrapedAt": {
    "label": "scrapedAt",
    "format": "string"
  },
  "urlKey": {
    "label": "urlKey",
    "format": "string"
  }
}
```

## About this Actor

This example demonstrates how to use [Sitemap Scraper - Find the Sitemap and Watch for New URLs](https://apify.com/neverempty/sitemap-finder-monitor.md) with a specific input configuration. Visit the [Actor detail page](https://apify.com/neverempty/sitemap-finder-monitor.md) to learn more, explore other use cases, and run it yourself.


## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
This Task's input is already configured above — use it as-is rather than inventing a new one.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For full API examples (JavaScript, Python, CLI, MCP, OpenAPI), see this Task's Actor page: https://apify.com/neverempty/sitemap-finder-monitor.md

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).
