# Help Center Sitemap Extractor (`junipr/help-center-sitemap-extractor`) Actor

Extract public help center sitemap URLs, article URLs, categories, locales, last-modified signals, crawl depth, and indexability clues into a clean help-center inventory.

- **URL**: https://apify.com/junipr/help-center-sitemap-extractor.md
- **Developed by:** [junipr](https://apify.com/junipr) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.90 / 1,000 sitemap file checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Help Center Sitemap Extractor

### Store Positioning

**Store title:** Help Center Sitemap Extractor

**Short description:** Extract public help center sitemap URLs, article URLs, categories, locales, last-modified signals, crawl depth, and indexability clues into a clean help-center inventory.

**SEO title:** Help Center Sitemap Extractor — technical SEO, web, and domain audit

**SEO description:** Extract public help center sitemap URLs, article URLs, categories, locales, last-modified signals, crawl depth, and indexability clues into a clean help-center inventory. Use it to find crawlability, indexability, security, metadata, and page-quality issues with evidence-backed rows and audit reports.

**Categories:** SEO\_TOOLS, DEVELOPER\_TOOLS

**Keywords:** help, center, sitemap, extractor, public data, local seo, web/domain audit

### Pay-Per-Event Pricing

This actor uses pay-per-event pricing. Event prices include Apify platform usage; users are not expected to pay a separate platform-usage pass-through charge for the configured pricing model.

- Tier: W1 — Web/domain audit
- Primary event: `sitemap-file-checked` at $0.00490 base
- Default max charge: $10.00
- Store discounts: FREE/BRONZE base, SILVER discounted, GOLD deepest approved discount

Event set:

- `actor-start`: base $0.00500, GOLD $0.00400. Help Center Sitemap Extractor: charged when actor start is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `sitemap-file-checked`: base $0.00490, GOLD $0.00392. Help Center Sitemap Extractor: charged when sitemap file checked is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `record-extracted`: base $0.00372, GOLD $0.00298. Help Center Sitemap Extractor: charged when record extracted is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `finding-emitted`: base $0.00372, GOLD $0.00298. Help Center Sitemap Extractor: charged when finding emitted is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `audit-report-generated`: base $0.05000, GOLD $0.04000. Help Center Sitemap Extractor: charged when audit report generated is completed. The price includes Apify platform usage; no separate usage pass-through is intended.

The actor accepts `actor-start` before work, accepts the primary event before each dataset row, and accepts the configured report event before writing report files. If `maxChargeUsd` or the live PPE limit blocks a charge, the corresponding row or report is not written.

### Public Task Concepts

- Audit Help Center Sitemap controls on a capped public sample
- Find high-priority Help Center Sitemap issues before release
- Validate Help Center Sitemap evidence from supplied pages
- Prioritize Help Center Sitemap fixes with severity and proof
- Export Help Center Sitemap QA rows for client review

Extract public help center sitemap URLs, article URLs, categories, locales, last-modified signals, crawl depth, and indexability clues into a clean help-center inventory.

### What it does

- Accept help center base URLs, sitemap URLs, robots.txt URLs, or raw sitemap XML.
- Discover sitemap indexes, article sitemaps, category pages, section pages, locale paths, and support article URLs.
- Extract URL type, title where available, locale, lastmod, changefreq, priority, depth, status code, canonical hints, and source sitemap.
- Deduplicate URLs and classify article/category/section/search/system URLs.
- Generate a support content inventory and sitemap coverage report.

### What it does not do

- No private support portals, authenticated crawling, full content scraping by default, SEO ranking promises, sitemap submission, or heavy site mirroring.

### Input fields

Primary inputs from the locked actor spec: `startUrls`, `sitemapUrls`, `robotsUrls`, `rawSitemapXml`, `allowedDomains`, `includeRobotsDiscovery`, `classifyUrlTypes`, `fetchPageTitles`, `includeStatusChecks`, `localeHints`, `maxUrls`, `timeoutMs`. `maxChargeUsd` keeps runs capped during production use.

### Output fields

Dataset rows include: `sourceSitemapUrl`, `url`, `urlType`, `title`, `locale`, `categoryHint`, `sectionHint`, `lastModified`, `changeFrequency`, `priority`, `statusCode`, `canonicalUrl`, `depth`, `isDuplicate`, `warning`.

### Starter example

Use `examples/input.tiny.json` as a small starter input. Keep the first run capped and review the dataset before increasing limits.

### Public task examples

- Run Help Center Sitemap Extractor on supplied sample data: Run Help Center Sitemap Extractor on supplied sample data using a small bounded input.
- Generate a Help Center Sitemap Extractor QA report: Generate a Help Center Sitemap Extractor QA report using a small bounded input.
- Find invalid rows with Help Center Sitemap Extractor: Find invalid rows with Help Center Sitemap Extractor using a small bounded input.
- Create a capped local endpoint readiness check for Help Center Sitemap Extractor: Create a capped local endpoint readiness check for Help Center Sitemap Extractor using a small bounded input.
- Prepare Help Center Sitemap Extractor output for downstream automation: Prepare Help Center Sitemap Extractor output for downstream automation using a small bounded input.

### Public source provenance

The starter input extracts a five-URL sample from Stripe's public Support sitemap at `https://support.stripe.com/sitemap.xml`. The source is fetched during each run, capped before enrichment, and restricted to the official `support.stripe.com` host.

### Reports

- `help-center-sitemap-inventory.md`
- `help-center-urls.json`
- `help-center-url-type-summary.csv`
- `locale-coverage-report.json`
- `sitemap-discovery-log.json`

### Limitations and safe use

Start with supplied-input runs, then enable live endpoints only with tight caps, domain allowlists, and no secrets in public examples.

# Actor input Schema

## `startUrls` (type: `array`):

Public pages to fetch and analyze. Keep first runs small and use allowed domains to constrain crawling.

## `sitemapUrls` (type: `array`):

Sitemap URLs to inspect for target pages and crawl candidates.

## `robotsUrls` (type: `array`):

Public robots URLs to fetch or inspect for Help Center Sitemap Extractor.

## `rawSitemapXml` (type: `array`):

Optional supplied sitemap XML documents to parse without fetching sitemap URLs.

## `allowedDomains` (type: `array`):

Optional domain allowlist that keeps fetched URLs constrained to approved hosts.

## `includeRobotsDiscovery` (type: `boolean`):

Include robots discovery in output rows or reports when available.

## `classifyUrlTypes` (type: `boolean`):

Classify URL Types controls Help Center Sitemap Extractor processing for the supplied inputs; keep values conservative for first runs.

## `fetchPageTitles` (type: `boolean`):

Fetch Page Titles controls Help Center Sitemap Extractor processing for the supplied inputs; keep values conservative for first runs.

## `includeStatusChecks` (type: `boolean`):

Include status checks in output rows or reports when available.

## `localeHints` (type: `array`):

Optional hints that improve Help Center Sitemap Extractor classification when source content is ambiguous.

## `maxUrls` (type: `number`):

Maximum URLs to process in one run; keep defaults low for safe first runs.

## `timeoutMs` (type: `number`):

Maximum time in milliseconds allowed for the Help Center Sitemap Extractor operation before it is treated as timed out.

## `maxChargeUsd` (type: `number`):

Maximum estimated PPE charge allowed for the run before the actor stops gracefully.

## Actor input object example

```json
{
  "startUrls": [],
  "sitemapUrls": [
    "https://support.stripe.com/sitemap.xml"
  ],
  "robotsUrls": [],
  "rawSitemapXml": [],
  "allowedDomains": [
    "support.stripe.com"
  ],
  "includeRobotsDiscovery": false,
  "classifyUrlTypes": true,
  "fetchPageTitles": false,
  "includeStatusChecks": false,
  "localeHints": [
    "en"
  ],
  "maxUrls": 5,
  "timeoutMs": 15000,
  "maxChargeUsd": 1
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("junipr/help-center-sitemap-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("junipr/help-center-sitemap-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call junipr/help-center-sitemap-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=junipr/help-center-sitemap-extractor",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/hhehmWhD8nLkg2n4k/builds/0qRFQx3dCZt9GklJA/openapi.json
