# Intercom Help Center Public Crawler (`junipr/intercom-help-center-public-crawler`) Actor

Crawl public Intercom-style help centers and extract collections, sections, article URLs, titles, descriptions, breadcrumbs, locale, update signals, and article content into structured inventory rows.

- **URL**: https://apify.com/junipr/intercom-help-center-public-crawler.md
- **Developed by:** [junipr](https://apify.com/junipr) (community)
- **Categories:** SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.90 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Intercom Help Center Public Crawler

### Store Positioning

**Store title:** Intercom Help Center Public Crawler

**Short description:** Crawl public Intercom-style help centers and extract collections, sections, article URLs, titles, descriptions, breadcrumbs, locale, update signals, and article content into structured inventory rows.

**SEO title:** Intercom Help Center Public Crawler — technical SEO, web, and domain audit

**SEO description:** Crawl public Intercom-style help centers and extract collections, sections, article URLs, titles, descriptions, breadcrumbs, locale, update signals, and article content into structured inventory rows. Use it to find crawlability, indexability, security, metadata, and page-quality issues with evidence-backed rows and audit reports.

**Categories:** SEO\_TOOLS

**Keywords:** intercom, help, center, public, crawler, structured data, public data, local seo, web/domain audit

### Pay-Per-Event Pricing

This actor uses pay-per-event pricing. Event prices include Apify platform usage; users are not expected to pay a separate platform-usage pass-through charge for the configured pricing model.

- Tier: W1 — Web/domain audit
- Primary event: `page-audited` at $0.00490 base
- Default max charge: $10.00
- Store discounts: FREE/BRONZE base, SILVER discounted, GOLD deepest approved discount

Event set:

- `actor-start`: base $0.00500, GOLD $0.00400. Intercom Help Center Public Crawler: charged when actor start is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `page-audited`: base $0.00490, GOLD $0.00392. Intercom Help Center Public Crawler: charged when page audited is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `record-extracted`: base $0.00372, GOLD $0.00298. Intercom Help Center Public Crawler: charged when record extracted is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `finding-emitted`: base $0.00372, GOLD $0.00298. Intercom Help Center Public Crawler: charged when finding emitted is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `audit-report-generated`: base $0.05000, GOLD $0.04000. Intercom Help Center Public Crawler: charged when audit report generated is completed. The price includes Apify platform usage; no separate usage pass-through is intended.

The actor accepts `actor-start` before work, accepts the primary event before each dataset row, and accepts the configured report event before writing report files. If `maxChargeUsd` or the live PPE limit blocks a charge, the corresponding row or report is not written.

### Public Task Concepts

- Audit Intercom Help Center Public Crawler controls on a capped public sample
- Find high-priority Intercom Help Center Public Crawler issues before release
- Validate Intercom Help Center Public Crawler evidence from supplied pages
- Prioritize Intercom Help Center Public Crawler fixes with severity and proof
- Export Intercom Help Center Public Crawler QA rows for client review

Crawl public Intercom-style help centers and extract collections, sections, article URLs, titles, descriptions, breadcrumbs, locale, update signals, and article content into structured inventory rows.

### What it does

- Accept public Intercom help center base URLs, collection URLs, article URLs, or sitemap URLs.
- Crawl bounded public help center pages.
- Extract collection names, collection URLs, article URLs, titles, subtitles/descriptions, article body text, breadcrumbs, locale, author/update signals where visible, and related article links.
- Preserve hierarchy and source evidence.
- Generate collection map and article inventory for migration, SEO, and support QA.

### What it does not do

- No Intercom admin API access, Messenger conversations, user/customer data, private help center crawling, access-control bypass, official certification, or heavy mirroring.

### Input fields

Primary inputs from the locked actor spec: `startUrls`, `helpCenterBaseUrl`, `sitemapUrls`, `allowedDomains`, `localeHints`, `maxPages`, `maxDepth`, `includeArticleBody`, `includeCollections`, `includeRelatedArticles`, `includeAuthorSignals`, `requestDelayMs`, `timeoutMs`. `maxChargeUsd` keeps runs capped during production use.

### Output fields

Dataset rows include: `sourceUrl`, `pageType`, `collectionName`, `collectionUrl`, `sectionName`, `articleTitle`, `articleUrl`, `articleDescription`, `breadcrumbs`, `locale`, `authorName`, `updatedAt`, `articleBodyText`, `relatedArticleUrls`, `statusCode`, `warning`.

### Starter example

Use `examples/input.tiny.json` as a small starter input. Keep the first run capped and review the dataset before increasing limits.

### Public task examples

- Run Intercom Help Center Public Crawler on supplied sample data: Run Intercom Help Center Public Crawler on supplied sample data using a small bounded input.
- Generate a Intercom Help Center Public Crawler QA report: Generate a Intercom Help Center Public Crawler QA report using a small bounded input.
- Find invalid rows with Intercom Help Center Public Crawler: Find invalid rows with Intercom Help Center Public Crawler using a small bounded input.
- Create a capped local endpoint readiness check for Intercom Help Center Public Crawler: Create a capped local endpoint readiness check for Intercom Help Center Public Crawler using a small bounded input.
- Prepare Intercom Help Center Public Crawler output for downstream automation: Prepare Intercom Help Center Public Crawler output for downstream automation using a small bounded input.

### Public source provenance

The starter input fetches one public Intercom-hosted Fieldwork article directly from `https://intercom.help/fieldwork/en/articles/927524-getting-started-welcome-to-fieldwork`, constrained to `intercom.help` and one page.

### Reports

- `intercom-help-center-public-crawl-report.md`
- `intercom-article-inventory.csv`
- `intercom-collection-map.json`
- `intercom-locale-summary.json`
- `intercom-crawl-errors.json`

### Limitations and safe use

Start with supplied-input runs, then enable live endpoints only with tight caps, domain allowlists, and no secrets in public examples.

# Actor input Schema

## `startUrls` (type: `array`):

Public pages to fetch and analyze. Keep first runs small and use allowed domains to constrain crawling.

## `helpCenterBaseUrl` (type: `string`):

Public help center base URL to fetch or inspect for Intercom Help Center Public Crawler.

## `sitemapUrls` (type: `array`):

Sitemap URLs to inspect for target pages and crawl candidates.

## `allowedDomains` (type: `array`):

Optional domain allowlist that keeps fetched URLs constrained to approved hosts.

## `localeHints` (type: `array`):

Optional hints that improve Intercom Help Center Public Crawler classification when source content is ambiguous.

## `maxPages` (type: `number`):

Maximum pages to process in one run; keep defaults low for safe first runs.

## `maxDepth` (type: `number`):

Maximum depth to process in one run; keep defaults low for safe first runs.

## `includeArticleBody` (type: `boolean`):

Include article body in output rows or reports when available.

## `includeCollections` (type: `boolean`):

Include collections in output rows or reports when available.

## `includeRelatedArticles` (type: `boolean`):

Include related articles in output rows or reports when available.

## `includeAuthorSignals` (type: `boolean`):

Include author signals in output rows or reports when available.

## `requestDelayMs` (type: `number`):

Request Delay ms controls Intercom Help Center Public Crawler processing for the supplied inputs; keep values conservative for first runs.

## `timeoutMs` (type: `number`):

Maximum time in milliseconds allowed for the Intercom Help Center Public Crawler operation before it is treated as timed out.

## `maxChargeUsd` (type: `number`):

Maximum estimated PPE charge allowed for the run before the actor stops gracefully.

## `pageHtmlByUrl` (type: `object`):

Optional URL-to-HTML map for processing supplied public page snapshots without fetching.

## Actor input object example

```json
{
  "startUrls": [
    "https://intercom.help/fieldwork/en/articles/927524-getting-started-welcome-to-fieldwork"
  ],
  "helpCenterBaseUrl": "https://intercom.help/fieldwork",
  "sitemapUrls": [],
  "allowedDomains": [
    "intercom.help"
  ],
  "localeHints": [
    "en"
  ],
  "maxPages": 1,
  "maxDepth": 1,
  "includeArticleBody": true,
  "includeCollections": true,
  "includeRelatedArticles": true,
  "includeAuthorSignals": true,
  "requestDelayMs": 0,
  "timeoutMs": 15000,
  "maxChargeUsd": 1,
  "pageHtmlByUrl": {}
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("junipr/intercom-help-center-public-crawler").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("junipr/intercom-help-center-public-crawler").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call junipr/intercom-help-center-public-crawler --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=junipr/intercom-help-center-public-crawler",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Wgx1lWkvdpj9obfwY/builds/phnbXQsKhBFa3sHEM/openapi.json
