# Workplace Program Detector ERG DEI Benefits Parental Leave (`mambalabs/workplace-program-detector`) Actor

Detects which people programs a company publishes: employee resource groups, DEI, wellbeing and mental health, learning and tuition support, parental and caregiver leave, volunteering. Reads careers, culture, benefits and ESG pages plus live job postings. One flat row per domain for Clay.

- **URL**: https://apify.com/mambalabs/workplace-program-detector.md
- **Developed by:** [Mamba Labs](https://apify.com/mambalabs) (community)
- **Categories:** Lead generation, Automation, Jobs
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $6.80 / 1,000 domain analyzeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

### 🔎 What can Workplace Program Detector do?

Give it a **company domain** and it returns **one flat row** naming the people programs
that company publishes: employee resource groups, DEI, wellbeing and mental health,
learning and tuition support, parental and caregiver policy, and volunteering and giving.

Every row also reports how much of the company the actor actually read, so you can tell a
real "no program" apart from a page it could not open.

| 📦 What you get | ⚙️ Features and integrations |
|---|---|
| 🏅 **Six program families**, one boolean each<br>🔗 **Evidence URL and phrase** per family, so you can check the finding<br>📊 **Coverage block**, what was reached, blocked or thin<br>🧾 **34 flat fields**, `snake_case`, one row per domain | 🌐 **Two retrieval paths**, website pages and live job postings<br>🤖 **Auto board discovery** for Greenhouse, Lever and Ashby<br>🧊 **14 day cache**, with a `skipCache` override<br>⬇️ **Export** to JSON, CSV, Excel, HTML or XML |

Bought by teams selling into HR, People Ops, benefits and learning and development, and by
employer brand and DEI consultancies building prospect lists from what companies publish.

> 🚫 **This is not an employer review or ratings source.** It does not score a company, rate
> its culture, or tell you whether a program works. It records what the company published on
> its own website and job board, and nothing else.

### 💡 Why use Workplace Program Detector?

| If you sell | Read these fields |
|---|---|
| ERG and community platforms | `has_erg`, `erg_evidence_url`, `erg_evidence_phrase` |
| DEI programs and consulting | `has_dei_program`, `dei_evidence_url` |
| Mental health and wellbeing benefits | `has_wellbeing_program`, `wellbeing_evidence_url` |
| Learning platforms and tuition benefits | `has_learning_program`, `learning_evidence_url` |
| Family and caregiver benefits | `has_parental_policy`, `parental_evidence_url` |
| Volunteering and giving platforms | `has_volunteering_program`, `volunteering_evidence_url` |
| Anything, as a disqualifier | `program_count`, `coverage`, `pages_reached` |

#### 🧭 Two retrieval paths, because one path misses half the answer

The actor reads company web pages and live job postings, and it ships both because they
find different things.

Measured during the build: across five real job boards covering **1,764 open roles**,
employee resource groups appeared in **zero** job bodies. Across **seven** company websites,
employee resource groups appeared on **two**. Benefits language runs the other way, and
shows up in job bodies that marketing pages leave out.

`signal_source` on every row tells you which path produced the finding.

### 📋 What data can Workplace Program Detector extract?

**34 fields** per domain. The ones buyers use:

| Field | What it holds |
|---|---|
| `has_erg` | Company publishes employee resource groups |
| `has_dei_program` | Published DEI program, excluding legal boilerplate |
| `has_wellbeing_program` | Wellbeing or mental health program |
| `has_learning_program` | Learning, development or tuition support |
| `has_parental_policy` | Parental or caregiver policy |
| `has_volunteering_program` | Volunteering or giving program |
| `program_count` | How many of the six were found |
| `*_evidence_url` | The page the finding came from |
| `*_evidence_phrase` | The wording that matched |
| `signal_source` | Which path fired, website or job postings |
| `coverage` | How much of the company was read |
| `pages_attempted`, `pages_reached`, `pages_reached_list` | What was tried and what opened |
| `pages_blocked`, `pages_thin` | Refused pages, and pages with almost no readable text |
| `ats_provider`, `ats_slug_used`, `jobs_scanned` | Which job board was read, and how many jobs |
| `render_mode` | `client_rendered` when the page needs a browser to show text |
| `fetch_status`, `fetch_error` | `ok`, `blocked`, or the error that stopped it |

> ⚠️ **`false` and `null` are not the same thing.** `false` means the pages were read and the
> language was not there. `null` means not enough was read to say. They are never collapsed
> into each other. A company that returns 403 on every path comes back with all six fields
> `null` and `fetch_status: "blocked"`, not six confident falses. Read `coverage` and
> `pages_reached` before you trust any `false`.

### 🛠️ How to find a company's employee programs

1. Open the **Input** tab and put a bare domain in `domain`, for example `hubspot.com`.
2. Leave `scan_web_pages` and `scan_job_postings` on `true` so both paths run.
3. Click **Start**.
4. Read `program_count` and the six `has_*` booleans for the answer.
5. Read `coverage` and `pages_reached` before you act on any `false`.
6. Export from the **Output** tab, or pull the row through the API.

#### 🧪 Using it in Clay

Add it as an Apify enrichment, map your domain column to `domain`, and the 34 fields land as
one flat row with no reshaping. Every input is accepted as a string, which is what Clay sends.

Gate the run on your ICP column so you only spend the event on accounts you would actually
work.

#### ⚡ Skipping board discovery

If you already know the company's job board slug, put it in `ats_slug` and the actor reads
that board directly instead of discovering it. That is the fastest path when you are running
a list you have already resolved.

### 💵 How much does it cost to detect workplace programs?

One `domain-analyzed` event per domain.

| Plan | Price per domain |
|---|---|
| Free | $0.008 |
| Bronze | $0.0076 |
| Silver | $0.0072 |
| Gold | $0.0068 |

There is also an Actor start event at $0.00005, charged once per run per GB of memory.

> 💳 **A domain that turns out to be unreachable is still billed.** The actor did the work of
> attempting it, and a `null` row that tells you the site blocked us is a real answer. What is
> never billed is a run that does not start.

### ⌨️ Input

Everything is on the **Input** tab. The options worth explaining:

| Field | Type | Default | What it does |
|---|---|---|---|
| `domain` | string | required | Bare domain, for example `hubspot.com`. Protocol and path are stripped. |
| `ats_slug` | string | empty | Skip board discovery and read this Greenhouse, Lever or Ashby slug directly. |
| `scan_job_postings` | boolean | `true` | Read the company's live job bodies. |
| `scan_web_pages` | boolean | `true` | Probe the careers, culture, benefits, DEI and ESG paths. |
| `max_pages` | string | `14` | Clamped to 1 to 25. Sent as a string for Clay. |
| `skipCache` | string | `false` | `true` forces a fresh crawl past the 14 day cache. |

### 📤 Output

One row per domain, exportable as **JSON, CSV, Excel, HTML or XML**. 34 flat `snake_case`
fields, with `null` rather than a missing key.

```json
{
  "domain": "hubspot.com",
  "has_erg": true,
  "erg_evidence_url": "https://www.hubspot.com/careers/culture",
  "erg_evidence_phrase": "employee resource groups",
  "has_dei_program": true,
  "has_wellbeing_program": true,
  "has_learning_program": true,
  "has_parental_policy": true,
  "has_volunteering_program": false,
  "program_count": 5,
  "signal_source": "website",
  "coverage": "high",
  "pages_attempted": 14,
  "pages_reached": 11,
  "pages_blocked": 0,
  "pages_thin": 1,
  "ats_provider": "greenhouse",
  "jobs_scanned": 132,
  "render_mode": "server_rendered",
  "fetch_status": "ok",
  "fetch_error": null
}
```

### 💡 Tips

- Run it monthly on the same list and diff two rows. That turns a snapshot into a trend.
- A documented absence is a pitch. `has_wellbeing_program: false` with `coverage: "high"` is a
  company whose competitors publish a program and it does not.
- Use `program_count` as a cheap sort. Companies publishing five or six families are already
  investing in this area and are usually the warmer conversation.
- Set `ats_slug` when you have it. It removes the discovery step.

### ⚠️ Known limits

**A `false` is only as good as the pages that opened.** On a company where one page was
reached, a `false` means very little. Check `pages_reached` and `coverage` on every row.

**JavaScript-only careers sites cannot be read.** Some large companies serve a careers page
that is an empty shell until a browser runs it. Measured during the build: one site's careers
page held 20 characters of readable text, and another held 1,141 characters inside 84,600
bytes of markup. Those rows say `render_mode: client_rendered` with `pages_thin` above zero.
Treat every `false` on them as unknown.

**Some sites refuse outright.** One of the seven domains in the build sample returned HTTP 403
on nine of fourteen paths. Those rows come back `fetch_status: "blocked"` with all six program
fields `null`.

**Some sites ask not to be read.** `robots.txt` is read and honored on every domain before
probing. A site that disallows crawling gets `pages_attempted: 0` and null program fields.
That is the site's decision recorded correctly, not a company without programs.

**Employee resource groups are almost never named in job postings.** If a company blocks its
website, `has_erg` stays null even when the other families resolve from job postings.

**Internal branding defeats keyword matching.** Companies name their groups whatever they
like. One large company in the build sample calls them Equality Groups and was invisible until
that phrase was added. Expect misses on unusual internal naming.

**Small companies often have no reachable website.** One test target has no working HTTPS
certificate and returns a network error on every path. That is a `none` coverage row, not a
company without programs.

**This is a snapshot, not a trend.** The actor does not say whether a program is growing or
being wound down. The archival sources that would answer that take between 2 and 18 seconds
per query, far more than this actor's whole runtime. Run it monthly and diff instead.

**Legal boilerplate is excluded on purpose.** "Equal opportunity employer" appears in almost
every US job posting and is not counted as a DEI program.

**The actor records that a company publishes employee resource groups, and deliberately does
not record which groups.** The ERG evidence phrase is capped tight and dropped entirely if
what survives names a protected characteristic. When that happens `has_erg` stays true and the
evidence URL stays populated, so you keep the finding and its source and lose only the
sentence. The same scrub runs on the DEI phrase.

### ❓ FAQ

##### Why is every program field null?

The site blocked the actor or could not be reached. Check `fetch_status` and `pages_blocked`.
That row is a coverage problem, not a company without programs.

##### Does it use a proxy?

No. Every path runs direct with full browser headers. The design answer to a block is an
honest `blocked` status rather than a workaround.

##### How fresh is the data?

Results are cached for 14 days per domain and input combination, because careers and benefits
pages change quarterly at best. Pass `skipCache: "true"` to force a fresh crawl.

##### Can I run a list instead of one domain?

Yes. Run it per row from Clay, or drive it through the Apify API and collect the dataset.

##### Does it tell me the names of a company's ERGs?

No, by design. It records that a company publishes employee resource groups and drops any
evidence phrase that names a protected characteristic.

### 🧩 Want other GTM data?

Mamba Labs builds custom actors for B2B go-to-market teams. The public versions
of that work live here on the Store, so our users get the same tooling we build
under contract.

| | |
|---|---|
| 🧑‍💼 [GTM Hiring Signal Scraper](https://apify.com/mambalabs/gtm-hiring-signal-scraper) | 🧱 [Tech Stack Detector](https://apify.com/mambalabs/gtm-tech-stack-signal-scraper) |
| 📡 [B2B Buying Signals Aggregator](https://apify.com/mambalabs/b2b-buying-signals-hiring-tech-stack-intent-for-clay) | 🔑 [Job Board Keyword Scanner](https://apify.com/mambalabs/job-board-keyword-signal-scanner) |
| 🔗 [Domain to LinkedIn URL Resolver](https://apify.com/mambalabs/domain-to-linkedin-url-resolver) | 🎯 [ICP Fit Scorer](https://apify.com/mambalabs/icp-account-lead-scoring-fit-scorer-0-100-for-clay) |
| 📋 [Job Posting Monitor](https://apify.com/mambalabs/gtm-job-discovery) | 📬 [Domain Deliverability Checker](https://apify.com/mambalabs/domain-deliverability-checker) |
| 🏢 [Company Firmographic Enricher](https://apify.com/mambalabs/company-firmographic-enricher) | 🌐 [Company Social Presence Mapper](https://apify.com/mambalabs/company-social-presence-mapper) |
| 🪪 [Company Identity Resolver](https://apify.com/mambalabs/company-identity-resolver) | 💰 [Funding and Press Signal Scanner](https://apify.com/mambalabs/funding-press-signal-scanner) |
| 🔄 [Company Change-Event Feed](https://apify.com/mambalabs/company-change-event-feed) | 👤 [People Finder and Email Verifier](https://apify.com/mambalabs/people-finder) |
| 🚀 [Prospect Engine](https://apify.com/mambalabs/b2b-prospect-engine) | 🤖 [AI Tooling Detector](https://apify.com/mambalabs/ai-tooling-detector) |
| 📮 [Outbound Stack Detector](https://apify.com/mambalabs/outbound-infrastructure-fingerprint) | 📝 [Publishing Frequency Tracker](https://apify.com/mambalabs/blog-publishing-frequency) |
| ✉️ [Work Email Waterfall Finder](https://apify.com/mambalabs/email-waterfall-orchestrator) | ⏩ [Sequencer Lead Push](https://apify.com/mambalabs/clay-to-instantly-smartlead-push) |
| 👥 [Team Page People Extractor](https://apify.com/mambalabs/team-page-people-extractor) | 🧭 [Company Discovery List Builder](https://apify.com/mambalabs/company-discovery-list-builder) |

> Every actor in the suite takes a domain or a company and returns one flat row,
> so they stack in the same Clay table without reshaping anything.

> 🛠️ **Need something custom built for you or your team?** Tell us what you are
> trying to find and we will build it. [Talk to Mamba Labs](https://mambabuilt.com/contact).

### 🆘 Support

Something wrong, or a company the actor reads incorrectly? Open an issue on the **Issues** tab
with the domain and the row, and we will look at it.

> ℹ️ **Sourcing and legal.** Every field comes from pages the company publishes itself, read
> directly, with `robots.txt` honored on every domain. The row describes what a company
> published on its own website and job board. It is not an assessment of the company, its
> workforce, or how well any program works. You are responsible for how you use the output,
> including under applicable employment and data protection law.

Built by [Mamba Labs](https://apify.com/mambalabs).

# Actor input Schema

## `domain` (type: `string`):

Bare domain, for example hubspot.com. Protocol and path are stripped.

## `ats_slug` (type: `string`):

Skip ATS discovery and read this Greenhouse, Lever or Ashby board slug directly.

## `scan_job_postings` (type: `boolean`):

Reads the company's live job bodies from its ATS. This path finds benefits language that marketing pages omit.

## `scan_web_pages` (type: `boolean`):

Probes the careers, culture, benefits, DEI and ESG paths. This path finds ERG and volunteering language that job postings omit.

## `max_pages` (type: `string`):

Sent as a string for Clay. Clamped to 1 to 25.

## `skipCache` (type: `string`):

false uses the 14 day cache, true forces a fresh crawl. Sent as a string for Clay, matching the fleet convention.

## Actor input object example

```json
{
  "domain": "hubspot.com",
  "scan_job_postings": true,
  "scan_web_pages": true,
  "max_pages": "14",
  "skipCache": "false"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domain": "hubspot.com",
    "scan_job_postings": true,
    "scan_web_pages": true,
    "max_pages": "14",
    "skipCache": "false"
};

// Run the Actor and wait for it to finish
const run = await client.actor("mambalabs/workplace-program-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domain": "hubspot.com",
    "scan_job_postings": True,
    "scan_web_pages": True,
    "max_pages": "14",
    "skipCache": "false",
}

# Run the Actor and wait for it to finish
run = client.actor("mambalabs/workplace-program-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domain": "hubspot.com",
  "scan_job_postings": true,
  "scan_web_pages": true,
  "max_pages": "14",
  "skipCache": "false"
}' |
apify call mambalabs/workplace-program-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mambalabs/workplace-program-detector"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Li6pPDhqbdzFz3h7G/builds/66ayzFv07k0vVwT0g/openapi.json
