# Company Leadership & Team Page Scraper: Names, Roles, LinkedIn (`datagleaner/team-page-contacts`) Actor

Company leadership team and team-page scraper. Give it domains; it reads each company's own team, about, leadership and Impressum pages and returns every named person with job title, seniority, LinkedIn link and published email. Never guesses emails. $0.002 per person.

- **URL**: https://apify.com/datagleaner/team-page-contacts.md
- **Developed by:** [Data Gleaner](https://apify.com/datagleaner) (community)
- **Categories:** Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 person founds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Company Leadership & Team Page Scraper: names, roles and LinkedIn from company websites

Get a company's leadership team, team members and team page contacts from the company website. Give it company domains and it returns every named person the company publishes on its team, about, leadership, management or Impressum pages, with **job title, seniority, the LinkedIn profile the site links and the email the site publishes**. It answers "who works at this company", "who runs this company" and "list the team members on this company's website" for a whole list of domains at once. $0.002 per person; domains with no team page are free.

**It never guesses or verifies emails.** An email is returned only when the company's own page shows it for that person (a `mailto:` link or the address written in the person's card). Shared inboxes such as `info@`, `sales@` or `hello@` are not attributed to anyone, and no `firstname@domain` pattern is ever invented.

**To try it:** paste one or more company domains into `domains` and click Start. With an empty input it runs a small built-in example (two domains, at most 5 people).

### What it does

For each domain it reads the home page, finds the links in the navigation and footer that lead to the team, about, leadership, management, founders or Impressum / legal-notice page (in English, German, French, Spanish, Dutch and Italian), and reads at most `maxPagesPerDomain` of them (default 4). An about page that links on to a separate team page is followed. If the home page links to none of them, it tries the usual paths (`/team`, `/about`, `/leadership`, `/impressum` and similar).

From each page it extracts the people three ways:

- **Structured data:** schema.org `Person` in JSON-LD or microdata, and an `Organization`'s `founder`, `member` or `employee`.
- **Team cards:** the repeated name-and-title blocks of a team grid, with the LinkedIn `/in/` link and personal email inside the same card. When a grid gives some people a department instead of a title ("Design", "Customer Success"), it is read from the same slot the other cards use.
- **Impressum / legal notice:** the people named after `Geschäftsführer`, `Vertreten durch`, `Vorstand`, `Aufsichtsrat`, `Inhaber`, `Managing Director`, `Represented by` and similar labels, with that label as the title.

Then it keeps out what is not the company's team: customer testimonials and quotes, blog and article bylines, investor and partner logos, menu items, people described as working elsewhere ("CTO at Initech", "former VP"), and navigation text that looks like a name.

Each person gets a **seniority** from the title: `founder`, `c-level` (CEO, CFO, President, Managing Director, Geschäftsführer), `vp`, `director` (including Head of), `board` (board members, Aufsichtsrat, advisors listed under a board heading), `manager` or `other`. Rows come out most senior first.

### Use cases

- Sales prospecting: the founders, executives and department heads of every company on a target account list, with their LinkedIn profiles.
- CRM enrichment: add the leadership team to company records you already have.
- Recruiting and market mapping: who runs engineering, product or marketing at a set of competitors.
- Due diligence and investor research: the management and supervisory board named in a German, Austrian or Swiss company's Impressum.
- AI agents that need "the leadership team of company X" or "who works at company X" as a tool call, paying only when people are found.

### Input

```json
{
  "domains": ["baremetrics.com", "mongodb.com", "sipgate.de", "https://thoughtbot.com/team"],
  "maxPeoplePerDomain": 25,
  "seniority": ["founder", "c-level", "vp"],
  "maxPagesPerDomain": 4,
  "maxConcurrency": 10,
  "delaySecs": 0.5
}
```

- `domains`: bare domains, full URLs, or a direct team page URL.
- `maxPeoplePerDomain`: keep at most this many people per company, most senior first (default 25; 0 means no limit).
- `seniority`: return only these levels. Empty returns everyone. People filtered out are not charged.
- `maxPagesPerDomain`: pages read per domain after the home page (default 4, at most 10).
- `maxConcurrency`, `delaySecs`, `requestTimeoutSecs`, `respectRobotsTxt` (default on) and `proxyConfiguration` (default none).

### Output

One row per person.

```json
{
  "domain": "baremetrics.com",
  "companyName": "Baremetrics",
  "personName": "Luke Marshall",
  "jobTitle": "CEO",
  "seniority": "c-level",
  "linkedinUrl": null,
  "email": "luke@baremetrics.com",
  "emailSource": "mailto-link",
  "sourceUrl": "https://baremetrics.com/about",
  "input": "baremetrics.com",
  "scrapedAt": "2026-10-09T14:56:34+00:00"
}
```

```json
{
  "domain": "mongodb.com",
  "companyName": "MongoDB",
  "personName": "Mike Berry",
  "jobTitle": "Chief Financial Officer",
  "seniority": "c-level",
  "linkedinUrl": "https://www.linkedin.com/in/mike-berry-59b7961",
  "email": null,
  "emailSource": null,
  "sourceUrl": "https://www.mongodb.com/company/leadership",
  "input": "mongodb.com",
  "scrapedAt": "2026-10-09T14:55:02+00:00"
}
```

- `linkedinUrl` is filled only when the company's page links that person's LinkedIn `/in/` profile. It is never searched for elsewhere.
- `email` and `emailSource` (`mailto-link` or `page-text`) are filled only when the page publishes a personal address for that person.
- `sourceUrl` is the page the person was found on, so every row can be checked.

A per-domain summary (status, pages read, people found and pushed) is saved in the key-value store as `DOMAIN_SUMMARY`. Its `status` is `ok`, `noPeople`, `noTeamPage`, `unreachable`, `blockedByRobots`, `notHtml`, `invalidDomain` or `error`.

### How accurate is it?

Measured on 56 real team, about, leadership and Impressum pages from SaaS companies, agencies, enterprises and German publishers, hand-labelled person by person (755 people): **precision 100%, recall 98%**. On the last 20 of those pages, labelled before the extractor ever saw them, it scored **precision 100% and recall 82%** before any tuning. Precision is held highest on purpose: a name in the output is a real person the company lists, not a customer quoting it or a menu item.

### How much does it cost?

Pay per event: **US$2.00 per 1,000 people** (event `person-found`, US$0.002 each). You pay once per named person with a role, published on the company's own site. Domains with no team page, no named people, or that are unreachable are free, and people removed by your `seniority` filter or `maxPeoplePerDomain` are not charged.

Worked example: 500 company domains averaging 6 people each is 3,000 x $0.002 = **$6**. To pay only for leaders, set `seniority` to `["founder", "c-level", "vp"]`: a typical company then returns 2 to 5 people. Platform usage is small because the Actor makes plain HTTP requests only.

### How to get the leadership team of a list of companies

1. Paste the company domains into `domains`, one per line.
2. Set `seniority` to the levels you want, or leave it empty for the whole team.
3. Click Start. Rows appear as each domain finishes, most senior person first.
4. Export the dataset as CSV, Excel or JSON, or read it through the API.

### Limits

- **Only what the company publishes.** A company with no team or about page that names people returns nothing (and costs nothing). It does not search LinkedIn, Crunchbase or anywhere else.
- **No JavaScript rendering.** Team pages that load their people with scripts after the page opens, and sites that block non-browser clients (HTTP 403 or 429), return nothing. Setting `proxyConfiguration` helps with some blocks.
- **Names need a first and last name.** Cards that show a first name only ("Anna, Support") are skipped.
- **People named only in running prose** ("our founder Jane started the company in 2015") are not extracted; titled cards, structured data and Impressum labels are.
- Up to 10 pages per domain; default 4.

### Use with Python

```python
## pip install apify-client
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("datagleaner/team-page-contacts").call(
    run_input={"domains": ["baremetrics.com", "mongodb.com"], "seniority": ["founder", "c-level"]}
)
for row in client.dataset(run.default_dataset_id).iterate_items():
    print(row["companyName"], row["personName"], row["jobTitle"], row["linkedinUrl"])
```

### Use with JavaScript / Node.js

```js
// npm install apify-client
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagleaner/team-page-contacts').call({
    domains: ['baremetrics.com', 'mongodb.com'],
    seniority: ['founder', 'c-level'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const row of items) console.log(row.companyName, row.personName, row.jobTitle);
```

### Use it from n8n, Make, Zapier or an AI agent

Actor ID: `datagleaner/team-page-contacts`

Minimal input:

```
{"domains": ["baremetrics.com", "mongodb.com"]}
```

Each tool below runs this Actor with your own Apify API token.

- **n8n:** add the **Apify** node (`@apify/n8n-nodes-apify`). On n8n Cloud you install it from the community node registry. Choose **Run an Actor and get dataset**, set Actor to `datagleaner/team-page-contacts` and paste the input above.

- **Make:** use the Apify app's **Run an Actor** module, then **Get Dataset Items** to read the results. **Watch Actor Runs** can trigger a scenario when a run finishes.

- **Zapier:** use the Apify action **Run Actor**, then the search **Fetch dataset items**. The trigger **Finished Actor run** starts a Zap when a run ends.

- **AI agents (MCP):** connect to `https://mcp.apify.com/?tools=datagleaner/team-page-contacts`. In Claude Code:

  ```
  claude mcp add --transport http apify "https://mcp.apify.com/?tools=datagleaner/team-page-contacts"
  ```

  Then run `/mcp` to sign in to Apify in your browser. Other clients can sign in with OAuth or send the header `Authorization: Bearer YOUR_APIFY_TOKEN`. Clients that run local MCP servers can use Apify's package (`@apify/actors-mcp-server`, run with `npx -y` and `APIFY_TOKEN` set) instead. Then ask the agent in plain words, for example:

  > Who is on the leadership team of mongodb.com, baremetrics.com and sipgate.de? Give me a table of name, title and LinkedIn, with the page each one came from.

- **LangChain (Python):**

```python
## pip install langchain-apify, then set APIFY_TOKEN in your environment
import json
from langchain_apify import ApifyActorsTool
tool = ApifyActorsTool("datagleaner/team-page-contacts")
result = tool.invoke({"run_input": json.loads('{"domains": ["baremetrics.com", "mongodb.com"]}')})
```

### Use over plain HTTP

Agents and scripts without MCP can call it in one HTTP request and get the rows back:

```bash
curl -X POST "https://api.apify.com/v2/acts/datagleaner~team-page-contacts/run-sync-get-dataset-items?token=<APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"domains": ["mongodb.com"], "seniority": ["founder", "c-level"]}'
```

### FAQ

**How do I find who works at a company from its website?** Put the company's domain in `domains`. The Actor finds the team, about and leadership pages from the site's own navigation and returns each named person with their title, seniority and the page they appear on.

**Does it find email addresses?** Only the ones the company itself publishes for that person, such as a `mailto:` link on the team card. It never guesses an address from a name pattern, never verifies addresses with mail servers, and never gives a person a shared `info@` or `sales@` inbox.

**Does it return LinkedIn profiles?** Yes, when the company's page links the person's LinkedIn profile. The link is normalised to `https://www.linkedin.com/in/<slug>`. It does not search LinkedIn.

**Will it include customers from testimonials?** No. Quotes, testimonials and case-study people are removed, as are people described as working at another company. On the labelled test pages, no customer or reviewer was returned as a team member.

**Does it work on German Impressum pages?** Yes. Managing directors (Geschäftsführer), the executive board (Vorstand), the supervisory board (Aufsichtsrat) and owners (Inhaber) are read from the legal notice with that role as the title.

**Can I get only the founders and executives?** Set `seniority` to `["founder", "c-level"]`. Everyone else is left out and not charged.

**I need the company's general contact details, not people.** Use [Website Contact Details Scraper](https://apify.com/datagleaner/website-contact-details-scraper) for emails, phones and social profiles, or [Company Enrichment by Domain](https://apify.com/datagleaner/company-enrichment-by-domain) for a one-row company profile.

### Related Actors

- [Website Contact Details Scraper](https://apify.com/datagleaner/website-contact-details-scraper): emails, phone numbers, social profiles and contact forms from a list of websites.
- [Company Enrichment by Domain](https://apify.com/datagleaner/company-enrichment-by-domain): domain to company profile (name, logo, official social profiles, contacts, founding date).

### Responsible use

This Actor reads publicly available pages that companies publish about themselves, and returns only what those pages show. Names, titles and work emails are still personal data: you are responsible for complying with the law that applies to you, including GDPR, CCPA and anti-spam rules, when you store the results or contact the people in them.

# Actor input Schema

## `domains` (type: `array`):

Company websites, one per line. `acme.com`, `https://www.acme.com/` and a direct team page such as `acme.com/team` all work. Only the same domain is crawled.

## `maxPeoplePerDomain` (type: `integer`):

Most senior first (founders, then C-level, VPs, directors, board, managers, others). 0 means no limit. Each person returned is one billed event.

## `seniority` (type: `array`):

Return only these levels. Leave empty for everyone. People filtered out are not billed.

## `maxPagesPerDomain` (type: `integer`):

Team, about, leadership and Impressum pages read per domain, after the home page. A 404 guess does not count.

## `maxConcurrency` (type: `integer`):

How many domains are crawled at the same time. Pages of one domain are always fetched one at a time.

## `delaySecs` (type: `number`):

Polite pause between two requests to the same website.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for each page before retrying.

## `respectRobotsTxt` (type: `boolean`):

Skip pages that the site's robots.txt disallows.

## `proxyConfiguration` (type: `object`):

Optional. Defaults to no proxy.

## Actor input object example

```json
{
  "domains": [
    "baremetrics.com",
    "mongodb.com"
  ],
  "maxPeoplePerDomain": 25,
  "maxPagesPerDomain": 4,
  "maxConcurrency": 10,
  "delaySecs": 0.5,
  "requestTimeoutSecs": 20,
  "respectRobotsTxt": true
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "baremetrics.com",
        "mongodb.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagleaner/team-page-contacts").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "baremetrics.com",
        "mongodb.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("datagleaner/team-page-contacts").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "baremetrics.com",
    "mongodb.com"
  ]
}' |
apify call datagleaner/team-page-contacts --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagleaner/team-page-contacts"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZO60PO7vSBsZCUXzZ/builds/bwoRX0Jz46aXrIuE0/openapi.json
