# Team Page Extractor — Names, Titles & LinkedIn from Websites (`inovaflow/team-page-extractor`) Actor

Give it company domains and get every person on their team, about and leadership pages: name, title, department, seniority, LinkedIn, X, photo, bio, plus company phone and e-mails. Optional pattern-based work e-mail with MX check. No login, dataset-only, MCP-ready.

- **URL**: https://apify.com/inovaflow/team-page-extractor.md
- **Developed by:** [inovaflow](https://apify.com/inovaflow) (community)
- **Categories:** Lead generation, Business, Automation
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 people

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Team Page Extractor — Names, Titles & LinkedIn from Websites

Give it company domains and get **every person listed on their team, about and leadership pages** — name, job title,
department, seniority, LinkedIn / X profile, headshot, bio, the e-mail printed next to them, plus the company's phone
and e-mails. Optional pattern-based work e-mail with a mail-server check. One row per person, one price per row.
No login, no API key, no browser — the site's own HTML, read structurally.

### Who it is for

- **List builders and SDR teams** who have accounts (domains) and need the people at them without a data vendor.
- **Agencies and recruiters** mapping who does what at a target company — with the source page for every row.
- **Investors, journalists, analysts** pulling leadership and partner lists from hundreds of firms at once.
- **AI agents** (MCP): give it domains, get flat rows with stable ids — the no-search fast path next to
  [Decision Maker Finder](https://apify.com/inovaflow/decision-maker-finder).

### What it does

For every domain:

1. reads the **homepage** (company name, phone, e-mails, LinkedIn page) and finds the pages most likely to list
   people: links whose path or label says *team / people / leadership / about / management / staff / attorneys /
   doctors …*, the well-known paths (`/team`, `/about`, `/about-us`, `/people`, `/leadership`, `/our-team`,
   `/company`, `/management`, `/staff`, `/who-we-are`, localized variants), and `sitemap.xml` hints;
2. reads the best candidates within your page budget (pagination of a productive team page included);
3. extracts people **structurally** — schema.org `Person` / `ItemList` JSON-LD, repeated cards (name element +
   role line + LinkedIn / X link + headshot + bio), hydration blobs of client-rendered sites, and "Name, Title"
   prose — and maps profile links to people by proximity when a site keeps them in a data blob;
4. rejects what is not a colleague: testimonials, blog authors, job postings, menu items, portfolio founders and
   customer quotes (people whose title names another company), product tiles, locations;
5. classifies the title into **department** (sales, marketing, engineering, …) and **seniority** (founder,
   c-level, vp, director, head, manager, individual);
6. optionally finds a **work e-mail**: learns the domain's address pattern from e-mails on its own pages, checks the
   mail server (and the mailbox where SMTP is reachable), and delivers the address only with that evidence.

### Why this one

- **Accuracy over coverage.** Every row carries the page it came from; a bare name never ships — a card needs a
  role line, a profile link or a headshot in a repeated grid. Titles are never invented, e-mails never guessed.
- **Any stack.** Tested on WordPress, Webflow, Framer, Next.js, Gatsby, Squarespace-style static sites, law-firm
  and clinic CMSs, VC sites: 40+ domains hand-checked (see `docs/DECISIONS.md` §3).
- **Cheap and fast.** Pure HTTP; 6 pages per site by default; ~5–15 s per domain; the run's own connection first and
  your proxy only when a site blocks it.
- **Dataset-only, MCP-ready.** Flat rows, stable `id` (`domain:name-slug`), `sources[]`, `scrapedAt`, a run summary
  in `OUTPUT`, a CSV in the key-value store.

### Fields

| Field | Type | Meaning |
| --- | --- | --- |
| `id` | string | `domain:first-last` — stable across runs |
| `fullName`, `firstName`, `lastName` | string | as printed (honorifics and post-nominals stripped) |
| `title` | string | null | the role line next to the name |
| `department` | enum | `executive`, `sales`, `marketing`, `engineering`, `product`, `data`, `design`, `finance`, `operations`, `hr`, `customer-success`, `it`, `legal`, `other` |
| `seniority` | enum | `founder`, `c-level`, `vp`, `director`, `head`, `manager`, `individual` |
| `company` | string | null | site name (og:site\_name / schema.org Organization / title) |
| `domain`, `website` | string | the company site (after redirects) |
| `linkedin`, `twitter` | string | null | personal profile URLs found on the page |
| `email` | string | null | the person's e-mail: printed on the page (`emailSource: "page"`) or pattern-built and mail-server-checked (`"pattern"`) |
| `emailConfidence` | number | null | 100 for a page e-mail; 70–95 for a pattern e-mail |
| `emailVerification` | enum | null | `smtp-valid`, `smtp-catch-all`, `mx-only`, `no-mx`, `syntax-only` |
| `emailCandidates[]` | array | unconfirmed guesses (`address`, `pattern`, `confidence`, `verification`) — free |
| `photoUrl` | string | null | headshot |
| `bio` | string | null | first 300 chars of the bio |
| `sourceUrl` | string | the page the person was read from |
| `extractedBy` | enum | `json-ld`, `card`, `blob`, `text` |
| `companyPhone`, `companyEmails[]`, `companyLinkedin` | — | company-level contacts from the pages read |
| `teamPages[]`, `sources[]`, `scrapedAt` | — | provenance |

### How to use

1. Paste domains (or team-page URLs you already know) into **Company domains**.
2. Optionally filter by **roles** (`CEO`, `VP Sales`, `Head of Marketing`, `Partner`), **departments** or
   **seniorities**, and require a title or a LinkedIn profile.
3. Keep **Find work e-mails** on to get pattern-based addresses ($0.01 each, only when confirmed).
4. Run. Rows land in the dataset (views *People* and *Leads & contacts*), `PEOPLE.csv` and the `OUTPUT` summary in
   the key-value store.

### Cost

Pay per event — see `PRICING.md`:

- **Person** $0.01 — one delivered person (charged once even if found on several pages).
- **Work e-mail** $0.01 — only when a pattern-built, mail-server-checked address is delivered. Page e-mails are free.
- **Actor start** $0.005.

A 10-domain run with ~200 people costs about $2.10 plus a few cents of platform usage.

### Input

```json
{ "domains": ["hubspot.com", "thoughtbot.com", "kleinerperkins.com"], "maxPagesPerSite": 6, "maxPeoplePerSite": 100, "findEmails": true }
```

### Output sample

```json
{
  "id": "hubspot.com:yamini-rangan",
  "fullName": "Yamini Rangan",
  "firstName": "Yamini",
  "lastName": "Rangan",
  "title": "Chief Executive Officer",
  "department": "executive",
  "seniority": "c-level",
  "company": "HubSpot",
  "domain": "hubspot.com",
  "website": "https://www.hubspot.com/",
  "linkedin": "https://www.linkedin.com/in/yaminirangan",
  "twitter": null,
  "email": null,
  "emailSource": null,
  "emailConfidence": null,
  "emailVerification": null,
  "emailCandidates": [{ "address": "yamini.rangan@hubspot.com", "pattern": "first.last", "confidence": 45, "verification": "mx-only" }],
  "photoUrl": "https://www.hubspot.com/hs-fs/hubfs/assets/hubspot.com/about/management%202025/Yamini-Rangan-Headshot.webp",
  "bio": null,
  "sourceUrl": "https://www.hubspot.com/company/management",
  "extractedBy": "card",
  "companyPhone": "+18884827768",
  "companyEmails": [],
  "companyLinkedin": "https://www.linkedin.com/company/hubspot",
  "teamPages": ["https://www.hubspot.com/company/management", "https://www.hubspot.com/company/board-of-directors"],
  "sources": ["https://www.hubspot.com/company/management"],
  "scrapedAt": "2026-09-26T15:20:11.000Z"
}
```

### FAQ

**Why is `title` null for some people?** The page shows only a name and a photo (many VC and agency grids). The
person still ships when the card sits in a repeated headshot grid on a listing page; turn on *Only people with a job
title* to drop them.

**Why did a domain return no people?** The site renders its team list in the browser (search apps of big law firms,
some React sites) or has no team page. The run summary (`OUTPUT.perDomain`) says `no_people` with the pages read.
Pasting the exact team-page URL as the domain often helps.

**Are board members and advisors included?** Yes when the site lists them on a leadership / board page; their
title keeps the other company ("General Partner, ICONIQ Capital").

**Does it charge for e-mail guesses?** No. Only addresses backed by the domain's own pattern (or an SMTP-confirmed
mailbox) are delivered and charged; the rest stay in `emailCandidates` with their confidence.

**Blocked sites?** The direct connection is tried first, then your proxy (residential by default). Sites that block
both are reported as `blocked` in the summary and never charged.

# Actor input Schema

## `domains` (type: `array`):

One per line: `hubspot.com`, `https://www.thoughtbot.com` or a team page you already know (`https://www.hubspot.com/company/management`). Each domain is crawled from its homepage to the pages most likely to list people.

## `maxPagesPerSite` (type: `integer`):

Pages read per domain (homepage included): the best-ranked team / about / leadership candidates from the site's links, well-known paths and sitemap, plus pagination of a productive team page.

## `maxPeoplePerSite` (type: `integer`):

Stop after this many people per domain (best evidence first). Only delivered rows are charged.

## `roles` (type: `array`):

Keep only people whose title matches one of these — abbreviations understood (`CEO`, `VP Sales`, `Head of Marketing`, `Partner`, `Attorney`). Empty = everyone.

## `departments` (type: `array`):

Keep only people classified into these departments (from the title).

## `seniorities` (type: `array`):

Keep only these seniority levels (from the title).

## `onlyWithTitle` (type: `boolean`):

Drop people whose card shows a name and photo but no role line.

## `onlyWithLinkedin` (type: `boolean`):

Drop people without a LinkedIn URL on the page.

## `findEmails` (type: `boolean`):

Learns the domain's address pattern from e-mails on its own pages (or confirms a mailbox by SMTP where possible) and delivers `first.last@domain`-style addresses only when that evidence exists — charged $0.01 per delivered address. Unconfirmed guesses stay in `emailCandidates` for free. E-mails printed next to a person on the page are always included, uncharged.

## `includeTextMatches` (type: `boolean`):

Besides structured cards and schema.org data, read "Name, Title" / "Title: Name" patterns from page text (strong titles only, lower confidence).

## `useSitemap` (type: `boolean`):

Read `/sitemap.xml` for team-page URLs when the homepage links to none.

## `maxConcurrency` (type: `integer`):

Domains read at the same time (capped to 4 at 1024 MB of memory).

## `proxyConfiguration` (type: `object`):

Sites are read through the run's own connection first; this proxy is used only for a site that blocks it. Residential by default.

## Actor input object example

```json
{
  "domains": [
    "hubspot.com",
    "thoughtbot.com",
    "kleinerperkins.com"
  ],
  "maxPagesPerSite": 6,
  "maxPeoplePerSite": 100,
  "onlyWithTitle": false,
  "onlyWithLinkedin": false,
  "findEmails": true,
  "includeTextMatches": true,
  "useSitemap": true,
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `people` (type: `string`):

One row per person: name, title, department, seniority, company, LinkedIn, X, e-mail, photo, bio, source page, company contacts.

## `csv` (type: `string`):

Spreadsheet / CRM-ready CSV of the people (first 5,000 rows).

## `summary` (type: `string`):

Counts per domain (status, pages read, people found and delivered), e-mail finding stats, request stats and notices.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "hubspot.com",
        "thoughtbot.com",
        "kleinerperkins.com"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("inovaflow/team-page-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "hubspot.com",
        "thoughtbot.com",
        "kleinerperkins.com",
    ],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("inovaflow/team-page-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "hubspot.com",
    "thoughtbot.com",
    "kleinerperkins.com"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call inovaflow/team-page-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,inovaflow/team-page-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Bz7i7iREEFlAdfygL/builds/NhMsQsljdFaA3snKD/openapi.json
