Team Page Extractor — Names, Titles & LinkedIn from Websites avatar

Team Page Extractor — Names, Titles & LinkedIn from Websites

Pricing

from $10.00 / 1,000 people

Go to Apify Store
Team Page Extractor — Names, Titles & LinkedIn from Websites

Team Page Extractor — Names, Titles & LinkedIn from Websites

Give it company domains and get every person on their team, about and leadership pages: name, title, department, seniority, LinkedIn, X, photo, bio, plus company phone and e-mails. Optional pattern-based work e-mail with MX check. No login, dataset-only, MCP-ready.

Pricing

from $10.00 / 1,000 people

Rating

0.0

(0)

Developer

inovaflow

inovaflow

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

1

Monthly active users

3 days ago

Last modified

Share

Give it company domains and get every person listed on their team, about and leadership pages — name, job title, department, seniority, LinkedIn / X profile, headshot, bio, the e-mail printed next to them, plus the company's phone and e-mails. Optional pattern-based work e-mail with a mail-server check. One row per person, one price per row. No login, no API key, no browser — the site's own HTML, read structurally.

Who it is for

  • List builders and SDR teams who have accounts (domains) and need the people at them without a data vendor.
  • Agencies and recruiters mapping who does what at a target company — with the source page for every row.
  • Investors, journalists, analysts pulling leadership and partner lists from hundreds of firms at once.
  • AI agents (MCP): give it domains, get flat rows with stable ids — the no-search fast path next to Decision Maker Finder.

What it does

For every domain:

  1. reads the homepage (company name, phone, e-mails, LinkedIn page) and finds the pages most likely to list people: links whose path or label says team / people / leadership / about / management / staff / attorneys / doctors …, the well-known paths (/team, /about, /about-us, /people, /leadership, /our-team, /company, /management, /staff, /who-we-are, localized variants), and sitemap.xml hints;
  2. reads the best candidates within your page budget (pagination of a productive team page included);
  3. extracts people structurally — schema.org Person / ItemList JSON-LD, repeated cards (name element + role line + LinkedIn / X link + headshot + bio), hydration blobs of client-rendered sites, and "Name, Title" prose — and maps profile links to people by proximity when a site keeps them in a data blob;
  4. rejects what is not a colleague: testimonials, blog authors, job postings, menu items, portfolio founders and customer quotes (people whose title names another company), product tiles, locations;
  5. classifies the title into department (sales, marketing, engineering, …) and seniority (founder, c-level, vp, director, head, manager, individual);
  6. optionally finds a work e-mail: learns the domain's address pattern from e-mails on its own pages, checks the mail server (and the mailbox where SMTP is reachable), and delivers the address only with that evidence.

Why this one

  • Accuracy over coverage. Every row carries the page it came from; a bare name never ships — a card needs a role line, a profile link or a headshot in a repeated grid. Titles are never invented, e-mails never guessed.
  • Any stack. Tested on WordPress, Webflow, Framer, Next.js, Gatsby, Squarespace-style static sites, law-firm and clinic CMSs, VC sites: 40+ domains hand-checked (see docs/DECISIONS.md §3).
  • Cheap and fast. Pure HTTP; 6 pages per site by default; ~5–15 s per domain; the run's own connection first and your proxy only when a site blocks it.
  • Dataset-only, MCP-ready. Flat rows, stable id (domain:name-slug), sources[], scrapedAt, a run summary in OUTPUT, a CSV in the key-value store.

Fields

FieldTypeMeaning
idstringdomain:first-last — stable across runs
fullName, firstName, lastNamestringas printed (honorifics and post-nominals stripped)
titlestring | nullthe role line next to the name
departmentenumexecutive, sales, marketing, engineering, product, data, design, finance, operations, hr, customer-success, it, legal, other
seniorityenumfounder, c-level, vp, director, head, manager, individual
companystring | nullsite name (og:site_name / schema.org Organization / title)
domain, websitestringthe company site (after redirects)
linkedin, twitterstring | nullpersonal profile URLs found on the page
emailstring | nullthe person's e-mail: printed on the page (emailSource: "page") or pattern-built and mail-server-checked ("pattern")
emailConfidencenumber | null100 for a page e-mail; 70–95 for a pattern e-mail
emailVerificationenum | nullsmtp-valid, smtp-catch-all, mx-only, no-mx, syntax-only
emailCandidates[]arrayunconfirmed guesses (address, pattern, confidence, verification) — free
photoUrlstring | nullheadshot
biostring | nullfirst 300 chars of the bio
sourceUrlstringthe page the person was read from
extractedByenumjson-ld, card, blob, text
companyPhone, companyEmails[], companyLinkedin—company-level contacts from the pages read
teamPages[], sources[], scrapedAt—provenance

How to use

  1. Paste domains (or team-page URLs you already know) into Company domains.
  2. Optionally filter by roles (CEO, VP Sales, Head of Marketing, Partner), departments or seniorities, and require a title or a LinkedIn profile.
  3. Keep Find work e-mails on to get pattern-based addresses ($0.01 each, only when confirmed).
  4. Run. Rows land in the dataset (views People and Leads & contacts), PEOPLE.csv and the OUTPUT summary in the key-value store.

Cost

Pay per event — see PRICING.md:

  • Person $0.01 — one delivered person (charged once even if found on several pages).
  • Work e-mail $0.01 — only when a pattern-built, mail-server-checked address is delivered. Page e-mails are free.
  • Actor start $0.005.

A 10-domain run with ~200 people costs about $2.10 plus a few cents of platform usage.

Input

{ "domains": ["hubspot.com", "thoughtbot.com", "kleinerperkins.com"], "maxPagesPerSite": 6, "maxPeoplePerSite": 100, "findEmails": true }

Output sample

{
"id": "hubspot.com:yamini-rangan",
"fullName": "Yamini Rangan",
"firstName": "Yamini",
"lastName": "Rangan",
"title": "Chief Executive Officer",
"department": "executive",
"seniority": "c-level",
"company": "HubSpot",
"domain": "hubspot.com",
"website": "https://www.hubspot.com/",
"linkedin": "https://www.linkedin.com/in/yaminirangan",
"twitter": null,
"email": null,
"emailSource": null,
"emailConfidence": null,
"emailVerification": null,
"emailCandidates": [{ "address": "yamini.rangan@hubspot.com", "pattern": "first.last", "confidence": 45, "verification": "mx-only" }],
"photoUrl": "https://www.hubspot.com/hs-fs/hubfs/assets/hubspot.com/about/management%202025/Yamini-Rangan-Headshot.webp",
"bio": null,
"sourceUrl": "https://www.hubspot.com/company/management",
"extractedBy": "card",
"companyPhone": "+18884827768",
"companyEmails": [],
"companyLinkedin": "https://www.linkedin.com/company/hubspot",
"teamPages": ["https://www.hubspot.com/company/management", "https://www.hubspot.com/company/board-of-directors"],
"sources": ["https://www.hubspot.com/company/management"],
"scrapedAt": "2026-09-26T15:20:11.000Z"
}

FAQ

Why is title null for some people? The page shows only a name and a photo (many VC and agency grids). The person still ships when the card sits in a repeated headshot grid on a listing page; turn on Only people with a job title to drop them.

Why did a domain return no people? The site renders its team list in the browser (search apps of big law firms, some React sites) or has no team page. The run summary (OUTPUT.perDomain) says no_people with the pages read. Pasting the exact team-page URL as the domain often helps.

Are board members and advisors included? Yes when the site lists them on a leadership / board page; their title keeps the other company ("General Partner, ICONIQ Capital").

Does it charge for e-mail guesses? No. Only addresses backed by the domain's own pattern (or an SMTP-confirmed mailbox) are delivered and charged; the rest stay in emailCandidates with their confidence.

Blocked sites? The direct connection is tried first, then your proxy (residential by default). Sites that block both are reported as blocked in the summary and never charged.