Team Page Extractor — Names, Titles & LinkedIn from Websites
Pricing
from $10.00 / 1,000 people
Team Page Extractor — Names, Titles & LinkedIn from Websites
Give it company domains and get every person on their team, about and leadership pages: name, title, department, seniority, LinkedIn, X, photo, bio, plus company phone and e-mails. Optional pattern-based work e-mail with MX check. No login, dataset-only, MCP-ready.
Pricing
from $10.00 / 1,000 people
Rating
0.0
(0)
Developer
inovaflow
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Give it company domains and get every person listed on their team, about and leadership pages — name, job title, department, seniority, LinkedIn / X profile, headshot, bio, the e-mail printed next to them, plus the company's phone and e-mails. Optional pattern-based work e-mail with a mail-server check. One row per person, one price per row. No login, no API key, no browser — the site's own HTML, read structurally.
Who it is for
- List builders and SDR teams who have accounts (domains) and need the people at them without a data vendor.
- Agencies and recruiters mapping who does what at a target company — with the source page for every row.
- Investors, journalists, analysts pulling leadership and partner lists from hundreds of firms at once.
- AI agents (MCP): give it domains, get flat rows with stable ids — the no-search fast path next to Decision Maker Finder.
What it does
For every domain:
- reads the homepage (company name, phone, e-mails, LinkedIn page) and finds the pages most likely to list
people: links whose path or label says team / people / leadership / about / management / staff / attorneys /
doctors …, the well-known paths (
/team,/about,/about-us,/people,/leadership,/our-team,/company,/management,/staff,/who-we-are, localized variants), andsitemap.xmlhints; - reads the best candidates within your page budget (pagination of a productive team page included);
- extracts people structurally — schema.org
Person/ItemListJSON-LD, repeated cards (name element + role line + LinkedIn / X link + headshot + bio), hydration blobs of client-rendered sites, and "Name, Title" prose — and maps profile links to people by proximity when a site keeps them in a data blob; - rejects what is not a colleague: testimonials, blog authors, job postings, menu items, portfolio founders and customer quotes (people whose title names another company), product tiles, locations;
- classifies the title into department (sales, marketing, engineering, …) and seniority (founder, c-level, vp, director, head, manager, individual);
- optionally finds a work e-mail: learns the domain's address pattern from e-mails on its own pages, checks the mail server (and the mailbox where SMTP is reachable), and delivers the address only with that evidence.
Why this one
- Accuracy over coverage. Every row carries the page it came from; a bare name never ships — a card needs a role line, a profile link or a headshot in a repeated grid. Titles are never invented, e-mails never guessed.
- Any stack. Tested on WordPress, Webflow, Framer, Next.js, Gatsby, Squarespace-style static sites, law-firm
and clinic CMSs, VC sites: 40+ domains hand-checked (see
docs/DECISIONS.md§3). - Cheap and fast. Pure HTTP; 6 pages per site by default; ~5–15 s per domain; the run's own connection first and your proxy only when a site blocks it.
- Dataset-only, MCP-ready. Flat rows, stable
id(domain:name-slug),sources[],scrapedAt, a run summary inOUTPUT, a CSV in the key-value store.
Fields
| Field | Type | Meaning |
|---|---|---|
id | string | domain:first-last — stable across runs |
fullName, firstName, lastName | string | as printed (honorifics and post-nominals stripped) |
title | string | null | the role line next to the name |
department | enum | executive, sales, marketing, engineering, product, data, design, finance, operations, hr, customer-success, it, legal, other |
seniority | enum | founder, c-level, vp, director, head, manager, individual |
company | string | null | site name (og:site_name / schema.org Organization / title) |
domain, website | string | the company site (after redirects) |
linkedin, twitter | string | null | personal profile URLs found on the page |
email | string | null | the person's e-mail: printed on the page (emailSource: "page") or pattern-built and mail-server-checked ("pattern") |
emailConfidence | number | null | 100 for a page e-mail; 70–95 for a pattern e-mail |
emailVerification | enum | null | smtp-valid, smtp-catch-all, mx-only, no-mx, syntax-only |
emailCandidates[] | array | unconfirmed guesses (address, pattern, confidence, verification) — free |
photoUrl | string | null | headshot |
bio | string | null | first 300 chars of the bio |
sourceUrl | string | the page the person was read from |
extractedBy | enum | json-ld, card, blob, text |
companyPhone, companyEmails[], companyLinkedin | — | company-level contacts from the pages read |
teamPages[], sources[], scrapedAt | — | provenance |
How to use
- Paste domains (or team-page URLs you already know) into Company domains.
- Optionally filter by roles (
CEO,VP Sales,Head of Marketing,Partner), departments or seniorities, and require a title or a LinkedIn profile. - Keep Find work e-mails on to get pattern-based addresses ($0.01 each, only when confirmed).
- Run. Rows land in the dataset (views People and Leads & contacts),
PEOPLE.csvand theOUTPUTsummary in the key-value store.
Cost
Pay per event — see PRICING.md:
- Person $0.01 — one delivered person (charged once even if found on several pages).
- Work e-mail $0.01 — only when a pattern-built, mail-server-checked address is delivered. Page e-mails are free.
- Actor start $0.005.
A 10-domain run with ~200 people costs about $2.10 plus a few cents of platform usage.
Input
{ "domains": ["hubspot.com", "thoughtbot.com", "kleinerperkins.com"], "maxPagesPerSite": 6, "maxPeoplePerSite": 100, "findEmails": true }
Output sample
{"id": "hubspot.com:yamini-rangan","fullName": "Yamini Rangan","firstName": "Yamini","lastName": "Rangan","title": "Chief Executive Officer","department": "executive","seniority": "c-level","company": "HubSpot","domain": "hubspot.com","website": "https://www.hubspot.com/","linkedin": "https://www.linkedin.com/in/yaminirangan","twitter": null,"email": null,"emailSource": null,"emailConfidence": null,"emailVerification": null,"emailCandidates": [{ "address": "yamini.rangan@hubspot.com", "pattern": "first.last", "confidence": 45, "verification": "mx-only" }],"photoUrl": "https://www.hubspot.com/hs-fs/hubfs/assets/hubspot.com/about/management%202025/Yamini-Rangan-Headshot.webp","bio": null,"sourceUrl": "https://www.hubspot.com/company/management","extractedBy": "card","companyPhone": "+18884827768","companyEmails": [],"companyLinkedin": "https://www.linkedin.com/company/hubspot","teamPages": ["https://www.hubspot.com/company/management", "https://www.hubspot.com/company/board-of-directors"],"sources": ["https://www.hubspot.com/company/management"],"scrapedAt": "2026-09-26T15:20:11.000Z"}
FAQ
Why is title null for some people? The page shows only a name and a photo (many VC and agency grids). The
person still ships when the card sits in a repeated headshot grid on a listing page; turn on Only people with a job
title to drop them.
Why did a domain return no people? The site renders its team list in the browser (search apps of big law firms,
some React sites) or has no team page. The run summary (OUTPUT.perDomain) says no_people with the pages read.
Pasting the exact team-page URL as the domain often helps.
Are board members and advisors included? Yes when the site lists them on a leadership / board page; their title keeps the other company ("General Partner, ICONIQ Capital").
Does it charge for e-mail guesses? No. Only addresses backed by the domain's own pattern (or an SMTP-confirmed
mailbox) are delivered and charged; the rest stay in emailCandidates with their confidence.
Blocked sites? The direct connection is tried first, then your proxy (residential by default). Sites that block
both are reported as blocked in the summary and never charged.