Linkedin Lead Scraper & Company Website Enrichment
Pricing
$19.99/month + usage
Linkedin Lead Scraper & Company Website Enrichment
Extract high-quality professional leads using the LinkedIn Lead Scraper. Collect names, job titles, company names, locations, and profile links automatically. Ideal for B2B prospecting, recruitment research, and building targeted outreach lists.
Pricing
$19.99/month + usage
Rating
5.0
(2)
Developer
Scraper Engine
Maintained by CommunityActor stats
0
Bookmarked
34
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
LinkedIn Lead Scraper — Emails, Company Domains and Tech Stack Data
LinkedIn Lead Scraper & Company Website Enrichment finds public LinkedIn profiles, posts, and pages that mention your keywords and expose a visible email address, then returns each lead as structured JSON with email, companyDomain, techStack, and a derived emailPattern — no LinkedIn login, cookies, or browser session required. Every field arrives typed and ready to load into a CRM, spreadsheet, or LLM pipeline. Start a run from the Actor's Apify Console page to watch leads land in the dataset in real time.
What is LinkedIn Lead Scraper & Company Website Enrichment?
It's an Apify Actor that searches Google for site:linkedin.com results matching your keywords, keeps only the results that carry a visible email address, and — optionally — fetches the homepage of each discovered company domain to read its public tech stack and any additional on-site contact emails. It never authenticates to LinkedIn: no account, cookie, or login is used or required at any point in the pipeline. It's built for B2B sales and growth teams, recruiters, and developers who want structured lead data (not raw HTML) flowing into a CRM or an AI agent.
What LinkedIn lead data is publicly available to scrape?
LinkedIn only indexes into Google what a logged-out visitor could already see on a public profile, post, or company page — a title, a text snippet, and a URL. The full page experience (connection graph, activity feed, contact-info panel, message inbox) sits behind LinkedIn's own login wall, and this Actor never visits the LinkedIn page itself — it reads only what Google's search result already shows.
| Data Category | Publicly indexed (Google result) | Restricted (LinkedIn login) |
|---|---|---|
| Result title / headline | ✅ | — |
| Snippet text (bio or post excerpt) | ✅ | — |
| Full profile bio / full post body | — | ✅ |
| Public profile, post, or page URL | ✅ | — |
| Email address, when visible in the indexed text | ✅ (only if present) | — |
| Connections list, follower/network graph | — | ✅ |
| Direct messages / InMail | — | ✅ |
| Full work history and education timeline | — | ✅ |
LinkedIn Lead Scraper & Company Website Enrichment only returns publicly visible data surfaced through Google's own index of LinkedIn pages — nothing behind a login wall.
What data can I extract with LinkedIn Lead Scraper & Company Website Enrichment?
Every run returns lead-identity fields, derived email intelligence, and — when enabled — company-domain enrichment data, all in one flat JSON row per lead.
| Field Name | Description |
|---|---|
network | Fixed source label, Linkedin.com. |
keyword | The search keyword that produced this row. |
title | The result's title/headline text. |
description | The snippet text (bio or post excerpt), truncated to 500 characters. |
url | The normalized LinkedIn profile/post/page URL (fragment stripped). |
email | The discovered email address, lowercased. |
companyDomain | The domain portion of email. |
isCorporateEmail | true when the email domain is not a known personal-email provider; null if no domain could be read. |
isPersonalEmail | true when the email domain is a known personal-email provider (Gmail, Yahoo, Outlook, etc.); null if no domain could be read. |
emailPattern | Generalized shape of real emails seen for this domain in this run, e.g. {first}.{last}@, {first}_{last}@, {first}-{last}@, or {first}@. null if pattern detection is off or no classifiable sample exists. |
emailPatternConfidence | high, medium, or low — how much the pattern agrees across the real samples seen so far. null when not computed. |
patternSampleCount | Number of real emails from this domain observed so far in the run. null when not computed. |
websiteEnrichmentStatus | enriched, fetch_failed, skipped_cap, or disabled — whether and why this domain's site was fetched. |
techStack | Detected technographic signatures for the domain (see below), or null if none detected or not enriched. |
additionalContactEmails | Extra on-site emails found on the company's own homepage that verifiably belong to that domain (or a personal-email provider) — or null. |
scrapedAt | ISO-8601 UTC timestamp of when the row was collected. |
Lead identity and source fields
network, keyword, title, description, url, email, companyDomain — what the lead is, where it came from, and how to reach them.
Derived email and domain intelligence
isCorporateEmail, isPersonalEmail, emailPattern, emailPatternConfidence, patternSampleCount — computed for free from the real emails already discovered this run, with zero extra requests.
Company-website enrichment and timestamps
websiteEnrichmentStatus, techStack, additionalContactEmails, scrapedAt — populated only when enrichCompanyWebsites is on and the target domain answered.
🤖 Add-on: Need additional LinkedIn data?
Pair this Actor with LinkedIn Company Profile Scraper to pull full company-page details (overview and full views) for the domains you've discovered, or LinkedIn Mass Company Profile Finder By Category Directory to expand from a keyword into a broader directory of matching company profiles before running lead discovery against them.
How does LinkedIn Lead Scraper & Company Website Enrichment differ from LinkedIn's official Marketing APIs?
LinkedIn's own Lead Sync and Community Management APIs only return leads your app already owns — form submissions from your own LinkedIn ad campaigns — not new prospects discovered from a keyword. Per LinkedIn's Marketing API quick-start documentation (checked 2026-07-30), using them requires creating a developer app, applying for product access, and passing LinkedIn's app review before your first call.
| Feature | LinkedIn Marketing APIs (Lead Sync / Community Management) | LinkedIn Lead Scraper & Company Website Enrichment |
|---|---|---|
| Access requirement | Developer app + per-product review and approval | No signup — runs directly on Apify |
| Data scope | Leads captured via your own native Lead Gen Forms | Any public profile, post, or page matching your keyword |
| New-prospect discovery | Not supported — returns only your own captured leads | Core function — finds new leads by keyword |
| Authentication | OAuth app + member authorization | None |
| Setup time | Application, review, and tier approval | Configure input fields and start a run |
| Company-domain enrichment | Not part of these APIs | Built in — tech stack and email-pattern intelligence per domain |
Use LinkedIn's official APIs when you need to sync leads already captured through your own LinkedIn ad campaigns. Use LinkedIn Lead Scraper & Company Website Enrichment when you need to discover new public leads by keyword without an app-review process.
How to use LinkedIn Lead Scraper & Company Website Enrichment
No account setup beyond signing in to Apify is required — the Actor runs entirely through the Apify Console or API.
- Open the Actor's page on Apify and click Try for free / Start.
- Provide at least one term in
searchKeywords— the run produces no output without one. - Optionally set
targetLocation,personalEmailDomainsFilter, and turn onenrichCompanyWebsitesif you want tech-stack and pattern data. - Start the run.
- Download results as JSON, CSV, or another Apify dataset export format, or stream them via the API as they're pushed.
How to scale to bulk lead extraction
searchKeywords accepts an array, so a single run processes every keyword in it sequentially, each collecting up to maxEmailsPerKeyword rows. There is no separate bulk URL-list input — scaling means adding more keywords to the array (up to the 5000-per-keyword ceiling) rather than supplying a list of target IDs. targetLocation and personalEmailDomainsFilter apply to the whole run, not per keyword; to vary location per segment, run the Actor once per location.
What can you do with LinkedIn lead data?
- 📈 Sales prospecting — a sales development rep building an outreach list uses
emailandisCorporateEmailto filter down to verified corporate contacts before loading them into a CRM. - 🏢 Account-based marketing — a growth marketer uses
techStackto spot which prospects already run Google Analytics or Shopify, then tailors the pitch to what they're missing. - ✉️ Cold outreach QA — an SDR uses
emailPatternandemailPatternConfidenceto sanity-check the format of a manually-collected address for the same company before sending, without ever emailing a guessed address. - 🔍 Recruitment sourcing — a recruiter uses
title,description, andurlto find keyword-matched public LinkedIn activity that also surfaces a contact email. - 🤖 AI agent research — an AI engineer feeds the typed JSON rows (
title,description,techStack) into a RAG pipeline or agent tool as structured account-research context ahead of a personalized outreach step.
How does LinkedIn Lead Scraper & Company Website Enrichment handle rate limits and blocking?
Every search request is routed through Apify Proxy (the GOOGLE_SERP group by default, or your own proxyConfiguration if you supply one) with a randomized user agent and Accept-Language header, and a randomized delay before each request. If a response looks blocked (non-200 status, or a small page containing phrases like "unusual traffic" or "captcha" — a size check keeps this from misreading a genuine full results page), the Actor rotates to a new proxy URL, waits, and retries, up to 3 attempts total before raising an error for that request. Company-website enrichment fetches use a separate proxy resolution, since the GOOGLE_SERP group only permits requests to Google itself. ⚠️ If a keyword returns 5 consecutive pages with no new email, the Actor stops that keyword and moves to the next one rather than retrying indefinitely — this is a real stopping condition, not a bug. The Actor does not solve CAPTCHAs; a persistently blocked request fails that request after its retries are exhausted.
⬇️ Input
LinkedIn Lead Scraper & Company Website Enrichment has no schema-required fields, but a run with no searchKeywords (or legacy keywords) produces zero results.
| Parameter | Required | Type | Description | Example Value |
|---|---|---|---|---|
searchKeywords | No | array | Keywords to search for on LinkedIn. Each keyword is searched for LinkedIn profiles/posts containing a public email address. | ["marketing", "founder"] |
targetLocation | No | string | Optional location to narrow the search. Leave empty to search globally. | "New York" |
personalEmailDomainsFilter | No | array | Only keep rows whose discovered email ends with one of these domains. Leave empty to keep every domain found. | ["@gmail.com"] |
maxEmailsPerKeyword | No | integer (min 1, max 5000) | Maximum number of email leads to collect per keyword. Default is 20. | 20 |
enrichCompanyWebsites | No | boolean (default false) | Fetches the homepage of each unique company domain discovered (once per domain) and reads its tech stack plus any additional on-site contact emails. | true |
maxDomainsToEnrich | No | integer (default 25, min 1, max 500) | Upper limit on how many unique company domains get an enrichment fetch in a single run. Only used when enrichCompanyWebsites is on. | 25 |
detectEmailPatternIntelligence | No | boolean (default true) | Derives companyDomain, isCorporateEmail/isPersonalEmail, and an inferred emailPattern per domain from real discovered emails — zero extra requests. Never generates a guessed email. | true |
keywords | No | array | Legacy alias of searchKeywords, kept for backward compatibility. | ["marketing"] |
location | No | string | Legacy alias of targetLocation, kept for backward compatibility. | "London" |
emailDomains | No | array | Legacy alias of personalEmailDomainsFilter, kept for backward compatibility. | ["@outlook.com"] |
maxEmails | No | integer (min 1, max 5000) | Legacy alias of maxEmailsPerKeyword, kept for backward compatibility. | 20 |
proxyConfiguration | No | object | Apify proxy configuration for the run. Defaults to the GOOGLE_SERP proxy group. | {"useApifyProxy": true, "apifyProxyGroups": ["GOOGLE_SERP"]} |
Example input
{"searchKeywords": ["marketing", "founder"],"targetLocation": "New York","personalEmailDomainsFilter": ["@gmail.com"],"maxEmailsPerKeyword": 20,"enrichCompanyWebsites": true,"maxDomainsToEnrich": 25,"detectEmailPatternIntelligence": true,"proxyConfiguration": {"useApifyProxy": true,"apifyProxyGroups": ["GOOGLE_SERP"]}}
⬆️ Output
Results push to the Apify dataset as typed JSON, one row per lead, with a consistent schema across runs. Export as JSON, CSV, Excel, or any other format the Apify dataset supports. Every pushed row is charged under the row_result event — there are no separate uncharged error or accounting rows, so the dataset's row count matches what you're billed for.
Example output
{"network": "Linkedin.com","keyword": "founder","title": "Jane Smith - Founder & CEO - Acme Robotics | LinkedIn","description": "Reach out at jane.smith@acmerobotics.com for partnership inquiries...","url": "https://www.linkedin.com/in/janesmith","email": "jane.smith@acmerobotics.com","companyDomain": "acmerobotics.com","isCorporateEmail": true,"isPersonalEmail": false,"emailPattern": "{first}.{last}@","emailPatternConfidence": "medium","patternSampleCount": 2,"websiteEnrichmentStatus": "enriched","techStack": {"analytics": ["Google Analytics", "Google Tag Manager"],"cms": ["Webflow"],"cdn": ["Cloudflare"]},"additionalContactEmails": ["hello@acmerobotics.com"],"scrapedAt": "2026-07-30T12:00:00+00:00"}
How does it work?
Search requests go to Google — not to LinkedIn directly — routed through Apify Proxy with a randomized user agent, headers, and a jittered delay before each request, retrying with a rotated proxy if the response looks blocked. The result page is parsed structurally: every anchor tag is scanned and resolved to its real destination, kept only when it's genuinely hosted on linkedin.com (and not one of LinkedIn's own help, business, or developer subdomains), rather than matching Google's own frequently-changing CSS class names — so the parser survives Google markup churn. Title, snippet, and email are read from the smallest container that still bounds a single result. When company-website enrichment is on, each unique discovered domain's homepage is fetched once, over a separate proxy pool, to read its HTML, script sources, and response headers for tech-stack signatures. Only publicly visible data is ever returned, and the output schema stays the same regardless of how Google's or LinkedIn's markup changes.
Integrations
LinkedIn Lead Scraper & Company Website Enrichment runs as a standard Apify Actor, so it works with anything that can call the Apify API.
Calling LinkedIn Lead Scraper & Company Website Enrichment programmatically
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_API_TOKEN>")run = client.actor("<YOUR_USERNAME>/linkedin-lead-scraper-and-company-website-enrichment").call(run_input={"searchKeywords": ["marketing", "founder"],"maxEmailsPerKeyword": 20,"enrichCompanyWebsites": True,})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item["email"], item["companyDomain"])
Works in Go, Ruby, Node.js, cURL — any language that can make an HTTP request.
No-code tools (n8n, Make)
In n8n, use the Apify node (or an HTTP Request node pointed at the run endpoint) to start a run and pull dataset items into your workflow. In Make, use Apify's published app to trigger a run from a scenario and route each lead row into your CRM or spreadsheet module.
Is it legal to scrape LinkedIn leads?
Scraping publicly available data is generally lawful, but the emails and profile snippets this Actor collects are personal data, so GDPR and CCPA govern how you store and use them, not whether you may collect them. LinkedIn Lead Scraper & Company Website Enrichment returns only publicly available data — content already visible in Google's own index of public LinkedIn pages, nothing gated behind a login. Consult legal counsel if your use case involves bulk storage of personal data.
Frequently asked questions
What LinkedIn lead fields does this Actor return?
The core fields are email, companyDomain, techStack, emailPattern, and title — see the full data fields table above for all 16.
Does this Actor require a LinkedIn account or login?
No. It never authenticates to LinkedIn — it searches Google's public index of linkedin.com pages and reads whatever a logged-out visitor would already see in the search result.
How many leads can I extract in one run?
Up to maxEmailsPerKeyword (max 5000) per keyword in searchKeywords, across as many keywords as you list — a keyword stops early if 5 consecutive result pages return no new email.
What happens if a keyword returns zero visible emails?
The Actor logs each empty page and stops that keyword after 5 consecutive empty pages, then moves on to the next keyword. It does not fail the run — a keyword with no visible emails simply contributes zero rows.
Can I scrape multiple LinkedIn keywords at once?
Yes. searchKeywords accepts an array, and every keyword in it is processed within the same run.
Does this Actor work with Claude, ChatGPT, and other AI agent tools?
It has no dedicated MCP server. It's callable as a standard Apify Actor run through the Apify API or apify-client SDK by any agent framework that can make an HTTP request.
How does this differ from a plain LinkedIn email finder?
Most keyword-based email finders stop at title, snippet, URL, and email. This Actor adds two derived layers on top from data it has already collected at no extra cost: emailPattern/emailPatternConfidence (the observed shape of real addresses per company domain) and, when enrichCompanyWebsites is on, a techStack read of each company's own homepage.
Does this Actor return data in a format LLMs can use directly?
Yes. Typed, normalized JSON with consistent field names across runs — no HTML parsing or selectors needed. Pass it straight to an LLM, index it into a vector store, or feed it to an agent tool.
What happens when LinkedIn or Google changes its layout or anti-bot system?
The SERP parser is structure-based (it scans every link and resolves real destinations rather than matching a hardcoded CSS class name), which is designed to tolerate exactly this kind of markup churn. No specific turnaround time is promised for any given change.
Can I use this Actor without managing proxies or browser infrastructure?
Yes. It runs on Apify Proxy automatically (the GOOGLE_SERP group by default) and makes plain HTTP requests — there's no headless browser or proxy pool for you to configure or maintain.
Which fields work best for AI training data and RAG indexing?
For RAG, index title and description — the highest-information free text. For training or enrichment pipelines, email, companyDomain, and isCorporateEmail return as consistently structured typed primitives across every row.
Related scrapers
| Scraper Name | What it extracts |
|---|---|
| LinkedIn Company Profile Scraper | Full company-page details (overview and full views) for a list of LinkedIn company URLs. |
| LinkedIn Mass Company Profile Finder By Category Directory | Discovers LinkedIn company profiles in bulk by keyword, business category, or domain across directory pages. |
Your feedback
Found a bug or missing a field? Let us know through the Issues tab on this Actor's Apify Console page — it's the fastest way to reach the maintainer directly.