Linkedin Lead Scraper & Company Website Enrichment avatar

Linkedin Lead Scraper & Company Website Enrichment

Pricing

$19.99/month + usage

Go to Apify Store
Linkedin Lead Scraper & Company Website Enrichment

Linkedin Lead Scraper & Company Website Enrichment

Extract high-quality professional leads using the LinkedIn Lead Scraper. Collect names, job titles, company names, locations, and profile links automatically. Ideal for B2B prospecting, recruitment research, and building targeted outreach lists.

Pricing

$19.99/month + usage

Rating

5.0

(2)

Developer

Scraper Engine

Scraper Engine

Maintained by Community

Actor stats

0

Bookmarked

34

Total users

1

Monthly active users

2 days ago

Last modified

Share

LinkedIn Lead Scraper — Emails, Company Domains and Tech Stack Data

LinkedIn Lead Scraper & Company Website Enrichment finds public LinkedIn profiles, posts, and pages that mention your keywords and expose a visible email address, then returns each lead as structured JSON with email, companyDomain, techStack, and a derived emailPattern — no LinkedIn login, cookies, or browser session required. Every field arrives typed and ready to load into a CRM, spreadsheet, or LLM pipeline. Start a run from the Actor's Apify Console page to watch leads land in the dataset in real time.

What is LinkedIn Lead Scraper & Company Website Enrichment?

It's an Apify Actor that searches Google for site:linkedin.com results matching your keywords, keeps only the results that carry a visible email address, and — optionally — fetches the homepage of each discovered company domain to read its public tech stack and any additional on-site contact emails. It never authenticates to LinkedIn: no account, cookie, or login is used or required at any point in the pipeline. It's built for B2B sales and growth teams, recruiters, and developers who want structured lead data (not raw HTML) flowing into a CRM or an AI agent.

What LinkedIn lead data is publicly available to scrape?

LinkedIn only indexes into Google what a logged-out visitor could already see on a public profile, post, or company page — a title, a text snippet, and a URL. The full page experience (connection graph, activity feed, contact-info panel, message inbox) sits behind LinkedIn's own login wall, and this Actor never visits the LinkedIn page itself — it reads only what Google's search result already shows.

Data CategoryPublicly indexed (Google result)Restricted (LinkedIn login)
Result title / headline
Snippet text (bio or post excerpt)
Full profile bio / full post body
Public profile, post, or page URL
Email address, when visible in the indexed text✅ (only if present)
Connections list, follower/network graph
Direct messages / InMail
Full work history and education timeline

LinkedIn Lead Scraper & Company Website Enrichment only returns publicly visible data surfaced through Google's own index of LinkedIn pages — nothing behind a login wall.

What data can I extract with LinkedIn Lead Scraper & Company Website Enrichment?

Every run returns lead-identity fields, derived email intelligence, and — when enabled — company-domain enrichment data, all in one flat JSON row per lead.

Field NameDescription
networkFixed source label, Linkedin.com.
keywordThe search keyword that produced this row.
titleThe result's title/headline text.
descriptionThe snippet text (bio or post excerpt), truncated to 500 characters.
urlThe normalized LinkedIn profile/post/page URL (fragment stripped).
emailThe discovered email address, lowercased.
companyDomainThe domain portion of email.
isCorporateEmailtrue when the email domain is not a known personal-email provider; null if no domain could be read.
isPersonalEmailtrue when the email domain is a known personal-email provider (Gmail, Yahoo, Outlook, etc.); null if no domain could be read.
emailPatternGeneralized shape of real emails seen for this domain in this run, e.g. {first}.{last}@, {first}_{last}@, {first}-{last}@, or {first}@. null if pattern detection is off or no classifiable sample exists.
emailPatternConfidencehigh, medium, or low — how much the pattern agrees across the real samples seen so far. null when not computed.
patternSampleCountNumber of real emails from this domain observed so far in the run. null when not computed.
websiteEnrichmentStatusenriched, fetch_failed, skipped_cap, or disabled — whether and why this domain's site was fetched.
techStackDetected technographic signatures for the domain (see below), or null if none detected or not enriched.
additionalContactEmailsExtra on-site emails found on the company's own homepage that verifiably belong to that domain (or a personal-email provider) — or null.
scrapedAtISO-8601 UTC timestamp of when the row was collected.

Lead identity and source fields

network, keyword, title, description, url, email, companyDomain — what the lead is, where it came from, and how to reach them.

Derived email and domain intelligence

isCorporateEmail, isPersonalEmail, emailPattern, emailPatternConfidence, patternSampleCount — computed for free from the real emails already discovered this run, with zero extra requests.

Company-website enrichment and timestamps

websiteEnrichmentStatus, techStack, additionalContactEmails, scrapedAt — populated only when enrichCompanyWebsites is on and the target domain answered.

🤖 Add-on: Need additional LinkedIn data?

Pair this Actor with LinkedIn Company Profile Scraper to pull full company-page details (overview and full views) for the domains you've discovered, or LinkedIn Mass Company Profile Finder By Category Directory to expand from a keyword into a broader directory of matching company profiles before running lead discovery against them.

How does LinkedIn Lead Scraper & Company Website Enrichment differ from LinkedIn's official Marketing APIs?

LinkedIn's own Lead Sync and Community Management APIs only return leads your app already owns — form submissions from your own LinkedIn ad campaigns — not new prospects discovered from a keyword. Per LinkedIn's Marketing API quick-start documentation (checked 2026-07-30), using them requires creating a developer app, applying for product access, and passing LinkedIn's app review before your first call.

FeatureLinkedIn Marketing APIs (Lead Sync / Community Management)LinkedIn Lead Scraper & Company Website Enrichment
Access requirementDeveloper app + per-product review and approvalNo signup — runs directly on Apify
Data scopeLeads captured via your own native Lead Gen FormsAny public profile, post, or page matching your keyword
New-prospect discoveryNot supported — returns only your own captured leadsCore function — finds new leads by keyword
AuthenticationOAuth app + member authorizationNone
Setup timeApplication, review, and tier approvalConfigure input fields and start a run
Company-domain enrichmentNot part of these APIsBuilt in — tech stack and email-pattern intelligence per domain

Use LinkedIn's official APIs when you need to sync leads already captured through your own LinkedIn ad campaigns. Use LinkedIn Lead Scraper & Company Website Enrichment when you need to discover new public leads by keyword without an app-review process.

How to use LinkedIn Lead Scraper & Company Website Enrichment

No account setup beyond signing in to Apify is required — the Actor runs entirely through the Apify Console or API.

  1. Open the Actor's page on Apify and click Try for free / Start.
  2. Provide at least one term in searchKeywords — the run produces no output without one.
  3. Optionally set targetLocation, personalEmailDomainsFilter, and turn on enrichCompanyWebsites if you want tech-stack and pattern data.
  4. Start the run.
  5. Download results as JSON, CSV, or another Apify dataset export format, or stream them via the API as they're pushed.

How to scale to bulk lead extraction

searchKeywords accepts an array, so a single run processes every keyword in it sequentially, each collecting up to maxEmailsPerKeyword rows. There is no separate bulk URL-list input — scaling means adding more keywords to the array (up to the 5000-per-keyword ceiling) rather than supplying a list of target IDs. targetLocation and personalEmailDomainsFilter apply to the whole run, not per keyword; to vary location per segment, run the Actor once per location.

What can you do with LinkedIn lead data?

  • 📈 Sales prospecting — a sales development rep building an outreach list uses email and isCorporateEmail to filter down to verified corporate contacts before loading them into a CRM.
  • 🏢 Account-based marketing — a growth marketer uses techStack to spot which prospects already run Google Analytics or Shopify, then tailors the pitch to what they're missing.
  • ✉️ Cold outreach QA — an SDR uses emailPattern and emailPatternConfidence to sanity-check the format of a manually-collected address for the same company before sending, without ever emailing a guessed address.
  • 🔍 Recruitment sourcing — a recruiter uses title, description, and url to find keyword-matched public LinkedIn activity that also surfaces a contact email.
  • 🤖 AI agent research — an AI engineer feeds the typed JSON rows (title, description, techStack) into a RAG pipeline or agent tool as structured account-research context ahead of a personalized outreach step.

How does LinkedIn Lead Scraper & Company Website Enrichment handle rate limits and blocking?

Every search request is routed through Apify Proxy (the GOOGLE_SERP group by default, or your own proxyConfiguration if you supply one) with a randomized user agent and Accept-Language header, and a randomized delay before each request. If a response looks blocked (non-200 status, or a small page containing phrases like "unusual traffic" or "captcha" — a size check keeps this from misreading a genuine full results page), the Actor rotates to a new proxy URL, waits, and retries, up to 3 attempts total before raising an error for that request. Company-website enrichment fetches use a separate proxy resolution, since the GOOGLE_SERP group only permits requests to Google itself. ⚠️ If a keyword returns 5 consecutive pages with no new email, the Actor stops that keyword and moves to the next one rather than retrying indefinitely — this is a real stopping condition, not a bug. The Actor does not solve CAPTCHAs; a persistently blocked request fails that request after its retries are exhausted.

⬇️ Input

LinkedIn Lead Scraper & Company Website Enrichment has no schema-required fields, but a run with no searchKeywords (or legacy keywords) produces zero results.

ParameterRequiredTypeDescriptionExample Value
searchKeywordsNoarrayKeywords to search for on LinkedIn. Each keyword is searched for LinkedIn profiles/posts containing a public email address.["marketing", "founder"]
targetLocationNostringOptional location to narrow the search. Leave empty to search globally."New York"
personalEmailDomainsFilterNoarrayOnly keep rows whose discovered email ends with one of these domains. Leave empty to keep every domain found.["@gmail.com"]
maxEmailsPerKeywordNointeger (min 1, max 5000)Maximum number of email leads to collect per keyword. Default is 20.20
enrichCompanyWebsitesNoboolean (default false)Fetches the homepage of each unique company domain discovered (once per domain) and reads its tech stack plus any additional on-site contact emails.true
maxDomainsToEnrichNointeger (default 25, min 1, max 500)Upper limit on how many unique company domains get an enrichment fetch in a single run. Only used when enrichCompanyWebsites is on.25
detectEmailPatternIntelligenceNoboolean (default true)Derives companyDomain, isCorporateEmail/isPersonalEmail, and an inferred emailPattern per domain from real discovered emails — zero extra requests. Never generates a guessed email.true
keywordsNoarrayLegacy alias of searchKeywords, kept for backward compatibility.["marketing"]
locationNostringLegacy alias of targetLocation, kept for backward compatibility."London"
emailDomainsNoarrayLegacy alias of personalEmailDomainsFilter, kept for backward compatibility.["@outlook.com"]
maxEmailsNointeger (min 1, max 5000)Legacy alias of maxEmailsPerKeyword, kept for backward compatibility.20
proxyConfigurationNoobjectApify proxy configuration for the run. Defaults to the GOOGLE_SERP proxy group.{"useApifyProxy": true, "apifyProxyGroups": ["GOOGLE_SERP"]}

Example input

{
"searchKeywords": ["marketing", "founder"],
"targetLocation": "New York",
"personalEmailDomainsFilter": ["@gmail.com"],
"maxEmailsPerKeyword": 20,
"enrichCompanyWebsites": true,
"maxDomainsToEnrich": 25,
"detectEmailPatternIntelligence": true,
"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyGroups": ["GOOGLE_SERP"]
}
}

⬆️ Output

Results push to the Apify dataset as typed JSON, one row per lead, with a consistent schema across runs. Export as JSON, CSV, Excel, or any other format the Apify dataset supports. Every pushed row is charged under the row_result event — there are no separate uncharged error or accounting rows, so the dataset's row count matches what you're billed for.

Example output

{
"network": "Linkedin.com",
"keyword": "founder",
"title": "Jane Smith - Founder & CEO - Acme Robotics | LinkedIn",
"description": "Reach out at jane.smith@acmerobotics.com for partnership inquiries...",
"url": "https://www.linkedin.com/in/janesmith",
"email": "jane.smith@acmerobotics.com",
"companyDomain": "acmerobotics.com",
"isCorporateEmail": true,
"isPersonalEmail": false,
"emailPattern": "{first}.{last}@",
"emailPatternConfidence": "medium",
"patternSampleCount": 2,
"websiteEnrichmentStatus": "enriched",
"techStack": {
"analytics": ["Google Analytics", "Google Tag Manager"],
"cms": ["Webflow"],
"cdn": ["Cloudflare"]
},
"additionalContactEmails": ["hello@acmerobotics.com"],
"scrapedAt": "2026-07-30T12:00:00+00:00"
}

How does it work?

Search requests go to Google — not to LinkedIn directly — routed through Apify Proxy with a randomized user agent, headers, and a jittered delay before each request, retrying with a rotated proxy if the response looks blocked. The result page is parsed structurally: every anchor tag is scanned and resolved to its real destination, kept only when it's genuinely hosted on linkedin.com (and not one of LinkedIn's own help, business, or developer subdomains), rather than matching Google's own frequently-changing CSS class names — so the parser survives Google markup churn. Title, snippet, and email are read from the smallest container that still bounds a single result. When company-website enrichment is on, each unique discovered domain's homepage is fetched once, over a separate proxy pool, to read its HTML, script sources, and response headers for tech-stack signatures. Only publicly visible data is ever returned, and the output schema stays the same regardless of how Google's or LinkedIn's markup changes.

Integrations

LinkedIn Lead Scraper & Company Website Enrichment runs as a standard Apify Actor, so it works with anything that can call the Apify API.

Calling LinkedIn Lead Scraper & Company Website Enrichment programmatically

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_API_TOKEN>")
run = client.actor("<YOUR_USERNAME>/linkedin-lead-scraper-and-company-website-enrichment").call(
run_input={
"searchKeywords": ["marketing", "founder"],
"maxEmailsPerKeyword": 20,
"enrichCompanyWebsites": True,
}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item["email"], item["companyDomain"])

Works in Go, Ruby, Node.js, cURL — any language that can make an HTTP request.

No-code tools (n8n, Make)

In n8n, use the Apify node (or an HTTP Request node pointed at the run endpoint) to start a run and pull dataset items into your workflow. In Make, use Apify's published app to trigger a run from a scenario and route each lead row into your CRM or spreadsheet module.

Scraping publicly available data is generally lawful, but the emails and profile snippets this Actor collects are personal data, so GDPR and CCPA govern how you store and use them, not whether you may collect them. LinkedIn Lead Scraper & Company Website Enrichment returns only publicly available data — content already visible in Google's own index of public LinkedIn pages, nothing gated behind a login. Consult legal counsel if your use case involves bulk storage of personal data.

Frequently asked questions

What LinkedIn lead fields does this Actor return?

The core fields are email, companyDomain, techStack, emailPattern, and title — see the full data fields table above for all 16.

Does this Actor require a LinkedIn account or login?

No. It never authenticates to LinkedIn — it searches Google's public index of linkedin.com pages and reads whatever a logged-out visitor would already see in the search result.

How many leads can I extract in one run?

Up to maxEmailsPerKeyword (max 5000) per keyword in searchKeywords, across as many keywords as you list — a keyword stops early if 5 consecutive result pages return no new email.

What happens if a keyword returns zero visible emails?

The Actor logs each empty page and stops that keyword after 5 consecutive empty pages, then moves on to the next keyword. It does not fail the run — a keyword with no visible emails simply contributes zero rows.

Can I scrape multiple LinkedIn keywords at once?

Yes. searchKeywords accepts an array, and every keyword in it is processed within the same run.

Does this Actor work with Claude, ChatGPT, and other AI agent tools?

It has no dedicated MCP server. It's callable as a standard Apify Actor run through the Apify API or apify-client SDK by any agent framework that can make an HTTP request.

How does this differ from a plain LinkedIn email finder?

Most keyword-based email finders stop at title, snippet, URL, and email. This Actor adds two derived layers on top from data it has already collected at no extra cost: emailPattern/emailPatternConfidence (the observed shape of real addresses per company domain) and, when enrichCompanyWebsites is on, a techStack read of each company's own homepage.

Does this Actor return data in a format LLMs can use directly?

Yes. Typed, normalized JSON with consistent field names across runs — no HTML parsing or selectors needed. Pass it straight to an LLM, index it into a vector store, or feed it to an agent tool.

What happens when LinkedIn or Google changes its layout or anti-bot system?

The SERP parser is structure-based (it scans every link and resolves real destinations rather than matching a hardcoded CSS class name), which is designed to tolerate exactly this kind of markup churn. No specific turnaround time is promised for any given change.

Can I use this Actor without managing proxies or browser infrastructure?

Yes. It runs on Apify Proxy automatically (the GOOGLE_SERP group by default) and makes plain HTTP requests — there's no headless browser or proxy pool for you to configure or maintain.

Which fields work best for AI training data and RAG indexing?

For RAG, index title and description — the highest-information free text. For training or enrichment pipelines, email, companyDomain, and isCorporateEmail return as consistently structured typed primitives across every row.

Scraper NameWhat it extracts
LinkedIn Company Profile ScraperFull company-page details (overview and full views) for a list of LinkedIn company URLs.
LinkedIn Mass Company Profile Finder By Category DirectoryDiscovers LinkedIn company profiles in bulk by keyword, business category, or domain across directory pages.

Your feedback

Found a bug or missing a field? Let us know through the Issues tab on this Actor's Apify Console page — it's the fastest way to reach the maintainer directly.