Lead Scraper & Email Finder - Decision Makers avatar

Lead Scraper & Email Finder - Decision Makers

Under maintenance

Pricing

Pay per usage

Go to Apify Store
Lead Scraper & Email Finder - Decision Makers

Lead Scraper & Email Finder - Decision Makers

Under maintenance

Upload a company list, get verified decision maker emails, phones, LinkedIn, and social profiles. 12-stage pipeline: website discovery, contact extraction, email finder, verification, social enrichment, lead scoring, and Excel export. For email marketing, cold outreach, and B2B prospecting.

Pricing

Pay per usage

Rating

5.0

(3)

Developer

Leadslogix LLC

Leadslogix LLC

Maintained by Community

Actor stats

3

Bookmarked

20

Total users

5

Monthly active users

0.89 hours

Issues response

a day ago

Last modified

Categories

Share

B2B Lead Generation & Sales Intelligence Platform — Extract Verified Decision Maker Emails at Scale

The most powerful B2B lead generation and contact enrichment tool on Apify. Upload a company list and get back verified decision maker emails, phone numbers, LinkedIn profiles, and company intelligence — all from a single 24-stage automated pipeline. No API keys required.

$2 per 1,000 results. First 20 free.

Keywords: B2B lead generation, email finder, email verifier, company scraper, contact extractor, lead scoring, sales intelligence, prospect enrichment, domain validator, website crawler, business data extraction, lead qualification, CRM export, email discovery, SMTP verification, B2B data enrichment, Apollo alternative, ZoomInfo alternative, Clearbit alternative, cold email tool, sales prospecting, LinkedIn scraper, decision maker finder


🔍 What Is This Tool?

This is an enterprise B2B sales intelligence platform that turns a simple list of company names into a complete, verified prospect database — ready for cold email, CRM import, or sales outreach.

Who it's for: Sales teams, SDRs, growth marketers, recruiters, agencies, and anyone who needs verified B2B contact data without paying $100-500/month for Apollo or ZoomInfo.

What it does:

  • 🌐 Discovers company websites from just a company name
  • 👤 Extracts decision makers (names, titles, emails, phones, LinkedIn) using 4 extraction methods
  • 📧 Finds emails through 5 discovery layers (DNS, website crawl, search engines, PDF mining, social)
  • Verifies every email with a 6-check pipeline and assigns B2B send tiers
  • 🧠 Scores and ranks contacts by seniority, authority, and email confidence
  • 📊 Exports CRM-ready data in CSV, Excel, JSON Lines, or via webhook

📋 Data You Get Back

CategoryFields
👤 ContactsFull name, job title, email, phone, LinkedIn URL, seniority level, persona type
🏢 CompaniesWebsite, domain, social profiles (8 platforms), tech stack, employee count, revenue signals
📧 Email IntelVerification status, B2B tier (TIER_1_SEND / TIER_2 / TIER_3 / SKIP), confidence score, auth records
🧠 Lead ScoreCombined priority (0-100), authority score, decision maker flag, persona classification
🔬 Company IntelTech stack fingerprint, SaaS detection, company maturity score, funding stage, SERP signals

⚡ Quick Start

1. Upload your company list

Provide companies as CSV/Excel upload, public URL, or inline JSON:

{
"companies": [
{"company_name": "Stripe", "website": "https://stripe.com"},
{"company_name": "Notion", "website": "https://notion.so"},
{"company_name": "Linear"},
{"company_name": "Vercel"},
{"company_name": "Figma"}
],
"maxResults": 20
}

💡 Tip: You only need the company_name column. The pipeline discovers websites automatically if none is provided.

2. Click Start

Watch progress in real-time: Stage 4/24: Enriching 45/100 companies...

3. Download results

Get your data from the Dataset tab (JSON/CSV/Excel) or KeyValueStore (multi-sheet Excel, JSON Lines).


⬆️ Output

Sample Output (JSON)

{
"company_name": "Acme Inc",
"domain": "acme.com",
"company_website": "https://acme.com",
"contact_name": "John Doe",
"contact_title": "VP Sales",
"contact_email": "john.doe@acme.com",
"contact_phone": "+1-555-0123",
"contact_linkedin": "https://linkedin.com/in/johndoe",
"extraction_method": "team_card",
"is_decision_maker": true,
"persona_type": "Champion",
"seniority": 4,
"lead_score": 85,
"combined_priority": 78,
"priority_band": "HIGH",
"verification_status": "valid",
"b2b_tier": "TIER_1_SEND",
"confidence_score": 92,
"correlation_confidence": 88,
"composite_confidence": 0.91,
"evidence_count": 4,
"data_freshness": "verified",
"auth_score": 85,
"linkedin_company": "https://linkedin.com/company/acme",
"twitter_url": "https://twitter.com/acme",
"tech_stack": "React, Next.js, AWS, Stripe",
"company_maturity_score": 72,
"employee_count_estimate": "50-200",
"has_mx": true,
"has_spf": true,
"has_dmarc": true,
"domain_score": 88,
"enrichment_status": "done"
}

📊 Dataset Views

The actor provides 6 pre-built dataset views in the Apify Console:

ViewWhat It Shows
All ContactsEvery contact with full scoring, verification, and evidence fields
High Priority Decision MakersFiltered to decision makers with correlation confidence and evidence
CompaniesCompany-level data: domain, social profiles, tech stack, maturity
Company IntelligenceTech stack, analytics tools, SaaS signals, maturity score
Funding IntelRevenue estimates, funding stage, employee counts, acquisition signals

📂 Export Formats

FormatLocationBest For
Apify DatasetDataset tabAPI access, JSON/CSV download
CSV (output.csv)KeyValueStoreCRM import (UTF-8 with BOM)
Excel (output.xlsx)KeyValueStore5-sheet workbook with Contacts, Companies, Locations, High_Priority, Audit
JSON Lines (output.jsonl)KeyValueStoreBigQuery, Snowflake, streaming ingestion
WebhookYour endpointReal-time delivery to CRM/Zapier

💰 Why Teams Switch from Apollo, ZoomInfo, and Lusha

Pain PointHow This Solves It
Apollo/ZoomInfo costs $100-500/mo for stale data$2 per 1,000 leads — fresh data scraped in real time, no subscription
Purchased lead lists have 30-50% bounce ratesBuilt-in 6-check email verification with B2B tier classification (TIER_1 = <5% bounce)
Contact databases miss small/mid-size companiesScrapes any company website directly — not limited to a pre-built database
LinkedIn Sales Navigator requires manual prospectingAutomated LinkedIn employee discovery via search engines (no login needed)
Generic web scrapers miss contacts in JavaScript4 extraction methods catch contacts in JSON-LD, JS bundles, and hydration payloads
No way to tell who's a decision makerAI lead scoring with seniority mapping, persona classification, and authority scoring
Exporting data requires manual cleanup14-rule junk removal, dedup, and CRM-ready export in CSV, Excel, JSON Lines
Running the same list twice wastes timeIncremental delta mode skips recently-enriched companies (~70% time savings)

💵 Cost Comparison

Solution1,000 Leads10,000 Leads100,000 Leads
This Actor~$3~$30~$250
Apollo.io$49/mo (limited)$99-399/moCustom pricing
ZoomInfo$250+/mo$500+/mo$1,000+/mo
Lusha$49/mo (limited)$199/moCustom pricing
Hunter.io$49/mo (500 lookups)$199/moCustom pricing

Includes per-event fees + estimated Apify platform charges. All stages, residential proxy.


🎯 Use Cases

Cold Email Outreach & Email Marketing

Upload your target company list and get verified decision maker emails with B2B tier classification. Filter by TIER_1_SEND for safest emails (<5% bounce rate). Import directly into Lemlist, Instantly, Smartlead, Apollo, Woodpecker, or Mailchimp.

Sales Prospecting & Lead List Building

Build targeted B2B lead lists from scratch. Start with just company names — the pipeline discovers websites, extracts leadership teams, finds and verifies emails, and scores every contact. Export the High_Priority sheet for your SDR team.

Account-Based Marketing (ABM)

Enrich your target account list with verified contacts, social profiles, tech stack data, and company intelligence. Decision maker mapping identifies Economic Buyers and Champions at each company.

CRM Data Enrichment

Have a CRM full of companies but missing contact details? Upload your list and the pipeline fills in emails, phones, LinkedIn URLs, social profiles, tech stack, and decision maker details. Incremental mode ensures you only pay for new enrichment.

Competitive Intelligence

Scrape company websites at scale to collect leadership teams, tech stacks, funding signals, and social presence. Company Intelligence shows tech stack fingerprinting, SaaS detection, and company maturity scores.

Recruitment & Talent Sourcing

Find hiring managers and leadership contacts. The pipeline extracts LinkedIn profiles alongside email addresses for combined outreach. Persona classification identifies Technical Evaluators and Champions.


🔑 Key Features

🌐 Multi-Source B2B Data Extraction

  • Website email extractor with 4-method contact extraction (JSON-LD, team cards, heuristic proximity, LinkedIn URLs)
  • 5-layer email discovery: DNS/OSINT, direct crawl, search engines, PDF mining, social platforms
  • LinkedIn employee discovery via multi-query search (CEO, CTO, VP, Director, Manager variations)
  • 8-platform social enrichment: LinkedIn, Twitter/X, Facebook, Instagram, YouTube, GitHub, Crunchbase, Glassdoor
  • SERP intelligence: revenue estimates, funding signals, employee counts, acquisition signals
  • File intelligence: PDF mining for contacts invisible to HTML scrapers
  • Hidden contact extraction: __NEXT_DATA__, __NUXT__, __INITIAL_STATE__, JS hydration payloads

🧠 AI-Powered Lead Scoring

  • Decision maker identification with 5-level seniority mapping (C-Suite → VP/Director → Manager → Staff → Unknown)
  • Persona classification: Economic Buyer, Champion, Technical Evaluator, Influencer
  • Combined priority score (0-100): 60% authority + 40% email confidence
  • Company intelligence profile: tech stack fingerprinting (18+ frameworks), SaaS detection, maturity scoring
  • Quality gate: configurable thresholds filter low-quality contacts before export

✅ Email Verification & Deliverability

  • 6-check pipeline: syntax, MX records, catch-all, disposable, role detection, DKIM/SPF/DMARC
  • B2B send tiers: TIER_1_SEND (safe) → TIER_2_LIKELY_GOOD → TIER_3_REVIEW → SKIP
  • 8-pattern email prediction for contacts missing emails: first.last@, flast@, firstlast@, first_last@, and more
  • Confidence scoring (0-100) with weighted components: SMTP +40, MX +20, auth +15, pattern +10

⚙️ Enterprise Infrastructure

  • Adaptive concurrency: auto-scales 4-32 workers based on success rate and response times
  • HTTP-first hybrid scraping: HTTP → Playwright → Playwright Stealth escalation (browser is last resort)
  • Cross-run cache: eliminates redundant DNS, SERP, LinkedIn, and verification lookups across runs
  • Incremental delta mode: skip companies enriched within freshness window (1-90 days)
  • Executive correlation: cross-source dedup with fuzzy Levenshtein name matching
  • Checkpoint/resume: large runs survive restarts and actor migrations

⬇️ Input

Data Input (choose one)

ParameterTypeDescription
inputFileFile uploadUpload a CSV or Excel file with company names and/or websites
inputUrlStringPublic URL to a CSV or Excel file
companiesJSON arrayInline company list as JSON objects

💡 Auto-detection: The actor recognizes 30+ column name aliases — company_name, organisation, business, exhibitor, firm, url, domain, web_address, and more. Any extra columns are preserved in output.

Settings & Pricing

ParameterTypeDefaultDescription
pipelineVersionStringv10Engine version: v10 (default), v9 (Intelligence OS), v8 (legacy)
maxResultsInteger20Max companies to process. Free: 20/run. Beyond: $2/1,000 results
workersInteger16Initial parallel workers (adaptive: auto-scales 4-32)
maxContactsPerCompanyInteger20Contact cap per company. Decision makers prioritized
maxCrawlPagesPerCompanyInteger25High-value pages crawled per company (5-60)

Incremental & Quality

ParameterTypeDefaultDescription
incrementalModeBooleanfalseSkip recently-enriched companies (~70% time savings)
incrementalFreshnessDaysInteger7Days before cached data is considered stale (1-90)
minLeadScoreInteger0Quality gate: min combined_priority to export
minConfidenceScoreInteger0Quality gate: min confidence score to export
targetConfidenceNumber0.80Goal-seeking enrichment loop confidence target (0.0-1.0)
maxPassesInteger5Max re-enrichment passes per company

Webhook & Export

ParameterTypeDefaultDescription
webhookUrlStringHTTP endpoint to receive results on completion
webhookSendFullResultsBooleanfalseInclude full data in webhook payload
exportJsonLinesBooleanfalseAlso export as .jsonl in KeyValueStore
pushWarehouseTablesBooleanfalsePush warehouse tables to dataset (increases PPE cost)

Pipeline Stage Controls

💡 Tip: Skip Google Boost + Social Enrichment for ~40% faster runs. The pipeline auto-adjusts downstream stages.

ParameterDefaultSkip Effect
skipGoogleBoostfalseSkip 8-step Google Discovery (~30% faster)
skipSocialEnrichmentfalseSkip 8-platform social discovery (~15% faster)
skipLinkedInDiscoveryfalseSkip LinkedIn employee discovery
skipSemanticPageDetectfalseSkip semantic page classification
skipSearchExpansionfalseSkip SERP intelligence (revenue/funding signals)
skipFileIntelligencefalseSkip PDF mining
skipDeepContactExtractfalseSkip 4-method deep re-extraction
skipHiddenContactExtractfalseSkip JS/JSON payload extraction
skipContactIntelligencefalseSkip decision maker mapping
skipCompanyIntelfalseSkip company intelligence profile
skipExecutiveCorrelationfalseSkip cross-source contact dedup
skipEmailDiscoveryfalseSkip 5-layer email discovery
skipEmailPredictionfalseSkip 8-pattern email prediction
skipVerificationfalseSkip 6-check email verification
skipQualityGatefalseSkip quality gate filtering

Parallel Processing

ParameterTypeDefaultDescription
parallelModeBooleantrueEnable parallel company processing
companyConcurrencyInteger5Min companies processed in parallel (floor for adaptive scaling)
crawlStopContactsInteger8Early-exit crawl after N titled contacts found
reuseBrowserContextsBooleantrueReuse browser contexts (rotated every 25 requests)

Proxy

ParameterTypeDescription
proxyConfigurationProxyApify Proxy config. Residential strongly recommended for best results

⚠️ Warning: Running without proxy is not recommended for batches over 20 companies. Datacenter proxies work for most sites but corporate sites may block them.


🏗️ How It Works — 24-Stage Intelligence Pipeline

INPUT: Company list (CSV / Excel / URL / JSON)
├── Stage 1: INGEST → Smart input parsing (30+ column aliases)
├── Stage 2: DISCOVER → Multi-engine website discovery
├── Stage 3: GOOGLE BOOST → 8-step search enhancement
├── Stage 4: ENRICH → Adaptive hybrid crawling (HTTP → Browser → Stealth)
├── Stage 5: GEO → Location intelligence
├── Stage 6: SOCIAL → 8-platform social discovery
├── Stage 7: LINKEDIN → Employee discovery via search engines
├── Stage 8: SEMANTIC PAGES → Intelligent page classification
├── Stage 9: SEARCH + SERP → Revenue, funding, employee signals
├── Stage 10: PDF MINING → File intelligence extraction
├── Stage 11: DEEP EXTRACT → 4-method contact re-extraction
├── Stage 12: HIDDEN EXTRACT → JS payload mining (Next.js, Nuxt, Vue)
├── Stage 13: CONTACT INTEL → Decision maker mapping & persona classification
├── Stage 14: COMPANY INTEL → Tech stack, maturity, SaaS detection
├── Stage 15: EXEC CORRELATION → Cross-source fuzzy dedup
├── Stage 16: EMAIL DISCOVER → 5-layer email discovery
├── Stage 17: EMAIL PREDICT → 8-pattern email prediction
├── Stage 18: VERIFY → 6-check email verification
├── Stage 19: SCORE → Lead scoring engine
├── Stage 20: CLEANUP → 14-rule junk removal
├── Stage 21: QUALITY GATE → Configurable threshold filtering
├── Stage 22: METRICS → Pipeline analytics
├── Stage 23: EXPORT → Multi-format output
└── Stage 24: WEBHOOK → Real-time delivery
OUTPUT: Verified leads → Dataset + CSV + Excel + JSON Lines + Webhook

Stage Details


🛡️ HTTP-First Hybrid Scraping Architecture

Every page is fetched with the cheapest method that works — a browser render is the last resort, not the default:

HTTP (pooled httpx, ~0.3s)Playwright (6s cap) → Playwright Stealth (15s cap)
LayerProxy TierWhen Used
HTTP (pooled keep-alive)DatacenterAlways first; JS-shell detection decides escalation
PlaywrightDatacenterOnly when HTTP returns a JS shell
Playwright StealthResidentialOnly when plain render is blocked; budget-capped per run

Efficiency

FeatureHow It Works
Compressed page storeCrawled HTML zlib-compressed and freed after last stage reads it
Pooled browser contextsOne per (browser, proxy tier), rotated every 25 requests
Resource blockingImages, media, fonts, CSS, and 40+ tracking domains blocked
Early-exit crawl gateStops low-priority pages once enough contacts found
Cross-run SERP cacheSearch queries hit network once per 7 days across all runs
LinkedIn + verification cacheSkip re-discovered profiles and re-verified emails
Crawl reuseEmail discovery reuses stage-4 crawl instead of re-fetching

Anti-Detection

FeatureHow It Works
Fingerprint rotationUA, viewport, locale, timezone, color scheme per context
Stealth hardeningnavigator/webdriver masking on stealth renders
Proxy tieringDatacenter for HTTP; residential reserved for stealth + search
Block detectionHTTP status + soft-block text markers trigger escalation
Adaptive concurrencyAuto-scales 4-32 workers based on success rate
Domain rate limitingPer-domain circuit breaker with recovery timeout

💻 API Examples

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("leadslogix/leadslogix-pipeline").call(run_input={
"inputUrl": "https://example.com/target-companies.csv",
"maxResults": 500,
"workers": 16,
"maxContactsPerCompany": 15,
"minLeadScore": 50,
"webhookUrl": "https://your-crm.com/webhook",
"proxyConfiguration": {"useApifyProxy": True},
})
# Get TIER_1 decision makers
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
if item.get("is_decision_maker") and item.get("b2b_tier") == "TIER_1_SEND":
print(f"{item['company_name']} | {item['contact_name']} | "
f"{item['contact_email']} | Score: {item['combined_priority']}")
# Download Excel from KeyValueStore
kv = client.key_value_store(run["defaultKeyValueStoreId"])
xlsx = kv.get_record("output.xlsx")
with open("leads.xlsx", "wb") as f:
f.write(xlsx["value"])

JavaScript

import { ApifyClient } from "apify-client";
const client = new ApifyClient({ token: "YOUR_API_TOKEN" });
const run = await client.actor("leadslogix/leadslogix-pipeline").call({
companies: [
{ company_name: "Datadog", website: "https://datadoghq.com" },
{ company_name: "Cloudflare", website: "https://cloudflare.com" },
{ company_name: "Twilio", website: "https://twilio.com" },
],
maxResults: 50,
workers: 16,
minLeadScore: 50,
proxyConfiguration: { useApifyProxy: true },
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
const tier1 = items.filter(
(i) => i.is_decision_maker && i.b2b_tier === "TIER_1_SEND"
);
console.log(`Found ${tier1.length} verified decision makers`);
for (const lead of tier1) {
console.log(`${lead.company_name} | ${lead.contact_name} | ${lead.contact_email}`);
}

cURL

curl -X POST "https://api.apify.com/v2/acts/leadslogix~leadslogix-pipeline/runs?token=YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"companies": [
{"company_name": "Figma", "website": "https://figma.com"},
{"company_name": "Canva", "website": "https://canva.com"}
],
"maxResults": 20,
"workers": 16,
"webhookUrl": "https://your-endpoint.com/webhook"
}'

Usage Examples

With Quality Gate & Webhook:

{
"inputUrl": "https://example.com/target-companies.csv",
"maxResults": 500,
"workers": 16,
"minLeadScore": 50,
"minConfidenceScore": 40,
"webhookUrl": "https://hooks.zapier.com/hooks/catch/123456/abcdef/",
"webhookSendFullResults": true,
"exportJsonLines": true,
"proxyConfiguration": {"useApifyProxy": true}
}

Incremental Mode (Repeat Runs):

{
"inputUrl": "https://example.com/same-companies.csv",
"maxResults": 1000,
"incrementalMode": true,
"incrementalFreshnessDays": 14,
"proxyConfiguration": {"useApifyProxy": true}
}

Fast Run (Skip Optional Stages):

{
"companies": [{"company_name": "Acme Corp"}],
"maxResults": 20,
"skipGoogleBoost": true,
"skipSocialEnrichment": true,
"skipLinkedInDiscovery": true,
"skipSearchExpansion": true,
"skipFileIntelligence": true
}

📊 Output Schema

Contact Fields

FieldTypeDescription
contact_nameStringFull name
contact_titleStringJob title
contact_emailStringEmail address
contact_phoneStringDirect phone number
contact_linkedinStringLinkedIn profile URL
extraction_methodStringHow found: jsonld, team_card, heuristic, linkedin, deep_extract, hidden_extract, file_intel, search
is_decision_makerBooleanHolds a leadership position
persona_typeStringEconomic Buyer, Champion, Technical Evaluator, Influencer
seniorityInteger (0-5)Title seniority level
lead_scoreInteger (0-100)Authority score
combined_priorityInteger (0-100)60% authority + 40% verification
priority_bandStringHIGH, MEDIUM, LOW, SKIP
verification_statusStringvalid, risky, invalid, unknown
b2b_tierStringTIER_1_SEND, TIER_2_LIKELY_GOOD, TIER_3_REVIEW, SKIP
confidence_scoreInteger (0-100)Email deliverability confidence
composite_confidenceFloat (0-1)Multi-signal composite confidence
correlation_confidenceInteger (0-100)Cross-source correlation
evidence_countIntegerNumber of independent evidence sources
data_freshnessStringverified, crawled, linkedin_only, predicted_only, search_derived
auth_scoreInteger (0-100)Domain authentication score

Company Fields

FieldTypeDescription
company_nameStringCompany name
company_websiteStringFull URL
domainStringNormalized domain
company_emailsStringSemicolon-separated company emails
company_phonesStringSemicolon-separated phones
linkedin_companyStringLinkedIn company page
twitter_url, facebook_url, instagram_urlStringSocial profiles
youtube_url, github_url, crunchbase_url, glassdoor_urlStringBusiness profiles
company_city, company_countryStringLocation
tech_stackStringDetected technologies
analytics_toolsStringDetected analytics platforms
company_maturity_scoreInteger (0-100)Business maturity index
is_saasBooleanSaaS company detection
employee_count_estimateStringEstimated employee count
estimated_revenue_mNumberRevenue estimate (millions USD)
funding_amount_mNumberFunding amount (millions USD)
funding_stageStringSeed, Series A-F
has_mx, has_spf, has_dkim, has_dmarcBooleanDNS validation
domain_scoreInteger (0-100)Domain trust score
website_quality_scoreInteger (0-100)Website quality index
pages_crawledIntegerPages successfully scraped
enrichment_statusStringdone, cached, failed, error

💵 Pricing

TierActor FeeResults Per RunBest For
Free$0Up to 20Testing the pipeline
Pay-Per-Event$2 per 1,000 resultsUnlimitedProduction lead generation

Apify platform compute charges (CPU, memory, proxy) are billed separately per your Apify subscription.

Cost Estimation

ScenarioCompaniesActor FeeEst. PlatformTotal
Quick test20$0 (free)~$0.05~$0.05
Small batch100$0.16~$0.15~$0.31
Medium batch500$0.96~$0.50~$1.46
Large batch1,000$1.96~$1.00~$2.96
Enterprise10,000$19.96~$10~$30

📈 Performance Benchmarks

MetricTypical Result
Companies per hour100-200 (all stages, residential proxy)
Contacts per company3-15 (varies by company size)
Email discovery rate60-80% of companies
Decision maker rate30-50% of contacts
TIER_1 email rate40-60% of verified emails
Cache hit rate30-70% on repeat runs

Estimated Run Times

CompaniesAll StagesSkip Google+SocialDiscovery Only
203-5 min2-3 min1-2 min
10015-25 min10-15 min5-8 min
5001-2 hours40-70 min20-30 min
1,0003-5 hours2-3 hours45-60 min
10,00024-48 hours16-30 hours6-10 hours

🔗 Integrations

PlatformHow to Connect
Google SheetsAuto-sync via Apify Google Sheets integration
HubSpotImport CRM-ready CSV, or webhook for real-time sync
SalesforceCSV import or connect via Zapier
PipedriveCSV import or webhook
Lemlist / Instantly / SmartleadExport TIER_1 emails as CSV
Apollo / Outreach / SalesLoftImport as prospect sequence
Zapier / MakeConnect to 5,000+ apps via Apify Zapier integration
BigQuery / SnowflakeIngest JSON Lines output
Custom APIFull REST API for scheduling and automation

🔄 Webhook Payload

When the pipeline completes, your webhook receives:

{
"event": "pipeline_complete",
"pipeline_version": "v10.0",
"timestamp": "2026-07-10T12:30:00.000Z",
"summary": {
"total_companies": 100,
"total_contacts": 450,
"high_priority": 85,
"decision_makers": 120,
"emails_found": 380,
"verified_emails": 310
},
"audit": {
"total_companies": 100,
"elapsed_seconds": 1200,
"pipeline_version": "v10.0 (24-stage Intelligence Platform)"
}
}

⏰ Scheduled Lead Generation

Automate recurring prospecting:

  1. Go to the actor page and click Schedules
  2. Create a schedule (e.g., 0 8 * * 1 for every Monday at 8 AM)
  3. Point the input to a URL that updates with new target companies
  4. Enable incrementalMode to skip previously-enriched companies
  5. Set a webhookUrl to receive results in your CRM automatically

🔧 Troubleshooting


❓ FAQ

How is this different from Apollo, ZoomInfo, or Lusha? Those tools maintain a pre-built database of contacts. This tool scrapes company websites and search engines in real time, finding contacts that static databases miss — especially at small/mid-size companies, international firms, and recently-hired executives. It's also 10-50x cheaper per lead.

Do I need API keys? No. This tool uses public web data, DNS records, and search engines. No paid API subscriptions required.

What input formats are supported? CSV, Excel (.xlsx, .xls), and inline JSON. Upload directly, provide a URL, or pass data via the API.

How does incremental mode work? When enabled, the pipeline checks its cross-run cache for each company. If enriched within the freshness window (default 7 days), it's skipped. Saves ~70% on repeat runs.

How does the quality gate work? Set minLeadScore and/or minConfidenceScore to filter contacts. Contacts below thresholds are excluded from export but tracked in metrics. Set both to 0 to export everything.

How accurate is the email verification? TIER_1_SEND emails typically have <5% bounce rate. The pipeline checks MX, SPF, DKIM, DMARC, catch-all, disposable, and role addresses. It does not perform SMTP-level mailbox verification.

Can I use this for a single company? Yes. Use inline JSON with one company and maxResults: 1. The API supports synchronous runs.

Does this work for non-English companies? Yes, but extraction rates are typically 30-50% lower for CJK and Arabic websites due to different page structures and email conventions.

What proxy should I use? Residential proxies give the best results. Datacenter proxies work for most sites but corporate sites may block them. No proxy is not recommended for 20+ companies.

Can I skip stages to save time? Yes. Toggle any of the 15 skip parameters. Skipping Google Boost + Social Enrichment saves ~40% runtime.

What's the maximum batch size? Up to 100,000 with maxResults. For 5,000+ companies, use 8-16 workers with residential proxy and incremental mode.

What pipeline version should I use? Use v10 (default) — it's the fastest and most efficient. v9 has a goal-seeking intelligence graph (more thorough but slower). v8 is the legacy fallback.


⚠️ Limitations

  • Email verification is DNS-based, not SMTP-based. Confirms the domain accepts mail but does not verify individual mailbox existence. For maximum accuracy, run TIER_2 emails through an additional SMTP service.
  • Websites behind login walls or with aggressive anti-bot measures may return limited contacts.
  • Non-English websites (Korean, Chinese, Japanese, Arabic) have lower extraction rates due to different page structures.
  • LinkedIn discovery uses search engines, not direct LinkedIn scraping. Results depend on profile visibility in search indexes.
  • SERP intelligence (revenue, funding) is regex-extracted from search snippets and may not be available for all companies.
  • Social enrichment depends on DuckDuckGo availability. Heavy usage may reduce discovery rates.

📜 Changelog

v10.1 (2026-07-04)

  • CU optimization: compressed HTML, pooled clients, cross-run caches, datacenter-first renders with stealth budget, early-exit crawl gate, per-stage error isolation, cooperative shutdown
  • Enrichment quality: careers/press/privacy pages crawled, LinkedIn company-match verification, cross-stage entity resolution, contacts ranked by composite confidence
  • Output change: warehouse _table rows no longer in dataset by default (opt in with pushWarehouseTables)
  • New inputs: crawlStopContacts, reuseBrowserContexts, pushWarehouseTables

v9.0 (2026-06-05)

  • Intelligence OS: graph-centric, goal-seeking engine with persistent intelligence graph
  • Evidence Engine: multi-source evidence with provenance tracking and contradiction detection
  • Signal Fusion: composite confidence from 5 dimensions (identity, deliverability, authority, relationship, evidence)
  • Crawl Budgeter: per-company budget allocation with ROI-based decisions

v8.1 (2026-06-01)

  • Higher-yield crawl (25 pages default), 5-layer email discovery, better contact retention, safer per-company dedup

v8.0 (2026-05-20)

  • Hybrid 7-engine scraping, smart fallback cascade, site auto-classification, enhanced anti-detection

v7.0 (2026-05-19)

  • Quality gate, webhook dispatcher, incremental delta mode, fuzzy dedup, JSON Lines export

v6.0 (2026-05-19)

  • SERP intelligence, PDF mining, executive correlation, adaptive concurrency, shared cache

v5.0 — v1.0

  • See full changelog in release notes