# Website Lead Intelligence — Find the Right Person, Verified (`citrine_venus/website-lead-intelligence-find-the-right-person-verified`) Actor

Website Lead Intelligence turns company websites into send-ready B2B leads — find the right person, verified emails, decision-makers, buying committee, lead scoring and a clear next action for cold email outreach, CRM enrichment and sales prospecting.

- **URL**: https://apify.com/citrine\_venus/website-lead-intelligence-find-the-right-person-verified.md
- **Developed by:** [Data Minds](https://apify.com/citrine_venus) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $100.00 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Website Lead Intelligence: Find the Right Person, Verified Emails & Send-Ready Leads

Turn a list of **company websites** into a scored, **send-ready outreach list** — **verified emails**, named **decision-makers**, a mapped **buying committee**, and one clear next action for every lead.

🔗 **Paste URLs → Start → watch verified contacts stream into Output in real time.**

***

### 📑 Table of Contents

- [What Is Website Lead Intelligence?](#-what-is-website-lead-intelligence)
- [What Data Can You Extract From Company Websites?](#-what-data-can-you-extract-from-company-websites)
- [How It Works](#-how-it-works)
- [How to Use (Apify Console)](#-how-to-use-apify-console)
- [Input Parameters](#-input-parameters)
- [Output Example](#-output-example)
- [Related Actors & Workflows](#-related-actors--workflows)
- [Frequently Asked Questions](#-frequently-asked-questions)
- [Support & Contact](#-support--contact)

***

### 🎁 What Is Website Lead Intelligence?

**Website Lead Intelligence** is an [Apify Actor](https://apify.com/actors) built for **B2B lead generation** and **sales prospecting**: you give it company domains (or business names), and it returns a prioritized **outreach list** with the **right person** to contact — not just a generic `info@` inbox.

It is a purpose-built **website lead intelligence** and **find the right person** tool: every domain becomes one scored lead with **website contact extraction**, **email finder** results, optional **email verification**, **lead scoring**, and a plain-English send decision (`SEND_NOW`, `VERIFY_FIRST`, `ENRICH_MORE`, or `SKIP`). Because it runs on the **Apify platform**, you get scheduling, run monitoring, a REST API for every dataset, webhooks, and export to JSON/CSV/Excel — so your **sales intelligence** pipeline stays automated instead of manual copy-paste from websites.

This is **not** a static contact database and **not** a generic dump of every email on a page. It is a **send-decision engine** for **account-based sales** and **cold email leads**: each website comes back with who to contact, how confident the data is, and what to do next.

***

### 📊 What Data Can You Extract From Company Websites?

For **CRM enrichment**, **lead enrichment**, and building a **target account list**, one run gives you six layers of data per website:

- 🌐 **The company itself** — domain, URL, company type, company meta, addresses, business hours, social links, and whether a contact form was detected.
- 📧 **Contact channels** — personal emails, generic emails, phones, and **verified emails** with status and confidence so your **email verification** step is already done before you send.
- 🙋 **Named contacts & decision-makers** — people found on team/about pages, ranked by seniority, with titles, a best-contact pick, and a suggested first-touch opening line.
- 🤝 **Buying committee & account intelligence** — decision-makers vs influencers vs champions, department coverage, pain signals, “why now” timing, and account readiness for **account-based sales**.
- 🚦 **Lead scoring & send decision** — a 0–100 lead score, A/B/C decision tier, bounce-risk bucket, and one next action so your **send-ready leads** are obvious at a glance.
- 🛟 **Recovery when a site is thin or blocked** — failed or sparse domains still return a full record with failure context and a recovery plan, instead of vanishing from your **B2B lead scraper** results.

Use it for **sales prospecting**, building a multi-threaded **buying committee**, turning domains into **cold email leads**, watching a **target account list** for changes, or feeding structured **sales intelligence** into Instantly, Smartlead, HubSpot, Salesforce, Notion, or Slack.

***

### 🔧 How It Works

1. **Add your targets.** Paste **company websites**, or discover sites from **business names** / niche **footer phrases**.
2. **Pick a Goal (optional).** Quick outreach, high deliverability, or max coverage — the Actor sets depth and verification dials for you.
3. **Each website is read for public signals.** Home, about, team, contact, and related pages are checked for people, emails, phones, and company cues.
4. **Contacts are organized.** Named people are ranked; personal vs generic inboxes are separated; optional gap-fill suggests likely emails for people without a published address.
5. **Emails can be verified.** Turn on verification to confirm deliverability and flag catch-all domains before you put addresses into outreach.
6. **Every lead is scored and decided.** You get a lead score, decision tier, buying-committee map, and one clear next action.
7. **Results stream live.** Rows appear in the Output dataset as soon as each site is ready — with optional CRM webhook, ready-to-import CSVs, Notion, and Slack delivery.
8. **Optional watchlist mode.** Re-run the same list later to spot new hires, new personal emails, or tier upgrades on your **target account list**.

***

### 🚀 How to Use (Apify Console)

1. Open this Actor on [Apify Console](https://console.apify.com) → **Actors**.
2. Paste your **website URLs** (or business names / footer phrases).
3. Optionally choose a **Goal** so presets and confidence mode configure themselves.
4. Click **Start** and open the **Output** tab — **send-ready leads** appear as each site finishes.
5. Switch dataset views (Overview, Contacts, Scoring, Decisioning, Account Intelligence, Monitoring).
6. Export JSON / CSV / Excel, or enable Instantly / Smartlead / HubSpot / Salesforce / Notion / Slack delivery for the next run.
7. Call the same run via the **Apify API** when you want **website lead intelligence** inside your own stack.

#### 🤖 Quick API example

```bash
curl -X POST "https://api.apify.com/v2/acts/<ACTOR_ID>/run-sync-get-dataset-items" \
     -H "Authorization: Bearer $APIFY_TOKEN" \
     -H "Content-Type: application/json" \
     -d '{"urls": ["https://apify.com", "https://stripe.com", "https://notion.so"]}'
```

***

### 🧩 Input Parameters

Only **Website URLs** is truly required — everything else has a sensible default. Fields below are grouped the same way they appear in the input form.

#### 🚀 Quick Start

| Field | Type | Description | Default |
|---|---|---|---|
| `goal` | string (`quick-outreach` | `high-deliverability` | `max-coverage`) | Outcome dial that auto-sets depth + verification for **lead generation** runs. | unset |
| `preset` | string (`auto` | `fast` | `balanced` | `maximum`) | How deep to scan each site. `auto` follows your Goal (or balanced). | `auto` |
| `confidenceMode` | string (`safe` | `balanced` | `aggressive`) | Risk appetite for which emails keep — use `safe` for careful cold outreach. | follows Goal / `balanced` |

#### 🔗 What to Process

| Field | Type | Description | Default |
|---|---|---|---|
| `urls` | array of strings | Company websites to process (bulk paste, up to 500). Core input for this **company website scraper**. | `[]` |
| `knownNames` | array of strings | Business names resolved to official websites via search when you do not have URLs. | `[]` |
| `footerPhrases` | array of strings | Distinctive niche phrases (exact-match search) to discover similar businesses at once. | `[]` |
| `nameSuffix` | string | Appended to name searches (e.g. `"plumbing"`, `"law firm"`) to disambiguate. | `""` |
| `discoveryCountry` | string (`US` | `UK` | `CA` | `AU` | `EU`) | Localizes discovery search / TLD filtering. | `US` |
| `maxResultsPerQuery` | integer | Search results kept per footer-phrase query. | `50` |
| `maxDiscoveredDomains` | integer | Cap on domains discovered from names/phrases (does not cap direct URLs). | `1000` |
| `excludeDomains` | array of strings | Permanent skip list for domains you never want in results. | `[]` |

#### 🕸️ Crawl Depth & Coverage

| Field | Type | Description | Default |
|---|---|---|---|
| `maxPagesPerDomain` | integer (1–20) | Max pages read per website. Blank = use preset. | preset |
| `deepScan` | boolean / unset | Also check legal / imprint / support-style pages (helpful for EU sites). | preset |
| `includeNames` | boolean | Extract named contacts — the core **find the right person** job. | `true` |
| `includeSocials` | boolean | Collect LinkedIn, X, Facebook, Instagram, YouTube and other social links. | `true` |
| `sitemapDiscovery` | boolean | Use sitemap hints to find About / Team / Contact pages with unusual paths. | `true` |
| `customProbePaths` | array of strings | Extra paths (e.g. `/meet-the-team`) checked on every site. | `[]` |
| `respectRobotsTxt` | boolean | Honor each site’s robots.txt rules. | `true` |
| `crawlDelayMs` | integer | Politeness delay between requests to the same domain. | `0` |

#### 📧 Email Verification & Gap-Filling

| Field | Type | Description | Default |
|---|---|---|---|
| `verifyEmails` | boolean / unset | Run **email verification** on found addresses and flag catch-all domains. | preset |
| `emailVerificationProvider` | string (`mx_smtp` | `zerobounce` | `none`) | Built-in check by default, or ZeroBounce if you supply its API key. | `mx_smtp` |
| `fillMissingEmails` | boolean / unset | For named people without a published email, suggest + verify likely pattern addresses. | preset |
| `maxEmailPatternsToVerify` | integer | Cap on generated address checks per site. | `3` |
| `enableProFallback` | boolean / unset | Re-render JS-heavy sites in a real browser when static HTML is empty. | preset |

#### 📈 Change Monitoring

| Field | Type | Description | Default |
|---|---|---|---|
| `compareToPrevRun` | boolean | Compare each domain to the previous run on the same watchlist. | `false` |
| `monitorStateKey` | string | Name for the watchlist history (auto-derived if blank). | auto |

#### 🗂️ Output Shaping

| Field | Type | Description | Default |
|---|---|---|---|
| `autoFilter` | string (`none` | `send-now-only` | `safe-only` | `max-leads`) | Keep only the **send-ready leads** (or safe / non-skip) you care about. | `none` |
| `outputProfile` | string (`full` | `standard` | `minimal`) | How many fields to keep per row. | `full` |
| `minLeadScore` | integer (0–100) | Drop sites below this **lead scoring** threshold. | none |
| `requirePersonalEmail` | boolean / unset | Keep only leads with a personal (non-generic) email. | confidence mode |
| `companyTypes` | array | Restrict to industries (SaaS, agency, legal, healthcare, etc.). | `[]` |
| `outputFormat` | string (`json` | `ndjson` | `csv`) | Format of the combined file in the key-value store. | `json` |

#### 🚚 Delivery & Integrations

| Field | Type | Description | Default |
|---|---|---|---|
| `exportFormats` | array | Ready-to-import CSVs: Instantly, Smartlead, Lemlist, Apollo, HubSpot, Salesforce, Mailshake, Outreach, Woodpecker, generic. | `[]` |
| `crmWebhookUrl` | string | POST each lead live for **CRM enrichment** as it is produced. | none |
| `crmFormat` | string (`generic-json` | `hubspot` | `salesforce`) | Shape of the webhook body. | `generic-json` |
| `crmOnlyTierA` | boolean | Only push top-tier (A) leads to the CRM webhook. | `false` |
| `notionConnector` | boolean | Post a digest / per-lead pages to Notion (`NOTION_API_KEY` required). | `false` |
| `notionDatabaseId` | string | Target Notion database ID. | none |
| `notionArchiveProfile` | string (`summary` | `per-lead`) | One summary page vs one page per lead. | `summary` |
| `slackConnector` | boolean | Post a run digest to Slack (`SLACK_BOT_TOKEN` required). | `false` |
| `slackChannel` | string | Channel for the digest (e.g. `#sales-leads`). | none |
| `deliverTopN` | integer | How many top-ranked leads appear in Notion/Slack digests. | `10` |

#### ⚡ Performance & Network

| Field | Type | Description | Default |
|---|---|---|---|
| `proxyConfiguration` | object | Optional. Starts direct by default; can escalate automatically if a site pushes back. | no proxy |
| `concurrency` | integer (1–50) | How many websites process in parallel. | `10` |
| `retryAttempts` | integer | Retries per request before giving up on that attempt path. | `3` |
| `requestTimeoutSeconds` | integer | Per-request timeout. | `20` |
| `searchProvider` | string | Provider used only for name/phrase discovery. | `duckduckgo` |
| `dryRun` | boolean | Validate input without outbound network calls. | `false` |

***

### 📦 Output Example

Each processed website becomes **one JSON object** in your dataset as soon as it is scored. Field names below match the Actor output.

```json
{
  "recordType": "lead",
  "domain": "apify.com",
  "url": "https://apify.com",
  "companyType": "saas",
  "emails": ["hello@apify.com"],
  "personalEmails": [],
  "genericEmails": ["hello@apify.com"],
  "verifiedEmails": [
    { "email": "hello@apify.com", "status": "valid", "confidence": 90 }
  ],
  "generatedEmails": [],
  "phones": [],
  "contacts": [],
  "bestContact": null,
  "leadScore": 30,
  "decision": {
    "tier": "C",
    "reason": "Generic emails only — use Email Pattern Finder for personal addresses"
  },
  "sendDecision": {
    "action": "ENRICH_MORE",
    "riskLevel": "high",
    "reasons": ["Only generic inboxes found on the public site"]
  },
  "isContactable": true,
  "isSendable": false,
  "companyIntelligence": {
    "maturityBand": "startup",
    "trustScore": 40
  },
  "buyingCommittee": {
    "decisionMakers": [],
    "size": 0
  },
  "pipelineValue": {
    "rankInBatch": 1,
    "relativeScore": 1.0
  },
  "plainEnglishSummary": "No contact data found on apify.com's public website. The recovery plan suggests next-best tools.",
  "firstTouch": null
}
```

A final ranked list is also stored in the run’s key-value store as **`OUTPUT`**, with a **`RUN_SUMMARY`** and any CSV exports you requested.

#### 🗂️ Dataset views (Output tab)

| View | What you see |
|---|---|
| 🔎 **Overview** | Domain, tier, lead score, best contact, next action — at a glance |
| 👤 **Contacts & Company** | Emails, phones, named contacts, socials, addresses, company info |
| 📊 **Scoring & Confidence** | Lead score, data quality, confidence, catch-all flags |
| 🚦 **Send Decision & Recovery** | Next action, bounce risk, recovery plan if the site was thin/blocked |
| 🧠 **Account Intelligence** | Buying committee, pain signals, timing, pipeline priority |
| 📈 **Change Monitoring** | What changed vs the previous watchlist run |

#### 🧩 Nested field shapes

| Object / array | Useful fields | Notes |
|---|---|---|
| `verifiedEmails[]` | `email`, `status`, `confidence` | Outcome of **email verification** for **verified emails** |
| `contacts[]` / `bestContact` | `name`, `title`, `email`, `score` | Named people for **decision-makers** outreach |
| `decision` | `tier`, `reason` | A / B / C priority for your **outreach list** |
| `sendDecision` | `action`, `riskLevel`, `reasons` | `SEND_NOW` / `VERIFY_FIRST` / `ENRICH_MORE` / `SKIP` |
| `buyingCommittee` | `decisionMakers`, `influencers`, `size` | Multi-threaded **account-based sales** map |
| `companyIntelligence` | maturity, trust, intent-style signals | Context for **sales intelligence** messaging |
| `recoveryPlan` | `method`, `nextBestTool` | What to try when **website contact extraction** was incomplete |

***

### 🔗 Related Actors & Workflows

If you need **website lead intelligence** — **find the right person**, **verified emails**, **lead scoring**, and a send decision from **company websites** — this Actor is the one you want.

It pairs well with:

| Need | How this Actor helps |
|---|---|
| **B2B lead scraper** for domains you already have | Paste `urls` and export **send-ready leads** |
| **Email finder** + **email verification** in one pass | Enable `verifyEmails` / `fillMissingEmails` |
| **CRM enrichment** into HubSpot / Salesforce / Instantly | Use `exportFormats` or `crmWebhookUrl` |
| Ongoing **target account list** monitoring | Turn on `compareToPrevRun` + an Apify Schedule |
| Niche discovery without a URL list | Use `knownNames` or `footerPhrases` as a **company website scraper** seed |

If you only need a raw dump of every email string on a page with no ranking, verification, or next action, a simpler **contact scraper** may be enough. This Actor is for teams that need **sales prospecting** quality: the right person, confidence, and a clear send decision.

***

### ❓ Frequently Asked Questions

#### 🎯 What is Website Lead Intelligence used for?

**Website Lead Intelligence** turns company domains into a prioritized **outreach list** for **B2B lead generation**, **cold email leads**, and **account-based sales**. Each site returns contacts, optional **verified emails**, **lead scoring**, and one next action — so **sales prospecting** starts from decisions, not raw scrapes.

#### 🙋 How does “find the right person” differ from a normal contact scraper?

A typical **contact scraper** lists every email it sees. This Actor ranks **named contacts**, separates personal vs generic inboxes, maps a **buying committee**, and picks a best contact with a first-touch suggestion — so **find the right person** is the product, not an afterthought.

#### 📧 Does it provide email verification for verified emails?

Yes. When verification is on, addresses are checked and returned under `verifiedEmails` with status and confidence, and catch-all domains are flagged. That keeps **email verification** inside the same **website lead intelligence** run as discovery.

#### 🧩 Can it work as an email finder when a person has no published address?

Optionally. With fill-missing enabled, the Actor can suggest likely pattern-based addresses for named people and (when verification is on) check them before you treat them as **send-ready leads**. Generated suggestions are labeled separately from published emails.

#### 🏢 Can I run lead generation without a URL list?

Yes. Use **business names** or niche **footer phrases** to discover websites, then process them like a normal **company website scraper** / **B2B lead scraper** run. Direct `urls` remain the fastest path when you already have a **target account list**.

#### 🤝 Does it map a buying committee for account-based sales?

Yes. **Account intelligence** groups people into decision-maker / influencer-style roles and reports department coverage so **account-based sales** teams can multi-thread instead of emailing a single generic inbox.

#### 📊 How does lead scoring help sales intelligence?

Each website gets a **lead score**, data-quality signals, and an A/B/C tier. Combined with `sendDecision`, that **sales intelligence** layer tells you who is worth sending today vs who needs more enrichment.

#### 🚚 Can I push results into my CRM or outreach stack?

Yes. Export ready-made CSVs for Instantly, Smartlead, Lemlist, Apollo, HubSpot, Salesforce, and more — or POST live rows to a **CRM enrichment** webhook. Notion and Slack digests are available for team visibility. Everything stays on the **Apify Actor** platform with scheduling and API access.

#### 📈 Can I monitor a target account list for changes?

Yes. Enable change monitoring with a watchlist key and schedule the Actor. Repeat runs flag new hires, new **personal emails**, tier upgrades, and related momentum on the same **target account list**.

#### 🚫 Why does a lead show thin data or a recovery plan?

Some sites publish almost no public contacts, rely heavily on client-side rendering, or block automated visits. Those domains still appear as full records with `failureType` / `recoveryPlan` guidance — nothing silently disappears from your **website contact extraction** results.

#### 🛡️ Is scraping company websites for leads legal?

This Actor works with **publicly available** information on company websites — the same pages a visitor can open without logging in. You remain responsible for how you use the data, for respecting each site’s terms, and for handling personal data under GDPR/CCPA and similar rules. This is not legal advice.

#### 🔌 Is there an API for website lead intelligence?

Yes — via **Apify**. Start runs and fetch dataset items with the Apify REST API, webhooks, or integrations. You get API-style access to **lead enrichment** without building and hosting your own crawler.

#### ⏱️ Can I schedule ongoing sales prospecting runs?

Yes. Apify Schedules let you re-run hourly, daily, or on any cron — ideal for refreshing **cold email leads**, re-verifying contacts, or tracking changes on a fixed domain list.

#### 🧹 How do I keep only send-ready leads?

Use `autoFilter: "send-now-only"` or `"safe-only"`, optionally raise `minLeadScore`, and/or require a personal email. That filters the dataset down to the **send-ready leads** your team will actually touch.

***

### 🙋 Support & Contact

Found a bug, need a new field, or want a **custom solution** built around this **Website Lead Intelligence** / **find the right person** workflow for your stack?

Reach out via the **Issues** tab on this Actor’s Apify Store page, or email **<hello.dataminds@gmail.com>** for custom builds, integrations, and feature requests.

Feedback is welcome — this Actor is actively maintained for teams that need reliable **B2B lead generation**, **verified emails**, and **sales intelligence** from real company websites.

# Actor input Schema

## `goal` (type: `string`):

Pick the outcome you care about and every technical dial below is set for you automatically. Leave unset to configure 'Preset' and 'Confidence mode' yourself.

## `preset` (type: `string`):

Controls crawl depth, email verification, missing-email fill-in, and the real-browser fallback all at once. 'Auto' (recommended) lets your Goal choose this for you.

## `confidenceMode` (type: `string`):

How conservative to be about which emails make it into your list. Independent from the preset above. Use 'Safe' for cold outreach where bounces hurt sender reputation. Leave unset to use your Goal's default (or 'balanced' if no Goal set).

## `urls` (type: `array`):

✨ The main input. One company website per line — paste as many as you like (up to 500). Every domain is deduplicated automatically and produces exactly one lead record. Leave empty if you'd rather discover websites from business names or a footer phrase below.

## `knownNames` (type: `array`):

Don't have a URL list? Paste business names and each one is resolved to its official website via web search before scraping. Combine with 'Name search suffix' below to disambiguate common names.

## `footerPhrases` (type: `array`):

Distinctive phrases that identify a niche (e.g. "we buy land in any state"). Each phrase is searched as an exact match and every organic result becomes a candidate website — great for finding an entire niche of similar businesses at once.

## `nameSuffix` (type: `string`):

Appended to every business-name search to disambiguate common names (e.g. "plumbing", "real estate", "law firm"). Leave blank to search names exactly as given.

## `discoveryCountry` (type: `string`):

Used both to localize the web search and to filter discovered domains by country-code TLD. Only relevant when using business names or footer phrases above — ignored for direct URLs.

## `maxResultsPerQuery` (type: `integer`):

How many search results to collect for each footer-phrase query.

## `maxDiscoveredDomains` (type: `integer`):

Hard cap on how many websites can be discovered from names/phrases in one run (doesn't limit websites pasted directly into 'Website URLs').

## `excludeDomains` (type: `array`):

A permanent blocklist — these domains are skipped even if they show up via URL input or discovery.

## `maxPagesPerDomain` (type: `integer`):

Ceiling on how many pages of one website are read (home, contact, about, team, and similar). Leave blank to use the selected preset's value. Higher finds more but takes longer per site.

## `deepScan` (type: `boolean`):

Also probes legal/imprint/impressum/support pages — useful for EU sites, which often publish contact details only there. Leave the toggle in its middle ('use preset') state to let the preset decide.

## `includeNames` (type: `boolean`):

Look for real people (name + job title) on team/about pages — this Actor's core job. Turn off only if you just want emails/phones and don't care who they belong to.

## `includeSocials` (type: `boolean`):

Collect LinkedIn, X/Twitter, Facebook, Instagram, YouTube and 8 other platform links found on each site.

## `sitemapDiscovery` (type: `boolean`):

Parses /sitemap.xml to find About/Team/Contact-shaped pages the standard probe list wouldn't guess — helps on sites with an unusual structure.

## `customProbePaths` (type: `array`):

Additional page paths (e.g. "/meet-the-team") checked on every website, on top of the built-in probe list.

## `respectRobotsTxt` (type: `boolean`):

Honor each site's robots.txt disallow rules. Recommended to leave on — it's both good etiquette and reduces the odds of being blocked.

## `crawlDelayMs` (type: `integer`):

Minimum wait between two requests to the *same* domain. Raise this for smaller sites you want to be extra polite to.

## `verifyEmails` (type: `boolean`):

Runs a real deliverability check (MX + mailbox probe, or a paid provider below) on every email found, and flags catch-all domains. Leave the toggle in its middle ('use preset') state to let the preset decide.

## `emailVerificationProvider` (type: `string`):

Free MX + mailbox probe by default. Switch to a paid provider for higher accuracy if you supply its API key as an environment variable.

## `fillMissingEmails` (type: `boolean`):

When a real person was found with no published email, generate their likely address from the company's detected naming pattern and verify it. Leave the toggle in its middle ('use preset') state to let the preset decide.

## `maxEmailPatternsToVerify` (type: `integer`):

Caps how many pattern-generated address guesses get a verification check on a single website, to keep verification calls bounded.

## `enableProFallback` (type: `boolean`):

When a React/Vue/Next.js/Angular site's static HTML comes back empty, re-renders that page in a real headless browser to recover contacts a plain HTTP fetch would miss. Leave the toggle in its middle ('use preset') state to let the preset decide.

## `compareToPrevRun` (type: `boolean`):

Compares every domain against its snapshot from a previous run with the same watchlist key and adds change flags (new hire, new personal email, tier upgrade, and more) plus an account-momentum score. Persists across runs automatically.

## `monitorStateKey` (type: `string`):

Name this watchlist so scheduled re-runs compare against the right history. Leave blank to auto-derive one from the input domain list.

## `autoFilter` (type: `string`):

Optionally trim the output down to only the leads that matter for your use case.

## `outputProfile` (type: `string`):

'Full' keeps every field (recommended for automation). 'Standard' drops agent-diagnostic fields. 'Minimal' keeps just the fields you need to decide who to contact.

## `minLeadScore` (type: `integer`):

Drop any website scoring below this. Leave blank for no minimum.

## `requirePersonalEmail` (type: `boolean`):

Only keep leads with a real person's email address, not just info@/hello@-style inboxes. Leave the toggle in its middle ('use preset') state to default to the confidence mode's setting.

## `companyTypes` (type: `array`):

Only keep websites classified as one of the selected industries. Leave empty to keep every industry.

## `outputFormat` (type: `string`):

Format of the single combined results file saved to this run's key-value store (in addition to the live dataset table).

## `exportFormats` (type: `array`):

Generate a ready-to-import CSV for one or more outreach tools, saved to this run's key-value store.

## `crmWebhookUrl` (type: `string`):

Every lead is POSTed here as it's produced (2 retries; auto-disabled after 5 consecutive failures so a dead endpoint can't stall the run).

## `crmFormat` (type: `string`):

Shape of the JSON body sent to the CRM webhook above.

## `crmOnlyTierA` (type: `boolean`):

Skip the CRM webhook for anything below the top decision tier.

## `notionConnector` (type: `boolean`):

Posts a run digest (or one page per lead) to a Notion database. Requires the NOTION\_API\_KEY environment variable and a Database ID below.

## `notionDatabaseId` (type: `string`):

Target database for Notion delivery.

## `notionArchiveProfile` (type: `string`):

How much detail to post to Notion — a single digest page for the whole run, or one page per lead.

## `slackConnector` (type: `boolean`):

Posts a run digest to Slack. Requires the SLACK\_BOT\_TOKEN environment variable.

## `slackChannel` (type: `string`):

Channel to post the digest to (e.g. "#sales-leads").

## `deliverTopN` (type: `integer`):

How many of your best leads (by pipeline rank) are included in the Notion/Slack digest.

## `proxyConfiguration` (type: `object`):

🚦 Default is NO proxy — every request goes straight to the target website. If a site ever rejects or blocks a request, the run automatically escalates step by step: 🚫 no proxy → 🖥️ datacenter proxy → 🏠 residential proxy (retrying at least 3× on residential) — then stays on residential for every remaining request for the rest of the run. Every escalation is logged clearly. Force a starting tier yourself here if you already know a target needs it.

## `concurrency` (type: `integer`):

How many domains are crawled at the same time.

## `retryAttempts` (type: `integer`):

How many times one failed/blocked request is retried, with exponential backoff, before giving up (per proxy tier).

## `retryBackoffBaseSeconds` (type: `number`):

Starting delay for exponential-jitter backoff between retries.

## `requestTimeoutSeconds` (type: `integer`):

How long to wait for a single response before treating it as failed.

## `searchProvider` (type: `string`):

Used only when resolving business names or footer phrases into websites.

## `jsRenderProvider` (type: `string`):

Engine used by the real-browser fallback above.

## `userAgent` (type: `string`):

Override the default User-Agent string sent with every request. Leave blank for the built-in default.

## `dryRun` (type: `boolean`):

Skips every outbound request — useful for validating your input configuration without spending run time or budget.

## Actor input object example

```json
{
  "preset": "auto",
  "urls": [
    "https://example.com",
    "www.another-company.com"
  ],
  "nameSuffix": "",
  "discoveryCountry": "US",
  "maxResultsPerQuery": 50,
  "maxDiscoveredDomains": 1000,
  "includeNames": true,
  "includeSocials": true,
  "sitemapDiscovery": true,
  "respectRobotsTxt": true,
  "crawlDelayMs": 0,
  "emailVerificationProvider": "mx_smtp",
  "maxEmailPatternsToVerify": 3,
  "compareToPrevRun": false,
  "autoFilter": "none",
  "outputProfile": "full",
  "outputFormat": "json",
  "crmFormat": "generic-json",
  "crmOnlyTierA": false,
  "notionConnector": false,
  "notionArchiveProfile": "summary",
  "slackConnector": false,
  "deliverTopN": 10,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "concurrency": 10,
  "retryAttempts": 3,
  "retryBackoffBaseSeconds": 0.5,
  "requestTimeoutSeconds": 20,
  "searchProvider": "duckduckgo",
  "jsRenderProvider": "playwright",
  "dryRun": false
}
```

# Actor output Schema

## `results` (type: `string`):

Every lead record, every field, in the order each website finished processing.

## `overview` (type: `string`):

The at-a-glance columns: domain, decision tier, lead score, best contact, next action.

## `contacts` (type: `string`):

Emails, phones, named contacts, social links, addresses and company metadata.

## `scoring` (type: `string`):

Lead score, data quality, confidence breakdown, coverage, decision tier and domain purity.

## `decisioning` (type: `string`):

Send decision, send plan, recovery plan, bounce-risk bucket and the risk/negative signals behind them.

## `accountIntelligence` (type: `string`):

Buying committee, department coverage, company maturity/trust signals, pain points, timing ('why now'), and pipeline priority.

## `monitoring` (type: `string`):

First/last-seen timestamps and what changed since the previous run on the same watchlist (only populated when 'Monitor for changes' is enabled).

## `outputJson` (type: `string`):

Every lead from this run in one combined JSON array, with pipeline-priority ranking finalized across the whole batch and every output filter already applied — the authoritative, ready-to-download result file.

## `runSummary` (type: `string`):

Counts and timing for this run: websites requested/processed, leads returned, send-ready leads, verified emails, proxy escalations.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://apify.com",
        "https://stripe.com",
        "https://notion.so"
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("citrine_venus/website-lead-intelligence-find-the-right-person-verified").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://apify.com",
        "https://stripe.com",
        "https://notion.so",
    ],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("citrine_venus/website-lead-intelligence-find-the-right-person-verified").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://apify.com",
    "https://stripe.com",
    "https://notion.so"
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call citrine_venus/website-lead-intelligence-find-the-right-person-verified --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,citrine_venus/website-lead-intelligence-find-the-right-person-verified"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/LblUrYnZyRsq6NitA/builds/YCGU8zoy3xJ6yMXmL/openapi.json
