Email Finder Bulk — Lists of Domains or People to Emails avatar

Email Finder Bulk — Lists of Domains or People to Emails

Pricing

from $4.75 / 1,000 email founds

Go to Apify Store
Email Finder Bulk — Lists of Domains or People to Emails

Email Finder Bulk — Lists of Domains or People to Emails

Bulk email finder, two modes in one run: give a list of **company domains** to harvest published emails, OR a list of **people** (name + domain) to find their work email. Per-email confidence, source, role flag, and cross-list dedup. No API key.

Pricing

from $4.75 / 1,000 email founds

Rating

0.0

(0)

Developer

Vitalii Bondarev

Vitalii Bondarev

Maintained by Community

Actor stats

0

Bookmarked

156

Total users

29

Monthly active users

3 days ago

Last modified

Share

Find emails for a whole list in one run — not 500 separate actor calls.

Email Finder Bulk is a hybrid lead-generation tool that runs two jobs in a single run, so you never have to chain two actors or run one address at a time:

  • Mode B — you have a list of DOMAINS (no names). Give it a list of company websites and it harvests the emails those sites publish — contact pages, mailto: links, page footers — and returns each one with a transparent source and a numeric confidence.
  • Mode A — you have a list of PEOPLE (name + company domain). Give it firstName, lastName, and domain and it builds frequency-ranked email patterns, checks the domain's mail records, and (when the run network allows it) verifies the mailbox over SMTP — returning the best email plus ranked alternatives, each with its own confidence.

Use either mode on its own, or both together in the same run. Everything is merged into one clean, deduplicated table.


Why this actor (what most bulk email tools get wrong)

Most email scrapers either return a single bare guess with no confidence, or they dump emails mixed with noise — tracking strings, image-filename artifacts, placeholder addresses — and leave you to clean it. Email Finder Bulk is built around three things those tools skip:

  1. A confidence score on every email. You always know how much to trust a result, and where it came from (mailto link, contact page, page body, name pattern, or SMTP-verified).
  2. A junk filter. Asset filenames (logo@2x.png), framework/tracking addresses, and placeholder domains (example.com, yourdomain.com) never reach your output.
  3. Cross-list deduplication. Run a list of 300 domains and the same email found twice is collapsed into one row by default — a clean export, not a pile of repeats.

No API key. No login. You pay only for emails actually found.


Input

You can fill in domains, people, or both.

Mode B — domains

A list of company domains or website URLs:

{
"domains": ["stripe.com", "notion.so", "https://www.diverxo.com"],
"maxPerDomain": 3
}

Mode A — people

A list of people, each with a name and company domain:

{
"people": [
{ "firstName": "Patrick", "lastName": "Collison", "domain": "stripe.com" },
{ "firstName": "Tobias", "lastName": "Lutke", "domain": "shopify.com" }
],
"verifySmtp": true
}

Options

FieldWhat it doesDefault
domainsMode B — websites to harvest emails from—
peopleMode A — people to find work emails for—
deduplicateEmailsCollapse the same email into one row; when off, repeats are marked duplicate_flag=truetrue
maxPerDomainCap emails returned per harvested domain (cost + noise control)3
verifySmtpMode A — attempt live SMTP verification; falls back to MX + pattern ranking if port 25 is blockedtrue
maxAlternativesMode A — ranked alternative candidates per person5
maxItemsTotal output cap across both modes (0 = unlimited)0
proxyConfigurationProxy used to reach websites; residential is the default and recommended for reliable accessresidential

Output

One flat row per email, merged across both modes:

{
"domain": "acme.com",
"email": "jane.doe@acme.com",
"first_name": "Jane",
"last_name": "Doe",
"email_source": "smtp_verified",
"confidence": 0.95,
"duplicate_flag": false,
"status": "verified",
"is_role_account": false,
"is_disposable": false,
"mx_found": true,
"verification_method": "smtp",
"on_domain": true,
"alternative_emails": [],
"parse_confidence": 1.0,
"input_mode": "people",
"scraped_at": "2026-06-15T12:00:00Z"
}

Field guide

  • email_source — where the email came from: mailto / contact_page / website_body (Mode B), or name_pattern / smtp_verified (Mode A).
  • confidence — 0–1 score. SMTP-verified mailboxes and mailto links score highest; body-text matches score lower.
  • duplicate_flag — true only when deduplication is off and the email already appeared earlier in the run.
  • is_role_account — true for shared addresses like info@, sales@, support@.
  • on_domain — true when the email's domain matches the site it was found on.
  • status — verified, accept_all, unverified_guess, no_mx (Mode A), or harvested (Mode B).

Pricing example

This actor is pay-per-result: you are charged once per email found, and never for failed lookups, low-confidence guesses, or collapsed duplicates.

Example: a run over 100 domains that finds 60 emails charges for 60 results. Domains that yield nothing cost you nothing in result fees.


Use cases

  • Outreach lists. Turn a list of company domains into a contact list ready for your sending tool.
  • CRM enrichment. Feed a list of accounts (people or domains) and fill the email column with scored, deduplicated results.
  • Lead research. Combine both modes — find decision-makers by name where you know them, and harvest general contacts where you don't.
  • Agency prospecting. Process a target list once, with confidence on every row so you can prioritize the strongest contacts.

FAQ

Do I need names? No. Mode B works from domains alone. Mode A is for when you do know the person's name.

Why do some emails have higher confidence? A mailbox confirmed over SMTP, or an address published in a mailto: link, is more trustworthy than one matched in page text — and the score reflects that.

Does SMTP verification always run? Mode A tries it, but many cloud networks block outbound port 25. When that happens the finder automatically falls back to MX + frequency-ranked patterns and labels the method accordingly — you still get a ranked best guess.

Will I get duplicates? Not by default. Deduplication collapses repeats into one row. Turn it off if you want to see every occurrence with a duplicate_flag.

Is a proxy required? For Mode B, a residential proxy is used by default for reliable access to target websites. You can adjust it in proxyConfiguration.


This actor collects publicly available contact information that companies and individuals choose to publish on their own websites and through public mail records. Use the results in compliance with applicable laws (including GDPR, CAN-SPAM, and local marketing regulations) and the terms of the sites you target. You are responsible for how you contact the people and businesses you find — always honor opt-outs and obtain consent where required.

Usage statistics

This Actor creates a small, content-free summary at the end of each run. It is used only to monitor reliability and improve this Actor. A copy is saved as USAGE_STATS in your own Apify key-value store, so you can see the exact record created for your run.

Set disableUsageStats to true in the input to opt out. Nothing is sent then; your USAGE_STATS record only says that statistics were disabled.

Only these fields are recorded:

  • schema version, Actor name and build number;
  • UTC start and finish hour (not a precise timestamp);
  • run duration, number of results and time to the first result, each as a coarse range;
  • whether the result was empty, the end status, and an error type from a fixed list;
  • memory setting and counts of charged events;
  • names of the input fields you set, never their values;
  • the selected option for input fields that offer a fixed list of choices (for example a sort order).

We do not collect input text, search terms, URLs, domains, usernames, email addresses, names, proxy credentials, tokens, scraped records, output items, raw error messages, stack traces, or hashes of any of those values. Records are kept for no longer than 13 months, used only as aggregated operational statistics, and never sold or shared.

Additional fields (Phase 2)

This Actor also records your Apify user ID, whether Apify marks the account as paying, the size range of list inputs, the selected country when the input offers a fixed list of countries, and one category from a fixed Actor taxonomy. We use these fields only for aggregate reliability, repeat-use and cross-Actor analysis; reports suppress any cell with fewer than five distinct users.

The same disableUsageStats: true input flag turns these fields off too. The user ID is removed after 13 months; we do not export, sell, share, or attempt to re-identify this data.

Run-outcome signals (v2)

To learn whether a run did what it was asked to do, the record also holds a few more coarse ranges and yes/no flags. None of them contains content:

  • the result limit you asked for (a range, when the input has one) and what share of it was delivered;
  • results delivered per input item you listed (a range);
  • output quality as ranges: how fully the result fields were filled, the share of rows that look like errors, the share of duplicate rows, and how many different fields appeared. These are counted in memory while results are saved; no result content is kept;
  • how the run was started (console, API, schedule, webhook, another Actor);
  • how it ended: stopped by you, timed out, reached the requested limit, stopped by the charge limit, and how many times the platform moved the run;
  • if this Actor reports it: how many items to process worked or failed (ranges) and one failure reason from a fixed list;
  • a short code made from the names of the input fields you set, never their values.

Repeat-run fingerprint (v2)

When your Apify user ID is recorded (see above), the record also holds an 8-character one-way code made from your input (proxy settings left out) and this Actor's name. It only lets us see that the same account ran the same input again soon after an unsatisfying run; we never see the input itself. It is stored only in the database, never published, and reports use it in aggregate with the same five-user minimum. It is the one exception to the statement above that no hashes are collected, and disableUsageStats: true turns it off.