# Bulk Email Verifier & Validator (`fanndev/bulk-email-verifier`) Actor

Clean an email list before you send: syntax, MX records, disposable and role mailboxes, typo repair and an SMTP mailbox probe with catch-all detection. Every address gets deliverable, undeliverable, risky or unknown - never a guess dressed up as a verdict.

- **URL**: https://apify.com/fanndev/bulk-email-verifier.md
- **Developed by:** [Faisal Ahdan naufal](https://apify.com/fanndev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.70 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk Email Verifier & Validator

Give it an email list. Get back one verdict per address — **deliverable, undeliverable, risky or unknown** — plus the exact reason, the MX host and the mail server's own words.

No API key to bring, no third-party verification service in the middle — syntax, DNS and the SMTP conversation all happen inside the run.

**On the Apify platform, read ["Port 25, stated plainly"](#port-25-stated-plainly) first:** outbound port 25 is blocked there, so hosted runs deliver everything except the live-mailbox check.

### Four statuses, and why `unknown` exists

Most verifiers report three outcomes and quietly file everything they could not establish under "valid" or "invalid". That is where a bounced campaign comes from. This one keeps a fourth:

| Status | What it means | What to do |
|---|---|---|
| `deliverable` | The mail server accepted the recipient and the domain does **not** accept everything | Send |
| `undeliverable` | Bad syntax, dead domain, no MX, or the server rejected the recipient | Remove |
| `risky` | It may accept and still hurt you: catch-all domain, disposable provider, role mailbox, typo-shaped domain | Your call — the reason says which |
| `unknown` | Greylisting, timeout, or SMTP was not available on this run | Re-run later; do not treat as valid |

`safeToSend` collapses this to one boolean, and it answers `false` for both `risky` and `unknown` on purpose.

### What proves each verdict

```json
{
  "email": "someone@example.com",
  "status": "undeliverable",
  "subStatus": "mailbox_not_found",
  "score": 0,
  "smtpCode": 550,
  "smtpMessage": "550 5.1.1 The email account that you tried to reach does not exist.",
  "mxHost": "aspmx.l.google.com",
  "reason": "The mail server rejected this recipient (SMTP 550)."
}
```

The server's reply is kept verbatim, so a surprising verdict can be audited in ten seconds instead of re-run through a second tool to see whether you believe the first one.

### The checks, cheapest first

1. **Syntax** — RFC-shaped parsing, length limits, the malformed shapes that survive CRM exports.
2. **Normalisation** — lowercased domain, plus-aliases stripped, Gmail dots removed. Two rows with the same `normalizedEmail` are one mailbox, not two sends.
3. **Classification** — role mailbox (`info@`, `sales@`), free provider, disposable provider (by domain *and* by the MX host behind it, which catches rotating alias domains a static list misses).
4. **Typo repair** — `gmial.com` → `gmail.com`. The edit budget scales with domain length, so `acme.com` is never "corrected" to `me.com`.
5. **MX lookup** — cached per domain for the run, with the RFC 5321 implicit-A fallback so small self-hosted domains are not written off. RFC 7505 null MX is read as "sends no mail".
6. **SMTP probe** — `EHLO` → `MAIL FROM` → `RCPT TO`, then `QUIT`. **`DATA` is never sent, so no mail is generated for the address being checked.**
7. **Catch-all detection** — after an acceptance, a random address on the same domain is offered. If that is accepted too, acceptance proved nothing, and the verdict drops to `risky`.

### Port 25, stated plainly

Mailbox-level verification needs outbound TCP 25, and **many hosts block it — including the Apify platform**. Measured on 2026-09-22, both routes time out:

| Route | Result |
|---|---|
| Direct connection from the run | blocked |
| Apify proxy, CONNECT tunnel to port 25 | blocked |

So on Apify this actor delivers checks 1–5 — syntax, normalisation, classification, typo repair and MX — and reports `unknown` rather than inventing a mailbox-level answer. Step 6 works when you run the image somewhere port 25 is open (your own VPS or container host), and the run says which world it was in:

```json
{ "recordType": "RUN_SUMMARY", "smtpRoute": "direct", "smtpPortReachable": false, "smtpEgressDetail": "TimeoutError" }
```

**Read `smtpPortReachable` before trusting any `unknown` in the dataset.** False means the limit was the network, not the address.

Where port 25 is open, two inputs measurably improve how often servers answer honestly: set `heloHostname` and `mailFromAddress` to a domain you actually control, with matching forward and reverse DNS.

### What you get per address

| Field | What it tells you |
|---|---|
| `status` / `subStatus` / `score` / `safeToSend` | The verdict, the machine-readable reason, 0–100 confidence, the one boolean |
| `reason` | One sentence in the terms a sender cares about |
| `normalizedEmail` / `isAlias` | The canonical mailbox behind plus-addressing and Gmail dots |
| `isRoleAccount` / `isFreeProvider` / `isDisposable` | List composition, independent of deliverability |
| `hasTypo` / `suggestedEmail` | The repaired address when a domain looks like a near-miss |
| `domainStatus` / `mxHost` / `mxRecords` / `implicitMx` | What DNS said |
| `smtpCode` / `smtpMessage` / `smtpStage` / `isCatchAll` | What the mail server said, and where the conversation ended |

One `RUN_SUMMARY` record closes every run: status breakdown, reason breakdown, deliverable rate, duplicates removed, domains resolved, and the port 25 verdict.

### Limits worth knowing

- **Yahoo, AOL and some Microsoft tenants accept every recipient at RCPT TO.** Catch-all detection flags this as `risky` rather than reporting a confident `deliverable` that is not.
- **Greylisting returns 450/451.** That is `unknown`, never `undeliverable` — re-run those rows later rather than deleting them.
- **The disposable list is a curated head, not an exhaustive one.** MX-host matching covers much of the long tail; a brand-new throwaway domain can still slip through as `unknown` or `deliverable`.
- **On Apify, `deliverable` is unreachable** — port 25 is blocked, so the best available verdict for a live mailbox is `unknown` with a good MX. Everything that removes addresses (bad syntax, dead domain, no MX, disposable, typo) still works exactly as documented.
- **One SMTP conversation at a time per domain.** A list of 10k addresses on one domain is slower than 10k addresses across 3k domains, on purpose: hammering one server turns honest answers into blanket rejections.
- Verification is for list hygiene on addresses you already hold. It is not an address-discovery tool and there is no mode that finds or guesses addresses.

### Failures are data

A DNS timeout, a closed connection or a refused probe never fails the run. Each becomes a row with its own `subStatus` and `smtpStage`, because "we could not tell, and here is why" is a fact the buyer needs — and a run that dies on row 40,000 of 50,000 wastes everything before it.

### Cost

One DNS lookup per **domain** (not per address) and at most one SMTP session per address. A 50k list of ~3k domains costs 3k lookups, not 50k. Syntax failures and dead domains never reach the network at all.

### Output shape

Dataset rows follow the repo envelope: `_input`, `_source` (`mx+smtp` or `mx-only`), `_scrapedAt`, `recordType` (`EMAIL` / `RUN_SUMMARY`). Optional JSON, NDJSON, CSV and XLSX exports land in the key-value store with the verdict columns first, ready to hand to an ESP.

### Development

```bash
python test_local.py          # syntax + scoring assertions, then MX-only over a sample list
python test_local.py --smtp   # same, plus the mailbox probe and the port 25 verdict
```

See `VERIFICATION_METHOD.md` for the protocol details and the reply codes behind each verdict.

# Actor input Schema

## `emails` (type: `array`):

One address per line. Display-name forms ('Jane Doe <jane@acme.com>') are accepted, duplicates are removed before anything is checked.

## `emailText` (type: `string`):

For pasting straight out of a spreadsheet or CRM export. Separate addresses by newline, comma or semicolon. Merged with the list above.

## `verifySmtp` (type: `boolean`):

Ask the mail server whether the mailbox exists (RCPT TO, never DATA - no mail is sent). Needs outbound port 25, which many networks block; the run tests this once and reports the answer in the summary record instead of failing.

## `detectCatchAll` (type: `boolean`):

After an address is accepted, ask the same server about a random address on that domain. If it accepts that too, the domain takes everything and the acceptance proved nothing - reported as risky rather than deliverable.

## `treatRoleAsRisky` (type: `boolean`):

info@, sales@, support@ and the like are real mailboxes but carry higher complaint rates on cold sends. Turn this off to score them purely on deliverability.

## `includeMxRecords` (type: `boolean`):

Keep the full MX list with preferences on each record. Turn off for a slimmer dataset; mxHost is always present.

## `onlySafeToSend` (type: `boolean`):

Output nothing but confirmed deliverable addresses. The run summary still counts everything, so you keep the real rates.

## `onlyStatuses` (type: `array`):

Keep only the statuses you select. Leave empty to return everything.

## `concurrency` (type: `integer`):

How many addresses to work on at once. SMTP conversations are serialised per domain regardless, so a list of many different domains goes fast while one big domain stays polite.

## `dnsTimeoutSeconds` (type: `integer`):

Per MX lookup. Results are cached per domain for the run.

## `smtpTimeoutSeconds` (type: `integer`):

Per SMTP step. Raise it for slow or greylisting servers; a timeout is reported as unknown, never as undeliverable.

## `heloHostname` (type: `string`):

The hostname announced in EHLO/HELO. Using a domain you actually control, with matching forward and reverse DNS, measurably reduces how often servers stonewall the probe.

## `mailFromAddress` (type: `string`):

The envelope sender used during the probe. No mail is ever sent to or from it; some servers simply refuse to answer an empty sender.

## `emitSummary` (type: `boolean`):

Append one RUN\_SUMMARY record with the status breakdown, deliverable rate, duplicates removed and whether port 25 was reachable from this run.

## `exportFormats` (type: `array`):

Also write the results to the key-value store in these formats. The dataset is always produced regardless.

## `proxyConfiguration` (type: `object`):

Outbound port 25 is blocked on the Apify platform, so mailbox-level checks only work if a proxy allows a CONNECT tunnel to port 25. Enable this to try; the run reports which route it used and falls back to MX-only when the tunnel is refused. DNS is never proxied.

## Actor input object example

```json
{
  "emails": [
    "support@apify.com",
    "this-mailbox-does-not-exist-9x7q@gmail.com",
    "info@shopify.com",
    "someone@mailinator.com",
    "typo@gmial.com",
    "not-an-email"
  ],
  "verifySmtp": true,
  "detectCatchAll": true,
  "treatRoleAsRisky": true,
  "includeMxRecords": true,
  "onlySafeToSend": false,
  "onlyStatuses": [],
  "concurrency": 20,
  "dnsTimeoutSeconds": 8,
  "smtpTimeoutSeconds": 12,
  "heloHostname": "mail.example.com",
  "mailFromAddress": "verify@example.com",
  "emitSummary": true,
  "exportFormats": [],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One verdict per address plus the run summary, including whether port 25 was reachable.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "emails": [
        "support@apify.com",
        "this-mailbox-does-not-exist-9x7q@gmail.com",
        "info@shopify.com",
        "someone@mailinator.com",
        "typo@gmial.com",
        "not-an-email"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("fanndev/bulk-email-verifier").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "emails": [
        "support@apify.com",
        "this-mailbox-does-not-exist-9x7q@gmail.com",
        "info@shopify.com",
        "someone@mailinator.com",
        "typo@gmial.com",
        "not-an-email",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("fanndev/bulk-email-verifier").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "emails": [
    "support@apify.com",
    "this-mailbox-does-not-exist-9x7q@gmail.com",
    "info@shopify.com",
    "someone@mailinator.com",
    "typo@gmial.com",
    "not-an-email"
  ]
}' |
apify call fanndev/bulk-email-verifier --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,fanndev/bulk-email-verifier"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bUsWp6pfRHtqdnmN9/builds/fMawpvuUJg8eSpbJD/openapi.json
