# Email Extractor: Website Emails, Phones & Socials, Verified (`pnda/email-extractor`) Actor

Email extractor for websites and domains: extract emails from website lists, plus phone numbers, social media links and address. A website email scraper where every email is SMTP-verified. $2.50 per 1,000 domains with contacts.

- **URL**: https://apify.com/pnda/email-extractor.md
- **Developed by:** [PNDA](https://apify.com/pnda) (community)
- **Categories:** Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 domain with contacts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Email Extractor: extract emails, phone numbers and social media links from any website, SMTP-verified

**Email Extractor** turns a list of websites or domains into a clean contact list: **emails, phone numbers, social media links (Facebook, Instagram, LinkedIn, X/Twitter, TikTok, YouTube, Pinterest), company name and address**, with **one row per domain**. Every email is **verified with a live SMTP mailbox check** before you get it, so you know which addresses are deliverable.

It is a domain email extractor, a website email scraper and a contact details scraper in one Actor, and you pay **per domain where contacts are found**, never per page crawled.

### Why this email extractor

| | Email Extractor (this Actor) | Typical website email scraper |
|---|---|---|
| Real SMTP verification of every email | **Included** (deliverable / risky / undeliverable / unknown, catch-all, role account) | None, or format / MX check only, or a paid add-on |
| Billing | **Per domain with contacts.** A domain with nothing found is free | Per page crawled (a 20-page site = 20 charges) or per row |
| Output | **One row per domain**, emails ranked (site's own domain first) | One row per page, to deduplicate yourself |
| Junk filtering | Image names, tracking IDs, placeholders, theme-vendor and plugin addresses removed | Often left in |
| Obfuscated emails | `name [at] domain.com`, `&#64;`, `%40`, Cloudflare email protection decoded | Partial |
| Phone format | E.164 (`+14155550123`) when the country is known | Raw text |

### What it extracts

- **Emails**: `mailto:` links, plain text, obfuscated forms (`[at]`, `(at)`, `&#64;`, `%40`) and Cloudflare-protected emails. Emails on the website's own domain come first, then free mailboxes (gmail...), then others.
- **SMTP verification** of each email: `deliverable`, `risky`, `undeliverable` or `unknown`, with a **catch-all** flag and a **role account** flag (info@, contact@, sales@...).
- **Phone numbers**: `tel:` links and numbers in the page text, normalised to **E.164** when the country is known (address on the site, country domain, or your default country).
- **Social media links**: Facebook, Instagram, LinkedIn (company or profile), X / Twitter, TikTok, YouTube, Pinterest. Share buttons and the website builder's own accounts are ignored.
- **schema.org data**: Organization / LocalBusiness email, telephone, address and `sameAs` profiles.
- **Company name** and **postal address** when the site publishes them.

### How it works

For each domain, the Actor reads **up to 5 pages**: the homepage, then the **contact, legal (impressum, mentions légales, legal notice) and about pages** found through the homepage links or the site's `sitemap.xml`. It **stops early** as soon as an email, a phone and a social link are found, so most sites take 1 to 3 pages.

Pages are fetched with plain HTTP first. Two optional fallbacks, billed **only when actually used**:

- **JavaScript-only sites** (React, Vue, some Wix pages) are opened in a real browser (`useBrowser`, on by default).
- **Protected sites** behind an anti-bot wall such as Cloudflare's "Just a moment..." page are unblocked (`unblockProtectedSites`, off by default). Without it, those domains come back with status `blocked`, free.

### Input

```json
{
  "urls": ["acme.com", "https://www.example-bakery.co.uk/contact", "contact@another-shop.fr"],
  "verifyEmails": true,
  "useBrowser": true,
  "unblockProtectedSites": false,
  "maxPagesPerDomain": 5,
  "defaultCountry": "US"
}
```

- `urls`: 1 to 10,000 website URLs, bare domains or even emails. `www.acme.com`, `acme.com/contact` and `https://acme.com` are the same domain: processed and billed once.
- `defaultCountry`: optional two-letter code used to format local phone numbers.

### Output (one row per domain)

```json
{
  "status": "ok",
  "input": "acme.com",
  "domain": "acme.com",
  "url": "https://www.acme.com/",
  "companyName": "Acme Inc.",
  "primaryEmail": "jane@acme.com",
  "emails": ["jane@acme.com", "info@acme.com"],
  "emailDetails": [
    { "email": "jane@acme.com", "ownDomain": true, "verdict": "deliverable", "catchAll": false, "roleAccount": false, "reason": "Mailbox exists", "charged": true },
    { "email": "info@acme.com", "ownDomain": true, "verdict": "risky", "catchAll": true, "roleAccount": true, "reason": "Catch-all domain: the server accepts any address", "charged": false }
  ],
  "phones": ["+14155550123"],
  "facebook": "https://www.facebook.com/acme",
  "instagram": "https://www.instagram.com/acme",
  "linkedin": "https://www.linkedin.com/company/acme",
  "twitter": "https://x.com/acme",
  "tiktok": null,
  "youtube": "https://www.youtube.com/@acme",
  "pinterest": null,
  "socialLinks": ["https://www.facebook.com/acme", "..."],
  "address": "1 Main St, 94105, San Francisco, US",
  "pagesRead": ["https://www.acme.com/", "https://www.acme.com/contact"],
  "fetchedWith": ["http"],
  "fromCache": false,
  "charged": { "domain": true, "verifiedEmails": 1, "browserPages": 0, "unlockerPages": 0 }
}
```

`status` is `ok` (contacts found), `no-contacts`, `blocked` (anti-bot wall), `unreachable` (DNS, TLS, timeout) or `error`. Only `ok` rows are charged. Undeliverable emails stay in `emailDetails` but are removed from `emails`. Export as JSON, CSV or Excel, or plug the dataset into Make, Zapier, n8n or Google Sheets.

### Pricing (pay per event)

| Event | Price | When |
|---|---|---|
| Domain with contacts | **$2.50 / 1,000 domains** | A domain where at least one email, phone or social link is found. **A domain with nothing found is free.** |
| Verified email | **$1.20 / 1,000 emails** | Only final SMTP answers: deliverable or undeliverable. Risky, catch-all and unknown are free. Up to 5 emails verified per domain. |
| Browser page | $2 / 1,000 pages | Only pages that needed a real browser (JavaScript-only sites). |
| Unblocked page | $6 / 1,000 pages | Only when "Unblock protected sites" is on and a page was actually unblocked. |
| Actor start | $0.005 per run | |

From October 24, 2026: a **catch-all email published on the website** (risky, `catchAll: true`) is charged **$0.60 / 1,000**, half the price of a verified email; other risky and unknown emails stay free.

Example: 1,000 small-business websites, 70% with contacts and 1.5 verified emails each on average: 700 × $0.0025 + 1,050 × $0.0012 = **about $3.00**, with SMTP verification included.

The Actor checks your **maximum cost per run** before every paid step: when it is reached, the remaining domains are listed in one `budget-too-low` row and are not charged.

**Works with a free Apify account too.** Free-plan accounts can use every feature at the same per-result prices, paid from their monthly Apify credit. The only limits are your credit and your maximum cost per run.

Results are kept for 30 days: running the same domain again within that window returns the stored result instantly, at the same price.

### Use cases

- **Lead generation**: turn a list of company websites into verified B2B emails and phone numbers.
- **Enrich a CRM**: add social media links, phones and verified emails to accounts you already have.
- **Local business outreach**: combine with a Google Maps search to get the websites, then extract and verify their contacts.
- **Influencer and partner research**: collect Instagram, TikTok and YouTube profiles of brands from their websites.
- **Clean your data**: find which published addresses actually accept mail before you send anything.

### FAQ

**Does it work with a free Apify account?** Yes, it works with a free Apify account too, at the same prices: you pay only for domains where contacts are found, from your monthly Apify credit.

**How is this different from an email extractor online or a Chrome extension?** It runs in the cloud on thousands of domains at once, reads the contact and legal pages for you, and verifies every email with a real SMTP check.

**Does it scrape the whole website?** No. It reads at most 5 pages per domain (homepage, contact, legal, about), the pages where businesses publish their contact details. That keeps runs fast and cheap.

**Why is an email "risky" or "unknown"?** Catch-all servers accept every address, so a single mailbox cannot be confirmed (`risky`, `catchAll: true`). Some mail servers do not answer clearly or rate-limit checks (`unknown`). Neither is charged.

**Can I use it from n8n, Make or Zapier?** Yes, through the Apify integrations or the API: send a list of URLs, read the dataset.

**Is it a Hunter.io alternative?** For domain-level contacts published on a website, yes: you get emails, phones and social media links with verification included, billed per domain.

### Related actors

- [Email Finder](https://apify.com/pnda/email-finder): No email on the site? Find a named person's work email from first name, last name and company domain.
- [Email Verifier](https://apify.com/pnda/email-verifier): Already have an email list from elsewhere? Bulk-verify it before sending, $0.95 per 1,000.
- [Google Maps API & Scraper](https://apify.com/pnda/google-maps-scraper): No list of websites yet? Pull local businesses by keyword and city, with their websites, then extract their contacts here.
- [Shopify & WooCommerce Stores Database](https://apify.com/pnda/shopify-woocommerce-stores): Want e-commerce leads? Filter Shopify and WooCommerce stores by country, category and apps, then extract their contacts here.
- [LinkedIn Scraper & People Search](https://apify.com/pnda/linkedin-scraper): Need the decision-maker, not the generic inbox? Find people by job title and company.

# Actor input Schema

## `urls` (type: `array`):

Website URLs or bare domains, one per line (1 to 10,000). Duplicates and www variants are merged: one row per domain.

## `verifyEmails` (type: `boolean`):

Check every email found with a live SMTP mailbox test: deliverable, risky, undeliverable or unknown, plus catch-all and role-account flags. Only deliverable and undeliverable answers are charged ($1.20 / 1,000).

## `useBrowser` (type: `boolean`):

When a page is an empty JavaScript shell (React, Vue, Wix...), open it in a real browser. Charged only for pages actually rendered ($2 / 1,000 pages).

## `unblockProtectedSites` (type: `boolean`):

Get through anti-bot walls (Cloudflare 'Just a moment...', 403). Charged only for pages actually unblocked ($6 / 1,000 pages). Off by default: protected sites come back with status 'blocked', free.

## `maxPagesPerDomain` (type: `integer`):

Homepage + contact, legal (impressum, mentions légales) and about pages. The crawl stops early once an email, a phone and a social link are found. Billing is per domain, never per page.

## `defaultCountry` (type: `string`):

Two-letter country code (US, GB, FR...) used to turn local phone numbers into E.164 when the site gives no country (address, country domain).

## Actor input object example

```json
{
  "urls": [
    "sparktoro.com",
    "https://www.nivon.com",
    "manufactum.de"
  ],
  "verifyEmails": true,
  "useBrowser": true,
  "unblockProtectedSites": false,
  "maxPagesPerDomain": 5
}
```

# Actor output Schema

## `dataset` (type: `string`):

All domains (JSON, CSV, Excel).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "sparktoro.com",
        "https://www.nivon.com",
        "manufactum.de"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pnda/email-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "sparktoro.com",
        "https://www.nivon.com",
        "manufactum.de",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("pnda/email-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "sparktoro.com",
    "https://www.nivon.com",
    "manufactum.de"
  ]
}' |
apify call pnda/email-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pnda/email-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XaU5KGYFpVhYMQaYF/builds/UPzOM4mDPmI6egqxl/openapi.json
