# Yellow Pages Email Scraper - Domain Filter & Deduplication (`insightflow/yellow-pages-email-scraper-powerful-and-intelligent`) Actor

📒 Yellow Pages Email Scraper — turn Yellow Pages searches into a clean local business email list. Filter by keyword, location and domain, merge alias duplicates and decode hidden addresses. 🏪 For SMB lead generation & cold email.

- **URL**: https://apify.com/insightflow/yellow-pages-email-scraper-powerful-and-intelligent.md
- **Developed by:** [InsightFlow](https://apify.com/insightflow) (community)
- **Categories:**
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.99 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

### Yellow Pages Email Scraper 🔍

**Yellow Pages Email Scraper** helps you extract emails from Yellow Pages listings without the manual hunting. It’s built for marketers, recruiters, sales pros, and researchers who need faster Yellow Pages lead generation, cleaner directory contact scraping, and a practical business directory email finder for public web data. 🚀

### 🌟 Key Features of Yellow Pages Email Scraper

| Feature | Benefit |
|---|---|
| ✅ **Targeted Keyword Search** | Use your keywords to find relevant Yellow Pages business listings and focused prospecting opportunities. |
| ✅ **Location Filtering** | Narrow your directory lead extraction by location for more relevant local business prospecting. |
| ✅ **Custom Email Domain Filters** | Filter for specific email domains such as `@gmail.com` or `@yahoo.com` to refine contact information extraction. |
| ✅ **Bulk Lead Collection** | Collect up to your chosen email limit for scalable business leads scraping. |
| ✅ **Real-Time Saving** | Results are saved as they’re found, helping prevent data loss during longer runs. |
| ✅ **Built-In Proxy Support** | Uses built-in proxy support for reliable scraping and better stability on public web data. |
| ✅ **Resume-Friendly Runs** | Progress is stored so you can continue large directory contact scraping jobs more smoothly. |

### 📥 Input — Yellow Pages Email Scraper Parameters

```json
{
  "keywords": [
    "manager",
    "founder"
  ],
  "location": "",
  "customDomains": [
    "@gmail.com",
    "@yahoo.com"
  ],
  "maxEmails": 20
}
```

| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| `keywords` | Array | Yes | `["manager", "founder"]` | A list of keywords or queries to search for relevant Yellow Pages listings. |
| `location` | String | No | `""` | Location used to filter search results. Leave it empty to search more broadly. |
| `customDomains` | Array | No | `["@gmail.com", "@yahoo.com"]` | Email domains to focus on when collecting contact details, such as common personal or business domains. |
| `maxEmails` | Integer | No | `20` | Maximum number of emails to collect before the actor stops. |

### 📤 Output — What Yellow Pages Email Scraper Returns

The actor saves each result as a JSON record in your Apify dataset. 📦

```json
{
  "keyword": "founder",
  "title": "Green Leaf Plumbing",
  "description": "Licensed plumbing services for residential and commercial clients in Austin, TX. Emergency repairs, water heater installation, and drain cleaning.",
  "url": "https://www.yellowpages.com/austin-tx/mip/green-leaf-plumbing-12345678",
  "email": "contact@greenleafplumbing.com"
}
```

| Field | Label | Format | Description |
|---|---|---|---|
| `keyword` | Keyword | text | The keyword that matched this result. |
| `title` | Title | text | The title of the Yellow Pages listing. |
| `description` | Description | text | The listing description or summary shown in the dataset. |
| `url` | Url | link | The URL of the Yellow Pages listing. |
| `email` | Email | text | The email address extracted from publicly available sources. |

### 💻 How to Use Yellow Pages Email Scraper — Step-by-Step

1. **Open the Actor** — Find **Yellow Pages Email Scraper** in the Apify Store and open the actor page.
2. **Enter Keywords** — Add the job titles, industries, or business terms you want to target.
3. **Set Location** — Optionally narrow results to a city, region, or country for better local business email harvesting.
4. **Choose Email Domains** — Add one or more domains to focus your email discovery software results.
5. **Set Max Emails** — Control how many contacts you want in one run.
6. **Run the Actor** — Start the actor and monitor the live logs while it collects leads.
7. **Export Results** — Download the dataset and use the leads in your CRM, spreadsheet, or outreach workflow.

*No coding required. Results are ready in minutes.*

### 💡 Best Use Cases for Yellow Pages Email Scraper

- 🎯 **B2B Lead Generation** — Build targeted prospect email database lists from Yellow Pages for outreach campaigns.
- 📣 **Email Marketing** — Gather public contact details for newsletters and drip sequences.
- 🤝 **Local Business Prospecting** — Find nearby businesses with publicly available contact information.
- 🔬 **Market Research** — Study business directory scraping tool results across industries and locations.
- 📊 **CRM Enrichment** — Add discovered emails to existing contact records for cleaner outreach.

### Disclaimer

This actor only accesses publicly available data on Yellow Pages. It does not scrape private profiles, authenticated content, or password-protected pages. Users are responsible for complying with applicable laws, platform terms, and anti-spam regulations. Use this tool for legitimate lead generation, research, and marketing purposes only. For data-removal requests, contact 📧 <insightflowofficial@gmail.com>.

### 🆘 Support & Feedback

Have a question or found an issue with the **Yellow Pages Email Scraper**? We’re here to help.

- 🐞 **Bug Reports:** Reach out with a clear description of the issue and what happened.
- ✨ **Custom Solutions & Feature Requests:** Share your idea if you need tailored data extraction support.
- 📧 **Email:** <insightflowofficial@gmail.com>

Your feedback helps improve this yellow pages email scraper for everyone.

### MX Lookup

Every address is checked at the DNS level: the actor resolves the mail domain's
MX records and reports what it found.

| Field | Meaning |
| --- | --- |
| `mxFound` | `true` when the domain publishes at least one mail server |
| `mxHost` | The lowest-preference (primary) mail server |
| `mxRecords` | Every MX record found, in preference order |
| `mxProvider` | Who runs the mail: Google Workspace, Microsoft 365, Zoho, Proton, ... |
| `mxStatus` | `found`, `no_records`, `no_such_domain`, `timeout`, `error`, or `skipped` |

**Inputs**

- **MX Lookup** - turn the check on or off (default: on).
- **Only keep contacts whose domain has an MX record** - drop unreachable domains.
  Only a definite negative (`no_records` / `no_such_domain`) drops a contact; a
  timeout or resolver error is treated as unknown and the lead is kept.
- **MX lookup timeout (seconds)** - per-domain DNS budget, 1-15s.

This is a domain-level check. It confirms the domain can receive mail; it does
not verify that an individual mailbox exists.

# Actor input Schema

## `keywords` (type: `array`):

A list of keywords or queries to search for.

## `location` (type: `string`):

Location to filter search results.

## `customDomains` (type: `array`):

List of custom email domains

## `maxEmails` (type: `integer`):

Maximum number of emails to collect. The scraper will stop once this limit is reached. Setting a higher limit allows for more potential results but doesn't guarantee reaching that number. This helps save costs by controlling scraping time.

## `minLeadScore` (type: `integer`):

Drop any contact whose Lead Score (0-100) falls below this. Leave at 0 for no filter.

## `mxLookup` (type: `boolean`):

Resolve each address's domain to its mail servers. Adds the resolved MX hosts and the mail provider (Google Workspace, Microsoft 365, ...) to every result, and feeds the Lead Score. DNS-level only - it confirms the domain can receive mail, not that the individual mailbox exists.

## `requireMxRecord` (type: `boolean`):

Drop any contact whose domain provably accepts no mail. A lookup that times out or errors is treated as unknown and kept, so a DNS hiccup never silently deletes good leads.

## `mxTimeoutSecs` (type: `integer`):

How long to wait for each DNS answer. Clamped to 1-15 seconds.

## Actor input object example

```json
{
  "keywords": [
    "manager",
    "founder"
  ],
  "location": "",
  "customDomains": [
    "@gmail.com",
    "@yahoo.com"
  ],
  "maxEmails": 20,
  "minLeadScore": 0,
  "mxLookup": true,
  "requireMxRecord": false,
  "mxTimeoutSecs": 3
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "manager",
        "founder"
    ],
    "location": "",
    "customDomains": [
        "@gmail.com",
        "@yahoo.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("insightflow/yellow-pages-email-scraper-powerful-and-intelligent").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": [
        "manager",
        "founder",
    ],
    "location": "",
    "customDomains": [
        "@gmail.com",
        "@yahoo.com",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("insightflow/yellow-pages-email-scraper-powerful-and-intelligent").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "manager",
    "founder"
  ],
  "location": "",
  "customDomains": [
    "@gmail.com",
    "@yahoo.com"
  ]
}' |
apify call insightflow/yellow-pages-email-scraper-powerful-and-intelligent --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,insightflow/yellow-pages-email-scraper-powerful-and-intelligent"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bNz4LQkg274iAxD1F/builds/gxbZVn0eUvgvYuaK4/openapi.json
