# Contact Info Extractor (`inestimable_zoysia/contact-info-extractor`) Actor

Extract emails, phone numbers, and social profiles from any list of websites or another scraper's dataset. Emails validated via DNS MX checks. Chain it after Google Maps Scraper to turn business listings into an outreach-ready lead list with a best-email pick per site.

- **URL**: https://apify.com/inestimable\_zoysia/contact-info-extractor.md
- **Developed by:** [Uncle Glooby](https://apify.com/inestimable_zoysia) (community)
- **Categories:** Lead generation, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Contact Info Extractor & Email Validator

Turn a list of websites into a clean contact list — emails, phone numbers, and social profiles — with every email domain validated so your outreach doesn't bounce on day one.

Built to slot directly after any scraper: Google Maps, business directories, or your own URL list. Those tools give you websites; this Actor turns websites into people you can actually reach.

### What it does

- **Two input modes:**
  - Paste a list of website URLs, or
  - Point it at a **dataset from another Actor run** (e.g. Google Maps Scraper output) and tell it which field holds the URL. Your original row is carried through in the output, so nothing gets disconnected.
- **Smart, shallow crawling** — fetches each homepage; if no email is found there, it follows up to a few contact-like pages (contact, about, impressum) per site. Capped so runs stay fast and costs stay predictable.
- **Extracts:**
  - **Emails** — from `mailto:` links and page text, with junk filtered out (image filenames, no-reply addresses, tracker noise)
  - **Phone numbers** — from `tel:` links and visible text, international formats supported
  - **Social profiles** — LinkedIn, Facebook, Instagram, X/Twitter, YouTube, TikTok
- **Email validation via DNS MX checks** — verifies each email's domain actually runs mail servers. Catches dead domains and typos before they hurt your sender reputation. (This is DNS-level validation; it confirms the domain can receive mail, not that a specific mailbox exists — full mailbox verification requires SMTP infrastructure that no ethical bulk tool can honestly promise.)
- **`bestEmail` per site** — prefers validated, personal addresses (jane@ beats info@) so you can use the output directly in outreach tools.

### Input example (dataset mode)

```json
{
    "inputDatasetId": "abc123",
    "urlField": "website",
    "maxPagesPerSite": 3,
    "validateEmails": true
}
```

### Output example (one row per site)

```json
{
    "website": "https://acmeplumbing.com",
    "domain": "acmeplumbing.com",
    "status": "ok",
    "pagesChecked": 2,
    "emails": [
        { "email": "jane@acmeplumbing.com", "mxValid": true },
        { "email": "info@acmeplumbing.com", "mxValid": true }
    ],
    "bestEmail": "jane@acmeplumbing.com",
    "phones": ["+1 (555) 123-4567"],
    "socials": {
        "facebook": "https://facebook.com/acmeplumbing",
        "instagram": "https://instagram.com/acmeplumbing"
    },
    "source": { "...your original dataset row..." : "..." }
}
```

A `SUMMARY` record in the key-value store reports totals: sites reached, sites with emails found, sites with phones.

### The lead-gen workflow this fits into

1. Run a scraper (Google Maps, directories, job boards...) → business listings with websites
2. **This Actor** → emails, phones, socials per business, validated
3. A dedupe/merge Actor → one clean list
4. Export to your CRM or outreach tool

### Tips

- Sites that block automated traffic come back with `status: "unreachable"` instead of failing the run. For large runs, enabling **Apify Proxy** in the input significantly improves reach.
- Generic addresses (info@, contact@) are still collected — `bestEmail` just prefers personal ones when available.
- Keep `maxPagesPerSite` at 3 unless you have a specific reason; going higher rarely finds more and slows the run.

### Pricing

Pay-per-result: billed per site in the output. Every site you feed in produces exactly one output row — reached or not — so costs are fully predictable from your input size.

# Actor input Schema

## `startUrls` (type: `array`):

List of website URLs to extract contact info from. Use this OR the dataset input below.

## `inputDatasetId` (type: `string`):

Optional. A dataset from another Actor run (e.g. Google Maps Scraper). Each row's URL field will be processed. If set, this is used instead of the URL list above.

## `urlField` (type: `string`):

Which field of the input dataset contains the website URL. Supports dot notation. Common values: 'website', 'url', 'domain'.

## `maxPagesPerSite` (type: `integer`):

The homepage is always fetched; if contact info is missing, up to this many additional contact-like pages (contact, about, impressum...) are checked. Higher = more thorough, slower.

## `validateEmails` (type: `boolean`):

Check that each email's domain has mail servers (MX records). Catches dead domains and typos. Adds a small amount of run time.

## `proxyConfiguration` (type: `object`):

Use Apify Proxy or custom proxies. Recommended for larger runs — some websites block datacenter IPs.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://apify.com"
    }
  ],
  "urlField": "website",
  "maxPagesPerSite": 3,
  "validateEmails": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://apify.com"
        }
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("inestimable_zoysia/contact-info-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://apify.com" }],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("inestimable_zoysia/contact-info-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://apify.com"
    }
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call inestimable_zoysia/contact-info-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,inestimable_zoysia/contact-info-extractor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/UTZSUpFy8KiHpLfbo/builds/XPtI49Pr9TXKP3CPk/openapi.json
