# Contact Details Scraper (`tactful_anvil/contact-details-scraper`) Actor

Extract emails, phone numbers and 13 social networks from any website list at a flat $2.40/1,000 sites. MX-verified emails, Cloudflare-obfuscation decoding, JSON-LD company data, smart contact-page crawling. Failed sites are never charged.

- **URL**: https://apify.com/tactful\_anvil/contact-details-scraper.md
- **Developed by:** [Mr Zack](https://apify.com/tactful_anvil) (community)
- **Categories:** Developer tools, Automation, Lead generation
- **Stats:** 4 total users, 3 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.40 / 1,000 website results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Contact Details Scraper — Email, Phone, Socials + MX Check

Turn any list of websites into a clean, **verified** contact database. Paste domains, get back emails (MX-checked so they don't bounce), phone numbers, 13 social networks, and structured company data — at a **flat $2.40 per 1,000 websites**, everything included.

### Why this one?

Most contact scrapers charge per *page*, per *contact*, or per *social profile* — you can't predict your bill, and they hand you raw strings that bounce in your outreach tool. This Actor is different:

- **Flat, predictable pricing** — one price per website, no matter how many pages we crawl or contacts we find. $2.40/1,000 websites vs. $3–$10.50/1,000 effective at the popular alternatives.
- **Failed websites are never charged.** Site down, DNS dead, non-HTML? You pay $0 for it. It's listed in the run `SUMMARY` so you know exactly what happened.
- **MX-verified emails, built in.** Every found email's domain is checked for valid MX records (DNS-over-HTTPS). Dead domains are flagged *before* they wreck your sender reputation. Competitors charge extra for this or skip it entirely.
- **Finds emails others miss**: `mailto:` links, plain text, **Cloudflare-protected emails** (`data-cfemail` — decoded, not the useless `[email protected]` placeholder), obfuscated patterns like `name [at] company [dot] com`, and JSON-LD structured data.
- **Smart crawling, not blind crawling.** We fetch the homepage, then spend your page budget on the pages that actually contain contacts: `/contact`, `/impressum`, `/about`, `/team`, `/support` — scored and prioritized, in 5 languages of URL patterns.
- **A `bestEmail` you can use immediately.** Ranked by quality: MX-valid corporate personal inbox → `info@` → free provider; disposable addresses are demoted. Plus a 0–100 `confidence` score per site.

### Who uses this

- **Sales & lead-gen teams** enriching lead lists (from Google Maps scrapers, directories, CSV exports) with emails that actually deliver.
- **Agencies** building outreach lists for clients — flat pricing means you can quote your client a fixed cost.
- **Recruiters** finding company contact channels at scale.
- **Data teams** appending contact + social columns to any company dataset.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `websites` | array | — | Domains (`acme.com`) or URLs (`https://acme.com`), one per line |
| `maxPagesPerWebsite` | integer | 8 | Crawl budget per site (1–25). More pages never costs more. |
| `verifyMx` | boolean | `true` | MX-check every email domain via DNS-over-HTTPS |
| `deobfuscateEmails` | boolean | `true` | Decode Cloudflare-protected + `[at]`/`[dot]` emails |
| `includeSocials` | boolean | `true` | Extract 13 social networks |
| `sameDomainOnly` | boolean | `true` | Stay on the exact input domain (off = allow subdomains) |
| `maxConcurrency` | integer | 10 | Websites processed in parallel |
| `navigationTimeoutSecs` | integer | 20 | Per-request timeout |

### Output (one row per website)

| Field | Description |
|---|---|
| `bestEmail` | Highest-quality email — MX-valid corporate personal inbox ranks first |
| `emails[]` | All emails with extraction `method` (mailto / text / cloudflare-decoded / deobfuscated / jsonld), `mxStatus` (valid / none / unknown), `isRole`, `isFree`, `isDisposable`, and source pages |
| `primaryPhone` | Best phone — `tel:` links and JSON-LD outrank text matches |
| `phones[]` | All phones with source + confidence |
| `socials` | LinkedIn, Facebook, Instagram, X, YouTube, TikTok, GitHub, Pinterest, Threads, Telegram, WhatsApp, Discord, Medium — profile links only, share/intent links filtered out |
| `organization` | Company name, type and address from JSON-LD structured data |
| `contactPageFound` | Whether a dedicated contact/impressum page was located |
| `confidence` | 0–100 contactability score |
| `pagesCrawled`, `crawledUrls` | Exactly what was fetched — full transparency |

A `SUMMARY` record in the key-value store gives you run-level stats: email hit rate, MX-valid %, and the list of failed (uncharged) websites.

### How to schedule this Actor

Contact data decays — people change jobs, domains die, companies rebrand. Re-verify your list on a schedule:

1. Open the Actor → **Schedules** tab → **Create new schedule**.
2. Pick a cadence — weekly or monthly works well for re-verifying an outreach list.
3. Set your `websites` input (or point your workflow at a dataset from a previous run).
4. Add a webhook or connect the run to Zapier/Make/n8n to push fresh contacts into your CRM automatically.

Because pricing is flat per website and dead sites are free, re-running a 1,000-domain list costs a predictable $2.41 — no surprises.

### Pricing

| Event | Price |
|---|---|
| Actor start | $0.01 |
| Website result (all contacts + MX verification included) | $0.0024 |

**Example:** 1,000 websites = **$2.41 total.** If 50 of them are unreachable, you pay for 950.

### FAQ

**Does it work on any website?** Any publicly accessible HTML website. It's HTTP-based (no browser), so heavily JavaScript-rendered SPAs may yield fewer contacts — but footers, contact pages and JSON-LD are almost always in the HTML.

**Is MX verification the same as SMTP verification?** MX checking confirms the domain can receive mail (kills hard-bounce domains). It does not probe individual mailboxes, which keeps it fast, cheap and included in the flat price.

**Where do the emails come from?** Only from the target website's own public pages. No third-party databases, no guessing patterns.

**GDPR note:** you're extracting publicly published business contact data; make sure your outreach complies with the laws of your jurisdiction.

### Related Actors

- **[Bulk Email Verifier & Validator](https://apify.com/tactful_anvil/bulk-email-verifier)** — pipe the emails this Actor finds straight into MX/SMTP verification for $0.48/1,000. Scrape → verify → outreach, all pay-per-event.

# Actor input Schema

## `websites` (type: `array`):

List of websites to extract contacts from. Accepts bare domains (acme.com) or full URLs (https://acme.com/about). One line per website.

## `maxPagesPerWebsite` (type: `integer`):

Crawl budget per site. The scraper always fetches the homepage first, then the highest-value pages (contact, impressum, about, team, support). 8 is enough for most sites. Pricing is flat per WEBSITE — more pages never costs you more.

## `verifyMx` (type: `boolean`):

Check every found email's domain for valid MX records via DNS-over-HTTPS. Filters dead domains before they bounce in your outreach tool. Included in the price.

## `deobfuscateEmails` (type: `boolean`):

Also decode Cloudflare-protected emails (data-cfemail) and patterns like "name \[at] domain \[dot] com" that plain scrapers miss.

## `includeSocials` (type: `boolean`):

Extract LinkedIn, Facebook, Instagram, X/Twitter, YouTube, TikTok, GitHub, Pinterest, Threads, Telegram, WhatsApp, Discord and Medium profile links.

## `sameDomainOnly` (type: `boolean`):

Only crawl pages on the exact same domain as the input website. Turn off to also allow subdomains (e.g. blog.acme.com).

## `maxConcurrency` (type: `integer`):

How many websites to process in parallel.

## `navigationTimeoutSecs` (type: `integer`):

Timeout per page request. Increase for slow sites.

## Actor input object example

```json
{
  "websites": [
    "apify.com",
    "www.smashingmagazine.com",
    "basecamp.com"
  ],
  "maxPagesPerWebsite": 8,
  "verifyMx": true,
  "deobfuscateEmails": true,
  "includeSocials": true,
  "sameDomainOnly": true,
  "maxConcurrency": 10,
  "navigationTimeoutSecs": 20
}
```

# Actor output Schema

## `contacts` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "apify.com",
        "www.smashingmagazine.com",
        "basecamp.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tactful_anvil/contact-details-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        "apify.com",
        "www.smashingmagazine.com",
        "basecamp.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("tactful_anvil/contact-details-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "apify.com",
    "www.smashingmagazine.com",
    "basecamp.com"
  ]
}' |
apify call tactful_anvil/contact-details-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tactful_anvil/contact-details-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/OjuIx87FSTLBx8NxO/builds/PGNUmnadBFQq3Gb7q/openapi.json
