# Website Contact Extractor - Emails, Phones & Socials (`plainsight/website-contact-extractor`) Actor

Extract emails, phone numbers and social profiles (LinkedIn, X, Facebook, Instagram, YouTube, TikTok...) from company websites in bulk. Checks the homepage plus contact, about, imprint and team pages. Decodes Cloudflare-protected emails. You only pay for sites where contacts are found.

- **URL**: https://apify.com/plainsight/website-contact-extractor.md
- **Developed by:** [kaleb ashton](https://apify.com/plainsight) (community)
- **Categories:** Lead generation, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 website with contacts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Extractor: emails, phones & social profiles from any website

Give it a list of company websites and get back their **email addresses, phone numbers and social media profiles** (LinkedIn, X/Twitter, Facebook, Instagram, YouTube, TikTok, GitHub and more), ready for your CRM, outreach tool or spreadsheet.

- 📬 Checks the **homepage plus the pages where contacts actually live**: Contact, About, Imprint/Impressum, Team, Support, Privacy (in English, German, French, Spanish, Italian and Portuguese)
- 🔓 Decodes **Cloudflare-protected emails** and `name [at] domain [dot] com` obfuscation that basic scrapers miss
- ☎️ Phone numbers **validated and normalised to international E.164 format** (+14155550123)
- 🧹 Filters junk: image filenames, tracking IDs, placeholder addresses like `you@example.com`
- 💸 **You only pay for websites where contact data is found**: $4 per 1,000

### What can you use it for?

- **Lead generation**: turn a list of company domains (from Google Maps, directories, trade-show lists, Shopify store lists...) into contactable leads.
- **CRM enrichment**: fill missing emails, phones and LinkedIn pages for accounts you already have.
- **Partnerships & PR**: find press, partnership or support contacts at scale.
- **Market research**: collect the social footprint of every company in a segment.
- **AI agents**: a reliable "find contact details for this company" tool via API or MCP.

### What data do you get?

| Field | Description |
|---|---|
| `primaryEmail` | Best general contact email (role address like info@/hello@/sales@ preferred) |
| `emails` | All valid emails found |
| `emailsOnDomain` | Emails on the company's own domain |
| `roleEmails` / `namedEmails` | Generic role addresses vs. addresses of named people |
| `primaryPhone`, `phones` | Validated phone numbers in E.164 format |
| `linkedin`, `twitter`, `facebook`, `instagram`, `youtube`, `tiktok` | First profile found per network (flat columns for CSV) |
| `socialProfiles` | All profiles per network, also including GitHub, Pinterest, Crunchbase, Threads, Bluesky, Discord, Telegram, WhatsApp, Yelp, Trustpilot, G2, Medium, Vimeo, Wellfound, Glassdoor |
| `companyName`, `address` | From the site's structured data and metadata |
| `pagesChecked` | Exactly which pages were read |

#### Sample output (trimmed)

Real output for `allbirds.com` (trimmed):

```json
{
  "domain": "allbirds.com",
  "companyName": "Allbirds",
  "primaryEmail": "help@allbirds.com",
  "primaryPhone": "+14243638064",
  "twitter": "https://twitter.com/allbirds",
  "facebook": "https://www.facebook.com/weareallbirds",
  "instagram": "https://www.instagram.com/allbirds",
  "youtube": "https://www.youtube.com/channel/UCnGErLCau5qNJ0Xwe6uEyTw",
  "roleEmails": ["help@allbirds.com"],
  "phones": ["+14243638064"],
  "pagesChecked": ["https://www.allbirds.com/", "https://www.allbirds.com/pages/help"]
}
```

### How to use it

1. Paste websites into **Websites**, one per line (`acme.com`, `https://www.acme.com/about` and so on).
2. Optionally set how many extra pages per site to check (default 4).
3. Click **Start** and export to CSV, Excel or JSON, or push straight to Google Sheets, HubSpot, Make, Zapier, n8n or Clay.

```bash
curl -X POST "https://api.apify.com/v2/acts/plainsight~website-contact-extractor/run-sync-get-dataset-items?token=YOUR_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"domains": ["basecamp.com", "mozilla.org"]}'
```

### Pricing

**$0.004 per website where at least one email, phone number or social profile is found** ($4 per 1,000). Websites with no contact data, unreachable sites and error pages are **free**, and they still appear in your results so you know they were checked.

**Try it free:** start a run with an empty input and it analyses 3 sample websites at no charge, so you can see the exact output format first.

### FAQ

**Does it find personal emails of employees?**
It only collects contact details that the company itself publishes on its own public website. It does not guess, generate or verify email addresses, and it does not visit social networks or log in anywhere.

**Is this GDPR compliant?**
The Actor collects publicly available business contact information. Whether and how you may use it depends on your jurisdiction and purpose (e.g. GDPR legitimate interest, CAN-SPAM). Make sure your outreach complies with the rules that apply to you.

**Why are some sites missing emails?**
Many companies only offer a contact form. In that case you'll still get phone numbers and social profiles when available, and you're only charged if something was found.

**Can it handle JavaScript-heavy sites?**
It reads the HTML the server sends, which covers the vast majority of company websites. Contact details that only appear after client-side JavaScript runs may be missed.

### Related tools by the same developer

- [Company Enrichment](https://apify.com/plainsight/company-enrichment): contacts **plus** tech stack, email provider, domain age and hosting in one row.
- [Tech Stack Detector](https://apify.com/plainsight/tech-stack-detector): 7,600+ technologies and internal SaaS tools from DNS.
- [Email Provider & DMARC Checker](https://apify.com/plainsight/email-security-audit): Google Workspace vs Microsoft 365, SPF/DKIM/DMARC grade.
- [Impressum Scraper](https://apify.com/plainsight/impressum-extractor): managing directors, register number and verified VAT ID from German/EU imprint pages.

# Actor input Schema

## `domains` (type: `array`):

Domains or URLs, one per line (e.g. acme.com or https://www.acme.com/about). Duplicates and www. prefixes are handled automatically.

## `maxPagesPerDomain` (type: `integer`):

Besides the homepage, how many contact-like pages (Contact, About, Imprint/Impressum, Team, Support, Privacy) to check per website. More pages find more contacts but take longer.

## `maxDomains` (type: `integer`):

Only process the first N domains (useful for a quick test).

## `maxConcurrency` (type: `integer`):

How many websites to process in parallel.

## Actor input object example

```json
{
  "domains": [
    "basecamp.com",
    "wordpress.org",
    "mozilla.org"
  ],
  "maxPagesPerDomain": 4,
  "maxConcurrency": 20
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "basecamp.com",
        "wordpress.org",
        "mozilla.org"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("plainsight/website-contact-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "basecamp.com",
        "wordpress.org",
        "mozilla.org",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("plainsight/website-contact-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "basecamp.com",
    "wordpress.org",
    "mozilla.org"
  ]
}' |
apify call plainsight/website-contact-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,plainsight/website-contact-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yCDJb2ucyMExcKpe3/builds/O8Wy1KN4zZFxo2EJB/openapi.json
