# Website Email & Contact Scraper - Pay Only for Results (`cancap/website-contact-scraper`) Actor

Get emails, phone numbers and social media links from any list of websites. One clean row per website with the best email and phone, LinkedIn, Facebook, Instagram and more. Junk is filtered out. Pay only for websites where an email or phone is found.

- **URL**: https://apify.com/cancap/website-contact-scraper.md
- **Developed by:** [CANCAP](https://apify.com/cancap) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 contact results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Scraper: Emails, Phones & Social Links

Paste a list of websites and get **emails, phone numbers and social media links** for each one, as **one clean row per website**. Built for lead lists: the best email and best phone are picked for you, junk is filtered out, and **you pay only for websites where an email or phone was found**.

No coding, no login and no API keys. Paste your websites and click Start.

### What you get for each website

- **Best email** and **all emails** found, ranked (own-domain and contact addresses first)
- **Best phone** and **all phones**, checked and written in international format (`+14155552671`)
- **Social links**: LinkedIn, Facebook, Instagram, X / Twitter, YouTube, TikTok, Pinterest, GitHub, Telegram, Threads
- **WhatsApp** number, when the website has a WhatsApp link
- **Company name**, **contact page** link, **address** and **country** when the website states them
- **Status** for every website, so you always know why a row is empty

### Who uses this

- **Sales and lead generation**: turn a list of company websites into a contact list.
- **Agencies**: enrich Google Maps, directory or Shopify store exports with emails and socials.
- **Recruiters and partnerships teams**: find the right inbox and LinkedIn page for a company.
- **Marketplaces and researchers**: check which businesses in a list are reachable, and how.

### How to use

1. Paste your websites, one per line. `https://www.example.com` and `example.com` both work.
2. Click **Start**.
3. Export the results to CSV, Excel or JSON, or send them to your CRM, Google Sheets, Make, Zapier or n8n.

It works well as the second step after a Google Maps or directory scraper: take the `website` column and paste it here.

```json
{
    "websites": ["https://www.eff.org", "fsf.org", "https://www.python.org"],
    "maxPagesPerWebsite": 5
}
```

### Output example

```json
{
    "inputUrl": "acme-plumbing.co.uk",
    "website": "https://www.acme-plumbing.co.uk",
    "domain": "acme-plumbing.co.uk",
    "status": "ok",
    "companyName": "Acme Plumbing Ltd",
    "email": "info@acme-plumbing.co.uk",
    "emails": ["info@acme-plumbing.co.uk", "office@acme-plumbing.co.uk"],
    "phone": "+441134960000",
    "phones": ["+441134960000", "+447400123456"],
    "linkedin": "https://www.linkedin.com/company/acme-plumbing",
    "facebook": "https://www.facebook.com/acmeplumbing",
    "instagram": "https://www.instagram.com/acmeplumbing",
    "twitter": null,
    "youtube": null,
    "tiktok": null,
    "whatsapp": "+447400123456",
    "contactPageUrl": "https://www.acme-plumbing.co.uk/contact-us/",
    "address": "1 High St, Leeds, LS1 1AA, GB",
    "country": "GB",
    "language": "en-GB",
    "emailDetails": [
        { "email": "info@acme-plumbing.co.uk", "type": "role", "sameDomain": true, "freeProvider": false }
    ],
    "phoneDetails": [
        { "phone": "+441134960000", "formatted": "+44 113 496 0000", "country": "GB", "type": "fixed_line" }
    ],
    "pagesScanned": 3,
    "scrapedAt": "2026-10-02T10:00:00.000Z"
}
```

### How it finds contacts

For each website the Actor reads the first page and then the pages most likely to show contacts: **contact, imprint, about and team** pages, in many languages (contact, kontakt, contacto, impressum, a-propos and more). You choose how many pages per website (default 5).

What keeps the results clean:

- **Emails**: image names (`logo@2x.png`), template placeholders (`name@example.com`), tracking keys, no-reply addresses and website-builder addresses are removed. Emails on domains that cannot receive mail are removed (mail server check).
- **Phones**: every number is validated for its country, so dates, prices and order numbers are not returned as phones. Fax numbers are skipped.
- **Social links**: share buttons, single posts and website-builder default links are skipped. Only profile pages are returned.

### Status column

| Status | Meaning | Charged |
| --- | --- | --- |
| `ok` | An email or a phone was found | Yes |
| `social_only` | Only social media links were found | No |
| `no_contacts` | The website was read, nothing was found | No |
| `unreachable` | The website does not exist or did not answer | No |
| `blocked` | The website refuses automated visitors | No |
| `robots_disallowed` | robots.txt asks robots not to read the website | No |
| `invalid_url` | The input line is not a website address | No |

### Pricing: pay only for results

- **Contact result**: charged once per website where at least one email or phone was found.
- Websites without contacts, unreachable websites and duplicates are **free**.

Set a **maximum cost per run** in the run options and the Actor stops cleanly at your limit.

### Good to know

- **Plain HTML only**: the Actor does not run a browser. This keeps it fast and cheap, but contacts that a website draws only with JavaScript can be missed.
- **Hidden addresses stay hidden**: emails that a website owner has deliberately obfuscated (for example `info [at] example [dot] com`) are not decoded.
- **No blocking workarounds**: if a website refuses automated visitors, the row says `blocked` and you are not charged.
- **Phone countries**: numbers written without a country code are read using the website's country (domain ending, language, address). If the country is unclear, only numbers with a country code are returned.
- **Email check**: the mail server check confirms that the domain can receive mail. It does not test whether the individual mailbox exists.
- **Large lists**: thousands of websites per run are fine (roughly 1,000 websites in 5 to 10 minutes; give the run more memory to go faster). If a long run is restarted by the platform, it continues where it stopped and nothing is charged twice.

### Responsible use

This Actor reads only pages that websites publish for everyone, respects robots.txt by default, uses no logins and bypasses no protection. Contact details can be personal data. Use them only where you have a lawful reason, for example business-to-business contact, and follow the privacy and anti-spam rules that apply to you (GDPR, CAN-SPAM and similar).

# Actor input Schema

## `websites` (type: `array`):

One website per line. Full addresses (https://www.example.com) and plain domains (example.com) both work. Use Bulk edit to paste a long list. Duplicate lines are skipped.

## `maxPagesPerWebsite` (type: `integer`):

The Actor reads the first page and then the pages most likely to show contacts: contact, imprint, about and team pages. 5 is enough for most websites. Use 1 to read only the first page.

## `saveOnlyWithContacts` (type: `boolean`):

Leave off to get one row for every website, including the ones without contacts or that could not be reached (these rows are free and show the reason in the status column).

## `onlySameDomainEmails` (type: `boolean`):

Turn on to keep only addresses like info@example.com for example.com. Leave off if your targets are small businesses, which often use Gmail or similar.

## `checkEmailDomains` (type: `boolean`):

Looks up the mail server (MX record) of every email's domain and removes addresses on domains without one. This catches typos and dead domains. It does not test the individual mailbox.

## `respectRobotsTxt` (type: `boolean`):

Skip pages that the website's robots.txt asks robots not to read. Recommended.

## `proxyConfiguration` (type: `object`):

Websites are read directly first. The proxy is only used as a second try when a website does not answer.

## Actor input object example

```json
{
  "websites": [
    "https://www.eff.org",
    "https://www.fsf.org",
    "https://www.python.org"
  ],
  "maxPagesPerWebsite": 5,
  "saveOnlyWithContacts": false,
  "onlySameDomainEmails": false,
  "checkEmailDomains": true,
  "respectRobotsTxt": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `contacts` (type: `string`):

One row per website: best email and phone, all emails and phones, social media links.

## `allFields` (type: `string`):

No description

## `summary` (type: `string`):

How many websites had contacts, how many had none, and how many could not be reached.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "https://www.eff.org",
        "https://www.fsf.org",
        "https://www.python.org"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("cancap/website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": [
        "https://www.eff.org",
        "https://www.fsf.org",
        "https://www.python.org",
    ],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("cancap/website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "https://www.eff.org",
    "https://www.fsf.org",
    "https://www.python.org"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call cancap/website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cancap/website-contact-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/7MQZ00cZmEm1nUXUy/builds/Kd8pbsEOIA1gF04rJ/openapi.json
