# Website Contact Scraper: Emails, Phones & Socials (`enisbodlli/website-contact-scraper`) Actor

Extracts company contact details from a list of websites: role email addresses such as info@ and sales@, phone numbers and social media links. One row per website, and you pay only for websites where a contact was found. Follows robots.txt and leaves out personal addresses.

- **URL**: https://apify.com/enisbodlli/website-contact-scraper.md
- **Developed by:** [Enis Bodlli](https://apify.com/enisbodlli) (community)
- **Categories:** Lead generation, Automation, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 website with contacts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Scraper

Turn a list of company websites into a contact list. This **website contact scraper** reads each
site's home page and its contact, about, imprint and legal pages, and returns one row per website
with the company's **email addresses**, **phone numbers** and **social media links**. To try it,
leave the three sample websites in the input and click **Start**: the run takes a few seconds.

**You pay only for websites where a contact was found.** A website with no contact, a site that
blocks the visit and a domain that does not exist still get a row, so you can see they were checked,
and cost nothing.

It returns company contact points, not people: shared mailboxes such as `info@`, `sales@` or
`kontakt@`, the company's phone numbers and its company pages on social networks. Addresses that
look like a person's (`firstname.lastname@`) are left out and counted.

### What you can do with a website contact scraper

- **Build a lead list from domains.** Feed in the websites of companies you want to reach and get
  the general email address, the phone number and the contact page of each.
- **Fill gaps in a CRM.** Add the missing email, phone or LinkedIn company page to accounts that
  only have a website.
- **Find the right channel.** `contactPageUrl` points to the contact form when a company publishes
  no email address.
- **Check a list of domains.** The `status` of each row tells a working site from one that is gone
  or that refuses automated visits.
- **Call it from code or an AI agent.** Start the Actor through the Apify API and read the rows as
  JSON.

### How to extract emails and phone numbers from websites

1. Paste the websites under **Websites**, one per line, as domains (`example.com`) or URLs. Through
   the API, pass them as an array.
2. Optionally change **Maximum pages per website**, or switch phone numbers or social links off.
3. Click **Start**. Each website's row is saved as soon as that website is done, so a run you stop
   keeps what it finished.
4. Open the **Output** tab and export the table as JSON, CSV or Excel, or read it through the API.

### Pricing

You pay per website **with at least one contact found** (an email address, a phone number or a
social media link):

| Apify plan | Per 1,000 websites with contacts | Per website |
|---|---|---|
| Free, Bronze | $3.00 | $0.003 |
| Silver | $2.50 | $0.0025 |
| Gold | $2.00 | $0.002 |

Plus $0.00005 per run start. Platform usage is included, so there is nothing else to pay. The price
does not depend on how many pages are read: a website is charged once.

- **1,000 websites, 640 of them with a contact:** 640 charges. That is $1.92 on the Free and Bronze
  plans, $1.60 on Silver and $1.28 on Gold. The other 360 websites are in the results at no cost.
- **50 websites, 31 of them with a contact:** 31 charges, $0.093 on the Free and Bronze plans.

Repeated domains are removed before the run, so no website is charged twice. A website is charged
at the moment its row is saved. A run that is stopped, fails or is restarted continues where it left
off: websites that already have a row are skipped, so they are not charged again.

You can set a maximum charge per run. The Actor then starts a website only while that limit has
room for it, stops when the limit is reached, and reports how many websites it did not check.

### Input

| Field | What it does |
|---|---|
| Websites (`websites`) | The websites to read. Required. Domains or URLs, up to 10,000 per run. Repeated domains are removed; the first entry counts. Only pages on the website's own domain (and its subdomains) are read. |
| Maximum pages per website (`maxPagesPerSite`) | The home page plus the most likely contact pages. Default 8, at most 20. |
| Include phone numbers (`includePhones`) | Default on. When off, `phones` is an empty list. |
| Include social media links (`includeSocialLinks`) | Default on. When off, every entry of `socialLinks` is `null`. |

Example:

```json
{
    "websites": ["hetzner.com", "basecamp.com", "g2.com"],
    "maxPagesPerSite": 5,
    "includePhones": true,
    "includeSocialLinks": true
}
```

### Output

One row per website, with the contacts from all its pages merged. `domain` identifies the row.
Below is the result of the example input, from a run on 2026-10-04: a website with contacts, one
without, and one that refuses automated visits. Only the first is charged.

```json
[
    {
        "website": "hetzner.com",
        "domain": "hetzner.com",
        "status": "ok",
        "statusReason": "Found 2 emails, 5 phone numbers and 5 social links on 4 pages. 3 personal-looking addresses left out.",
        "emails": ["info@hetzner.com", "contact-fi@hetzner.com"],
        "personalEmailsSkipped": 3,
        "phones": [
            { "text": "+49 (0)9831 505-0", "digits": "4998315050" },
            { "text": "+358 (0)753259-0", "digits": "3587532590" },
            { "text": "+49 (0) 911 234226-100", "digits": "49911234226100" },
            { "text": "+49 (0) 3745 74447-100", "digits": "49374574447100" },
            { "text": "+358 753259-100", "digits": "358753259100" }
        ],
        "socialLinks": {
            "linkedin": "https://www.linkedin.com/company/hetzner-online",
            "facebook": "https://www.facebook.com/hetzner.de",
            "instagram": "https://www.instagram.com/hetzner.online",
            "x": "https://x.com/Hetzner_Online",
            "youtube": "https://www.youtube.com/user/HetznerOnline",
            "tiktok": null,
            "github": null
        },
        "contactPageUrl": "https://www.hetzner.com/support/",
        "pageTitle": "Günstige Dedicated Server, Cloud & Hosting aus Deutschland",
        "pagesCrawled": 4,
        "checkedAt": "2026-10-04T13:12:09.893Z"
    },
    {
        "website": "basecamp.com",
        "domain": "basecamp.com",
        "status": "no_contacts_found",
        "statusReason": "No company email, phone number or social link on 2 pages. 1 personal-looking address left out.",
        "emails": [],
        "personalEmailsSkipped": 1,
        "phones": [],
        "socialLinks": { "linkedin": null, "facebook": null, "instagram": null, "x": null, "youtube": null, "tiktok": null, "github": null },
        "contactPageUrl": null,
        "pageTitle": "Basecamp",
        "pagesCrawled": 2,
        "checkedAt": "2026-10-04T13:12:09.639Z"
    },
    {
        "website": "g2.com",
        "domain": "g2.com",
        "status": "blocked",
        "statusReason": "The home page could not be read: HTTP 403 (access denied to automated visits).",
        "emails": [],
        "personalEmailsSkipped": 0,
        "phones": [],
        "socialLinks": { "linkedin": null, "facebook": null, "instagram": null, "x": null, "youtube": null, "tiktok": null, "github": null },
        "contactPageUrl": null,
        "pageTitle": null,
        "pagesCrawled": 0,
        "checkedAt": "2026-10-04T13:12:09.449Z"
    }
]
```

| Field | Meaning |
|---|---|
| `website` | The website as you entered it. |
| `domain` | Its host name in lower case, without `www.`. One row per domain. |
| `status` | `ok`, `no_contacts_found`, `blocked` or `unreachable`. See below. |
| `statusReason` | One sentence saying why. |
| `emails` | Shared and role addresses of the company, lower-cased, without repeats, addresses on the website's own domain first. At most 25. |
| `personalEmailsSkipped` | How many addresses were left out because they are not recognisably a shared mailbox. |
| `phones` | Up to 10 phone numbers: `text` as the site shows it, `digits` with digits only. An international number starts with its country code. |
| `socialLinks` | The company's page on LinkedIn, Facebook, Instagram, X, YouTube, TikTok and GitHub, or `null`. |
| `contactPageUrl` | The website's contact page, or `null`. |
| `pageTitle` | Title of the home page. |
| `pagesCrawled` | How many pages were read. |
| `checkedAt` | When the website was checked (UTC). |

Websites are read in parallel, so rows can be in a different order than your list. Match them on
`domain`. A summary of the run (websites per status, charges, entries that were skipped) is saved
as `RUN_SUMMARY` in the run's key-value store.

### What each status means

| Status | Meaning | Charged |
|---|---|---|
| `ok` | At least one email address, phone number or social link was found. | Yes |
| `no_contacts_found` | The site was read and shows no company contact point on the pages visited. | No |
| `blocked` | The site refused the visit: HTTP 401, 403 or 429, a CAPTCHA or JavaScript challenge, or a robots.txt rule. | No |
| `unreachable` | The domain does not exist, the server did not answer in time, or the home page returned an error. | No |

### Limits

- **Personal-looking addresses are left out.** An address is returned only when the part before the
  `@` is a function or a department, in English and the major European languages. A name
  (`jane.smith@`), initials, a single first name, or any word the Actor does not know is not
  returned, only counted in `personalEmailsSkipped`. The purpose is company contact points, not
  personal data about individuals. This also drops some harmless addresses with unusual names.
- **Sites that block automated visits are reported, not forced.** The Actor sends plain requests
  from Apify's servers under its own name (`website-contact-scraper`). It uses **no proxies and no
  browser** and solves no CAPTCHAs. A refusal is final, with one exception: when a site answers "too
  many requests" and names a wait of up to 30 seconds, the Actor waits that long and asks again.
  A site that refuses comes back as `blocked`.
- **JavaScript-only sites may show nothing.** Pages that build their content in the browser arrive
  empty. They come back as `no_contacts_found`, with a note in `statusReason`.
- **Addresses a site hides on purpose are not decoded**, for example Cloudflare email protection or
  `info (at) example.com`.
- **Phone numbers** are taken from phone links and from numbers next to a label such as "Tel" or
  "Phone", or written in international form. Fax and mobile numbers, and numbers on team pages, are
  left out. A number without a label in a national format can be missed.
- **Pages read:** the home page and the contact, about, imprint, legal and team pages linked from
  it, on the same domain. Privacy policies and terms are skipped, because the addresses there
  mostly belong to others. After 90 seconds no further page of a website is requested, and a page
  larger than 5 MB is not read.
- **Social links** are the profiles a site links to. Personal profiles such as LinkedIn `/in/` are
  skipped; on other networks a company page and a personal one look the same.

### FAQ

#### Is this legal?

The Actor reads public pages of the websites you list, the same pages a visitor sees. It follows
each site's robots.txt, does not log in anywhere and does not get around blocks. It returns company
contact points and leaves out addresses that look like a person's. What you do with the results is
your responsibility: whether and how you may contact a company depends on the laws that apply to
you and to them, such as rules on unsolicited email.

#### Why is there no email for a website?

Many companies publish only a contact form. Look at `contactPageUrl`. If `personalEmailsSkipped` is
above zero, the site lists addresses that were left out. If `statusReason` mentions JavaScript, the
site needs a browser.

#### Why is a website "blocked"?

The site answered a plain, identified request with a refusal. Large consumer sites often do. The
Actor reports this and moves on; the row costs nothing.

#### Can I get the addresses of named people?

No. This Actor returns shared mailboxes only.

#### How long does a run take?

A website takes a few seconds, and ten are read at a time. In a test on 2026-10-04, 15 company
websites from twelve countries took 10 seconds. A site that does not answer takes about a minute
before it is reported as unreachable.

#### Can I check whether the addresses work?

Yes, with [Email Validator & List Cleaner](https://apify.com/enisbodlli/email-validator), which
checks syntax and whether the domain receives mail.

#### Where do I report a problem?

Open an issue on the **Issues** tab of this Actor. Name the website and say what you expected to
find on it. Issues are answered there.

### More Actors from this developer

Contacts and lists:

- [Email Validator & List Cleaner](https://apify.com/enisbodlli/email-validator)

Jobs:

- [Company Jobs Search](https://apify.com/enisbodlli/company-jobs-search)
- [ATS Job Postings: Workday, Greenhouse, Lever & Ashby](https://apify.com/enisbodlli/ats-job-postings)
- [Workday Jobs Scraper](https://apify.com/enisbodlli/workday-jobs-scraper)
- [Greenhouse Jobs Scraper](https://apify.com/enisbodlli/greenhouse-jobs-scraper)
- [Lever Jobs Scraper](https://apify.com/enisbodlli/lever-jobs-scraper)
- [Ashby Jobs Scraper](https://apify.com/enisbodlli/ashby-jobs-scraper)

Company registers:

- [Handelsregister Scraper: German Company Register](https://apify.com/enisbodlli/handelsregister-scraper)
- [North Data Scraper: German & European Companies](https://apify.com/enisbodlli/northdata-company-scraper)
- [European Company Registry Search](https://apify.com/enisbodlli/eu-company-registry-search)
- [US Business Entity Search & New Business Filings](https://apify.com/enisbodlli/us-business-registry-search)
- [Brazil CNPJ Scraper: Company Search & Lookup](https://apify.com/enisbodlli/brazil-cnpj-company-search)

# Changelog

This Actor's version history is a separate document: https://apify.com/enisbodlli/website-contact-scraper/changelog.md

# Actor input Schema

## `websites` (type: `array`):

The company websites to read, one per line, as a domain (example.com) or a URL (https://www.example.com). Each website gives one result row, with the contacts from all its pages merged. Only pages on the website's own domain are read. Repeated domains are removed. Up to 10,000 websites per run.

## `maxPagesPerSite` (type: `integer`):

How many pages to read on each website at most: the home page plus the contact, about, imprint, legal and team pages it links to. More pages find more contact details and take longer. The price does not depend on it: a website is charged once, however many pages are read.

## `includePhones` (type: `boolean`):

Return the company's phone numbers, taken from phone links and from numbers that stand next to a label such as "Tel" or "Phone". Fax and mobile numbers are left out. When switched off, "phones" is an empty list.

## `includeSocialLinks` (type: `boolean`):

Return the company's pages on LinkedIn, Facebook, Instagram, X, YouTube, TikTok and GitHub. Personal profiles such as a LinkedIn /in/ address are left out. When switched off, every entry of "socialLinks" is null.

## Actor input object example

```json
{
  "websites": [
    "apify.com",
    "hetzner.com",
    "automattic.com"
  ],
  "maxPagesPerSite": 5,
  "includePhones": true,
  "includeSocialLinks": true
}
```

# Actor output Schema

## `results` (type: `string`):

One row per website checked, in the run's default dataset: status, emails, phone numbers, social profiles and the contact page.

## `runSummary` (type: `string`):

Totals for the run: websites with contacts (charged), without contacts, blocked and unreachable (not charged).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "apify.com",
        "hetzner.com",
        "automattic.com"
    ],
    "maxPagesPerSite": 5
};

// Run the Actor and wait for it to finish
const run = await client.actor("enisbodlli/website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": [
        "apify.com",
        "hetzner.com",
        "automattic.com",
    ],
    "maxPagesPerSite": 5,
}

# Run the Actor and wait for it to finish
run = client.actor("enisbodlli/website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "apify.com",
    "hetzner.com",
    "automattic.com"
  ],
  "maxPagesPerSite": 5
}' |
apify call enisbodlli/website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,enisbodlli/website-contact-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XQ2JSJnaSepMiPuaS/builds/bopfb6vzC0zrPKPPh/openapi.json
