# Website Contact Extractor: Emails, Phones & Socials (`redfoxscout/website-contact-extractor`) Actor

Extract business emails, phone numbers, address and social media profiles (LinkedIn, Facebook, Instagram, X, YouTube, TikTok) from any list of company websites. Checks the home page plus contact, about and imprint pages. Works with Google Maps results.

- **URL**: https://apify.com/redfoxscout/website-contact-extractor.md
- **Developed by:** [Red Fox Scout](https://apify.com/redfoxscout) (community)
- **Categories:** Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.13 / 1,000 websites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

Get **one clean row per company with emails you can send to**. For each company website it checks the **home page, contact, about and imprint pages** (found through the links, then the sitemap), checks that each address's domain can receive mail (**MX check**), drops no-reply and junk addresses, and puts the best address first in `primaryEmail`. Phone numbers come as written and in **international format (E.164)**, with the source page for each email and social profiles from 17 networks. For sales teams, agencies and anyone building B2B lead lists.

**Price: $1.50 per 1,000 websites.** Sites that fail to load or block the visit are free. Apify's free plan ($5 monthly credit) covers about 3,300 websites a month.

[![Find contact emails for a list of company websites: real output preview (Company websites in, one row per company with emails checked for a working mail server)](https://raw.githubusercontent.com/TikTop-Data/apify-store-assets/main/images/website-contact-extractor-output.png)](https://apify.com/redfoxscout/website-contact-extractor)

### Questions it answers

- What's the best email to send to for each company on my list?
- Which addresses are on domains that can't receive mail, so I don't waste a send?
- What's each company's phone number in international format?
- Which of these businesses have a TikTok, Pinterest, WhatsApp or Telegram account?
- Where is each company's contact page?
- Who runs each company, and how do I reach them?

### Who uses this

- 📈 **Sales and lead generation:** turn a list of company websites into a contact list you can use.
- 🏢 **Agencies:** enrich prospect lists before outreach or ads targeting.
- 🗂️ **CRM and data teams:** fill missing emails, phones and social links on account records.
- 🗺️ **Local business research:** pair it with Google Maps results to reach local businesses.

### Why this Website Contact Extractor

- 📬 **Emails you can send to:** every email domain is checked for a working mail server, dead addresses are dropped and the best one comes first, with the page it was found on.
- 📄 **Contact pages included:** it follows the site's contact, about and imprint links, and the sitemap when the links are not enough.
- 🧹 **Clean results:** removes placeholder addresses, image file names, tracking addresses and no-reply senders; decodes Cloudflare-protected emails.
- 🔗 **Socials in the same row** as emails and phones, from links and from JSON-LD `sameAs` data.
- 👥 **People on team pages:** names and job titles from each company's own team, leadership and about pages, with the email, phone or LinkedIn shown beside them.
- 🗂️ **Reads other datasets:** pick a Google Maps or lead-list dataset and it uses the website column.
- 🔔 **Change alerts:** give a monitor name and get only sites whose contacts changed since the last run.

### How to use the Website Contact Extractor

1. Click **Try for free** (a free Apify account is enough).
2. Paste company websites or domains, one per line, or pick a dataset with a website field.
3. Click **Start**. Each site takes a few seconds.
4. Download as JSON, CSV or Excel, or send the rows to Google Sheets or your CRM.

### Input

| Field | Default | What it does |
|---|---|---|
| `urls` | – | Company websites or domains, one per line. |
| `startUrls` | – | The same list in Apify's standard start-URL format. |
| `datasetId` | – | Read websites from another Actor's dataset (e.g. Google Maps results). |
| `datasetUrlField` | `website` | Field in that dataset that holds the website. |
| `maxDatasetItems` | 10000 | Read at most this many rows from that dataset. |
| `monitorName` | – | Remember results under this name to detect changes next run. |
| `onlyChanges` | `false` | Output only sites whose contacts changed (needs a monitor name). |
| `maxPagesPerSite` | 4 | Contact, about, imprint or team pages checked per site besides the home page. |
| `includePeople` | `true` | List the people on the company's team, leadership and about pages. Off: no team pages are read. |
| `maxPeople` | 50 | Most people kept per website (up to 1,000). |
| `maxTeamPages` | 5 | Extra pages of a people directory to follow ("Next", page 2, 3...), up to 50. Same price per website. |
| `includeUrlPatterns` | – | Check only pages whose URL matches one of these (see above). |
| `excludeUrlPatterns` | – | Never check pages whose URL matches one of these. |
| `checkEmailDomains` | `true` | Look up each email domain's mail servers (MX). Turn off to keep every address found. |
| `maxConcurrency` | 20 | Sites checked in parallel. |
| `requestTimeoutSecs` | 30 | Give up on a page after this long. |
| `proxyConfiguration` | off | Use your own proxy group if some sites block or rate-limit the checks. |

Example (the input behind the output below):

```json
{ "urls": ["kyokocoffee.com", "terriblelovecoffee.com", "joscoffee.com"] }
```

### Output

One row per website. The real row for kyokocoffee.com from that run, shortened:

```json
{
  "inputUrl": "kyokocoffee.com",
  "url": "https://kyokocoffee.com/",
  "domain": "kyokocoffee.com",
  "httpStatus": 200,
  "companyName": "Kyoko Coffee",
  "primaryEmail": "info@kyokocoffee.com",
  "emails": ["info@kyokocoffee.com"],
  "emailCount": 1,
  "emailsChecked": [
    { "email": "info@kyokocoffee.com", "type": "general", "mx": "valid-mx", "sourceUrl": "https://kyokocoffee.com/" }
  ],
  "emailSources": { "info@kyokocoffee.com": "https://kyokocoffee.com/" },
  "primaryPhone": "(512) 387-1373",
  "phones": ["(512) 387-1373", "+1-512-387-0131"],
  "primaryPhoneE164": "+15123871373",
  "phonesE164": ["+15123871373", "+15123870131"],
  "address": "6401 Airport Boulevard Austin, TX, 78752 United States",
  "linkedin": null,
  "facebook": null,
  "instagram": "https://www.instagram.com/kyoko.coffee",
  "x": null,
  "youtube": null,
  "tiktok": null,
  "socialProfiles": { "instagram": "https://www.instagram.com/kyoko.coffee" },
  "contactPageUrl": "https://kyokocoffee.com/about",
  "pagesChecked": 2,
  "people": [],
  "peopleCount": 0,
  "error": null,
  "checkedAt": "2026-10-08T20:42:22.448Z"
}
```

`primaryEmail` and `emails` are the sendable addresses; `emailsChecked` lists every address found, with its status. `people` and `peopleCount` list the people on team pages (see People on team pages).

### Example tasks

Ready-made inputs you can open, change and run:

- [Find contact emails for a list of company websites](https://apify.com/redfoxscout/website-contact-extractor/examples/contact-emails-checked-one-row-per-company)
- [Phone number extractor for business websites](https://apify.com/redfoxscout/website-contact-extractor/examples/phone-number-extractor-websites)
- [Emails plus Instagram, Facebook and LinkedIn from websites](https://apify.com/redfoxscout/website-contact-extractor/examples/social-media-and-email-from-websites)

### People on team pages

The team, leadership, staff and about pages are read within the same page budget as the contact pages (`maxPagesPerSite`). For each person a page shows, the row gets `name` and `title`, plus `email`, `phone` or `linkedin` only when the page shows them next to that person. Person entries in the page's JSON-LD (founders, employees, board members) are included too. Navigation, footers, testimonials, blog authors and form options are skipped, and a phone or email shared by several people (a switchboard or shared inbox) is left out of each of them.
Big directories that spread people over many pages (law firms, agencies) are followed page by page with **Directory pages per site** (`maxTeamPages`). On hinshawlaw.com/en/professionals, 10 directory pages took a run from 27 to 209 people, 180 of them with their direct phone, in 45 seconds, at the price of one website.

Real output from `{ "urls": ["https://www.hinshawlaw.com/en/professionals"], "maxPeople": 3 }`:

```json
{
  "peopleCount": 3,
  "people": [
    { "name": "Eliot C. Abbott", "title": "Senior Counsel", "email": null, "phone": "305-358-7747", "linkedin": null, "sourceUrl": "https://www.hinshawlaw.com/en/professionals" },
    { "name": "Michael P. Adams", "title": "Partner", "email": null, "phone": "312-704-3184", "linkedin": null, "sourceUrl": "https://www.hinshawlaw.com/en/professionals" },
    { "name": "Jerri C. Adams Belcher", "title": "Associate", "email": null, "phone": "312-704-3046", "linkedin": null, "sourceUrl": "https://www.hinshawlaw.com/en/professionals" }
  ]
}
```

Titles are copied as the page shows them, so some sites add a city or a label to a title. Turn off `includePeople` to skip the team pages.

### Emails you can send to

- **MX check:** each email domain is looked up once for its mail servers. Addresses on a domain that does not exist (`domain-not-found`) or has no mail server (`no-mx`) are left out of `emails` and `primaryEmail`. They stay in `emailsChecked` with that status, so nothing is lost. If the DNS lookup does not answer, the address is kept as `unchecked`.
- **Best address first:** the company's own domain comes first. Then sales@, info@ and contact@, then general inboxes (hello@, enquiries@), then named roles (support@, press@, careers@), then other addresses. Privacy, abuse and webmaster addresses come last, and no-reply addresses are dropped.
- **Where each address came from:** `emailSources` gives the page for every address in `emails`, and `emailsChecked` gives it for every address found.
- **Phones in E.164:** `phonesE164` and `primaryPhoneE164` (e.g. `+14155550100`) when the number has a `+` or `00` prefix, or when the site's country is clear from its JSON-LD address, its domain or its page language. The number as written stays in `phones`.
- **No guessed addresses:** only addresses printed on the company's own pages are returned.

### Choose which pages to check

- **Contact pages per site** (`maxPagesPerSite`, default 4): the number of contact, about, imprint and (with `includePeople`) team pages read besides the home page.
- **Only check pages matching** (`includeUrlPatterns`): check only pages whose URL matches, e.g. `team`, `locations` or `*/offices/*`. A `*` matches any text; a word without `*` matches anywhere in the URL.
- **Skip pages matching** (`excludeUrlPatterns`): never check matching pages, e.g. `blog` or `*/news/*`.
- Without patterns, the contact, about and imprint pages linked from the home page are read first, then pages from `sitemap.xml`.

### Work with other datasets

Pick a dataset under **Websites from another Actor**, e.g. Google Maps Scraper results. The Actor reads its `website` field (or the field you name, such as `url`, `domain` or `contact.website`) and checks every site. For an automatic run after each scrape, add an Actor-to-Actor integration with `{ "datasetId": "{{resource.defaultDatasetId}}", "datasetUrlField": "website" }`.

### Change alerts

Give the run a **Change alert name** (e.g. `clients-weekly`) and put it on an Apify schedule. Each row then gets `changeStatus` (`new`, `changed` or `unchanged`) and `changedFields`, compared with the previous run of the same name. Turn on **Only new and changed sites** and each run's dataset is just the change report. Unchanged sites are still checked and charged as usual.

### What this Actor does not do

It does not guess emails or phones from name patterns, look people up on other sites, do people-search (home addresses, personal phones or emails), or log in to sites. It does not connect to mail servers or send anything, so it shows that a domain accepts mail, not that a mailbox exists.

### How much does it cost?

$1.50 per 1,000 websites checked, contact and team pages included. People add nothing to the price.

- 100 websites: about $0.15
- 10,000 websites: about $15
- Sites that fail to load or block the visit: free

Apify's free plan gives $5 of credit every month, enough for about 3,300 websites.

### FAQ

**Is this legal?** It reads publicly published business contact details from company websites, including the names and job titles companies publish for their own people, and returns them as shown. You are responsible for how you use them, including email marketing laws (CAN-SPAM, GDPR) in your market.

**Why is a site's email empty?** Some sites only offer a contact form or show the address as an image. Those can't be read as text. An address on a domain with no mail server is also left out of `emails`; check `emailsChecked` to see it.

**What does "domain-not-found" or "no-mx" mean?** The address's domain does not exist, or exists but has no mail server, so mail to it will bounce. It is kept in `emailsChecked` with that status.

**Can it read JavaScript-only sites?** It reads the HTML the site sends. Sites that build everything in the browser may return fewer contacts.

### Related Actors

- [Company Social Links Finder](https://apify.com/redfoxscout/company-social-links): official social profiles and key pages for any list of companies.
- [Email Security Checker](https://apify.com/redfoxscout/email-security-checker): mail provider, SPF, DMARC and more for each domain.
- [Tech Stack Detector](https://apify.com/redfoxscout/tech-stack-detector): what technologies each website uses.

# Actor input Schema

## `urls` (type: `array`):

Company websites or domains, one per line. Or pick a dataset below (e.g. Google Maps results). Or use "Websites from another Actor" below.

## `startUrls` (type: `array`):

Same as the list above, in Apify's standard start-URL format (handy when another tool or Actor passes startUrls).

## `datasetId` (type: `string`):

Check the websites in another Actor run's results, e.g. Google Maps Scraper or any lead list. Pick the dataset here, or in an Actor-to-Actor integration set it to {{resource.defaultDatasetId}}. Its websites are added to the list above.

## `datasetUrlField` (type: `string`):

The dataset field that holds the website or URL, e.g. website (Google Maps Scraper), url or domain. Dot paths like contact.website work. If it is empty, website, url and domain are tried.

## `maxDatasetItems` (type: `integer`):

Read at most this many rows from that dataset.

## `monitorName` (type: `string`):

Name this list (e.g. "clients-weekly") to compare each run with the previous run of the same name. Every row then gets changeStatus (new, changed or unchanged) and changedFields. Best on an Apify schedule.

## `onlyChanges` (type: `boolean`):

Needs a change alert name. Unchanged sites are still checked and charged as usual, but left out of the results, so each run's dataset is your change report.

## `maxPagesPerSite` (type: `integer`):

Besides the home page, how many contact, about or imprint pages to check on each site.

## `includePeople` (type: `boolean`):

List the people the company shows on its own team, about, leadership and staff pages: name, title, and the email, phone or LinkedIn shown beside them. Team pages count toward Contact pages per site. Off: no team pages are read and people is null.

## `maxPeople` (type: `integer`):

Most people returned per website (default 50, up to 1,000). Large firms list hundreds; raise this together with Directory pages.

## `maxTeamPages` (type: `integer`):

People directories split over several pages (law firms, agencies) are followed page by page ('Next'), up to this many extra pages per site, until the people limit is reached. Same price per site.

## `includeUrlPatterns` (type: `array`):

Check only pages whose URL matches one of these, e.g. team, locations or */team*. A \* matches any text; a word without \* matches anywhere in the URL. Empty = the contact, about and imprint pages found through links or the sitemap.

## `excludeUrlPatterns` (type: `array`):

Never check pages whose URL matches one of these, e.g. blog or */news/*. Same matching as above.

## `checkEmailDomains` (type: `boolean`):

Look up each email domain's mail servers (MX). Addresses on domains that cannot receive mail are left out of emails and primaryEmail but stay in emailsChecked with their status. Free: no mail is sent.

## `maxConcurrency` (type: `integer`):

How many sites to check at once.

## `requestTimeoutSecs` (type: `integer`):

Give up on a site after this long.

## `proxyConfiguration` (type: `object`):

Off by default. Turn on if some sites block or rate-limit the checks.

## Actor input object example

```json
{
  "urls": [
    "kyokocoffee.com",
    "terriblelovecoffee.com",
    "joscoffee.com"
  ],
  "datasetUrlField": "website",
  "maxDatasetItems": 10000,
  "onlyChanges": false,
  "maxPagesPerSite": 4,
  "includePeople": true,
  "maxPeople": 50,
  "maxTeamPages": 5,
  "includeUrlPatterns": [],
  "excludeUrlPatterns": [],
  "checkEmailDomains": true,
  "maxConcurrency": 20,
  "requestTimeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "kyokocoffee.com",
        "terriblelovecoffee.com",
        "joscoffee.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("redfoxscout/website-contact-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "kyokocoffee.com",
        "terriblelovecoffee.com",
        "joscoffee.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("redfoxscout/website-contact-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "kyokocoffee.com",
    "terriblelovecoffee.com",
    "joscoffee.com"
  ]
}' |
apify call redfoxscout/website-contact-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,redfoxscout/website-contact-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/z1yCTmMnvGTvAgkg0/builds/LcZGRHOk5KEJVs68S/openapi.json
