# Website Contact Info Extractor (Emails, Phones, Socials) (`zhucl1006/website-contact-info-extractor`) Actor

Extract emails, phone numbers, social media profiles (LinkedIn, Facebook, Instagram, X, YouTube, TikTok...), contact page, contact form and company address from any list of websites. Checks homepage + contact/about/impressum pages. Pay only for sites read.

- **URL**: https://apify.com/zhucl1006/website-contact-info-extractor.md
- **Developed by:** [leo zhu](https://apify.com/zhucl1006) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.01 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Info Extractor (Emails, Phones, Socials)

Turn a list of websites into a **B2B contact list**: for every site the Actor reads the homepage plus the most relevant contact pages (Contact, About, Impressum / Imprint, Team, Support, Legal notice) and extracts **email addresses, phone numbers, social media profiles (LinkedIn, Facebook, Instagram, X/Twitter, YouTube, TikTok, Pinterest, GitHub, WhatsApp, Telegram, Threads), the contact page URL, whether there is a contact form, the company name and postal address** (from schema.org data), plus the **email provider** (Google Workspace, Microsoft 365, ...) from MX records.

- **Pay only for websites that could be read.** Unreachable sites, sites blocked by anti-bot protection, sites whose robots.txt disallows crawling, invalid URLs and duplicates are free.
- **Same price no matter how many pages** (up to 10) are checked per website.
- **Fast and cheap**: plain HTTP requests, no headless browser - dozens of websites per minute on 512 MB memory.
- **Clean output**: one row per website, with a primary email and phone, emails sorted by relevance (own-domain role addresses like info@ / sales@ first), phone numbers normalised to international E.164 format, flags for role/free-mail/own-domain emails, and the page where each email was found.

### Use cases

- **Lead generation & sales prospecting** - enrich a list of company websites (from Google Maps, directories, trade-show exhibitor lists, CRM exports) with emails, phones and LinkedIn pages.
- **CRM data enrichment & cleanup** - fill in missing emails, phone numbers and social links for existing accounts; detect Google Workspace vs Microsoft 365 users.
- **Local business outreach** - restaurants, clinics, agencies, shops: find the contact email, phone, contact form and Instagram/Facebook page.
- **Influencer & partnership research** - collect social profiles of brands and publishers in bulk.
- **Market research** - which companies in a niche offer a contact form, WhatsApp chat or public phone line.

### Input example

```json
{
  "websites": ["joesstonecrab.com", "https://www.python.org", "apify.com"],
  "maxPagesPerDomain": 4,
  "includeDns": true,
  "sameDomainEmailsOnly": false
}
```

Tips:

- Domains, full URLs or `www.` hosts all work; each website is processed and charged once.
- Set `sameDomainEmailsOnly` to keep only addresses on the site's own domain (drops agency, free-mail and third-party emails).
- Use it after a Google Maps / directory scraper: feed the `website` column in and join the results back by `domain`.

### Output example

Real output from a test run on 2026-09-27:

```json
{
  "input": "joesstonecrab.com",
  "domain": "joesstonecrab.com",
  "url": "https://www.joesstonecrab.com/",
  "status": "ok",
  "companyName": "Joe's Stone Crab",
  "primaryEmail": "info@joesstonecrab.com",
  "emails": [
    "info@joesstonecrab.com"
  ],
  "emailCount": 1,
  "primaryPhone": "+13056730365",
  "phones": [
    "+13056730365",
    "+13056734611",
    "+13056739035"
  ],
  "contactPageUrl": "https://www.joesstonecrab.com/contact/",
  "hasContactForm": true,
  "facebook": "https://www.facebook.com/JoesStoneCrab",
  "instagram": "https://www.instagram.com/joesstonecrab",
  "linkedin": null,
  "x": null,
  "youtube": null,
  "tiktok": "https://www.tiktok.com/@joesstonecrab_",
  "pinterest": null,
  "github": null,
  "whatsapp": null,
  "telegram": null,
  "threads": null,
  "socialProfiles": {
    "facebook": [
      "https://www.facebook.com/JoesStoneCrab",
      "https://www.facebook.com/people/Joes-Take-Away/61559644092931",
      "https://www.facebook.com/JoesStoneCrab",
      "https://www.facebook.com/JoesStoneCrab"
    ],
    "instagram": [
      "https://www.instagram.com/joesstonecrab",
      "https://www.instagram.com/joestakeaway__"
    ],
    "tiktok": [
      "https://www.tiktok.com/@joesstonecrab_"
    ]
  },
  "address": null,
  "emailProvider": [
    "Microsoft 365"
  ],
  "hasMx": true,
  "emailDetails": [
    {
      "email": "info@joesstonecrab.com",
      "sameDomain": true,
      "role": true,
      "freemail": false,
      "sourceUrl": "https://www.joesstonecrab.com/location/joes-stone-crab/"
    }
  ],
  "phoneDetails": [
    {
      "phone": "+13056730365",
      "source": "tel-link"
    },
    {
      "phone": "+13056734611",
      "source": "tel-link"
    },
    {
      "phone": "+13056739035",
      "source": "tel-link"
    }
  ],
  "emailsObfuscated": false,
  "pagesVisited": [
    "https://www.joesstonecrab.com/",
    "https://www.joesstonecrab.com/contact/",
    "https://www.joesstonecrab.com/location/joes-take-away/",
    "https://www.joesstonecrab.com/location/joes-stone-crab/"
  ],
  "pagesSkippedByRobots": [],
  "robotsTxt": "found",
  "checkedAt": "2026-09-27T12:22:27+00:00",
  "errors": []
}
```

Every dataset can be exported as **CSV, Excel, JSON or XML**, or sent to Google Sheets, HubSpot, Zapier, Make and webhooks via Apify integrations. The Overview table is flat (one column per social network) and CSV-friendly. A run summary is stored in the key-value store as `RUN_SUMMARY`.

### Pricing

Pay per event, no monthly fee:

- a small fee per run start, and
- a fee per **website successfully read** (status `ok`), regardless of how many pages were checked or how many contacts were found.

Not charged: `unreachable`, `blocked`, `robotsDisallowed`, `httpError`, `noHtml`, `invalid`, `error` and duplicate inputs. See the Pricing tab for current prices.

### How it works & limitations

- Pages are fetched with plain HTTP (no JavaScript rendering). Sites that build their content only with JavaScript may return fewer contacts.
- **robots.txt is respected** (RFC 9309) and requests to one website are made one after another with a short pause.
- **No anti-bot bypass**: pages behind Cloudflare / captcha challenges are reported as `blocked` and not charged. Emails hidden by Cloudflare email obfuscation are **not decoded**; `emailsObfuscated: true` tells you they exist.
- Phone numbers come from `tel:` links, schema.org data and visible text (validated with Google's libphonenumber); `phoneDetails[].source` shows which. Text-found numbers can occasionally be fax or registry numbers.
- Placeholder and technical addresses (example.com, sentry, image file names such as logo@2x.png, hashes) are filtered out.

### Data protection

The Actor only collects information that the website owners publish on their own public pages. Business contact data can still be personal data (e.g. `firstname@company.com`). **You are responsible for processing the results lawfully** - for example under GDPR (legitimate interest assessment, information duties, opt-out) and anti-spam rules such as CAN-SPAM and PECR - and for respecting each website's terms. Role addresses are flagged with `role: true` to help you prefer generic company contacts.

### FAQ

**How many pages per website are checked?** By default the homepage plus up to 3 contact-related pages (`maxPagesPerDomain`: 4, max 10). The price is the same.

**Does it find personal emails of employees?** Only if they are published on the website. It does not guess or generate email addresses and does not verify mailboxes.

**Can I schedule it or call it from my app?** Yes - use Apify schedules, the REST API, or integrations (Zapier, Make, n8n, webhooks).

**Why is a website `blocked`?** It served an anti-bot challenge. We do not bypass such protections; you are not charged.

### Keywords

website contact scraper, email extractor, email finder from website, contact info scraper, phone number extractor, social media links extractor, LinkedIn company page finder, Facebook page finder, Instagram profile finder, lead generation, B2B leads, sales prospecting, CRM enrichment, local business leads, contact page finder, impressum scraper, bulk website email scraper.

# Actor input Schema

## `websites` (type: `array`):

Website URLs or domains, one per line (e.g. https://www.example.com or example.com). Duplicates are processed and charged once.

## `maxPagesPerDomain` (type: `integer`):

Homepage plus the most relevant contact pages (contact, about, impressum/imprint, team, support, legal). Same price regardless of this setting.

## `includeDns` (type: `boolean`):

Look up the domain's MX records to tell whether it uses Google Workspace, Microsoft 365, etc. No extra charge.

## `sameDomainEmailsOnly` (type: `boolean`):

Drop emails on other domains (e.g. agency, freemail or third-party addresses).

## `maxConcurrency` (type: `integer`):

Websites processed in parallel. Pages of the same website are always fetched one after another.

## `requestTimeoutSecs` (type: `integer`):

Timeout for each page request.

## Actor input object example

```json
{
  "websites": [
    "apify.com",
    "https://www.python.org",
    "basecamp.com"
  ],
  "maxPagesPerDomain": 4,
  "includeDns": true,
  "sameDomainEmailsOnly": false,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 15
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "apify.com",
        "https://www.python.org",
        "basecamp.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("zhucl1006/website-contact-info-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        "apify.com",
        "https://www.python.org",
        "basecamp.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("zhucl1006/website-contact-info-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "apify.com",
    "https://www.python.org",
    "basecamp.com"
  ]
}' |
apify call zhucl1006/website-contact-info-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zhucl1006/website-contact-info-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/1Srb4IG6HdnCe5xXr/builds/8ElXq22JmoS2RAsL8/openapi.json
