# Website Email Scraper - Contact Details, Phones & Socials (`tidytools/website-contact-extractor`) Actor

Contact details scraper and email extractor for company websites: emails (incl. obfuscated), phone numbers, social media links, address and contact forms. $2/1,000 sites; nothing found = free.

- **URL**: https://apify.com/tidytools/website-contact-extractor.md
- **Developed by:** [Yukai Lin](https://apify.com/tidytools) (community)
- **Categories:** Lead generation, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.60 / 1,000 processed websites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Website Email & Contact Details Scraper do?

Give it company or business websites (URLs or bare domains) and get one clean record per website:

- 📧 **Emails**: `mailto:` links, addresses written in the text, **obfuscated addresses** ("info \[at] firm \[dot] de", "info(at)firm.de") and **Cloudflare-protected emails** (decoded). Lower-cased, de-duplicated, and junk is filtered out: image file names such as `logo@2x.png`, placeholder and example addresses (`you@example.com`, `name@domain.com`), error-tracking IDs and no-reply addresses.
- 📞 **Phone and fax numbers** in international format (`+12122542246`), from `tel:` links, schema.org data and the page text. Every number is checked with Google's libphonenumber rules for its country, and numbers in text are only taken after a label ("Tel:", "Phone", "電話"...), in international format, or in a typical phone layout, so order numbers, dates and prices are not picked up.
- 👥 **Social profiles**: Facebook, X, LinkedIn, Instagram, YouTube, TikTok, GitHub, Pinterest, Threads, Trustpilot, Discord, Telegram and WhatsApp. Only real profile URLs count (share buttons, posts and videos are ignored), and default profiles from website templates (e.g. `facebook.com/wix`) are skipped.
- 🏢 **Organization name, logo and postal address** from the site's schema.org data (JSON-LD or Microdata: Organization, LocalBusiness, Restaurant, Dentist...).
- 📝 **Contact form URL**: the page with a contact form (HTML forms with a message box, or HubSpot, Contact Form 7, Typeform, Jotform and other embeds).
- 🔎 **Where each value was found**: the page and the method (`mailto`, `text`, `obfuscated`, `cloudflare`, `jsonld`, `tel-link`), plus the list of pages checked.

For each website it reads the **home page plus likely contact pages**: contact, imprint/Impressum, locations, about, team, support and legal pages (in English, German, French, Spanish, Italian, Dutch, Japanese, Chinese and Korean wording), up to 5 pages by default.

### Use cases

- **Lead generation**: turn a list of company domains into emails, phone numbers and social profiles
- **Enrich Google Maps or directory results**: add emails and socials to businesses that only list a website
- **CRM clean-up**: find current contact details and social accounts for your accounts
- **Outreach research**: find the press, sales or support address of each company

### How much does it cost?

| Event | Price |
|---|---|
| Processed website | **$2.00 / 1,000 websites** ($0.002 each) |

**Example:** 2,000 websites, of which 1,500 have at least one contact, cost 1,500 × $0.002 = **$3.00**; the other 500 are free.

**No start fee.** One price per website, whether 1 or 20 pages are read. **You pay only when at least one email, phone number or social profile is found.** Websites with nothing found, unreachable websites (they carry an `errorType`: `blocked`, `network`, `not_found`, `timeout`…), invalid lines and duplicates are free. Every row has `charged: true` or `false`. Higher Apify plans get volume discounts (see the *Pricing* tab).

#### Control your cost

- **Before anything is read**, the run logs its plan: the number of websites × the price = the most the run can cost, compared with your maximum charge per run.
- **Charged**: a website where at least one email, phone number or social profile was found ($0.002).
- **Free**: websites with nothing found, unreachable or blocked websites, lines that are not a website (e.g. a company name such as "Acme Inc"), and duplicates: `acme.com`, `www.acme.com` and `acme.com/en` are one website, charged once. There is no extra charge from us for more pages, the browser fallback or our own network; only an Apify Proxy you choose yourself is billed by Apify.
- **Maximum charge per run**: set it in the run options (Console) or with `maxTotalChargeUsd` (API). When the next website could exceed it, no new website is started; the run ends with status `LIMIT_REACHED`, and the `SUMMARY` record lists the websites not processed (`notProcessed`: the count and up to 100 inputs) so you can run them again with a higher limit.
- **If Apify restarts the run** (server migration or Resurrect), items already finished are skipped and not charged again (`SUMMARY.resumedSkipped`).

#### Price comparison (checked September 2026)

| Actor | Price | Cost for 1,000 websites at 5 pages each |
|---|---|---|
| **Website Contact Details & Social Media Extractor (this Actor)** | **$0.002 per website with contacts found** | **$2 at most** (websites with nothing found are free) |
| vdrmota/contact-info-scraper | $0.002 per page scraped | $10 |
| compass/crawler-google-places, contacts add-on | $0.002 per place, only inside a new Google Maps scrape (plus $0.004 per place scraped) | $2 for contacts (+ $4 for the places) |

vdrmota/contact-info-scraper is the most used Actor in this category (about 2,260 users in the last 30 days when we checked) and can follow links more deeply; it charges per page. This Actor charges per website where something was found and picks the likely contact pages for you.

### How to use it

1. Paste domains or website URLs into **Domains or websites**, one per line (a column copied from a spreadsheet works). Blank lines are ignored; a line that is not a website, such as a company name, gets one `invalid_input` row and is not charged. Or pick another run's dataset (see below).
2. Optionally change **Pages per website** (1–20, default 5).
3. Click **Start** and export the results as JSON, CSV or Excel.

#### Output example (real result, September 2026)

```json
{
    "input": "stumptowncoffee.com",
    "inputIndex": 0,
    "url": "https://stumptowncoffee.com/",
    "domain": "stumptowncoffee.com",
    "finalUrl": "https://www.stumptowncoffee.com/",
    "success": true,
    "organizationName": "Stumptown Coffee Roasters",
    "primaryEmail": "info@stumptowncoffee.com",
    "primaryPhone": "+18777113385",
    "emailsText": "info@stumptowncoffee.com, legal@stumptown.com, customerservice@stumptown.com",
    "phonesText": "+18777113385, +18332672844, +18557113385",
    "emails": ["info@stumptowncoffee.com", "legal@stumptown.com", "customerservice@stumptown.com"],
    "phones": ["+18777113385", "+18332672844", "+18557113385"],
    "faxes": [],
    "socialProfiles": {
        "instagram": "https://www.instagram.com/stumptowncoffee",
        "facebook": "https://www.facebook.com/stumptowncoffee",
        "x": "https://twitter.com/stumptowncoffee",
        "youtube": "https://www.youtube.com/channel/UC0yk_H5Np-uGvy98UM-jj6w",
        "tiktok": "https://www.tiktok.com/@stumptowncoffee",
        "pinterest": "https://www.pinterest.com/stumptowncoffee"
    },
    "address": "700 SW 5th Ave, 3rd Floor, #400, 97214 Portland, OR",
    "logo": "https://customers.seomanager.com/knowledgegraph/logo/stumptowncoffee_myshopify_com_logo.png",
    "contactFormUrl": null,
    "emailDetails": [
        { "email": "info@stumptowncoffee.com", "page": "https://www.stumptowncoffee.com/", "method": "jsonld", "sameDomain": true, "mailbox": "role" },
        { "email": "legal@stumptown.com", "page": "https://www.stumptowncoffee.com/pages/terms-and-conditions", "method": "text", "sameDomain": false, "mailbox": "role" }
    ],
    "phoneDetails": [
        { "number": "+18777113385", "display": "877-711-3385", "national": "(877) 711-3385", "country": "US", "type": "phone", "page": "https://www.stumptowncoffee.com/", "method": "jsonld" }
    ],
    "pagesChecked": [
        { "url": "https://www.stumptowncoffee.com/", "kind": "home", "via": "direct", "httpStatus": 200 },
        { "url": "https://www.stumptowncoffee.com/pages/contact-us", "kind": "contact", "via": "direct", "httpStatus": 200 },
        { "url": "https://www.stumptowncoffee.com/pages/locations", "kind": "locations", "via": "direct", "httpStatus": 200 },
        { "url": "https://www.stumptowncoffee.com/pages/our-story", "kind": "about", "via": "direct", "httpStatus": 200 },
        { "url": "https://www.stumptowncoffee.com/pages/terms-and-conditions", "kind": "legal", "via": "direct", "httpStatus": 200 }
    ],
    "charged": true,
    "via": "direct"
}
```

A German restaurant (hofbraeuhaus.de) returned `willkommen@hofbraeuhaus.de` (mailto link), `+4989290136100` (written as "+49 (0)89 – 290 136 100" in the text) and its Facebook and Instagram profiles. A Munich food retailer (dallmayr.com) returned `info@dallmayr.de`, decoded from Cloudflare's email protection.

For spreadsheets and CRMs, `primaryEmail` is an email on the website's own domain (else the first email), a role mailbox such as info@, contact@, hello@ or sales@ before a person's address, `primaryPhone` the first phone number (not a fax), and `emailsText` / `phonesText` hold all of them separated by commas. `input` is the line you gave; when several lines were the same website they are merged into one row and listed in `inputs` (a page path such as `hofbraeuhaus.de/en` is read as an extra page, `kind: "given"`). Emails on the website's own domain come first (`sameDomain: true`); others (a parent company, an agency, a data-protection officer) follow. `mailbox` in `emailDetails` is `role` for known shared inboxes (info@, sales@, support@, press@…) and `other` for the rest (often a person's address, but not always), so you can filter shared inboxes out or in. Profiles linked only once, e.g. from a blog post, are not reported as the company's own. `via` says where each page was read from: `direct` (Apify's network), `backend` (our servers) or `browser`. A `SUMMARY` record in the key-value store has the run `status` (`SUCCESS`, `PARTIAL_RESULTS`, `FAILED`, `NO_RESULTS`, `LIMIT_REACHED`) and counts of websites with emails, phones and social profiles, charged websites, invalid lines and merged duplicates.

#### Find contacts for a Google Maps (or any) dataset

Set **Dataset ID** to the dataset of another run, **Dataset field** to the field with the website (e.g. `website`) and, optionally, **Dataset fields to keep** (e.g. `title`, `placeId`). Each output row then carries `datasetItemIndex` and a `source` object with those fields, so it maps back to the place it came from.

```json
{ "datasetId": "YOUR_DATASET_ID", "datasetField": "website", "datasetKeepFields": ["title", "placeId"] }
```

Every item of the dataset gets exactly one output row, so a Google Maps export joins back 1:1 (by `datasetItemIndex` or a kept field such as `placeId`):

- A place **without a website** gets a row with `errorType: "no_website"` (not charged).
- A place whose "website" is an **Instagram, Facebook, Yelp, Linktree or similar page** is not read as a company website (that would return the platform's own details): its row has `errorType: "not_a_website"`, the profile in `socialProfiles`, and is not charged.
- **Branches of a chain** that share one website are read and charged once; each other branch gets a copy of the result with its own `source`, `charged: false` and `duplicateOf`.
- When a website **redirects to a different domain** (for example an expired domain now parked on a directory), the row has `redirectedTo` and a `warning`: the details found may belong to that other site.

**Real example (September 30, 2026).** 30 Google Maps places from compass/crawler-google-places (10 coffee shops in Portland, 10 dentists in Austin, 10 bakeries in Munich), with `datasetKeepFields: ["title", "placeId"]`:

| Result | Places |
|---|---|
| Output rows (one per place, each with its `placeId`) | 30 of 30 |
| No website on Google Maps | 1 (free) |
| Instagram or Facebook page instead of a website | 3 (free, profile returned) |
| Website no longer exists | 2 (free) |
| **Websites read with contacts found (charged)** | **24** |
| … with an email | 21 (70% of all places, 88% of the websites read) |
| … with a social profile | 23 |

Cost: 24 × $0.002 = **$0.048** (about **$1.60 per 1,000 places**); run time 12 seconds. Google Maps already lists a phone number for almost every place, so the email and the social profiles are what this step adds.

### Use with AI agents (MCP)

Connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/website-contact-extractor) to Claude, Cursor or any MCP client, then ask e.g. "Find the emails, phone numbers and social profiles of these 30 companies."

```json
{ "domains": ["stumptowncoffee.com", "hofbraeuhaus.de"] }
```

Websites with nothing found and unreachable websites are not charged; failed rows carry an `errorType`.

### Use it from code

```bash
curl -X POST "https://api.apify.com/v2/acts/tidytools~website-contact-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"domains":["stumptowncoffee.com"]}'
```

Send websites in `domains` (plain strings; bare domains work). The URL-list field `startUrls` is also accepted, but Apify rejects the whole run (HTTP 400) when it holds a bare domain or a blank line.

Need industry, a one-line description and a country guess as well? See **Company Website Enrichment** by tidytools.

**Schedules and integrations:** run it daily or weekly with Apify Schedules, get a webhook when a run finishes, or send the results to Zapier, Make, n8n, Google Sheets, Slack and other apps with Apify integrations. Results can be exported as JSON, CSV, Excel or XML.

### Advanced settings

- **Plain HTTP requests from**: *Auto* (recommended) reads pages from Apify's network first, because the raw HTML is needed for Cloudflare-protected emails and contact forms, then from our servers, then with a real browser. JavaScript-only home pages are read with a browser automatically.
- **Time limit per website** (default 120 seconds): after it, no further pages of that website are started and the pages read so far are returned (the skipped ones are listed in `timedOutPages`). If even the home page could not be read in time, the website fails with `errorType: "timeout"` (not charged).
- **Proxy**: optional Apify Proxy for requests sent from Apify's network (billed to your Apify account).

### Limitations

- Only public pages; nothing behind a login. Sites with strong bot protection can stay unreachable (not charged); a residential proxy may help.
- Values are what the website shows. Phone numbers written without a country code are read with the site's likely country (address, international numbers on the site, domain ending, language), which can be wrong for multi-country sites.
- Email addresses shown only as images, or only inside a contact form, cannot be found.
- Store directories can list many numbers (e.g. one per branch); all of them are returned, each with its page.
- The Actor reads a handful of pages per site, not the whole site.

### Responsible use

- The Actor reads only information that a company publishes on its own public website (business emails, phone numbers, social profiles, address). It does not guess, generate or verify personal email addresses, and it does not log in anywhere.
- You are responsible for how you use the results: follow the data-protection and anti-spam laws that apply to you and to the people you contact (for example GDPR in the EU/UK, CAN-SPAM in the US, CASL in Canada), and respect each website's terms.
- Business contact details can still be personal data (e.g. a named employee's email). Keep only what you need and honour opt-out requests.

### FAQ

**Which pages does it read on each website?** The home page plus likely contact pages: contact, imprint/Impressum, locations, about, team, support and legal pages, up to 5 pages by default.

**Do I pay for websites where nothing is found?** No. You pay only when at least one email, phone number or social profile is found. Unreachable websites, invalid lines and duplicates are free too.

**Can it add emails to Google Maps results?** Yes. Pass the dataset ID of a Google Maps (or any) run and the field that holds the website; the fields you choose, such as `placeId`, are copied to each row.

**Can it find emails that are hidden from bots?** Often, yes: obfuscated addresses such as "info \[at] firm \[dot] de" and Cloudflare-protected emails are decoded. Addresses shown only as images or only inside a contact form cannot be found.

### Support

Open an issue in the **Issues** tab with the website. Issues are checked regularly.

# Actor input Schema

## `domains` (type: `array`):

Main input (fill this, or `startUrls` / `datasetId`). Bare domains or URLs as plain text, one per line (a column pasted from a spreadsheet works), e.g. acme.com or www.bakery.co.uk/en (https:// is added). Each website is processed and charged once: acme.com, www.acme.com and acme.com/en are one website (a page path you give is read as an extra page). Lines that are not a domain, such as a company name, are reported as invalid and not charged. Up to 50,000 per run.

## `startUrls` (type: `array`):

The same as "Domains or websites", as a URL list; you can also upload or link a text/CSV file of websites. Bare domains belong in the list above (this editor rejects them).

## `datasetId` (type: `string`):

Find contacts for the websites in another Actor run's dataset, e.g. a Google Maps or directory scraper.

## `datasetField` (type: `string`):

Field with the website or domain, e.g. "website" (dot paths such as "company.url" work). Empty = the first URL field found (url, website, link, domain...).

## `datasetKeepFields` (type: `array`):

Fields copied from each dataset item into the output's "source" object, e.g. title, placeId, address, so every row maps back to the business it came from.

## `maxPagesPerSite` (type: `integer`):

The home page plus likely contact pages (contact, imprint/Impressum, about, team, support, legal). The price per website is the same however many pages are read.

## `maxConcurrency` (type: `integer`):

How many websites are processed at the same time (each reads up to 2 pages at once).

## `siteTimeoutSecs` (type: `integer`):

After this time no further pages of a website are started and the pages read so far are returned (skipped pages are listed in timedOutPages). If even the home page is not read by then, the website fails (not charged).

## `httpVia` (type: `string`):

Some sites block requests from data centers. Auto retries from a second network before falling back to a real browser. If those routes are refused, Auto also tries our second server (Oracle, different IP) before a real browser.

## `proxyConfiguration` (type: `object`):

Only used for requests sent from Apify's network. Apify proxy usage is billed to your Apify account.

## Actor input object example

```json
{
  "domains": [
    "katzsdelicatessen.com",
    "hofbraeuhaus.de"
  ],
  "maxPagesPerSite": 5,
  "maxConcurrency": 5,
  "siteTimeoutSecs": 120,
  "httpVia": "auto"
}
```

# Actor output Schema

## `contacts` (type: `string`):

No description

## `full` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "katzsdelicatessen.com",
        "hofbraeuhaus.de"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tidytools/website-contact-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "katzsdelicatessen.com",
        "hofbraeuhaus.de",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("tidytools/website-contact-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "katzsdelicatessen.com",
    "hofbraeuhaus.de"
  ]
}' |
apify call tidytools/website-contact-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tidytools/website-contact-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kFhBWk7SCYTakF6Fn/builds/uy4pJNb0vZsD4SPb2/openapi.json
