# Website Contact Extractor: Email & Phone Scraper, Socials $2/1k (`transparent_meteorite/website-contact-extractor`) Actor

Bulk-extract emails, phone numbers and social profiles from business websites. Crawls home + contact, about, team, impressum and privacy pages; one clean row per site with role/named email tags, E.164 phones, 8 socials and address. $0.002 per site.

- **URL**: https://apify.com/transparent_meteorite/website-contact-extractor.md
- **Developed by:** [Open Data Actors](https://apify.com/transparent_meteorite) (community)
- **Categories:** Lead generation, Social media, Developer tools
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Extractor: Emails, Phones, Socials

**Video walkthrough:** [Watch the 1-minute demo on YouTube](https://www.youtube.com/watch?v=-0hrPhp9xTM)

![Website Contact Extractor on Apify](https://api.apify.com/v2/key-value-stores/9BAn0msnxToXlD9TO/records/banner-website-contact-extractor.png)

![Website Contact Extractor sample output table](https://api.apify.com/v2/key-value-stores/9BAn0msnxToXlD9TO/records/output-website-contact-extractor.png)

- **In short:** Website Contact Extractor (Apify actor `transparent_meteorite/website-contact-extractor`) crawls each website's home, contact, about, team, impressum and privacy pages and returns one clean row of contact details per site.
- **Who it is for:** lead-generation agencies, sales teams enriching domain lists, marketers building outreach lists from their own site lists.
- **Input:** start URLs or domains, max pages per site, include named (personal) emails, same domain only, only with contacts, default country for phones.
- **Output:** emails tagged role or named, phones in E.164, LinkedIn, Facebook, Instagram, X, TikTok, YouTube and other socials, address, source page for each value.
- **Price:** $0.002 per site crawled plus $0.00005 per run start; sites that fail to load are not charged. Pay per result, no subscription; Apify's free plan credit covers a first test.
- **Limits:** finds only contacts published on the site; no guessing or verification of email addresses.

**Key facts**

- Actor name: Website Contact Extractor
- Actor ID: `transparent_meteorite/website-contact-extractor`
- Store page: https://apify.com/transparent_meteorite/website-contact-extractor
- Data source: the public website itself
- Pricing model: pay per event (`apify-actor-start` $0.00005, `site-crawled` $0.002)
- Output formats: JSON, CSV, Excel, XML, HTML table, RSS (Apify dataset)
- Access: Apify Console, REST API, JavaScript/Python clients, schedules, webhooks, Apify MCP server
- Login or third-party API key needed: no, only an Apify account
- Also known as: email extractor, contact details scraper, website email scraper, phone number extractor, social media links finder
- Maintainer: transparent_meteorite (independent developer)
- Last updated: 2026-10-07

Turn a list of business websites into a clean contact sheet: **one row per site** with every email (tagged role or named, with the page it was found on), phone numbers normalized to E.164, 8 social profiles, the contact page URL, company name and street address. **$0.002 per site, all pages included.**

### What does it do

For each website you give it, the actor:

1. Fetches the homepage (plain HTTP, no browser, robots.txt respected).
2. Finds the pages where businesses actually publish contact details, by link text and URL in 10+ languages: **contact, impressum/imprint, about, team, privacy**, and crawls the best ones (default 5 pages per site, configurable 1-20).
3. Extracts and merges everything into one row:
   - **Emails** from `mailto:` links, visible text, schema.org data, **Cloudflare-protected addresses (cfemail decoded)** and **obfuscated forms** like `jane [at] acme [dot] com`, `info(at)acme.de`, `sales AT acme DOT com`.
   - Each email is tagged **`role`** (info@, sales@, support@, bookings@...) or **`named`** (jane.doe@...), with `foundOn` page URL, page type and how it was found.
   - **Phones** from `tel:` links, schema.org and page text, validated and normalized to **E.164** (`+13135550142`) with country.
   - **Socials**: LinkedIn, Facebook, Instagram, X/Twitter, YouTube, TikTok, Pinterest, GitHub (share buttons and post links filtered out).
   - **Company name** (schema.org, og:site_name or cleaned page title), **address** (schema.org PostalAddress), **contact page URL**.
4. Flags dead, blocked and parked sites with a clear `status` so you never pay for them.

### Why use it

- **Priced per site, not per page.** Contact scrapers that charge per page cost 5x more for the same 5-page crawl. Here a site is $0.002 no matter how many of its pages are crawled.
- **One row per site.** No stitching per-page rows back together; CSV/Excel ready with flat `emailsList`, `phonesList`, `linkedin`, `facebook`... columns.
- **Role vs named emails** so you can route info@ to cold outreach and personal addresses to careful, compliant handling (or switch named emails off entirely).
- **Source page for every email and phone** so you can verify where it came from.
- **E.164 phones** ready for dialers, CRMs and SMS tools.
- **Fair billing:** unreachable, blocked (403/captcha), parked-domain and robots-disallowed sites are returned with a status and **never charged**.
- Fast and light: plain fetch + cheerio, 5 sites in about 12 seconds on 512 MB.

### How to use

1. Click **Try for free** and paste website URLs or bare domains into **Websites** (or upload a CSV/TXT file of URLs).
2. Optionally set **Max pages per site**, turn **Include named emails** off for role-only output, or set **Default phone country** for local numbers.
3. Click **Start**. Each site appears as a row as soon as it finishes.
4. Download as CSV, Excel or JSON, or open the **Emails** view for one row per email with its type and source page.

### Input example

```json
{
  "startUrls": [
    { "url": "https://www.shinola.com" },
    { "url": "https://buddyspizza.com" },
    { "url": "https://slowsbarbq.com" }
  ],
  "domains": ["batchbrewingcompany.com", "greektownchicago.org"],
  "maxPagesPerSite": 5,
  "includeNamedEmails": true,
  "sameDomainOnly": true,
  "defaultCountry": "US"
}
```

### Sample output

Real rows from a platform run (5 sites, 12 s, default input):

| Domain | Company | Emails | Phones | Socials | Contact page | Pages |
|---|---|---|---|---|---|---|
| shinola.com | Shinola | customerservice@shinola.com, privacy@shinola.com | | linkedin, facebook, instagram, x, youtube, tiktok, pinterest | | 4 |
| buddyspizza.com | Buddy's Pizza | buddyspizza@buddyspizza.com | +18009650505, +18334510774 | facebook, instagram, x | buddyspizza.com/contact/ | 4 |
| slowsbarbq.com | Slows Bar BQ | events@slowsbarbq.com, manager@slowsbarbq.com | +13139629828 | facebook, instagram | | 1 |
| batchbrewingcompany.com | Batch Brewing Company | events@batchbrewingcompany.com, contact@batchbrewingcompany.com | +13133388008 | facebook, instagram, tiktok | batchbrewingcompany.com/contact | 2 |
| greektownchicago.org | Greektown Chicago | contact@greektownchicago.org | +13122852508 | facebook, instagram, x | greektownchicago.org/contact/ | 4 |

One item (shortened):

```json
{
  "domain": "batchbrewingcompany.com",
  "status": "ok",
  "companyName": "Batch Brewing Company",
  "companyNameSource": "schema.org",
  "primaryEmail": "events@batchbrewingcompany.com",
  "emails": [
    { "email": "events@batchbrewingcompany.com", "type": "role", "sameDomain": true, "foundOn": "https://www.batchbrewingcompany.com/contact", "pageType": "contact", "method": "mailto" },
    { "email": "contact@batchbrewingcompany.com", "type": "role", "sameDomain": true, "foundOn": "https://www.batchbrewingcompany.com/contact", "pageType": "contact", "method": "mailto" }
  ],
  "primaryPhone": "+13133388008",
  "phones": [{ "phone": "+13133388008", "national": "(313) 338-8008", "country": "US", "foundOn": "https://www.batchbrewingcompany.com/", "pageType": "home", "method": "schema.org" }],
  "socials": { "linkedin": null, "facebook": "https://www.facebook.com/batchbrewingcompany", "instagram": "https://www.instagram.com/batchbrewing", "x": null, "youtube": null, "tiktok": "https://www.tiktok.com/@batchbrewingcompany", "pinterest": null, "github": null },
  "contactPageUrl": "https://www.batchbrewingcompany.com/contact",
  "address": { "street": "1400 Porter Street", "city": "Detroit", "region": "MI", "postalCode": "48226-2409", "country": "US", "full": "1400 Porter Street, Detroit, MI 48226-2409, US" },
  "pagesCrawled": 2,
  "emailsList": "events@batchbrewingcompany.com, contact@batchbrewingcompany.com",
  "roleEmails": "events@batchbrewingcompany.com, contact@batchbrewingcompany.com",
  "namedEmails": null,
  "phonesList": "+13133388008"
}
```

`status` values: `ok`, `blocked` (403/429/bot challenge), `http_error`, `unreachable` (DNS/timeout), `parked` (domain for sale), `robots_disallowed`, `not_html`. Only `ok` sites with at least one crawled page are charged.

### Pricing

Pay per event, no subscription:

| Event | Price |
|---|---|
| Actor start | $0.00005 per run |
| Site crawled | **$0.002 per site** (up to 20 pages included) |

**Example:** 1,000 websites = 1,000 x $0.002 + $0.00005 = **about $2.00**, whether you crawl 1 or 5 pages per site. If 120 of them are dead, parked or blocked you pay about $1.76. Set a max charge in the run options and the actor stops cleanly when it is reached.

### Use from AI agents (MCP, ChatGPT, Claude, Perplexity)

Website Contact Extractor works as a tool for AI assistants through the official Apify MCP server. Add this server URL to any MCP client (Claude Desktop, Claude Code, Cursor, ChatGPT connectors, VS Code):

```
https://mcp.apify.com/?actors=transparent_meteorite/website-contact-extractor
```

Claude Desktop / Cursor config:

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com/?actors=transparent_meteorite/website-contact-extractor",
      "headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }
    }
  }
}
```

Then ask in plain language, for example: "How do I get emails and phone numbers from a list of websites?"

Call it directly over HTTP (runs the actor and returns the dataset items in one request):

```bash
curl -X POST "https://api.apify.com/v2/acts/transparent_meteorite~website-contact-extractor/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"startUrls": [{"url": "https://www.shinola.com"}, {"url": "https://buddyspizza.com"}, {"url": "https://slowsbarbq.com"}, {"url": "https://www.batchbrewingcompany.com"}, {"url": "https://greektownchicago.org"}], "maxPagesPerSite": 5, "includeNamedEmails": true, "sameDomainOnly": true}'
```

Python:

```python
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("transparent_meteorite/website-contact-extractor").call(run_input={"startUrls": [{"url": "https://www.shinola.com"}, {"url": "https://buddyspizza.com"}, {"url": "https://slowsbarbq.com"}, {"url": "https://www.batchbrewingcompany.com"}, {"url": "https://greektownchicago.org"}], "maxPagesPerSite": 5, "includeNamedEmails": true, "sameDomainOnly": true})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

The same actor works in n8n, Make, Zapier, LangChain, LlamaIndex and CrewAI through their Apify integrations.

#### Questions people ask

**How do I get emails and phone numbers from a list of websites?**
Paste the URLs into startUrls; you get one row per site with emails, phones and social profiles.

**Does it guess or verify emails?**
No, it only returns emails published on the site, with the page it was found on.

**Am I charged for sites that are down?**
No, only sites that were crawled are charged.

### Integrations

**API (cURL)**

```bash
curl -X POST "https://api.apify.com/v2/acts/transparent_meteorite~website-contact-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"domains":["acme.com","example-bakery.com"],"maxPagesPerSite":5}'
```

**Python**

```python
from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("transparent_meteorite/website-contact-extractor").call(run_input={"domains": ["acme.com"]})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(row["domain"], row["primaryEmail"], row["primaryPhone"])
```

**JavaScript**

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('transparent_meteorite/website-contact-extractor').call({ domains: ['acme.com'] });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
```

**No-code:** Make, Zapier, n8n, Google Sheets and Airtable via Apify integrations; **webhooks** on run finish; **schedules** to re-check a lead list weekly. A common chain: Google Maps or directory scraper -> this actor -> CRM.

### FAQ

**Is it legal to extract contact details from websites?**
The actor only reads public pages that businesses publish themselves, without logins, captcha bypass or proxies, and it obeys robots.txt. Business contact data is generally public, but named personal emails can be personal data under GDPR, CCPA and similar laws. You are responsible for having a lawful basis for how you store and use the results (e.g. B2B legitimate interest, honoring opt-outs, CAN-SPAM). Turn **Include named emails** off if you only want role addresses.

**Why did a site return no emails?**
Many sites only offer a contact form, or render contacts with JavaScript. The row still includes phones, socials and the contact page URL when found. Sites that load everything with JavaScript may need a browser-based scraper.

**What does role vs named mean?**
`role` = a function mailbox such as info@, sales@, support@, bookings@ or brand@brand.com. `named` = an address that looks like a person (jane.doe@, jsmith@).

**Which pages are crawled?**
The homepage, then links classified as contact, impressum/imprint, about, team and privacy (English, German, French, Spanish, Italian, Dutch, Portuguese), in that priority. If no contact link exists, `/contact` and `/contact-us` are tried.

**Do I pay for failed sites?**
No. Invalid input, DNS failures, timeouts, 403/429/captcha pages, parked domains and robots.txt-disallowed sites are returned with a status and are not charged.

**How are phones normalized?**
With libphonenumber. International numbers are parsed as written; local numbers use the site's country-code domain (.de, .co.uk...) or the **Default phone country** input. Numbers that cannot be validated are dropped from text, and kept unnormalized when they come from an explicit `tel:` link.

**Can it crawl a whole website?**
It is built for contact discovery, not full-site crawling: up to 20 targeted pages per site keeps it fast and cheap.

**Does it use proxies?**
No. Requests go out directly with an honest user agent at polite speed (pages of one site are fetched one by one).

### Limits

- JavaScript-only sites and sites behind Cloudflare/bot challenges return `blocked` or few results.
- Up to 50 emails and 15 phones per site are kept (most relevant first: same-domain, mailto/schema, role).
- Addresses are taken from schema.org markup only (no free-text address guessing).

### Changelog

- 2026-10-07: added plain-language summary, key facts, AI-agent (MCP) section and question-style FAQ; refreshed Store SEO metadata.
- **0.1 (2026-10-07)**: first release. Emails with role/named tags and source page, cfemail and \[at]/(dot) de-obfuscation, E.164 phones, 8 socials, schema.org address, impressum/team discovery, per-site pricing, no charge for failed sites.

### Support

Found a site that should have worked, or need a field added? Open an issue on the actor's **Issues** tab with the URL and what you expected; we usually reply within a day.

# Actor input Schema

## `startUrls` (type: `array`):

Business websites to extract contacts from. Paste URLs or bare domains (acme.com). One row is returned per site; duplicates are merged. Upload a CSV/text file of URLs for bulk runs.

## `domains` (type: `array`):

Optional plain list of domains, one per line (e.g. acme.com). Merged with Websites above.

## `maxPagesPerSite` (type: `integer`):

Homepage plus the best contact, impressum, about, team and privacy pages found by link text and URL (in that priority). All pages of a site are included in the one per-site price.

## `includeNamedEmails` (type: `boolean`):

Every email is tagged 'role' (info@, sales@, support@...) or 'named' (jane.doe@...). Turn off to return role addresses only.

## `sameDomainOnly` (type: `boolean`):

Only follow links on the site's own domain and subdomains. Emails on other domains (e.g. a gmail address shown on the site) are still returned and flagged sameDomain=false.

## `onlyWithContacts` (type: `boolean`):

Skip rows that have no email, phone or social profile. Sites still count as crawled.

## `defaultCountry` (type: `string`):

ISO 2-letter country used to normalize local-format phone numbers to E.164 (+13135550142). Country-code domains (.de, .co.uk...) override it automatically.

## `maxSites` (type: `integer`):

Cap on sites crawled (and charged) this run.

## `maxConcurrency` (type: `integer`):

How many sites to crawl at once. Pages inside one site are fetched one after another to stay polite.

## `pageTimeoutSecs` (type: `integer`):

Timeout per page request.

## `siteTimeoutSecs` (type: `integer`):

Total time budget per site; whatever was found so far is returned.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.shinola.com"
    },
    {
      "url": "https://buddyspizza.com"
    },
    {
      "url": "https://slowsbarbq.com"
    },
    {
      "url": "https://www.batchbrewingcompany.com"
    },
    {
      "url": "https://greektownchicago.org"
    }
  ],
  "maxPagesPerSite": 5,
  "includeNamedEmails": true,
  "sameDomainOnly": true,
  "onlyWithContacts": false,
  "defaultCountry": "US",
  "maxSites": 1000,
  "maxConcurrency": 8,
  "pageTimeoutSecs": 15,
  "siteTimeoutSecs": 60
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `emails` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.shinola.com"
        },
        {
            "url": "https://buddyspizza.com"
        },
        {
            "url": "https://slowsbarbq.com"
        },
        {
            "url": "https://www.batchbrewingcompany.com"
        },
        {
            "url": "https://greektownchicago.org"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("transparent_meteorite/website-contact-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [
        { "url": "https://www.shinola.com" },
        { "url": "https://buddyspizza.com" },
        { "url": "https://slowsbarbq.com" },
        { "url": "https://www.batchbrewingcompany.com" },
        { "url": "https://greektownchicago.org" },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("transparent_meteorite/website-contact-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.shinola.com"
    },
    {
      "url": "https://buddyspizza.com"
    },
    {
      "url": "https://slowsbarbq.com"
    },
    {
      "url": "https://www.batchbrewingcompany.com"
    },
    {
      "url": "https://greektownchicago.org"
    }
  ]
}' |
apify call transparent_meteorite/website-contact-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,transparent_meteorite/website-contact-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/zxuR66xK9DndMEUtr/builds/jgX1APrZB34V3X0pF/openapi.json
