# Company Contact Details & Website Email Finder (`datagleaner/website-contact-details-scraper`) Actor

Find emails for a list of company websites: $0.002 per site with an email found, free when none is. Get each company's emails, phones, own social accounts and contact form, each with its source page. Plain HTTP, no browser.

- **URL**: https://apify.com/datagleaner/website-contact-details-scraper.md
- **Developed by:** [Data Gleaner](https://apify.com/datagleaner) (community)
- **Categories:** Lead generation, Automation, Developer tools
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 website with an emails

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Company Contact Details & Website Email Finder: emails, phones and social profiles from websites

Find emails for a list of company websites: $0.002 per site with an email found, free when none is. Get each company's emails, phones, own social accounts and contact form, each with its source page. Plain HTTP, no browser.

Extract emails, phone numbers, social profiles and contact forms from a list of websites. You get one row per website: paste company domains and get each site's public contact details, with the page each value came from. Phones are normalized to E.164, addresses come from schema.org data, and company names from JSON-LD, Open Graph and meta tags. It uses plain HTTP (no browser), so thousands of sites run fast and cheap, and you pay only for sites where an email was found.

**To try it:** paste one or more website addresses into `websites` and click Start.

### What it does

For each website it fetches the home page, then the pages most likely to hold contact details: contact, about, impressum / legal notice, team, company profile and footer links, including Japanese (お問い合わせ, 会社概要, 特定商取引法), Chinese (聯絡我們, 關於我們), German (Kontakt, Impressum), French, Spanish and Italian equivalents. It stays on the same site (a subdomain only when a link clearly names a contact page, such as help.example.com/contact), follows redirects (http to https, www or not), respects `robots.txt` and does at most `maxPagesPerSite` pages per site (default 8).

Extracted:

- **Emails**: `mailto:` links, plain text, obfuscated forms (`name [at] domain [dot] com`, `name(at)domain.com`), Cloudflare email protection, and JSON-LD. Addresses on the site's own domain (or a brand-named domain, or webmail such as Gmail) go in `emails`; addresses on unrelated domains (partners, quoted customers, demo data) are kept apart in `otherEmails` and are not counted as the site's contacts or billed. On gnu.org, for example, 38 other-domain addresses landed in `otherEmails` and none in `emails`.
- **Phone numbers**: `tel:` links, text and JSON-LD, validated and normalized to E.164 with Google's libphonenumber. Local numbers are read using the site's country, inferred from the domain, page language and structured data (or your `defaultCountry`).
- **Social profiles** on 14 platforms: LinkedIn, X (Twitter), Facebook, Instagram, YouTube, TikTok, GitHub, Pinterest, Threads, Telegram, WhatsApp, Discord, Reddit and Snapchat. `socials` holds the site's own accounts (linked from its header, footer or nav, or matching its brand); other people's profiles it merely links (testimonials, team, press) go in `mentionedSocials`. Share buttons and post links are ignored.
- **Contact forms**: pages with a message form, plus embedded Google Forms, Typeform, HubSpot, Marketo, Pardot, Jotform, Calendly and similar, including forms a script draws after the page loads (recognized from the embed code, since pages are not rendered).
- **Addresses**: schema.org `PostalAddress` in JSON-LD.
- **Company name and description** from JSON-LD, Open Graph and meta tags.

### Use cases

Lead generation and sales prospecting from a company list, enriching a CRM, building supplier or partner directories, outreach research, checking that a company's published contact details are current.

### Input

```json
{
  "websites": ["https://www.sakura.ad.jp", "appier.com", "https://stripe.com"],
  "maxPagesPerSite": 8,
  "respectRobotsTxt": true,
  "defaultCountry": "",
  "maxConcurrency": 10,
  "maxConcurrencyPerDomain": 2,
  "requestTimeoutSecs": 20
}
```

`websites` takes full URLs or bare domains. `maxPagesPerSite` is 1 to 30 (default 8). `maxConcurrency` (sites in parallel, 1 to 50, default 10), `maxConcurrencyPerDomain` (1 to 5, default 2), `requestTimeoutSecs` (3 to 120, default 20) and `proxyConfiguration` (default: no proxy) tune speed and politeness. With an empty `websites` list the Actor runs a built-in example on apify.com.

### Output

One item per input website.

```json
{
  "website": "https://www.appier.com",
  "domain": "appier.com",
  "finalUrl": "https://www.appier.com/en/",
  "status": "ok",
  "error": null,
  "companyName": "Appier",
  "description": "Appier is a software-as-a-service company ...",
  "emailList": ["contactus-hk@appier.com"],
  "phoneList": ["+886287802800"],
  "linkedin": ["https://www.linkedin.com/company/appier"],
  "emails": [{"value": "contactus-hk@appier.com", "foundOn": "https://www.appier.com/en/contact"}],
  "otherEmails": [],
  "phones": [{"value": "+886287802800", "raw": "+886-2-8780-2800", "country": "TW", "foundOn": "https://www.appier.com/en/contact"}],
  "socials": {
    "linkedin": [{"url": "https://www.linkedin.com/company/appier", "foundOn": "https://www.appier.com/en/"}],
    "twitter": [], "facebook": [], "instagram": [], "youtube": [], "tiktok": [], "github": [], "pinterest": [], "threads": [],
    "telegram": [], "whatsapp": [], "discord": [], "reddit": [], "snapchat": []
  },
  "mentionedSocials": {"linkedin": [], "twitter": []},
  "contactForms": [{"url": "https://www.appier.com/en/contact", "foundOn": "https://www.appier.com/en/contact"}],
  "addresses": [{"streetAddress": "7 Xinyi Rd", "addressLocality": "Taipei", "addressRegion": null, "postalCode": "110", "addressCountry": "TW", "formatted": "7 Xinyi Rd, Taipei, 110, TW", "foundOn": "https://www.appier.com/en/"}],
  "pagesCrawled": 8,
  "pageUrls": ["https://www.appier.com/en/"],
  "scrapedAt": "2026-10-07T15:40:00+00:00"
}
```

Every value has a `foundOn` page. For CSV, Sheets, Clay or Zapier use the flat columns: `emailList` and `phoneList` (plain strings), one list per platform (`linkedin`, `twitter`, ... `snapchat`, URLs only), and `domain` (host without `www.`) as a join key for CRM matching. `mentionedSocials` has the same 14 platform keys as `socials` (shortened above).

`status` is `ok` (at least one contact found), `noContacts`, `unreachable` (network error or HTTP error such as 403/429), `blockedByRobots`, `invalidUrl` or `error`.

### How to extract emails and phone numbers from a list of websites

1. Open the Actor and paste your company websites into `websites` (full URLs or bare domains such as `appier.com`).
2. Leave `maxPagesPerSite` at 8 for most sites. Raise it (up to 30) for large sites with deep contact or team pages.
3. Click Start. Each site produces one item as soon as it is done.
4. Open the Output tab and export as CSV, Excel or JSON, or fetch the dataset through the API.

### How much does it cost to scrape contact details?

Pay per event: **US$2.00 per 1,000 websites with an email** (event `website-with-email`, US$0.002 each). You are charged once per website where at least one email was found on the site's own domain (addresses on other domains and other people's profiles do not count). Sites with only a phone, social profile or contact form, and sites that are unreachable or have nothing, are returned free.

Worked example: 1,000 websites, of which about 60% have an email, costs 600 x $0.002 = **$1.20**. Platform usage is small because the Actor makes plain HTTP requests only.

Per-page scrapers such as vdrmota/contact-info-scraper charge per page crawled (about $2 per 1,000 pages at the time of writing, with a 20-page default), which comes to roughly $20 to $40 per 1,000 sites. Here you pay $2 per 1,000 sites with an email, and every other site is free. Check the other Actor's current pricing before comparing.

### Limits

- No JavaScript rendering: details injected only by scripts, or on sites that block non-browser clients (HTTP 403/429), are not found.
- Text addresses are not parsed; addresses come from schema.org JSON-LD only.
- Emails and phones shown as images are not read.
- Up to 30 pages per site; default 8.
- Only the same site is crawled (plus a subdomain when a link clearly names a contact page); the Actor does not follow links to other sites.

### Use with Python

```python
## pip install apify-client
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("datagleaner/website-contact-details-scraper").call(
    run_input={"websites": ["https://www.appier.com", "stripe.com"], "maxPagesPerSite": 5}
)
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item["website"], item["status"], [e["value"] for e in item["emails"]])
```

### Use with JavaScript / Node.js

```js
// npm install apify-client
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagleaner/website-contact-details-scraper').call({
    websites: ['https://www.appier.com', 'stripe.com'],
    maxPagesPerSite: 5,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) console.log(item.website, item.status, item.emails.map((e) => e.value));
```

### Use it from n8n, Make, Zapier or an AI agent

Actor ID: `datagleaner/website-contact-details-scraper`

Minimal input:

```
{"websites": ["https://www.appier.com", "stripe.com"]}
```

Each tool below runs this Actor with your own Apify API token.

- **n8n:** add the **Apify** node (`@apify/n8n-nodes-apify`). On n8n Cloud you install it from the community node registry. Choose **Run an Actor and get dataset**, set Actor to `datagleaner/website-contact-details-scraper` and paste the input above.

- **Make:** use the Apify app's **Run an Actor** module, then **Get Dataset Items** to read the results. **Watch Actor Runs** can trigger a scenario when a run finishes.

- **Zapier:** use the Apify action **Run Actor**, then the search **Fetch dataset items**. The trigger **Finished Actor run** starts a Zap when a run ends.

- **AI agents (MCP):** connect to `https://mcp.apify.com?tools=datagleaner/website-contact-details-scraper`. In Claude Code:

  ```
  claude mcp add --transport http apify "https://mcp.apify.com?tools=datagleaner/website-contact-details-scraper"
  ```

  Then run `/mcp` to sign in to Apify in your browser. Other clients can sign in with OAuth or send the header `Authorization: Bearer YOUR_APIFY_TOKEN`. Clients that run local MCP servers can use Apify's package (`@apify/actors-mcp-server`, run with `npx -y` and `APIFY_TOKEN` set) instead. Then ask the agent in plain words, for example:

  > Find the contact email, phone number and LinkedIn page for each of these companies: appier.com, stripe.com, sakura.ad.jp. Put the results in a table with the page each one was found on.

- **LangChain (Python):**

```python
## pip install langchain-apify, then set APIFY_TOKEN in your environment
import json
from langchain_apify import ApifyActorsTool
tool = ApifyActorsTool("datagleaner/website-contact-details-scraper")
result = tool.invoke({"run_input": json.loads('{"websites": ["https://www.appier.com", "stripe.com"]}')})
```

### FAQ

**How do I extract emails from a website?** Put the website in `websites`. The Actor reads the home page and the contact, about, legal-notice and team pages, and returns every email it finds with the page it came from, including obfuscated forms such as `name [at] domain [dot] com`.

**Can I scrape emails from a list of websites in bulk?** Yes. Give it thousands of domains in one run; it works through them in parallel (`maxConcurrency` sites at once, at most `maxConcurrencyPerDomain` requests per site), and each site becomes one row.

**Does it also scrape phone numbers and social media links?** Yes. Phones are validated and normalized to E.164 with libphonenumber, and social links cover 14 platforms (LinkedIn, X, Facebook, Instagram, YouTube, TikTok, GitHub, Pinterest, Threads, Telegram, WhatsApp, Discord, Reddit, Snapchat).

**Is there a free way to try it?** Sites with no email are not charged, and a run of a few sites costs a fraction of a cent, which Apify's free plan credit covers.

**Why is a site `unreachable`?** The site timed out, refused (403) or rate-limited (429) the request. Retrying later or setting `proxyConfiguration` often helps; the default is no proxy.

### Responsible use

This Actor reads publicly available web pages only. You are responsible for complying with the law that applies to you, including data protection rules (GDPR, CCPA and similar) and anti-spam law, when you store or contact the people and companies in the results.

# Actor input Schema

## `websites` (type: `array`):

Company websites to scan, one per line. `example.com` and `https://www.example.com/` both work. Only the same domain is crawled.

## `maxPagesPerSite` (type: `integer`):

Home page plus the most promising contact, about, legal/impressum and team pages, up to this many pages per site.

## `defaultCountry` (type: `string`):

Optional ISO country code (e.g. `US`, `TW`, `JP`) used to read local phone numbers when the site's own country cannot be inferred. Leave empty to infer from the domain, page language and structured data.

## `respectRobotsTxt` (type: `boolean`):

Skip pages that the site's robots.txt disallows.

## `maxConcurrency` (type: `integer`):

How many websites are crawled at the same time.

## `maxConcurrencyPerDomain` (type: `integer`):

Polite limit on simultaneous requests to a single website.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for each page before retrying.

## `proxyConfiguration` (type: `object`):

Optional. Defaults to no proxy.

## Actor input object example

```json
{
  "websites": [
    "https://www.sakura.ad.jp",
    "https://www.appier.com"
  ],
  "maxPagesPerSite": 8,
  "respectRobotsTxt": true,
  "maxConcurrency": 10,
  "maxConcurrencyPerDomain": 2,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "https://www.sakura.ad.jp",
        "https://www.appier.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagleaner/website-contact-details-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        "https://www.sakura.ad.jp",
        "https://www.appier.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("datagleaner/website-contact-details-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "https://www.sakura.ad.jp",
    "https://www.appier.com"
  ]
}' |
apify call datagleaner/website-contact-details-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagleaner/website-contact-details-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yJXjLQnyvW8sI550o/builds/QNHcCBnsK5UB69ENz/openapi.json
