# Website Contact Scraper: Emails, Phones & Socials (`digital_influx/website-contacts`) Actor

Paste company websites, get one row per site: emails checked in DNS, phones, social profiles, company name, address, key people with their roles (from the legal notice), careers page and the job board it uses. Respects robots.txt. Pay only for sites with results.

- **URL**: https://apify.com/digital\_influx/website-contacts.md
- **Developed by:** [Bruno Petrelli](https://apify.com/digital_influx) (community)
- **Categories:** Lead generation, Marketing, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 website with results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Scraper: Emails, Phones & Socials

Paste a list of company websites and get **one clean row per website**: emails, phone numbers, social profiles, company name and address, **the people who run the company** (managing directors, board members, founders) with their roles, the careers page, and the **job board the company uses** (Greenhouse, Lever, Workday, Ashby, BambooHR and more). Every email is checked in DNS, so addresses on domains that cannot receive mail are left out.

**You pay only for websites where something was found.** Sites with no contact data, sites that can't be reached and sites whose robots.txt says no are still listed, for free.

### Why this one

- **One row per website, not one per page.** Everything found on the homepage, contact, imprint/Impressum, about and careers pages is merged and deduplicated.
- **Fast and cheap.** No browser, no proxies: plain HTTP requests, about 1 to 2 seconds per site. Stripe, Userpilot, ottonova, Basecamp and one made-up domain took 9 seconds in our test, 5 pages each.
- **Emails other scrapers miss:** Cloudflare "email protection" (`[email protected]`) is decoded, as are `\u0040`/`&#64;` escapes and `info [at] company [dot] com`.
- **Less junk:** image names like `logo@2x.png`, form placeholders (`your.email@company.com`, `beispiel@email.de`, `nombre@ejemplo.com`), `example.com` and tracking IDs are dropped.
- **Best emails first:** the site's own domain, then shared inboxes (`info@`, `sales@`, `support@`), then linked addresses before names that only appear in text. Each email says where it came from.
- **Decision makers, from sources meant to name them.** In Germany, Austria and Switzerland the legal notice (Impressum) must name the managing directors, and many sites mark up their founders and team as schema.org Person. The Actor reads both: `Geschäftsführer`, `Vorstand`, `Vertreten durch`, `Management board`, `Directors`, `Gérant`, `Directeur de la publication`, `Administrador único` and more, each with the role in English. No guessing from running text. On n26.com it finds the management and supervisory boards; on userpilot.com the CEO and CTO.
- **Emails that can receive mail.** Each email domain is looked up in DNS (MX records, RFC 5321 and RFC 7505). Typos, placeholders and dead domains are dropped from the list. No test email is sent and no mail server is contacted.
- **Hiring signal included:** the careers page and the job board URL, ready to paste into our **ATS Jobs Scraper** to get the open jobs.
- **Polite by default:** robots.txt is respected, pages of one site are read one at a time, and the crawler names itself (`website-contacts`).

### Use it for

- **Lead generation:** turn a list of domains (from a CRM, a directory or Google Maps) into emails, phones and LinkedIn pages.
- **CRM enrichment:** fill in missing social profiles, phone numbers and addresses.
- **Sales intelligence:** see which companies are hiring and on which applicant tracking system.
- **AI agents:** plain JSON, one object per company, easy for an LLM to read. Agents can call it through the Apify MCP server.

### Input

| Field | What it does |
|---|---|
| **Websites** (required) | One per line: `stripe.com`, `www.example.de`, or any URL of the site. |
| Max pages per website | Homepage included, default 5. The Actor picks contact, imprint, about, careers and team pages. |
| Max emails per website | Default 20, best first. `emailsTotal` tells you how many were found. |
| Only emails on the site's own domain | Drops regulators, agencies and personal addresses the site mentions. |
| Drop emails whose domain cannot receive mail | On by default. DNS check of every email domain. |
| People and their roles | On by default. Managing directors, board members, owners and founders the site names. |
| Respect robots.txt | On by default. |
| Websites in parallel | Default 5. |

```json
{
  "websites": ["stripe.com", "userpilot.com", "https://www.ottonova.de"],
  "maxPagesPerSite": 5
}
```

### Output

One item per website:

```json
{
  "input": "userpilot.com",
  "domain": "userpilot.com",
  "url": "https://userpilot.com/",
  "status": "ok",
  "error": null,
  "companyName": "Userpilot",
  "description": "Userpilot is the AI-powered, no-code product growth platform...",
  "emails": ["sales@userpilot.com", "accounting@userpilot.com", "support@userpilot.com"],
  "emailsTotal": 3,
  "phones": ["(702) 830-7422", "+17373452881", "+17028190599", "+1-702-839-7422"],
  "linkedin": ["https://www.linkedin.com/company/teamuserpilot"],
  "linkedinPeople": [],
  "twitter": ["https://x.com/teamuserpilot"],
  "facebook": ["https://www.facebook.com/userpilot"],
  "instagram": [],
  "youtube": ["https://www.youtube.com/@userpilot"],
  "tiktok": [],
  "github": [],
  "pinterest": [],
  "threads": [],
  "careersPage": "https://userpilot.bamboohr.com/careers",
  "jobBoards": [{ "platform": "bamboohr", "url": "https://userpilot.bamboohr.com/careers" }],
  "address": "7200 N Mopac Expy, Suite 300, 78731 Austin, TX, US",
  "logo": "https://userpilot.com/wp-content/uploads/2026/04/userpilot-logo-2026-dark.svg",
  "language": "en-us",
  "people": [
    { "name": "Yazan Sehwail", "role": "Co-Founder & CEO", "roleAsWritten": "Co-Founder & CEO", "source": "schema.org", "page": "https://userpilot.com/" },
    { "name": "Thabet Gharabah", "role": "Co-Founder & CTO", "roleAsWritten": "Co-Founder & CTO", "source": "schema.org", "page": "https://userpilot.com/" }
  ],
  "emailDetails": [
    { "email": "sales@userpilot.com", "type": "role", "source": "cloudflare", "page": "https://userpilot.com/contact-us/", "ownDomain": true, "mailServer": true }
  ],
  "pagesVisited": ["https://userpilot.com/", "https://userpilot.com/contact-us/"],
  "scrapedAt": "2026-09-24T19:00:00.000Z"
}
```

Field notes:

- `status`: `ok` (something found, charged), `no_data` (nothing found, free), `blocked_by_robots` (free), `unreachable` or `http_error` (free, with the reason in `error`).
- `emailDetails[].source`: `mailto` link, `cloudflare` (decoded), `schema.org` data, `obfuscated` ("\[at]"), or `text`. `type` is `role` for shared inboxes and `personal` otherwise.
- `phones` come only from `tel:` links and schema.org data, which the site itself marks as phone numbers. Digits in running text are too often prices or IDs.
- `people[]`: `name` as the site writes it (academic titles kept), `role` in English (`Managing director`, `Executive board`, `Supervisory board`, `Legal representative`, `Owner`, `Founder`, or the job title from schema.org; a note like `(Chair)` is kept), `roleAsWritten` in the site's language, `source` (`legal notice` or `schema.org`) and the `page`. One entry per person: roles found on several pages are joined with `; `. At most 25 per site.
- `emailDetails[].mailServer`: `true` when the domain can receive mail, `false` when it cannot (no such domain, or a "null MX"), `null` when DNS did not answer. Addresses with `false` are not in `emails`.
- `linkedin` holds company pages; `linkedinPeople` holds personal profiles linked from the site.
- Social lists put handles that contain the site's name first (`stripe`, `stripehq`) before other accounts the site links to.
- `jobBoards` lists the job boards the pages link to. Each URL can be pasted as is into the ATS Jobs Scraper.
- The run's **OUTPUT** record in the key-value store has the status of every input, including invalid and duplicate ones.

### Pricing

Pay per event: **one event per website with results** (`status: ok`). People and the email check are included in that price. A site where nothing was found, or that could not be read, costs nothing. The number of pages read does not change the price. Set a maximum cost for the run and the Actor stops cleanly when it is reached.

### Good to know

- **Static HTML only.** Contact data that a site loads with JavaScript after the page opens (some contact widgets, some careers pages) is not visible without a browser. In our tests most sites put emails, phones and social links in the HTML, but not all.
- **The job board** is found when a page links to it. Careers pages that load their jobs with JavaScript show the careers page but may not show the board.
- **robots.txt** is read for the site and for any domain it redirects to. A server error on robots.txt means "do not crawl" (RFC 9309), so such sites come back as `blocked_by_robots`.
- Pages are read on the host the homepage lands on (for example www.example.com), not on other subdomains like blog. or press.
- **Personal data:** names, emails and phone numbers of people are personal data under the GDPR and similar laws. Use the results only for purposes you have a lawful basis for, such as B2B outreach with a legitimate interest, and honor opt-out requests. `onlyOwnDomainEmails` and the email type help you keep to company inboxes.

### Support

A site that should have data but came back empty, or an email that looks wrong? Open an issue on the Actor's Issues tab with the website. Issues get an answer within a few days.

# Actor input Schema

## `websites` (type: `array`):

One per line: a domain (stripe.com) or any URL of the site (https://www.example.de/impressum). The Actor starts at the given address, reads the homepage and the best contact pages it links to, and returns one row per website.

## `maxPagesPerSite` (type: `integer`):

Homepage included. The Actor picks the pages most likely to hold contact data: contact, imprint/Impressum, about, careers, team. 5 is enough for most sites.

## `maxEmailsPerSite` (type: `integer`):

Best first: the site's own domain, shared inboxes (info@, sales@) and linked addresses before names that only appear in text. Keeps a staff directory from flooding the row. The total found is in emailsTotal.

## `onlyOwnDomainEmails` (type: `boolean`):

Drop addresses on other domains (a regulator, an agency, a personal Gmail) that the site happens to mention.

## `checkEmailDomains` (type: `boolean`):

Checks the DNS mail records (MX) of every email domain found. Addresses on domains that do not exist or take no mail (typos, placeholders, dead domains) leave the emails list; emailDetails still shows them with mailServer false. No email is sent and no mail server is contacted.

## `includePeople` (type: `boolean`):

Names the site itself publishes with a role: managing directors, board members and owners from the legal notice (Impressum, mentions légales, aviso legal), and founders or employees marked up as schema.org Person. Names are personal data: use them only with a legal basis (e.g. GDPR legitimate interest).

## `respectRobotsTxt` (type: `boolean`):

Recommended. Pages the site's robots.txt disallows are not read, and a site that disallows everything is reported as blocked\_by\_robots (free).

## `maxConcurrency` (type: `integer`):

How many websites are read at the same time (1 to 20). Pages of one website are always read one after another.

## Actor input object example

```json
{
  "websites": [
    "stripe.com",
    "userpilot.com",
    "https://www.ottonova.de"
  ],
  "maxPagesPerSite": 5,
  "maxEmailsPerSite": 20,
  "onlyOwnDomainEmails": false,
  "checkEmailDomains": true,
  "includePeople": true,
  "respectRobotsTxt": true,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "stripe.com",
        "userpilot.com",
        "https://www.ottonova.de"
    ],
    "maxPagesPerSite": 5
};

// Run the Actor and wait for it to finish
const run = await client.actor("digital_influx/website-contacts").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": [
        "stripe.com",
        "userpilot.com",
        "https://www.ottonova.de",
    ],
    "maxPagesPerSite": 5,
}

# Run the Actor and wait for it to finish
run = client.actor("digital_influx/website-contacts").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "stripe.com",
    "userpilot.com",
    "https://www.ottonova.de"
  ],
  "maxPagesPerSite": 5
}' |
apify call digital_influx/website-contacts --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,digital_influx/website-contacts"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/aDVyFTgu1cMHBy4S1/builds/qCk7D65stl3mtOcsc/openapi.json
