# Company Contact Extractor: Emails, Phones & Social Links (`rod_analytics/company-contact-extractor`) Actor

Enrich company domains with contact data: role emails (info@, sales@), main phone in E.164, LinkedIn, X, Facebook, Instagram, YouTube, TikTok and GitHub links, address, VAT IDs and contact page. GDPR-aware by default. One row per domain, batch or real-time API for AI agents.

- **URL**: https://apify.com/rod\_analytics/company-contact-extractor.md
- **Developed by:** [Rod Services](https://apify.com/rod_analytics) (community)
- **Categories:** Lead generation, Business
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event + usage

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Company Contact Extractor do?

**Company Contact Extractor** turns a list of **company domains into contact data**: generic **business emails** (info@, sales@, office@, support@), the **main phone number in E.164 format**, the company's **LinkedIn, X (Twitter), Facebook, Instagram, YouTube, TikTok and GitHub** links, **postal address**, **EU VAT IDs**, company register numbers, the contact page and the contact form URL. You get **one clean row per company**, ready for your CRM.

It visits the homepage and the few pages that matter (contact, impressum, imprint, about, legal) in many European languages, with plain HTTP requests and no browser. That makes it fast and cheap: about **$3 per 1,000 domains**, no matter how many pages it reads. It is **GDPR-aware by default**: personal addresses such as jane.doe@ are left out unless you switch them on.

Run it on the Apify platform with API access, scheduling, integrations (Make, Zapier, n8n, HubSpot, Google Sheets, webhooks) and monitoring, or call it **one domain at a time as a real-time API** from your app or AI agent.

Try it now: press **Start** with the prefilled example (apify.com, hetzner.com, pipedrive.com). It finishes in under 30 seconds.

### Why use Company Contact Extractor?

- **B2B lead enrichment.** You have a list of company websites from a trade fair, a directory, Google Maps or your CRM. Add emails, phone, LinkedIn page, address and VAT ID in one pass.
- **Sales prospecting.** Build account lists with the company's own published contact channels, not guessed addresses.
- **CRM enrichment and data hygiene.** Fill missing phone numbers, fix phone formats (E.164 works everywhere), attach LinkedIn company pages, verify that email domains still receive mail (MX check).
- **KYC and supplier onboarding.** Pull the legal name, register number (HRB, Company No., KvK, KRS...) and VAT ID from the impressum or legal notice.
- **AI agents and LLM tools.** A single `GET /?domain=acme.com` endpoint returns predictable JSON in a few seconds. Ideal as a tool for agents that research companies.
- **Market research.** Measure which companies publish a phone, a contact form or a TikTok channel.

### How to extract company contact details from a website

1. Open the **Input** tab.
2. Paste company domains or URLs into **Company domains or URLs**, one per line. `acme.com` is fine.
3. Optional: change **Max pages per domain** (default 5) or the **Pages to prioritise** keywords.
4. Press **Start**.
5. Open the **Output** tab. Pick the **Overview**, **Emails**, **Social profiles** or **Company details** view. Download as JSON, CSV, Excel or HTML, or fetch by API.

### Input

All fields are on the Input tab. Only `domains` is required.

| Field                     | Type             | Default                                                    | Description                                                                         |
| ------------------------- | ---------------- | ---------------------------------------------------------- | ----------------------------------------------------------------------------------- |
| `domains`                 | array of strings |                                                            | Company domains or URLs. Duplicates and www variants are merged.                    |
| `maxPagesPerDomain`       | integer          | `5`                                                        | Homepage plus best matching priority pages (1-50). Price per domain stays the same. |
| `priorityPages`           | array of strings | `contact, about, impressum, imprint, kontakt, team, legal` | Keywords for links worth following. Built-in ones also match translations.          |
| `includePersonalEmails`   | boolean          | `false`                                                    | Also return addresses of named people. Read the GDPR section first.                 |
| `includeAllPhones`        | boolean          | `false`                                                    | Return every number found (max 20), not only the main company number.               |
| `checkMx`                 | boolean          | `true`                                                     | DNS MX lookup for every email domain (`mxFound`).                                   |
| `respectRobotsTxt`        | boolean          | `true`                                                     | Skip pages disallowed by robots.txt.                                                |
| `maxConcurrency`          | integer          | `10`                                                       | Websites crawled in parallel.                                                       |
| `maxConcurrencyPerDomain` | integer          | `2`                                                        | Pages of one website fetched at once.                                               |
| `timeoutSecs`             | integer          | `15`                                                       | Timeout per page. One domain is capped at 4x this value (min. 60 s).                |
| `proxyConfiguration`      | object           | off                                                        | Optional. Apify datacenter proxy or your own proxy URLs. No residential.            |

```json
{
    "domains": ["apify.com", "https://www.hetzner.com", "pipedrive.com"],
    "maxPagesPerDomain": 5,
    "includePersonalEmails": false
}
```

### Output

One item per domain. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. The **Emails** view gives one row per address, handy for CSV import into a CRM.

```json
{
    "domain": "hetzner.com",
    "inputUrl": "hetzner.com",
    "url": "https://www.hetzner.com/",
    "companyName": "Hetzner Online GmbH",
    "emails": [
        {
            "email": "info@hetzner.com",
            "type": "role",
            "sourceUrl": "https://www.hetzner.com/unternehmen/ueber-uns/",
            "mxFound": true
        }
    ],
    "phones": [
        {
            "e164": "+4998315050",
            "raw": "+49 (0)9831 505-0",
            "sourceUrl": "https://www.hetzner.com/unternehmen/ueber-uns/"
        }
    ],
    "socials": {
        "linkedin": "https://www.linkedin.com/company/hetzner-online",
        "x": "https://x.com/hetzner_online",
        "facebook": "https://www.facebook.com/hetzner.de",
        "instagram": "https://www.instagram.com/hetzner.online",
        "youtube": "https://www.youtube.com/user/HetznerOnline",
        "tiktok": null,
        "github": null
    },
    "address": null,
    "vatIds": ["DE812871812"],
    "registrationNumbers": ["HRB 6089"],
    "contactPageUrl": "https://www.hetzner.com/support/",
    "contactFormUrl": null,
    "personalEmailsHidden": 3,
    "pagesCrawled": 5,
    "httpStatus": 200,
    "error": null,
    "scrapedAt": "2026-09-27T12:09:22.452Z"
}
```

### Data fields

| Field                  | Description                                                                                            |
| ---------------------- | ------------------------------------------------------------------------------------------------------ |
| `domain`               | Company domain without www.                                                                            |
| `companyName`          | From schema.org Organization, the copyright line, og:site\_name or the page title.                      |
| `emails`               | `{ email, type: role or personal, sourceUrl, mxFound }`. Company domains first, third parties removed. |
| `phones`               | `{ e164, raw, sourceUrl }`. Main company number by default. Fax numbers are never returned.            |
| `socials`              | Company profile per network: `linkedin`, `x`, `facebook`, `instagram`, `youtube`, `tiktok`, `github`.  |
| `address`              | `{ street, postalCode, city, region, country, formatted }` from schema.org Organization/LocalBusiness. |
| `vatIds`               | Validated EU VAT formats (DE123456789, ATU12345678, NL123456789B01...), UK VAT and Swiss UID.          |
| `registrationNumbers`  | Register entries as printed: HRB 12345, Company No. 01234567, KvK, KRS, SIREN, IČO, CVR and more.      |
| `contactPageUrl`       | The contact page that was crawled.                                                                     |
| `contactFormUrl`       | Page with a contact form (HTML form with a message field, or HubSpot/Typeform/Jotform embeds).         |
| `personalEmailsHidden` | How many personal addresses were seen but not returned.                                                |
| `pagesCrawled`         | HTML pages read for this domain.                                                                       |
| `error`                | Why a domain failed (DNS, timeout, bot protection, robots.txt). Failed domains are free.               |

#### How the extraction works

- **Emails** come from `mailto:` links, visible text, schema.org and Cloudflare-protected addresses. Obfuscated forms like `info [at] acme [dot] com`, `info(at)acme.de` or `kontakt (at) firma (punkt) de` are decoded. Image names like `logo@2x.png` and tracker IDs are ignored.
- **Role vs personal**: the local part is compared with a multilingual list of shared mailboxes (info, sales, office, support, kontakt, vertrieb, datenschutz, pardavimai, myynti...).
- **Phones** are normalised with libphonenumber to E.164. The country is guessed from the domain TLD, the schema.org address or the page language. Numbers in national format count only next to a "phone" keyword, so order numbers and dates are not mistaken for phones.
- **Social links** are only collected from links on the company website. Share buttons, single posts and personal LinkedIn profiles (`/in/`) are ignored. When several profiles exist, the one in the header/footer that matches the brand wins.

### Use it as an API for AI agents (Standby mode)

The Actor also runs as an always-ready HTTP endpoint. Each request enriches one domain and returns the JSON row:

```text
GET https://rod-analytics--company-contact-extractor.apify.actor/?domain=acme.com
Authorization: Bearer <APIFY_TOKEN>
```

Optional query parameters: `maxPages`, `includePersonalEmails`, `includeAllPhones`, `respectRobotsTxt`, `checkMx`. Calls are billed per successful domain, like batch runs. Add it to your agent as a tool through the Apify MCP server or any HTTP tool.

### How much does it cost to extract company contacts?

The Actor uses **pay per event** pricing: you pay for results, not for compute time.

| Event            | Price                    |
| ---------------- | ------------------------ |
| Actor start      | $0.001 per run           |
| Domain processed | $0.003 ($3.00 per 1,000) |

- **1,000 domains cost about $3.00**, whether the Actor reads 1 page or 5 pages per site.
- **Failed domains are free**: DNS errors, timeouts, bot walls and robots.txt blocks are listed but not charged.
- Set **Maximum cost per run** in the run options to cap spending. The Actor stops when the limit is reached.
- With the Apify free plan's monthly credit you can enrich well over a thousand companies.

### GDPR, ePrivacy and responsible use

This Actor is built for **B2B use** and follows privacy by design:

- **Default mode returns only generic role mailboxes** (info@, sales@, contact@, hello@, office@, support@ and similar) and the **company's main phone number**. Addresses and phone numbers of named employees are **not returned**; only their count is (`personalEmailsHidden`).
- **`includePersonalEmails` is off by default.** If you switch it on, the output can contain personal data under the GDPR (for example jane.doe@company.com). **You are the data controller** for that data. You need a **lawful basis** (for example legitimate interest for B2B outreach, documented in a balancing test), must inform the people concerned (Art. 14 GDPR), honour objections and deletion requests, and follow **ePrivacy / national marketing rules** (in many EU countries cold emails to individuals need prior consent). If you cannot meet these duties, keep the option off.
- **No guessing.** The Actor never generates, permutes or guesses email addresses (no `firstname.lastname@` patterns). It only reports what the company itself publishes on its website.
- **No social network scraping.** LinkedIn, Facebook, X, Instagram, TikTok, YouTube and GitHub are never visited. The Actor only reads links that appear on the company's own site.
- **Polite crawling.** robots.txt is respected by default, at most 2 parallel requests per site, a few pages per domain.

This is not legal advice. Check your use case with your data protection officer.

### Tips and advanced options

- **Speed**: raise `maxConcurrency` to 20-30 for big lists. The job is network bound; 1 GB of memory is plenty.
- **Depth**: 5 pages find the impressum and contact page on most sites. Raise it for sites with many languages.
- **Different languages**: add your own keywords to `priorityPages`, for example `kundenservice` or `ansprechpartner`.
- **Blocked sites**: some big-brand sites use bot protection and answer 403 to datacenter IPs (about 10% in our tests of large corporate sites). Try Apify datacenter proxy or your own proxy URLs for those domains.
- **Personal data minimisation**: keep `includePersonalEmails` and `includeAllPhones` off unless you really need them.

### FAQ, disclaimers and support

**Does it work on JavaScript-only websites?** It reads the server HTML without a browser. Sites that render their menu only with JavaScript may return fewer pages. Contacts in the footer or in schema.org data are usually still found.

**Why is an email I can see on the site missing?** Addresses on other companies' domains (agencies, regulators, partners) are removed on purpose. Personal addresses need `includePersonalEmails`.

**Which proxies can I use?** No proxy (the default), Apify datacenter proxy, or your own proxy URLs. Residential and SERP proxies are not supported. A run that asks for them stops at the start with a clear message and does no work.

**Is scraping company websites legal?** Reading publicly available business contact information is generally allowed, but you are responsible for how you store and use the data. Respect the websites' terms, robots.txt and privacy laws.

**Found a bug or need a custom field?** Open an issue on the **Issues** tab. Custom enrichment pipelines are available on request.

# Actor input Schema

## `domains` (type: `array`):

One per line. Bare domains like acme.com are fine, full URLs too. Duplicates (also www vs non-www) are removed. Each domain gives one output row.

## `maxPagesPerDomain` (type: `integer`):

Homepage plus the best matching priority pages. 5 covers contact, impressum and about pages on most sites. The price per domain does not depend on this number.

## `priorityPages` (type: `array`):

Keywords matched against link URLs and link texts, in priority order. Built-in keywords also match translations (contact matches kontakt, contacto, contatti, yhteystiedot, kontaktai...). Links that match none of them are not followed.

## `includePersonalEmails` (type: `boolean`):

Off (default): only generic role mailboxes such as info@, sales@, office@, support@ are returned. On: addresses of named people found on the site (jane.doe@) are returned too. You become the data controller for that personal data and need a lawful basis (for example B2B legitimate interest) and must follow ePrivacy rules for marketing. See the README.

## `includeAllPhones` (type: `boolean`):

Off (default): only the company's main phone number. On: every number found, including numbers on team pages (up to 20). Fax numbers are never returned.

## `checkMx` (type: `boolean`):

Look up MX records of each email domain and set mxFound. One cached DNS query per domain. Does not contact any mail server.

## `respectRobotsTxt` (type: `boolean`):

Skip pages disallowed by the site's robots.txt. Recommended.

## `maxConcurrency` (type: `integer`):

How many websites are crawled at the same time.

## `maxConcurrencyPerDomain` (type: `integer`):

How many pages of the same website are fetched at once. Keep it low to be polite.

## `timeoutSecs` (type: `integer`):

Maximum wait for one page. A whole domain is capped at 4x this value (at least 60 s).

## `proxyConfiguration` (type: `object`):

Optional. Most company websites work without a proxy. Use Apify datacenter proxy or your own proxy URLs if many sites block you. Residential and SERP proxies are not supported.

## Actor input object example

```json
{
  "domains": [
    "apify.com",
    "hetzner.com",
    "pipedrive.com"
  ],
  "maxPagesPerDomain": 5,
  "priorityPages": [
    "contact",
    "about",
    "impressum",
    "imprint",
    "kontakt",
    "team",
    "legal"
  ],
  "includePersonalEmails": false,
  "includeAllPhones": false,
  "checkMx": true,
  "respectRobotsTxt": true,
  "maxConcurrency": 10,
  "maxConcurrencyPerDomain": 2,
  "timeoutSecs": 15,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `emails` (type: `string`):

No description

## `socials` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "apify.com",
        "hetzner.com",
        "pipedrive.com"
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("rod_analytics/company-contact-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "apify.com",
        "hetzner.com",
        "pipedrive.com",
    ],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("rod_analytics/company-contact-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "apify.com",
    "hetzner.com",
    "pipedrive.com"
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call rod_analytics/company-contact-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,rod_analytics/company-contact-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QAt3pYw4CAWb69n8G/builds/XfHaLyvSLHe1nv0J2/openapi.json
