# Company Contact Finder: Website Emails, Deliverability Checked (`accountable_eel/company-contact-finder`) Actor

Find company emails from a website, each one checked for deliverability (MX, SPF, DMARC, disposable) before you pay. A contact scraper with email check plus a team page email finder that infers addresses from the company's own format. Fixed price per website, to Google Sheets or Clay.

- **URL**: https://apify.com/accountable\_eel/company-contact-finder.md
- **Developed by:** [Adrian Voss](https://apify.com/accountable_eel) (community)
- **Categories:** Lead generation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $12.50 / 1,000 website with a deliverable emails

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Company Contact Finder: Website Emails, Deliverability Checked

Give it a list of company domains. Get back the contact emails those companies
publish on their own websites, each one checked against its domain's mail
records before you are charged, plus (if you ask for it) the people the site
names on its team pages with an email built from the address format that
company's own published addresses demonstrate.

One flat price per website that returns a deliverable address, whatever the
page count. A website that returns nothing costs nothing.

### Who it's for

- **Outbound and agency teams** building a company list from domains, who want
  the general inbox and the named decision-maker in one pass instead of two
  tools.
- **RevOps and data teams enriching a CRM**, who need a per-row source URL they
  can show to anyone who asks where an address came from.
- **Anyone who got a surprise bill from a crawler** that charged per page and
  went four levels deep on a 4,000-page site.
- **Clay, n8n, Make and AI-agent workflows** that need one HTTP call per company
  and a fixed, predictable unit cost.

### Why this one

- **Fixed price per website, not per page.** The crawl is not a crawl: the
  homepage plus only the links that look like contact, legal, about, team or
  support pages, same host, hard-capped (10 by default, 30 maximum). Reading 3
  pages and reading 30 cost the same.
- **Every address is checked before you pay.** MX records (with the RFC 5321
  A-record fallback), SPF, DMARC, a DKIM lookup at the selectors real providers
  document, and a community-maintained disposable-domain list. An address whose
  domain cannot receive mail at all is dropped by default and is never billed.
- **A learned address format, not a guessed one.** Other pattern finders assume
  `first.last` because it is the most common convention in general. This one
  only ever uses a format that **this** company's own published addresses
  demonstrate. If the site publishes no personal address, no format is learned
  and no person email is produced.
- **You are never charged for a miss.** No deliverable address, nothing
  suppressed left over, nothing matching your filters: the row comes back with
  `"found": false` and the run bills nothing for that website.
- **An opt-out list that actually works.** Addresses and whole domains on your
  suppression list are dropped before anything is returned and before anything
  is counted, so honouring an opt-out never costs you money.
- **A hard ceiling on people.** At most 25 inferred person emails per website,
  flagged when the cap bites, so a 400-person staff directory cannot turn into a
  surprise bill.

### Price

- **Website with a deliverable email**: $25 per 1,000 email addresss
- **Person email worked out**: $30 per 1,000 email addresss

Plus a $0.00005 start fee per run. Each event above is billed independently, only when it actually returns data — misses (`found:false`) are never charged.

Two units, and only two:

| Unit | When it is charged |
| --- | --- |
| `site` | Once per website that returned at least one deliverable address, however many pages were read and however many addresses came back. |
| `person-email` | Once per **inferred** person email that passed the deliverability check. Only possible with people turned on. |

Never charged: a website with no result, an address on your suppression list, an
address whose domain has no mail route or is a disposable domain, an address
filtered out by your `emailTypes` choice, and every published address beyond the
first (they are all covered by the one flat `site` price).

### How to use

1. **In the Apify Console.** Open the actor page and click **Start** — the `domains` field is already pre-filled with a working example. Results land in the run's dataset as soon as each item is found.
2. **Via the API.** Call it directly with a POST request — no Console needed once you have an API token:
   ```bash
   curl "https://api.apify.com/v2/acts/accountable_eel~company-contact-finder/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
     -X POST \
     -H "Content-Type: application/json" \
     -d '{"domains":["hetzner.com","fastly.com"]}'
   ```
3. **On a schedule.** Save this actor as an Apify **Task** with the input you want, then add a **Schedule** (hourly, daily, weekly) so it runs on its own — no server of your own required.

### Input

```json
{
  "domains": [
    "hetzner.com",
    "fastly.com"
  ]
}
```

One company per line: acme.com, www.acme.com or https://acme.com/en/ all work. Each site is read as its homepage plus only its contact-looking pages (contact, impressum, about, team, leadership, support, legal, privacy), capped by 'Max pages per site' below. Every address found is checked against its domain's mail records. One flat price per website that returns a deliverable email; a website that returns nothing is free. Accepted formats: hetzner.com, www.fastly.com, https://basecamp.com/.

Everything else is optional:

| Field | Default | What it does |
| --- | --- | --- |
| `findPeople` | `false` | Also read team, about, leadership and people pages for names and job titles, and build one email per person from the company's own address format. Billed per person email that passes the check. |
| `emailTypes` | `"both"` | `"role"` for shared inboxes only (info@, sales@, support@), `"personal"` for a person's address only. Narrowing drops the other kind before billing. An array (`["role"]`) works too. |
| `onlyDeliverable` | `true` | Drop addresses whose domain has no mail route, or is a disposable domain. Turn off to see those rows; they are free either way. |
| `suppressionList` | empty | Your opt-out list: full addresses, or whole domains (`acme.com` or `@acme.com`, subdomains included). Case-insensitive, applied before anything is returned or billed. |
| `maxPagesPerSite` | `10` | 1 to 30, homepage included. Does not change the price. |
| `expandRows` | `true` | One row per email address. Turn off for one row per website with every address in an `emails` array. The bill is identical either way. |

### Output

One row per email address by default, so the dataset drops straight into Google
Sheets, a CRM import, or a Clay table:

| query | found | status | domain | companyName | email | emailType | origin | sourceUrl | personName | personTitle | pattern | patternConfidence | mxFound | provider | spf | dmarc | dkim | disposable | deliverability | personEmailBilled | emails | emailsText | emailCount | publishedEmailCount | inferredEmailCount | peopleFound | peopleCapReached | phones | socials | contactFormUrl | addressText | website | pagesFetched | pagesRead | pagesSkippedByRobots | pageCapReached | checkedAt | scrapedAt |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| hetzner.com | true | OK | <domain> | <company name> | <email> | <kind of address> | <where it came from> | <source page> | <person> | <job title> | <address format> | <format confidence> | <mx record found> | <mailbox host> | <spf published> | <dmarc published> | <dkim found at a common selector> | <disposable domain> | <deliverability> | <billed as a person email> | \<all addresses (grouped)> | \<all addresses (comma-separated)> | <addresses returned> | <published addresses> | <inferred person addresses> | <people named on the site> | <people cap reached> | <phones> | <social profiles> | <contact form page> | <postal address> | <homepage read> | <pages fetched> | <pages read> | \<pages skipped (robots.txt)> | <page cap reached> | <checked at> | 1970-01-01T00:00:00.000Z |

The fields that matter most:

| Field | What it tells you |
| --- | --- |
| `email`, `emailType` | The address, and whether it is a shared role inbox or a person's. |
| `origin` | `published` if the company printed this exact address on its own site, `inferred` if it was built for a named person from the company's own format. |
| `sourceUrl` | For a published address, the page it was printed on. For an inferred one, the page(s) the format was learned from. |
| `personName`, `personTitle` | Who the address belongs to, when the site names them. |
| `pattern`, `patternConfidence` | The learned format, and how much site evidence backed it (0.9 two or more published addresses back it, 0.7 one does, 0.6 two match by shape alone, 0.45 one does, capped at 0.5 for a person whose accented name the site never spelled out in an address). |
| `deliverability` | `deliverable-domain`, `no-mail-route`, or `disposable`. See the FAQ: this is a domain-level check. |
| `mxFound`, `provider`, `spf`, `dmarc`, `dkim`, `disposable` | The mail records the verdict was read from, so you can judge it yourself. |
| `phones`, `socials`, `contactFormUrl`, `addressText` | The rest of the company's published contact surface, repeated on each of its rows. |
| `personEmailBilled` | `true` on exactly the rows charged as a `person-email`. |

A miss comes back as a row with `"found": false`, a status
(`NOT_FOUND` / `BLOCKED` / `BAD_FORMAT` / `REQUEST_FAILED`) and a message saying
which addresses were found and why none survived. Misses are never charged.

### Example runs

Paste either of these into the Input tab and press Start. Both return rows with
no key, no token and no account of your own beyond an Apify login.

**Role inboxes for two companies, no people.** The default shape, and the
cheapest: two `site` events, nothing else.

```json
{
  "domains": ["hetzner.com", "fastly.com"],
  "emailTypes": "role",
  "findPeople": false
}
```

**One company, everything it publishes, still no people.** Role and personal
addresses, deliverability on each, phones and socials repeated per row.

```json
{
  "domains": ["hetzner.com"],
  "maxPagesPerSite": 10,
  "findPeople": false
}
```

**Honouring an opt-out.** The same run with two addresses permanently excluded.
They are dropped before the run counts anything, so suppressing them makes the
run cheaper, never more expensive.

```json
{
  "domains": ["hetzner.com"],
  "findPeople": false,
  "suppressionList": ["info@hetzner.com", "@example.com"]
}
```

### Use it from Clay, n8n, Make, or an AI agent

This actor runs synchronously over plain HTTP — call it directly from a script, a workflow tool, or an AI agent, no Apify Console needed once you have an API token.

```bash
curl "https://api.apify.com/v2/acts/accountable_eel~company-contact-finder/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
  -X POST \
  -H "Content-Type: application/json" \
  -d '{"domains":["hetzner.com","fastly.com"]}'
```

**n8n.** Add an HTTP Request node: Method `POST`, URL `https://api.apify.com/v2/acts/accountable_eel~company-contact-finder/run-sync-get-dataset-items?token=<YOUR_TOKEN>`, Body Content Type `JSON`, JSON Body `{"domains":["hetzner.com","fastly.com"]}` (swap in an expression from an earlier node for a real value).

**Clay.** Add an "HTTP API" column: Method `POST`, URL `https://api.apify.com/v2/acts/accountable_eel~company-contact-finder/run-sync-get-dataset-items?token=<YOUR_TOKEN>`, Body `{"domains":["{{company}}"]}`, mapping the row's company into the `domains` array.

**MCP.** In Claude, Cursor, or any MCP client with the Apify MCP server, ask for "Company Email Finder: Website Emails, Checked" — the agent will find and run this actor.

### vs. alternatives

| | This actor | Per-page contact crawlers | Decision-maker finders |
| --- | --- | --- | --- |
| Unit you pay for | One website with a result | Each page crawled | Each domain row, plus each contact |
| Cost of a deep site | Unchanged | Grows with the crawl | n/a |
| Cost of a miss | Nothing | You still paid for the pages | Often the domain row anyway |
| Email check | Included in the price | A separate paid add-on | Included, priced per contact |
| Where names come from | The company's own team pages | n/a | Third-party databases |
| Address format | Learned from this company's own published addresses | n/a | A general convention prior |
| Mailbox existence | Never claimed | Sometimes implied | Often claimed as confirmed |

If what you actually need is one page rendered with JavaScript, or a list of
companies you do not have domains for yet, this is the wrong tool: it is
HTTP-only and domain-in.

### Data & privacy

- **Only the company's own domain is read.** The crawl never leaves the site's
  own host. No third-party site, data broker, or social platform is ever
  visited: the `socials` field is only the links the company itself publishes.
- **robots.txt is honoured for every path.** It is fetched before any page, so a
  `Disallow: /` site costs one robots.txt request and is never crawled. Paths
  its rules disallow are skipped and counted in `pagesSkippedByRobots`.
- **Free personal webmail is dropped** (gmail.com, hotmail.com, gmx.de and the
  rest) unless the site publishes it as its business contact, on a contact or
  legal-notice page, in its own schema.org data, or as the only address the
  whole site offers. A sole trader whose only listed contact is a Gmail address
  is a real business contact; a reader's address quoted in a blog post is not.
  When it is kept, `emailType` is `personal`.
- **Your suppression list wins over everything**, and is applied before anything
  is returned or billed.
- **You are the data controller for what you do with the output.** This actor
  returns business contact information a company chose to publish about itself;
  what happens next is your decision, in your systems.
- **Business-contact use.** Keep a lawful basis for your outreach, such as
  legitimate interest, keep a record of where each address came from (the
  `sourceUrl` field is there for exactly that), identify yourself in your first
  message, and honour every opt-out by adding it to `suppressionList` on the
  next run. None of this is legal advice; if your situation is unusual, ask
  someone qualified.

### FAQ

**Do you verify that the mailbox exists?**
No, and nothing here says otherwise. `deliverability` is a **domain**-level
verdict: whether the domain accepts mail at all, whether it is a throwaway
domain, and what its published mail policy looks like. Confirming one specific
mailbox requires an SMTP `RCPT TO` handshake on port 25, which cloud platforms
block to prevent spam abuse. Any tool that claims to have confirmed a mailbox
from HTTP alone is claiming something it cannot know.

**So what does `deliverable-domain` actually buy me?**
It removes the two biggest sources of hard bounces before you send: domains that
cannot receive mail at all, and throwaway domains. It does not promise that
`jane.doe@` is Jane's real address.

**How is an inferred email different from a guess?**
A guess picks the most common convention in general. This actor only uses a
format this company's own website demonstrates. `patternConfidence` tells you
how much evidence there was, and `pattern` tells you which format was used, so
you can decide for yourself. If the site publishes no personal address, you get
no inferred emails at all rather than a confident-looking guess.

**What happens to a name with an umlaut or an accent?**
Two things are learned, not one: the format (`first.last`) and how the company
spells an accent in an address. `Müller` becomes `mueller` at a company whose own
published addresses show that convention, and `muller` at one that shows the
other. When nothing the site published settles it, the address is still returned
but `pattern` says `(accented name, spelling unproven)` and
`patternConfidence` is capped at 0.5 for that person only. A sharp s is always
`ss`.

**Why is `dkim` false on a domain I know signs its mail?**
DNS has no mechanism for listing DKIM selectors: the record lives at
`<selector>._domainkey.<domain>` and the selector is chosen by whatever sends
the mail. This actor checks the selectors real providers document (google,
selector1, selector2, k1, default, mail, s1). `false` means none of those
answered, never "this domain has no DKIM".

**Does raising the page cap cost more?**
No. The `site` price is flat per website. The cap only decides how thoroughly a
site is read before the actor stops.

**A site has 300 people on its team page. What happens?**
At most 25 inferred person emails are produced, and `peopleCapReached` comes
back `true`. Published addresses are not capped.

**Nothing came back for a domain I know has a contact page.**
Check the `status` and `message` on the row. Common causes: the site's
navigation is rendered by JavaScript, so no links exist in the raw HTML (the
actor tries `/contact` and `/about` anyway); robots.txt disallows the path; or
the site answered HTTP 403 to bot traffic. All three are free.

**Can I use this to find someone's personal email address?**
No. It reads what a company publishes about how to reach it, on its own domain,
for business contact. It does not search people, look up private addresses, or
touch anything outside the site you gave it.

### Related actors

- **Website Contact Scraper** — the same crawl without the deliverability check
  or the people inference: emails, phones, socials, contact form and address at
  a fixed price per site.
- **Email Deliverability Check** — the same domain-level checks for a list of
  addresses you already have.
- **Email Pattern Finder** — ranked candidates for one named person at one
  domain, when you have the name but not the website.
- **Company Domain Enrichment** — RDAP, DNS, tech stack and TLS for a domain,
  when you need the company's technical footprint rather than its inbox.

# Actor input Schema

## `domains` (type: `array`):

One company per line: acme.com, www.acme.com or https://acme.com/en/ all work. Each site is read as its homepage plus only its contact-looking pages (contact, impressum, about, team, leadership, support, legal, privacy), capped by 'Max pages per site' below. Every address found is checked against its domain's mail records. One flat price per website that returns a deliverable email; a website that returns nothing is free. Accepted formats: hetzner.com, www.fastly.com, https://basecamp.com/. You're only charged for the ones we actually find — a miss costs nothing.

## `testRun` (type: `boolean`):

Turn this on to test your input on a small sample before running the full list. Turn it off to process everything.

## `onlyFound` (type: `boolean`):

Only keep rows where something was actually found. Misses are always free, whether or not you show them here.

## `includeKeywords` (type: `array`):

Optional. Only keep results that mention at least one of these words (e.g. a job title, a city, a product name). Leave empty to keep everything.

## `excludeKeywords` (type: `array`):

Optional. Drop any result that mentions one of these words. Leave empty to skip nothing.

## `maxResults` (type: `integer`):

Optional. Stop the run once this many results have been found — useful for a quick, cheap sample. Leave blank for no limit.

## `findPeople` (type: `boolean`):

Off by default. When on, the site's team, about, leadership and people pages are also read for names and job titles, and each person gets one email built from the address format this company's own published addresses demonstrate. If the site publishes no personal address, no format is learned and no person email is produced: a format is never guessed. Billed separately per person email that passes the deliverability check.

## `emailTypes` (type: `string`):

Leave on both for everything. Picking one kind drops the other before anything is billed, so you are never charged for addresses you filtered out. Personal-only still finds people when you turn people on; role-only never does.

## `onlyDeliverable` (type: `boolean`):

On by default: an address is dropped when its domain has no mail route at all, or is a known disposable/temp-mail domain. Turn off to see those rows too; they are never billed either way. Mailbox existence is not checked by anyone without an SMTP handshake, and this actor does not claim to.

## `suppressionList` (type: `array`):

Your opt-out list. One entry per line: a full address (jane@acme.com) or a whole domain (acme.com or @acme.com, which also covers its subdomains). Matched ignoring upper/lower case, and applied before anything is returned or billed. Add anyone who asks you to stop contacting them.

## `columns` (type: `array`):

Choose which pieces of information to include in each result row. All are included by default.

## `expandRows` (type: `boolean`):

When on, each email address found gets its own row instead of being grouped under its company. You're still only charged once per company, no matter how many rows it produces.

## `maxPagesPerSite` (type: `integer`):

1 to 30, homepage included. Only the homepage and links that look like contact, legal, about or team pages are ever fetched, so most sites stop well before the cap. The price per website is the same whatever this is set to.

## `maxConcurrency` (type: `integer`):

Parallel requests. Keep conservative — this target has no browser fallback, so getting blocked costs more than slow-and-steady.

## `proxyConfiguration` (type: `object`):

Apify Proxy config. Residential recommended for anti-bot-sensitive targets.

## Actor input object example

```json
{
  "domains": [
    "hetzner.com",
    "fastly.com"
  ],
  "testRun": false,
  "onlyFound": false,
  "includeKeywords": [],
  "excludeKeywords": [],
  "findPeople": false,
  "emailTypes": "both",
  "onlyDeliverable": true,
  "suppressionList": [],
  "columns": [
    "domain",
    "companyName",
    "email",
    "emailType",
    "origin",
    "sourceUrl",
    "personName",
    "personTitle",
    "pattern",
    "patternConfidence",
    "mxFound",
    "provider",
    "spf",
    "dmarc",
    "dkim",
    "disposable",
    "deliverability",
    "personEmailBilled",
    "emails",
    "emailsText",
    "emailCount",
    "publishedEmailCount",
    "inferredEmailCount",
    "peopleFound",
    "peopleCapReached",
    "phones",
    "socials",
    "contactFormUrl",
    "addressText",
    "website",
    "pagesFetched",
    "pagesRead",
    "pagesSkippedByRobots",
    "pageCapReached",
    "checkedAt"
  ],
  "expandRows": true,
  "maxPagesPerSite": 10,
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "hetzner.com",
        "fastly.com"
    ],
    "includeKeywords": [],
    "excludeKeywords": [],
    "suppressionList": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("accountable_eel/company-contact-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "hetzner.com",
        "fastly.com",
    ],
    "includeKeywords": [],
    "excludeKeywords": [],
    "suppressionList": [],
}

# Run the Actor and wait for it to finish
run = client.actor("accountable_eel/company-contact-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "hetzner.com",
    "fastly.com"
  ],
  "includeKeywords": [],
  "excludeKeywords": [],
  "suppressionList": []
}' |
apify call accountable_eel/company-contact-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,accountable_eel/company-contact-finder"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/njPApfjGuTLb37zbo/builds/YvownAZ1MaXT0UOfP/openapi.json
